Polarization spectrum three-dimensional imaging method and device using intelligent algorithm

By combining the physical properties of polarization spectroscopy with a multimodal fusion network, the problem of normal reconstruction accuracy in complex scenes of polarization 3D imaging was solved, achieving higher accuracy and adaptability in 3D normal reconstruction.

CN122023732APending Publication Date: 2026-05-12NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHWESTERN POLYTECHNICAL UNIV
Filing Date
2026-02-04
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing polarization 3D imaging technology has low accuracy in normal reconstruction when dealing with complex scenes, especially those with high reflectivity and multiple materials. Furthermore, existing deep learning methods suffer from data loss or distortion in complex scenes.

Method used

By combining the physical properties of polarization spectroscopy with polarization information, normal priors, and spectral information, a physically constrained three-dimensional normal reconstruction model based on polarization spectroscopy is constructed. A neural network is used for three-dimensional normal reconstruction, and a multi-modal fusion network architecture is introduced to improve reconstruction accuracy and adaptability.

Benefits of technology

It significantly improves the accuracy and generalization ability of 3D normal reconstruction, and can better handle complex scenes such as high reflectivity and mixed materials, providing more complete physical information guidance and multi-dimensional light field information interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122023732A_ABST
    Figure CN122023732A_ABST
Patent Text Reader

Abstract

The invention provides a polarization spectrum three-dimensional imaging method and device using an intelligent algorithm, and relates to the field of three-dimensional imaging, a three-dimensional model and a polarization spectrum image are collected, and the three-dimensional model and the polarization spectrum image are processed to obtain a real normal, combined polarization information, normal prior information and spectral information; constructing a polarization spectrum three-dimensional normal reconstruction model based on physical constraints by using the combined polarization information, the normal prior information and the spectral information, training the polarization spectrum three-dimensional normal reconstruction model by using a real normal to obtain an optimal polarization spectrum three-dimensional normal reconstruction model, and setting a loss function and an evaluation index to evaluate the model; and inputting the polarization spectrum information of the target scene into the polarization spectrum three-dimensional normal reconstruction model to generate a prediction normal of the target scene, and reconstructing a three-dimensional shape of the target according to the prediction normal. According to the polarization spectrum three-dimensional imaging method and device using the intelligent algorithm, the reconstruction precision of the three-dimensional normal is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of three-dimensional imaging with monocular cameras, and in particular to a polarization spectral three-dimensional imaging method and apparatus utilizing intelligent algorithms. Background Technology

[0002] Traditional polarization-based 3D imaging techniques are limited by azimuth ambiguity and zenith angle multi-value problems, resulting in low accuracy of 3D normal reconstruction. Although combining multi-view, photometric stereo, binocular vision, and depth camera technologies with polarization-based 3D imaging can effectively improve the accuracy of 3D normal reconstruction, these methods involve complex imaging processes and face challenges such as image matching difficulties and detection distance limitations.

[0003] With the continuous development of deep learning in the field of polarization imaging, data-driven polarization 3D imaging methods have gradually attracted widespread attention. One study proposed a deep learning-based polarization 3D face reconstruction method. This method combines a pre-trained 3D shape model (3DMM) to directly estimate a coarse depth map of the face from the polarization image, and then corrects the blurred face normals calculated by physical methods. However, this method is mainly suitable for 3D reconstruction of facial features and is difficult to adapt to other complex objects. Another study introduced viewing encoding into the normal prediction network model to handle the non-orthogonal projection problem in polarization 3D imaging; however, it uses a dual-device imaging system with a polarization camera and a depth camera for normal data acquisition, resulting in a cumbersome processing flow. A third study proposed a design based on physical confidence priors to address the noise interference differences between transmission and reflection components in transparent objects, and combined this with a polarization angle loss function to mitigate the influence of the transmission component. This method uses four blurred normal priors under specular reflection, along with images of polarization degree and polarization angle, as input to a neural network model. While this design has achieved significant results in transparent object scenes, the performance of normal reconstruction may degrade to some extent when dealing with target scenes dominated by diffuse reflection. Another study designed a polarization 3D normal reconstruction method based on a U-shaped generative adversarial network; however, this method inputs four blurred normals under the specular reflection model and two blurred normals under the diffuse reflection model into the neural network for training. Excessive redundant normal prior information may make the network overly reliant on specific features. A method proposes a polarization 3D imaging system combining a polarizer and an event camera, then uses deep learning for 3D normal reconstruction. However, relying on rotating polarizers still has limitations. In high-speed motion or dynamic scenes, rotating polarizers may not be able to capture all polarization information in real time, leading to data loss or distortion.

[0004] In recent years, deep learning-based polarization 3D imaging has significantly improved the accuracy of 3D normal reconstruction by introducing neural networks to deeply explore polarization characteristics under single-view conditions. However, due to the lack of sufficient physical prior information, neural networks still face difficulties in handling complex scenes such as high reflectivity and multiple materials. Therefore, exploring an effective balance between establishing physical prior information about the target's reflected light field and the powerful modeling capabilities of neural networks can provide new ideas for achieving more accurate and adaptable single-view high-dimensional imaging schemes. Summary of the Invention

[0005] The purpose of this invention is to provide a polarization spectral three-dimensional imaging method and device using intelligent algorithms. By utilizing the physical characteristics of polarization spectroscopy, the fuzzy prior information of normals in specular reflection and diffuse reflection is corrected. The physical characteristics of light intensity, polarization, and spectrum of the reflected light field are fully integrated to improve the reconstruction accuracy of the three-dimensional normals. By combining the inputs of polarization information, prior normals, and spectral information, the reconstruction method is more conducive to handling complex scenes such as high reflectivity and mixed materials.

[0006] To achieve the above objectives, this invention provides a polarization spectral three-dimensional imaging method utilizing intelligent algorithms, comprising the following steps: Acquire 3D models and polarization spectral images, and process the 3D models and polarization spectral images to obtain true normals, combined polarization information, prior normal information, and spectral information; A physically constrained three-dimensional normal reconstruction model based on polarization spectrum is constructed by combining polarization information, prior normal information, and spectral information. The optimal three-dimensional normal reconstruction model is obtained by training the model with real normals. A loss function and evaluation index are set to evaluate the model. The polarization spectral information of the target scene is input into the polarization spectral 3D normal reconstruction model to generate the predicted normal of the target scene, and the 3D shape of the target is reconstructed based on the predicted normal.

[0007] Preferably, a physically constrained three-dimensional polarization-spectral normal reconstruction model is constructed using the true normal, combined polarization information, prior normal information, and spectral information, including... The combined polarization information, normal prior information, and spectral information are encoded separately, and normal prior features, spectral features, and polarization information features are extracted. The prior features of the normal, the spectral features, and the polarization information features are blurred multiple times; The features obtained from each blurring, the prior features of the normal, the spectral features, and the polarization information features are fused and then sharpened to obtain the predicted normal features. The predicted normal features are then decoded to obtain the predicted normal.

[0008] Preferably, the features obtained from each blurring process, multi-scale features, prior normal features, spectral features, and polarization information features are fused and then sharpened to obtain the predicted normal features, including... The features obtained by blurring the prior features, spectral features and polarization information features each time are fused to obtain the first normal feature. The features of the last layer of blurring are fused through the maximum fusion layer to obtain the second normal feature. The second normal feature and the first normal feature obtained from the previous layer are sharpened to obtain the third normal feature. The third normal feature is successively fused with the first normal feature obtained from the previous layer to obtain the final predicted normal feature.

[0009] Preferably, the expression for the loss function is: ; In the formula, W and H These represent the width and height of the image, respectively. n i,j Indicates the image in pixels ( i , j The true normal vector at point (), and Let be the reconstructed normal vector of the pixel. The cosine loss function quantifies the accuracy of the prediction result by measuring the cosine value of the angle between the true normal vector and the reconstructed normal vector.

[0010] Preferably, the evaluation indicators include the average angle error, the median error angle, and the root mean square error; And the ratio of the number of samples with an error angle ≤11.25° to the total number of samples evaluated, the ratio of the number of samples with an error angle ≤22.5° to the total number of samples evaluated, and the ratio of the number of samples with an error angle ≤30° to the total number of samples evaluated.

[0011] Preferably, the blurring process includes The prior features of normal lines, spectral features, and polarization information are first compressed using a max pooling layer, and then multiple convolutional layers are set up for feature extraction.

[0012] Preferably, the sharpening process includes The features from the previous layer's sharpening process are processed using bilinear interpolation to match the resolution of the current fusion layer's output. Then, these features are integrated again with the fusion features obtained from the current layer.

[0013] Preferably, the features from the last layer of blurring are fused through a maximum fusion layer to obtain the second normal features, including... After normalizing the prior features of normals, spectral features, and polarization information features processed by the last layer of fuzzification, they are calculated independently and in different subspaces. The results of independent calculations, calculations in different subspaces, and normalization are then concatenated and fused. Finally, the global information of each branch is effectively fused through two convolutional layers to obtain the second normal feature.

[0014] Preferably, the three-dimensional model and polarization spectral image are processed to obtain the true normal, combined polarization information, prior normal information, and spectral information, including... Align the 3D model with the 3D mesh and the 2D image to render realistic normals; The polarization spectral image is analyzed to obtain combined polarization information, prior normal information, and spectral information.

[0015] A polarization spectral three-dimensional imaging device utilizing intelligent algorithms, comprising: The data acquisition module is used to acquire 3D models and polarization spectral images; The data processing module is used to process the 3D model and polarization spectrum image to obtain the real normal, and combine polarization information, normal prior information and spectral information; The model building module is used to construct a physically constrained three-dimensional polarization-spectral normal reconstruction model by combining polarization information, prior normal information, and spectral information. The model training module is used to train the polarization spectrum 3D normal reconstruction model using real normals to obtain the optimal polarization spectrum 3D normal reconstruction model. The evaluation module is used to set the loss function and evaluation metrics to evaluate the model; The prediction module is used to input the polarization spectrum information of the target scene into the polarization spectrum three-dimensional normal reconstruction model to generate the predicted normal of the target scene, and reconstruct the three-dimensional shape of the target based on the predicted normal.

[0016] Therefore, this invention employs the aforementioned polarization spectral three-dimensional imaging method and apparatus utilizing intelligent algorithms, combining the guidance of prior physical information about the target's reflected light field with the powerful modeling capabilities of neural networks, effectively improving the reconstruction accuracy of three-dimensional normals. Compared to traditional polarization three-dimensional normal reconstruction methods, this invention adopts a data-driven approach, significantly improving normal accuracy and generalization ability. Compared to existing deep learning-based polarization three-dimensional reconstruction methods, this invention has the following technical advantages: The dataset has a higher information dimension. The dataset constructed in this invention further introduces spectral information acquisition on the basis of existing polarization 3D reconstruction methods, expanding the data dimension and providing more comprehensive and solid data support for subsequent 3D normal reconstruction based on polarization spectrum.

[0017] More comprehensive physical information guidance. This invention acquires polarization spectral information through a polarization spectral imaging device, and then uses the characteristics of polarization spectral analysis to correct the prior information of the physical model's normals. The corrected specular reflection normal information and diffuse reflection normal information are used as physical constraints, enabling the network model to better cope with complex scenes.

[0018] More dimensions for interpreting light field information. This invention utilizes neural networks to synthesize the intensity, polarization, and spectral characteristics of the light field. It designs a combination of polarization information, normal priors, and spectral information as inputs to the network model, enabling 3D normal reconstruction to consider multi-dimensional information of the light field and improving the accuracy of 3D normal reconstruction.

[0019] Multimodal Information Fusion Network. Unlike previous deep learning-based polarization 3D normal reconstruction methods, this approach treats polarization-spectral 3D normal reconstruction as a multimodal fusion problem, introducing a fusion physical neural network architecture. Three independent encoders are employed to extract features from prior normal information, spectral information, and polarization information, thereby enhancing the robustness of the 3D normal reconstruction model. Attached Figure Description

[0020] Figure 1 For physical constraint-based 3D normal reconstruction network; Figure 2 Structural design for downsampling Figure 3 Structural design for MSAFusion; Figure 4 For upsampling structural design; Figure 5 For comparison of normal reconstruction results; Figure 5 (a) is the test scenario; Figure 5 (b) is the true normal result; Figure 5 (c) shows the Mahmoud normal reconstruction result; Figure 5 (d) shows the result of the Miyazaki normal reconstruction; Figure 5 (e) shows the DeepSfP normal reconstruction result; Figure 5 (f) SfPW represents the result of normal reconstruction; Figure 5 (g) TransSfP is the truth value of the normal; Figure 5 (h) represents the test results of this method; Figure 6 This is a diagram showing the error in normal reconstruction. Figure 6 (a) is a map showing the reconstruction error of the Mahmoud normal; Figure 6 (b) shows the error map of the Miyazaki normal reconstruction; Figure 6 (c) is the DeepSfP normal reconstruction error map; Figure 6 (d) is the SfPW normal reconstruction error map; Figure 6(e) is the TransSfP normal reconstruction error map; Figure 6 (f) is the error map of normal reconstruction using this method; Figure 7 For 3D information reconstruction based on polarization spectrum; Figure 7 (a) The three-dimensional normal reconstruction result of an apple-shaped object and the integral of the reconstructed normal are shown in sequence when the angle difference is 7.15°; Figure 7 (b) The three-dimensional reconstruction results of the front-to-back positional relationship between the target arm and abdomen and the reconstruction normal integral are shown in sequence when the angle difference is 6.76°. Figure 8 For 3D information reconstruction based on polarization spectrum; Figure 8 (c) The three-dimensional reconstruction results of different animals and the reconstruction normal integral are shown in turn when the angle difference is 9.00°; Figure 8 (d) The three-dimensional reconstruction results of different animals and the integral of the reconstruction normal are shown in turn when the angle difference is 8.31°; Figure 9 3D information analysis for transparent materials: Figure 9 (a) is a transparent object; Figure 9 (b) Reconstruct the normal when the angle difference is 11.64°; Figure 9 (c) is the reconstructed point cloud map; Figure 10 for Figure 9 The corresponding reconstructed point cloud map and 3D topography map; Figure 10 (d) is the reconstructed point cloud map; Figure 10 (e) is the reconstructed three-dimensional topography. Detailed Implementation

[0021] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0022] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0023] Example 1 A polarization spectral three-dimensional imaging method utilizing intelligent algorithms includes the following steps: Step 1: Acquire the 3D model and polarization spectral images. Process the 3D model and polarization spectral images to obtain the true normal, combined polarization information, prior normal information, and spectral information; process the 3D model and polarization spectral images to obtain the true normal, combined polarization information, prior normal information, and spectral information, including... Align the 3D model with the 3D mesh and the 2D image to render realistic normals; The polarization spectral image is analyzed to obtain combined polarization information, prior normal information, and spectral information. In this embodiment, specifically... To obtain realistic normal information as a training dataset, this method uses a 3D scanner to perform 3D mesh scanning on the training samples to obtain 3D models, and then acquires polarization spectral images using a polarization spectral imaging device.

[0024] Align the 3D mesh and the 2D image in Blender software, and then render the realistic normals.

[0025] The polarization spectral information is analyzed to calculate the combined polarization information, normal prior information, and spectral information required for the network input.

[0026] Step 2: Construct a physically constrained three-dimensional normal reconstruction model of polarization spectrum using combined polarization information, prior normal information, and spectral information. Train the three-dimensional normal reconstruction model of polarization spectrum using real normals to obtain the optimal three-dimensional normal reconstruction model of polarization spectrum. Set loss function and evaluation index to evaluate the model. A physically constrained three-dimensional polarization-spectral normal reconstruction model is constructed using real normals, combined polarization information, prior normal information, and spectral information. The combined polarization information, normal prior information, and spectral information are encoded separately, and normal prior features, spectral features, and polarization information features are extracted to enhance the robustness of the 3D normal reconstruction model. The prior features of the normal, the spectral features, and the polarization information features are blurred multiple times; After fusing the features obtained from each blurring step, the prior normal features, the spectral features, and the polarization information features, a deblurring process is performed to obtain the predicted normal features. These predicted normal features are then decoded to obtain the predicted normal, including... The features obtained by blurring the prior features, spectral features and polarization information features each time are fused to obtain the first normal feature. The features of the last layer of blurring are fused through the maximum fusion layer to obtain the second normal feature. The second normal feature and the first normal feature obtained from the previous layer are sharpened to obtain the third normal feature. The third normal feature is successively fused with the first normal feature obtained from the previous layer to obtain the final predicted normal feature.

[0027] In this embodiment, as Figure 1As shown, three independent encoders are employed, each with independent parameters, to ensure that each feature extraction channel can be independently optimized based on its respective modality information. The InConv module first upsamples the input information to 64 dimensions through convolutional layers. Subsequently, the Down modules in each branch downsample the input information at each level, gradually reducing the resolution to 256, 128, 64, 32, and 16, corresponding to feature dimensions of 128, 256, 512, 1024, and 1024, respectively. The Fusion module fuses downsampled information from different modalities at the same level, enhancing the complementarity of information. In the final downsampling layer, the MSAFusion module employs a multi-head self-attention mechanism to fully capture the global feature information of each modality before fusion. This fusion mechanism, through multi-level information interaction among these three branches, enhances the model's feature representation capability. Furthermore, this method introduces the skip connection structure of U-Net to introduce features from different levels into the upsampling module, improving feature transfer efficiency and detail recovery capability. During the upsampling stage, the Up module integrates fused features from different levels, restoring the resolution to 32, 64, 128, 256, and 512 respectively. Finally, the feature dimensions are adjusted through the OutConv convolutional layer to output the predicted normals.

[0028] The blurring process involves first compressing the prior features of the normal vectors, spectral features, and polarization information using a max-pooling layer, and then extracting the features using multiple convolutional layers. In this embodiment, the structure of the Down module is as follows: Figure 2 As shown, this module uses a max-pooling layer for downsampling, which not only reduces the image resolution but also decreases computational cost, thus improving the model's training efficiency. The two convolutional layers that follow are used to enhance feature extraction capabilities.

[0029] The sharpening process includes processing the features from the previous sharpening layer using bilinear interpolation to match the resolution of the current fusion layer's output, and then integrating it again with the fused features obtained from the current layer. In this embodiment, the Up module is designed as follows: Figure 4 As shown, this module first upsamples the output of the previous Up module using bilinear interpolation, increasing its resolution to match the output of the current fusion layer. Then, the feature information from the two different layers is concatenated and integrated through two convolutional layers, completing the upsampling operation.

[0030] The features from the last layer of blurring are fused through a maximum fusion layer to obtain the second normal features, including... After normalizing the prior normal features, spectral features, and polarization information features obtained from the final fuzzification layer, each feature is calculated independently and in different subspaces. The results of independent calculations, subspace calculations, and normalization are then concatenated and fused. Finally, two convolutional layers are used to effectively fuse the global information of each branch to obtain the second normal feature. In this embodiment, five Fusion modules using residual connections are set before the final fuzzification layer to promote local information interaction between different modalities, enhancing the coherence of information flow and fusion efficiency. The final MSAFusion structure (final fuzzification layer) is as follows: Figure 3 As shown, after Down 5 processing, the feature information of the three modalities is input into the multi-head self-attention module for computation. Each self-attention module contains eight heads, each responsible for processing different subspaces of the input information and independently performing self-attention computation. The computation results are then concatenated and output, and two convolutional layers are used to effectively fuse the global information of each branch, addressing the problem of insufficient feature fusion that may occur in local information fusion. Through this series of operations, MSAFusion can more comprehensively capture the correlation between different modalities, improving the expressive power of features and the overall fusion performance.

[0031] The expression for the loss function is: ; In the formula, W and H These represent the width and height of the image, respectively. n i,j Indicates the image in pixels ( i , j The true normal vector at point (), and Let be the reconstructed normal vector of the pixel. The cosine loss function quantifies the accuracy of the prediction result by measuring the cosine value of the angle between the true normal vector and the reconstructed normal vector.

[0032] The evaluation metrics include mean angle error (Mean), median angle error (Median), and root mean square error (RMSE), and the specific calculation formulas are shown below: ; ; In the formula, θ k This represents the angle of error between the reconstructed normal vector and the true normal vector at a certain pixel. M This represents the total number of evaluation samples, and `median()` represents the operation of taking the median from the statistical data. In quantitative comparisons, the smaller the values ​​of the mean angle error, the median error angle, and the root mean square error, the better the 3D normal reconstruction effect.

[0033] The ratios of samples with an error angle ≤11.25° to the total number of evaluation samples, the ratios of samples with an error angle ≤22.5° to the total number of evaluation samples, and the ratios of samples with an error angle ≤30° to the total number of evaluation samples are also included. These three distribution indicators reflect the degree of agreement between the reconstructed normal vector and the true normal vector; the higher the value, the better the 3D normal reconstruction effect.

[0034] Step 3: Input the polarization spectrum information of the target scene into the polarization spectrum 3D normal reconstruction model to generate the predicted normal of the target scene, and reconstruct the 3D shape of the target based on the predicted normal.

[0035] A polarization spectral three-dimensional imaging device utilizing intelligent algorithms, comprising: The data acquisition module is used to acquire 3D models and polarization spectral images; The data processing module is used to process the 3D model and polarization spectrum image to obtain the real normal, and combine polarization information, normal prior information and spectral information; The model building module is used to construct a physically constrained three-dimensional polarization-spectral normal reconstruction model by combining polarization information, prior normal information, and spectral information. The model training module is used to train the polarization spectrum 3D normal reconstruction model using real normals to obtain the optimal polarization spectrum 3D normal reconstruction model. The evaluation module is used to set the loss function and evaluation metrics to evaluate the model; The prediction module is used to input the polarization spectrum information of the target scene into the polarization spectrum three-dimensional normal reconstruction model to generate the predicted normal of the target scene, and reconstruct the three-dimensional shape of the target based on the predicted normal.

[0036] Example 2 The proposed polarization spectrum 3D normal reconstruction model was implemented on the PyTorch platform, and an NVIDIA GeForce RTX 4090 GPU was used for training. To reduce GPU memory usage, the image size was adjusted to 512×512 during data processing. The network was optimized using a stochastic gradient descent (SGD) optimizer with an initial learning rate of 10. -6 The momentum factor is 0.9, and the weight decay coefficient is 5×10. -4 In addition, a step-down learning rate strategy was introduced, where the learning rate is dynamically adjusted by multiplying by 0.1 every 10 training cycles.

[0037] To verify the accuracy of the polarization spectrum three-dimensional normal reconstruction algorithm, this method incorporates two traditional physical calculations and three deep learning-based polarization three-dimensional normal reconstruction algorithms for comprehensive testing and analysis. Figure 5The normal reconstruction results under different scenarios intuitively show the performance of each method in terms of normal reconstruction accuracy and adaptability.

[0038] Next, mainly analyze and evaluate four deep learning-based 3D normal reconstruction algorithms.

[0039] Figure 5 The first column shows a decorative item made of gypsum. However, the raised "Fu" character on its surface is painted with acrylic paint, which has a significant difference from the gypsum material. At the same time, there are drastic gradient changes and complex detail features in areas such as the "Fu" character and the target's mouth, which also pose challenges to 3D normal reconstruction. Due to gradient jumps and material interference, DeepSfP fails to reconstruct the normal information of the "Fu" character area, and the complex features of the target's mouth cannot be effectively presented. Although SfPW and TransSfP have improved in the normal reconstruction of mouth details, the interference of the mixed material still has a greater impact, resulting in an unsatisfactory normal reconstruction effect in the "Fu" character area. The proposed polarization spectroscopy 3D normal reconstruction method in this paper integrates polarization, spectroscopy, and normal prior constraints. Guided by multiple physical prior information, it significantly improves the normal reconstruction accuracy of complex details such as the target's mouth and effectively suppresses the interference of the mixed material at the edge of the "Fu" character.

[0040] The polyethylene apple in the second column, the resin alien in the sixth column, and the decorative item made of a mixture of ceramic and glass in the seventh column are respectively affected by obvious high-reflection interference in the center of the apple, the face of the alien, and the waist area of the decorative item. These high-reflection scenarios cause distortions in the 3D normal reconstruction results of the DeepSfP, SfPW, and TransSfP algorithms. In addition, some complex detail areas with drastic gradient changes also inhibit the normal reconstruction accuracy of these methods. The algorithm proposed in this paper corrects the zenith angle ambiguity during specular reflection by considering polarization spectroscopy information, thus achieving more accurate 3D normal reconstruction in high-reflection and complex detail areas.

[0041] The surface colors of the scenes in the third, fourth, and fifth columns vary greatly, resulting in uneven reflectivity of the target features. The Squirtle decorative item in the fourth column, as a representative scene of mixed colors, has relatively smooth surfaces in the eye and mouth areas, and the normal changes should be relatively gentle. However, due to the interference of non-uniform reflectivity in the eye area, obvious errors occur in the normal reconstruction of DeepSfP and SfPW in this area, with gradient jumps. Although TransSfP has improved the reconstruction result in the eye area, the interference in areas such as the black whiskers and abdominal patterns of the mouth still affects the normal reconstruction accuracy. The proposed multi-modal neural network in this paper effectively addresses the reflectivity interference in the eye and mouth areas by learning the differences in polarization spectroscopy characteristics, enabling the reconstruction algorithm to maintain a high normal reconstruction accuracy in complex texture areas.

[0042] To further visualize the 3D normal reconstruction errors of different algorithms, this method calculates... Figure 5 The error angle between the reconstructed normal and the true normal in each scene is shown in the following figures. Figure 6 As shown in the normal reconstruction error map, the darker the purple area, the greater the normal reconstruction error.

[0043] Error visualization plots are helpful in showing the differences in normal reconstruction accuracy of different algorithms in various regions, providing an important evaluation basis for the performance analysis of 3D normal reconstruction algorithms. Overall, the reconstruction results of this method are presented in lighter colors in the error plot, indicating that its normal reconstruction quality is relatively high. Taking the seventh column as an example, it can be observed that all algorithms have 3D normal reconstruction errors in the head region of the target. Excluding the large errors of traditional algorithms, SfPW has the deepest purple color, indicating that the reconstruction error in this region is more serious. TransSfP has the widest purple area coverage, indicating that the number of pixels with errors in this region is the largest. The reconstruction results of DeepSfP and the proposed algorithm are lighter purple in the head region, indicating relatively low reconstruction errors. However, due to the interference of high reflectivity, DeepSfP, SfPW, and TransSfP all show obvious purple areas near the waist, with SfPW being the most severely affected by interference. In contrast, the reconstruction algorithm proposed in this method has almost no purple area distribution in the waist, maintaining high 3D normal reconstruction accuracy.

[0044] To quantitatively evaluate the performance of different algorithms in 3D normal reconstruction, this method uses the six evaluation metrics proposed in the previous section to conduct a comprehensive quantitative analysis of the reconstructed normals of each algorithm. The test samples cover complex structures, mixed materials, and highly reflective scenes, ensuring the breadth and effectiveness of the test objects. Table 1 shows the quantitative results of 3D normal reconstruction of different algorithms in the test samples. The polarization spectral 3D normal reconstruction proposed in this method shows the best reconstruction accuracy in the quantitative test. Compared with the five existing algorithms, the proposed reconstruction algorithm shows significant improvements in all accuracy metrics. The mean angle error (Mean) is reduced by 71.64%, 67.02%, 24.33%, 19.41%, and 27.97% compared to Mahmoud, Miyazaki, DeepSfP, SfPW, and TransSfP, respectively. In terms of the median angle error (Median), the proposed algorithm reduces the performance of these five algorithms by 74.52%, 64.81%, 25.75%, 20.16%, and 27.21%, respectively. Furthermore, the root mean square error (RMSE) was also reduced by 68.34%, 67.41%, 21.55%, 18.41%, and 27.57%, respectively. Taking the distribution index of the lowest error angle of 11.25° as an example, the algorithm of this method improved by 635.67%, 134.88%, 26.24%, 15.30%, and 24.58% compared with Mahmoud, Miyazaki, DeepSfP, SfPW, and TransSfP, respectively. The comparison results of various indicators show that the polarization spectrum three-dimensional normal reconstruction algorithm exhibits excellent performance in terms of reconstruction accuracy.

[0045] Table 1 Comparison of 3D Normal Reconstruction Accuracy

[0046] 3D reconstruction imaging applications Polarized optical 3D reconstruction technology can effectively acquire the 3D normal information of the target surface, and the normal information directly characterizes the geometric shape features of the target. Next, this method will reconstruct the 3D shape features of the target object by integrating the normal data. The Shaplets algorithm, based on wavelet reconstruction, has certain filtering capabilities and exhibits good anti-interference ability in normal integration reconstruction. Therefore, this method uses the Shaplets algorithm to globally integrate the normal information to reconstruct the 3D shape of the target. Figure 7 and Figure 8 The results of 3D reconstruction for different scenarios are shown. For each scenario, the original image, normal information, 3D topography, and the relative height change curve of the 3D topography corresponding to the dashed axial part of the original 2D image are displayed.

[0047] Figure 7(a) shows the 3D reconstruction result of an apple-shaped object. Within the light-colored box in the lower left corner of the object, there is a dent of a certain depth. Due to interference from environmental factors such as lighting, the 2D image cannot clearly display the geometric features of the dent. However, the gradient change in this area and its difference from the surrounding smooth surface can be effectively perceived from the reconstructed normal information. The 3D coordinate system displays the 3D topography information reconstructed through normal integration, which essentially achieves a visual representation of the target's geometric features. Notably, the relative depth change of the dent is clearly presented in the lower left corner of the 3D topography. This visualization imaging effect provides important reference value for fields such as 3D defect detection.

[0048] Figure 7 (b) The imaging study focused on the anterior-posterior positional relationship between the target arm and abdomen. Due to lighting and viewing angle limitations, it is difficult to distinguish the anterior-posterior positions of the arm and abdomen in two-dimensional images. However, in the three-dimensional topography after normal integration, the surface geometry of the object is clearly presented, successfully revealing the anterior-posterior positional difference between the arm and abdomen, effectively improving the recognizability of the target's positional features. At the same time, the rightmost axial curve accurately depicts the relative height changes of the abdominal region. Figure 8 (c) and Figure 8 (d) shows the imaging results of different animal shapes. After 3D information reconstruction, the geometric shape information of the target object is fully displayed. These reconstruction results not only verify the effectiveness of the polarization spectral 3D information reconstruction technology used in this method, but also lay a solid foundation for subsequent single-view high-dimensional imaging applications based on polarization spectroscopy.

[0049] Figure 9 and Figure 10This paper demonstrates the process of 3D information analysis for transparent objects. A deep pit defect exists on the left side of the original object, but due to the interference of the transparent material, the intensity image failed to effectively identify the defect features. After reconstructing the object's normal information using polarization spectroscopy, the contrast of the pit region's features in the normal map is significantly enhanced, effectively revealing the geometric information of this region. Further analysis of the reconstructed point cloud map reveals significant changes in the overall contour information. In the magnified view of the red-framed area of ​​the point cloud map, the point cloud data of the target region clearly presents the pit feature, validating the application potential of polarization spectroscopy 3D information reconstruction technology in transparent material scenes. Subsequently, by integrating the reconstructed normals, the resulting 3D topography map further reveals the geometric morphology of the pit, significantly improving the detection and recognition capabilities of target features. The relative height variation curve on the right is extracted from the 3D topography along the pink axis. In the curve, the dark red segment represents the depth variation of the pit region, showing a trend of first decreasing and then increasing. This depth variation information effectively characterizes the geometric properties of the pit, providing important evidence for accurate defect analysis.

[0050] Therefore, the present invention adopts the above-mentioned polarization spectrum three-dimensional imaging method and device using intelligent algorithms. By utilizing the physical characteristics of polarization spectrum, the blurred normal prior information of specular reflection and diffuse reflection is corrected. The physical characteristics of light intensity, polarization, and spectrum of reflected light field are fully integrated to improve the reconstruction accuracy of three-dimensional normal. By combining the input of polarization information, normal prior and spectral information, the reconstruction method is more conducive to handling complex scenes such as high reflectivity and mixed materials.

[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A polarization spectral three-dimensional imaging method utilizing intelligent algorithms, characterized in that, Includes the following steps: Acquire 3D models and polarization spectral images, and process the 3D models and polarization spectral images to obtain true normals, combined polarization information, prior normal information, and spectral information; A physically constrained three-dimensional normal reconstruction model based on polarization spectrum is constructed by combining polarization information, prior normal information, and spectral information. The optimal three-dimensional normal reconstruction model is obtained by training the model with real normals. A loss function and evaluation index are set to evaluate the model. The polarization spectral information of the target scene is input into the polarization spectral 3D normal reconstruction model to generate the predicted normal of the target scene, and the 3D shape of the target is reconstructed based on the predicted normal.

2. The polarization spectral three-dimensional imaging method using intelligent algorithms according to claim 1, characterized in that, A physically constrained three-dimensional polarization-spectral normal reconstruction model is constructed using real normals, combined polarization information, prior normal information, and spectral information. The combined polarization information, normal prior information, and spectral information are encoded separately, and normal prior features, spectral features, and polarization information features are extracted. The prior features of the normal, the spectral features, and the polarization information features are blurred multiple times; The features obtained from each blurring, the prior features of the normal, the spectral features, and the polarization information features are fused and then sharpened to obtain the predicted normal features. The predicted normal features are then decoded to obtain the predicted normal.

3. The polarization spectral three-dimensional imaging method using intelligent algorithms according to claim 2, characterized in that, After fusing the features obtained from each blurring process, multi-scale features, prior normal features, spectral features, and polarization information features, a sharpening process is performed to obtain the predicted normal features, including... The features obtained by blurring the prior features, spectral features and polarization information features each time are fused to obtain the first normal feature. The features of the last layer of blurring are fused through the maximum fusion layer to obtain the second normal feature. The second normal feature and the first normal feature obtained from the previous layer are sharpened to obtain the third normal feature. The third normal feature is successively fused with the first normal feature obtained from the previous layer to obtain the final predicted normal feature.

4. The polarization spectral three-dimensional imaging method using intelligent algorithms according to claim 1, characterized in that, The expression for the loss function is: ; In the formula, W and H These represent the width and height of the image, respectively. n i,j Indicates the image in pixels ( i , j The true normal vector at point (), and Let be the reconstructed normal vector of the pixel. The cosine loss function quantifies the accuracy of the prediction result by measuring the cosine value of the angle between the true normal vector and the reconstructed normal vector.

5. The polarization spectral three-dimensional imaging method using intelligent algorithms according to claim 1, characterized in that, Evaluation metrics include mean angle error, median error angle, and root mean square error. And the ratio of the number of samples with an error angle ≤11.25° to the total number of samples evaluated, the ratio of the number of samples with an error angle ≤22.5° to the total number of samples evaluated, and the ratio of the number of samples with an error angle ≤30° to the total number of samples evaluated.

6. The polarization spectral three-dimensional imaging method using intelligent algorithms according to claim 3, characterized in that, The blurring process includes The prior features of normal lines, spectral features, and polarization information are first compressed using a max pooling layer, and then multiple convolutional layers are set up for feature extraction.

7. The polarization spectral three-dimensional imaging method using intelligent algorithms according to claim 3, characterized in that, The sharpening process includes The features from the previous layer's sharpening process are processed using bilinear interpolation to match the resolution of the current fusion layer's output. Then, these features are integrated again with the fusion features obtained from the current layer.

8. The polarization spectral three-dimensional imaging method using intelligent algorithms according to claim 3, characterized in that, The features from the last layer of blurring are fused through a maximum fusion layer to obtain the second normal features, including... After normalizing the prior features of normals, spectral features, and polarization information features processed by the last layer of fuzzification, they are calculated independently and in different subspaces. The results of independent calculations, calculations in different subspaces, and normalization are then concatenated and fused. Finally, the global information of each branch is effectively fused through two convolutional layers to obtain the second normal feature.

9. A polarization spectral three-dimensional imaging method using intelligent algorithms according to claim 1, characterized in that, Processing the 3D model and polarization spectral image yields the true normal, combined polarization information, prior normal information, and spectral information, including... Align the 3D model with the 3D mesh and the 2D image to render realistic normals; The polarization spectral image is analyzed to obtain combined polarization information, prior normal information, and spectral information.

10. A polarization spectral three-dimensional imaging device utilizing intelligent algorithms, characterized in that, include The data acquisition module is used to acquire 3D models and polarization spectral images; The data processing module is used to process the 3D model and polarization spectrum image to obtain the real normal, and combine polarization information, normal prior information and spectral information; The model building module is used to construct a physically constrained three-dimensional polarization-spectral normal reconstruction model by combining polarization information, prior normal information, and spectral information. The model training module is used to train the polarization spectrum 3D normal reconstruction model using real normals to obtain the optimal polarization spectrum 3D normal reconstruction model. The evaluation module is used to set the loss function and evaluation metrics to evaluate the model; The prediction module is used to input the polarization spectrum information of the target scene into the polarization spectrum three-dimensional normal reconstruction model to generate the predicted normal of the target scene, and reconstruct the three-dimensional shape of the target based on the predicted normal.