Target polarization three-dimensional imaging method under water scattering medium condition based on deep learning
By applying a deep learning-based polarization three-dimensional imaging method in an underwater environment, combining convolutional neural networks and polarization images, the problem of three-dimensional imaging in complex underwater environments is solved, and a higher precision surface normal prediction is achieved.
Patent Information
- Application Number
- CN202510044797.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
AI Technical Summary
In complex underwater environments, it is difficult for the prior art to achieve high-precision and high-quality three-dimensional imaging, especially under different underwater depths and ambient lighting conditions, image acquisition and processing caused by noise and scattering media becomes more complex.
A deep learning-based method is adopted, combining polarization images and three-dimensional models, a convolutional neural network applied to underwater target polarization three-dimensional imaging is designed, and the surface normals of underwater targets are predicted through training and testing.
It effectively improves the accuracy of surface normal prediction, reduces texture loss during underwater target reconstruction, and obtains a higher normal prediction effect.
Smart Images

Figure CN119941994A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of optical imaging technology, and in particular relates to a method for three-dimensional polarization imaging of a target under scattering medium conditions in water based on deep learning. Background Art
[0002] Underwater high-precision optical 3D imaging has always been a focus of marine engineering research. The scattering and absorption characteristics in different underwater environments have a great influence on the propagation of light, and reconstructing high-precision 3D imaging results from low-quality and severely degraded underwater images faces great challenges. In addition, the ambient lighting and noise at different underwater depths also make the acquisition, processing and 3D reconstruction of underwater images more complicated. Under such complex conditions, it is particularly difficult to obtain high-quality images and accurate 3D information of the target. Polarized light provides a good solution to this problem with its unique propagation characteristics and information carrying capacity. Analyzing the polarization characteristics of light captured by the imaging system can provide redundant information that traditional imaging methods cannot provide, providing a new idea for achieving high-precision and high-quality 3D reconstruction underwater.
[0003] In complex underwater environments, polarization imaging has obvious advantages in enhancing image contrast and improving object recognition capabilities. Polarization imaging has also attracted attention in 3D shape reconstruction applications, using polarized light reflected from the target surface to infer surface normals and achieve accurate 3D imaging. Although polarization 3D imaging technology and the corresponding imaging datasets have been widely studied, its application in underwater environments remains a major challenge and lacks in-depth research. Summary of the invention
[0004] The purpose of the present invention is to solve the shortcomings existing in the background technology and provide a method for polarization three-dimensional imaging of targets under scattering medium conditions in water based on deep learning. Polarization images and polarization representations are input into the constructed network input port, and polarization three-dimensional imaging is performed using polarization images of four polarization states and their polarization information, which can effectively improve the accuracy of surface normal prediction.
[0005] To achieve the above object, the technical solution of the present invention is: a method for three-dimensional polarization imaging of a target under scattering medium conditions in water based on deep learning, comprising:
[0006] S1, using a polarization camera to collect a polarization image of an underwater target, and using a three-dimensional scanner to obtain a three-dimensional model of the underwater target;
[0007] S2. Align the polarization image and the three-dimensional model of the underwater target, preprocess the acquired polarization image, construct an underwater polarization image dataset, and divide it into a training set, a validation set, and a test set;
[0008] S3. Design a convolutional neural network for polarization 3D imaging of underwater targets, train the designed network with a training set to obtain the training weights of the network, and test it with a test set to obtain the predicted surface normal;
[0009] S4. Compare the predicted surface normal with the true value of the target normal, and use evaluation indicators to measure the effect of underwater target three-dimensional imaging.
[0010] In one embodiment of the present invention, the specific step of S1 is:
[0011] S11, using a polarization camera to collect polarization images of underwater targets at different angles and in underwater environments of different concentrations;
[0012] S12. Use a high-precision 3D scanner to obtain a high-quality 3D model of an underwater target.
[0013] In one embodiment of the present invention, the specific steps of S2 are:
[0014] S21, using a 3D shape-image alignment method, aligning the scanned 3D shape of the underwater target from a scanner coordinate system to an image coordinate system of a polarization camera;
[0015] S22, using a renderer to calculate the true value of the surface normal of the aligned three-dimensional shape;
[0016] S23, preprocessing the original polarization image data, calculating the polarization representation input of the network, the polarization representation consists of the total light intensity of the polarization image, the encoding space of the azimuth angle and the improved polarization degree image, and constructing an underwater polarization image dataset;
[0017] S24. Divide the underwater polarization image dataset into a training set, a validation set, and a test set in proportion.
[0018] In one embodiment of the present invention, the specific steps of S3 are:
[0019] S31. Use an encoder that reduces spatial resolution and a decoder that restores spatial resolution, combined with a channel attention mechanism module to form a network structure of a convolutional neural network applied to polarization 3D imaging of underwater targets, and use the original polarization image and polarization representation as input to the network;
[0020] S32, training the designed network using a training set in the underwater polarization image data set to obtain a training weight of the network;
[0021] S33. Use the test set in the underwater polarization image dataset to test the trained network to obtain the final underwater target surface normal prediction result.
[0022] In one embodiment of the present invention, the evaluation index adopts the average angle error, which is expressed as follows:
[0023]
[0024] Among them, p GT is the angle of the normal line of the true normal map, p * is the angle of the predicted normal map normal.
[0025] In one embodiment of the present invention, the total light intensity of the polarization image is:
[0026]
[0027] Among them, I 0 ,I 45 ,I 90 ,I 135 They are the polarization components of the polarization images captured by the polarization camera in the directions of 0°, 45°, 90°, and 135°, respectively.
[0028] In one embodiment of the present invention, the coding space of the azimuth angle is:
[0029] φ e =(cos2φ,sin2φ)
[0030] Where f is the polarization phase angle.
[0031] In one embodiment of the present invention, the improved polarization degree image is:
[0032]
[0033] Among them, S 1 is the light intensity difference between the polarization component in the 0° direction and the polarization component in the 90° direction; S 2 It is the light intensity difference between the polarization component in the 45° direction and the polarization component in the 135° direction.
[0034] In one embodiment of the present invention, in step S3, the convolutional neural network applied to polarized three-dimensional imaging of underwater targets uses an encoder module for reducing spatial resolution and a decoder module for restoring spatial resolution to form a basic network architecture. Each encoder module and decoder module uses residual U blocks of different depths, and the residual U blocks are used to replace a single convolutional layer or deconvolution layer operation; jump connections are used between blocks at the same level; and a channel attention mechanism is introduced in the decoder module structure.
[0035] The present invention also provides a computer-readable storage medium, on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the method steps described above can be implemented.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] 1. The present invention constructs a polarization image dataset of underwater targets, including polarization images of four polarization states in five different depths of underwater environments simulating Jerlov type I water bodies and real surface normal maps of corresponding underwater targets obtained using a high-precision three-dimensional scanner;
[0038] 2. The present invention proposes a method combining deep learning and polarization 3D imaging, which combines the U2Net network framework with the channel attention mechanism and effective polarization feature representation, which can deeply mine the features of underwater targets, improve the surface normal recovery capability and better preserve the target texture details;
[0039] 3. To address the π ambiguity problem of polarization azimuth, the present invention introduces an azimuth coding space and improves the original polarization degree formula to form a polarization representation input into the proposed network model, thereby improving the extraction rate of texture information during underwater target reconstruction and obtaining a higher normal prediction effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 A flowchart of a method for polarized three-dimensional imaging of a target under scattering medium conditions in water based on deep learning provided by an embodiment of the present invention;
[0041] Figure 2 A schematic diagram of the convolutional neural network structure provided by an embodiment of the present invention;
[0042] Figure 3 A schematic diagram of a residual U block provided in an embodiment of the present invention;
[0043] Figure 4 A schematic diagram of a channel attention module provided in an embodiment of the present invention;
[0044] Figure 5 A comparison chart of the three-dimensional imaging effects of the method provided in an embodiment of the present invention and the existing method. DETAILED DESCRIPTION
[0045] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.
[0046] The present invention provides a method for three-dimensional polarization imaging of a target under scattering medium conditions in water based on deep learning, comprising:
[0047] S1, using a polarization camera to collect a polarization image of an underwater target, and using a three-dimensional scanner to obtain a three-dimensional model of the underwater target;
[0048] S2. Align the polarization image and the three-dimensional model of the underwater target, preprocess the acquired polarization image, construct an underwater polarization image dataset, and divide it into a training set, a validation set, and a test set;
[0049] S3. Design a convolutional neural network for polarization 3D imaging of underwater targets, train the designed network with a training set to obtain the training weights of the network, and test it with a test set to obtain the predicted surface normal;
[0050] S4. Compare the predicted surface normal with the true value of the target normal, and use evaluation indicators to measure the effect of underwater target three-dimensional imaging.
[0051] The following is the specific implementation process of the present invention.
[0052] Example 1
[0053] like Figure 1 As shown, the present invention provides a method for polarization three-dimensional imaging of a target under scattering medium conditions in water based on deep learning, comprising the following steps:
[0054] S1, using a polarization camera to collect a polarization image of an underwater target, and using a three-dimensional scanner to obtain a three-dimensional model of the underwater target;
[0055] Specifically, a polarization camera is used to collect clear polarization images of the target underwater. In order to obtain underwater environments of different depths simulating Jerlov type I water bodies, 10ml, 20ml, 40ml, and 60ml of blue pigment are added to the water in sequence. Then, polarization images of the target at different angles in underwater environments simulating different depths are collected, and a high-precision 3D scanner based on structured light is used to obtain a high-quality 3D model of the underwater target.
[0056] S2, aligning the polarization image and the three-dimensional model of the underwater target, and preprocessing the acquired polarization image to construct an underwater polarization image dataset;
[0057] Specifically, the specific process of constructing an underwater polarization image dataset in the present invention is as follows:
[0058] S21, using a 3D shape-image alignment method, aligning the scanned 3D shape of the underwater target from a scanner coordinate system to an image coordinate system of a polarization camera;
[0059] S22, using a renderer to calculate the true value of the surface normal of the aligned three-dimensional shape;
[0060] S23, preprocessing the original polarization image data, calculating the polarization representation input of the network model, the polarization representation consists of the total light intensity of the polarization image, the encoding space of the azimuth angle and the improved polarization degree image, and constructing the underwater target polarization data set;
[0061] Specifically, the total light intensity of the polarization image used in the present invention is:
[0062]
[0063] Among them, I 0 ,I 45 ,I 90 ,I 135 They are the polarization components of the polarization images captured by the polarization camera in the directions of 0°, 45°, 90°, and 135° respectively;
[0064] Specifically, the coding space used in the present invention is:
[0065] φ e =(cos2φ,sin2φ),
[0066] Where, f is the polarization phase angle;
[0067] Specifically, the polarization degree image used in the present invention is:
[0068]
[0069] Among them, S 1 is the light intensity difference between the polarization component in the 0° direction and the polarization component in the 90° direction; S 2 is the light intensity difference between the polarization component in the 45° direction and the polarization component in the 135° direction;
[0070] S24. Divide the dataset into training set, validation set and test set in a ratio of 4:1:1.
[0071] S3. Design a convolutional neural network for polarized 3D imaging of underwater targets, train the network with a training set to obtain the training weights of the network model, and finally test the test set to obtain the predicted surface normal;
[0072] Specifically, the convolutional neural network framework constructed by the present invention is as follows: Figure 2 As shown, the specific method flow is:
[0073] S31, using an encoder that reduces the spatial resolution and a decoder that restores the spatial resolution, each encoder and decoder module uses a residual U block with different depths, such as Figure 3 As shown in Figure 1, the residual U-block mainly consists of three components: the first is the input convolution layer, which converts the input feature map into an intermediate feature map. The second is the U-Net encoder-decoder structure, which is used to extract and encode multi-scale context information. The third is the number of channels M in the internal layer, which encodes the depth and rich local and global features of the residual U-block. The residual U-block is used to replace a single convolutional layer or deconvolution layer operation, combined with the following Figure 4The channel attention mechanism module shown constitutes the network structure, taking the original polarization image and polarization representation as the input of the convolutional neural network;
[0074] S32, using the training set in the data set to train the designed convolutional neural network to obtain the training weights of the network model;
[0075] S33, using the test set in the data set to test the trained convolutional neural network model, obtain the final underwater target surface normal prediction result, and compare the results of this method with those of other methods, such as Figure 5 As shown in the figure, this method can effectively solve the azimuth ambiguity problem, reduce the texture loss in the underwater target reconstruction process, and improve the accuracy of surface normal prediction;
[0076] S4. Compare the predicted surface normal with the true value of the target normal, and use the average angle error index to measure the effect of underwater target three-dimensional imaging.
[0077] Specifically, the average angle error used in the present invention is:
[0078]
[0079] Among them, p GT is the angle of the normal line of the true normal map, p * is the angle of the predicted normal map normal.
[0080] Furthermore, the comparison results of the average angle error index of the underwater target three-dimensional imaging method of the present invention and other methods in the test data set are shown in Table 1.
[0081] Table 1
[0082] Method MAE Miyazaki 54.18° Mahmoud 48.00° Lei 10.03° Ba 9.84° Wu 9.51° This method 8.43°
[0083] In Table 1, MAE (Mean Angular Error) represents the average angular error. The smaller the index is, the closer the predicted surface normal map is to the true value of the surface normal.
[0084] The present invention also provides a computer-readable storage medium, on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the method steps described above can be implemented.
[0085] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.
Claims
1. A method for polarization three-dimensional imaging of a target in an underwater scattering medium based on deep learning, characterized in that: include: S1, using a polarization camera to collect a polarization image of an underwater target, and using a three-dimensional scanner to obtain a three-dimensional model of the underwater target; S2. Align the polarization image and the three-dimensional model of the underwater target, preprocess the acquired polarization image, construct an underwater polarization image dataset, and divide it into a training set, a validation set, and a test set; S3. Design a convolutional neural network for polarization 3D imaging of underwater targets, train the designed network with a training set to obtain the training weights of the network, and test it with a test set to obtain the predicted surface normal; S4. Compare the predicted surface normal with the true value of the target normal, and use evaluation indicators to measure the effect of underwater target three-dimensional imaging.
2. According to the method of claim 1, the method of three-dimensional polarization imaging of a target in an underwater scattering medium based on deep learning is characterized in that: The specific steps of S1 are: S11, using a polarization camera to collect polarization images of underwater targets at different angles and in underwater environments of different concentrations; S12. Use a high-precision 3D scanner to obtain a high-quality 3D model of an underwater target.
3. According to the method of claim 1, wherein the method comprises: The specific steps of S2 are: S21, using a 3D shape-image alignment method, aligning the scanned 3D shape of the underwater target from a scanner coordinate system to an image coordinate system of a polarization camera; S22, using a renderer to calculate the true value of the surface normal of the aligned three-dimensional shape; S23, preprocessing the original polarization image data, calculating the polarization representation input of the network, the polarization representation consists of the total light intensity of the polarization image, the encoding space of the azimuth angle and the improved polarization degree image, and constructing an underwater polarization image dataset; S24. Divide the underwater polarization image dataset into a training set, a validation set, and a test set in proportion.
4. The method for polarization three-dimensional imaging of a target in an underwater scattering medium based on deep learning according to claim 3, characterized in that: The specific steps of S3 are: S31. Use an encoder that reduces spatial resolution and a decoder that restores spatial resolution, combined with a channel attention mechanism module to form a network structure of a convolutional neural network applied to polarization 3D imaging of underwater targets, and use the original polarization image and polarization representation as input to the network; S32, training the designed network using a training set in the underwater polarization image data set to obtain a training weight of the network; S33. Use the test set in the underwater polarization image dataset to test the trained network to obtain the final underwater target surface normal prediction result.
5. The method for polarization three-dimensional imaging of a target in an underwater scattering medium based on deep learning according to claim 1, characterized in that: The evaluation index adopts the average angle error.
6. The method for polarization three-dimensional imaging of a target in an underwater scattering medium based on deep learning according to claim 3, characterized in that: The total light intensity of the polarization image is: Among them, I0, I 45 ,I 90 ,I 135 They are the polarization components of the polarization images captured by the polarization camera in the directions of 0°, 45°, 90°, and 135°, respectively.
7. The method for polarization three-dimensional imaging of a target in an underwater scattering medium based on deep learning according to claim 3, characterized in that: The encoding space of the azimuth is: f e =(cos2φ,sin2φ) Where f is the polarization phase angle.
8. The method for polarization three-dimensional imaging of a target in an underwater scattering medium based on deep learning according to claim 3, characterized in that: The improved polarization degree image is: Among them, S1 is the light intensity difference between the polarization component in the 0° direction and the polarization component in the 90° direction; S2 is the light intensity difference between the polarization component in the 45° direction and the polarization component in the 135° direction.
9. The method for polarization three-dimensional imaging of a target in an underwater scattering medium based on deep learning according to claim 1, characterized in that: In step S3, the convolutional neural network applied to polarized three-dimensional imaging of underwater targets uses an encoder module for reducing spatial resolution and a decoder module for restoring spatial resolution to form a basic network architecture. Each encoder module and decoder module uses residual U blocks of different depths, and the residual U blocks are used to replace a single convolutional layer or deconvolution layer operation; jump connections are used between blocks at the same level; and a channel attention mechanism is introduced in the decoder module structure.
10. A computer-readable storage medium having stored thereon computer program instructions that can be executed by a processor, and when the processor executes the computer program instructions, the method steps according to any one of claims 1 to 9 can be implemented.