A large scene polarization three-dimensional imaging method based on monocular depth estimation
By combining a monocular depth estimation network with polarization images, the problem of normal vector singularity in large-scene polarization 3D reconstruction is solved, achieving efficient 3D imaging with a simplified system structure.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIDIAN UNIV
- Filing Date
- 2023-02-27
- Publication Date
- 2026-05-19
AI Technical Summary
In existing technologies, when using a single-camera system for large-scene polarization 3D reconstruction, the singularity of the normal vector causes distortion, and methods based on the Kinect depth sensor increase system complexity.
By acquiring polarization and color images of the target object at different angles, a pre-trained monocular depth estimation network is used to predict the depth image. The scene surface normal vector is determined by combining the polarization image and the preset viewing direction, and then 3D reconstruction is performed after correction.
It realizes large-scene polarization 3D imaging in a monocular system, solves the problem of normal vector singularity, simplifies the system structure and reduces hardware complexity.
Smart Images

Figure CN116363301B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of 3D reconstruction technology, specifically relating to a large-scene polarization 3D imaging method based on monocular depth estimation. Background Technology
[0002] With societal progress and advancements in computer hardware performance, 3D imaging is increasingly widely used in daily life. Among the many 3D imaging technologies available, polarization-based 3D imaging, as a passive 3D imaging technique, offers advantages such as long-range capability, high precision, low cost, and the ability to modulate without a light source. Simultaneously, with the continuous development of computer vision, the ability of neural networks to solve various nonlinear problems is gradually improving, and using neural networks to achieve monocular depth estimation of scenes can demonstrate excellent performance. However, when using a single-camera system to perform polarization-based 3D reconstruction of large scenes, the singularity of the target normal vector interpreted through polarization information leads to severe distortion after integral reconstruction. Therefore, it is necessary to use other prior information to correct the target normal. However, combining other 3D imaging methods or depth sensors to obtain prior depth information has other limitations, such as light source restrictions or limited effective range. For example, polarization-based 3D reconstruction based on the Kinect depth sensor, developed by Microsoft's Kinect imaging system, employs unique PrimeSense ranging technology, similar to some structured light technologies, utilizing light with specific patterns whose patterns can be lines, points, surfaces, and other shapes. First, structured light is projected onto the object's surface, and then a camera receives the reflected structured light pattern. Since the received pattern will inevitably be deformed due to the object's three-dimensional shape, the spatial information of the object's surface can be calculated by analyzing the pattern's position on the camera and the degree of deformation. This method, based on a polarization camera, adds a binocular Kinect imaging system to simultaneously acquire polarization and prior depth information of the target scene. Because the multi-camera system acquires information from different angles, the camera pose is used to align the polarization and depth images. Then, coarse prior depth information is used to correct the singularity normal vectors generated during polarization 3D imaging. Finally, an integral algorithm is used to achieve 3D reconstruction of the target. The system structure is as follows: Figure 1 As shown.
[0003] In other words, currently, for large-scale polarization 3D reconstruction, the optical detection system using a single camera cannot solve the problem of normal vector singularity, which will lead to distortion of the 3D reconstruction results. Furthermore, the method of correcting normal vectors by obtaining depth prior information based on the Kinect depth sensor greatly increases the complexity of the system. Summary of the Invention
[0004] To address the aforementioned problems in related technologies, this invention provides a large-scene polarization 3D imaging method based on monocular depth estimation. The technical problem to be solved by this invention is achieved through the following technical solution:
[0005] This invention provides a method for large-scene polarization 3D imaging based on monocular depth estimation, comprising:
[0006] Acquire polarization images of the target object from different angles, as well as color images of the target object;
[0007] The color image is input into a pre-trained monocular depth estimation network to predict the depth image of the color image; the pre-trained monocular depth estimation network includes an encoder and a decoder;
[0008] The depth image is converted into depth data, and based on the depth data, the three-dimensional position information of each pixel in the depth image is obtained;
[0009] Based on the three-dimensional position information and the preset viewing direction, the surface normal vector of the first scene is obtained;
[0010] Based on the polarization image and the preset target surface refractive index, the surface normal vector of the second scene is determined;
[0011] Based on the first scene surface normal vector, the second scene surface normal vector is corrected to obtain the corrected scene surface normal vector;
[0012] Based on the corrected scene surface normal vector, the three-dimensional contour of the target object is reconstructed to obtain a three-dimensional image of the target object.
[0013] In some embodiments, the encoder includes a backbone network with fully connected layers that are fully convolutional neural networks; the decoder includes multiple transposed convolutional layers for upsampling.
[0014] In some embodiments, before inputting the color image into a pre-trained monocular depth estimation network to predict the depth image of the color image, the method includes:
[0015] The initial monocular depth estimation network is trained iteratively using training samples. In each training iteration, when the predicted depth image output by the monocular depth estimation network is obtained, the depth gradient and intensity gradient are determined based on the predicted depth image. The training samples include: multiple color images and the corresponding real depth image for each color image.
[0016] Based on the depth gradient, the intensity gradient, the predicted depth image, and the real depth image corresponding to the predicted depth image, the loss value is determined for each step. The network parameters of the monocular depth estimation network are updated based on the loss value until the pre-trained monocular depth estimation network is obtained at the end of training.
[0017] In some embodiments, the formula for the loss value each time is as follows:
[0018] L = L δ +L D ;
[0019]
[0020]
[0021] Where L represents the loss value for each iteration, δ is a preset parameter, f(x) represents the predicted depth image for each iteration, y' represents the true depth image corresponding to f(x), and N represents the total number of pixels in f(x). and This represents the depth gradient. This represents the intensity gradient. This represents the pixel value of the pixel in the i-th row and j-th column, where i and j represent the pixel in the i-th row and j-th column.
[0022] In some embodiments, the second scene surface normal vector is determined based on the incident angle and azimuth angle of the incident light on the surface of the target object; the step of correcting the second scene surface normal vector based on the first scene surface normal vector to obtain the corrected scene surface normal vector includes:
[0023] The correction value is determined based on the point ratio between the first scene surface normal vector and the second scene surface normal vector;
[0024] The corrected azimuth angle is obtained based on the correction value and the azimuth angle.
[0025] The corrected scene surface normal vector is obtained based on the incident angle and the corrected azimuth angle.
[0026] In some embodiments, the formula for the corrected azimuth angle is as follows:
[0027]
[0028]
[0029] in, Indicates the corrected azimuth angle, n pol Let n represent the surface normal vector of the second scene. depThis represents the surface normal vector of the first scene.
[0030] In some embodiments, the polarization images at different angles include: a first polarization image at 0°, a second polarization image at 45°, a third polarization image at 90°, and a fourth polarization image at 135°; determining the surface normal vector of the second scene based on the polarization images and the preset target surface refractive index includes:
[0031] Based on the first polarization image, the second polarization image, the third polarization image, and the fourth polarization image, the azimuth angle of the incident light on the surface of the target object, as well as the I parameter, Q parameter, and U parameter in the Stokes vector, are determined respectively.
[0032] The degree of polarization is determined based on the I, Q, and U parameters;
[0033] Based on the degree of polarization and the preset target surface refractive index, the incident angle of the incident light on the target object surface is determined;
[0034] The surface normal vector of the second scene is determined based on the azimuth angle and the incident angle.
[0035] In some embodiments, the formula for the azimuth angle is as follows:
[0036]
[0037] in, I' represents the azimuth angle, I'0 represents the first polarization image, and I' 45 Indicates the second polarization image, I' 90 This refers to the third polarization image.
[0038] In some embodiments, the formula for the angle of incidence is as follows:
[0039]
[0040] Where θ represents the incident angle, p t The polarization degree is represented by c, and the refractive index of the preset target surface is represented by c.
[0041] In some embodiments, obtaining the first scene surface normal vector based on the three-dimensional position information and a preset viewing direction includes:
[0042] Based on the three-dimensional position information, the neighborhood of each pixel is constructed, as well as the plane equation and constraints of each pixel;
[0043] The plane equations are solved using the Lagrange multiplier method to obtain the initial scene surface normal vectors.
[0044] The initial scene surface normal vector is adjusted according to the preset viewing direction to obtain the first scene surface normal vector.
[0045] The present invention has the following beneficial technical effects:
[0046] On the one hand, by using a trained deep neural network for monocular depth estimation, monocular depth estimation is performed on the input large scene image, thus providing prior information for large scene polarization 3D imaging in a simple way, realizing large scene polarization 3D imaging under a monocular system; on the other hand, by combining deep learning with polarization 3D imaging, the problem of polarization normal vector singularity in large scene is solved, thus realizing large scene polarization 3D imaging under a monocular system, making its application possible on hardware platforms with simple structure and low performance.
[0047] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0048] Figure 1 A schematic diagram of an exemplary Kinect-based imaging system provided for embodiments of the present invention;
[0049] Figure 2 A flowchart of a large-scene polarization 3D imaging method based on monocular depth estimation provided in an embodiment of the present invention;
[0050] Figure 3 The present invention provides an exemplary network structure diagram of a monocular depth estimation network and a flowchart illustrating the process of processing an input color image using a pre-trained monocular depth estimation network.
[0051] Figure 4 A diagram illustrating the relationship between the normal vector of the target object's surface and the incident angle and azimuth angle of the incident light on the target object's surface, provided for an embodiment of the present invention.
[0052] Figure 5 Another flowchart of an exemplary large-scene polarization 3D imaging method based on monocular depth estimation provided in an embodiment of the present invention. Detailed Implementation
[0053] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.
[0054] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0055] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0056] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, disclosure, and appended claims in carrying out the claimed invention. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0057] Figure 2 This is an optional flowchart of the control method for covertly approaching a target aircraft provided in the embodiments of the present invention, such as... Figure 2 As shown, the method includes the following steps:
[0058] S101. Acquire polarization images of the target object at different angles, as well as color images of the target object.
[0059] In some embodiments, a color image I0 of the object scene can be obtained by using an imaging detector CMOS camera to capture reflected light from the surface of the target object in a natural light environment; and polarization images I'0, I'0, I'0, and I'0 of the target object scene can be obtained by using a polarization detector at 0°, 45°, 90°, and 135°. 45 、I' 90 、I' 135 .
[0060] In other embodiments, polarization images I'0, I'0, I'135° of the target object scene can be acquired using a polarization detector. 45 、I' 90 、I' 135 Then, the color image I0 is calculated using the following formula:
[0061] Here, the target object can be any object, and this embodiment of the invention does not limit it.
[0062] S102. Input the color image into the pre-trained monocular depth estimation network to predict the depth image of the color image; the pre-trained monocular depth estimation network includes an encoder and a decoder.
[0063] Here, the encoder includes a backbone network with fully convolutional networks (FCNs) as its fully connected layers; the decoder is connected after the encoder, and the decoder includes multiple (e.g., 5) transposed convolutional layers for upsampling, fusing multi-scale features through an add operation during upsampling. For example, the backbone network can be a ResNet-50, so the encoder's network structure can be a network structure obtained by replacing the fully connected layers in ResNet-50 with FCNs.
[0064] Here, FCN is used instead of the fully connected layer in the backbone network, which can overcome the problem of the network having a fixed image input size and realize multi-scale monocular depth estimation.
[0065] For example, Figure 3 This diagram shows the network structure of a pre-trained monocular depth estimation network and a flowchart illustrating the process of using the pre-trained monocular depth estimation network to process the input color image. Figure 3 As shown, 11 is the first convolutional layer of the encoder (7*7 kernel), 12 is the first pooling layer of the encoder (3*3 kernel), modules 13-16 are the parts of ResNet-50 excluding fully connected layers, and each module in modules 13-16 is obtained by stacking different numbers of network layers 10; 17-19 are the FCN part, with 17 being a convolutional layer with a 7*7 kernel, and 18 and 19 being convolutional layers with a 1*1 kernel; 20-24 are the five transposed convolutional layers in the encoder (all with a 7*7 kernel), and the dashed lines represent the add operation. Figure 3 As shown, each network layer 10 consists of a convolutional layer 111 and a convolutional layer 112, where the convolutional kernel of convolutional layer 111 is 1*1 and the convolutional kernel of convolutional layer 112 is 3*3. Figure 3As shown, a color image a is input into a pre-trained monocular depth estimation network. After a series of processing steps, a depth image a' with a size of 112*112*64 can be obtained.
[0066] In some embodiments, prior to S102, the method further includes:
[0067] S201. The initial monocular depth estimation network is trained iteratively using training samples. In each training iteration, when the predicted depth image output by the monocular depth estimation network is obtained, the depth gradient and intensity gradient are determined based on the predicted depth image. The training samples include: multiple color images and the real depth image corresponding to each color image.
[0068] Here, the multiple color images in the training samples can be RGB color images of large scenes such as city roads contained in the KITTI dataset. The ground truth depth images in the training samples can be augmented depth images obtained by augmenting the sparse depth images in the KITTI dataset (the augmentation method is as follows: the depth values are saved as uint16 PNG images, where a pixel value of 0 represents no label value, and non-zero pixel values are converted to floating-point types and then divided by 256.0 to obtain the depth values in meters). During training, the color RGB images are used as input, and the augmented depth images are used as ground truth labels to train the network. Due to the existence of the depth image, the data augmentation part does not perform affine transformations, but only rotation and cropping operations.
[0069] Here, during each (e.g., the z-th) training of the initial monocular depth estimation network, at least one color image can be input into the network to obtain at least one corresponding predicted depth image. Then, for each predicted depth image, the derivatives of the predicted depth image along the x-axis and y-axis are calculated to obtain the depth gradient of the predicted depth image, and the intensity gradient of the predicted depth image is obtained by intensity processing of the predicted depth image.
[0070] S202. Based on the depth gradient, intensity gradient, predicted depth image, and the corresponding real depth image, determine the loss value for each step. Update the network parameters of the monocular depth estimation network based on the loss value until the pre-trained monocular depth estimation network is obtained at the end of training.
[0071] Here, after obtaining the predicted depth image for the z-th time, and the depth gradient and intensity gradient of the predicted depth image, the loss value for the z-th time can be calculated using the following formula:
[0072] L = L δ +L D ;
[0073]
[0074]
[0075] Where L represents the loss value at the z-th iteration, δ is a preset parameter, f(x) represents the predicted depth image at the z-th iteration, y' represents the true depth image corresponding to f(x), and N represents the total number of pixels in f(x). and Represents the depth gradient. Indicates the intensity gradient. This represents the pixel value of the pixel in the i-th row and j-th column, where i and j represent the pixel in the i-th row and j-th column.
[0076] Here, since depth discontinuities often occur on the image gradient, the image intensity gradient is also taken into account to keep the depth locally smooth, thus achieving better training results.
[0077] Here, after obtaining the loss value at the z-th time, the network parameters of the monocular depth estimation network used for the z-th training can be updated based on the loss value at the z-th time, thereby obtaining the monocular depth estimation network used for the z+1-th training. During the z+1-th training, the above training principle is continued to be used for training until the preset number of training times is reached or the network converges, at which point the training ends, and the monocular depth estimation network obtained from the last training is used as the pre-trained monocular depth estimation network.
[0078] S103. Convert the depth image into depth data, and based on the depth data, obtain the three-dimensional position information of each pixel in the depth image.
[0079] Here, the depth map can be divided by 256 to convert it into depth data in meters. Then, using the conversion formula from pixel coordinates to world coordinates, the depth data can be represented as 3D point position information in the world coordinate system. Thus, the 3D position information of each pixel in the depth image is obtained. Specifically, the conversion formula from pixel coordinates to world coordinates is as follows:
[0080] Thus, the three-dimensional position information of each pixel is obtained as follows: Represents depth data, x w y w z w This represents the 3D point position information of each pixel in the world coordinate system. u and v are the pixel coordinates of each pixel in the pixel coordinate system, u0 and v0 are the origin coordinates of the pixel coordinate system, f is the focal length, and dx and dy represent the length and width of each pixel, respectively.
[0081] S104. Based on the three-dimensional position information and the preset viewing direction, the surface normal vector of the first scene is obtained.
[0082] Here, based on the three-dimensional position information, the neighborhood of each pixel, as well as the plane equation and constraints of each pixel, can be constructed. Then, the plane equation is solved by the Lagrange multiplier method to obtain the initial scene surface normal vector. The initial scene surface normal vector is adjusted according to the preset viewing direction to obtain the first scene surface normal vector.
[0083] Specifically, the normal vector is calculated using the point cloud map, and the neighborhood points are fitted to a plane. The normal vector of this plane is then used as the normal vector of the point. The process is as follows:
[0084] (1) Construct the neighborhood of each pixel. The neighborhood formula is as follows:
[0085] Where m is a preset value, representing the total number of pixels;
[0086] (2) Construct the plane equations and constraints as follows:
[0087] Among them, A, B, C, and D are unknown parameters;
[0088] (3) Differentiating the formula in step (2) yields:
[0089]
[0090] Where M is the covariance matrix. And so on; let have to:
[0091] (4) Using the Lagrange multiplier method, a new optimization function is constructed as follows:
[0092] f(n',λ)=(Mn') T Mn'+λ(1-n'Tn');
[0093] (5) Find the normalized eigenvector corresponding to the smallest eigenvalue of the covariance matrix M. That is, the plane normal vector (the normal vector of the point cloud);
[0094] (6) Unify the viewing direction according to the preset line of sight to ensure the consistency of the plane normal vector direction. Taking the origin as the viewpoint position, the preset line of sight direction is: Normal vector of point cloud Adjustments were made so that: This yields the adjusted plane normal vector (i.e., the first scene surface normal vector mentioned above).
[0095] S105. Based on the polarization image and the preset target surface refractive index, determine the surface normal vector of the second scene.
[0096] Here, it can be based on the first polarization image I'0 and the second polarization image I' 45 Third polarization image I' 90 and the fourth polarization image I' 135 The azimuth angle of the incident light on the target object surface and the I, Q, and U parameters in the Stokes vector are determined respectively; the degree of polarization is determined based on the I, Q, and U parameters; the incident angle of the incident light on the target object surface is determined based on the degree of polarization and the preset target surface refractive index; and the surface normal vector of the second scene is determined based on the azimuth angle and the incident angle.
[0097] Specifically, the Stokes vector representation is a commonly used method for representing polarization characteristics. It refers to the fact that the polarization state of a beam of light can be completely described by four fixed parameters, called the Stokes vector. Since each Stokes parameter is represented by light intensity, it can be directly measured using certain photoelectric instruments. The Stokes vector can be expressed as:
[0098] Among them, E x and E y These represent the components of the electric field vector of the light reflected from the surface of the target object along the x and y axes, respectively. Using the Stokes vector representation, the degree of polarization P... t The calculation formula is:
[0099] Specifically, the formula for calculating the azimuth angle is as follows:
[0100]
[0101] Specifically, the formula for calculating the angle of incidence is as follows:
[0102]
[0103] Where θ represents the angle of incidence, p t The value represents the degree of polarization, and c represents the preset target surface refractive index (for example, it can be 1.5).
[0104] Here, as Figure 4 As shown, when the azimuth and incident angle of the incident light on the target object's surface are obtained, the polar coordinates of the target object's surface normal vector are obtained. And θ, thus the surface normal vector of the target object can be obtained. (The surface normal vector of the second scene mentioned above).
[0105] S106. Based on the surface normal vector of the first scene, correct the surface normal vector of the second scene to obtain the corrected surface normal vector.
[0106] Here, because the light intensity of the polarized images obtained from two rotation angles spaced 180° apart is the same during the polarizer rotation process, there is a 180° uncertainty between the incident azimuth angle of the incident light on the target object surface to be reconstructed and the actual incident azimuth angle in the calculation results. This leads to uncertainty in the direction of the object surface normal vector obtained from the polarization information, so it is necessary to correct the object surface normal vector obtained from the polarization information.
[0107] Here, according to and The points between them are used to determine the correction value. Based on the correction value and the azimuth angle, the corrected azimuth angle is obtained. Then, based on the incident angle θ and the corrected azimuth angle... The corrected scene surface normal vector is obtained.
[0108] Specifically, the corrected azimuth angle The formula is as follows:
[0109]
[0110]
[0111] Where, n pol for n dep for
[0112] The principle behind correcting the normal vector using the above formula is as follows: when the angle between the normal vector estimated by the network and the normal vector obtained by polarization is less than 180°, the normal vector obtained by polarization is considered accurate. Conversely, when the angle between the normal vector estimated by the network and the normal vector obtained by polarization is greater than 180°, a phase difference of π needs to be added to the azimuth angle of the normal vector obtained by polarization for correction. Using this method, the normal vector of the target surface micro-element can be uniquely solved.
[0113] S107. Based on the corrected scene surface normal vector, reconstruct the three-dimensional contour of the target object to obtain a three-dimensional image of the target object.
[0114] Here, the three-dimensional contour of the target object can be reconstructed using the following formula: Among them, Z F (u,v), P(u,v), and Q(u,v) are the Fourier transforms of Z(x,y), p(x,y), and q(x,y), respectively, where p(x,y) and q(x,y) are the obtained normal gradient data.
[0115] For example, Figure 5 Another flowchart illustrating an exemplary large-scene polarization 3D imaging method based on monocular depth estimation provided in embodiments of the present invention. Figure 5 As shown, firstly, a color image (I0) and four polariton images (I'0, I ... 45 、I' 90 、I' 135 Then, on the one hand, the color image I0 can be input into a pre-trained monocular depth estimation network to obtain a coarse depth map (the predicted depth image). The depth map is then converted into a point cloud (i.e., the three-dimensional position information of each pixel). The normal vector of the target object's surface is then calculated using the point cloud. On the other hand, based on the four polariton images, the degree of polarization of the scene and the azimuth angle of the incident light on the target object's surface are calculated using Stokes vectors. The angle of incidence of the incident light on the target object's surface is calculated based on the degree of polarization, and the normal vector of the target object's surface is calculated based on the angle of incidence and the azimuth angle. After that, adopt right The scene surface normal vector is corrected, and the 3D shape of the target object is recovered by scene surface integration based on the corrected scene surface normal vector, thus obtaining the 3D information (3D image) of the target object.
[0116] This invention, on the one hand, uses a trained deep neural network for monocular depth estimation to perform monocular depth estimation on input large-scene images, thereby providing prior information for large-scene polarization 3D imaging in a simple way, realizing large-scene polarization 3D imaging under a monocular system; on the other hand, by combining deep learning with polarization 3D imaging, it solves the problem of polarization normal vector singularity in large scenes, thus realizing large-scene polarization 3D imaging under a monocular system, making its application possible on hardware platforms with simple structure and low performance.
[0117] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A method for large-scene polarization 3D imaging based on monocular depth estimation, characterized in that, include: Acquire polarization images of the target object from different angles, as well as color images of the target object; The color image is input into a pre-trained monocular depth estimation network to predict the depth image of the color image; the pre-trained monocular depth estimation network includes an encoder and a decoder; The depth image is converted into depth data, and based on the depth data, the three-dimensional position information of each pixel in the depth image is obtained; Based on the three-dimensional position information and the preset viewing direction, the surface normal vector of the first scene is obtained; Based on the polarization image and the preset target surface refractive index, the surface normal vector of the second scene is determined; Based on the first scene surface normal vector, the second scene surface normal vector is corrected to obtain the corrected scene surface normal vector; Based on the corrected scene surface normal vector, the three-dimensional contour of the target object is reconstructed to obtain a three-dimensional image of the target object; The second scene surface normal vector is determined based on the incident angle and azimuth angle of the incident light on the surface of the target object; the step of correcting the second scene surface normal vector based on the first scene surface normal vector to obtain the corrected scene surface normal vector includes: The correction value is determined based on the dot product between the surface normal vector of the first scene and the surface normal vector of the second scene; The corrected azimuth angle is obtained based on the correction value and the azimuth angle. The corrected scene surface normal vector is obtained based on the incident angle and the corrected azimuth angle. The formula for the corrected azimuth angle is as follows: ; ; in, Indicates the corrected azimuth angle. This represents the surface normal vector of the second scene. This represents the surface normal vector of the first scene.
2. The large-scene polarization 3D imaging method based on monocular depth estimation according to claim 1, characterized in that, The encoder includes a backbone network with fully connected layers that are fully convolutional neural networks; the decoder includes multiple transposed convolutional layers for upsampling.
3. The large-scene polarization 3D imaging method based on monocular depth estimation according to claim 1, characterized in that, Before inputting the color image into a pre-trained monocular depth estimation network to predict the depth image of the color image, the method includes: The initial monocular depth estimation network is trained iteratively using training samples. In each training iteration, when the predicted depth image output by the monocular depth estimation network is obtained, the depth gradient and intensity gradient are determined based on the predicted depth image. The training samples include: multiple color images and the corresponding real depth image for each color image. Based on the depth gradient, the intensity gradient, the predicted depth image, and the real depth image corresponding to the predicted depth image, the loss value is determined for each step. The network parameters of the monocular depth estimation network are updated based on the loss value until the pre-trained monocular depth estimation network is obtained at the end of training.
4. The large-scene polarization 3D imaging method based on monocular depth estimation according to claim 3, characterized in that, The formula for the loss value each time is as follows: ; ; ; in, This represents the loss value for each instance. These are preset parameters. This represents the predicted depth image for each iteration. express The corresponding true depth image, express The total number of pixels in the image. and This represents the depth gradient. This represents the intensity gradient. Indicates the first Line 1 The pixel values of the columns of pixels. and Indicates the first Line 1 The number of pixels in a column.
5. The large-scene polarization 3D imaging method based on monocular depth estimation according to claim 1, characterized in that, The polarization images at different angles include: a first polarization image at 0°, a second polarization image at 45°, a third polarization image at 90°, and a fourth polarization image at 135°; the determination of the second scene surface normal vector based on the polarization images and the preset target surface refractive index includes: Based on the first polarization image, the second polarization image, the third polarization image, and the fourth polarization image, the azimuth angle of the incident light on the surface of the target object, as well as the I parameter, Q parameter, and U parameter in the Stokes vector, are determined respectively. The degree of polarization is determined based on the I, Q, and U parameters; Based on the degree of polarization and the preset target surface refractive index, the incident angle of the incident light on the target object surface is determined; The surface normal vector of the second scene is determined based on the azimuth angle and the incident angle.
6. The large-scene polarization 3D imaging method based on monocular depth estimation according to claim 5, characterized in that, The formula for the azimuth angle is as follows: ; in, Indicates azimuth. Indicates the first polarization image, Represents the second polarization image, This refers to the third polarization image.
7. The large-scene polarization 3D imaging method based on monocular depth estimation according to claim 5, characterized in that, The formula for the incident angle is as follows: ; in, Indicates the incident angle, Indicates the degree of polarization, This represents the refractive index of the preset target surface.
8. The large-scene polarization 3D imaging method based on monocular depth estimation according to claim 1, characterized in that, The process of obtaining the first scene surface normal vector based on the three-dimensional position information and the preset viewing direction includes: Based on the three-dimensional position information, the neighborhood of each pixel is constructed, as well as the plane equation and constraints of each pixel; The plane equations are solved using the Lagrange multiplier method to obtain the initial scene surface normal vectors. The initial scene surface normal vector is adjusted according to the preset viewing direction to obtain the first scene surface normal vector.