A multi-target absolute depth estimation method based on monocular polarization 3D imaging
Through monocular polarization three-dimensional imaging combined with camera calibration and monocular ranging model, the problem that polarization three-dimensional imaging technology cannot obtain absolute depth is solved, and high-precision absolute depth reconstruction in multi-objective scenarios is achieved, which broadens the application range.
Patent Information
- Application Number
- CN202310251809.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-15
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-03-15
AI Technical Summary
The existing polarization three-dimensional imaging technology cannot obtain the absolute depth information of the target, and the three-dimensional reconstruction effect of the target at different depths in multi-objective scenarios is unclear. The existing methods are complex and costly, which limits their application scope.
Through monocular polarization three-dimensional imaging combined with camera calibration and monocular ranging model, an adaptive internal reference prediction method is constructed to restore the absolute depth information of the target in multi-objective scenes, solve the problem of camera internal reference changes, and achieve absolute depth recovery under dynamic focus.
High-precision absolute depth reconstruction of multi-objective scenarios under monocular conditions is achieved, the application scenarios of polarization three-dimensional imaging technology are broadened, and the low cost and high robustness are maintained.
Smart Images

Figure CN116485869B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of optical imaging, and in particular relates to a multi-target absolute depth estimation method based on monocular polarization three-dimensional imaging. Background Art
[0002] In recent years, polarization 3D imaging, a new technology for 3D reconstruction based on the polarization information of reflected light from a target, has become a hot topic in research. Its unique imaging mechanism offers advantages over other 3D imaging methods, including simple equipment, long-range detection, high precision, and robustness. This technology is expected to significantly improve the overall performance of 3D reconstruction in numerous applications, such as facial recognition in security, product defect detection in industry, and autonomous driving in machine vision. However, existing polarization 3D imaging techniques, lacking the ability to measure true distance, can only obtain relative depth information from the target surface, severely limiting their application. Combining methods for obtaining prior true information with polarization 3D imaging can achieve absolute depth information recovery, but such methods require more equipment and complex data processing, increasing the time and cost of the overall 3D imaging system. Furthermore, existing methods do not consider the 3D reconstruction of targets at different depths in multi-target scenes, and the effectiveness of polarization 3D reconstruction for such scenes remains unclear.
[0003] Existing 3D reconstruction methods include those based on polarization binocular vision. This method recovers the target's relative depth information in the pixel coordinate system by acquiring the polarization information of the target's reflected light. It also integrates the rough absolute depth information obtained by binocular stereo vision technology with the camera parameters calibrated using binocular stereo. This method converts the point cloud data in the polarization-acquired image pixel coordinate system into absolute data in the world coordinate system, achieving high-precision polarization 3D imaging of the target. The figure below shows the 3D reconstruction method based on polarization binocular vision.
[0004] First, the left camera and polarizer are combined to capture polarization images at different angles, and the dense point cloud data of the target in the pixel coordinate system is obtained using a polarization-based three-dimensional imaging method. Then, the SURF feature point detection algorithm is used to detect the target surface feature points and match the feature point pairs of the left and right camera images. Then, the three-dimensional coordinates in the world coordinate system, i.e., the sparse point cloud, are solved based on binocular stereo vision technology. The camera parameters obtained by binocular positioning are combined to calculate the translation transformation parameters between the image pixel coordinate system and the world coordinate system, and the scale change parameters are solved using the least squares method. Finally, the relative depth information of the target obtained by polarization is converted into absolute depth information.
[0005] The 3D reconstruction method based on polarization binocular vision uses the absolute depth information obtained by binocular stereo vision technology to realize the conversion of the relative depth information of the target obtained by the polarization 3D imaging method into absolute depth information. However, the direct introduction of binocular stereo vision technology will bring its inherent defects. Therefore, this method will limit the original advantages of polarization 3D imaging technology such as low cost, high precision and strong robustness. At the same time, this method only conducts experimental verification and analysis on a single target, and does not explore the polarization 3D absolute depth reconstruction effect of targets at different depths in multi-target scenes, which narrows the application scope of polarization 3D imaging technology. Summary of the Invention
[0006] In order to solve the above problems existing in the prior art, the present invention provides a multi-target absolute depth estimation method based on monocular polarization 3D imaging. The technical problem to be solved by the present invention is achieved through the following technical solutions:
[0007] The present invention provides a multi-target absolute depth estimation method based on monocular polarization three-dimensional imaging, comprising:
[0008] S1: Perform camera pre-calibration to obtain the linear variation coefficient of the principal point during camera focusing.
[0009] S2: Based on the fuzziness evaluation value, the nearest area of the target object is focused until it is clear and the camera intrinsic parameters of the current position are obtained;
[0010] S3: Obtain polarization images of the clear target area at different angles, and calculate the polarization degree of the reflected light from the clear target surface, as well as the azimuth angle and zenith angle of the target surface normal;
[0011] S4: constructing the target surface normal and integrating to reconstruct the three-dimensional surface contour of the clear target area;
[0012] S5: constructing a monocular ranging model, and obtaining a true distance between two points in the target surface space according to the monocular ranging model;
[0013] S6: Recover the absolute depth information of the target based on the true distance between two points on the target surface space;
[0014] S7: focusing other blurred target areas to make them clear based on the blur evaluation value and obtaining adaptive prediction values of the camera intrinsic parameters after focusing;
[0015] S8: Using the adaptive prediction value of the camera intrinsic parameter after focusing, repeat steps S3-S6 to restore the absolute depth three-dimensional information of the blurred target after dynamic focusing.
[0016] In one embodiment of the present invention, the S1 includes:
[0017] Focus the camera to different positions multiple times, and obtain d1, d2, ..., d by camera calibration.k The principal point coordinates at the position are fitted based on the least squares method to obtain the linear variation coefficient η of the principal point coordinates in the x and y directions. x and η y , where the principal point is the intersection of the camera optical axis and the imaging plane.
[0018] In one embodiment of the present invention, the S2 includes:
[0019] The grayscale variance operator is used to represent the fuzziness evaluation value of the target image. The size of the grayscale variance operator of the target image is changed by moving the lens. When the grayscale variance operator is the largest, it means that the target image is in focus.
[0020] After the focus is clear, the camera internal parameters at this moment are obtained through the camera calibration method, which is recorded as:
[0021]
[0022] Where dx and dy represent the pixel sizes on the abscissa and ordinate of the detector, respectively; f′ represents the focal length of the camera when the image is in focus; and (u0′, v0′) represents the coordinates of the principal point when the image is in focus.
[0023] In one embodiment of the present invention, the S3 includes:
[0024] S3.1: Under natural light, obtain polarization images I0, I1 at four target angles of 0°, 45°, 90°, and 135°, respectively. 45 , I 90 , I 135 ;
[0025] S3.2: Obtaining Stokes vectors using polarization images at different angles, and obtaining polarization degrees using the Stokes vectors;
[0026] S3.3: Calculate the azimuth of the target surface normal using polarization images at different angles and the zenith angle θ.
[0027] In one embodiment of the present invention, the S4 includes:
[0028] S4.1: Characterize the surface normal of the target
[0029]
[0030] Where p and q represent the gradients of the target microfacet normal vector in the x and y directions respectively;
[0031] S4.2: Based on the Cartesian surface assumption, obtain the target polarization 3D imaging model:
[0032] cos(z)=∫∫((z x -p) 2 +(z y -q) 2 )dxdy
[0033] Among them, z = f (x, y) represents the three-dimensional surface function of the target, z x Represents the gradient of the target three-dimensional surface in the x-axis direction, z y Indicates the gradient of the target 3D surface in the y-axis direction.
[0034] In one embodiment of the present invention, the S5 includes:
[0035] Place a single camera horizontally, and use the triangular similarity relationship between the imaging coordinates of two points on the reference ground on the camera imaging plane, the distance between the two points and the imaging plane, the height of the camera from the reference ground, and the focal length of the camera to calculate the spatial distance between the two points.
[0036] In one embodiment of the present invention, the S5 includes:
[0037] S5.1: Obtain the depth distance between two points on the reference ground using the similarity relationship in the monocular ranging model:
[0038]
[0039]
[0040]
[0041] Among them, D1 and D2 are the vertical absolute depths of the two points p1 and p2 of the target on the reference ground to the camera, respectively. v =f′ / dy represents the effective focal length of the camera in the y direction, v1 and v2 represent the longitudinal coordinates of the two points in the imaging plane, v0′ represents the longitudinal coordinate of the current principal point, Δz represents the distance between the two points in the z direction in the camera coordinate system, H c The height of the camera from the reference ground;
[0042] S5.2: According to the pinhole imaging model, the relationship between the target's world coordinate system and the pixel coordinate system is:
[0043]
[0044] Among them, R represents a 3*3 rotation matrix, T represents a 3*1 translation matrix, M0 represents the camera's intrinsic parameter matrix, M1 represents the camera's extrinsic parameter matrix, f' represents the current camera focal length, X w ,Y w ,Z wIndicates the x, y, and z axis coordinates of the target in the world coordinate system. c Indicates the z-axis coordinate value in the target camera coordinate system;
[0045] S5.3: Obtain the actual distance between points p1 and p2 in the x-direction in the camera coordinate system:
[0046]
[0047] Among them, F u =f′ / dx represents the effective focal length of the camera in the x direction.
[0048] In one embodiment of the present invention, the S6 includes:
[0049] S6.1: Based on the real distance between the two target points in the x-direction in the camera coordinate system measured in step S5, obtain the real scale factor of the target polarization 3D imaging model:
[0050] x s =Δx / |x p2 -x p1 |
[0051] y s =x s ·d y / d x
[0052] z s =Δz / |z p2 -z p1 |
[0053] Among them, x s 、y s 、z s Respectively represent the scale factors of the target polarization 3D imaging model in the x, y, and z directions relative to the true value, z p1 、z p2 Respectively represent the coordinate values of points p1 and p2 in the z direction in the target polarization three-dimensional imaging model, x p1 、x p2 Respectively represent the coordinate values of points p1 and p2 in the x-direction in the target polarization 3D imaging model;
[0054] The absolute depth information of the target is restored according to the true scale factor to obtain a 3D absolute depth model of the target:
[0055] x t =x p ·x s ,y t =y p ·y s ,z t =zp ·z s
[0056] Among them, x t 、y t 、z t Respectively represent the coordinate values in the x, y, and z directions of the reconstructed target 3D absolute depth model, x p 、y p 、z p They respectively represent the coordinate values in the x, y, and z directions in the target polarization three-dimensional imaging model.
[0057] In one embodiment of the present invention, the S7 includes:
[0058] S7.1: Focus the object in the blurred area until it is sharp based on the blur evaluation value, and measure the lens displacement Δf during the focusing process using a micrometer screw;
[0059] S7.2: Based on the lens displacement Δf during the focusing process, the offset of the principal point coordinates in the x and y directions is obtained as follows:
[0060] H x =Δf·η x
[0061] H y =Δf·η y
[0062] Where Δf is the focus displacement, η x ,η y is the linear variation coefficient;
[0063] S7.3: Obtain the coordinates of the principal point after focusing:
[0064]
[0065] S7.4: Obtain the effective focal length of the camera in the x and y directions after focusing
[0066] S7.5: Obtain the adaptive prediction value M0′ of the camera intrinsic parameter after focusing:
[0067]
[0068] Compared with the prior art, the present invention has the following beneficial effects:
[0069] 1. This multi-target absolute depth estimation method based on monocular polarization 3D imaging builds upon traditional polarization 3D imaging methods. By combining the construction of a monocular ranging model with camera calibration, it can recover target absolute depth information, effectively overcoming the application limitations of traditional polarization 3D imaging technology. Furthermore, the monocular ranging model constructed in this invention enables 3D reconstruction of the target's absolute depth using a single camera, providing a theoretical foundation and technical solution for a convenient and low-cost polarization 3D imaging method.
[0070] 2. The post-focus camera intrinsic parameter adaptive prediction model proposed in the present invention can effectively solve the problem of camera intrinsic parameter changes caused by the need to focus in multi-target scenes at different depths; by constructing a camera intrinsic parameter adaptive prediction model, the present invention combines a monocular ranging model with a camera calibration method to achieve monocular absolute depth three-dimensional information recovery in multi-target scenes, effectively broadening the application scenarios of polarization three-dimensional imaging technology.
[0071] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 This is a flow chart of a multi-target absolute depth estimation method based on monocular polarization three-dimensional imaging provided by an embodiment of the present invention;
[0073] Figure 2 This is a schematic diagram of the positional relationship between a lens displacement direction and an imaging plane provided by an embodiment of the present invention;
[0074] Figure 3 is a schematic diagram of a normal representation model of a target surface provided by an embodiment of the present invention;
[0075] Figure 4 is a schematic diagram of a monocular ranging model provided by an embodiment of the present invention;
[0076] Figure 5 This is a schematic diagram of a camera focusing process provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0077] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the following is a detailed description of a multi-target absolute depth estimation method based on monocular polarization three-dimensional imaging proposed in accordance with the present invention, in combination with the accompanying drawings and specific implementation methods.
[0078] The aforementioned and other technical contents, features, and effects of the present invention are clearly presented in the following detailed description of the specific embodiments in conjunction with the accompanying drawings. Through the description of the specific embodiments, a deeper and more specific understanding of the technical means and effects adopted by the present invention to achieve the intended purpose can be obtained. However, the accompanying drawings are provided for reference and illustration purposes only and are not intended to limit the technical solutions of the present invention.
[0079] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations are intended to cover non-exclusive inclusion, such that an article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the article or device comprising the element.
[0080] Traditional polarization 3D imaging technology performs 3D information inversion based on the polarization information of the target's reflected light. It can only obtain the relative depth information of the target surface, and it is difficult to reconstruct the target's true 3D contour, which seriously restricts its further development and practical application in the field of vision. As for other methods that can restore the absolute depth information of the target, such as the 3D reconstruction method based on polarization binocular vision, there are problems such as complex operation and high cost, which limit the original advantages of polarization 3D imaging technology such as low cost, high precision and strong robustness. At the same time, these methods do not consider the reconstruction problem of different depths in multi-target scenes, and the effect of polarization 3D reconstruction of targets in such scenes is still unclear. Therefore, the present invention provides a monocular polarization 3D imaging method that can restore the absolute depth of multiple targets. Based on the surface shape obtained by polarization 3D imaging, by combining camera calibration and monocular ranging model, an adaptive internal parameter prediction method is proposed to solve the problem of camera internal parameter changes caused by the need to focus on blurred targets outside the depth of field in multi-target scenes, and ultimately achieve high-precision recovery of absolute depth information under dynamic focus of targets at different depths.
[0081] See Figure 1 , Figure 1 This is a flowchart of a multi-target absolute depth estimation method based on monocular polarization three-dimensional imaging provided by an embodiment of the present invention, the method comprising:
[0082] S1: Perform camera pre-calibration to obtain the linear variation coefficient of the principal point during camera focusing.
[0083] The principal point is the intersection of the camera optical axis and the imaging plane, and its coordinates are expressed as (u0, v0). During the focusing process, the principal point coordinates change linearly due to the slight non-perpendicularity between the camera barrel direction, i.e., the lens displacement direction, and the imaging plane. For example, Figure 2 shown.
[0084] Focus the camera to different positions multiple times, and obtain d1, d2, ..., d by camera calibration. k The coordinates of the principal point at the position are recorded as (u 01 ,v 01 )、(u 02 ,v 02 )、……、(u 0k ,v 0k ), the linear variation coefficient η of the principal point coordinates in the x and y directions is fitted based on the least squares method x and η y .
[0085] In this embodiment, camera pre-calibration mainly includes the following steps:
[0086] S1a: Prepare calibration images
[0087] Calibration images are taken using a calibration plate at different positions, angles, and postures. A minimum of three images are taken, and 10-20 images are ideal. The calibration plate uses a checkerboard pattern consisting of alternating black and white rectangles.
[0088] S1b: Extract corner information for each calibration image
[0089] The corner points are the inner corner points on the calibration board, and the find Chessboard Corners function is used to extract the corner point information.
[0090] S1c: Extract sub-pixel information for each calibration image
[0091] By using the cornerSubPix function, sub-pixel information is further extracted based on the initially extracted corner information to reduce the camera calibration error.
[0092] S1d: Draw the successfully calibrated corner points on the captured calibration image
[0093] Use the drawChessboardCorners function to draw the successfully calibrated corner points.
[0094] S1e: Camera Calibration
[0095] Use the calibrateCamera function to calibrate and calculate the camera's intrinsic and extrinsic parameter coefficients.
[0096] Through the above steps, the camera intrinsic parameter matrix is obtained, which is recorded as:
[0097]
[0098] Where dx and dy represent the pixel size of the detector (i.e., the camera's CMOS sensor) on the horizontal and vertical axes, respectively, in mm, f represents the camera focal length, and (u0, v0) represents the principal point coordinates.
[0099] Focus the camera to different positions multiple times, and obtain the camera internal parameter matrix at different positions to obtain the principal point coordinates at different positions. Then, the linear variation coefficient η of the principal point coordinates in the x and y directions is fitted based on the least squares method. x and η y .
[0100] S2: Based on the fuzziness evaluation value, the nearest area of the target object is focused until it is clear and the camera intrinsic parameters of the current position are obtained.
[0101] Specifically, the grayscale variance operator is used to represent the blur evaluation value of the target image. By moving the lens to change the blur of the target image, that is, changing the size of the grayscale variance operator, when the grayscale variance operator is maximized, it indicates that the target is in focus and clear, and the target image is the image of the target object. The calculation formula of the grayscale variance operator is as follows:
[0102]
[0103] Among them, N x 、N y They represent the number of row pixels and column pixels in the target area respectively, and I(x,y) represents the grayscale value at the pixel (x,y) position.
[0104] After the focus is clear, the camera internal parameters at this moment are obtained through the camera calibration method in step S1, which is recorded as:
[0105]
[0106] Where dx and dy represent the pixel sizes on the abscissa and ordinate of the detector, respectively; f′ represents the focal length of the camera when the image is in focus; and (u0′, v0′) represents the coordinates of the principal point when the image is in focus.
[0107] S3: Obtain polarization images of the clear target area at different angles and calculate the polarization degree P of the reflected light from the clear target surface and the azimuth angle of the target surface normal and the zenith angle θ.
[0108] In this embodiment, step S3 includes:
[0109] S3.1: Under natural light, place a polarizer in front of the CMOS camera and rotate the polarizer at 0°, 45°, 90°, and 135° to obtain polarized images I0, I1, and I2 of the target scene at four angles. 45 , I 90 , I 135 .
[0110] S3.2: Use the Stokes vector to characterize the polarization state of the reflected light. The Stokes parameters can be obtained from the four polarization images I0, I 45 , I 90 and I 135 The Stokes vector is expressed as:
[0111]
[0112] Among them, E x and E y They represent the components of the electric field vector of the light reflected from the surface of the object on the x and y axes, respectively, x and δ y Indicates the phase in that direction.
[0113] The specific calculation formula of the polarization degree P based on Stokes vector representation is:
[0114]
[0115] S3.3: Calculate the azimuth of the target surface normal using polarization images at different angles and zenith angle θ, the normal characterization model is as follows Figure 3 shown.
[0116] Specifically, firstly, polarization images I0, I 45 , I 90 , I 135 Calculate the azimuth The calculation formula is as follows:
[0117]
[0118] Then, the zenith angle θ can be calculated based on the relationship between the polarization degree P and the zenith angle θ. The calculation formula is:
[0119]
[0120] Here, n represents the refractive index of the surface of the object. In this embodiment, the surface refractive index is 1.5.
[0121] S4: constructing the target surface normal and integrating to reconstruct the three-dimensional surface contour of the clear target area.
[0122] Specifically, step S4 of this embodiment includes:
[0123] S4.1: First, characterize the surface normal of the target according to the following formula
[0124]
[0125] Where p and q represent the gradients of the target microfacet normal vector in the x and y directions, respectively.
[0126] S4.2: Based on the Cartesian surface assumption (continuous and closed surface), the target surface integral 3D reconstruction formula is obtained, that is, the polarization 3D imaging model of the target:
[0127] cos(z)=∫∫((z x -p) 2 +(z y -q) 2 )dxdy (9)
[0128] Among them, z = f (x, y) represents the three-dimensional surface function of the target, z x Represents the gradient of the target three-dimensional surface in the x-axis direction, z y Indicates the gradient of the target 3D surface in the y-axis direction.
[0129] S5: Construct a monocular distance measurement model, and obtain the true distance between two points in the target surface space according to the monocular distance measurement model.
[0130] In this embodiment, a monocular ranging model is proposed. Figure 4 , Figure 4 This is a schematic diagram of a monocular distance measurement model provided by an embodiment of the present invention. A single camera is placed horizontally, and the spatial distance between two points is calculated using the triangular similarity relationship between the imaging coordinates of two points on the reference ground on the camera imaging plane, the distance between the two points and the imaging plane, the height of the camera from the reference ground, and the focal length of the camera. c is the height of the camera from the reference ground, Oc-XcYcZc represents the camera coordinate system, O c is the camera optical center, F v =f′ / dy represents the effective focal length of the camera in the y direction, v h Expressed O c The horizontal line is located on the imaging plane, p1 and p2 are two points on the reference ground, that is, two points randomly selected on the reference ground, v1 and v2 are the longitudinal coordinates of the two points in the imaging plane, D1 and D2 are the longitudinal absolute depths from the target point to the camera.
[0131] According to the similarity relationship in the monocular ranging model, the depth distance between two points on the reference ground can be expressed as follows:
[0132]
[0133] According to the pinhole imaging model, the relationship between the target's world coordinate system and pixel coordinate system is:
[0134]
[0135] Among them, R represents a 3*3 rotation matrix, T represents a 3*1 translation matrix, M0 represents the camera's intrinsic parameter matrix, M1 represents the camera's extrinsic parameter matrix, f' represents the current camera focal length, X w ,Y w ,Z w Indicates the x, y, and z axis coordinates of the target in the world coordinate system. c Indicates the z-axis coordinate value in the target camera coordinate system.
[0136] Combining equations (10) and (11), we can derive the expression of the actual distance between points p1 and p2 in the x direction in the camera coordinate system:
[0137]
[0138] Among them, F u =f′ / dx represents the effective focal length of the camera in the x direction.
[0139] According to the monocular ranging model constructed by equations (10) and (12), the real spatial distance between two points of the target at different depths of the reference ground is obtained through the camera principal point coordinates and focal length obtained in step S2, as well as the measured camera height and target point imaging coordinates.
[0140] S6: Recover the absolute depth information of the target based on the true distance between two points on the target surface space;
[0141] In this embodiment, S6 includes:
[0142] S6.1: Obtain a true scale factor of the polarization 3D imaging model based on the true distance between two points of the target at different depths of the reference ground measured in step S5:
[0143]
[0144] Among them, x s 、y s 、z s Respectively represent the scale factors of the target polarization 3D imaging model in the x, y, and z directions relative to the true value, z p1 、z p2Respectively represent the coordinate values of points p1 and p2 in the z direction in the target polarization three-dimensional imaging model, x p1 、x p2 They respectively represent the coordinate values of points p1 and p2 in the x direction in the target polarization 3D imaging model.
[0145] The absolute depth information of the target is restored according to the true scale factor to obtain a 3D absolute depth model of the target:
[0146] x t =x p ·x s , y t =y p ·y s , z t =z p ·z s (14)
[0147] Among them, x t 、y t 、z t Respectively represent the coordinate values in the x, y, and z directions in the 3D absolute depth model of the target reconstruction, x p 、y p 、z p They respectively represent the coordinate values in the x, y, and z directions in the target polarization 3D imaging model restored in step 6.
[0148] S7: Based on the blur evaluation value, the other blurred target areas are focused until they are clear and the camera intrinsic parameters after focusing are adaptively obtained.
[0149] In this step, the other blurred areas are focused until they are clear based on the blur evaluation value according to the method of step S2, and the lens displacement during the focusing process is measured by a micrometer screw and recorded as Δf.
[0150] Combined with the analysis of the principal point coordinate change during the focusing process in step S1 and Figure 5 The schematic diagram of camera focal length change during focusing is shown to construct a post-focus adaptive intrinsic parameter prediction model.
[0151] Specifically, step S7 of this embodiment includes:
[0152] S7.1: Focus the object in the blurred area until it is sharp based on the blur evaluation value, and measure the lens displacement Δf during the focusing process using a micrometer screw;
[0153] S7.2: The change in the principal point coordinates is a linear process. Based on the lens displacement Δf during the focusing process, the offset of the principal point coordinate values in the x and y directions is obtained as:
[0154]
[0155] Where Δf is the focus displacement, η x ,η y is the linear variation coefficient obtained in step 1.
[0156] S7.3: Obtain the coordinates of the principal point after focusing:
[0157]
[0158] S7.4: Since the focal length error caused by a slight tilt of the lens barrel is much smaller than the camera focal length, the measured camera focus displacement Δf is directly substituted into equation (17) to obtain the effective focal length of the camera after focusing. The effective focal length of the camera in the x-direction (and the same for the y-direction) after focusing is:
[0159]
[0160] S7.5: Construct an adaptive prediction model for the camera intrinsic parameters after focusing to obtain the adaptive prediction value M0′ of the camera intrinsic parameters after focusing:
[0161]
[0162] S8: Using the adaptive prediction value of the camera intrinsic parameter after focusing, repeat steps S3-S6 to restore the absolute depth three-dimensional information of the blurred target after dynamic focusing.
[0163] This multi-target absolute depth estimation method based on monocular polarization 3D imaging builds upon traditional polarization 3D imaging methods. By combining the construction of a monocular ranging model with camera calibration, it can recover target absolute depth information, effectively overcoming the application limitations of traditional polarization 3D imaging techniques. Furthermore, the monocular ranging model constructed in this invention enables 3D reconstruction of the target's absolute depth using a single camera, providing a theoretical foundation and technical solution for a convenient and low-cost polarization 3D imaging method.
[0164] In addition, the post-focus camera intrinsic parameter adaptive prediction model proposed in the present invention can effectively solve the problem of camera intrinsic parameter changes caused by the need to focus in multi-target scenes at different depths; by constructing a camera intrinsic parameter adaptive prediction model, the present invention combines a monocular ranging model with a camera calibration method to achieve monocular absolute depth three-dimensional information recovery in multi-target scenes, effectively broadening the application scenarios of polarization three-dimensional imaging technology.
[0165] Another embodiment of the present invention provides a storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the steps of the multi-target absolute depth estimation method based on monocular polarization three-dimensional imaging described in the above embodiment. Another aspect of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor calls the computer program in the memory, the steps of the multi-target absolute depth estimation method based on monocular polarization three-dimensional imaging as described in the above embodiment are implemented. Specifically, the above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including several instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) or a processor to execute some steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0166] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A multi-target absolute depth estimation method based on monocular polarization 3D imaging, characterized in that: include: S1: Perform camera pre-calibration to obtain the linear variation coefficient of the principal point during camera focusing. S2: Based on the fuzziness evaluation value, the nearest area of the target object is focused until it is clear and the camera intrinsic parameters of the current position are obtained; S3: Obtain polarization images of the clear target area at different angles, and calculate the polarization degree of the reflected light from the clear target surface, as well as the azimuth angle and zenith angle of the target surface normal; S4: constructing the target surface normal and integrating to reconstruct the three-dimensional surface contour of the clear target area; S5: constructing a monocular ranging model, and obtaining a true distance between two points in the target surface space according to the monocular ranging model; S6: Recover the absolute depth information of the target based on the true distance between two points on the target surface space; S7: focusing other blurred target areas to make them clear based on the blur evaluation value and obtaining adaptive prediction values of the camera intrinsic parameters after focusing; S8: Using the adaptive prediction value of the camera intrinsic parameter after focusing, repeat steps S3-S6 to restore the absolute depth three-dimensional information of the blurred target after dynamic focusing.
2. The multi-target absolute depth estimation method based on monocular polarization 3D imaging according to claim 1, characterized in that: Said S1 comprises: Focus the camera to different positions multiple times, and obtain d1, d2, ..., d by camera calibration. k The principal point coordinates at the position are fitted based on the least squares method to obtain the linear variation coefficient η of the principal point coordinates in the x and y directions. x and η y , where the principal point is the intersection of the camera optical axis and the imaging plane.
3. The multi-target absolute depth estimation method based on monocular polarization 3D imaging according to claim 1, characterized in that: The S2 includes: The grayscale variance operator is used to represent the fuzziness evaluation value of the target image. The size of the grayscale variance operator of the target image is changed by moving the lens. When the grayscale variance operator is the largest, it means that the target image is in focus. After the focus is clear, the camera internal parameters at this moment are obtained through the camera calibration method, which is recorded as: Where dx and dy represent the pixel sizes on the abscissa and ordinate of the detector, respectively; f′ represents the focal length of the camera when the image is in focus; and (u0′, v0′) represents the coordinates of the principal point when the image is in focus.
4. The multi-target absolute depth estimation method based on monocular polarization 3D imaging according to claim 1, characterized in that: The S3 includes: S3.1: Under natural light, obtain polarization images I0, I1 at four target angles of 0°, 45°, 90°, and 135°, respectively. 45 , I 90 , I 135 ; S3.2: Obtaining Stokes vectors using polarization images at different angles, and obtaining polarization degrees using the Stokes vectors; S3.3: Calculate the azimuth of the target surface normal using polarization images at different angles and the zenith angle θ.
5. The multi-target absolute depth estimation method based on monocular polarization 3D imaging according to claim 4, characterized in that: The S4 includes: S4.1: Characterize the surface normal of the target Where p and q represent the gradients of the target microfacet normal vector in the x and y directions respectively; S4.2: Based on the Cartesian surface assumption, obtain the target polarization 3D imaging model: Among them, z = f (x, y) represents the three-dimensional surface function of the target, z x Represents the gradient of the target three-dimensional surface in the x-axis direction, z y Indicates the gradient of the target 3D surface in the y-axis direction.
6. The multi-target absolute depth estimation method based on monocular polarization 3D imaging according to claim 5, characterized in that: The S5 includes: Place a single camera horizontally, and use the triangular similarity relationship between the imaging coordinates of two points on the reference ground on the camera imaging plane, the distance between the two points and the imaging plane, the height of the camera from the reference ground, and the focal length of the camera to calculate the spatial distance between the two points.
7. The multi-target absolute depth estimation method based on monocular polarization 3D imaging according to claim 6, characterized in that: The S5 includes: S5.1: Obtain the depth distance between two points on the reference ground using the similarity relationship in the monocular ranging model: Among them, D1 and D2 are the vertical absolute depths of the two points p1 and p2 of the target on the reference ground to the camera, respectively. v =f′ / dy represents the effective focal length of the camera in the y direction, v1 and v2 represent the longitudinal coordinates of the two points in the imaging plane, v0′ represents the longitudinal coordinate of the current principal point, Δz represents the distance between the two points in the z direction in the camera coordinate system, H c The height of the camera from the reference ground; S5.2: According to the pinhole imaging model, the relationship between the target's world coordinate system and the pixel coordinate system is: Among them, R represents a 3*3 rotation matrix, T represents a 3*1 translation matrix, M0 represents the camera's intrinsic parameter matrix, M1 represents the camera's extrinsic parameter matrix, f' represents the current camera focal length, X w ,Y w ,Z w Indicates the x, y, and z axis coordinates of the target in the world coordinate system. c Indicates the z-axis coordinate value in the target camera coordinate system; S5.3: Obtain the actual distance between points p1 and p2 in the x-direction in the camera coordinate system: Among them, F u =f′ / dx represents the effective focal length of the camera in the x direction.
8. The multi-target absolute depth estimation method based on monocular polarization 3D imaging according to claim 7, characterized in that: The S6 includes: S6.1: Based on the real distance between the two target points in the x-direction in the camera coordinate system measured in step S5, obtain the real scale factor of the target polarization 3D imaging model: x s =Δx / |x p2 -x p1 | y s =x s ·d y / d x with s =Δz / |z p2 -with p1 | Among them, x s 、y s 、z s Respectively represent the scale factors of the target polarization 3D imaging model in the x, y, and z directions relative to the true value, z p1 、z p2 Respectively represent the coordinate values of points p1 and p2 in the z direction in the target polarization three-dimensional imaging model, x p1 、x p2 Respectively represent the coordinate values of points p1 and p2 in the x-direction in the target polarization 3D imaging model; The absolute depth information of the target is restored according to the true scale factor to obtain a 3D absolute depth model of the target: x t =x p ·x s ,y t =y p ·y s ,z t =z p ·z s Among them, x t 、y t 、z t Respectively represent the coordinate values in the x, y, and z directions of the reconstructed target 3D absolute depth model, x p 、y p 、z p They respectively represent the coordinate values in the x, y, and z directions in the target polarization three-dimensional imaging model.
9. The multi-target absolute depth estimation method based on monocular polarization 3D imaging according to claim 8, characterized in that: The S7 includes: S7.1: Focus the object in the blurred area until it is sharp based on the blur evaluation value, and measure the lens displacement Δf during the focusing process using a micrometer screw; S7.2: Based on the lens displacement Δf during the focusing process, the offset of the principal point coordinates in the x and y directions is obtained as follows: H x =Δf·η x H y =Δf·η y Where Δf is the focus displacement, η x ,η y is the linear variation coefficient; S7.3: Obtain the coordinates of the principal point after focusing: S7.4: Obtain the effective focal length of the camera in the x and y directions after focusing S7.5: Obtain the adaptive prediction value M0′ of the camera intrinsic parameter after focusing:
Citation Information
Patent Citations
Multi-target scene polarization three-dimensional imaging method based on deep learning
CN114663578A
Polarization three-dimensional imaging method capable of representing absolute depth of target
CN114758060A