Depth estimation method based on binocular polarization camera and atmospheric scattering model
By combining binocular polarization cameras and atmospheric scattering models, and calculating and fusing the near- and long-distance depth information, the problems of low accuracy and high cost of long-distance depth estimation in the prior art are solved, and efficient depth estimation in complex environments is achieved.
Patent Information
- Application Number
- CN202510503317.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The prior art has low accuracy and high cost in long-distance depth estimation, especially in complex environments such as haze and rainy days, which significantly affect the accuracy of depth estimation.
The depth estimation method based on binocular polarization camera and atmospheric scattering model is adopted, and the full-scene dense depth map is calculated by obtaining the information of the binocular polarization camera, and the full-scene dense depth map is generated by multi-scale fusion of the alignment, normalization and depth estimation network model.
It improves the robustness of depth estimation in complex environments, achieves full-scene depth continuity optimization, reduces hardware costs, and performs well in environments with haze and complex lighting.
Smart Images

Figure CN120014013A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a depth estimation method based on a binocular polarization camera and an atmospheric scattering model. Background Art
[0002] With the continuous development of autonomous driving, robot navigation and 3D modeling technology, depth estimation has become one of the key issues. Traditional binocular cameras obtain depth information of close-range scenes through parallax calculation, but they face great challenges in long-distance depth estimation, mainly reflected in baseline length limitations and parallax attenuation. In order to overcome this problem, researchers have proposed methods using lidar, monocular vision and other sensors for long-distance depth estimation in recent years. However, these methods usually require high equipment costs or cannot effectively deal with atmospheric scattering effects in complex environments. The atmospheric scattering effect, especially in environmental conditions such as haze and rainy days, significantly affects the attenuation of light propagation, further reducing the accuracy of long-distance depth estimation. Summary of the invention
[0003] The present invention proposes a depth estimation method based on a binocular polarization camera and an atmospheric scattering model to solve the problems of low accuracy and high cost of long-distance depth estimation in the prior art.
[0004] The present invention achieves the above-mentioned purpose through the following technical solutions:
[0005] The present invention provides a depth estimation method based on a binocular polarization camera and an atmospheric scattering model, comprising:
[0006] Get the information of binocular polarization camera;
[0007] According to the information of the binocular polarization camera, a first close-range dense depth map and a first far-scene sparse depth map are calculated based on an atmospheric scattering model;
[0008] Aligning and normalizing the first close-range dense depth map and the first far-scene sparse depth map to obtain a second close-range dense depth map and a second far-scene sparse depth map;
[0009] Verifying the sparse depth value of the atmospheric scattering model according to the second close-range dense depth map to mark the credible depth points in the second far-scene sparse depth map, and generating a confidence mask;
[0010] The second close-range dense depth map, the second far-scene sparse depth map and the confidence mask are input into a depth estimation network model, and a full-scene dense depth map is output. The depth estimation network model is used to first extract features of the second close-range dense depth map and the second far-scene sparse depth map, and is also used to use the confidence mask to complete the features of the second far-scene sparse depth map to obtain the features of the far-scene dense depth map, and is also used to perform multi-scale fusion of the features of the second close-range dense depth map and the features of the far-scene dense depth map to obtain a depth map.
[0011] Furthermore, the information of the binocular polarization camera is obtained, including:
[0012] Obtaining a camera baseline length, a camera focal length, and a parallax of a binocular polarization camera, wherein the binocular polarization camera is a binocular polarization camera with a short baseline;
[0013] Obtain the polarization angle of the binocular polarization camera, the total light intensity of the scene, and the polarization image when the polarization direction is 0°.
[0014] Further, a first close-range dense depth map and a first far-scene sparse depth map are calculated based on information of the binocular polarization camera and an atmospheric scattering model, including:
[0015] The first close-range dense depth map and the first far-scene sparse depth map are calculated based on information of the binocular polarization camera and an atmospheric scattering model, including:
[0016] The formula for calculating the first close-range dense depth map is as follows:
[0017]
[0018] Where B is the camera baseline length, f is the camera focal length, and d is the parallax. is the first close-range dense depth map;
[0019] The formula for calculating the first far scene sparse depth map is as follows:
[0020]
[0021]
[0022]
[0023] Where β is the atmospheric scattering coefficient, T(x) is the transmission function, A is the global atmospheric light intensity, is the pixel intensity received by the camera, is the real scene intensity in the absence of fog, α is the polarization angle, DoP is the degree of polarization, S0 represents the total light intensity of the scene, and I0 represents the polarization image when the polarization direction is 0°.
[0024] Furthermore, aligning and normalizing the first close-range dense depth map and the first far-scene sparse depth map includes:
[0025] Performing pixel alignment on the first close-range dense depth map and the first far-scene sparse depth map;
[0026] The depth values of the first close-range dense depth map and the first far-scene sparse depth map are normalized.
[0027] Furthermore, the depth estimation network model is also used to perform detail enhancement on the full-scene dense depth map based on U-Net.
[0028] Furthermore, the second close-range dense depth map, the second far-scene sparse depth map, and the confidence mask are input into a depth estimation network model to output a full-scene dense depth map, including:
[0029] Construct multimodal feature extraction module, sparse depth completion module, fusion module and detail enhancement module;
[0030] The multimodal feature extraction module is used to extract features of the second close-range dense depth map and the second far-scene sparse depth map using a convolutional neural network;
[0031] The sparse depth completion module is used to calculate the attention weight in the high confidence area of the second far scene sparse depth map based on the Transformer network and the confidence mask, and mark the credible area in the second far scene sparse depth map to obtain the features of the far scene dense depth map;
[0032] The fusion module is used to project the features of the distant scene dense depth map and the features of the second close-range dense depth map to a unified dimension, and then calculate the attention weights of the features of the distant scene dense depth map and the features of the second close-range dense depth map through a multi-head self-attention mechanism, concatenate and globally average the outputs of all attention heads to obtain sparse depth weights and dense depth weights, and then fuse the features of the distant scene dense depth map and the features of the second close-range dense depth map through a multi-scale fusion operation according to the sparse depth weights and dense depth weights to obtain a depth map;
[0033] The detail enhancement module is used to perform detail enhancement on the depth map using U-Net to obtain a dense depth map of the entire scene.
[0034] Furthermore, the optimization goal of the depth estimation network model is to minimize the loss function, which is:
[0035]
[0036]
[0037]
[0038]
[0039] in, is the predicted depth, is the dense depth truth, is the deep regression loss, is the sparse supervision loss, is the structure preservation loss, is the confidence mask, I is the original RGB image taken by the binocular polarization camera, N is the total number of pixels of the predicted depth and the true depth corresponding to its position, i is the position index of the predicted depth and the true depth, M is the total number of pixels of the predicted depth and the sparse depth corresponding to its position, j is the position index of the predicted depth and the sparse depth, ▽xZpred represents the gradient of Zpred in the x direction, ▽yZpred represents the gradient of Zpred in the y direction, ▽xI represents the gradient of I in the x direction, and ▽yI represents the gradient of I in the y direction.
[0040] The beneficial effects of the present invention are:
[0041] 1) The technical solution of the present invention innovatively combines a short-baseline binocular polarization camera with an atmospheric scattering model, enhancing the robustness of depth estimation in complex environments. A new depth estimation framework is provided by using a progressive extrapolation strategy that drives long-range estimation using close-range depth.
[0042] 2) This invention achieves deep continuity optimization of the entire scene through the fusion of multiple models of geometry and physical depth. It improves the adaptability and scalability of complex scenes, especially in haze and complex lighting environments, which can improve the accuracy of unmanned platform motion estimation and reduce hardware costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a flow chart of the depth estimation method based on a binocular polarization camera and an atmospheric scattering model in this application. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0045] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0047] In the description of the present invention, it should be understood that the terms "upper", "lower", "inside", "outside", "left", "right", etc. indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, or are directions or positional relationships in which the product of the invention is usually placed when in use, or are directions or positional relationships commonly understood by those skilled in the art. These directions or positional relationships are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as a limitation on the present invention.
[0048] Furthermore, the terms “first”, “second”, etc. are merely used for distinguishing descriptions and should not be understood as indicating or implying relative importance.
[0049] In the description of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms such as "setting" and "connection" should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0050] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.
[0051] like Figure 1 As shown, a depth estimation method based on a binocular polarization camera and an atmospheric scattering model includes:
[0052] S1: Get the information of binocular polarization camera;
[0053] S2: Calculate a first close-range dense depth map and a first far-scene sparse depth map based on the atmospheric scattering model according to the information of the binocular polarization camera;
[0054] S3: aligning and normalizing the first close-range dense depth map and the first far-scene sparse depth map to obtain a second close-range dense depth map and a second far-scene sparse depth map;
[0055] S4: verifying the sparse depth value of the atmospheric scattering model according to the second close-range dense depth map to mark the credible depth points in the second far-scene sparse depth map, and generating a confidence mask;
[0056] S5: Input the second close-range dense depth map, the second far-scene sparse depth map and the confidence mask into the depth estimation network model, and output the full-scene dense depth map. The depth estimation network model is used to first extract the features of the second close-range dense depth map and the second far-scene sparse depth map, and to use the confidence mask to complete the features of the second far-scene sparse depth map to obtain the features of the far-scene dense depth map, and to perform multi-scale fusion of the features of the second close-range dense depth map and the features of the far-scene dense depth map to obtain a depth map.
[0057] In some embodiments, a close-range depth estimation is performed first:
[0058] Using a short baseline (2 cm) binocular polarization camera, dense depth information within the range of 1-2 meters is obtained by the following steps:
[0059] The disparity map is calculated using the Semi-Global Matching (SGM) method. The disparity is converted into depth using formula (1) by combining the camera's intrinsic and extrinsic parameters.
[0060] (1)
[0061] Where B is the camera baseline length, f is the camera focal length, d is the parallax, and Z is the required scene depth, thereby obtaining the first close-range dense depth map.
[0062] In some embodiments, atmospheric scattering model parameter estimation is then performed, including:
[0063] Based on the close-range depth data, the key parameters of the atmospheric scattering model are calibrated:
[0064] Based on the polarization information, the atmospheric light intensity can be calculated by the following formula:
[0065] (2)
[0066] Where α is the polarization angle, d is the degree of polarization, S0 represents the total light intensity of the scene, and I0 represents the polarization image when the polarization direction is 0°.
[0067] According to the atmospheric physical scattering model, the transfer function T(x) is calculated by the global atmospheric light intensity A and the image intensity I(x):
[0068] (3)
[0069] in is the pixel intensity received by the camera, is the real scene intensity without fog, is the transmittance, which decreases with increasing distance z, A is the global atmospheric light intensity, and β is the atmospheric scattering coefficient.
[0070] Assuming that the atmospheric scattering coefficient β is a constant value, the depth information of the entire scene can be calculated :
[0071] (4)
[0072] Thus, a first far scene sparse depth map is obtained.
[0073] In some embodiments, depth compensation and optimization are then performed, including:
[0074] (1) Depth map alignment and normalization: Pixel alignment is performed on the first close-range dense depth map and the first far-scene sparse depth map to unify the resolution, normalize the depth values, and eliminate scale differences (such as depth range differences) to obtain the second close-range dense depth map and the second far-scene sparse depth map.
[0075] (2) Sparse depth confidence calculation:
[0076] Using the second closest dense depth map Sparse depth map of the second distant scene to verify the atmospheric scattering model The credible depth points in the second distant scene sparse depth map are marked to generate a confidence mask.
[0077] (5)
[0078] Where ɛ is the tolerance error threshold.
[0079] (3) Network structure
[0080] Input: Second closest dense depth map , the second distant scene sparse depth map , RGB original image I, confidence mask .
[0081] Output: dense depth map of the entire scene .
[0082] 1) Multimodal feature extraction module:
[0083] A convolutional neural network is used to extract features of the second close-range dense depth map and the second far-scene sparse depth map.
[0084] 2) Sparse depth completion module:
[0085] Based on the Transformer network and the confidence mask, the attention weight is calculated in the high confidence area of the second distant scene sparse depth map, and the credible area in the second distant scene sparse depth map is marked to obtain the features of the distant scene dense depth map. The attention weight is calculated only in the high confidence area, and the credible area (high-precision point) in the sparse depth map is marked to suppress the interference of noise or low-quality points.
[0086]
[0087] in represents the total number of pixels in the high confidence area. Q, K, and V are the three core vectors of the sparse depth map, which are used to calculate the attention weights. is the scaling factor.
[0088] 3) Fusion module:
[0089] The features of the distant scene dense depth map and the features of the second close-range dense depth map are projected to a unified dimension, and then the attention weights of the features of the distant scene dense depth map and the features of the second close-range dense depth map are calculated through a multi-head self-attention mechanism, and the outputs of all attention heads are concatenated and globally averaged to obtain sparse depth weights and dense depth weights, and then the features of the distant scene dense depth map and the features of the second close-range dense depth map are fused through a multi-scale fusion operation according to the sparse depth weights and dense depth weights to obtain a depth map.
[0090] Generate dynamic weights by computing the correlation between sparse depth (far field) and dense depth (near field) features and :
[0091] 1. Feature Projection:
[0092] Sparse deep features and dense deep features Project to uniform dimension:
[0093] ;
[0094] ;
[0095] ;
[0096] , and is a learnable weight matrix, Q is used to query the correlation between sparse depth and dense depth, and K and V generate key-value pairs based on dense depth.
[0097] 2. Multi-head self-attention calculation:
[0098] The attention weights of sparse depth to dense depth are calculated through the multi-head self-attention mechanism:
[0099]
[0100] Where: h: attention head number (H heads in total), : scaling factor, : The learnable weight of the h-th head.
[0101] 3. Global weight aggregation:
[0102] The outputs of all attention heads are concatenated and globally averaged to obtain sparse and dense weights:
[0103]
[0104]
[0105] According to the importance of local features, the weights of dense and sparse depths are adaptively assigned, so that the network relies on dense depth in the near field and sparse depth in the far field.
[0106] The final fusion features are:
[0107] (4) Loss function
[0108] 1) Deep regression loss: (6)
[0109] 2) Sparse supervision loss:
[0110] (7)
[0111] 3) Structure preservation loss: Combined with the original RGB image gradient, ensure that the generated depth map is consistent with the scene structure:
[0112] (8)
[0113] 4) Total loss
[0114] (9)
[0115] in, is the predicted depth, is the dense depth truth, is the deep regression loss, is the sparse supervision loss, is the structure preservation loss, , , They are the weight of the deep regression loss, the weight of the sparse supervision loss, and the weight of the structure preservation loss. is the confidence mask, I is the original RGB image taken by the binocular polarization camera, N is the total number of pixels of the predicted depth and the true depth corresponding to its position, i is the position index of the predicted depth and the true depth, M is the total number of pixels of the predicted depth and the sparse depth corresponding to its position, j is the position index of the predicted depth and the sparse depth, ▽xZpred represents the gradient of Zpred in the x direction, ▽yZpred represents the gradient of Zpred in the y direction, ▽xI represents the gradient of I in the x direction, and ▽yI represents the gradient of I in the y direction.
[0116] The beneficial effects of the present invention are:
[0117] 1) The technical solution of the present invention innovatively combines a short-baseline binocular polarization camera with an atmospheric scattering model, enhancing the robustness of depth estimation in complex environments. A new depth estimation framework is provided by using a progressive extrapolation strategy that drives long-range estimation using close-range depth.
[0118] 2) This invention achieves deep continuity optimization of the entire scene through the fusion of multiple models of geometry and physical depth. It improves the adaptability and scalability of complex scenes, especially in haze and complex lighting environments, which can improve the accuracy of unmanned platform motion estimation and reduce hardware costs.
[0119] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A depth estimation method based on a binocular polarization camera and an atmospheric scattering model, characterized in that: include: Get the information of binocular polarization camera; According to the information of the binocular polarization camera, a first close-range dense depth map and a first far-scene sparse depth map are calculated based on an atmospheric scattering model; Aligning and normalizing the first close-range dense depth map and the first far-scene sparse depth map to obtain a second close-range dense depth map and a second far-scene sparse depth map; Verifying the sparse depth value of the atmospheric scattering model according to the second close-range dense depth map to mark the credible depth points in the second far-scene sparse depth map, and generating a confidence mask; The second close-range dense depth map, the second far-scene sparse depth map and the confidence mask are input into a depth estimation network model, and a full-scene dense depth map is output. The depth estimation network model is used to first extract features of the second close-range dense depth map and the second far-scene sparse depth map, and is also used to use the confidence mask to complete the features of the second far-scene sparse depth map to obtain the features of the far-scene dense depth map, and is also used to perform multi-scale fusion of the features of the second close-range dense depth map and the features of the far-scene dense depth map to obtain a depth map.
2. The depth estimation method based on a binocular polarization camera and an atmospheric scattering model according to claim 1, characterized in that: Get the information of the binocular polarization camera, including: Obtaining a camera baseline length, a camera focal length, and a parallax of a binocular polarization camera, wherein the binocular polarization camera is a binocular polarization camera with a short baseline; Obtain the polarization angle of the binocular polarization camera, the total light intensity of the scene, and the polarization image when the polarization direction is 0°.
3. The depth estimation method based on a binocular polarization camera and an atmospheric scattering model according to claim 2, characterized in that: The first close-range dense depth map and the first far-scene sparse depth map are calculated based on information of the binocular polarization camera and an atmospheric scattering model, including: The formula for calculating the first close-range dense depth map is as follows: , Where B is the camera baseline length, f is the camera focal length, and d is the parallax. is the first close-range dense depth map; The formula for calculating the first far scene sparse depth map is as follows: , , , Where β is the atmospheric scattering coefficient, T(x) is the transmission function, A is the global atmospheric light intensity, is the pixel intensity received by the camera, is the real scene intensity in the absence of fog, α is the polarization angle, DoP is the degree of polarization, S0 represents the total light intensity of the scene, and I0 represents the polarization image when the polarization direction is 0°.
4. The depth estimation method based on a binocular polarization camera and an atmospheric scattering model according to claim 1, characterized in that: Aligning and normalizing the first close-range dense depth map and the first far-scene sparse depth map, including: Performing pixel alignment on the first close-range dense depth map and the first far-scene sparse depth map; The depth values of the first close-range dense depth map and the first far-scene sparse depth map are normalized.
5. The depth estimation method based on a binocular polarization camera and an atmospheric scattering model according to claim 1, characterized in that: The depth estimation network model is also used to perform detail enhancement on the full-scene dense depth map based on U-Net.
6. The depth estimation method based on a binocular polarization camera and an atmospheric scattering model according to claim 1, characterized in that: Inputting the second close-range dense depth map, the second far-scene sparse depth map, and the confidence mask into a depth estimation network model, and outputting a full-scene dense depth map, including: Construct multimodal feature extraction module, sparse depth completion module, fusion module and detail enhancement module; The multimodal feature extraction module is used to extract features of the second close-range dense depth map and the second far-scene sparse depth map using a convolutional neural network; The sparse depth completion module is used to calculate the attention weight in the high confidence area of the second far scene sparse depth map based on the Transformer network and the confidence mask, and mark the credible area in the second far scene sparse depth map to obtain the features of the far scene dense depth map; The fusion module is used to project the features of the distant scene dense depth map and the features of the second close-range dense depth map to a unified dimension, and then calculate the attention weights of the features of the distant scene dense depth map and the features of the second close-range dense depth map through a multi-head self-attention mechanism, concatenate and globally average the outputs of all attention heads to obtain sparse depth weights and dense depth weights, and then fuse the features of the distant scene dense depth map and the features of the second close-range dense depth map through a multi-scale fusion operation according to the sparse depth weights and dense depth weights to obtain a depth map; The detail enhancement module is used to perform detail enhancement on the depth map using U-Net to obtain a dense depth map of the entire scene.
7. The depth estimation method based on a binocular polarization camera and an atmospheric scattering model according to claim 1, characterized in that: The optimization goal of the depth estimation network model is to minimize the loss function, which is: , , , , in, is the predicted depth, is the dense depth truth, is the deep regression loss, is the sparse supervision loss, is the structure preservation loss, , , They are the weight of the deep regression loss, the weight of the sparse supervision loss, and the weight of the structure preservation loss. is the confidence mask, I is the original RGB image taken by the binocular polarization camera, N is the total number of pixels of the predicted depth and the true depth corresponding to its position, i is the position index of the predicted depth and the true depth, M is the total number of pixels of the predicted depth and the sparse depth corresponding to its position, j is the position index of the predicted depth and the sparse depth, ▽xZpred represents the gradient of Zpred in the x direction, ▽yZpred represents the gradient of Zpred in the y direction, ▽xI represents the gradient of I in the x direction, and ▽yI represents the gradient of I in the y direction.
Citation Information
Patent Citations
Sparse depth map densing method and device
CN104346608A
360-degree environment depth completion and map reconstruction method based on cross-modal fusion
CN114119889A
Image processing method and device
CN114511778A
Depth completion method based on sparse representation
CN116862965A
Depth completion training method of sparse depth and related equipment
CN117557887A
Cited By
Degree quantization underwater depth estimation method based on confidence guiding fusion and monocular backspacing mechanism
CN121527149A