Underwater binocular three-dimensional reconstruction method based on refractive geometry modeling and adaptive matching
By using a method based on refractive geometry modeling and adaptive matching, the problems of information transmission and target consistency in underwater 3D reconstruction were solved, achieving high-precision underwater 3D reconstruction results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-05
- Publication Date
- 2026-05-01
AI Technical Summary
Existing underwater 3D reconstruction methods lack information transfer and target consistency between various stages, resulting in imaging errors and insufficient matching accuracy.
A method based on refraction geometry modeling and adaptive matching is adopted. The camera is calibrated and geometrically corrected by constructing a refraction equivalent pinhole model. Image enhancement is performed by combining local texture and spectral characteristics. A pixel-level image quality evaluation index is constructed to achieve adaptive fusion and filtering optimization of disparity maps, and finally, a 3D point cloud is generated.
It improves the accuracy and stability of underwater 3D reconstruction, especially in complex underwater environments, and can more accurately recover parallax-depth relationships and epipolar constraints to generate high-precision 3D point clouds.
Smart Images

Figure CN121788729B_ABST
Abstract
Description
Underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching Technical Field
[0001] This invention belongs to the technical field of underwater 3D reconstruction, specifically relating to an underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching. Background Technology
[0002] Traditional underwater 3D reconstruction work directly uses the pinhole camera model from the air: either the intrinsic and extrinsic parameters are calibrated in the air and then used directly after deployment; or even when a calibration board is photographed underwater, a "single-medium pinhole + simple distortion" model is still used, simply attributing the refraction effect to a traditional distortion model for empirical fitting. This approach does not physically distinguish between multi-medium imaging of "air-glass-water," resulting in systematic errors in the parallax-depth geometric relationship itself.
[0003] In terms of the goals and methods of image preprocessing (enhancement), traditional underwater image enhancement mostly serves the human eye's perception: by using methods such as white balance, color correction, histogram equalization, Retinex, and dehazing, the image is made "clearer and better looking," and then the enhanced image is directly fed into a general stereo matching algorithm. The enhancement strategies of these methods are often global or locally adaptive and unrelated to matching, without explicitly modeling "whether this area is suitable for matching."
[0004] In terms of stereo matching and cost aggregation strategies, traditional underwater binocular stereo reconstruction mostly directly applies matching algorithms used on land, such as block matching, SGM, and ELAS. At most, adaptive support weights based on color difference and spatial distance are used during cost aggregation. However, no specific modeling is done for the unique imaging statistics of underwater imaging, such as "color attenuation, turbidity, scattering noise, and low texture." Some existing technologies perform block matching, SGM, ELAS, etc., after enhancement, but the enhancement information does not enter the matching process; it only performs some general enhancement / denoising / dehazing operations at the input image level.
[0005] For underwater binocular 3D reconstruction, traditional algorithms make local replacements or simple stacking on the basis of existing modules: for example, adding a CLAHE on the basis of air calibration parameters, and then connecting a standard SGM; or simply dehazing the underwater image and then using a publicly implemented stereo matching library. However, the drawback of such design is that there is a lack of information transmission and target consistency between the various links. Summary of the Invention
[0006] The purpose of this invention is to address the aforementioned shortcomings in the existing technology by providing an underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching, thereby solving the problems of lack of information transmission and target consistency between various stages of the existing underwater binocular 3D reconstruction.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0008] An underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching includes the following steps:
[0009] S1. Based on the multi-medium refraction effect of air-glass-water, a refraction equivalent pinhole model is constructed, and the binocular camera system is calibrated by minimizing the reprojection error. Then, the calibration parameters are used to perform geometric correction on the underwater binocular images.
[0010] S2. For the geometrically corrected underwater binocular image, the adaptive window radius is calculated based on the contrast of local region texture, and a weighted dark channel is constructed in combination with underwater spectral characteristics to estimate transmittance and recover scene irradiance, thereby obtaining the enhanced underwater binocular image.
[0011] S3. Based on the local texture intensity, local contrast and residual turbidity of the enhanced underwater binocular images, a pixel-level joint evaluation index for image quality is constructed.
[0012] S4. Two disparity maps are calculated using two sets of matching parameters respectively, and the two disparity maps are fused at the pixel level based on the joint image quality evaluation index to obtain a fused disparity map.
[0013] S5. Calculate the disparity confidence of the fused disparity map, construct a comprehensive weight based on the disparity confidence and the joint image quality evaluation index, filter and optimize the fused disparity map based on the comprehensive weight, and back-project the optimized disparity map to the three-dimensional space through the calibration parameters in step S1 to generate a three-dimensional point cloud of the underwater scene.
[0014] Furthermore, S1 specifically includes:
[0015] Based on the theory of equivalent virtual pinhole camera, the multi-medium of air-glass-water is equivalent to a virtual camera, and then a refractive equivalent pinhole model related to depth is constructed.
[0016] By acquiring underwater calibration plate images within a preset depth range and constructing a reprojection error objective function that includes depth-related intrinsic parameters and distortion parameters in the refraction equivalent pinhole model, the optimal parameters of the reprojection error objective function are iteratively solved using a nonlinear least squares algorithm. This completes the calibration of the binocular camera system, and the calibration parameters are used to perform geometric correction on the underwater binocular images to restore accurate epipolar constraints and disparity-depth relationships.
[0017] Furthermore, in step S1, within a preset depth range, the equivalent intrinsic parameters and distortion parameters of the refractive equivalent pinhole model are fitted with the depth, as expressed as:
[0018]
[0019]
[0020] In the formula, For equivalent internal reference, For a given depth, For the intrinsic parameter fitting constant term, , These are the intrinsic parameter fitting coefficients. For distortion parameters, These are the linear fitting coefficients. This is the distortion fitting constant term.
[0021] Furthermore, in S2, the adaptive window radius is calculated based on the contrast of the local region texture, which includes:
[0022] The local texture features of underwater stereo images are calculated and represented as follows:
[0023]
[0024] In the formula, For pixels Local texture features, To fix the small window, To fix the total number of pixels within a small window, For pixels grayscale value, For window The mean;
[0025] Pixels Local texture features Normalization is performed to obtain normalized local texture features. and based on The adaptive window radius is calculated as follows:
[0026]
[0027] In the formula, To adapt to the window radius, Minimum window radius, This represents the maximum window radius.
[0028] Furthermore, in S2, a weighted dark channel is constructed by combining underwater spectral characteristics, which is represented as follows:
[0029]
[0030] In the formula, For weighted dark channels, In pixels Centered on, with radius In the field, For channel weights, For pixels In color channels Pixel values; Indicates color channels, For red channels, As a green channel, For the blue channel;
[0031] Combined with weighted dark channel estimation of transmittance:
[0032]
[0033] In the formula, Transmittance, To maintain a small amount of fog-like adjustment factor, For ambient background light;
[0034] Using guided filtering to improve transmittance The transmittance is refined to obtain the refined transmittance. Based on the refined transmittance To restore scene irradiance, a color compensation term is introduced into the red channel, which is expressed as:
[0035]
[0036] In the formula, The scene irradiance after red channel compensation. The initial scene irradiance for the red channel recovery. These are empirical parameters.
[0037] Furthermore, regarding channel weights Based on the water color and local color distribution, the channel weights are adaptively reduced. The red channel weight in the equation is expressed as:
[0038]
[0039] In the formula, For pixels In the weight of the red channel, , This is an empirical coefficient. For pixels Local eigenvalues.
[0040] Furthermore, in step S3, a pixel-level joint image quality evaluation index is constructed based on the local texture intensity, local contrast, and residual turbidity of the enhanced underwater binocular image, which is expressed as:
[0041]
[0042] In the formula, As a joint evaluation index for image quality, To normalize local texture intensity, To normalize local contrast, To normalize residual turbidity, , , These are the weighting coefficients.
[0043] Furthermore, S4 specifically includes:
[0044] Configure two sets of matching parameters, one of which uses a small window and weak smoothing constraints to obtain a disparity map that is biased towards detail preservation. ;
[0045] Another set of matching parameters, using a large window and strong smoothing constraints, yields a more robust disparity map. ;
[0046] Based on joint evaluation index of image quality disparity map Parallax diagram Pixel-level fusion is represented as follows:
[0047]
[0048] In the formula, To fuse disparity maps, For pixels Parallax maps that prioritize detail preservation. For pixels Parallax maps with a bias towards robustness.
[0049] Furthermore, S5 specifically includes:
[0050] For the fused disparity map, a confidence index based on the cost interval is calculated according to the minimum and second minimum matching costs of each pixel. This index is then combined with a left-right consistency check to obtain the final disparity confidence map. ;
[0051] Based on joint evaluation index of image quality With disparity confidence map The comprehensive weights are constructed as follows:
[0052]
[0053] In the formula, For comprehensive weighting;
[0054] And based on this comprehensive weight For fused disparity maps Adaptive filtering and interpolation are performed, and the equivalent focal length and baseline length of the refraction equivalent pinhole model are combined to fuse the disparity map. By converting triangulation into depth representation and combining it with the calibration parameters in step S1, the back projection of pixels to the world coordinate system is completed, thereby obtaining the three-dimensional point cloud of the underwater scene.
[0055] Furthermore, based on the comprehensive weighting For fused disparity maps Adaptive filtering and interpolation processing is performed, including:
[0056] For the overall weight Pixels with weights below the threshold are not used directly for their disparity values. Instead, they are corrected by plane fitting or weighted interpolation using high-weight pixels in the local neighborhood. For large water background areas with weights consistently below the threshold, they are directly removed during the 3D reconstruction stage.
[0057] The underwater binocular 3D reconstruction method based on refraction geometry modeling and adaptive matching provided by this invention has the following beneficial effects:
[0058] This invention, starting from the underwater refraction imaging mechanism, constructs a complete 3D reconstruction framework of "geometric correction—adaptive enhancement—adaptive matching." First, in the calibration stage, the air-glass-water multi-medium refraction effect is explicitly considered. An underwater-specific refractive equivalent pinhole model (imaging model) is established using an equivalent pinhole method with distortion compensation. This model is used for joint calibration and geometric correction of the binocular camera system, restoring as accurate an epipolar constraint and disparity-depth relationship as possible. Second, on the geometrically corrected image, imaging quality indicators are defined based on local gradient, variance, and contrast. The window size and contrast stretching intensity are adaptively adjusted to implement adaptive window enhancement "for stereo matching," improving the matching compatibility of weak textures and degraded areas. Finally, in the cost aggregation stage of stereo matching, priors such as spatial proximity, enhanced brightness similarity, local imaging quality, and underwater turbidity are incorporated into the adaptive weight design. Within frameworks such as SGM, robust disparity estimation is achieved for complex underwater degradation conditions, resulting in more accurate and stable underwater 3D reconstruction results.
[0059] This invention tightly couples enhancement design with stereo matching: first, a "matchability / imaging quality" index is constructed using local gradient, variance, and contrast; then, the enhancement window size and contrast stretching intensity are adaptively adjusted based on this index. Areas with rich texture undergo only slight enhancement to preserve details, while areas with weak texture or severe turbidity have larger statistical windows and increased contrast stretching, directly improving the discriminative power during cost calculation. This "matching-oriented adaptive enhancement" essentially uses matching requirements to guide preprocessing, rather than treating preprocessing as an independent module disconnected from matching. Attached Figure Description
[0060] Figure 1 is a flowchart of the underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching in this embodiment.
[0061] Figure 2 is a schematic diagram of the binocular camera system in this embodiment. Detailed Implementation
[0062] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0063] The underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching in this embodiment, referring to Figure 1, specifically includes the following:
[0064] S1. Based on the multi-medium refraction effect of air-glass-water, a refraction equivalent pinhole model is constructed, and the binocular camera system is calibrated by minimizing the reprojection error. Then, the calibration parameters are used to perform geometric correction on the underwater binocular images.
[0065] In the underwater stereo camera system, referring to Figure 2, in the figure: C L C R These are the centers of the left and right cameras, respectively; Q1 and R1 are the refraction points at the water-glass interface, and Q2 and R2 are the refraction points at the glass-air interface; R in R is the incident direction vector; out d is the outgoing direction vector; d is the distance from the camera's optical center to the inner surface of the glass; t is the glass thickness; P ’ c It is the equivalent virtual image point after refraction.
[0066] Cameras are typically encased in a waterproof housing, isolated from external water through a flat glass window. The imaging light path passes sequentially through three media: water, glass, and air, resulting in refractive distortion that significantly deviates from the traditional pinhole camera model. Let the refractive indices of water, glass, and air be... The normal vector of the outer surface of the planar glass is Camera optical center The distance to the inner surface is For a point in the water , , , Representing points in water The three axial components in the camera coordinate system
[0067] The imaging rays reaching the camera must satisfy Snell's law. Represented as vectors, the incident direction... and the direction of launch satisfy:
[0068]
[0069] In the formula, The refractive index of two adjacent media , These are the calculated coefficients in the vector form of Snell's law;
[0070] Light rays refract sequentially at the water-glass and glass-air interfaces. The direction vector incident on the camera coordinate system can be obtained by applying Snell's law twice. Ultimately, like a dot , , They are respectively, and they satisfy:
[0071]
[0072] In the formula, For the camera intrinsic parameter matrix, As a scale factor, is the homogeneous pixel coordinate vector of the image point on the image plane;
[0073] Traditional underwater calibration methods often approximate the entire system as a pinhole model in a single medium, compensating only through radial / tangential distortion terms, which fails to accurately describe refractive distortion varying with the incident angle. Therefore, this embodiment introduces explicit refractive geometry modeling and models the medium plane parameters and equivalent refractive distortion parameters separately. Directly solving point-by-point on the three-medium model is costly; therefore, this embodiment adopts the concept of an "equivalent virtual pinhole camera": the "water-glass-air" system is equivalent to a virtual camera in the water, thus constructing a depth-dependent refractive equivalent pinhole model, whose equivalent intrinsic parameters... and distortion parameters With imaging depth Slow change can be achieved through the following two steps:
[0074] (1) For a given depth Using the physical refraction model, a set of spatial direction vectors The direction in the water was traced through two refractions. Then, align the projection results with the ideal underwater pinhole model to obtain the equivalent intrinsic parameters at that depth. and distortion parameters ;
[0075] (2) Within a certain depth range ~ Internal, for equivalent internal references and distortion parameters Polynomial or piecewise linear fitting can be expressed as:
[0076]
[0077] In the formula, For a given depth, For the intrinsic parameter fitting constant term, , These are the intrinsic parameter fitting coefficients. For distortion parameters, These are the linear fitting coefficients. This is the distortion fitting constant term.
[0078] The depth-dependent refractive equivalent pinhole model in this embodiment is geometrically closer to the real refractive light path, while maintaining the analytical form of the pinhole projection, which facilitates subsequent stereo calibration and reconstruction.
[0079] During the calibration phase, several underwater calibration plate images were acquired at preset depth locations. A reprojection error objective function was constructed, which included the depth-related intrinsic parameters and distortion parameters in the refraction equivalent pinhole model. The parameters of the refraction equivalent pinhole model were then solved through the following optimization.
[0080] The objective function for the reprojection error is expressed as:
[0081]
[0082] In the formula, For the calibration board Three-dimensional point coordinates; For the first The pose corresponding to the image; These are the measured image points; This is a projection function based on the refraction equivalent model; This represents the parameters to be solved, including camera intrinsic parameters, medium plane parameters, refraction equivalent model parameters, etc. This represents the total number of underwater calibration board images acquired. This represents the total number of calibration plate feature points used in the calculation for each image.
[0083] This embodiment uses nonlinear least squares methods such as Levberg–Marquardt to iteratively solve the problem. Because the model incorporates depth-related terms, in practical applications, the working distance can be adjusted according to the task scenario. right and Adaptive updates are performed to improve projection accuracy near this depth. Finally, refractive equivalent pinhole models are established for the left and right cameras respectively. Then, stereo geometric calibration is completed through epipolar constraints. The calibration parameters are used to perform geometric correction on the underwater binocular images to restore accurate epipolar constraints and disparity-depth relationships, providing a geometric basis for subsequent accurate 3D reconstruction.
[0084] This embodiment S1 starts from the refractive imaging mechanism and explicitly introduces an "equivalent pinhole + refractive distortion compensation" model during the calibration stage to jointly fit and correct the imaging geometry in the underwater environment. Instead of simply treating the underwater environment as "air with some distortion," it incorporates multi-medium effects into the imaging model through additional refractive parameters, and then completes epipolar correction and parallax-depth relationship calibration on this basis. The direct result of this difference is that, under the same binocular hardware conditions, the method of this invention can obtain more accurate and consistent geometric reconstruction, especially when the working distance varies greatly or the field of view is wide, the depth error is more easily controlled within a quantifiable range.
[0085] S2. For the geometrically corrected underwater binocular image, the adaptive window radius is calculated based on the contrast of local region texture, and a weighted dark channel is constructed in combination with underwater spectral characteristics to estimate transmittance and recover scene irradiance, thereby obtaining the enhanced underwater binocular image.
[0086] Underwater image degradation can be addressed by referencing atmospheric scattering models, while also considering wavelength-dependent absorption characteristics. For color channels c∈{R,G,B}, we have:
[0087]
[0088] in, For red channels, As a green channel, For the blue channel, For observation images; Irradiance of the scene to be restored; For ambient background light; For pixels The transmittance satisfies:
[0089]
[0090] In the formula, This represents the attenuation coefficient of water in color channel c. Optical path length;
[0091] Since water attenuates red light much more than green and blue channels, directly using the land dark channel prior (treating RGB equally) would result in the red channel in distant regions being approximately zero, thus severely underestimating transmittance. Therefore, two mechanisms, channel weighting and adaptive windowing, are introduced within the dark channel framework.
[0092] Traditional dark passage Defined as:
[0093]
[0094] In the formula, For The neighborhood centered at r; The image to be processed in pixels Color channel The pixel brightness value below;
[0095] This embodiment is based on underwater spectral characteristics, Rewritten as a weighted dark channel:
[0096]
[0097] In the formula, In pixels Centered on, with radius In the field, For channel weights, satisfying .
[0098] Based on the water color and local color distribution, the red channel weight is adaptively reduced, which can be expressed as:
[0099]
[0100] in, For pixels The local eigenvalues can be the average value of the local blue-green channels or the "distance" estimated through color attenuation priors; For pixels Weight in the red channel; , This is an empirical coefficient; the farther away the scene / the more blue-green it appears, The smaller; For pixels In terms of the weight of the green channel, For pixels Weighting in the blue channel;
[0101] radius Adaptive strategy adopted; ambient background light It can still be estimated using several pixels corresponding to the weighted maximum value of the dark channel in the image. The difference lies in considering factors when selecting candidate pixels. Instead of a traditional dark channel, this approach addresses the issue in underwater scenes where nearby objects exhibit rich textures and high contrast, while distant areas have weak textures and noticeable fog. A fixed-size dark channel window struggles to balance detail and robustness in both types of areas. Therefore, this embodiment introduces an adaptive window radius based on the contrast definition of local texture regions. .
[0102] First, calculate the local variance or gradient energy of the grayscale image:
[0103]
[0104] In the formula, For pixels Local texture features, To fix the small window, To fix the total number of pixels within a small window, For pixels grayscale value, For window The mean.
[0105] Pixels Local texture features Normalization is performed to obtain normalized local texture features. Therefore, the adaptive window radius is defined as:
[0106]
[0107] In the formula, To adapt to the window radius, Minimum window radius, The maximum window radius;
[0108] Strong texture ( If the texture is large, select a small window to avoid blurry edges; if the texture is weak... If the window size is small, a larger window is selected to improve the stability of the estimation. Based on this, transmittance is estimated using a weighted dark channel:
[0109]
[0110] In the formula, Transmittance, To maintain a small amount of fog-like adjustment factor;
[0111] To reduce noise caused by texture, this embodiment further employs guided filtering to refine the transmittance while maintaining alignment with the structure edges. Guided filter radius. It can also be based on Adaptive adjustment: A smaller radius is used in the edge region and a larger radius is used in the flat region, thus obtaining a refined transmittance. The scene irradiance is restored using the following formula:
[0112]
[0113] In the formula, This refers to the restored scene irradiance (i.e., the clear pixel value after dehazing / enhancement). This is the lower limit of transmittance;
[0114] Considering that red light attenuates too much underwater, a simple color compensation term is introduced into the red channel after restoration:
[0115]
[0116] In the formula, The scene irradiance after red channel compensation. The initial scene irradiance for the red channel recovery. These are empirical parameters;
[0117] The final combined image is an enhanced color image, which has significantly better contrast and color than the original image, providing more stable texture information for subsequent stereo matching.
[0118] S3. Based on the local texture intensity, local contrast and residual turbidity of the enhanced underwater binocular images, a pixel-level joint evaluation index for image quality is constructed.
[0119] Given enhanced left and right images , By fusing multiple local features, a pixel-level image quality and joint image quality evaluation index are constructed. This evaluation map comprehensively considers multiple factors such as local texture intensity, local contrast, and residual turbidity to characterize pixels. The quality of the matching conditions at a given location. Local texture intensity is measured by the gradient magnitude or the eigenvalues of the structure tensor, denoted by its normalized result. Local contrast is measured by the luminance variance or the increase in contrast before and after enhancement, and normalized to... The residual turbidity is then determined using the obtained transmittance. The reaction was carried out, and the normalized turbidity was defined. It is inversely proportional to transmittance. Combining the above three factors with a weighted linear ratio, it can be expressed as:
[0120]
[0121]
[0122] In the formula, , , These are the weighting coefficients;
[0123] when When the value is close to 1, it indicates that the texture of the area is clear, the contrast is high, and the scattering is weak, so the expected matching result is more reliable; when When the size is small, it generally corresponds to a low-texture or highly turbid water background, making matching more difficult.
[0124] S4. Two disparity maps are calculated using two sets of matching parameters respectively, and the two disparity maps are fused at the pixel level based on the joint image quality evaluation index to obtain a fused disparity map.
[0125] In obtaining quality evaluation maps (joint image quality evaluation indicators) After that, based on The stereo matching process is adaptively adjusted.
[0126] First, in constructing the pixel-disparity cost volume At that time, the spatial scale of the supporting window changes with Dynamic variation: A smaller window is used in high-quality regions to preserve high-frequency geometric details; a larger window is used in low-quality regions to improve the stability of cost estimation. To balance implementation complexity and efficiency, this embodiment does not explicitly change the window size for each pixel. Instead, it configures two sets of representative matching parameters: one set uses a small window and weaker smoothing constraints to obtain a disparity map that prioritizes detail preservation. Another set uses a large window and strong smoothing constraints to obtain a more robust disparity map. Based on this, pixel-level fusion of the two disparity maps is performed using a quality evaluation map:
[0127]
[0128] In the formula, To fuse disparity maps, For pixels Parallax maps that prioritize detail preservation. For pixels Robust disparity maps;
[0129] In this way, in high-quality areas, the fusion result is closer to detailed parallax; in low-quality areas, it relies more on robust parallax, achieving spatial adaptation of the matching strategy. Simultaneously, when using semi-global matching (SGM) for cost aggregation, through... The penalty coefficient in the smoothing term is adjusted regionally: the smoothing coefficient is reduced in high-quality edge regions to allow the disparity to change rapidly at depth abrupt changes; the smoothing coefficient is increased in low-quality, weak-texture regions to make the disparity field smoother and suppress noise overall.
[0130] S5. Calculate the disparity confidence of the fused disparity map, construct a comprehensive weight based on the disparity confidence and the joint image quality evaluation index, perform filtering optimization on the fused disparity map based on the comprehensive weight, and backproject the optimized disparity map to three-dimensional space through the calibration parameters in step S1 to generate a three-dimensional point cloud of the underwater scene. Specifically, this includes the following:
[0131] After obtaining the fused disparity map Then, disparity confidence is further introduced to filter out unreliable matches. For each pixel, the minimum and second minimum matching costs are denoted as follows: and Then, a confidence index based on the cost interval can be defined:
[0132]
[0133] In the formula, As a confidence level indicator, To prevent tiny constants with a denominator of zero;
[0134] By combining this with a left-right consistency check, the final disparity confidence map is obtained. ;
[0135] Image quality assessment and matching confidence reflect two aspects: "whether it can be seen clearly" and "whether it matches well," respectively. Multiplying the two together yields a comprehensive weight:
[0136]
[0137] In the formula, This is the overall weight.
[0138] And based on this weight, adaptive filtering and interpolation are performed on the disparity map: For Pixels with lower weights (below the threshold weight) are not used directly for their parallax. Instead, they are corrected by plane fitting or weighted interpolation using high-weight pixels in the local neighborhood. For large water background areas with consistently low weights (below the threshold weight), they are directly removed during the 3D reconstruction stage to avoid generating a large number of isolated noise points.
[0139] Finally, the equivalent focal length given by the refraction equivalent pinhole model is... Given a known baseline length, fuse disparity maps Transformed into depth representation using triangulation:
[0140]
[0141] In the formula, The calculated scene depth value, The baseline distance for the binocular camera system;
[0142] Then, combining the extrinsic parameter relationship (calibration parameters) established in step S1, the back projection from pixels to the world coordinate system is completed, thereby obtaining the 3D point cloud of the underwater scene. Because an adaptive mechanism based on image quality and visibility evaluation is introduced throughout the matching and reconstruction process, the method in this embodiment effectively suppresses matching artifacts in the water background and low-contrast areas while preserving the surface details of targets such as rocks, significantly improving the overall accuracy and stability of underwater stereo reconstruction.
[0143] Although specific embodiments of the invention have been described in detail with reference to the accompanying drawings, this should not be construed as limiting the scope of protection of this patent. Various modifications and variations that can be made by a person skilled in the art without inventive effort within the scope described in the claims still fall within the scope of protection of this patent.
Claims
1. An underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching, characterized in that, Includes the following steps: S1. A refractive equivalent pinhole model is constructed based on the multi-medium refraction effect of air-glass-water. The binocular camera system is calibrated by minimizing the reprojection error, and then the underwater binocular image is geometrically corrected using the calibration parameters. S2. For the geometrically corrected underwater binocular image, an adaptive window radius is calculated based on the contrast of local texture, and a weighted dark channel is constructed in combination with underwater spectral characteristics to estimate transmittance and recover scene irradiance, thereby obtaining an enhanced underwater binocular image. S3. Based on the local texture intensity, local contrast, and residual turbidity of the enhanced underwater binocular image, a pixel-level image quality joint evaluation index is constructed. S4. Two disparity maps are calculated using two sets of matching parameters, and the two disparity maps are fused at the pixel level based on the image quality joint evaluation index to obtain a fused disparity map. S5. Calculate the disparity confidence score of the fused disparity map, construct a comprehensive weight based on the disparity confidence score and the joint image quality evaluation index, perform filtering optimization on the fused disparity map based on the comprehensive weight, and then pass the optimized disparity map through step S1. The calibration parameters in the image are back-projected into three-dimensional space to generate a three-dimensional point cloud of the underwater scene; S1 specifically includes: based on the equivalent virtual pinhole camera theory, the multi-medium of air-glass-water is equivalent to a virtual camera, thereby constructing a depth-related refraction equivalent pinhole model; by acquiring underwater calibration plate images within a preset depth range, and constructing a reprojection error objective function containing the depth-related equivalent intrinsic parameters and distortion parameters in the refraction equivalent pinhole model, the optimal parameters of the reprojection error objective function are iteratively solved using a nonlinear least squares algorithm, thereby completing the calibration of the binocular camera system, and using the calibration parameters to perform geometric correction on the underwater binocular images to restore accurate epipolar constraints and disparity-depth relationships; in S1, within the preset depth range, the equivalent intrinsic parameters and distortion parameters of the refraction equivalent pinhole model are fitted with the depth, which are expressed as: In the formula, For equivalent internal reference, For a given depth, For the intrinsic parameter fitting constant term, 、 These are the internal parameter fitting coefficients. For distortion parameters, These are the linear fitting coefficients. The initial distortion parameters are at zero depth; in S2, the adaptive window radius is calculated based on the contrast of the local region texture, which includes: calculating the local texture features of the underwater binocular image, expressed as: In the formula, For pixels Local texture features, To fix the small window, To fix the total number of pixels within a small window, For pixels grayscale value, For window The mean of the pixels; Local texture features Normalization is performed to obtain normalized local texture features. and based on The adaptive window radius is calculated as follows: In the formula, To adapt to the window radius, Minimum window radius, This represents the maximum window radius.
2. The underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching according to claim 1, characterized in that, In S2, a weighted dark channel is constructed by combining underwater spectral characteristics, which is represented as follows: In the formula, For weighted dark channels, In pixels Centered on, with radius In the field, For channel weights, For the image to be processed in pixels Color channel The pixel brightness value below; Indicates color channels, For red channels, As a green channel, For the blue channel; transmittance is estimated by combining the weighted dark channel: In the formula, Transmittance, To maintain a small amount of fog-like adjustment factor, For ambient background light; guided filtering is used to adjust transmittance. The transmittance is refined to obtain the refined transmittance. Based on the refined transmittance To restore scene irradiance, a color compensation term is introduced into the red channel, which is expressed as: In the formula, The scene irradiance after red channel compensation. The initial scene irradiance for the red channel recovery. These are empirical parameters.
3. The underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching according to claim 2, characterized in that, For channel weights Based on the water color and local color distribution, the channel weights are adaptively reduced. The red channel weight in the equation is expressed as: In the formula, For pixels In the weight of the red channel, 、 This is an empirical coefficient. For pixels Local eigenvalues.
4. The underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching according to claim 1, characterized in that, In step S3, a pixel-level joint image quality evaluation index is constructed based on the local texture intensity, local contrast, and residual turbidity of the enhanced underwater binocular image, which is expressed as follows: In the formula, As a joint evaluation index for image quality, To normalize local texture intensity, To normalize local contrast, To normalize residual turbidity, 、 、 These are the weighting coefficients.
5. The underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching according to claim 1, characterized in that, S4 specifically includes: configuring two sets of matching parameters, one set of which uses a small window and weak smoothing constraints to obtain a disparity map that is biased towards detail preservation. Another set of matching parameters uses a large window and strong smoothing constraints to obtain a more robust disparity map. Based on joint evaluation index of image quality disparity map Parallax diagram Pixel-level fusion is represented as follows: In the formula, To fuse disparity maps, For pixels Parallax maps that prioritize detail preservation. For pixels Parallax maps with a bias towards robustness.
6. The underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching according to claim 1, characterized in that, S5 specifically includes: for the fused disparity map, calculating a confidence index based on the cost interval according to the minimum and second minimum matching costs of each pixel, and then combining it with left-right consistency checks to obtain the final disparity confidence map. Based on joint evaluation index of image quality With disparity confidence map The comprehensive weights are constructed as follows: In the formula, This is the overall weight; and based on this overall weight... For fused disparity maps Adaptive filtering and interpolation are performed, and the equivalent focal length and baseline length of the refraction equivalent pinhole model are combined to fuse the disparity map. By converting triangulation into depth representation and combining it with the calibration parameters in step S1, the back projection of pixels to the world coordinate system is completed, thereby obtaining the three-dimensional point cloud of the underwater scene.
7. The underwater binocular 3D reconstruction method based on refractive geometry modeling and adaptive matching according to claim 6, characterized in that, Based on comprehensive weighting For fused disparity maps Adaptive filtering and interpolation processing are performed, including: for the comprehensive weights Pixels with weights below the threshold are not used directly for their disparity values. Instead, they are corrected by plane fitting or weighted interpolation using high-weight pixels in the local neighborhood. For large water background areas with weights consistently below the threshold, they are directly removed during the 3D reconstruction stage.
Citation Information
Patent Citations
Monocular and binocular cooperative positioning and mapping method and device for underwater refraction compensation
CN121033151A
Underwater multi-medium refraction parameter calibration method based on binocular camera
CN121259101A