Three-dimensional reconstruction method for underwater concrete based on three-dimensional point cloud domain binocular stereo vision

By employing a binocular stereo vision method based on a 3D point cloud domain, the problems of identifying fine surface features and noise interference in underwater environments were solved, achieving high-precision 3D reconstruction of underwater concrete and ensuring the geometric consistency and detail preservation of the structure.

CN122134938APending Publication Date: 2026-06-02SOUTHWEAT UNIV OF SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHWEAT UNIV OF SCI & TECH
Filing Date
2026-03-05
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively distinguish between static surface features and dynamic suspended noise in underwater environments, leading to issues such as filter misclassification of surface features, difficulty in identifying dynamic noise in a single spatial dimension, point cloud pseudo-tremors caused by water flow disturbances, and geometric distortion in long sequence stitching.

Method used

A binocular stereo vision method based on the three-dimensional point cloud domain is adopted. Through frame-by-frame dense reconstruction, spatiotemporal neighborhood model, dual threshold adaptive filtering and image texture gradient filtering, combined with temporal weighted average fusion and global ICP optimization, a high-precision underwater concrete three-dimensional point cloud is generated.

Benefits of technology

It effectively preserves the surface details of underwater concrete structures, improves the integrity and accuracy of 3D reconstruction, suppresses noise interference, enhances the signal-to-noise ratio and smoothness of point clouds, and ensures the geometric consistency of long-distance structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134938A_ABST
    Figure CN122134938A_ABST
Patent Text Reader

Abstract

This invention discloses a binocular stereo vision-based method for 3D reconstruction of underwater concrete based on a 3D point cloud domain. The method includes the following steps: acquiring underwater images; processing the underwater images using a frame-by-frame dense reconstruction method to generate point clouds for all frames; unifying all frame point clouds to a global coordinate system and establishing a spatiotemporal neighborhood model for each point, analyzing its motion characteristics and neighborhood co-operation, obtaining the position change rate and neighborhood motion consistency coefficient for each point, and optimizing the points in all frame point clouds using a dual-threshold adaptive filtering and image texture gradient filtering method; performing temporal weighted average fusion and global optimization processing to generate a 3D point cloud of underwater concrete. This invention analyzes the motion characteristics of points in 3D space and uses spatiotemporal consistency filtering for noise reduction. The advantage of this scheme is that it directly processes 3D structural data, preserves the integrity of spatial geometric information, is less likely to lose fine features such as concrete cracks, and achieves high reconstruction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of binocular stereo vision underwater concrete 3D reconstruction technology, specifically involving a binocular stereo vision underwater concrete 3D reconstruction method based on a 3D point cloud domain. Background Technology

[0002] Binocular stereo vision technology, as a non-contact, high-precision measurement method, is widely used in the 3D reconstruction and surface defect detection of underwater concrete structures such as dams and bridge piers. However, the complex underwater environment, strong scattering effects, and interference from massive suspended particles pose severe challenges to 3D reconstruction, and different application requirements face drastically different technical bottlenecks.

[0003] For scenarios that emphasize high-precision quantitative detection, existing technologies struggle to effectively distinguish between "static surface micro-features" and "dynamic floating noise" in three-dimensional space. Furthermore, the cumulative error from frame-by-frame registration during long-sequence scans leads to significant geometric distortions in the model. Existing technologies suffer from the following drawbacks: (1) Traditional filtering "falsely kills" fine surface features: Existing technologies (such as single statistical filtering) mainly rely on spatial distance distribution for noise reduction, which makes it difficult to distinguish between "suspended particle noise" and "fine texture or defects on the concrete surface" (both of which are spatially discrete points). In order to remove noise, it is often necessary to increase the filtering intensity, which leads to the removal of key surface details.

[0004] (2) It is difficult to accurately identify dynamic noise in a single spatial dimension: Traditional denoising methods lack verification in the time dimension and cannot effectively distinguish between "spatially isolated but temporally stable" structural details and "spatially isolated and temporally discrete" suspended objects, resulting in a high denoising misjudgment rate.

[0005] (3) Point cloud “pseudo-shock” caused by water flow disturbance: Due to the impact of water flow and light refraction, static concrete structures will produce slight positional jitter in continuous frames. Existing technology directly superimposes the points, which will result in rough point cloud surface and reduced accuracy, forming “pseudo-shock” noise.

[0006] (4) Geometric distortion in long sequence stitching: In large-scale scene reconstruction, the small errors generated by frame-by-frame registration will continue to accumulate, causing the final global model to bend or twist and deform, and cannot truly reflect the macroscopic shape of long-distance structures. Summary of the Invention

[0007] To address the aforementioned shortcomings in existing technologies, the binocular stereo vision-based underwater concrete 3D reconstruction method based on 3D point cloud domain provided by this invention solves the following problems existing in the prior art: (1) The problem of traditional filtering "falsely killing" fine surface features; (2) The problem that dynamic noise is difficult to accurately identify in a single spatial dimension; (3) The problem of point cloud “pseudo-trembling” caused by water flow disturbance; (4) Geometric distortion problem in long sequence splicing.

[0008] To achieve the aforementioned objectives, the technical solution adopted by this invention is: a binocular stereo vision-based three-dimensional reconstruction method for underwater concrete based on a three-dimensional point cloud domain, comprising the following steps: S1. Acquire underwater images and process them using a frame-by-frame dense reconstruction method to generate point clouds for all frames; S2. Unify all frame point clouds to the global coordinate system, and establish a spatiotemporal neighborhood model for each point. Analyze its motion characteristics and neighborhood cooperation to obtain the position change rate and neighborhood motion consistency coefficient of each point. S3. Based on the position change rate and neighborhood motion consistency coefficient of each point, optimize the points in all frame point clouds through dual threshold adaptive filtering and image texture gradient filtering methods. S4. Based on the points in all optimized frame point clouds, perform time-weighted average fusion and global optimization to generate a three-dimensional point cloud of underwater concrete.

[0009] Furthermore, S1 includes the following sub-steps: S11. Acquire underwater images and process them using an improved SGBM stereo matching algorithm to generate a dense disparity map; S12. Using the intrinsic parameters, extrinsic parameters and baseline distance obtained from the binocular camera calibration, back-project each pixel in the dense parallax map from the two-dimensional image plane to the three-dimensional world coordinate system to generate a single-frame dense point cloud. S13. Perform statistical filtering on the dense point cloud of a single frame to remove isolated outliers and obtain the point cloud of all frames.

[0010] The beneficial effects of the above-mentioned further solutions are as follows: The present invention generates a single-frame dense point cloud by back-projecting the camera imaging model, and makes the three-dimensional world points, the camera optical center, and the image pixels collinear. The mapping relationship between the two-dimensional image and the three-dimensional world is established through perspective transformation. After epipolar correction, the binocular stereo matching is simplified from a two-dimensional search to a one-dimensional horizontal search, which greatly improves the matching efficiency and accuracy.

[0011] Furthermore: In S12, each pixel in the dense parallax map is back-projected from the two-dimensional image plane to its coordinates in the three-dimensional world coordinate system. The formula is: In the formula, The x-axis coordinate is... The vertical axis is the coordinate. The vertical axis coordinates are... The baseline distance of the binocular cameras. These are the pixel coordinates of the image plane. The horizontal coordinate of the pixel. The vertical coordinate of the pixel. Let these be the coordinates of the camera's principal point. Let be the horizontal coordinates of the camera's principal point. Let be the longitudinal coordinates of the camera's principal point. The effective focal length of the camera along the horizontal axis. , is the pixel scale factor along the horizontal axis. is the pixel scale factor along the vertical axis. The physical focal length of the camera. The disparity value of a pixel; In S13, the statistical filtering process is specifically as follows: Calculate the distance distribution between each point and its neighbors, where the target point to be detected is calculated. To its k The average Euclidean distance of the neighborhood points The specific expression is: In the formula, for The Neighboring points, It is an L2 norm; Calculate the standard deviation of the distance distribution for all points to obtain the average standard deviation of the distance for all points. Then, calculate the target point to be detected. Standard deviation of distance distribution The specific expression is; All points are divided into normally distributed points and outliers. The specific method for determining whether a target point to be detected is an outlier is as follows: Determine whether the target point to be detected satisfies the outlier determination formula. If it does, the target point to be detected is considered an outlier deviating from the normal distribution; otherwise, the target point to be detected is considered a point in the normal distribution. The outlier determination formula is as follows: In the formula, The mean distance of all points. The standard deviation of the average distance to all points. To preset the information threshold; Remove all outliers from a single frame of dense point cloud, retain the normally distributed points, and generate point clouds for all frames.

[0012] Furthermore, S2 includes the following sub-steps: S21. Using all frame point clouds as source point clouds, extract SIFT feature points of the left and right eyes corresponding to the point clouds of adjacent frames, obtain corresponding point pairs through feature descriptor matching, then use the RANSAC algorithm to remove erroneous matching pairs, solve the rigid body transformation matrix between adjacent frames, transform the source point cloud to the target point cloud coordinate system, and obtain coarsely registered point clouds. S22. Using the coarsely registered point cloud as the initial value, the ICP algorithm is executed on the point clouds of adjacent frames. By iteratively minimizing the Euclidean distance error between point pairs, the rigid body transformation matrix is ​​optimized, and finally the point clouds of all key frames are unified to the same global coordinate system to generate a registered global point cloud sequence. S23. Establish a spatiotemporal neighborhood for each point based on the global point cloud sequence, calculate the position change rate of each point, and calculate the neighborhood motion consistency coefficient through neighborhood motion consistency verification.

[0013] The beneficial effects of the above-mentioned further scheme are as follows: Based on the optimal estimation of "minimizing the Euclidean distance of corresponding point pairs", the present invention solves the rigid body transformation matrix through singular value decomposition, avoiding the complexity of directly solving nonlinear equations. The iterative process ensures that the registration accuracy is gradually improved, and finally realizes the global coordinate system of all point clouds.

[0014] Furthermore: In S21, the expression for transforming the source point cloud to the target point cloud coordinate system is specifically as follows: In the formula, For the point in the target point cloud coordinate system, For points in the source point cloud, For rotation matrix, It is a translation vector; In S22, the ICP algorithm solves for the optimal rigid body transformation by iteratively minimizing the sum of squared Euclidean distances between corresponding point pairs in the source and target point clouds. The specific expression for the objective function of the ICP algorithm is: In the formula, M This represents the number of corresponding point pairs. The optimal rotation matrix is... To be the optimal translation vector, For the target point cloud One corresponding point, The first point cloud in the source cloud One corresponding point, The independent variable refers to the value at which the objective function reaches its minimum. and A set; In S23, calculate Time point Rate of change of position The specific expression is: In the formula, The number of frames within the time window. The time interval between two adjacent frames. for Time point , for Time point, This represents the number of frames between the current frame and the target frame. calculate Time point Neighborhood motion consistency coefficient The specific expression is: In the formula, The standard deviation of the rate of change of the location of neighboring points. For point The average rate of change of position of its spatial neighboring points.

[0015] Furthermore, S3 includes the following sub-steps: S31. Dual-threshold adaptive screening method: Set high motion threshold Low motion threshold High consistency threshold and low consistency threshold ; Response to the rate of change of position at any point Consistency coefficient of motion with neighborhood satisfy: , If the point is not a noise point, it will be removed from all frame point clouds. Response to the rate of change of position at any point Consistency coefficient of motion with neighborhood satisfy: , If so, then that point will be considered a valid point and retained. For the remaining points in all frame point clouds that do not belong to noise points and valid points, they are regarded as points to be determined by auxiliary judgment, and further optimized by the image texture gradient filtering method. S32, Image Texture Gradient Filtering Method: The Sobel operator is used to calculate the texture gradient of the image region corresponding to the point to be assisted in the decision. ; In the formula, for Sobel gradient in the axial direction, for Sobel gradient along the axial direction; In the formula, The grayscale value of the pixel in the image region corresponding to the point to be assisted in the determination is... This is a convolution operation; Calculate the neighborhood average texture gradient of the image region corresponding to the point to be assisted in the decision. ; In the formula, The total number of pixels in the neighborhood. The image neighborhood corresponding to the blurred point; If the average texture gradient of the neighborhood of the image region corresponding to the point to be assisted in determination is greater than a preset texture threshold, then the point is regarded as a valid point and retained.

[0016] The beneficial effects of the above-mentioned further solutions are as follows: This invention divides point clouds into three categories through dual thresholds, achieving rapid removal of most noise and retention of effective points; the adaptability of the thresholds is achieved through a Gaussian mixture model, fitting the rate of change of point positions to two Gaussian distributions, and automatically solving for the optimal threshold. and This improves the algorithm's adaptability to different underwater environments.

[0017] Furthermore, S4 includes the following sub-steps: S41. Based on the points in all optimized frame point clouds, the neighborhood motion consistency coefficient is used as the basis for weight allocation, and time-weighted average fusion is performed according to the time window to generate the fused point cloud. S42. Perform global ICP optimization based on the fused point cloud to minimize the total registration error between all point clouds, correct geometric distortion, and generate a three-dimensional point cloud of underwater concrete.

[0018] Furthermore: In S41, generation is achieved through time-weighted average fusion. Optimization points at specific times The specific method is as follows: In the formula, The time window is half the width. for Time point The neighborhood motion consistency coefficient, for Time point ; In S42, global ICP optimization aims to minimize the total registration error between the average point cloud of all frames and the average point cloud. The specific expression of the global ICP optimization objective function is as follows: In the formula, This represents the total number of keyframes. For the first The number of points in a frame point cloud. The average point cloud for all frames. For the first The rotation matrix corresponding to the frame point cloud. For the first The translation vector corresponding to the frame point cloud. For the first The first frame in the point cloud One point, For the first The optimal rotation matrix corresponding to the frame point cloud. For the first The optimal translation vector corresponding to the frame point cloud.

[0019] The beneficial effects of the above-mentioned further solutions are as follows: Based on the assumption that "points with stable motion have higher credibility", this invention assigns weights through motion consistency coefficients, allowing stable points to dominate the fusion process, weakening the influence of jittery points, thereby suppressing residual noise and improving the accuracy and smoothness of point clouds; weighted average fusion is a linear fusion method, which is simple to calculate, has good real-time performance, and can preserve the spatiotemporal continuity of point clouds.

[0020] The beneficial effects of this invention are as follows: (1) This invention introduces a dual-threshold adaptive filtering and image texture gradient filtering and judgment mechanism. For points whose motion characteristics are in the fuzzy area, the algorithm automatically backtracks to the two-dimensional image domain to calculate the texture gradient. It can accurately identify the surface detail edge points with high texture features and forcibly retain them. This mechanism effectively avoids the accidental removal of key surface information, significantly improves the integrity and recognizability of the three-dimensional reconstruction of the surface details of underwater concrete structures, solves the problem of traditional filtering "accidentally killing" surface micro features, and achieves a balance between noise reduction and detail preservation.

[0021] (2) This invention establishes a spatiotemporal neighborhood model and uses the rate of change of position of a point in consecutive frames and the motion consistency coefficient for discrimination. This method utilizes the laws of physical motion and can effectively distinguish between "suspended particles moving at high speed with water flow" (high rate of change, low consistency) and "relatively stationary or co-moving concrete structures" (low rate of change, high consistency), fundamentally improving the accuracy of noise identification, solving the problem that it is difficult to accurately identify dynamic noise in a single spatial dimension, resulting in a high rate of false noise removal, and improving the accuracy of noise removal.

[0022] (3) This invention proposes a time-weighted average fusion algorithm that dynamically allocates weights using a motion consistency coefficient. The algorithm automatically reduces the weight of jittering points and increases the weight of stable points, thereby effectively suppressing high-frequency jittering noise caused by the water flow environment. This makes the reconstructed concrete surface smoother and more continuous, significantly improves the signal-to-noise ratio of the point cloud, solves the problem of "pseudo-juddering" of the point cloud caused by water flow disturbance, and improves the smoothness of the reconstructed surface.

[0023] (4) Based on frame-by-frame processing, this invention adds a global ICP optimization step. By jointly optimizing all frame point clouds as a whole and minimizing the global registration error, this scheme effectively corrects the geometric distortion caused by accumulated errors. This ensures the macroscopic geometric consistency of the reconstruction model of long-distance underwater structures (such as dam walls), meets the needs of high-precision engineering measurement, solves the geometric distortion problem in long sequence stitching, and guarantees the geometric accuracy of the global model.

[0024] In summary, this invention analyzes the motion characteristics of points in three-dimensional space after stereo matching and uses spatiotemporal consistency filtering for noise reduction. The advantage of this method is that it directly processes three-dimensional structural data, preserves the integrity of spatial geometric information, is less prone to losing fine features such as concrete cracks, and achieves high reconstruction accuracy. Attached Figure Description

[0025] Figure 1 This is a flowchart of the binocular stereo vision method for three-dimensional reconstruction of underwater concrete based on a three-dimensional point cloud domain, according to the present invention. Detailed Implementation

[0026] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0027] like Figure 1As shown, in one embodiment of the present invention, a binocular stereo vision method for three-dimensional reconstruction of underwater concrete based on a three-dimensional point cloud domain includes the following steps: S1. Acquire underwater images and process them using a frame-by-frame dense reconstruction method to generate point clouds for all frames; S2. Unify all frame point clouds to the global coordinate system, and establish a spatiotemporal neighborhood model for each point. Analyze its motion characteristics and neighborhood cooperation to obtain the position change rate and neighborhood motion consistency coefficient of each point. S3. Based on the position change rate and neighborhood motion consistency coefficient of each point, optimize the points in all frame point clouds through dual threshold adaptive filtering and image texture gradient filtering methods. S4. Based on the points in all optimized frame point clouds, perform time-weighted average fusion and global optimization to generate a three-dimensional point cloud of underwater concrete.

[0028] S1 includes the following steps: S11. Acquire underwater images and process them using an improved SGBM stereo matching algorithm to generate a dense disparity map; S12. Using the intrinsic parameters, extrinsic parameters and baseline distance obtained from the binocular camera calibration, back-project each pixel in the dense parallax map from the two-dimensional image plane to the three-dimensional world coordinate system to generate a single-frame dense point cloud, and add a timestamp to each point to mark the keyframe to which it belongs. S13. Perform statistical filtering on the dense point cloud of a single frame to remove isolated outliers and obtain the point cloud of all frames.

[0029] In S11, to address the issues of low contrast and high scattering noise in underwater images, the traditional semi-global block matching (SGBM) algorithm is improved (e.g., by adding underwater image preprocessing, adaptive cost calculation, and optimized path aggregation weights). An improved SGBM stereo matching algorithm is designed to perform stereo matching on the left and right eyes of each key frame, obtain the disparity value corresponding to each pixel, and form a dense disparity map.

[0030] The core of the improved SGBM stereo matching algorithm consists of four steps: cost calculation, path aggregation, disparity selection, and disparity optimization. Cost calculation employs the Census transform (which is highly resistant to illumination variations and suitable for underwater images), encoding the neighborhood information of each pixel into a binary string, and calculating the matching cost between left and right eye pixels using Hamming distance. Path aggregation uses multi-directional (typically 8-directional) dynamic programming to minimize the global matching cost function and avoid local optima. Finally, uniqueness constraints and sub-pixel interpolation are used to optimize the disparity map, improving density and accuracy.

[0031] In S12, each pixel in the dense parallax map is back-projected from the two-dimensional image plane to its point coordinates in the three-dimensional world coordinate system. The formula is: In the formula, The x-axis coordinate is... The vertical axis is the coordinate. The vertical axis coordinates are... The baseline distance between the binocular cameras (the straight-line distance between the optical centers of the left and right cameras, obtained through calibration). These are the pixel coordinates of the image plane (obtained by calibrating the origin of the image coordinate system, usually the image center). The horizontal coordinate of the pixel. The vertical coordinate of the pixel. Let these be the coordinates of the camera's principal point. Let be the horizontal coordinates of the camera's principal point. Let be the longitudinal coordinates of the camera's principal point. The effective focal length of the camera along the horizontal axis. , is the pixel scale factor along the horizontal axis. is the pixel scale factor along the vertical axis. The physical focal length of the camera. This represents the disparity value of a pixel (after epipolar correction, disparity exists only in the horizontal direction, and the vertical disparity is 0). , The x-coordinate of the pixel in the left eye image. The x-coordinate of the pixel in the right eye image; In this embodiment, the present invention generates a single-frame dense point cloud by back-projecting the camera imaging model, and establishes a mapping relationship between the two-dimensional image and the three-dimensional world by collinearizing the three-dimensional world points, the camera optical center, and the image pixels through perspective transformation; after epipolar correction, the binocular stereo matching is simplified from a two-dimensional search to a one-dimensional horizontal search, which greatly improves the matching efficiency and accuracy.

[0032] In S13, the core of statistical filtering is to calculate the distance distribution between each point and its neighboring points, and to remove outliers that deviate from the normal distribution. The specific method of statistical filtering is as follows: Calculate the distance distribution between each point and its neighbors, where the target point to be detected is calculated. To its k The average Euclidean distance of the neighborhood points The specific expression is: In the formula, for The A number of neighboring points (usually using k-nearest neighbor search, (Preset number of neighboring points). It is an L2 norm; Calculate the standard deviation of the distance distribution for all points to obtain the average standard deviation of the distance for all points. Then, calculate the target point to be detected. Standard deviation of distance distribution The specific expression is; All points are divided into normally distributed points and outliers. The specific method for determining whether a target point to be detected is an outlier is as follows: Determine whether the target point to be detected satisfies the outlier determination formula. If it does, the target point to be detected is considered an outlier deviating from the normal distribution; otherwise, the target point to be detected is considered a point in the normal distribution. The outlier determination formula is as follows: In the formula, The mean distance of all points. The standard deviation of the average distance to all points. Set a preset threshold (usually 2-3, to control the strictness of filtering); Remove all outliers from a single frame of dense point cloud, retain the normally distributed points, and generate point clouds for all frames.

[0033] In this embodiment, since the generated single-frame dense point cloud contains mixed data of concrete structure main body (static, low motion) and underwater suspended particles (dynamic, high motion), the present invention uses statistical filtering to process the single-frame dense point cloud, selects points that conform to normal distribution, removes isolated points that are obviously deviated from the main body, and does not destroy the overall structure of the target point cloud, thereby reducing data redundancy and interference for subsequent time series analysis.

[0034] S2 includes the following steps: S21. Using all frame point clouds as source point clouds, extract SIFT feature points of the left and right eyes corresponding to the point clouds of adjacent frames, obtain corresponding point pairs through feature descriptor matching, then use the RANSAC algorithm to remove erroneous matching pairs, solve the rigid body transformation matrix between adjacent frames, eliminate large-scale pose deviations, transform the source point cloud to the target point cloud coordinate system, and obtain coarsely registered point clouds. S22. Using the coarsely registered point cloud as the initial value, perform the ICP algorithm on the point clouds of adjacent frames. By iteratively minimizing the Euclidean distance error between point pairs, optimize the rigid body transformation matrix, and finally unify the point clouds of all key frames to the same global coordinate system, forming a spatiotemporally ordered point cloud sequence. , For the first The point coordinates of the frame, Given the total number of keyframes, generate the registered global point cloud sequence; S23. Establish a spatiotemporal neighborhood for each point based on the global point cloud sequence, calculate the position change rate of each point, and calculate the neighborhood motion consistency coefficient through neighborhood motion consistency verification.

[0035] In S21, the expression for transforming the source point cloud to the target point cloud coordinate system is as follows: In the formula, For the point in the target point cloud coordinate system, For points in the source point cloud, For rotation matrix, It is a translation vector; In this embodiment, SIFT (Scale Invariant Feature Transform) possesses scale, rotation, and illumination invariance, making it suitable for feature matching in underwater images. This invention achieves corresponding point matching by generating 128-dimensional feature descriptors and calculating the Euclidean distance between descriptors; the RANSAC algorithm improves the robustness of rigid body transformation solutions by eliminating erroneous matching pairs through random sampling consistency.

[0036] In S22, the ICP algorithm solves for the optimal rigid body transformation by iteratively minimizing the sum of squared Euclidean distances between corresponding point pairs in the source and target point clouds. The specific expression for the objective function of the ICP algorithm is: In the formula, M This represents the number of corresponding point pairs. The optimal rotation matrix is... To be the optimal translation vector, For the target point cloud One corresponding point, The first point cloud in the source cloud One corresponding point, The independent variable refers to the value at which the objective function reaches its minimum. and A set; In this embodiment, the present invention is based on the optimal estimation of "minimizing the Euclidean distance between corresponding point pairs", and solves the rigid body transformation matrix through singular value decomposition (SVD) to avoid the complexity of directly solving nonlinear equations. The iterative process ensures that the registration accuracy is gradually improved, and finally realizes the global coordinate system of all point clouds.

[0037] Preferably, the iterative steps of the ICP algorithm are as follows: ① Nearest neighbor search (find the nearest neighbor in the target point cloud for each source point to form corresponding point pairs); ② Solve the optimal rigid body transformation (solve the above minimization problem through singular value decomposition SVD); ③ Transform the source point cloud; ④ Determine convergence. If the distance error is less than the preset threshold or the maximum number of iterations is reached, stop the iteration; otherwise, return to step ①. In S23, calculate Time point Rate of change of position The specific expression is: In the formula, The number of frames within the time window. The time interval between two adjacent frames. for Time point , for Time point, This represents the number of frames between the current frame and the target frame. A higher value indicates more intense exercise; calculate Time point Neighborhood motion consistency coefficient The specific expression is: In the formula, The standard deviation of the rate of change of the location of neighboring points. For point and the average rate of change of position of its spatial neighborhood points, The closer the value is to 1, the more coordinated the movement of the point with its neighboring points, and the more likely it is to be a static point in the concrete structure; the closer the value is to 0, the more discrete the movement, and the more likely it is to be noise from suspended particles.

[0038] In this embodiment, for the registered global point cloud sequence, a spatiotemporal neighborhood is established for each point: within a preset spatial radius. Inside, search continuously Frame (time window) The corresponding point set in ) A time-series correlation graph is constructed; the rate of change of position of each point is calculated, and a neighborhood motion consistency check is introduced to avoid misjudging small vibrations on the concrete surface (such as slight displacements caused by water flow disturbance) as noise.

[0039] The principle of the spatiotemporal neighborhood model is to integrate "spatial proximity" and "temporal continuity," assuming that "static points of the concrete structure are spatially adjacent and temporally continuous, exhibiting consistent motion characteristics; while suspended particulate noise is spatially isolated and temporally discrete, exhibiting random motion characteristics." Spatial radius constraints reduce interference from irrelevant points, and temporal window constraints capture the temporal motion trajectories of points, providing reliable motion characteristic data for subsequent denoising.

[0040] S3 includes the following steps: S31. Dual-threshold adaptive screening method: Set high motion threshold Low motion threshold High consistency threshold and low consistency threshold Among them, high motion threshold and low motion threshold can be solved by adaptive algorithms, such as Gaussian Mixture Model (GMM) clustering based on point cloud motion characteristics, which automatically divides the threshold intervals of static and dynamic points, and achieves high consistency threshold. Typically, a value of 0.3 is used as the low consistency threshold. The value is usually 0.7, but can be adjusted according to the underwater environment.

[0041] Response to the rate of change of position at any point Consistency coefficient of motion with neighborhood satisfy: , If the point is considered a noise point, such as a high-speed underwater suspended particulate matter, it will be removed from all frame point clouds. Response to the rate of change of position at any point Consistency coefficient of motion with neighborhood satisfy: , If so, then that point will be considered a valid point in the concrete structure and retained. For the remaining points in all frame point clouds that do not belong to noise points and valid points, they are regarded as points to be determined by auxiliary judgment, and further optimized by the image texture gradient filtering method. In this embodiment, based on the prior knowledge that "in underwater scenarios, concrete structure points move at low speeds in a coordinated manner, while suspended particles move at high speeds in a discrete manner," the point cloud is divided into three categories using a dual threshold, achieving rapid removal of most noise and retention of effective points. The adaptiveness of the threshold is achieved through a Gaussian mixture model (GMM), fitting the rate of change of point positions to two Gaussian distributions (corresponding to static points and dynamic noise points, respectively), and automatically solving for the optimal threshold. and This improves the algorithm's adaptability to different underwater environments.

[0042] S32, Image Texture Gradient Filtering Method: The Sobel operator is used to calculate the texture gradient of the image region corresponding to the point to be assisted in the decision. ; In the formula, for Sobel gradient in the axial direction, for Sobel gradient along the axial direction; In the formula, The grayscale value of the pixel in the image region corresponding to the point to be assisted in the determination is... This is a convolution operation; In this embodiment, areas with defects such as concrete cracks and edges appear as high texture gradients in two-dimensional images, while the image areas corresponding to low-velocity suspended particles are usually flat. By tracing back the texture information of the two-dimensional image, the deficiencies of the three-dimensional point cloud motion characteristic analysis are made up for, avoiding the loss of key defect information and achieving the goal of "denoising without losing details". The Sobel operator calculates the image gradient through neighborhood convolution, which has the characteristics of simple calculation and strong noise resistance, and is suitable for texture extraction of low-contrast underwater images.

[0043] Calculate the neighborhood average texture gradient of the image region corresponding to the point to be assisted in the decision. ; In the formula, The total number of pixels in the neighborhood. The image neighborhood corresponding to the blurred point; If the average texture gradient of the neighborhood of the image region corresponding to the point to be assisted in determination is greater than a preset texture threshold, then the point is regarded as a valid point and retained.

[0044] The principle of the image texture gradient screening method is as follows: calculate the image texture gradient. If the texture gradient is significant, it indicates that it may be located in a region rich in detail such as cracks or structural edges, and should be retained first. If the texture is flat, it indicates that it may be low-speed suspended particles or image noise, and should be removed.

[0045] S4 includes the following steps: S41. Based on the points in all optimized frame point clouds, the neighborhood motion consistency coefficient is used as the basis for weight allocation, and time-weighted average fusion is performed according to the time window to generate the fused point cloud. S42. Perform global ICP optimization based on the fused point cloud to minimize the total registration error between all point clouds, correct geometric distortion, and generate a high-precision, high-purity underwater concrete 3D point cloud.

[0046] In S41, the time-weighted average fusion is used to generate... Optimization points at specific times The specific method is as follows: In the formula, The time window is half the width. for Time point The neighborhood motion consistency coefficient, for Time point In this embodiment, the motion consistency coefficient is used. As a basis for weight allocation, the more stable the motion ( Points closer to 1) receive higher weights, thereby suppressing residual jitter noise (such as small displacements of point clouds caused by water flow disturbances) and improving the smoothness and stability of point clouds.

[0047] In this embodiment, based on the assumption that "points with stable motion have higher credibility", weights are assigned through motion consistency coefficients to allow stable points to dominate the fusion process, weaken the influence of jittery points, thereby suppressing residual noise and improving the accuracy and smoothness of the point cloud. Weighted average fusion is a linear fusion method that is simple to calculate, has good real-time performance, and can preserve the spatiotemporal continuity of the point cloud.

[0048] In S42, unlike frame-by-frame ICP, global ICP optimization aims to minimize the total registration error between the point clouds of all frames and the average point cloud. The specific expression of the objective function for global ICP optimization is as follows: In the formula, This represents the total number of keyframes. For the first The number of points in a frame point cloud. The average point cloud for all frames. For the first The rotation matrix corresponding to the frame point cloud. For the first The translation vector corresponding to the frame point cloud. For the first The first frame in the point cloud One point, For the first The optimal rotation matrix corresponding to the frame point cloud. For the first The optimal translation vector corresponding to the frame point cloud.

[0049] Frame-by-frame ICP registration generates "cumulative error" (the registration error of each frame is propagated to subsequent frames, causing global point cloud distortion). Global ICP optimizes all frame point clouds as a whole, using the average point cloud as a benchmark, minimizing the total global error, thereby correcting the cumulative error and improving the global consistency and geometric accuracy of the point cloud. The core of global ICP is "global optimization", which is different from the "local optimization" of frame-by-frame ICP, and is suitable for global optimization of large-scale temporal point clouds.

[0050] In the description of this invention, the above are merely preferred embodiments and are not intended to limit the scope of protection of this invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A binocular stereo vision-based method for 3D reconstruction of underwater concrete based on a 3D point cloud domain, characterized in that, Includes the following steps: S1. Acquire underwater images and process them using a frame-by-frame dense reconstruction method to generate point clouds for all frames; S2. Unify all frame point clouds to the global coordinate system, and establish a spatiotemporal neighborhood model for each point. Analyze its motion characteristics and neighborhood cooperation to obtain the position change rate and neighborhood motion consistency coefficient of each point. S3. Based on the position change rate and neighborhood motion consistency coefficient of each point, optimize the points in all frame point clouds through dual threshold adaptive filtering and image texture gradient filtering methods. S4. Based on the points in all optimized frame point clouds, perform time-weighted average fusion and global optimization to generate a three-dimensional point cloud of underwater concrete.

2. The method for three-dimensional reconstruction of underwater concrete based on binocular stereo vision using a three-dimensional point cloud domain according to claim 1, characterized in that, S1 includes the following steps: S11. Acquire underwater images and process them using an improved SGBM stereo matching algorithm to generate a dense disparity map; S12. Using the intrinsic parameters, extrinsic parameters and baseline distance obtained from the binocular camera calibration, back-project each pixel in the dense parallax map from the two-dimensional image plane to the three-dimensional world coordinate system to generate a single-frame dense point cloud. S13. Perform statistical filtering on the dense point cloud of a single frame to remove isolated outliers and obtain the point cloud of all frames.

3. The method for three-dimensional reconstruction of underwater concrete based on binocular stereo vision using a three-dimensional point cloud domain according to claim 2, characterized in that, In S12, each pixel in the dense parallax map is back-projected from the two-dimensional image plane to its point coordinates in the three-dimensional world coordinate system. The formula is: In the formula, The x-axis coordinate is... The vertical axis is the coordinate. The vertical axis coordinates are... The baseline distance of the binocular cameras. These are the pixel coordinates of the image plane. The horizontal coordinate of the pixel. The vertical coordinate of the pixel. Let these be the coordinates of the camera's principal point. Let be the horizontal coordinates of the camera's principal point. Let be the longitudinal coordinates of the camera's principal point. The effective focal length of the camera along the horizontal axis. , is the pixel scale factor along the horizontal axis. is the pixel scale factor along the vertical axis. The physical focal length of the camera. The disparity value of a pixel; In S13, the statistical filtering process is specifically as follows: Calculate the distance distribution between each point and its neighbors, where the target point to be detected is calculated. To its k The average Euclidean distance of the neighborhood points The specific expression is: In the formula, for The Neighboring points, It is an L2 norm; Calculate the standard deviation of the distance distribution for all points to obtain the average standard deviation of the distance for all points. Then, calculate the target point to be detected. Standard deviation of distance distribution The specific expression is; All points are divided into normally distributed points and outliers. The specific method for determining whether a target point to be detected is an outlier is as follows: Determine whether the target point to be detected satisfies the outlier determination formula. If it does, the target point to be detected is considered an outlier deviating from the normal distribution; otherwise, the target point to be detected is considered a point in the normal distribution. The outlier determination formula is as follows: In the formula, The mean distance of all points. The standard deviation of the average distance to all points. To preset the information threshold; Remove all outliers from a single frame of dense point cloud, retain the normally distributed points, and generate point clouds for all frames.

4. The method for three-dimensional reconstruction of underwater concrete based on binocular stereo vision using a three-dimensional point cloud domain according to claim 1, characterized in that, S2 includes the following steps: S21. Using all frame point clouds as source point clouds, extract SIFT feature points of the left and right eyes corresponding to the point clouds of adjacent frames, obtain corresponding point pairs through feature descriptor matching, then use the RANSAC algorithm to remove erroneous matching pairs, solve the rigid body transformation matrix between adjacent frames, transform the source point cloud to the target point cloud coordinate system, and obtain coarsely registered point clouds. S22. Using the coarsely registered point cloud as the initial value, the ICP algorithm is executed on the point clouds of adjacent frames. By iteratively minimizing the Euclidean distance error between point pairs, the rigid body transformation matrix is ​​optimized, and finally the point clouds of all key frames are unified to the same global coordinate system to generate a registered global point cloud sequence. S23. Establish a spatiotemporal neighborhood for each point based on the global point cloud sequence, calculate the position change rate of each point, and calculate the neighborhood motion consistency coefficient through neighborhood motion consistency verification.

5. The method for three-dimensional reconstruction of underwater concrete based on binocular stereo vision using a three-dimensional point cloud domain according to claim 4, characterized in that, In S21, the expression for transforming the source point cloud to the target point cloud coordinate system is as follows: In the formula, For the point in the target point cloud coordinate system, For points in the source point cloud, Let be a rotation matrix. It is a translation vector; In S22, the ICP algorithm solves for the optimal rigid body transformation by iteratively minimizing the sum of squared Euclidean distances between corresponding point pairs in the source and target point clouds. The specific expression for the objective function of the ICP algorithm is: In the formula, M This represents the number of corresponding point pairs. To be the optimal rotation matrix, To be the optimal translation vector, For the target point cloud One corresponding point, The first point cloud in the source cloud One corresponding point, The independent variable refers to the value at which the objective function reaches its minimum. and A set; In S23, calculate Time point Rate of change of position The specific expression is: In the formula, The number of frames within the time window. The time interval between two adjacent frames. for Time point , for Time point, This represents the number of frames between the current frame and the target frame. calculate Time point Neighborhood motion consistency coefficient The specific expression is: In the formula, The standard deviation of the rate of change of the location of neighboring points. For point The average rate of change of position of its spatial neighboring points.

6. The method for three-dimensional reconstruction of underwater concrete based on binocular stereo vision using a three-dimensional point cloud domain according to claim 5, characterized in that, S3 includes the following steps: S31. Dual-threshold adaptive screening method: Set high motion threshold Low motion threshold High consistency threshold and low consistency threshold ; Response to the rate of change of position at any point Consistency coefficient of motion with neighborhood satisfy: , If the point is not a noise point, it will be removed from all frame point clouds. Response to the rate of change of position at any point Consistency coefficient of motion with neighborhood satisfy: , If so, then that point will be considered a valid point and retained. For the remaining points in all frame point clouds that do not belong to noise points and valid points, they are regarded as points to be determined by auxiliary judgment, and further optimized by the image texture gradient filtering method. S32, Image Texture Gradient Filtering Method: The Sobel operator is used to calculate the texture gradient of the image region corresponding to the point to be assisted in the decision. ; In the formula, for Sobel gradient in the axial direction, for Sobel gradient along the axial direction; In the formula, The grayscale value of the pixel in the image region corresponding to the point to be assisted in the determination is... This is a convolution operation; Calculate the neighborhood average texture gradient of the image region corresponding to the point to be assisted in the decision. ; In the formula, The total number of pixels in the neighborhood. The image neighborhood corresponding to the blurred point; If the average texture gradient of the neighborhood of the image region corresponding to the point to be assisted in determination is greater than a preset texture threshold, then the point is regarded as a valid point and retained.

7. The binocular stereo vision method for three-dimensional reconstruction of underwater concrete based on a three-dimensional point cloud domain as described in claim 6, characterized in that, S4 includes the following steps: S41. Based on the points in all optimized frame point clouds, the neighborhood motion consistency coefficient is used as the basis for weight allocation, and time-weighted average fusion is performed according to the time window to generate the fused point cloud. S42. Perform global ICP optimization based on the fused point cloud to minimize the total registration error between all point clouds, correct geometric distortion, and generate a three-dimensional point cloud of underwater concrete.

8. The method for three-dimensional reconstruction of underwater concrete based on binocular stereo vision using a three-dimensional point cloud domain according to claim 7, characterized in that, In S41, the time-weighted average fusion is used to generate... Optimization points at specific times The specific method is as follows: In the formula, The time window is half the width. for Time point The neighborhood motion consistency coefficient, for Time point ; In S42, global ICP optimization aims to minimize the total registration error between the average point cloud of all frames and the average point cloud. The specific expression of the global ICP optimization objective function is as follows: In the formula, This represents the total number of keyframes. For the first The number of points in a frame point cloud. The average point cloud for all frames. For the first The rotation matrix corresponding to the frame point cloud. For the first The translation vector corresponding to the frame point cloud. For the first The first frame in the point cloud One point, For the first The optimal rotation matrix corresponding to the frame point cloud. For the first The optimal translation vector corresponding to the frame point cloud.