Underwater concrete scene three-dimensional reconstruction method, system and device and storage medium

By extracting feature points from enhanced video images, calculating the camera's movement distance and translation vector, and combining pixel depth values ​​for 3D reconstruction, the problem of accuracy and efficiency in 3D reconstruction of underwater concrete structures in turbid water and environments with severe light absorption was solved, achieving high-precision and high-efficiency underwater 3D reconstruction.

CN121120996AActive Publication Date: 2025-12-12TIANFU YONGXING LAB

Patent Information

Application Number
CN202511679316.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2025-12-12
Estimated Expiration
2045-11-17

AI Technical Summary

Technical Problem

In turbid water and environments with severe light absorption, the image quality of 3D reconstruction of underwater concrete structures is poor, resulting in low accuracy and efficiency.

Method used

By extracting feature points from the enhanced video images, calculating the camera's movement distance and translation vector, and combining pixel depth values ​​for 3D reconstruction, a 3D mesh model is generated using epipolar geometry alignment and multi-view fusion.

Benefits of technology

It improves the quality of underwater image acquisition, saves the time required to build 3D point cloud data, and enables large-scale, high-precision, and high-efficiency 3D reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120996A_ABST
    Figure CN121120996A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an underwater concrete scene three-dimensional reconstruction method, system and device and a storage medium, and relates to the technical field of underwater three-dimensional reconstruction, and the method comprises the steps: extracting feature points in each key frame image from a plurality of key frame images of video data, and obtaining feature matching point pairs; calculating relative poses of adjacent key frame images and three-dimensional coordinates of the matched feature points according to the pixel coordinates of the same matched feature point; calculating a pixel depth value of each pixel point according to the relative pose; and performing multi-view fusion according to the relative pose, the three-dimensional coordinates of the matched feature points and the pixel depth value of each pixel point, and generating a three-dimensional grid model of the underwater concrete scene. According to the method, the feature points in the enhanced video image are obtained, the moving distance and the translation vector of the shooting camera are calculated, and three-dimensional reconstruction is carried out in combination with the pixel depth value, so that the image acquisition quality is improved, point cloud data are saved, and the precision and efficiency of three-dimensional reconstruction of a large-range underwater scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of underwater three-dimensional reconstruction, in particular to an underwater concrete scene three-dimensional reconstruction method, system, device and storage medium. BACKGROUND

[0002] The long-term safety and durability of underwater structures of hydropower hubs such as concrete gravity dams and arch dams are the core of protecting the safety of life and property downstream and the stability of the power system. They are long-term in complex water environments, subject to the coupling of multiple factors such as hydraulic load, temperature change, freeze-thaw cycle, and are prone to defects such as cracks, spalling, and steel corrosion.

[0003] In recent years, through the deep integration of underwater robot (ROV / AUV) technology and high-precision optical sensing technology, underwater robots are equipped with high-definition camera equipment, high-brightness lighting systems and laser scanners to carry out close-range, high-stability data collection underwater, and simultaneously obtain high-definition images and three-dimensional point cloud information of the structure surface, and then realize accurate quantitative analysis and detection of concrete defects.

[0004] However, due to the turbidity of the water body and the serious light absorption in the underwater environment, the optical images obtained by shooting will have problems such as color distortion, blurring and glare, which makes the quality of the underwater images collected poor, and it takes a lot of time to construct three-dimensional point cloud data, resulting in low accuracy and efficiency of three-dimensional reconstruction of large-scale underwater scenes. SUMMARY

[0005] Therefore, the present application aims to provide an underwater concrete scene three-dimensional reconstruction method, system, device and storage medium, which obtains feature points from the enhanced video images, calculates the moving distance and translation vector of the shooting camera based on the feature points, and then combines the pixel depth value in each key frame image to perform three-dimensional reconstruction, realizing large scene, high precision and high efficiency underwater three-dimensional reconstruction. In the case of turbidity of the water body and serious light absorption in the underwater environment, the quality of underwater image acquisition can be improved, and the construction of three-dimensional point cloud data can be saved, thereby improving the accuracy and efficiency of three-dimensional reconstruction of large-scale underwater scenes.

[0006] To achieve the above object, in a first aspect, embodiments of the present application provide a method for three-dimensional reconstruction of an underwater concrete scene, the method comprising: obtaining video data of an underwater concrete surface; obtaining a plurality of key frame images from the video data; extracting feature points in each of the key frame images, the feature points representing salient parts of the underwater concrete surface; performing epipolar geometry comparison on the feature points in adjacent key frame images to obtain a feature matching point pair in the adjacent key frame images, the feature matching point pair comprising pixel coordinates of a same matched feature point in the adjacent key frame images; calculating a relative pose of the adjacent key frame images based on the pixel coordinates of the same matched feature point in the adjacent key frame images, the relative pose representing a rotation angle and a translation vector of a camera corresponding to the adjacent key frame images; calculating a three-dimensional coordinate of the matched feature point based on the pixel coordinates of the same matched feature point in the adjacent key frame images and the relative pose; calculating a pixel depth value of each pixel point in the key frame images based on the relative pose; and performing multi-view fusion based on the relative pose, the three-dimensional coordinate of the matched feature point, and the pixel depth value of each pixel point to generate a three-dimensional mesh model of the underwater concrete scene.

[0007] In the present embodiment, a plurality of key frame images are obtained from the video data, the plurality of key frame images representing pictures taken from a plurality of perspectives, and feature points of salient parts of the underwater concrete surface are extracted from each key frame image. Then, a matching feature point pair in adjacent two key frame images is compared, so as to calculate a relative pose based on pixel coordinates of the matching feature point pair in the key frame images, and further obtain a rotation angle and a translation vector of a camera corresponding to the adjacent key frame images. Then, a three-dimensional coordinate of the matched feature point can be calculated based on the relative pose and the pixel coordinates of the feature points, so as to calculate a pixel depth in each key frame image and combine the three-dimensional coordinate of the matched feature point to perform multi-view fusion, thereby generating a three-dimensional mesh model of the underwater concrete scene. In this way, by obtaining feature points from enhanced video images, calculating a moving distance and a translation vector of a camera based on the feature points, and further combining a pixel depth value in each key frame image to perform three-dimensional reconstruction, three-dimensional reconstruction of an underwater scene with a large scene, high precision, and high efficiency is achieved, which can improve the quality of underwater image acquisition in a turbid water environment and serious light absorption, and save the construction of three-dimensional point cloud data, thereby improving the precision and efficiency of three-dimensional reconstruction of a large-scale underwater scene.

[0008] In some embodiments, performing epipolar geometric alignment on feature points in adjacent keyframe images to obtain feature matching point pairs in adjacent keyframe images includes: obtaining the pixel coordinates of feature points in adjacent keyframe images; generating a binary descriptor based on the pixel coordinates of the feature points; aligning an initial feature matching point pair based on the binary descriptor; and performing epipolar geometric alignment on the initial feature matching point pair to obtain feature matching point pairs in adjacent keyframe images.

[0009] This setup performs epipolar geometry comparison on feature points in adjacent keyframe images to identify feature point pairs that actually belong to the same feature point in adjacent keyframe images. This reduces the number of feature point calculations and provides accurate feature point data for subsequent camera pose estimation and 3D reconstruction.

[0010] In some embodiments, calculating the relative pose of adjacent keyframe images based on the pixel coordinates of the same matched feature point in adjacent keyframe images includes: establishing an epipolar geometric constraint equation and solving for the essential matrix based on the pixel coordinates of the same matched feature point in adjacent keyframe images and preset parameters of the camera; and decomposing the essential matrix to obtain the relative pose of adjacent keyframe images.

[0011] With this setup, when the camera travels along a straight trajectory, there is a stable relative motion relationship between consecutive frames. By accurately capturing its minute pose changes, even in low-texture areas, a rigorous geometric modeling and analytical solution process can be used to accurately estimate the relative pose of the camera, thereby enhancing the geometric consistency of 3D reconstruction.

[0012] In some embodiments, calculating the pixel depth value of each pixel in the keyframe image based on the relative pose includes: calculating the depth value of feature points in the keyframe image based on the relative pose; assigning a preset number of initial depth values ​​to each pixel in the keyframe image, wherein the initial depth values ​​are close to the depth value of the feature point closest to the pixel; performing photometric consistency calculations on the preset number of initial depth values ​​and the depth values ​​of each target feature point corresponding to the pixel to obtain scores for different initial depth values ​​relative to the depth value of each target feature point, and selecting the initial depth value with the highest score from all the scores of the initial depth values ​​corresponding to each target feature point as the target depth value; the target depth value corresponds one-to-one with the target feature point; the target feature point is a feature point in the keyframe image whose distance from the pixel is within a preset range; and determining the pixel depth value of the pixel based on the target depth values ​​corresponding to all target feature points.

[0013] This setup, by assigning an initial depth value to achieve the local plane assumption and photometric consistency constraint, helps to accurately restore the concave shape of underwater concrete crater areas when reconstructing them, even though some smooth surfaces lack obvious features. This is achieved through depth guidance from nearby edge feature points and multi-view color consistency verification, effectively solving the problem of missing depth in non-feature areas and significantly improving the integrity and detail restoration capabilities of the reconstruction results.

[0014] In some embodiments, the step of generating a 3D mesh model of an underwater concrete scene by multi-view fusion based on the relative pose, the 3D coordinates of the matched feature points, and the pixel depth value of each pixel includes: performing multi-view fusion based on the relative pose, the 3D coordinates of the matched feature points, and the pixel depth value of each pixel to generate a sparse 3D point cloud; constructing a voxel mesh for the sparse 3D point cloud, wherein each voxel corresponds to a pixel; calculating the symbolic distance of each voxel in multiple views based on the depth value of each pixel, wherein the symbolic distance represents the distance of the current voxel from the boundary of the sparse 3D point cloud; removing the target symbolic distance from the multiple symbolic distances corresponding to each voxel, and obtaining the effective distance corresponding to each voxel based on the remaining symbolic distances; wherein the number of differences between the target symbolic distance and other symbolic distances corresponding to the voxel that exceed a preset value is greater than a preset number; and generating a 3D mesh model of the surface of the underwater concrete scene based on the effective distance corresponding to each voxel.

[0015] This setup, by inverting the gradient field of the divergence field through the effective distances corresponding to all voxels, generates a watertight triangular mesh model. Through a rigorous voxel mesh fusion and surface reconstruction mechanism, it achieves high-quality 3D modeling from discrete observations to continuous entities, significantly improving the topological integrity and usability of the reconstruction results.

[0016] In some embodiments, obtaining multiple keyframe images from the video data includes: selecting multiple video images from the video data; performing background light stripping on each video image to obtain a background light stripped image; performing cast rate compensation on the background light stripped image to obtain a preliminary clear image; and performing adaptive color correction on the preliminary clear image to obtain an image-enhanced keyframe image.

[0017] This setup, through enhanced processes such as background light stripping and cast rate compensation, significantly improves image quality, providing high-quality input for subsequent 3D reconstruction and enhancing applicability in high-turbidity environments.

[0018] In some embodiments, the step of performing transmittance compensation on the background light stripped image to obtain a preliminary clear image includes: performing transmittance compensation on the background light stripped image to obtain a compensated image; and performing regularization optimization on the compensated image to obtain a preliminary clear image.

[0019] This configuration, through a combination of transmittance compensation and regularization optimization, improves the accuracy and stability of transmittance estimation.

[0020] Secondly, embodiments of the present invention provide a three-dimensional reconstruction system for an underwater concrete scene. The system includes: a data acquisition module for acquiring video data of the underwater concrete surface and acquiring multiple keyframe images from the video data; a three-dimensional reconstruction module for extracting feature points from each keyframe image, wherein the feature points represent conspicuous parts of the underwater concrete surface; performing epipolar geometric comparison on feature points in adjacent keyframe images to obtain feature matching point pairs in adjacent keyframe images, wherein the feature matching point pairs include the pixel coordinates of the same matched feature point in the adjacent keyframe images; and performing a three-dimensional reconstruction based on the same matched feature point. The relative pose of adjacent keyframe images is calculated based on the pixel coordinates of the same matched feature point in the adjacent keyframe images. The relative pose represents the rotation angle and translation vector of the camera corresponding to the adjacent keyframe images. The three-dimensional coordinates of the feature point are calculated based on the pixel coordinates of the same matched feature point in the adjacent keyframe images and the relative pose. The pixel depth value of each pixel in the keyframe image is calculated based on the relative pose. Multi-view fusion is performed based on the relative pose, the three-dimensional coordinates of the matched feature point, and the pixel depth value of each pixel to generate a three-dimensional mesh model of the underwater concrete scene.

[0021] Thirdly, embodiments of the present invention provide an electronic device, including a processor and a memory, wherein the memory stores a computer program executable by the processor, and the processor can execute the computer program to implement the underwater concrete scene three-dimensional reconstruction method as described in the first aspect.

[0022] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the three-dimensional reconstruction method for an underwater concrete scene as described in the first aspect.

[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0024] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 A flowchart of three-dimensional reconstruction of an underwater concrete scene is provided as an embodiment of the present invention; Figure 2 A comparison image of the original image and the enhanced image provided for an embodiment of the present invention; Figure 3 These are comparison images of different image enhancement algorithms provided in embodiments of the present invention. Figure 4 for Figure 1 Flowchart of sub-steps S201~S204 of step S200; Figure 5 The keyframe extraction images and robot-acquired images provided in the embodiments of the present invention; Figure 6 for Figure 1 Flowcharts of sub-steps S401~S404 in step S400; Figure 7 for Figure 1 The flowchart of sub-steps S501~S502 in step S500; Figure 8 for Figure 1 Flowcharts of sub-steps S701~S704 in step S700; Figure 9 for Figure 1 Flowcharts of sub-steps S801~S805 of step S800; Figure 10 The 3D reconstructed and manually stitched images provided in the embodiments of the present invention; Figure 11 This is a schematic diagram of the measurement of a 3D reconstructed concrete crater provided in an embodiment of the present invention; Figure 12 These are comparison images showing the effects of image enhancement provided in embodiments of the present invention. Figure 13 Large-scale 3D reconstruction image of underwater concrete structure provided in the embodiments of the present invention; Figure 14 A schematic diagram of the functional modules of the underwater concrete scene three-dimensional reconstruction system provided in an embodiment of the present invention; Figure 15 A block diagram of an electronic device provided in an embodiment of the present invention.

[0026] Icons: 1000 - Underwater concrete scene 3D reconstruction system; 1100 - Data acquisition module; 1200 - Image enhancement module; 1300 - 3D reconstruction module; 2000 - Electronic equipment; 2100 - Processor; 2200 - Memory; 2300 - Bus; 2400 - Communication interface. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0028] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0029] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0030] The deep integration of underwater robot (ROV / AUV) technology with high-precision optical sensing technology has propelled a qualitative leap in underwater structure inspection methods for hydropower projects. Equipped with high-definition cameras, high-brightness lighting systems, and laser scanners, underwater robots can conduct close-range, highly stable data acquisition underwater, simultaneously acquiring high-definition images and 3D point cloud information of the structural surface. This enables precise quantitative analysis and digital recording of concrete defects, laying a crucial technological foundation for the subsequent construction of an intelligent underwater structure inspection system.

[0031] Recent advancements have been made in applying underwater robots to underwater concrete structure damage detection, effectively improving the level of intelligent detection in this field. Although the method has been successfully applied in underwater concrete inspection, significant challenges remain under complex conditions: first, the concrete surface is often covered with biological deposits and sediments, obscuring the true structural morphology; second, underwater optical images are prone to degradation issues such as color distortion, blurring, and glare; and third, large-scale underwater scene 3D reconstruction requires a balance between accuracy and efficiency, with global defect localization and quantification remaining unsolved problems.

[0032] As described in the background section, due to the turbidity of the underwater environment and severe light absorption, the captured optical images will suffer from color distortion, blurring, and glare, resulting in poor quality of the acquired underwater images. It takes a lot of time to construct three-dimensional point cloud data, which leads to low accuracy and efficiency in the three-dimensional reconstruction of large-scale underwater scenes.

[0033] To this end, embodiments of the present invention provide a dual-modal underwater robot equipped with a dredging function and a clear water observation component. The dredging component cleans away biological deposits and sediments adhering to the concrete surface, which helps the clear water observation component to observe a clearer and more realistic concrete surface structure. Furthermore, embodiments of the present invention also provide a method for three-dimensional reconstruction of underwater concrete scenes, see reference. Figure 1 , Figure 1 This invention provides a flowchart for the 3D reconstruction of an underwater concrete scene. It obtains feature points from enhanced video images, calculates the camera's movement distance and translation vector based on these feature points, and then combines this with pixel depth values ​​from each keyframe image for 3D reconstruction. This achieves large-scale, high-precision, and high-efficiency underwater 3D reconstruction. It improves the quality of underwater image acquisition even in turbid water environments with severe light absorption, and saves on the construction of 3D point cloud data, thereby improving the accuracy and efficiency of large-scale underwater scene 3D reconstruction. The underwater concrete scene 3D reconstruction method also includes steps S100-S800: S100: Acquire video data of the underwater concrete surface.

[0034] In this embodiment, video data acquisition relies on a vision system mounted on a dual-modal underwater robot. This robot possesses dredging capabilities and a clear water observation component, enabling it to adapt to harsh underwater environments such as high turbidity, low light levels, and complex currents. During video data acquisition, the robot's dredging component first removes biological deposits and sediment covering the underwater concrete structure surface to expose its true structural morphology. Subsequently, the clear water observation component is lowered to fit closely to the concrete surface, effectively reducing the scattering and absorption of light by the water, thereby improving image quality. Specifically, the vision system continuously acquires high-definition video streams during the robot's movement, with an image resolution of 1920×1080 and a frame rate of 30FPS. The camera used is an LBF-C50HD2, and the raw video data is transmitted in real-time to a surface base station via a buoyancy cable. Furthermore, the camera's intrinsic parameters have been pre-calibrated, and lens distortion has been corrected to ensure the accuracy of subsequent geometric calculations. In practical applications, the robot moves at a constant speed of 20 cm / s along the water-stop joint area of ​​the dam, ensuring that the video sequence has good spatiotemporal continuity and sufficient parallax coverage, so as to achieve stable and reliable acquisition of raw video data that can be used for 3D reconstruction in complex engineering scenarios.

[0035] For example, see Figure 2 , Figure 2 A comparison image of the original image and the enhanced image provided for an embodiment of the present invention. Original image Figure 2 (a) There is significant blurring and low contrast, which will be detrimental to feature matching; the enhanced image Figure 2 (b) It possesses richer image details, which will play a crucial role in feature matching. To quantify the image enhancement quality, a cross-sectional comparative experiment was introduced, comparing the proposed method with two existing underwater image enhancement algorithms. Four image quality evaluation metrics were introduced to evaluate the enhanced image: Underwater Color Image Quality Evaluation (UCIQE), Underwater Image Quality Measure (UIQM), Natural Image Quality Evaluator (NIQE), and Peak Signal-to-Noise Ratio (PSNR). The comparative effects of different image enhancement algorithms are shown below. Figure 3 As shown, Figure 3The figures show a comparison of different image enhancement algorithms provided in embodiments of the present invention. In the figures, Originalimage refers to the original image; Ref[x] refers to the image enhancement method proposed by Lin et al.; L2uwe[x] refers to an underwater image enhancement algorithm proposed by Marques and Albu at CVPRW 2020; and Proposed Method refers to the image enhancement method proposed in embodiments of the present invention. Figure 3 The image enhancement method of the present invention can be used to demonstrate the advantages of visual effects, such as more natural color restoration, higher texture clarity and better contrast.

[0036] Furthermore, since underwater concrete image data falls under the category of no-reference image quality assessment, three no-reference image quality assessment methods were employed here. In addition, to further ensure the rigor of the experiment, images captured by the robot at a stationary underwater moment (with the sediment having naturally settled) were used as reference images for the PSNR index. The evaluation results of images enhanced by different algorithms are shown in Table 1.

[0037] Table 1

[0038] In Table 1, "Original image" represents the original image, "Ref[x] our method" represents the method according to the embodiments of this invention, "Ref[x]" represents the reference method, and "Propose method" represents the method proposed in the embodiments of this invention. It can be seen that a smaller NIQE value indicates better image quality, while for other evaluation indicators, a larger value indicates better image quality. The statistical results show that our proposed image enhancement method based on wavelength compensation and depth optimization scored 0.49, 3.56, 17.58, and 20.88 in the four evaluation indicators, respectively, achieving the best score in the entire evaluation. It is worth noting that the PSNR image evaluation is only for reference, because in practical engineering applications, we cannot wait for the sediment in the water to settle naturally at each image acquisition point before collecting data.

[0039] S200: Obtain multiple keyframe images from video data.

[0040] In this embodiment, the video data is a raw image stream continuously acquired by the LBF-C50HD2 high-definition camera mounted on a dual-modal underwater robot during underwater operations. The resolution is 1920×1080, and the frame rate is 30 FPS, recording the dynamic observation process of the underwater concrete surface of the dam after dredging. During the above steps, a temporal analysis is first performed on the entire video data. Several frames are extracted as candidate keyframes according to a preset time interval strategy. Specifically, one frame is selected every 0.5 seconds to ensure appropriate parallax coverage between adjacent keyframes while the robot moves at a constant speed of 20 cm / s, while avoiding redundant calculations due to excessive frame overlap. It should be noted that the keyframe images here are defined as representative static image frames selected from the original video sequence that possess independent spatial observation information and can be used for subsequent 3D reconstruction processing. The selection principle considers both temporal uniformity and the degree of motion change, ensuring scene coverage integrity while meeting the basic requirements of the SFM algorithm for parallax input. Furthermore, considering the slight shaking and motion blur present in the underwater environment, the selected keyframes were all taken from video clips during periods of stable robot posture. Frames with significant blurring due to rapid movement or mechanical vibration were excluded, thus achieving the transformation from continuous dynamic video to discrete static images. This provides a well-structured and temporally coherent basic data unit for subsequent feature extraction and matching. For example, in practical engineering applications, by extracting keyframes from a video stream lasting several minutes, hundreds of evenly distributed keyframe images were obtained, effectively supporting the large-scale 3D reconstruction task of a 30-meter-long watertight joint area. This not only reduced the total amount of data processing but also ensured the rationality and effectiveness of the input images in terms of spatial distribution and temporal sequence.

[0041] S300. Extract feature points from each keyframe image. These feature points represent prominent parts of the underwater concrete surface.

[0042] In this embodiment, feature points are defined as pixel locations in an image with significant local changes. They possess good repeatability and descriptive stability, making them suitable for cross-view matching. Feature point extraction employs the FAST corner detection algorithm to process each enhanced keyframe image. This algorithm identifies corners by determining whether the grayscale difference between the center pixel and several sampling points on its surrounding circumference exceeds a set threshold. It boasts advantages such as fast computation speed and strong robustness, making it particularly suitable for underwater operations with high real-time requirements. Specifically, the input is an enhanced grayscale image, and the output is a set of two-dimensional pixel coordinates, each coordinate corresponding to a potential corner location. Further, to improve rotation invariance, a binary descriptor set is generated using rotation-aware binary Robust Independent Elementary Features (rBRIEF) for similarity measurement in subsequent matching stages. It should be noted that prominent areas here specifically include textured regions such as crack edges on concrete surfaces, aggregate boundaries, construction joints, and uneven structures formed by erosion. These regions appear in the image as locations with drastic gradient changes or high local contrast. Therefore, the quality of feature point extraction directly affects the success rate of subsequent matching and the accuracy of 3D reconstruction. In practical applications, although the overall texture of underwater concrete surfaces is relatively simple, local details are restored after image enhancement, allowing the FAST algorithm to still stably extract a sufficient number of effective feature points. Based on the above description, this step completes the task of locating key structural information from a 2D image, providing a basic unit for establishing multi-view geometric associations.

[0043] S400. Perform epipolar geometry comparison on the feature points in adjacent keyframe images to obtain feature matching point pairs in adjacent keyframe images. The feature matching point pairs contain the pixel coordinates of the same matched feature point in the adjacent keyframe images.

[0044] In this embodiment, feature points extracted from adjacent keyframe images and their corresponding rBRIEF binary descriptors are first obtained. Then, based on the Hamming distance metric for the similarity between different descriptors, a matching threshold is set to filter out preliminary candidate matching point pairs, forming an initial set of feature matching point pairs. However, relying solely on descriptor similarity may lead to a large number of mismatches, especially in cases of drastic lighting changes or the presence of repetitive textures. Therefore, a geometric consistency verification mechanism must be introduced. Specifically, the epipolar geometric constraint relationship expressed by the fundamental matrix is ​​used to verify the initial matching results: for a pair of matching points (x... a ,x b If x satisfies b Fx aIf the Sampson distance is approximately 0 and less than a preset threshold, the matching point pair is retained; otherwise, it is discarded. Furthermore, the fundamental matrix F is robustly estimated using the RANSAC algorithm, enabling accurate fitting of the epipolar constraint model even with numerous outliers. It should be noted that the same matched feature point here refers to the imaging position of a fixed physical point in space under two adjacent viewpoints, with the two establishing a projection correspondence through camera motion. Thus, this step achieves a screening process from potential matching to geometrically consistent matching, significantly improving the reliability of the matching results. For example, in the underwater environment of a dam, slight blurring and displacement between images are caused by water flow disturbances and minor robot vibrations, yet this method still maintains a high matching accuracy. Based on the above description, this step provides a high-confidence set of corresponding points for camera pose estimation, serving as a crucial link connecting two-dimensional observation and three-dimensional structure.

[0045] S500: Calculate the relative pose of adjacent keyframe images based on the pixel coordinates of the same matched feature point in adjacent keyframe images. The relative pose represents the rotation angle and translation vector of the camera corresponding to the adjacent keyframe images.

[0046] In this embodiment, the camera intrinsic parameter matrix K has been obtained through prior calibration and is used as a known parameter in the calculation. However, the relative pose between two adjacent views cannot be directly measured and must be solved indirectly through matching point pairs. During the above steps, the pixel coordinates of the feature matching point pairs are first converted into normalized image coordinates, i.e., x = K. -1 p; then based on the polar geometry constraint equation x b T Ex a = 0 Construct a solution model for the essential matrix E, where E = [t]×R, [t]× is the antisymmetric matrix of the translation vector t, and R is the rotation matrix. Specifically, the fundamental matrix F is estimated using the eight-point method or multi-point least squares method, and then E = K -T FK derives the essential matrix E; after performing singular value decomposition (SVD) on E, four possible (R, t) solution combinations are obtained. Furthermore, by triangulating the matching point pairs and determining the physical rationality constraint of whether their 3D coordinate depth is positive, the correct relative pose solution is uniquely determined. Thus, this method fully utilizes the geometric characteristics of rigid body motion, avoiding ambiguity in the solution space. In practical applications, when a dual-modal underwater robot moves slowly along a straight trajectory, a stable relative motion relationship exists between consecutive frames. This method can accurately capture its pose changes, maintaining a certain estimation capability even in low-texture regions. Based on the above description, this step achieves accurate estimation of the camera's motion state, providing reliable motion parameter support for 3D structure reconstruction.

[0047] S600: Calculate the three-dimensional coordinates of the matched feature point based on the pixel coordinates and relative poses of the same matched feature point in adjacent keyframe images.

[0048] In this embodiment, the three-dimensional coordinates of spatial points are solved using a triangulation method based on the known relative poses of the cameras and the pixel coordinates of the matching point pairs. Specifically, the projection matrices P1 and P2 of the two cameras are compared with the corresponding image points x... a and x b Substituting the perspective projection equation λx = PX, an overdetermined system of equations is constructed, and the optimal 3D point X is solved by minimizing the reprojection error. Furthermore, considering the presence of measurement noise, Direct Linear Transformation (DLT) or SVD methods are typically used for numerical solutions to ensure stable and reliable results. It should be noted that the matched feature points here refer to reliable matching point pairs verified by epipolar geometry, and their 3D coordinate calculation depends on the precise relative pose between the two views. Thus, this process realizes the mapping transformation from 2D image observation to 3D spatial coordinates, and is the core step in generating sparse point clouds. For example, when reconstructing a concrete crater area, multiple corner points are triangulated to form local point cloud clusters, initially outlining the spatial contour of the defect. Further, as more views are added, the scale of the sparse point cloud can be continuously expanded through incremental SFM, and error accumulation can be suppressed through local bundle adjustment. Based on the above description, this step completes the initial 3D structure construction task, providing geometric skeleton support for subsequent dense reconstruction.

[0049] S700: Calculate the pixel depth value of each pixel in the keyframe image based on the relative pose.

[0050] In this embodiment, only a few feature points have 3D coordinates, and most pixels do not participate in matching. Therefore, a densitying strategy is needed to derive the depth information of each pixel. During the above steps, firstly, based on the relative pose and the position of a pixel in the reference image, it is assumed to be located in a certain depth plane, and homography is used to project it onto other adjacent views. Then, the photometric consistency between the projected area and the actual image patch is evaluated. A commonly used metric is zero-mean normalized cross-correlation (ZNCC), with a value closer to 1 indicating higher consistency. Specifically, the PatchMatch algorithm is used to efficiently search for the optimal depth hypothesis, iteratively optimizing the depth estimate of each pixel within the depth search space. Furthermore, to improve robustness, only a few views with the highest consistency scores are selected for fusion calculation, and a smoothing regularization term is introduced to suppress noise and occlusion effects. It should be noted that the pixel depth value here is defined as the distance from the camera optical center to the corresponding spatial point along the line of sight, and its calculation depends on the consistency assumption of multi-view observations and the local plane model. Thus, this method achieves an extension from sparse 3D points to dense depth maps. In practical applications, even in areas with poor texture, as long as there is sufficient parallax and lighting consistency, reasonable depth information can still be recovered, providing pixel-by-pixel depth support for generating complete surface geometry, which is a key step in achieving high-resolution reconstruction.

[0051] S800 performs multi-view fusion based on relative pose, the 3D coordinates of matched feature points, and the pixel depth value of each pixel to generate a 3D mesh model of the underwater concrete scene.

[0052] In this embodiment, the relative poses of all keyframes, the coordinates of sparse 3D points, and the pixel depth map of each frame are first used as input for multi-view fusion processing. Specifically, the 3D space is divided into a uniform voxel grid, and the symbolic distance (SDF) from each voxel to the surface is calculated along the line of sight of each camera. The observation results from different perspectives are then integrated by weighted averaging to form a truncated symbolic distance field (TSDF). Furthermore, to suppress the influence of noise and outliers, the multiple symbolic distances received by each voxel are filtered: if the number of differences between a distance and other distances exceeding a preset threshold is greater than a preset number, it is considered an outlier and is removed. The remaining distances are used to calculate the effective distance. It should be noted that multi-view fusion here refers to unifying multiple local observations into a global coordinate system to construct a consistent implicit surface representation. Thus, this process effectively reduces the uncertainty caused by single-view errors and occlusion. Furthermore, based on the constructed TSDF volume field, the Poisson reconstruction algorithm is used to solve the gradient field inversion problem of the indicator function, transforming the implicit function into explicit geometry; that is, extracting the isosurface with an indicator function value of 0.5 as the object surface, ultimately generating a watertight triangular mesh model. For example, in engineering applications, this method successfully reconstructed a 30-meter-long dam cutoff joint area, clearly presenting the erosion zone and crater structure, completing a high-quality 3D model from discrete observation to a continuous entity, significantly improving the topological integrity and usability of the reconstruction results, and meeting the needs of subsequent damage quantification analysis.

[0053] In some embodiments, the obtained keyframe images need to undergo image enhancement processing, see [reference]. Figure 4 , Figure 4 for Figure 1 The flowchart of sub-steps S201~S204 of step S200, wherein steps S201~S204 include: S201. Select multiple video images from the video data.

[0054] In this embodiment, the video data is a raw image stream continuously acquired by the LBF-C50HD2 high-definition camera mounted on a dual-modal underwater robot during the detection process. It contains a temporally continuous but highly redundant sequence of frames. During the above steps, the entire video stream is first traversed temporally, and a representative set of video images is selected as the basic input for subsequent processing based on a preset time interval strategy or motion change threshold. Specifically, a fixed time interval sampling method is used, extracting one frame every 0.5 seconds to ensure that the selected images are evenly distributed in spatial coverage while meeting the parallax range requirements of the incremental SfM algorithm. It should be noted that the video images here are defined as single frames from the raw video stream, without any enhancement or correction processing, retaining the degradation characteristics of the underwater imaging process, such as bluish color, low contrast, and background light interference. Furthermore, the selection process prioritizes excluding blurred frames caused by robot startup, turning, or water current disturbance to ensure that the candidate images possess basic clarity and stability. Therefore, this step achieves the goal of extracting effective observation frames from a highly redundant video stream, providing a structurally complete and computationally efficient image set for subsequent image enhancement and 3D reconstruction. In practical applications, this method can extract hundreds of effective images from several minutes of continuous video, supporting large-scale 3D reconstruction tasks. As the starting point of the image enhancement process, it directly affects the quality and efficiency of subsequent processing.

[0055] For example, when a robot is working underwater, video data may suffer from motion blur and image blurring caused by suspended particles, such as... Figure 5 As shown, Figure 5 The keyframe extraction images and robot-acquired images provided in this embodiment of the invention. Figure 5 (a) shows that the keyframes extracted from the video exhibit image blurring; Figure 5 (b) High-resolution (sharp) images acquired under static robot conditions. Although the quality of keyframe images extracted from video is relatively low, this is beneficial for reducing detection time in detection engineering tasks. Furthermore, video data has temporal continuity, thus enabling 3D reconstruction based on the SFM method. During underwater imaging, optical cameras experience absorption and scattering of light by the water. According to the Jaffe-McGlamery underwater optics theory, underwater image degradation is mainly caused by the exponential decay of light reflected from objects during propagation and backscattering caused by suspended particles in the water. These two processes lead to color distortion, decreased contrast, and blurring in the image. Based on the principle of light transmission underwater, the physical model of underwater imaging can be described as:

[0056] in, This represents a degraded image captured by the camera. Represents the color channels of an image; This indicates the image that needs to be enhanced; Indicates wavelength-dependent transmittance; It is the attenuation coefficient. It is the absorption coefficient. Represents the scattering coefficient; It is the global background light, i.e., the backscattered component; Indicates the depth of the scene. Different wavelengths attenuate differently underwater, satisfying... .

[0057] S202. Perform background light stripping on each video image to obtain a background light stripped image.

[0058] In this embodiment, for each video image selected in S201, background light component estimation and removal operations are performed to eliminate the impact of backscattered light caused by suspended particles in the water on image quality. Specifically, based on the Jaffe-McGlamery underwater optical imaging model, the degraded image is decomposed into the sum of object reflected light and background scattered light, where the background light mainly originates from non-target light scattering from the water in front of the camera. To accurately estimate the background light, a "brightness-gradient" dual-constraint strategy is adopted: first, the set of pixels with brightness values ​​in the top 2% of the image is identified, and regions with gradient amplitudes below an adaptive threshold are selected from this set to avoid interference caused by direct light sources, specular reflection areas, and locally bright concrete materials; then, the pixel grayscale values ​​in this set are statistically analyzed, and their mean or weighted average is taken as the global background light intensity estimate. Further, the background light component is subtracted from the original image to obtain a background light stripped image that retains only the reflection information of the object surface. It should be noted that the background light stripping here is defined as a physically prior-based image preprocessing technique aimed at restoring the radiation of the real scene and improving image contrast and detail visibility. Therefore, this step effectively alleviates image haze caused by scattering in high-turbidity environments. For example, in the experiment, the original image exhibited obvious whitening and low contrast, such as... Figure 5 As shown in (a), after background light stripping, the boundaries between cracks and aggregates on the concrete surface become clearer, as shown in (a). Figure 5 As shown in (b). Based on the above description, this step, as the first stage of the image enhancement chain, provides a cleaner input basis for subsequent transmittance estimation and color restoration.

[0059] For example, the background light originates from the physical properties of water scattering. To avoid interference from bright objects (light sources, specular reflections from concrete materials), gradient constraints will be used to exclude interference from bright objects and ensure the extraction of true background scattered light. For instance, a brightness-gradient joint constraint method will be used to estimate the background light, expressed as:

[0060] in, It is the set of pixels with the highest brightness (top 2%). grayscale image It is the 98th percentile of the cumulative luminance distribution; It is the gradient magnitude; It uses an adaptive threshold. The main purpose is to retain only pixels with gradient values ​​less than the adaptive threshold (these pixels have gradual brightness changes, consistent with the uniformity of the background light), while completely filtering out bright, high-gradient interference. Finally, the median of these filtered pure background light candidate pixels is taken as the final background light, which represents the true brightness of water scattering while avoiding the influence of a few residual outliers. This completely separates interfering bright spots from the true background light, clearing the way for subsequent image enhancement.

[0061] S203. Perform transmittance compensation on the background light stripped image to obtain a preliminary clear image.

[0062] In this embodiment, transmittance compensation is performed on the background light stripping image to obtain a compensated image; regularization optimization is then performed on the compensated image to obtain a preliminary clear image.

[0063] Furthermore, transmittance refers to the proportion of light that is not absorbed or scattered during its propagation from the light source to the target surface and back to the camera; its spatial distribution directly determines the image sharpness. In the above steps, firstly, based on an improved dark channel prior method combined with a wavelength compensation factor, the transmittance of the image after background light stripping is initially estimated. Local minimum filtering is used to extract the image's dark channel, and a wavelength attenuation ratio is introduced to adjust the transmittance differences between different color channels, making the recovery of the red channel more consistent with actual attenuation patterns. Furthermore, to overcome the block artifacts and spatial discontinuities caused by traditional dark channel priors, transmittance estimation is transformed into a depth optimization problem. A dark channel prior method is introduced into the image transmittance estimation process, combined with wavelength characteristics, improving the transmittance estimation model, which can be expressed as:

[0064] in, Indicates a local neighborhood; It is a deep retention factor; This represents the wavelength compensation ratio. In underwater images, apart from a few bright areas such as light sources and reflections, most areas have localized dark pixels with very low brightness (such as shadows in concrete cracks or dark parts of the water). These dark pixels are mainly caused by light scattered by the water. The dark channel prior utilizes this characteristic, finding the darkest local pixel to preliminarily estimate the degree of light attenuation due to scattering, essentially providing a basic reference value for transmittance. Generally, underwater, red light attenuates the fastest, followed by green light, and blue light the slowest. This difference can cause severe color casts in the image (e.g., lack of red and blue tint). Therefore, based on the dark channel prior, wavelength characteristics (using the formula...) are added. This allows for customized corrections for different colors: for example, if red light attenuates strongly, its estimated transmittance value is lowered using this factor to match the actual attenuation pattern of red light; if blue light attenuates weakly, it is appropriately increased to avoid over-correction. In essence, the dark channel prior is responsible for the overall direction (identifying the approximate impact of scattering), while wavelength characteristics are responsible for calibration (correcting deviations based on color differences). Combining these two factors allows for the calculation of transmittance that better reflects the actual underwater conditions. Subsequently, this transmittance is used to remove scattered light and supplement missing colors, making blurry or discolored underwater images clearer and more natural in color.

[0065] Furthermore, since the transmittance map generated by the dark channel prior may contain block artifacts, the physical constraint relationship of the Beer-Lambert law is introduced here. Transmittance estimation is transformed into a depth optimization problem, and the depth optimization model is expressed as:

[0066] in, Scene depth to be optimized; It is a smoothing weight; It's the depth gradient. Physical consistency is achieved through the depth map. Transmittance can then be calculated back. This will make the red light channel transmittance more consistent with actual attenuation; there will be no blocky abrupt changes at the target edge; and the continuous region of the depth step can maintain a sharp transition.

[0067] Furthermore, in order to overcome the noise amplification problem and the ill-conditioned inverse problem during scene reconstruction (when... When the equation solution is unstable, a variational optimization framework is introduced for regularization modeling:

[0068] in, This indicates the optimized transmittance; The TV norm represents the non-convex form. To balance quality and efficiency, a closed-form approximation is used to represent the sharp image. :

[0069] in, Indicates the critical threshold, when Directly using the physical inverse transform, when Enable regularization protection to avoid noise explosion. Regularization modeling can find a restored image that both closely matches the original image information and is sufficiently natural and clear, constrained by two parts: the first part is... The first part can be understood as calculating the fit between the restored image and the original image, to avoid the restored image appearing clear but not matching the actual scene (for example, replacing concrete cracks with other textures); the second part is the penalty item. Its function is to limit the area in the image. Sudden changes, such as abrupt changes in brightness in a certain area (not a real concrete edge or crack, but noise or random color blocks), will result in a penalty.

[0070] S204. Perform adaptive color correction on the initially clear image to obtain the enhanced keyframe image.

[0071] In this embodiment, to address the color distortion problem in the preliminary clear image output by S203, a dual-channel adaptive color correction operation is performed to restore the color distribution under natural visual perception. Specifically, the statistical characteristics of each color channel (R, G, B) are first analyzed, including the mean and standard deviation, to identify channel imbalance caused by selective absorption by water (typically manifested as an overly strong blue channel and severely attenuated red channel). Then, brightness compensation processing is performed, adjusting the gain coefficients of each channel to align their means and eliminate overall color cast. Based on this, dynamic contrast adjustment is implemented, using non-linear stretching based on the standard deviation and attenuation difference coefficient of each channel to restore saturation and color levels. It should be noted that adaptive color correction is defined as an automatic white balance and color restoration method based on channel statistical characteristics. Its parameters are dynamically adjusted according to the local features of each frame, rather than using a fixed transformation matrix. This effectively corrects the common blue-green bias problem in underwater imaging, making the color of the concrete body closer to the observation effect in air. For example, in the enhanced image (such as...), Figure 5 As shown in (b), the originally bluish concrete surface now appears grayish-white, and the distinction between cracks and deposits is more pronounced, which is beneficial for subsequent feature extraction and manual interpretation. Furthermore, the image obtained after this step is the final enhanced keyframe image, which will be used as input to the SFM algorithm for 3D reconstruction. Based on the above description, this step completes the final correction stage of the image enhancement process, ensuring that the output image meets the standards required for high-quality reconstruction in terms of brightness, contrast, and color.

[0072] Furthermore, to correct the distortion of the image color space, an adaptive color compensation model is designed based on the imbalance characteristics of the statistical distribution of color channels:

[0073] in, This represents the channels of the reconstructed input image; Indicates the channel mean; Indicates the standard deviation of the channel; This represents the attenuation difference coefficient.

[0074] It's understandable that red, green, and blue light diverge differently underwater. Red light doesn't spread far and is easily absorbed by water, while blue light can spread very far, causing photos to often appear bluish and lacking in red, resulting in colors that differ significantly from the actual scene—this is color distortion. Furthermore, the degree of color cast varies from photo to photo: some underwater photos are taken in murky water, resulting in more severe color casts; others, shooting close-ups, have slightly better colors, making it difficult to correct with a fixed set of color adjustment parameters. The adaptive color compensation model calculates the average brightness of each of the red, green, and blue color channels, then examines the color distribution of each channel. Next, it uses these two data points, combined with the color decay patterns of different colors underwater, to bring the brightness and distribution of the three colors back to equilibrium. This makes the originally color-distorted image look more natural and makes it easier to see details such as cracks and textures in concrete later on.

[0075] In some embodiments, to accurately locate matching feature points in adjacent keyframe images, see [reference]. Figure 6 , Figure 6 for Figure 1 The flowchart of sub-steps S401~S404 of step S400, wherein steps S401~S404 include: S401. Obtain the pixel coordinates of feature points in adjacent keyframe images.

[0076] In this embodiment, the pixel coordinates of the feature points refer to the two-dimensional image positions with significant local grayscale changes extracted from the enhanced keyframe images by the FAST corner detection algorithm. These coordinates constitute the basic input for subsequent descriptor generation and matching. During the execution of the above steps, feature detection operations are first performed on two adjacent keyframe images to identify the candidate corner point sets on their respective image planes. Specifically, several sampling points are selected around a pixel in a circle. If the grayscale values ​​of multiple consecutive points are all higher or lower than the center pixel plus a certain threshold, then the position is determined to be a corner point. The output is a set of two-dimensional coordinates, each corresponding to the precise position of a detected feature point in the image. It should be noted that adjacent keyframe images here specifically refer to two video frames that are consecutive in time or have sufficient parallax, and there is an estimable relative motion relationship between them, suitable for establishing geometric correspondence. Furthermore, the obtained pixel coordinates are not normalized and are directly represented based on the original image resolution (1920×1080) for subsequent descriptor calculation and matching verification, completing the feature space location task and providing basic data units for constructing cross-viewpoint correspondences. In practical applications, although the surface texture of underwater concrete is relatively simple, the edges and aggregate structure are clearly visible after image enhancement, which enables the FAST algorithm to still extract a sufficient number of effective feature points stably.

[0077] For example, the FAST algorithm is used to perform corner detection and feature point extraction through local brightness contrast, which is used to quickly locate stable feature points in underwater concrete images:

[0078] in, Indicates the input number of the first... Frame image; This represents the FAST corner detection algorithm, and its output... It is a set of coordinates of key pixels in the image. These coordinates precisely correspond to stable feature points with significant characteristics in the underwater concrete image (such as the corners of cracks, the raised edges of the concrete surface, the intersections of edges, etc.). Each coordinate marks the specific location of the feature point in the i-th frame of the image. .

[0079] S402. Generate a binary descriptor based on the pixel coordinates of the feature points.

[0080] In this embodiment, for each feature point with acquired pixel coordinates, a local neighborhood region is defined centered on that feature point. The rotation-aware BRIEF descriptor calculation is generated through grayscale comparison, and its output is a set of binary descriptors. ,in The method generates binary descriptors using the rBRIEF (rotation-aware Binary Robust Independent Elementary Features) algorithm. Specifically, this method compares the grayscale values ​​at preset positions within a local image patch to form a binary string composed of "0"s and "1"s. For example, if the grayscale value of the first point in a pair of sampling points is greater than that of the second point, it is recorded as "1"; otherwise, it is recorded as "0". Furthermore, to improve rotation invariance, the dominant direction of the feature point is estimated based on the gradient distribution around it, and then the sampling template is rotated and aligned according to this direction, thus ensuring that the same feature under different poses has similar descriptors. It should be noted that the binary descriptor here is defined as a fixed-length bit string (e.g., 256 bits), which has the advantages of high storage efficiency and fast matching speed. It is particularly suitable for resource-constrained underwater robot platforms, converting local image information into quantifiable and comparable digital features, enhancing the robustness of feature representation. For example, under uneven lighting or slight blurring conditions, even if pixel values ​​fluctuate slightly, as long as the relative grayscale relationship remains stable, the descriptor can still maintain consistency. Based on the above description, this step realizes the transformation from spatial location to semantic features, laying the technical foundation for efficient matching.

[0081] S403. Based on binary descriptors, the initial feature matching point pairs are determined.

[0082] In this embodiment, the initial feature matching aims to quickly filter out potentially corresponding feature point pairs as a candidate set for subsequent geometric verification. During the above steps, the binary descriptor of each feature point in the current frame is compared one by one with the descriptors of all feature points in the next frame, and the Hamming distance is used to measure the degree of difference between the two. Specifically, the Hamming distance is equal to the number of bits that differ in the same position between two binary strings of equal length; the smaller the distance, the higher the similarity. A matching threshold is set, and when the Hamming distance of a pair of descriptors is less than the threshold and it is the best match for the current point, it is recorded as an initial matching point pair. Furthermore, to improve matching accuracy, a ratio test can be introduced, that is, the match is only accepted when the ratio of the distance between the best match and the second best match is less than a set ratio, so as to exclude candidates with strong ambiguity. It should be noted that the initial feature matching point pairs here have not yet undergone geometric consistency verification and may contain a large number of mismatches, especially in areas with repetitive textures or low contrast.

[0083] For example, feature points in different images need to be correlated to support 3D reconstruction. Feature points extracted from consecutive frames need further feature matching to find pairs of points with similar descriptors in the two images. The feature matching set can be represented as:

[0084] in, , It is an image and Key points in; and A binary descriptor representing the key points of two consecutive frames of an image; This represents the Hamming distance function, used to calculate the difference between binary strings; This indicates the matching threshold.

[0085] For example, in a feature matching scenario where the steps are feature point extraction, descriptor generation, feature point matching, and geometric verification, image A captures a "crack corner A" of an underwater concrete block, generating the descriptor "01011001..."; image B captures the same "crack corner B" (slightly moved), generating the descriptor "01011010...". The Hamming distance between the two images is calculated to be 1 (only one bit different), which is less than a set threshold of 3. Therefore, the pair (corner A, corner B) is added to the corresponding point pair list M. a If the descriptor of another convex point C in image B is "10100111…", and its Hamming distance from corner A is 10, which is greater than the threshold, then it is determined that it is not a corresponding point.

[0086] S404. Perform epipolar geometry comparison on the initial feature matching point pairs to obtain feature matching point pairs in adjacent keyframe images.

[0087] In this embodiment, the initial feature matching point pairs obtained in S403 are geometrically validated using epipolar geometric constraints to eliminate erroneous matches that do not conform to the epipolar relationship. Specifically, the fundamental matrix F is first estimated based on the matching point set, and the RANSAC (Random Sample Consensus) algorithm is used to robustly fit the model in the presence of a large number of outliers: each time, the smallest sample set (such as the eight-point method) is randomly selected to calculate F, and then the results satisfying x are statistically analyzed. b Fx aThe number of interior points ≤ ε (Sampson distance) is iterated multiple times, and the F with the most supporters is retained as the optimal solution. Subsequently, all initial matching point pairs are substituted into this basic matrix for verification, and only point pairs that satisfy the epipolar constraint are retained as the final feature matching point pairs. It should be noted that the epipolar geometric comparison here refers to the geometric constraint established based on the camera projection relationship between two views. It requires that the projection of the same spatial point in the two images must lie on the corresponding epipolar line, thereby eliminating erroneous correspondences caused by descriptor similarity, effectively suppressing mismatches caused by lighting changes, texture repetition, or noise interference, and significantly improving the reliability of the matching results. For example, in the underwater environment of a dam, there are often regularly arranged construction joints or aggregate distributions on the concrete surface, which can easily cause confusion at the descriptor level. However, epipolar geometric verification can accurately retain the true corresponding points. Furthermore, the filtered feature matching point pairs will be used for solving the essential matrix and estimating the camera pose, directly affecting the geometric accuracy of 3D reconstruction.

[0088] For example, mismatches may occur during the matching process. Therefore, mismatches can be eliminated by comparing the epipolar geometry of the preceding and following images. The geometric verification relationship for feature point matching can be expressed as:

[0089] in, and The second coordinates of the key points; The fundamental matrix has epipolar geometric constraints. ; This represents the Huber robust loss function. Finally, Sampson distance is introduced to evaluate the matching quality:

[0090] Relying on descriptor similarity can introduce a large number of false matches. By using epipolar geometric constraints, false matches can be eliminated, and finally, geometrically consistent matching point pairs will be output.

[0091] For example, the core logic of using epipolar geometry constraints to eliminate mismatches is that the projection points of the same real-world point in two images must follow the geometric law that light rays and baselines are coplanar. Matching pairs that do not satisfy this law are incorrect, and the goal is to retain truly geometrically consistent matching point pairs. Suppose that in underwater dam detection, feature point A (the corner of a crack) is extracted from the previous image. Through descriptor similarity comparison, two candidate matching points B (the corner of the same crack) and C (the protrusion of another unrelated crack) are found in the subsequent image. In this case, using epipolar geometry constraints to eliminate mismatches can first calculate the fundamental matrix F (which contains the camera pose, intrinsic parameters, and other geometric relationships of the two images, and is the core rule carrier of epipolar geometry) using a small number of known correct matching points in the two images (such as manually marked obvious crack endpoints). Then, based on point A and the fundamental matrix F of the previous frame, calculate the corresponding epipolar line l2 in the next frame; then determine whether it is on the line: check whether candidate points B and C fall on the epipolar line l2: point B (true matching point): falls exactly on the epipolar line l2, which conforms to the geometric law of "coplanar light rays", and is retained as a correct matching pair; point C (incorrect matching point): deviates far from the epipolar line l2 (for example, the distance is more than 5 pixels), which violates the physical law, and is judged as an incorrect match and removed.

[0092] In some embodiments, to clearly know the angle of rotation and translation vector of the camera during the capture of two adjacent keyframe images, refer to... Figure 7 , Figure 7 for Figure 1 The flowchart of sub-steps S501~S502 of step S500, wherein steps S501~S502 include: S501. Based on the pixel coordinates of the same matched feature point in adjacent keyframe images and the preset parameters of the shooting camera, establish the epipolar geometric constraint equation and solve for the essential matrix.

[0093] In this embodiment, the preset parameters of the camera refer to the camera intrinsic parameter matrix obtained through prior calibration, which includes information such as focal length, principal point coordinates, and pixel scale, and are known prior conditions. During the execution of the above steps, the pixel coordinates of the feature matching point pairs verified by epipolar geometry in S404 are first transformed from the image coordinate system to the normalized image plane to eliminate the influence of lens distortion and imaging scale. Specifically, based on epipolar geometry theory, for matching point pairs in two images... and Satisfying the epipolar geometric constraint equations: ;in, and The intrinsic parameter matrix of the camera can be obtained through calibration; Represents the essential matrix, It is an antisymmetric matrix of the translation vector. This is the rotation matrix. In other words, by substituting all matching point pairs into this equation to construct an overdetermined linear system of equations, and solving it using the eight-point method or multi-point least squares method, the preliminary estimation results are regularized using singular value decomposition (SVD) to satisfy the inherent constraint that two non-zero singular values ​​must be equal. It should be noted that the epipolar geometric constraint equation here is defined as the geometric dependence of the projection points between the two views caused by camera motion, and its mathematical form strictly depends on the projective invariance under rigid body transformation. Therefore, the establishment and solution process of this equation realizes the transition from two-dimensional matching observation to three-dimensional moving structures, completing the core modeling task of the geometric relationship between cameras and providing the necessary input for subsequent pose decomposition.

[0094] S502. Decompose the essential matrix to obtain the relative pose of adjacent keyframe images.

[0095] In this embodiment, the singular value decomposition (SVD) of the essential matrix E can be expressed as E=Udiag(1,1,0). T Since the rotation matrix R has two rotation directions and the translation vector t has two directions (positive and negative), the combination of the two naturally produces four theoretical pose solutions. Each solution consists of the rotation matrix R and the unit-length translation vector t. A unique solution can be determined by the physical constraint that the depth of the triangulation point is positive. ;in, express The third column needs to satisfy the following conditions for solving: Camera pose transformation can be solved using the essential matrix of epipolar geometry, but this only determines the relative pose and sparse 3D point cloud, making it difficult to directly obtain the absolute scale of the scene. Here, triangulation is used to establish the minimum reprojection error. Triangulation utilizes the geometric consistency of multi-view observations and calculates the depth value of spatial points through the principle of ray convergence, thereby transforming the relative pose into a 3D structure with physical scale. The 3D coordinates optimized by triangulation can be represented as:

[0096] in, Indicates a three-dimensional point with estimation; Represents the perspective projection function; This represents the camera projection matrix. The 3D points are the initial 3D coordinates of key features on the underwater concrete surface obtained through triangulation calculations. Since the initial solution for the 3D coordinates has accumulated errors, it is necessary to minimize the reprojection error during global optimization. A minimum reprojection error optimization function is established by jointly optimizing all camera poses and 3D coordinates using Bundle Adjustment (BA):

[0097] in, Indicates the camera's pose; The covariance matrix represents the observation noise. SFM decomposes motion parameters through epipolar geometric constraints and then solves for structure and motion through a combination of triangulation and nonlinear optimization. Its accuracy depends on the robustness of feature matching and the convergence of bundle adjustment, while scale uncertainty needs to be addressed by introducing objects of known size or by sensor fusion.

[0098] It should be noted that the relative pose here is defined as the rotation angle and translation vector of the subsequent keyframe relative to the previous keyframe, fully describing the rigid body motion state of the camera during continuous shooting. Thus, this decomposition process fully utilizes the algebraic structure of the essential matrix and the physical constraints of 3D space, excluding mathematically valid but practically unrealizable solutions. For example, when a dual-modal underwater robot moves at a constant speed along the surface of a dam, its motion direction is mainly a forward vector, and the corresponding translation direction should be roughly consistent with the robot's heading. This prior knowledge helps to further verify the rationality of the solution. Furthermore, the obtained relative pose is used as the initial estimate for incremental SFM, and is used for pose solving of subsequent newly added views and global bundle adjustment optimization, realizing the analytical restoration from the essential matrix to specific motion parameters, providing an accurate attitude reference for 3D structure reconstruction.

[0099] In some embodiments, in order to calculate the depth value of each pixel in the keyframe image, refer to Figure 8 , Figure 8 for Figure 1 The flowchart of sub-steps S701~S704 of step S700, wherein steps S701~S704 include: S701. Calculate the depth values ​​of feature points in the keyframe image based on the relative pose.

[0100] In this embodiment, relative pose refers to the rotation matrix and translation vector of the camera between adjacent keyframes, obtained through step S500, which provides the necessary geometric relationship for triangulation. For feature point pairs that have been matched in S400, the coordinates of the feature point in 3D space are recovered using triangulation methods, utilizing their pixel coordinates in two or more views and the corresponding camera projection matrix. Specifically, the normalized coordinates of the matching points are substituted into the perspective projection equation to construct a system of linear equations about the 3D points, and the optimal solution is obtained by minimizing the reprojection error. Commonly used algorithms include DLT (Direct Linear Transformation) or numerical optimization methods based on SVD. Furthermore, the obtained 3D coordinates are used to calculate the depth value of the feature point relative to the reference frame camera, i.e., the distance of the spatial point along the camera's principal axis, by inversely calculating the camera's optical center position and viewing direction. It should be noted that the depth value of the feature point here is defined as the Euclidean distance between the optical center of the shooting camera and its corresponding physical surface point. This is the initial seed data for subsequent dense depth estimation, realizing the transformation from sparse matching to 3D structural information and providing a reliable benchmark for depth propagation in non-feature regions. In practical applications, underwater imaging is subject to slight blurring and noise interference, which may introduce errors into the triangulation results. Therefore, iterative optimization using local bundle adjustment is typically employed to improve depth accuracy. Based on the above description, this step lays the geometric foundation for pixel-level depth estimation, ensuring spatial consistency in subsequent dense reconstruction.

[0101] For example, suppose a pixel The neighborhood of lies on a small plane Above, the plane is formed by the normal. and depth Definition. Projecting pixel blocks from a reference image onto other views using a homography matrix:

[0102] in, and These represent the camera's first... Each intrinsic parameter matrix and the reference camera matrix; Indicates the estimation from the reference to the first The rotation matrix and translation vector of each camera. The main task is to calculate the depth value for each pixel (pixels without paired feature points) to form a 2.5D depth map. Each concrete area in the keyframe image now has information about its distance from the camera. For example, the depth values ​​of all pixels between A and B are between 1.2 and 1.25 meters, and the transition is smooth.

[0103] S702. Assign a preset number of initial depth values ​​to each pixel in the keyframe image. The initial depth values ​​are close to the depth values ​​of the feature points closest to the pixel.

[0104] In this embodiment, for ordinary pixels in the keyframe image that do not participate in feature matching, a set of candidate depth hypotheses is set as the initial range for subsequent optimization search. Specifically, firstly, several feature points whose spatial distance from the current pixel is within a preset neighborhood range are determined. These feature points have obtained reliable depth values ​​through S701. Then, several of these feature points with the closest depth values ​​are selected as references, and a preset number of initial depth values ​​are generated around their depth values. For example, N depth hypotheses are sampled with a fixed step size within the range of ±Δd. It should be noted that the preset number here is usually 8 to 16, which ensures sufficient coverage of the search space while avoiding excessive computational complexity. The feature point closest to the pixel not only refers to the one with the smallest Euclidean pixel distance, but also includes neighboring feature points located on the same local planar structure, in accordance with the prior assumption of local surface continuity, effectively reducing the search space for subsequent optimization and improving algorithm efficiency. For example, in a relatively flat area of ​​a concrete surface, the depths of neighboring feature points are highly correlated, and the initial depth set generated based on this is more likely to contain the true depth values. Furthermore, this process establishes a set of depth hypotheses for each pixel to drive efficient iterative optimization of PatchMatch-like algorithms, completing the initial expansion from sparse depth to dense depth and providing a reasonable starting point for achieving pixel-by-pixel depth estimation.

[0105] S703. Perform photometric consistency calculations on the preset number of initial depth values ​​and the depth values ​​of each target feature point corresponding to the pixel, obtain the score of different initial depth values ​​relative to the depth value of each target feature point, and select the initial depth value with the highest score from all the initial depth values ​​corresponding to each target feature point as the target depth value; the target depth value corresponds one-to-one with the target feature point; the target feature point is the feature point in the keyframe image whose distance from the pixel is within a preset range.

[0106] In this embodiment, photometric consistency refers to the requirement that the grayscale or color of the same physical surface point should remain similar under different viewing angles. This is the core assumption of multi-view stereo matching. During the execution of the above steps, for each initial depth value, it is regarded as the assumed depth of the current pixel. The image block containing the pixel is projected into other adjacent views using the relative pose obtained in S500, forming a reprojection region. Subsequently, the similarity score between the reference image block and the projected image block is calculated. The commonly used index is zero-mean normalized cross-correlation (ZNCC), and the closer the value is to 1, the higher the consistency. Specifically, for each target feature point (i.e., a known depth feature point located in the neighborhood of the current pixel), the projection consistency of all initial depth values ​​under multiple views is evaluated, and the initial depth value with the highest score is recorded as the target depth value corresponding to the target feature point. It should be noted that the target feature point here is limited to feature points whose distance from the current pixel on the image plane is less than a certain threshold. The purpose is to establish local structural associations and avoid introducing misleading information from distant irrelevant points. Furthermore, due to potential uneven lighting and color shifts in the underwater environment, ZNCC typically performs mean normalization on image patches before calculation to enhance robustness. Thus, this step verifies the depth hypothesis through multi-view observations, selecting the one that best conforms to the physical laws of imaging. In practical applications, even if consistency decreases in a single view due to glare or occlusion, the correct match can still be preserved through multi-view fusion. Based on the above description, this step achieves refined screening of the initial depth hypothesis, improving the accuracy of local depth estimation.

[0107] For example, since the underwater concrete reconstruction surface reflects light diffusely, the same physical point should have a similar color from multiple viewpoints. Therefore, cross-correlation is normalized by zero mean. Assessing projection consistency:

[0108] in, and Representing the reference image and the first Grayscale values ​​of each view; and This represents the average grayscale value of the reference image block and the projected image block; and This represents the standard deviation of gray levels between the reference image patch and the projected image patch; Indicated by A neighborhood window centered on the center; A value closer to 1 indicates a more accurate depth / hairline hypothesis; a value closer to 0 indicates occlusion or failure to meet the Lambert hypothesis. By continuously adjusting the preset depth value, the optimal depth value for each pixel can be quickly found through consistency verification.

[0109] S704. Determine the pixel depth value of each pixel based on the target depth values ​​corresponding to all target feature points.

[0110] In this embodiment, the target depth values ​​selected by each target feature point in S703 are aggregated and comprehensively judged to determine the final pixel depth value of the current pixel. Specifically, multiple target depth values ​​are fused using methods such as weighted averaging, median filtering, or multi-model fitting: if the target depth values ​​are concentrated, their mean or median is taken as the final result; if there is significant dispersion, weighted fusion is performed based on the visibility weight or consistency score of the corresponding view to suppress the influence of outliers. Furthermore, to improve surface smoothness and geometric continuity, a regularization term can be introduced to constrain the depth gradient changes between adjacent pixels to prevent drastic jumps. It should be noted that the pixel depth value here is defined as the scene depth corresponding to each pixel in the keyframe image, constituting the basic input for subsequent multi-view fusion to generate the symbolic distance field, effectively integrating depth cues from multiple neighboring feature points, and enhancing the stability and completeness of the estimation results. For example, in the edge area of ​​a concrete crater, although some pixels lack direct matching information, a reasonable depth transition can still be recovered through the collaborative guidance of multiple surrounding feature points. Furthermore, the obtained pixel depth map will serve as the core output of the MVS stage, used for TSDF voxel fusion and 3D mesh reconstruction. Based on the above description, this step completes the integration task from locally optimal depth to a globally consistent depth map, providing complete geometric support for generating high-quality 3D models.

[0111] For example, since the search space of the entire image is very large, an energy function is defined here and the PatchMatch algorithm is used for efficient solution:

[0112] in, Indicates the number of optimal views used for evaluation; Indicates smoothing weights; This indicates the discovery of a long spatial gradient penalty function. This will be selected during the view selection process. highest Each view participates in the calculation, thus obtaining the pixel depth value of each pixel in the image.

[0113] In some embodiments, to fuse data from multiple perspectives to obtain a 3D mesh model of the underwater concrete scene, refer to... Figure 9 , Figure 9 for Figure 1 The flowchart of sub-steps S801~S805 of step S800, wherein steps S801~S805 include: S801. Based on the relative pose, the 3D coordinates of the matched feature points, and the pixel depth value of each pixel, multi-view fusion is performed to generate a sparse 3D point cloud.

[0114] In this embodiment, the relative pose is derived from the camera rotation and translation parameters between adjacent keyframes calculated in step S500. The 3D coordinates of the matched feature points are obtained by the triangulation process in S600, and the pixel depth value of each pixel is derived by the dense depth estimation method in S700. During the execution of the above steps, the camera poses corresponding to all keyframe images are first unified to the global coordinate system to ensure that the observation data of each view has a consistent spatial reference benchmark. Specifically, a world coordinate system is established with the first frame as the origin, and the poses of subsequent frames are gradually added using an incremental SfM strategy, and the accumulated error is optimized by local bundle adjustment. Subsequently, the pixel depth map obtained by S700 in each keyframe is combined with the camera intrinsic parameters and pose corresponding to that frame, and each pixel is back-projected into 3D space to form a dense point cloud cluster; at the same time, the sparse feature points reconstructed in S600 are included in the fusion system as high-confidence control points. Furthermore, spatial registration and deduplication are performed on the 3D point sets from multiple perspectives, retaining common viewpoints within overlapping areas and removing outliers based on depth consistency principles. It should be noted that the sparse 3D point cloud here is defined as the initially integrated 3D point set, realizing the transformation from a single-frame depth map to a multi-view joint point cloud, constructing a preliminary 3D representation of the scene. In practical applications, this fusion process effectively improves point cloud coverage, especially demonstrating good detail reproduction capabilities in texture-rich areas. Based on the above description, this step provides a complete input data foundation for subsequent voxelization representation and surface reconstruction.

[0115] S802. Construct a voxel mesh for the sparse 3D point cloud, where each voxel corresponds to a pixel.

[0116] In this embodiment, the spatial region containing the sparse 3D point cloud generated in S801 is divided into a uniformly distributed 3D voxel grid. Each voxel represents a cubic volume unit, and its side length can be preset according to the reconstruction accuracy requirements (e.g., 1 cm or less). Specifically, firstly, the spatial bounding box of the entire point cloud is determined, and the grid is divided along the x, y, and z directions according to resolution, forming a discretized spatial index structure. Then, the mapping relationship between voxels and image pixels is established: for each pixel in each keyframe, if the ray formed by its back projection passes through a voxel, then the voxel is considered to be associated with that pixel. It should be noted that the correspondence between each voxel and one pixel here is not a one-to-one relationship, but rather means that in multi-view observation, each voxel can receive observation information from multiple pixels, that is, one voxel may correspond to pixels from multiple viewpoints. Furthermore, this voxel grid serves as the carrier of the symbolic distance function (SDF) to record the geometric distance from each spatial location to the surface. Thus, this step completes the transformation from disordered point cloud to structured spatial representation, providing topological support for subsequent symbolic distance calculation and weighted fusion. For example, when reconstructing the water-stop joint area of ​​a dam, a voxel mesh covering a detection path up to 30 meters long can accurately capture the spatial morphology of craters and abrasion zones. Based on the above description, this step, as a fundamental step in implicit surface modeling, achieves the dual functions of spatial discretization and observation-related modeling.

[0117] For example, the space is divided into a voxel grid, the signed distance function (SDF) from each voxel to the surface is calculated, and multi-view data is integrated through weighted fusion to construct a global signed distance field (TSDF) function. For voxels... and camera pose The effective distance along the ray direction is expressed as:

[0118] in, Indicates the position of the camera's optical center; Represents the direction of the line of sight and the vector The included angle; Indicates the depth along the ray direction.

[0119] S803. Based on the depth value of each pixel, calculate the symbolic distance of each voxel in multiple viewpoints. The symbolic distance represents the distance of the current voxel from the boundary of the sparse 3D point cloud.

[0120] In this embodiment, the signed distance is a key geometric quantity that measures whether a voxel is located inside or outside the surface of an object and its distance. The sign indicates inside or outside, and the absolute value represents the distance. During the above steps, for each voxel, all camera views that can observe it are traversed. Using the depth value of the corresponding pixel and the camera pose at that view, a ray is constructed along the line of sight, and the directed distance from the voxel's center to the nearest surface point is calculated. Specifically, if the voxel's center is in front of the surface (visible side), the signed distance is positive; if it is behind the surface (occluded side), it is negative; and if it is close to the surface, it approaches zero. This process is based on the TSDF (Truncated Signed Distance Function) model, considering only valid observations within the truncation distance range; distances outside this range are not included. It should be noted that the multiple views here include all keyframes that meet the visibility condition, i.e., the voxel is not occluded and is within the camera's field of view. Furthermore, the signed distance provided by each view is calculated independently, forming a set of candidate values ​​for subsequent fusion processing. Therefore, this step transforms two-dimensional depth information into a three-dimensional implicit representation through physical imaging relationships, enhancing the ability to describe complex geometric structures. In practical applications, due to slight vibrations and local blurring in the underwater environment, some observations may be biased. Therefore, a truncation mechanism is introduced to suppress the propagation of errors over long distances. Based on the above description, this step completes the crucial transformation from explicit point clouds to implicit range fields, laying the mathematical foundation for generating watertight surfaces.

[0121] S804. Remove the target symbolic distance from the multiple symbolic distances corresponding to each voxel, and obtain the effective distance corresponding to each voxel based on the remaining symbolic distances; among the differences between the target symbolic distance and the other symbolic distances corresponding to the voxel, the number of differences exceeding a preset value is greater than the preset number.

[0122] In this embodiment, outlier removal is performed on the multiple symbolic distance values ​​corresponding to each voxel obtained in S803 to improve the robustness of distance estimation. Specifically, for a set of symbolic distances received for a given voxel, the absolute value of the difference between each value and other values ​​is checked one by one. If the number of symbolic distances whose differences from other symbolic distances exceed a preset threshold (e.g., 2 cm) is greater than a preset number (e.g., more than half), then the value is determined to be an outlier, i.e., the target symbolic distance, and is removed. The remaining unremoved symbolic distances participate in a weighted average calculation, with the weights typically related to the observation angle, confidence level, or distance cutoff factor, ultimately yielding the effective distance for the voxel. It should be noted that the target symbolic distance here specifically refers to anomaly measurements that significantly deviate from the true surface due to occlusion, motion blur, mismatch, or noise interference. Their existence severely affects the smoothness and accuracy of the reconstruction results. Thus, this screening mechanism effectively suppresses the impact of poor-quality observations on the overall model. For example, when local deposits or air bubbles are present on a concrete surface, the surface location may be misjudged from certain perspectives, but the correct information can still be preserved through multi-view consistency comparison. Furthermore, this process demonstrates the core advantage of TSDF fusion: improving reconstruction reliability through redundant observations. Based on the above description, this step achieves refined purification of the implicit distance field, ensuring the quality and stability of subsequent isosurface extraction.

[0123] For example, to suppress the effects of noise and occlusion, these observations are truncated and weighted to obtain the global TSDF volume:

[0124] in, Indicates the cutoff distance; The weights are represented. The TSDF volume field is a discrete implicit representation and needs to be converted into an explicit triangular mesh for subsequent use.

[0125] S805. Generate a 3D mesh model of the underwater concrete scene surface based on the effective distance corresponding to each voxel.

[0126] In this embodiment, the effective distances corresponding to all voxels calculated in S804 are integrated into a globally consistent truncated signed distance field (TSDF Volume) as the implicit function representation of the 3D scene. Specifically, the Poisson reconstruction algorithm is used to process this TSDF field: First, according to the gradient divergence theorem, the surface reconstruction problem is transformed into solving the gradient field inversion problem of the indicator function, that is, finding a scalar function whose gradient direction is consistent with the observation normal and satisfies the Poisson equation. The partial differential equation is discretized on the voxel mesh using the finite difference method to obtain a continuous implicit surface representation. Then, the isosurface with a function value of 0.5, i.e., the position where the signed distance is zero, is extracted as the actual boundary of the object. Further, algorithms such as Marching Cubes or Dual Contouring are used to discretize the isosurface into a triangular mesh to generate a watertight, pore-free 3D surface model. It should be noted that the 3D mesh model here is defined as an explicit geometric expression composed of vertices, edges, and triangular faces, which can be directly used for visualization, measurement, and defect quantification analysis. Therefore, this step completes the final transformation from implicit field to explicit surface, achieving high-quality 3D modeling. In practical application, this method successfully reconstructed a 30-meter-long dam cutoff joint area, clearly presenting the concrete abrasion zone and crater structure, verifying its engineering practicality. Based on the above description, this step, as the endpoint of the entire reconstruction process, outputs standardized 3D digital assets that can be used for subsequent intelligent detection and maintenance decisions.

[0127] For example, the gradient field of the indicator function is transformed into a continuous surface through Poisson reconstruction, and the isosurface is extracted. Based on the Gaussian divergence theorem, the surface reconstruction problem is transformed into solving the Poisson equation:

[0128] in, Indicates an indicator function; This represents the divergence of the hairline field. Finally, extracting the isosurface with an indicator function value of 0.5 as the object surface will yield the reconstructed triangular network:

[0129] Received The table shows a watertight triangulation network, which can be directly printed as a 3D image. MVS can be summarized as starting from 2D pixels, obtaining 3D depth through geometric modeling and physical constraints, fusing multi-view data to construct a global geometric field, and finally transforming it into a usable 3D model.

[0130] In some embodiments, see Figure 10 , Figure 10 The 3D reconstructed and manually stitched images provided in the embodiments of the present invention Figure 10(a) is an image of the enhanced concrete image arranged in the order of key frame extraction based on SFM, and then the concrete is reconstructed in three dimensions using the above method. (b) is an image presented by manual stitching in previous work in order to intuitively present the damage of the concrete in the detection area.

[0131] For example, 2D images show concrete damage, but without depth data, this damage cannot be effectively quantified. Reconstruction provides a clear view of concrete erosion and scour, which is beneficial for subsequent concrete damage analysis. Since the underwater concrete of the dam has been subjected to prolonged water erosion, the entire area exhibits erosion; therefore, this study focuses on the scour of the concrete, such as... Figure 11 As shown, Figure 11 This is a schematic diagram of 3D reconstruction of a concrete crater measurement provided in an embodiment of the present invention. From... Figure 11 As can be seen, the concrete crater area is an irregular cube, so there are multiple three-dimensional measurement methods. Here, the minimum circumscribed cube measurement method is used to represent the three-dimensional dimensions of the crater. The measured dimensions of the crater in this area are 6.23cm*4.21cm*2.19cm.

[0132] For example, to verify that image enhancement is beneficial for reconstruction, the same data was used for verification, such as... Figure 12 As shown, Figure 12 The image enhancement effect comparison diagram is provided for the embodiments of the present invention. Figure 12 (a) is a direct reconstruction of video keyframes. Figure 12 (b) is the data after 3D reconstruction following image enhancement. The reconstructed point cloud data reveals missing points when directly reconstructing using video keyframe data. Figure 12 (b) The comparison also reveals that areas with indistinct features failed to participate in the reconstruction, resulting in incomplete reconstruction areas.

[0133] Finally, a complete reconstruction of an area where the dual-mode underwater robot operated underwater was performed, obtaining large-scale 3D data of a water-stop joint in the underwater concrete dam, such as... Figure 13 As shown, Figure 13 This is a large-scale 3D reconstruction image of an underwater concrete structure provided for an embodiment of the present invention. Figure 13 (a) represents the underwater trajectory of the dual-mode robot. The robot's trajectory follows the distribution of the underwater concrete sealing joints of the dam, so the entire movement process is approximately a straight line. Figure 13 (b) is a keyframe of the video. Figure 13 (c) is a large-scale 3D point cloud based on SFM technology. Figure 13 (d) shows a two-dimensional grid of a large-scale concrete structure.

[0134] For example, due to the differences in the performance and reconstruction principles of underwater robots, quantitative comparisons are difficult. Therefore, we adopted a qualitative comparison method, comparing the 3D reconstruction capabilities of underwater robots in several existing papers, as shown in Table 2. For instance, Yu et al. designed VNRF, accelerating rendering through voxel meshes and using label propagation to solve crack occlusion; Fan et al.'s method and our proposed method consider 3D reconstruction of large underwater scenes; Rho et al.'s method improves reconstruction efficiency by optimizing the AUV scanning path through multi-sonar fusion.

[0135] Table 2

[0136] Table 2 presents a qualitative comparison of the 3D reconstruction capabilities of different robots underwater, comparing them in four aspects: the size of the underwater reconstruction scene; reconstruction accuracy; dredging capability; and reconstruction efficiency. The method proposed by Fan et al. and the method proposed in this embodiment of the invention both consider 3D reconstruction of large underwater scenes; except for the method proposed by Rho et al., which has meter-level accuracy, the reconstruction accuracy of the other methods is millimeter-level; only the dual-modal robot provided in this embodiment of the invention possesses dredging capability. Therefore, through comparison, it can be found that the dual-modal underwater robot provided in this embodiment of the invention has a greater advantage in operating in complex underwater scenes.

[0137] Based on the above method, embodiments of the present invention also provide a system corresponding to the above method, such as... Figure 14 As shown, Figure 14 This is a functional module diagram of the underwater concrete scene 3D reconstruction system 1000 provided in an embodiment of the present invention. It should be noted that the basic principle and technical effects of the underwater concrete scene 3D reconstruction system 1000 provided in this embodiment are the same as those in the above method embodiments. For the sake of brevity, parts not mentioned in this embodiment can be referred to the corresponding content in the method embodiments. The underwater concrete scene 3D reconstruction system 1000 includes a data acquisition module 1100, an image enhancement module 1200, and a 3D reconstruction module 1300.

[0138] In this embodiment, the data acquisition module 1100 is used to acquire video data of the underwater concrete surface and obtain multiple keyframe images from the video data. It can be understood that the data acquisition module 1100 is used to perform the above-described step S100.

[0139] The image enhancement module 1200 is used to select multiple video images from the video data; perform background light stripping on each video image to obtain a background light stripped image; perform transmittance compensation on the background light stripped image to obtain a preliminary clear image; and perform adaptive color correction on the preliminary clear image to obtain an enhanced keyframe image. It can be understood that the image enhancement module 1200 is used to perform the above steps S201~S204.

[0140] The image enhancement module 1200 is also used to perform transmittance compensation on the background light stripping image to obtain a compensated image; and to perform regularization optimization on the compensated image to obtain a preliminary clear image.

[0141] The 3D reconstruction module 1300 is used to extract feature points from each keyframe image, where the feature points represent conspicuous parts of the underwater concrete surface; perform epipolar geometry comparison on the feature points in adjacent keyframe images to obtain feature matching point pairs in adjacent keyframe images, where each feature matching point pair contains the pixel coordinates of the same matched feature point in the adjacent keyframe images; calculate the relative pose of the adjacent keyframe images based on the pixel coordinates of the same matched feature point in the adjacent keyframe images, where the relative pose represents the rotation angle and translation vector of the camera corresponding to the adjacent keyframe images; calculate the 3D coordinates of the feature points based on the pixel coordinates of the same matched feature point in the adjacent keyframe images and the relative pose; calculate the pixel depth value of each pixel in the keyframe image based on the relative pose; and perform multi-view fusion based on the relative pose, the 3D coordinates of the matched feature points, and the pixel depth value of each pixel to generate a 3D mesh model of the underwater concrete scene. It can be understood that the 3D reconstruction module 1300 is used to perform the above steps S300~S800.

[0142] In some embodiments, the 3D reconstruction module 1300 is further configured to obtain the pixel coordinates of feature points in adjacent keyframe images; generate binary descriptors based on the pixel coordinates of the feature points; align initial feature matching point pairs based on the binary descriptors; and perform epipolar geometry alignment on the initial feature matching point pairs to obtain feature matching point pairs in adjacent keyframe images. It can be understood that the 3D reconstruction module 1300 is also configured to perform the above steps S401~S404.

[0143] In some embodiments, the 3D reconstruction module 1300 is further configured to establish epipolar geometric constraint equations and solve for the essential matrix based on the pixel coordinates of the same matched feature point in adjacent keyframe images and the preset parameters of the camera; and decompose the essential matrix to obtain the relative pose of adjacent keyframe images. It can be understood that the 3D reconstruction module 1300 is used to perform the above steps S501~S502.

[0144] In some embodiments, the 3D reconstruction module 1300 is further configured to calculate the depth values ​​of feature points in the keyframe image based on relative pose; assign a preset number of initial depth values ​​to each pixel in the keyframe image, wherein the initial depth values ​​are close to the depth values ​​of the feature points closest to the pixel; perform photometric consistency calculations on the preset number of initial depth values ​​and the depth values ​​of each target feature point corresponding to the pixel, thereby obtaining a score for each initial depth value relative to the depth value of each target feature point, and selecting the initial depth value with the highest score from all the initial depth values ​​corresponding to each target feature point as the target depth value; the target depth value corresponds one-to-one with the target feature point; the target feature point is a feature point in the keyframe image whose distance from the pixel is within a preset range; and determine the pixel depth value of the pixel based on the target depth values ​​corresponding to all the target feature points. It can be understood that the 3D reconstruction module 1300 is also configured to perform the above steps S701 to S704.

[0145] In some embodiments, the 3D reconstruction module 1300 is further configured to perform multi-view fusion based on relative pose, the 3D coordinates of matched feature points, and the pixel depth value of each pixel to generate a sparse 3D point cloud; construct a voxel mesh for the sparse 3D point cloud, where each voxel corresponds to a pixel; calculate the symbolic distance of each voxel in multiple views based on the depth value of each pixel, where the symbolic distance represents the distance of the current voxel from the boundary of the sparse 3D point cloud; remove the target symbolic distance from the multiple symbolic distances corresponding to each voxel, and obtain the effective distance corresponding to each voxel based on the remaining symbolic distances; the number of differences between the target symbolic distance and other symbolic distances corresponding to the voxel that exceed a preset value is greater than a preset number; and generate a 3D mesh model of the underwater concrete scene surface based on the effective distance corresponding to each voxel. It can be understood that the 3D reconstruction module 1300 is also configured to perform the above steps S801~S805.

[0146] Based on the same inventive concept disclosed above, the present invention also provides a block diagram of an electronic device 2000 performing the above method. Please refer to... Figure 15 , Figure 15 This is a block diagram of an electronic device 2000 provided in an embodiment of the present invention. The electronic device 2000 includes a processor 2100, a memory 2200, a bus 2300, and a communication interface 2400. The processor 2100 and the memory 2200 are connected via the bus 2300, and the processor 2100 communicates with external devices via the communication interface 2400.

[0147] Processor 2100 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed through integrated logic circuits in the hardware of processor 2100 or through software instructions. The processor 2100 may be a general-purpose processor 2100, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0148] The memory 2200 is used to store computer programs. For example, the underwater concrete scene three-dimensional reconstruction system 1000 in this embodiment of the invention includes at least one software function module that can be stored in the memory 2200 in the form of software or firmware. After receiving the execution instruction, the processor 2100 executes the program to implement the underwater concrete scene three-dimensional reconstruction method in this embodiment of the invention.

[0149] The memory 2200 may include high-speed random access memory (RAM) or non-volatile memory. Optionally, the memory 2200 may be a storage device built into the processor 2100 or a storage device independent of the processor 2100.

[0150] Bus 2300 can be ISA bus 2300, PCI bus 2300 or EISA bus 2300, etc. Figure 15 It is indicated by only one double-headed arrow, but does not mean that there is only one bus 2300 or one type of bus 2300.

[0151] Electronic devices 2000 can be mobile phones, tablets, laptops, desktop computers, and other computer devices.

[0152] Based on the same inventive concept, embodiments of the present invention also provide a computer-readable storage medium storing a computer program thereon. When executed by a processor 2100, this computer program implements the aforementioned method for three-dimensional reconstruction of an underwater concrete scene. This computer-readable storage medium may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0153] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for three-dimensional reconstruction of an underwater concrete scene, characterized in that, The method includes: Acquire video data of underwater concrete surfaces; Multiple keyframe images are obtained from the video data; Feature points are extracted from each of the keyframe images, and the feature points represent conspicuous parts of the underwater concrete surface. Perform epipolar geometry comparison on feature points in adjacent keyframe images to obtain feature matching point pairs in adjacent keyframe images. The feature matching point pairs contain the pixel coordinates of the same matched feature point in adjacent keyframe images. The relative pose of adjacent keyframe images is calculated based on the pixel coordinates of the same matched feature point in adjacent keyframe images. The relative pose represents the rotation angle and translation vector of the camera corresponding to the adjacent keyframe images. The three-dimensional coordinates of the matched feature point are calculated based on the pixel coordinates of the same matched feature point in adjacent keyframe images and the relative pose. Calculate the pixel depth value of each pixel in the keyframe image based on the relative pose; A three-dimensional mesh model of the underwater concrete scene is generated by multi-view fusion based on the relative pose, the three-dimensional coordinates of the matched feature points, and the pixel depth value of each pixel.

2. The method according to claim 1, characterized in that, The step of performing epipolar geometric comparison on feature points in adjacent keyframe images to obtain feature matching point pairs in adjacent keyframe images includes: Obtain the pixel coordinates of feature points in adjacent keyframe images; Generate a binary descriptor based on the pixel coordinates of the feature points; Initial feature matching point pairs are determined based on the binary descriptor; The initial feature matching point pairs are subjected to epipolar geometry alignment to obtain feature matching point pairs in adjacent keyframe images.

3. The method according to claim 1, characterized in that, The step of calculating the relative pose of adjacent keyframe images based on the pixel coordinates of the same matched feature point in adjacent keyframe images includes: Based on the pixel coordinates of the same matched feature point in adjacent keyframe images and the preset parameters of the camera, an epipolar geometric constraint equation is established and the essential matrix is ​​solved. The relative poses of adjacent keyframe images are obtained by decomposing the essential matrix.

4. The method according to claim 1, characterized in that, The step of calculating the pixel depth value of each pixel in the keyframe image based on the relative pose includes: Calculate the depth values ​​of feature points in the keyframe image based on the relative pose; Each pixel in the keyframe image is assigned a preset number of initial depth values, and the initial depth values ​​are close to the depth values ​​of the feature points closest to the pixel. The preset number of initial depth values ​​are respectively compared with the depth value of each target feature point corresponding to the pixel to calculate photometric consistency, so as to obtain the score of different initial depth values ​​relative to the depth value of each target feature point. The initial depth value with the highest score is selected from all the scores of the initial depth values ​​corresponding to each target feature point as the target depth value. The target depth value corresponds one-to-one with the target feature point. The target feature point is the feature point in the keyframe image whose distance from the pixel is within a preset range. The pixel depth value of the pixel is determined based on the target depth values ​​corresponding to all target feature points.

5. The method according to claim 1, characterized in that, The step of generating a 3D mesh model of the underwater concrete scene by performing multi-view fusion based on the relative pose, the 3D coordinates of the matched feature points, and the pixel depth value of each pixel includes: Multi-view fusion is performed based on the relative pose, the 3D coordinates of the matched feature points, and the pixel depth value of each pixel to generate a sparse 3D point cloud. A voxel mesh is constructed for the sparse 3D point cloud, wherein each voxel in the voxel mesh corresponds to a pixel. Based on the depth value of each pixel, the symbolic distance of each voxel in multiple viewpoints is calculated. The symbolic distance represents the distance of the current voxel from the boundary of the sparse 3D point cloud. The target symbolic distance is removed from the multiple symbolic distances corresponding to each voxel, and the effective distance corresponding to each voxel is obtained based on the remaining symbolic distances; among the differences between the target symbolic distance and the other symbolic distances corresponding to the voxel, the number of differences exceeding a preset value is greater than a preset number; A three-dimensional mesh model of the underwater concrete scene surface is generated based on the effective distance corresponding to each voxel.

6. The method according to claim 1, characterized in that, The step of obtaining multiple keyframe images from the video data includes: Select multiple video images from the video data; Perform background light stripping on each of the video images to obtain a background light stripped image; By performing castability compensation on the background light stripped image, a preliminary clear image is obtained; Adaptive color correction is performed on the initially clear image to obtain the enhanced keyframe image.

7. The method according to claim 6, characterized in that, The step of performing cast rate compensation on the background light stripped image to obtain a preliminary clear image includes: Transmittance compensation is performed on the background light stripping image to obtain a compensated image; The compensated image is then regularized to obtain a preliminary clear image.

8. A three-dimensional reconstruction system for underwater concrete scenes, characterized in that, The system includes: The data acquisition module is used to acquire video data of the underwater concrete surface and to obtain multiple keyframe images from the video data; A 3D reconstruction module is used to extract feature points from each keyframe image, where the feature points represent conspicuous parts of the underwater concrete surface; perform epipolar geometry comparison on the feature points in adjacent keyframe images to obtain feature matching point pairs in adjacent keyframe images, where each feature matching point pair contains the pixel coordinates of the same matched feature point in the adjacent keyframe images; calculate the relative pose of the adjacent keyframe images based on the pixel coordinates of the same matched feature point in the adjacent keyframe images, where the relative pose represents the rotation angle and translation vector of the camera corresponding to the adjacent keyframe images; calculate the 3D coordinates of the feature points based on the pixel coordinates of the same matched feature point in the adjacent keyframe images and the relative pose; calculate the pixel depth value of each pixel in the keyframe image based on the relative pose; and perform multi-view fusion based on the relative pose, the 3D coordinates of the matched feature points, and the pixel depth value of each pixel to generate a 3D mesh model of the underwater concrete scene.

9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a computer program that can be executed by the processor to implement the three-dimensional reconstruction method for underwater concrete scenes according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the three-dimensional reconstruction method for underwater concrete scenes as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Non-cooperative target three-dimensional space pose estimation method and system, equipment and medium

    CN118196188A

  • End-to-end three-dimensional reconstruction method and device based on unmanned aerial vehicle image

    CN118781280A

  • Three-dimensional reconstruction method and apparatus for monocular endoscope image, and terminal device

    WO2021115071A1

Cited By

  • AI-driven aerial suspension imaging method, device, equipment and medium

    CN121582482A

  • Mapping method for curved surface and irregular surface

    CN122090006A