Underwater concrete scene three-dimensional reconstruction method, system, device and storage medium
By extracting feature points from enhanced video images and performing 3D reconstruction, the problem of low accuracy and efficiency in 3D reconstruction of underwater concrete structures in complex environments has been solved, achieving high-precision and high-efficiency 3D reconstruction results.
Patent Information
- Application Number
- CN202511679316.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-11-17
AI Technical Summary
In underwater concrete structure 3D reconstruction, under conditions of turbid water and severe light absorption, existing technologies suffer from poor image quality, low reconstruction accuracy, and low efficiency.
By extracting feature points from the enhanced video images, calculating the camera's movement distance and translation vector, and combining pixel depth values for 3D reconstruction, a 3D mesh model is generated using epipolar geometry alignment and multi-view fusion.
It achieves high-precision and high-efficiency large-scale 3D reconstruction in complex underwater environments, improves the quality of underwater image acquisition, and saves the time required to build 3D point cloud data.
Smart Images

Figure CN121120996B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of underwater three-dimensional reconstruction, in particular to a method, system and device for three-dimensional reconstruction of underwater concrete scenes and a storage medium. BACKGROUND
[0002] The long-term safety and durability of underwater structures of hydropower hubs such as concrete gravity dams and arch dams are the core of protecting the safety of life and property downstream and the stability of the power system. They are long-term in complex water environments, subject to the coupling of multiple factors such as hydraulic load, temperature change, freeze-thaw cycle, and are prone to defects such as cracks, spalling, and steel corrosion.
[0003] In recent years, through the deep integration of underwater robot (ROV / AUV) technology and high-precision optical sensing technology, underwater robots are equipped with high-definition camera equipment, high-brightness lighting systems and laser scanners to carry out close-range, high-stability data collection underwater, and simultaneously obtain high-definition images and three-dimensional point cloud information of the structure surface, and then realize accurate quantitative analysis and detection of concrete defects.
[0004] However, due to the turbidity of the water body and the serious absorption of light in the underwater environment, the optical images obtained by shooting will have problems such as color distortion, blurring and glare, which makes the quality of the underwater images collected poor, and it takes a lot of time to construct three-dimensional point cloud data, resulting in low accuracy and efficiency of three-dimensional reconstruction of large-scale underwater scenes. SUMMARY
[0005] Therefore, the present application aims to provide a method, system and device for three-dimensional reconstruction of underwater concrete scenes and a storage medium, which obtains feature points from enhanced video images, calculates the moving distance and translation vector of the shooting camera based on the feature points, and then performs three-dimensional reconstruction combined with the pixel depth value in each key frame image, realizes large-scale, high-precision and efficient underwater three-dimensional reconstruction, improves the quality of underwater image acquisition in the case of turbidity of the water body and serious absorption of light in the underwater environment, and saves the construction of three-dimensional point cloud data, thereby improving the accuracy and efficiency of three-dimensional reconstruction of large-scale underwater scenes.
[0006] To achieve the above object, in a first aspect, embodiments of the present application provide a method for three-dimensional reconstruction of an underwater concrete scene, the method comprising: obtaining video data of an underwater concrete surface; obtaining a plurality of key frame images from the video data; extracting feature points in each of the key frame images, the feature points representing salient parts of the underwater concrete surface; performing epipolar geometry comparison on the feature points in adjacent key frame images to obtain a feature matching point pair in the adjacent key frame images, the feature matching point pair comprising pixel coordinates of a same matched feature point in the adjacent key frame images; calculating a relative pose of the adjacent key frame images based on the pixel coordinates of the same matched feature point in the adjacent key frame images, the relative pose representing a rotation angle and a translation vector of a camera corresponding to the adjacent key frame images; calculating a three-dimensional coordinate of the matched feature point based on the pixel coordinates of the same matched feature point in the adjacent key frame images and the relative pose; calculating a pixel depth value of each pixel point in the key frame images based on the relative pose; and performing multi-view fusion based on the relative pose, the three-dimensional coordinate of the matched feature point, and the pixel depth value of each pixel point to generate a three-dimensional mesh model of the underwater concrete scene.
[0007] In the present embodiment, a plurality of key frame images are obtained from the video data, the plurality of key frame images representing pictures taken from a plurality of perspectives, and feature points of salient parts of the underwater concrete surface are extracted from each key frame image. Then, a matching feature point pair in adjacent two key frame images is compared, so as to calculate a relative pose based on pixel coordinates of the matching feature point pair in the key frame images, and further obtain a rotation angle and a translation vector of a camera corresponding to the adjacent key frame images. Then, a three-dimensional coordinate of the matched feature point can be calculated based on the relative pose and the pixel coordinates of the feature points, so as to calculate a pixel depth in each key frame image and combine the three-dimensional coordinate of the matched feature point to perform multi-view fusion, thereby generating a three-dimensional mesh model of the underwater concrete scene. In this way, by obtaining feature points from enhanced video images, calculating a moving distance and a translation vector of a camera based on the feature points, and further combining a pixel depth value in each key frame image to perform three-dimensional reconstruction, three-dimensional reconstruction of an underwater scene with a large scene, high precision, and high efficiency is achieved, which can improve the quality of underwater image acquisition in a turbid water environment and serious light absorption, and save the construction of three-dimensional point cloud data, thereby improving the precision and efficiency of three-dimensional reconstruction of a large-scale underwater scene.
[0008] In some embodiments, the epipolar geometry comparison of the feature points in the adjacent key frame images comprises: obtaining pixel coordinates of the feature points in the adjacent key frame images; generating binary descriptors according to the pixel coordinates of the feature points; comparing the binary descriptors to obtain initial feature matching point pairs; and performing epipolar geometry comparison on the initial feature matching point pairs to obtain the feature matching point pairs in the adjacent key frame images.
[0009] In this way, the feature points in the adjacent key frame images are compared by epipolar geometry to compare the feature point pairs in the adjacent key frame images that actually belong to the same feature point, so as to reduce the number of feature point calculations and provide accurate feature point data for subsequent camera pose estimation and three-dimensional reconstruction.
[0010] In some embodiments, the relative poses of the adjacent key frame images are calculated according to the pixel coordinates of the same matched feature points in the adjacent key frame images, which comprises: establishing epipolar geometry constraint equations and solving essential matrices according to the pixel coordinates of the same matched feature points in the adjacent key frame images and the preset parameters of the shooting camera; and decomposing the essential matrices to obtain the relative poses of the adjacent key frame images.
[0011] In this way, when the shooting camera travels along a straight line trajectory, there is a stable relative motion relationship between the front and rear frames, and accurate capture of the small attitude changes can achieve accurate estimation of the relative poses of the camera and enhance the geometric consistency of three-dimensional reconstruction through strict geometric modeling and analytical solving process, even in low-texture areas.
[0012] In some embodiments, the pixel depth value of each pixel point in the key frame image is calculated according to the relative pose, which comprises: calculating the depth values of the feature points in the key frame image according to the relative pose; assigning a preset number of initial depth values to each pixel point in the key frame image, the initial depth values being close to the depth values of the nearest feature points; performing photometric consistency calculation on the preset number of initial depth values and the depth values of each target feature point corresponding to the pixel point to obtain scores of different initial depth values relative to the depth values of each target feature point, and selecting the initial depth value with the highest score from the scores of all initial depth values corresponding to each target feature point as a target depth value; the target depth value corresponds to the target feature point one-to-one; the target feature point is a feature point in the key frame image that is within a preset range from the pixel point; and the pixel depth value of the pixel point is determined according to the target depth values of all target feature points.
[0013] In this way, by virtue of the initial depth value, the local plane assumption and the photometric consistency constraint are realized, which helps to accurately restore the concave shape of the smooth surface in the reconstructed underwater concrete pit area, effectively solves the depth missing problem of the non-feature area, and significantly improves the integrity and detail restoration capability of the reconstruction result.
[0014] In some embodiments, the multi-view fusion according to the relative pose, the three-dimensional coordinates of the matched feature points, and the pixel depth value of each pixel point to generate a three-dimensional grid model of the underwater concrete scene includes: multi-view fusion according to the relative pose, the three-dimensional coordinates of the matched feature points, and the pixel depth value of each pixel point to generate a sparse three-dimensional point cloud; constructing a voxel grid for the sparse three-dimensional point cloud, each voxel in the voxel grid corresponding to a pixel point; based on the depth value of each pixel point, calculating the signed distance of each voxel under multiple views, the signed distance representing the distance of the current voxel from the boundary of the sparse three-dimensional point cloud; removing the target signed distance from the multiple signed distances corresponding to each voxel, and obtaining the effective distance corresponding to each voxel according to the remaining signed distances; the number of difference values exceeding the preset value in the difference values between the target signed distance and the other signed distances corresponding to the voxel is greater than the preset number; generating a three-dimensional grid model of the surface of the underwater concrete scene based on the effective distance corresponding to each voxel.
[0015] In this way, by solving the gradient field of the divergence field corresponding to all voxels, a watertight triangular mesh model is generated, which realizes high-quality three-dimensional modeling from discrete observation to continuous entity through a rigorous voxel grid fusion and surface reconstruction mechanism, and significantly improves the topological integrity and availability of the reconstruction result.
[0016] In some embodiments, the obtaining a plurality of key frame images from the video data includes: selecting a plurality of video images from the video data; performing background light stripping on each video image to obtain a background light stripped image; performing projection rate compensation on the background light stripped image to obtain a preliminary clear image; and performing adaptive color correction on the preliminary clear image to obtain an image-enhanced key frame image.
[0017] In this way, through the enhanced process of background light stripping and projection rate compensation, the image quality is significantly improved, providing high-quality input for subsequent three-dimensional reconstruction and enhancing the applicability in high turbidity environments.
[0018] In some embodiments, the step of performing transmittance compensation on the background light stripped image to obtain a preliminary clear image includes: performing transmittance compensation on the background light stripped image to obtain a compensated image; and performing regularization optimization on the compensated image to obtain a preliminary clear image.
[0019] This configuration, through a combination of transmittance compensation and regularization optimization, improves the accuracy and stability of transmittance estimation.
[0020] Secondly, embodiments of the present invention provide a three-dimensional reconstruction system for an underwater concrete scene. The system includes: a data acquisition module for acquiring video data of the underwater concrete surface and acquiring multiple keyframe images from the video data; a three-dimensional reconstruction module for extracting feature points from each keyframe image, wherein the feature points represent conspicuous parts of the underwater concrete surface; performing epipolar geometric comparison on feature points in adjacent keyframe images to obtain feature matching point pairs in adjacent keyframe images, wherein the feature matching point pairs include the pixel coordinates of the same matched feature point in the adjacent keyframe images; and performing a three-dimensional reconstruction based on the same matched feature point. The relative pose of adjacent keyframe images is calculated based on the pixel coordinates of the same matched feature point in the adjacent keyframe images. The relative pose represents the rotation angle and translation vector of the camera corresponding to the adjacent keyframe images. The three-dimensional coordinates of the feature point are calculated based on the pixel coordinates of the same matched feature point in the adjacent keyframe images and the relative pose. The pixel depth value of each pixel in the keyframe image is calculated based on the relative pose. Multi-view fusion is performed based on the relative pose, the three-dimensional coordinates of the matched feature point, and the pixel depth value of each pixel to generate a three-dimensional mesh model of the underwater concrete scene.
[0021] Thirdly, embodiments of the present invention provide an electronic device, including a processor and a memory, wherein the memory stores a computer program executable by the processor, and the processor can execute the computer program to implement the underwater concrete scene three-dimensional reconstruction method as described in the first aspect.
[0022] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the three-dimensional reconstruction method for an underwater concrete scene as described in the first aspect.
[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be regarded as a limitation on the scope, and for those of ordinary skill in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0025] Figure 1 A flow chart of a three-dimensional reconstruction process of an underwater concrete scene provided by the embodiment of the present application;
[0026] Figure 2 A contrast diagram of an original image and an enhanced image provided by the embodiment of the present application;
[0027] Figure 3 A contrast effect diagram of different image enhancement algorithms provided by the embodiment of the present application;
[0028] Figure 4 A flow chart of sub-steps S201-S204 of step S200 in the embodiment of the present application; Figure 1 A flow chart of sub-steps S401-S404 of step S400 in the embodiment of the present application;
[0029] Figure 5 A key frame extraction image and a robot collected image provided by the embodiment of the present application;
[0030] Figure 6 A flow chart of sub-steps S501-S502 of step S500 in the embodiment of the present application; Figure 1 A flow chart of sub-steps S401-S404 of step S400 in the embodiment of the present application;
[0031] Figure 7 A flow chart of sub-steps S701-S704 of step S700 in the embodiment of the present application; Figure 1 A flow chart of sub-steps S501-S502 of step S500 in the embodiment of the present application;
[0032] Figure 8 A flow chart of sub-steps S801-S805 of step S800 in the embodiment of the present application; Figure 1 A flow chart of sub-steps S701-S704 of step S700 in the embodiment of the present application;
[0033] Figure 9 A flow chart of sub-steps S801-S805 of step S800 in the embodiment of the present application; Figure 1 A flow chart of sub-steps S801-S805 of step S800 in the embodiment of the present application;
[0034] Figure 10 A 3D reconstruction and a manually spliced image provided by the embodiment of the present application;
[0035] Figure 11 A concrete pit measurement schematic diagram after 3D reconstruction provided by the embodiment of the present application;
[0036] Figure 12 A contrast effect diagram of image enhancement provided by the embodiment of the present application;
[0037] Figure 13 A large-scale 3D reconstruction map of an underwater concrete structure provided by an embodiment of the present application;
[0038] Figure 14 A functional module schematic diagram of an underwater concrete scene 3D reconstruction system provided by an embodiment of the present application;
[0039] Figure 15 A block schematic diagram of an electronic device provided by an embodiment of the present application.
[0040] Icon: 1000-underwater concrete scene 3D reconstruction system; 1100-data acquisition module; 1200-image enhancement module; 1300-3D reconstruction module; 2000-electronic device; 2100-processor; 2200-memory; 2300-bus; 2400-communication interface. DETAILED DESCRIPTION
[0041] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. The components of the embodiments of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations.
[0042] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0043] It should be noted that the relational terms such as "first" and "second" and the like are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that these entities or operations have any such actual relationship or order. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the processes, methods, articles or devices including the element.
[0044] The deep integration of underwater robot (ROV / AUV) technology and high-precision optical sensing technology promotes the underwater structure detection means of the hydropower hub to achieve a qualitative leap. By carrying high-definition camera equipment, high-brightness lighting systems and laser scanners, underwater robots can carry out close-range, high-stability data collection underwater, synchronously obtain high-definition images and three-dimensional point cloud information of the structure surface, and then realize accurate quantitative analysis and digital recording of concrete defects, laying a key technical foundation for the construction of the subsequent intelligent detection system of underwater structure.
[0045] The application of underwater robots to underwater concrete structure damage detection has made the latest progress, effectively improving the intelligent detection level in this field. Although the method has been successfully applied in underwater concrete detection, it still faces significant challenges under complex working conditions: first, biological silt and sediments are often attached to the concrete surface, obscuring the true structure morphology; second, underwater optical images are prone to color distortion, blurring and glare; third, three-dimensional reconstruction of large-scale underwater scenes needs to balance accuracy and efficiency, and global positioning and quantification of defects are still problems to be solved.
[0046] As described in the background, due to the turbidity of the underwater environment and the serious absorption of light, the optical images obtained by shooting will have problems such as color distortion, blurring and glare, resulting in poor quality of the collected underwater images, which requires a lot of time to construct three-dimensional point cloud data, thus leading to low accuracy and efficiency of three-dimensional reconstruction of large-scale underwater scenes.
[0047] Therefore, the embodiment of the present application provides a dual-mode underwater robot, which carries a dredging function and a clear water observation component. The biological silt and sediments attached to the concrete surface are cleaned by the components of the dredging function, which helps the clear water observation component to observe the relatively clear and true concrete surface structure morphology. The embodiment of the present application also provides an underwater concrete scene three-dimensional reconstruction method, which is described in detail below. Figure 1 , Figure 1 The underwater concrete scene three-dimensional reconstruction flowchart provided by the embodiment of the present application obtains feature points from the enhanced video images, calculates the moving distance and translation vector of the shooting camera based on the feature points, and then combines the pixel depth value in each key frame image to perform three-dimensional reconstruction, realizing large-scale, high-precision and high-efficiency underwater three-dimensional reconstruction. Under the condition of turbidity of the underwater environment and serious absorption of light, the quality of underwater image acquisition can be improved, and the construction of three-dimensional point cloud data can be saved, thereby improving the accuracy and efficiency of three-dimensional reconstruction of large-scale underwater scenes. The underwater concrete scene three-dimensional reconstruction method further includes steps S100-S800:
[0048] S100, obtain video data of the underwater concrete surface.
[0049] In this embodiment, the collection of video data relies on the vision system mounted on the dual-mode underwater robot, which has the functions of dredging and clean water observation components, and can adapt to harsh underwater environments such as high turbidity, low light and complex water flow. In the process of acquiring video data, first, the biological attachments and sediments covering the surface of the underwater concrete structure are removed by the dredging component of the robot to expose the true structure form; then the clean water observation component is lowered to adhere to the concrete surface, effectively reducing the scattering and absorption of light propagation by water, thereby improving the imaging quality. Specifically, the vision system continuously collects high-definition video streams during the robot's travel, with an image resolution of 1920x1080 and a frame rate of 30FPS. The camera used is LBF-C50HD2, and the raw video data is transmitted in real time to the water surface base station through the buoyancy cable. Further, the camera intrinsic parameters have been pre-calibrated, and the lens distortion has been corrected to ensure the accuracy of subsequent geometric calculations. In actual application, the robot moves at a speed of 20 cm / s along the dam sealing joint area, ensuring that the video sequence has good spatiotemporal continuity and sufficient parallax coverage, and realizes stable and reliable acquisition of raw video data for three-dimensional reconstruction in complex engineering scenarios.
[0050] Exemplarily, refer to Figure 2 , Figure 2 The contrast chart of the raw image and the enhanced image provided by the embodiment of the present application. The raw image Figure 2 (a) has obvious blur and low contrast, which is not conducive to feature matching; the enhanced image Figure 2 (b) has more image details, which plays a crucial role in feature matching. To quantify the quality of image enhancement, a transverse contrast experiment is introduced to compare the proposed method with two existing underwater image enhancement algorithms, and four image quality evaluation indexes are introduced, including Underwater Color Image Quality Evaluation (UCIQE), Underwater Image Quality Measure (UIQM), Natural Image Quality Evaluator (NIQE), and Peak Signal-to-Noise Ratio (PSNR) to evaluate the enhanced image. The comparison results of different image enhancement algorithms are shown in Figure 3 , Figure 3The contrast effect diagram of different image enhancement algorithms provided by the embodiment of the present application is shown in the figure. In the figure, Original image refers to the original image; Ref[x] refers to the image enhancement method proposed by Lin et al., L2uwe[x] refers to an underwater image enhancement algorithm proposed by Marques and Albu CVPR W 2020, and Proposed Method refers to the image enhancement method proposed by the embodiment of the present application. From Figure 3 The image enhancement method of the embodiment of the present application can demonstrate the advantages in visual effect, such as more natural color recovery, higher texture clarity and better contrast.
[0051] Further, since the underwater concrete image data belongs to the no-reference image quality evaluation, three no-reference image quality evaluation methods are adopted here. In addition, in order to further ensure the rigor of the experiment, we take the image captured by the robot when it is stationary in water as the reference image (the sand in water has naturally settled) as the reference image of the PSNR index, and the evaluation results of the enhanced images of different algorithms are shown in Table 1.
[0052] Table 1
[0053]
[0054] In Table 1, Original image represents the original image, Ref[x] our method represents the method of the embodiment of the present application, Ref[x] represents the reference method, and Propose method represents the method proposed by the embodiment of the present application. As can be seen, the smaller the NIQE value, the better the image quality, and the larger the other evaluation indexes, the better the image quality. From the statistical results of the data, it can be seen that the scores of the image enhancement method based on wavelength compensation and depth optimization proposed by us in the four evaluation indexes are 0.49, 3.56, 17.58 and 20.88 respectively, and the best score is obtained in the whole evaluation. It is worth noting that the PSNR image evaluation is only used as a reference, because in actual engineering application, we cannot wait for the sand in water to naturally settle after collecting data at each image acquisition point.
[0055] S200, acquire a plurality of key frame images from video data.
[0056] In the embodiment, the video data is a raw image stream continuously collected by an LBF-C50HD2 high-definition camera mounted on a dual-mode underwater robot during underwater operation. The resolution is 1920x1080, the frame rate is 30 FPS, and the dynamic observation process of the dam underwater concrete surface after dredging treatment is recorded. During the execution of the above steps, first, the whole video data is subjected to time series analysis, and a plurality of frames are extracted as candidate key frames according to a preset time interval strategy. The specific sampling period is to select one frame every 0.5 seconds, so as to ensure that the adjacent key frames have appropriate parallax coverage in the process of the robot moving at a speed of 20 cm / s, and at the same time, avoid redundant calculation due to too high frame overlap. It should be noted that the key frame image is defined as a representative static image frame selected from the original video sequence, which has independent spatial observation information and can be used for subsequent three-dimensional reconstruction processing. The selection principle takes into account the time uniformity and the degree of motion change, which not only ensures the completeness of scene coverage, but also meets the basic requirements of SFM algorithm for parallax input. Further, considering the existence of slight shaking and motion blur in underwater environment, the selected key frames are all from the video segments of the robot posture stable period, excluding the frames obviously blurred due to rapid movement or mechanical vibration, realizing the transformation from continuous dynamic video to discrete static image, and providing a clear structure and time sequence coherent basic data unit for subsequent feature extraction and matching. For example, in actual engineering application, by extracting key frames from a video stream lasting for several minutes, hundreds of uniformly distributed key frame images are finally obtained, effectively supporting the large-scale three-dimensional reconstruction task of a 30-meter long water stop joint area, not only reducing the total amount of data processing, but also guaranteeing the rationality and effectiveness of the input image in spatial distribution and time sequence.
[0057] S300, extracting feature points in each key frame image, the feature points representing salient parts of the underwater concrete surface.
[0058] In the embodiment, the feature points are defined as the pixel positions with significant local changes in the image, which have good repeatability and description stability, and are suitable for cross-view matching. The FAST corner detection algorithm is used to process each enhanced key frame image to extract the feature points. The algorithm identifies the corner points by judging whether the gray level difference between the center pixel and the sampling points on the circumference exceeds the set threshold, and has the advantages of fast calculation speed and strong robustness, and is particularly suitable for underwater operation scenarios with high real-time requirements. Specifically, the input is the enhanced gray image, and the output is a set of two-dimensional pixel coordinates, each coordinate corresponding to a potential corner position. Further, to improve the rotation invariance, a set of binary descriptors is generated in combination with the descriptor (rotation-aware Binary Robust Independent Elementary Features, rBRIEF), which is used for similarity measurement in the subsequent matching stage. It should be noted that the conspicuous positions specifically include the crack edges, aggregate boundaries, construction joints, and textured areas such as concave-convex structures formed by erosion on the concrete surface. These areas are represented as positions with sharp gradient changes or high local contrast in the image. As can be seen, the extraction quality of the feature points directly affects the subsequent matching success rate and three-dimensional reconstruction accuracy. In actual application, although the overall texture of the underwater concrete surface is relatively simple, the local details are restored after image enhancement, so that the FAST algorithm can still stably extract a sufficient number of effective feature points. Based on the above description, this step completes the task of locating key structural information from two-dimensional images, providing a basic unit for establishing multi-view geometric correlation.
[0059] S400, the feature points in the adjacent key frame images are compared by epipolar geometry, to obtain feature matching point pairs in the adjacent key frame images, and the feature matching point pairs contain pixel coordinates of the same matched feature points in the adjacent key frame images.
[0060] In the embodiment, first, the feature points extracted from the adjacent key frame images and their corresponding rBRIEF binary descriptors are obtained, then the similarity between different descriptors is measured based on the Hamming distance, and a matching threshold is set to screen out the preliminary candidate matching point pairs to form an initial feature matching point pair set. However, relying only on the descriptor similarity may lead to a large number of false matches, especially in the case of severe changes in illumination or the presence of repeated textures. Therefore, a geometric consistency verification mechanism must be introduced. Specifically, the epipolar geometric constraint relationship expressed by the fundamental matrix is used to verify the initial matching results: for a matching point pair (x a ,x b ), if x b Fx aIf the Sampson distance is less than a preset threshold, the matching point pair is retained; otherwise, it is rejected. Further, the fundamental matrix F is robustly estimated by the RANSAC algorithm, which can accurately fit the epipolar constraint model even in the presence of many outliers. It should be noted that the same matched feature point herein refers to the imaging position of a fixed physical point in space under two adjacent viewing angles, and the two are projected corresponding relationship through camera motion. As can be seen, this step realizes the screening process from potential matching to geometrically consistent matching, significantly improving the reliability of the matching result. For example, in the underwater environment of the dam, due to the disturbance of water flow and the slight shaking of the robot, there is slight blur and displacement between images, but this method can still maintain a high matching accuracy. Based on the above description, this step provides a high-confidence corresponding point set for camera pose estimation, which is a key link connecting two-dimensional observation and three-dimensional structure.
[0061] S500, calculate the relative pose of the adjacent key frame images according to the pixel coordinates of the same matched feature points in the adjacent key frame images, the relative pose representing the rotation angle and translation vector of the cameras corresponding to the adjacent key frame images.
[0062] In this embodiment, the camera intrinsic matrix K has been obtained through pre-dating and is used as a known parameter in the calculation, while the relative pose between two adjacent views cannot be directly measured and must be indirectly solved through matching points. During the execution of the above step, first, the pixel coordinates of the feature matching points are converted into normalized image coordinates, i.e. x = K -1 p; then based on the epipolar geometry constraint equation x b T Ex a = 0 to construct a solving model of the essential matrix E, where E = [t] x R, [t] x is the anti-symmetric matrix of the translation vector t, and R is the rotation matrix. Specifically, the eight-point method or the multi-point least squares method is used to estimate the fundamental matrix F, and then the essential matrix E is derived from E = K -T FK. After singular value decomposition (SVD) of E, four possible (R, t) solution combinations can be obtained. Further, by triangulating the matching points and judging whether the depth of the three-dimensional coordinates is positive, a physically reasonable constraint is determined, and the correct relative pose solution is uniquely determined. As can be seen, this method makes full use of the geometric characteristics of rigid body motion and avoids ambiguity in the solution space. In actual application, when the dual-mode underwater robot moves slowly along a straight line, there is a stable relative motion relationship between the front and rear frames, and this method can accurately capture the attitude change, and can maintain a certain estimation ability even in low-texture areas. Based on the above description, this step realizes accurate estimation of the camera motion state, providing reliable motion parameter support for three-dimensional structure recovery.
[0063] S600, calculate the three-dimensional coordinates of the matched feature points based on the pixel coordinates and relative poses of the same matched feature points in adjacent key frame images.
[0064] In this embodiment, the three-dimensional coordinates of the space points are solved by using the known camera relative poses and the pixel coordinates of the matched point pairs by using the triangulation method. Specifically, the projection matrices P1 and P2 of the two cameras and the corresponding image points x a and x b are substituted into the perspective projection equation λx = PX to construct an overdetermined equation set and solve the optimal three-dimensional point X by minimizing the re-projection error. Further, considering the existence of measurement noise, the DLT (Direct Linear Transformation) or SVD method is usually used for numerical solution to ensure the stability and reliability of the results. It should be noted that the matched feature points here refer to reliable matched point pairs verified by epipolar geometry, and the calculation of their three-dimensional coordinates depends on the accurate relative poses between the two views. As can be seen, this process realizes the mapping conversion from two-dimensional image observation to three-dimensional space coordinates, and is the core step of generating sparse point clouds. For example, when reconstructing the concrete pit area, multiple corner points form local point cloud clusters after triangulation, preliminarily outlining the spatial profile of the defect. Further, with the addition of more views, the sparse point cloud size can be continuously expanded by incremental SFM, and error accumulation can be suppressed by local bundle adjustment. Based on the above description, this step completes the task of constructing the initial three-dimensional structure, providing a geometric skeleton support for subsequent dense reconstruction.
[0065] S700, calculate the pixel depth value of each pixel point in the key frame image according to the relative pose.
[0066] In the embodiment, only a few feature points have three-dimensional coordinates, and most pixel points do not participate in matching, so the depth information of each pixel needs to be derived through a densification strategy. During the execution of the above steps, first, based on the relative pose and the position of a pixel point in the reference image, it is assumed that it is located on a certain depth plane, and it is projected into other adjacent views using homographic transformation; then the photometric consistency between the projected region and the actual image block is evaluated, and the commonly used index is zero-mean normalized cross correlation (ZNCC), and the closer the value is to 1, the higher the consistency. Specifically, the PatchMatch algorithm is used to efficiently search for the best depth hypothesis, and the depth estimation of each pixel is iteratively optimized within the depth search space. Further, to improve robustness, only a few views with the highest consistency score are selected to participate in fusion calculation, and a smoothing regularization term is introduced to suppress noise and occlusion effects. It should be noted that the pixel depth value here is defined as the distance from the camera optical center to the corresponding spatial point of the pixel along the line of sight, and its calculation depends on the consistency assumption and local plane model of multi-view observation. As can be seen, this method realizes the expansion from sparse three-dimensional points to dense depth map. In actual application, even in the texture-poor area, as long as there is enough parallax and light consistency, reasonable depth information can still be recovered, which provides pixel-by-pixel depth support for generating complete surface geometry, and is a key link to realize high-resolution reconstruction.
[0067] S800, multi-view fusion is performed according to the relative pose, the three-dimensional coordinates of the matched feature points, and the pixel depth value of each pixel point, to generate a three-dimensional grid model of the underwater concrete scene.
[0068] In the embodiment, firstly, the relative poses of all key frames, sparse three-dimensional point coordinates, and pixel depth maps of each frame are taken as inputs for multi-view fusion processing. Specifically, the three-dimensional space is divided into a uniform voxel grid, the signed distance from each voxel to the surface along the line of sight of each camera is calculated, and the observation results from different perspectives are integrated by weighted average to form a truncated signed distance field (TSDF). Further, to suppress the influence of noise and outliers, multiple signed distances received by each voxel are screened: if the number of distances whose difference from other distances exceeds a predetermined threshold is greater than a predetermined number, the distance is considered to be an outlier and is removed, and the remaining distances are used to calculate the effective distance. It should be noted that the multi-view fusion here refers to unifying multiple local observations in a global coordinate system to construct a consistent implicit surface representation. As can be seen, this process effectively reduces the uncertainty caused by single-view errors and occlusions. Further, based on the constructed TSDF volume field, the gradient field inversion problem of the indicator function is solved using the Poisson reconstruction algorithm to convert the implicit function to an explicit geometry; that is, the isosurface of the indicator function value of 0.5 is extracted as the object surface, and a watertight triangular mesh model is finally generated. For example, in engineering applications, this method successfully reconstructs a 30-meter-long dam water stop joint area, clearly presents the erosion zone and scour pit structure, completes high-quality three-dimensional modeling from discrete observations to continuous entities, significantly improves the topological integrity and usability of the reconstruction results, and meets the needs of subsequent damage quantification analysis.
[0069] In some embodiments, the obtained key frame images need to be subjected to image enhancement processing, referring to Figure 4 , Figure 4 For Figure 1 the flowchart of sub-steps S201-S204 of step S200, steps S201-S204 include:
[0070] S201, selecting a plurality of video images from video data.
[0071] In this embodiment, the video data is the original image stream continuously collected by the LBF-C50HD2 high-definition camera carried by the dual-mode underwater robot during the detection process, which contains a sequence of frames that are continuous in time but have high redundancy. During the execution of the above steps, first, the entire video is time-sequentially traversed, and a set of representative video images is selected as the basis input for subsequent processing according to the preset time interval strategy or motion change threshold. Specifically, a fixed time interval sampling method is used, and a frame image is extracted every 0.5 seconds to ensure that the selected images are uniformly distributed in space and meet the requirements of the incremental SfM algorithm for the parallax range. It should be noted that the video image here is defined as a single frame in the original video stream, which has not been subjected to any enhancement or correction processing and retains the degradation characteristics of the underwater imaging process, such as color bias, low contrast, and background light interference. Further, the selection process prioritizes the exclusion of blurred frames caused by robot startup, turning, or water flow disturbance to ensure that the candidate images have basic clarity and stability. As can be seen, this step achieves the purpose of extracting effective observation frames from high-redundancy video streams, providing a complete structure and computationally efficient image set for subsequent image enhancement and three-dimensional reconstruction. In actual application, hundreds of effective images can be extracted from minutes of continuous video through this method, supporting three-dimensional reconstruction tasks in a large scene range, and directly affecting the quality and efficiency of subsequent processing as the starting link of the image enhancement process.
[0072] For example, when the robot works underwater, the video data has motion blur and image blur caused by suspended particles, as shown in Figure 5 , Figure 5 The key frame extraction images provided by the embodiment of the present application and the images collected by the robot, Figure 5 (a) is the key frame extracted from the video, which shows image blur; Figure 5 (b) is a high-definition image (sharp image) collected under the condition of a stationary robot. Although the quality of the key frame image extracted from the video is relatively low, it is beneficial to reduce the detection time in the detection engineering task, and in addition, the video data has temporal continuity and can therefore be used for 3D reconstruction based on the SFM method. In the process of underwater imaging by an optical camera, the water absorbs and scatters light. According to the Jaffe-McGlamery underwater optical theory, underwater image degradation is mainly caused by exponential attenuation of object reflected light during propagation; and backscattered light caused by suspended particles in water. These two processes cause color distortion, contrast reduction, and blurring in the image. According to the transmission principle of light underwater, the underwater imaging physical model can be described as:
[0073]
[0074] wherein, represents the degraded image collected by the camera, a color channel representing an image; representing an image to be enhanced; representing a wavelength-dependent transmittance; is an attenuation coefficient, is an absorption coefficient, representing a scattering coefficient; is a global background light, i.e. a backscattering component; representing a scene depth. Different wavelengths have different degrees of attenuation under water, satisfying .
[0075] S202, background light stripping is performed on each video image to obtain a background light stripped image.
[0076] In this embodiment, for each video image selected in S201, a background light component estimation and removal operation is performed to eliminate the influence of backscattering light caused by suspended particles in water on the imaging quality. Specifically, based on the Jaffe-McGlamery underwater optical imaging model, the degraded image is decomposed into the sum of an object reflection light term and a background scattering light term, where the background light is mainly derived from the non-target light scattering of the water in front of the camera. To accurately estimate the background light, a "brightness-gradient" double constraint strategy is adopted: first, identify the pixel set whose brightness value is in the top 2% of the image, and select the region whose gradient amplitude is lower than the adaptive threshold from the set to avoid interference caused by direct light area, mirror reflection area and local high-brightness materials of concrete; then, statistical analysis is performed on the pixel gray value in the set, and the mean or weighted average thereof is taken as the estimated value of the global background light intensity. Further, the original image is subtracted by the background light component to obtain a background light stripped image which only retains the object surface reflection information. It should be noted that the background light stripping defined here is a kind of image preprocessing technology based on physical prior, aiming to restore the true scene radiation and improve the image contrast and detail visibility. As can be seen, this step effectively alleviates the image fogging phenomenon caused by scattering in a high turbidity environment. For example, in the experiment, the original image shows obvious whitening and low contrast, as shown in Figure 5 (a). After background light stripping, the concrete surface cracks and aggregate boundaries become clearer, as shown in Figure 5 (b). Based on the above description, this step as the first stage of the image enhancement link provides a purer input basis for subsequent transmittance estimation and color restoration.
[0077] Exemplarily, the background light is derived from the physical properties of water scattering. In order to avoid the interference of high-brightness objects (light source, partial mirror reflection of concrete material), gradient constraint is adopted here to exclude the interference of high-brightness objects to ensure the extraction of real background scattering light, for example, a brightness-gradient joint constraint method is used to estimate the background light, represented as:
[0078]
[0079] wherein, is the pixel set of the first 2% of the luminance, a grayscale image, is the 98th quantile of the luminance cumulative distribution; is the gradient magnitude; is the adaptive threshold. The main purpose is to only keep the pixels with small gradient values (these pixels have a gentle change in brightness, which meets the uniformity of the background light), and completely filter out the high-light interference with large gradient. Finally, the median of these filtered pure background light candidates is taken as the final background light, which can represent the true brightness of the water scattering and avoid the influence of a few residual outliers. In this way, the interference bright spots and the real background light are completely separated, which clears the obstacles for subsequent image enhancement.
[0080] S203, transmittance compensation is performed on the background light stripped image to obtain a preliminary clear image.
[0081] In this embodiment, the background light stripped image is subjected to transmittance compensation to obtain a compensated image; the compensated image is subjected to regularization optimization to obtain a preliminary clear image.
[0082] Further, the transmittance refers to the proportion of light that is not absorbed or scattered during the propagation from the light source to the target surface and back to the camera, and its spatial distribution directly determines the clarity of the image. During the execution of the above steps, first, the image stripped of the background light is preliminarily estimated for transmittance based on the improved dark channel prior method combined with a wavelength compensation factor, the dark channel of the image is extracted using local minimum filtering, and a wavelength attenuation ratio is introduced to adjust the transmittance difference of different color channels, so that the red light channel is restored to be more consistent with the actual attenuation law. Further, to overcome the block artifacts (Block Artifacts) and spatial discontinuity problems caused by the traditional dark channel prior, the transmittance estimation is converted into a depth optimization problem, the dark channel prior method is introduced in the image transmittance estimation process, and the wavelength characteristics are combined to improve the transmittance estimation model, which can be expressed as:
[0083]
[0084] wherein, denotes a local neighborhood; is a depth preservation factor; represents the wavelength compensation ratio. In underwater images, except for a few high-light areas such as light sources and reflections, most areas have dark pixels with low brightness (such as the shadows of concrete cracks and dark parts of water bodies). These dark pixels are mainly caused by water scattering. The dark channel prior is to estimate the degree of light attenuation caused by scattering by finding the locally darkest pixels, which is equivalent to providing a basic reference value for the transmittance. Generally, in underwater environments, red light decays the fastest, green light is second, and blue light is the slowest. This difference will cause the image to be severely color cast (such as red deficiency and blue bias). Therefore, by adding the wavelength characteristics (represented by the formula ) to the dark channel prior, a certain correction can be made for different colors: for example, if the red light attenuation is strong, the transmittance estimate value of the red light can be adjusted lower to make the result consistent with the actual attenuation of the red light; if the blue light attenuation is weak, it can be adjusted appropriately to avoid excessive correction. It can be understood that the dark channel prior is responsible for the general direction (finding the general impact of scattering), and the wavelength characteristics are responsible for calibration (correcting the deviation according to the color difference). The combination of the two can calculate a transmittance that is more consistent with the actual situation underwater. The transmittance can then be used to remove scattered light and supplement missing colors, which can make the blurred and color cast underwater images clearer and more natural.
[0085] Further, the transmittance map generated by the dark channel prior may have block artifacts. Here, the physical constraint relationship of Beer-Lambert's law is introduced to convert the transmittance estimation into a depth optimization problem, and the depth optimization model is represented as:
[0086]
[0087] wherein, is the scene depth to be optimized; is a smoothing weight; is a depth gradient. The depth map obtained is physically consistent, and the transmittance can be calculated back, which makes the red channel transmittance more consistent with the actual attenuation; the target edge has no block mutation; and the depth step continuous region can remain sharp.
[0088] Further, in order to overcome the noise amplification problem and the ill-posed problem (when , the equation is unstable to solve) in the scene reconstruction process, a variational optimization framework is introduced for regularization modeling:
[0089]
[0090] wherein, represents the optimized transmittance; represents a non-convex TV norm. To balance quality and efficiency, a closed-form approximate solution is used to represent the clear image :
[0091]
[0092] wherein, represents a critical threshold, when directly using the physical inverse transform, when regularization protection is enabled to avoid noise explosion. By regularization modeling, a repaired image that not only fits the original image information but also is natural and clear can be found, relying on two parts of joint constraints: the first part is , which can be understood as calculating the fitness of the repaired image and the original image, to avoid the repaired image looking clear but not matching the actual scene (such as repairing the concrete crack into another texture); the second part is the penalty term , which functions to limit the sudden change in the image. If the brightness of a certain region jumps too sharply (not a real concrete edge, crack, but noise or a messy color block), it will be deducted by this penalty term.
[0093] S204, performing adaptive color correction on the preliminary clear image to obtain an image-enhanced key frame image.
[0094] In this embodiment, in view of the color distortion problem existing in the preliminary clear image output by S203, a two-channel adaptive color correction operation is performed to restore the color distribution under natural visual perception. Specifically, first, the statistical characteristics of each color channel (R, G, B) are analyzed, including the mean and standard deviation, to identify the channel imbalance phenomenon caused by selective absorption of water (usually manifested as excessive blue channel and serious attenuation of red channel). Then, brightness compensation processing is performed to adjust the gain coefficients of each channel to align the mean values and eliminate overall color deviation; on this basis, dynamic contrast adjustment is implemented to restore saturation and color levels according to the standard deviation and attenuation difference coefficient of each channel. It should be noted that adaptive color correction is defined as an automatic white balance and color restoration method based on channel statistical characteristics, and its parameters are dynamically adjusted according to the local features of each image, rather than using a fixed transformation matrix. The blue-green color deviation commonly existing in underwater imaging is effectively corrected, making the concrete body color closer to the observation effect in the air. For example, in the enhanced image (as shown in Figure 5 (b), the originally blue concrete surface presents a grayish tone, and the cracks and sediments are more distinct, which is conducive to subsequent feature extraction and manual interpretation. Further, the image obtained after this step is the final image-enhanced key frame image, which will be used as the input of the SFM algorithm for three-dimensional reconstruction. Based on the above description, this step completes the last modification link of the image enhancement process, ensuring that the output image meets the standards required for high-quality reconstruction in terms of brightness, contrast, and color.
[0095] Further, in order to correct the distortion of the image color space, an adaptive color compensation model is designed according to the color channel statistical distribution imbalance characteristics:
[0096]
[0097] wherein, represents the reconstructed input image channel; represents the channel mean value; represents the channel standard deviation; represents the attenuation difference coefficient.
[0098] It can be understood that the divergence of red, green and blue light is not the same in the underwater environment; red light does not diverge far and is easily absorbed by water, while blue light can diverge far, resulting in the photos taken being often blue and lacking red, and the color being far from the actual scene, which is color distortion. Moreover, the color deviation degree of each underwater photo is different: some have turbid points and the color deviation is more serious; some take close-up shots and the color is slightly better, and it is difficult to solve with a fixed color adjustment parameter. The adaptive color compensation model calculates the average brightness of the red, green and blue color channels respectively, then looks at the color distribution of each channel, and then combines the attenuation law of different colors under water according to the two data to make the brightness and distribution of the three colors return to balance. The originally deviated image looks more natural, and it is also convenient to see the details such as cracks and textures of concrete.
[0099] In some embodiments, in order to accurately find the matching feature points in adjacent key frame images, refer to Figure 6 , Figure 6 is a flowchart of sub-steps S401-S404 of step S400 in Figure 1 , steps S401-S404 include:
[0100] S401, acquire the pixel coordinates of the feature points in the adjacent key frame image.
[0101] In the embodiment, the pixel coordinates of the feature points refer to the two-dimensional image positions with significant local gray level changes extracted from the enhanced key frame images by the FAST corner detection algorithm, which constitute the basic input for the subsequent descriptor generation and matching. During the execution of the above steps, first, the feature detection operation is performed on the adjacent two key frame images to identify the candidate corner point set on the image plane of each image. Specifically, a number of sampling points are selected on the circumference around the pixel as the center, and if the gray values of the continuous multiple points are all higher or lower than the center pixel plus a certain threshold, it is determined that the position is a corner point. The output result is a set of two-dimensional coordinate sets, each coordinate corresponding to the accurate position of a detected feature point in the image. It should be noted that the adjacent key frame images herein specifically refer to two video frames that are continuous in time sequence or have sufficient parallax, and there is an estimable relative motion relationship between them, which is suitable for establishing geometric correspondence. Further, the obtained pixel coordinates are not normalized and are directly represented based on the original image resolution (1920x1080) for subsequent descriptor calculation and matching verification, which completes the positioning task of the feature space position and provides basic data units for constructing the cross-view correspondence relationship. In actual application, although the texture of the underwater concrete surface is relatively single, the edge and aggregate structure are clear after image enhancement, so that the FAST algorithm can still stably extract a sufficient number of effective feature points.
[0102] Exemplarily, the FAST algorithm is used to extract feature points by local brightness contrast for image corner detection, which is used for fast positioning of stable feature points in underwater concrete images:
[0103]
[0104] wherein, represents the input i-th frame image; represents the FAST corner detection algorithm, and the output is a set of coordinates of key pixels on the image, which accurately correspond to stable feature points (such as the corners of cracks, convex edges on the concrete surface, edge intersection points, etc.) with significant characteristics in the underwater concrete image, and each coordinate marks the specific position of the feature point in the i-th frame image; .
[0105] S402, generating a binary descriptor according to the pixel coordinates of the feature points.
[0106] In the embodiment, for each feature point with an obtained pixel coordinate, a local neighborhood region is defined around it, and the above represents the rotation-aware BRIEF description calculation, which is generated by gray comparison, and the output is a set of binary descriptors , wherein ; and a binary descriptor is generated by the rBRIEF (rotation-aware Binary Robust Independent Elementary Features) algorithm. Specifically, the method forms a binary string composed of "0" and "1" by comparing the gray value at a preset position within a local image block, for example: if the gray value of the first point is greater than that of the second point in a pair of sampling points, it is recorded as "1", otherwise it is recorded as "0". Further, to improve the rotation invariance, the dominant direction of the feature point is estimated according to the gradient distribution around the feature point, and then the sampling template is rotated and aligned according to the direction, so as to ensure that the same feature under different poses has similar descriptors. It should be noted that the binary descriptor here is defined as a fixed-length bit string (such as 256 bits), which has the advantages of high storage efficiency and fast matching speed, and is particularly suitable for resource-limited underwater robot platforms. It converts local image information into quantifiable digital features, and enhances the robustness of feature expression. For example, under uneven lighting or slight blur conditions, even if the pixel value fluctuates slightly, as long as the relative gray relationship remains stable, the descriptor can still maintain consistency. Based on the above description, this step realizes the conversion from spatial position to semantic feature, and lays a technical foundation for efficient matching.
[0107] S403、Based on the binary descriptor, the initial feature matching point pair is matched.
[0108] In this embodiment, the initial feature matching aims to quickly screen out potential corresponding feature point pairs as a candidate set for subsequent geometric verification. During the execution of the above step, the binary descriptor of each feature point in the current frame is compared with the descriptors of all feature points in the next frame one by one, and the Hamming distance is used to measure the difference between them. Specifically, the Hamming distance is equal to the number of bits with different values in the same bit of two equal-length binary strings, and the smaller the distance, the higher the similarity. Set a matching threshold, when the Hamming distance of a pair of descriptors is less than the threshold and is the best match for the current point, it is recorded as an initial matching point pair. Further, to improve the matching accuracy, a proportion test can be introduced, that is, only when the distance ratio of the best match to the second best match is lower than the set proportion, the match is accepted, so as to exclude ambiguous candidates. It should be noted that the initial feature matching point pair has not been verified for geometric consistency, and may contain a large number of false matches, especially in repetitive texture or low contrast areas.
[0109] Illustratively, the feature points in different images need to establish a corresponding relationship to support three-dimensional reconstruction, and the feature points extracted from the front and rear frame images need to be further matched, and the point pairs with similar descriptors in the two images are found. The feature matching set can be represented as:
[0110]
[0111] wherein, , is the image with key points; with denotes the binary descriptor corresponding to the key point pair of the two consecutive images; denotes the Hamming distance function for computing the difference of binary strings; denotes the matching threshold.
[0112] For example, in the case where the step of feature matching is extracting feature points, generating descriptors, matching feature points, and geometric verification, image A takes a "crack corner A" of a piece of concrete underwater, and generates a descriptor "01011001…"; image B takes a "crack corner B" of the same place (slightly shifted), and generates a descriptor "01011010…". The Hamming distance between the two is 1 (only 1 bit is different), which is less than the set threshold 3, so the pair of (corner A, corner B) is put into the corresponding point pair list M a If the descriptor of another protrusion C in image B is "10100111…", and the Hamming distance with corner A is 10, which is greater than the threshold, it is determined that it is not a corresponding point.
[0113] S404, perform epipolar geometry comparison on the initial feature matching point pairs to obtain feature matching point pairs in adjacent key frame images.
[0114] In this embodiment, the initial feature matching point pairs obtained in S403 are verified for geometric consistency using epipolar geometric constraints, and false matches that do not conform to the epipolar relationship are removed. Specifically, first, estimate the fundamental matrix F based on the matching point set, and use the RANSAC (Random Sample Consensus) algorithm to robustly fit the model in the presence of a large number of outliers: each time, randomly select a minimum sample set (such as the eight-point method) to calculate F, and then count the number of points that satisfy x b Fx aThe number of inner points of < ε (Sampson distance) is iterated multiple times, and the F with the most supporters is retained as the optimal solution. Subsequently, all initial matching point pairs are substituted into the basis matrix for verification, and only the point pairs that meet the epipolar constraint are retained as the final feature matching point pairs. It should be noted that the epipolar geometry comparison here refers to the geometric constraint established according to the camera projection relationship between two views, which requires the projections of the same spatial point in two images to be located on the corresponding epipolar line, thereby excluding false correspondences caused by descriptor similarity, effectively suppressing the false matching phenomenon caused by changes in lighting, texture repetition or noise interference, and significantly improving the reliability of the matching result. For example, in the underwater environment of a dam, regular construction joints or aggregate distribution often exist on the concrete surface, which can easily cause confusion at the descriptor level, and the epipolar geometric verification can accurately retain the true corresponding points. Further, the screened feature matching point pairs will be used for essential matrix solving and camera pose estimation, directly affecting the geometric accuracy of three-dimensional reconstruction.
[0115] Exemplarily, false matching may occur during the matching process, so the false matching is removed through the epipolar geometric relationship between the front and rear images, and the geometric verification relationship of the feature point matching can be expressed as:
[0116]
[0117] wherein, and represent the second coordinates of the key points; represents the basis matrix, and there is an epipolar geometric constraint relationship ; represents the Huber robust loss function. Finally, the Sampson distance is introduced to evaluate the matching quality:
[0118]
[0119] False matching may be introduced by descriptor similarity, and the false matching is removed by using the epipolar geometric constraint, and finally the geometrically consistent matching point pairs are output.
[0120] For example, the core logic of eliminating false matches by epipolar geometry constraint is that the projection points of the same real space point in two images must follow the coplanar geometry rule of light rays and the baseline. The matching pair that does not meet this rule is false, and the purpose is to leave the matching point pair that is truly geometrically consistent. Assuming that in the underwater detection of a dam, a feature point A (the corner of a crack) is extracted from the previous frame image, and through descriptor similarity comparison, two candidate matching points B (the corner of the same crack) and C (the protrusion of another unrelated crack) are found in the next frame image. At this time, the elimination of false matches by epipolar geometry constraint can first calculate the fundamental matrix F (which contains the geometric relationship of camera pose, intrinsic parameters, etc. of two frames of images, and is the core rule carrier of epipolar geometry) through a small number of correct matching points (such as obvious crack endpoints marked by hand) in the two images. Then, according to the point A of the previous frame and the fundamental matrix F, the corresponding epipolar line l2 in the next frame is calculated. Then, it is judged whether the candidate points B and C fall on the epipolar line l2: point B (the real matching point): just falls on the epipolar line l2, meets the geometric rule of “light coplanar”, and is kept as a correct matching pair; point C (false matching point): deviates from the epipolar line l2 far away (such as more than 5 pixels), violates the physical law, and is determined as a false match and eliminated.
[0121] In some embodiments, in order to clearly know the angle and translation vector of the shooting camera in the process of shooting the adjacent two key frame images, refer to Figure 7 , Figure 7 For Figure 1 the step flowchart of sub-steps S501-S502 of step S500, steps S501-S502 include:
[0122] S501, according to the pixel coordinates of the same matched feature point in the adjacent key frame images and the preset parameters of the shooting camera, an epipolar geometry constraint equation is established and an essential matrix is solved.
[0123] In this embodiment, the preset parameters of the shooting camera refer to the camera intrinsic matrix obtained through pre-calibration, which contains focal length, principal point coordinates, pixel size and other information, and belongs to known prior conditions. During the execution of the above steps, first, the pixel coordinates of the feature matching points pair verified by the epipolar geometry in S404 are converted from the image coordinate system to the normalized image plane to eliminate the influence of lens distortion and imaging scale. Specifically, based on the epipolar geometry theory, for the matching point pair in the two images and satisfy the epipolar geometry constraint equation: ; wherein, and indicate that the intrinsic matrix of the camera can be obtained by calibration; indicates the essential matrix, is the skew-symmetric matrix of the translation vector, is a rotation matrix. That is, substituting all the matched point pairs into this equation constructs an over-determined linear equation group, which is solved by the eight-point method or the multi-point least squares method, and the preliminary estimated result is regularized by singular value decomposition (SVD) to meet the internal constraint that the two non-zero singular values are equal. It should be noted that the epipolar geometry constraint equation defined here is the geometric dependence relationship of the projection points between the two views caused by the camera motion, and its mathematical form strictly depends on the projective invariance under the rigid body transformation. As can be seen, the establishment and solving process of this equation realizes the transition from two-dimensional matching observation to three-dimensional motion structure, completes the core modeling task of the geometric relationship between cameras, and provides necessary input for subsequent pose decomposition.
[0124] S502, decompose the essential matrix to obtain the relative pose of the adjacent key frame image.
[0125] In this embodiment, the singular value decomposition SVD of the essential matrix E can be expressed as E = Udiag(1, 1, 0) T Because the rotation matrix R has two rotation directions, and the translation vector t has two directions, the combination of the two naturally produces four theoretical pose solutions, each of which is composed of a rotation matrix R and a unit length translation vector t, and the unique solution can be determined by the physical constraint that the depth of the triangular point is positive: ; wherein, represents the third column of , and the solution needs to satisfy . The camera motion pose transformation can be solved by the essential matrix of epipolar geometry, but this can only determine the relative pose and sparse 3D point cloud, and it is difficult to directly obtain the absolute scale of the scene. Here, the minimum re-projection error is established by triangulation, which uses the geometric consistency of multi-view observation to calculate the depth value of the space point through the light ray intersection principle, thereby converting the relative pose into a 3D structure with physical scale. The three-dimensional coordinates optimized by triangulation can be expressed as:
[0126]
[0127] wherein, represents the estimated three-dimensional point; represents a perspective projection function; represents a camera projection matrix. The three-dimensional point is the initial three-dimensional coordinate point of the key feature of the underwater concrete surface calculated by triangulation, and the initial solution of the three-dimensional coordinate point has accumulated error, so when doing global optimization, it is necessary to minimize the re-projection error, and the camera pose and three-dimensional coordinate point are optimized jointly to establish a minimum re-projection error optimization function by bundle adjustment (BA):
[0128]
[0129] wherein, represents a pose of the camera; represents a covariance matrix of the observation noise. SFM is to solve the structure and motion by decomposing the motion parameters from epipolar geometry constraints, and then jointly solving the structure and motion by triangulation and non-linear optimization. Its accuracy depends on the robustness of feature matching and the convergence of bundle adjustment, and the scale uncertainty needs to be solved by introducing known size objects or sensor fusion.
[0130] It should be noted that the relative pose here is defined as the rotation angle and translation vector of the next key frame relative to the previous key frame, which fully describes the rigid motion state of the camera in the continuous shooting process. As can be seen, this decomposition process fully utilizes the algebraic structure of the essential matrix and the physical constraints of three-dimensional space, and excludes solutions that are mathematically correct but not achievable in reality. For example, when a dual-mode underwater robot travels at a constant speed along the surface of a dam, its motion direction is mainly the forward vector, and the corresponding translation direction should be roughly consistent with the heading direction of the robot. This prior helps to further verify the reasonableness of the solution. Further, the obtained relative pose is used as the initial estimate of the incremental SFM for the pose solving and global bundle adjustment optimization of the newly added view, realizing the analytical reduction from the essential matrix to the specific motion parameters, and providing an accurate attitude reference for three-dimensional structure recovery.
[0131] In some embodiments, to calculate the depth value of each pixel point in the key frame image, refer to Figure 8 , Figure 8 for Figure 1 the flowchart of sub-steps S701-S704 of step S700, steps S701-S704 include:
[0132] S701, calculating the depth value of the feature point in the key frame image according to the relative pose.
[0133] In this embodiment, relative pose refers to the rotation matrix and translation vector of the camera between adjacent keyframes, obtained through step S500, which provides the necessary geometric relationship for triangulation. For feature point pairs that have been matched in S400, the coordinates of the feature point in 3D space are recovered using triangulation methods, utilizing their pixel coordinates in two or more views and the corresponding camera projection matrix. Specifically, the normalized coordinates of the matching points are substituted into the perspective projection equation to construct a system of linear equations about the 3D points, and the optimal solution is obtained by minimizing the reprojection error. Commonly used algorithms include DLT (Direct Linear Transformation) or numerical optimization methods based on SVD. Furthermore, the obtained 3D coordinates are used to calculate the depth value of the feature point relative to the reference frame camera, i.e., the distance of the spatial point along the camera's principal axis, by inversely calculating the camera's optical center position and viewing direction. It should be noted that the depth value of the feature point here is defined as the Euclidean distance between the optical center of the shooting camera and its corresponding physical surface point. This is the initial seed data for subsequent dense depth estimation, realizing the transformation from sparse matching to 3D structural information and providing a reliable benchmark for depth propagation in non-feature regions. In practical applications, underwater imaging is subject to slight blurring and noise interference, which may introduce errors into the triangulation results. Therefore, iterative optimization using local bundle adjustment is typically employed to improve depth accuracy. Based on the above description, this step lays the geometric foundation for pixel-level depth estimation, ensuring spatial consistency in subsequent dense reconstruction.
[0134] For example, suppose a pixel The neighborhood of lies on a small plane Above, the plane is formed by the normal. and depth Definition. Projecting pixel blocks from a reference image onto other views using a homography matrix:
[0135]
[0136] in, and These represent the camera's first... Each intrinsic parameter matrix and the reference camera matrix; Indicates the estimation from the reference to the first The rotation matrix and translation vector of each camera. The main task is to calculate the depth value for each pixel (pixels without paired feature points) to form a 2.5D depth map. Each concrete area in the keyframe image now has information about its distance from the camera. For example, the depth values of all pixels between A and B are between 1.2 and 1.25 meters, and the transition is smooth.
[0137] S702, assign a preset number of initial depth values to each pixel point in the key frame image, and the initial depth value is close to the depth value of the nearest feature point to the pixel point.
[0138] In the embodiment, for the ordinary pixel points in the key frame image which do not participate in feature matching, a set of candidate depth hypotheses is set for them as the initial range of subsequent optimization search. Specifically, first, a plurality of feature points which are within a preset neighborhood range in spatial distance from the current pixel point are determined, and these feature points have obtained reliable depth values through S701; then, a plurality of nearest depth values are selected as references, and a preset number of initial depth values are generated around the depth values, for example, N depth hypotheses are sampled in a fixed step within a range of ±Δd. It should be noted that the preset number is usually 8-16, which not only ensures that the search space is covered sufficiently, but also avoids high computational complexity; and the nearest feature point not only refers to the one with the smallest Euclidean pixel distance, but also includes adjacent feature points on the same local plane structure, so as to meet the prior assumption of local surface continuity, effectively reduce the search space of subsequent optimization, and improve the algorithm efficiency. For example, in the area where the concrete surface is relatively flat, the depth of the adjacent feature points has high correlation, and the initial depth set generated on this basis is more likely to contain the true depth value. Further, the process establishes a depth hypothesis set for each pixel, which is used to drive the PatchMatch type algorithm for efficient iterative optimization, completes the preliminary expansion from sparse depth to dense depth, and provides a reasonable starting point for realizing pixel-by-pixel depth estimation.
[0139] S703, perform photometric consistency calculation on the preset number of initial depth values and the depth value of each target feature point corresponding to the pixel point respectively, obtain the scores of different initial depth values relative to the depth value of each target feature point, and select the initial depth value with the highest score from the scores of all initial depth values corresponding to each target feature point as the target depth value; the target depth value corresponds to the target feature point one by one; the target feature point is a feature point in the key frame image which is within a preset range in distance from the pixel point.
[0140] In this embodiment, photometric consistency refers to the requirement that the grayscale or color of the same physical surface point should remain similar under different viewing angles. This is the core assumption of multi-view stereo matching. During the execution of the above steps, for each initial depth value, it is regarded as the assumed depth of the current pixel. The image block containing the pixel is projected into other adjacent views using the relative pose obtained in S500, forming a reprojection region. Subsequently, the similarity score between the reference image block and the projected image block is calculated. The commonly used index is zero-mean normalized cross-correlation (ZNCC), and the closer the value is to 1, the higher the consistency. Specifically, for each target feature point (i.e., a known depth feature point located in the neighborhood of the current pixel), the projection consistency of all initial depth values under multiple views is evaluated, and the initial depth value with the highest score is recorded as the target depth value corresponding to the target feature point. It should be noted that the target feature point here is limited to feature points whose distance from the current pixel on the image plane is less than a certain threshold. The purpose is to establish local structural associations and avoid introducing misleading information from distant irrelevant points. Furthermore, due to potential uneven lighting and color shifts in the underwater environment, ZNCC typically performs mean normalization on image patches before calculation to enhance robustness. Thus, this step verifies the depth hypothesis through multi-view observations, selecting the one that best conforms to the physical laws of imaging. In practical applications, even if consistency decreases in a single view due to glare or occlusion, the correct match can still be preserved through multi-view fusion. Based on the above description, this step achieves refined screening of the initial depth hypothesis, improving the accuracy of local depth estimation.
[0141] For example, since the underwater concrete reconstruction surface reflects light diffusely, the same physical point should have a similar color from multiple viewpoints. Therefore, cross-correlation is normalized by zero mean. Assessing projection consistency:
[0142]
[0143] in, and Representing the reference image and the first Grayscale values of each view; and This represents the average grayscale value of the reference image block and the projected image block; and This represents the standard deviation of gray levels between the reference image patch and the projected image patch; Indicated by A neighborhood window centered on the center; A value closer to 1 indicates a more accurate depth / hairline hypothesis; a value closer to 0 indicates occlusion or failure to meet the Lambert hypothesis. By continuously adjusting the preset depth value, the optimal depth value for each pixel can be quickly found through consistency verification.
[0144] S704. Determine the pixel depth value of each pixel based on the target depth values corresponding to all target feature points.
[0145] In this embodiment, the target depth values selected by each target feature point in S703 are aggregated and comprehensively judged to determine the final pixel depth value of the current pixel. Specifically, multiple target depth values are fused using methods such as weighted averaging, median filtering, or multi-model fitting: if the target depth values are concentrated, their mean or median is taken as the final result; if there is significant dispersion, weighted fusion is performed based on the visibility weight or consistency score of the corresponding view to suppress the influence of outliers. Furthermore, to improve surface smoothness and geometric continuity, a regularization term can be introduced to constrain the depth gradient changes between adjacent pixels to prevent drastic jumps. It should be noted that the pixel depth value here is defined as the scene depth corresponding to each pixel in the keyframe image, constituting the basic input for subsequent multi-view fusion to generate the symbolic distance field, effectively integrating depth cues from multiple neighboring feature points, and enhancing the stability and completeness of the estimation results. For example, in the edge area of a concrete crater, although some pixels lack direct matching information, a reasonable depth transition can still be recovered through the collaborative guidance of multiple surrounding feature points. Furthermore, the obtained pixel depth map will serve as the core output of the MVS stage, used for TSDF voxel fusion and 3D mesh reconstruction. Based on the above description, this step completes the integration task from locally optimal depth to a globally consistent depth map, providing complete geometric support for generating high-quality 3D models.
[0146] For example, since the search space of the entire image is very large, an energy function is defined here and the PatchMatch algorithm is used for efficient solution:
[0147]
[0148] in, Indicates the number of optimal views used for evaluation; Indicates smoothing weights; This indicates the discovery of a long spatial gradient penalty function. This will be selected during the view selection process. highest Each view participates in the calculation, thus obtaining the pixel depth value of each pixel in the image.
[0149] In some embodiments, to fuse data from multiple perspectives to obtain a 3D mesh model of the underwater concrete scene, refer to... Figure 9 , Figure 9 for Figure 1 The flowchart of sub-steps S801~S805 of step S800, wherein steps S801~S805 include:
[0150] S801, multi-view fusion is performed according to the relative pose, the three-dimensional coordinates of the matched feature points, and the pixel depth value of each pixel point, and a sparse three-dimensional point cloud is generated.
[0151] In the embodiment, the relative pose is derived from the camera rotation and translation parameters between adjacent key frames calculated in step S500, the three-dimensional coordinates of the matched feature points are obtained by the triangulation process in S600, and the pixel depth value of each pixel point is derived by the dense depth estimation method of S700. During the execution of the above steps, first, the camera poses corresponding to all key frame images are unified to the global coordinate system to ensure that the observation data of each view has a consistent spatial reference basis. Specifically, a world coordinate system is established with the first frame as the origin, the poses of subsequent frames are gradually added using an incremental SfM strategy, and the accumulated error is optimized by local bundle adjustment. Subsequently, the pixel depth map obtained by S700 is combined with the camera intrinsic parameters and the pose corresponding to each key frame, each pixel point is back projected to the three-dimensional space, and a dense point cloud cluster is formed; at the same time, the sparse feature points reconstructed in S600 are included in the fusion system as high-confidence control points. Further, the three-dimensional point sets from multiple views are spatially registered and processed to remove duplicates, the co-visible points in the overlapping area are retained, and the abnormal points are removed according to the depth consistency principle. It should be noted that the sparse three-dimensional point cloud defined here is the preliminary integrated three-dimensional point set, which realizes the conversion from a single frame depth map to a multi-view joint point cloud and constructs a preliminary three-dimensional representation of the scene. In actual application, the fusion process effectively improves the point cloud coverage, especially in texture-rich areas, showing good detail restoration capability. Based on the above description, this step provides a complete input data basis for the subsequent voxelized representation and surface reconstruction.
[0152] S802, a voxel grid is constructed for the sparse three-dimensional point cloud, and each voxel in the voxel grid corresponds to a pixel point.
[0153] In this embodiment, the space region where the sparse three-dimensional point cloud generated in S801 is located is divided into a uniformly distributed three-dimensional voxel grid, and each voxel represents a cubic volume unit, and the edge length can be preset according to the reconstruction accuracy requirement (such as 1 cm or less). Specifically, first, the space bounding box of the entire point cloud is determined, and the grid is divided along the x, y and z directions according to the resolution to form a discretized space index structure. Then, the mapping relationship between the voxels and the image pixels is established: for each pixel point in each key frame, if the ray formed by the back projection thereof passes through a certain voxel, it is considered that the voxel is associated with the pixel. It should be noted that each voxel corresponding to a pixel point here is not a one-to-one relationship, but means that in multi-view observation, each voxel can receive observation information from multiple pixel points, that is, one voxel can correspond to multiple pixel points under multiple views. Further, the voxel grid serves as a carrier of a signed distance function (SDF) and is used to record the geometric distance from each spatial position to the surface. As can be seen, this step completes the conversion from unordered point cloud to structured space representation, and provides topological support for subsequent signed distance calculation and weighted fusion. For example, when reconstructing the dam water stop joint area, the voxel grid covers a detection path of up to 30 meters, and can accurately capture the spatial morphology of the pit and the erosion zone. Based on the above description, this step, as a basic link of implicit surface modeling, realizes the dual functions of spatial discretization and observation association modeling.
[0154] Exemplarily, the space is divided into a voxel grid, the signed distance function (SDF) of each voxel to the surface is calculated, and the multi-view data is integrated by weighted fusion to construct a global truncated SDF (TSDF) function. For each voxel and camera pose , the effective distance along the ray direction is represented as:
[0155]
[0156] wherein, represents the camera optical center position; represents the included angle between the line-of-sight direction and the vector ; and represents the depth along the ray direction.
[0157] S803, based on the depth value of each pixel point, the signed distance of each voxel under multiple views is calculated, and the signed distance represents the distance of the current voxel to the boundary of the sparse three-dimensional point cloud.
[0158] In the embodiment, the signed distance is a key geometric quantity for measuring whether a voxel is inside or outside the surface of an object and the distance. During the execution of the above steps, for each voxel, all camera perspectives that can observe the voxel are traversed, and the directed distance from the center of the voxel to the nearest surface point is calculated by using the depth value of the corresponding pixel point under the perspective and the camera pose to construct a ray along the line of sight. Specifically, if the center of the voxel is in front of the surface (visible side), the signed distance is positive; if it is behind the surface (occluded side), the signed distance is negative; if it is close to the surface, the signed distance tends to zero. The process is implemented based on the TSDF (Truncated Signed Distance Function) model, and only effective observations within the truncated distance range are considered, and distances beyond the range are not taken into account. It should be noted that the multiple perspectives here include all key frames that meet the visibility conditions, i.e., the voxel is not occluded and is within the camera field of view. Further, the signed distance provided by each perspective is independently calculated to form a set of candidate values for subsequent fusion processing. As can be seen, this step converts two-dimensional depth information into three-dimensional implicit representation through physical imaging relationship, enhancing the description ability of complex geometric structure. In actual application process, due to the existence of slight jitter and local blur in underwater environment, part of the observation may have deviation, therefore, the truncation mechanism is introduced to suppress the propagation of long distance error. Based on the above description, this step completes the key conversion from explicit point cloud to implicit distance field, and lays a mathematical foundation for generating watertight surface.
[0159] S804, remove the target signed distance from the multiple signed distances corresponding to each voxel, and obtain an effective distance corresponding to each voxel according to the remaining signed distances; the number of difference values that exceed a preset value in the difference values between the target signed distance and other signed distances corresponding to the voxel is greater than a preset number.
[0160] In this embodiment, for each voxel corresponding to a plurality of symbol distance values obtained in S803, an outlier rejection operation is performed to improve the robustness of distance estimation. Specifically, for a given voxel, a set of symbol distances is received, and each value is checked for the absolute value of the difference between it and the other values: if the number of differences between a symbol distance and the remaining symbol distances exceeds a predetermined threshold (such as 2 cm) by more than a predetermined number (such as more than half), the value is determined to be an outlier, i.e. the target symbol distance, and is rejected. The remaining symbol distances that are not rejected participate in a weighted average operation, and the weights are usually related to the observation angle, confidence or distance cutoff factor, and the effective distance of the voxel is finally obtained. It should be noted that the target symbol distance specifically refers to an abnormal measurement value that deviates significantly from the true surface due to occlusion, motion blur, mismatch or noise interference, and its existence can seriously affect the smoothness and accuracy of the reconstruction result. As can be seen, this screening mechanism effectively suppresses the influence of poor observations on the overall model. For example, when there are local deposits or bubbles attached to the concrete surface, some angles may misjudge the surface position, but correct information can still be retained through multi-view consistency comparison. Further, this process embodies the core advantage of TSDF fusion: improving reconstruction reliability through redundant observations. Based on the above description, this step realizes the refinement and purification of the implicit distance field, ensuring the quality and stability of the subsequent isosurface extraction.
[0161] Exemplarily, to suppress the influence of noise and occlusion, these observations are truncated and weighted averaged to obtain a global TSDF volume:
[0162]
[0163] wherein, represents the truncated distance; represents the weight. The TSDF volume field is a discrete implicit representation, which needs to be converted into an explicit triangular mesh for subsequent use.
[0164] S805, generating a three-dimensional mesh model of the underwater concrete scene surface based on the effective distance corresponding to each voxel.
[0165] In this embodiment, all the effective distances corresponding to the voxels calculated in S804 are integrated into a globally consistent truncated signed distance field (TSDF Volume) as an implicit function representation of the three-dimensional scene. Specifically, the TSDF field is processed using a Poisson reconstruction algorithm: first, according to the gradient divergence theorem, the surface reconstruction problem is converted into a gradient field inversion problem of the indicator function, that is, to find a scalar function whose gradient direction is consistent with the observed normal and satisfies the Poisson equation form. By finite difference method, the partial differential equation is discretized and solved on the voxel grid to obtain a continuous implicit surface representation. Subsequently, the isosurface of the function value of 0.5, i.e. the position where the signed distance is zero, is extracted as the actual boundary of the object. Further, algorithms such as Marching Cubes or Dual Contouring are used to discretize the isosurface into a triangular mesh to generate a watertight, hole-free three-dimensional surface model. It should be noted that the three-dimensional mesh model here is defined as an explicit geometric representation consisting of vertices, edges and triangular facets, which can be directly used for visualization, measurement and defect quantification analysis. As can be seen, this step completes the final conversion from implicit field to explicit surface, realizing high-quality three-dimensional modeling. In actual application, this method successfully reconstructs a 30-meter-long dam water stop joint area, clearly presenting the abrasion zone and pit structure of the concrete, verifying its engineering practicability. Based on the above description, this step serves as the end point of the entire reconstruction process, outputting standardized three-dimensional digital assets that can be used for subsequent intelligent detection and maintenance decision-making.
[0166] Illustratively, the gradient field of the indicator function is converted into a continuous surface by Poisson reconstruction, and the isosurface is extracted. Based on the Gauss divergence theorem, the surface reconstruction problem is converted into a Poisson equation solution:
[0167]
[0168] wherein, represents the indicator function; represents the divergence of the field of lines. Finally, the isosurface of the indicator function value of 0.5 is extracted as the object surface to obtain the reconstructed triangular network:
[0169]
[0170] The obtained is a watertight triangular network that can be directly printed as a 3D image. MVS can be summarized as starting from 2D pixels, obtaining 3D depth through geometric modeling and physical constraints, fusing multi-view data to construct a global geometric field, and finally converting into a usable 3D model.
[0171] In some embodiments, reference is made to Figure 10 , Figure 10The 3D reconstructed and manually stitched images provided in the embodiments of the present invention Figure 10 (a) is an image of the enhanced concrete image arranged in the order of key frame extraction based on SFM, and then the concrete is reconstructed in three dimensions using the above method. (b) is an image presented by manual stitching in previous work in order to intuitively present the damage of the concrete in the detection area.
[0172] For example, 2D images show concrete damage, but without depth data, this damage cannot be effectively quantified. Reconstruction provides a clear view of concrete erosion and scour, which is beneficial for subsequent concrete damage analysis. Since the underwater concrete of the dam has been subjected to prolonged water erosion, the entire area exhibits erosion; therefore, this study focuses on the scour of the concrete, such as... Figure 11 As shown, Figure 11 This is a schematic diagram of 3D reconstruction of a concrete crater measurement provided in an embodiment of the present invention. From... Figure 11 As can be seen, the concrete crater area is an irregular cube, so there are multiple three-dimensional measurement methods. Here, the minimum circumscribed cube measurement method is used to represent the three-dimensional dimensions of the crater. The measured dimensions of the crater in this area are 6.23cm*4.21cm*2.19cm.
[0173] For example, to verify that image enhancement is beneficial for reconstruction, the same data was used for verification, such as... Figure 12 As shown, Figure 12 The image enhancement effect comparison diagram is provided for the embodiments of the present invention. Figure 12 (a) is a direct reconstruction of video keyframes. Figure 12 (b) is the data after 3D reconstruction following image enhancement. The reconstructed point cloud data reveals missing points when directly reconstructing using video keyframe data. Figure 12 (b) The comparison also reveals that areas with indistinct features failed to participate in the reconstruction, resulting in incomplete reconstruction areas.
[0174] Finally, a complete reconstruction of an area where the dual-mode underwater robot operated underwater was performed, obtaining large-scale 3D data of a water-stop joint in the underwater concrete dam, such as... Figure 13 As shown, Figure 13 This is a large-scale 3D reconstruction image of an underwater concrete structure provided for an embodiment of the present invention. Figure 13 (a) represents the underwater trajectory of the dual-mode robot. The robot's trajectory follows the distribution of the underwater concrete sealing joints of the dam, so the entire movement process is approximately a straight line. Figure 13 (b) is a keyframe of the video. Figure 13 (c) is a large-scale 3D point cloud based on SFM technology.Figure 13 (d) shows a two-dimensional grid of a large-scale concrete structure.
[0175] For example, due to the differences in the performance and reconstruction principles of underwater robots, quantitative comparisons are difficult. Therefore, we adopted a qualitative comparison method, comparing the 3D reconstruction capabilities of underwater robots in several existing papers, as shown in Table 2. For instance, Yu et al. designed VNRF, accelerating rendering through voxel meshes and using label propagation to solve crack occlusion; Fan et al.'s method and our proposed method consider 3D reconstruction of large underwater scenes; Rho et al.'s method improves reconstruction efficiency by optimizing the AUV scanning path through multi-sonar fusion.
[0176] Table 2
[0177]
[0178] Table 2 presents a qualitative comparison of the 3D reconstruction capabilities of different robots underwater, comparing them in four aspects: the size of the underwater reconstruction scene; reconstruction accuracy; dredging capability; and reconstruction efficiency. The method proposed by Fan et al. and the method proposed in this embodiment of the invention both consider 3D reconstruction of large underwater scenes; except for the method proposed by Rho et al., which has meter-level accuracy, the reconstruction accuracy of the other methods is millimeter-level; only the dual-modal robot provided in this embodiment of the invention possesses dredging capability. Therefore, through comparison, it can be found that the dual-modal underwater robot provided in this embodiment of the invention has a greater advantage in operating in complex underwater scenes.
[0179] Based on the above method, embodiments of the present invention also provide a system corresponding to the above method, such as... Figure 14 As shown, Figure 14 This is a functional module diagram of the underwater concrete scene 3D reconstruction system 1000 provided in an embodiment of the present invention. It should be noted that the basic principle and technical effects of the underwater concrete scene 3D reconstruction system 1000 provided in this embodiment are the same as those in the above method embodiments. For the sake of brevity, parts not mentioned in this embodiment can be referred to the corresponding content in the method embodiments. The underwater concrete scene 3D reconstruction system 1000 includes a data acquisition module 1100, an image enhancement module 1200, and a 3D reconstruction module 1300.
[0180] In this embodiment, the data acquisition module 1100 is used to acquire video data of the underwater concrete surface and obtain multiple keyframe images from the video data. It can be understood that the data acquisition module 1100 is used to perform the above-described step S100.
[0181] The image enhancement module 1200 is configured to select a plurality of video images from the video data, perform background light stripping on each video image to obtain a background light stripped image, perform transmittance compensation on the background light stripped image to obtain a preliminary clear image, and perform adaptive color correction on the preliminary clear image to obtain an image-enhanced key frame image. It can be understood that the image enhancement module 1200 is configured to perform steps S201-S204 described above.
[0182] The image enhancement module 1200 is further configured to perform transmittance compensation on the background light stripped image to obtain a compensated image, and perform regularization optimization on the compensated image to obtain the preliminary clear image.
[0183] The three-dimensional reconstruction module 1300 is configured to extract a feature point in each key frame image, the feature point representing a conspicuous part of the underwater concrete surface, perform epipolar geometry comparison on the feature points in adjacent key frame images to obtain a feature matching point pair in the adjacent key frame images, the feature matching point pair including pixel coordinates of a same matched feature point in the adjacent key frame images, calculate a relative pose of the adjacent key frame images according to the pixel coordinates of the same matched feature point in the adjacent key frame images, the relative pose representing a rotation angle and a translation vector of a camera corresponding to the adjacent key frame images, calculate a three-dimensional coordinate of the feature point based on the pixel coordinates of the same matched feature point in the adjacent key frame images and the relative pose, calculate a pixel depth value of each pixel point in the key frame image according to the relative pose, and perform multi-view fusion based on the relative pose, the three-dimensional coordinate of the matched feature point, and the pixel depth value of each pixel point to generate a three-dimensional mesh model of the underwater concrete scene. It can be understood that the three-dimensional reconstruction module 1300 is configured to perform steps S300-S800 described above.
[0184] In some embodiments, the three-dimensional reconstruction module 1300 is further configured to obtain pixel coordinates of the feature points in the adjacent key frame images, generate a binary descriptor based on the pixel coordinates of the feature points, compare initial feature matching point pairs based on the binary descriptor, and perform epipolar geometry comparison on the initial feature matching point pairs to obtain the feature matching point pair in the adjacent key frame images. It can be understood that the three-dimensional reconstruction module 1300 is further configured to perform steps S401-S404 described above.
[0185] In some embodiments, the three-dimensional reconstruction module 1300 is further configured to establish an epipolar geometry constraint equation and solve an essential matrix based on the pixel coordinates of the same matched feature point in the adjacent key frame images and preset parameters of the camera, and decompose the essential matrix to obtain the relative pose of the adjacent key frame images. It can be understood that the three-dimensional reconstruction module 1300 is configured to perform steps S501-S502 described above.
[0186] In some embodiments, the three-dimensional reconstruction module 1300 is further configured to calculate a depth value of a feature point in the key frame image according to the relative pose; assign a preset number of initial depth values to each pixel in the key frame image, the initial depth values being close to a depth value of a feature point closest to the pixel; perform photometric consistency calculation on the preset number of initial depth values and a depth value of each target feature point corresponding to the pixel respectively to obtain a score of each initial depth value relative to the depth value of each target feature point, and select an initial depth value with the highest score from scores of all initial depth values corresponding to each target feature point as a target depth value; the target depth value corresponds to the target feature point one-to-one; the target feature point is a feature point in the key frame image having a distance within a preset range to the pixel; and determine a pixel depth value of the pixel according to target depth values of all target feature points. It can be understood that the three-dimensional reconstruction module 1300 is further configured to perform steps S701-S704.
[0187] In some embodiments, the three-dimensional reconstruction module 1300 is further configured to perform multi-view fusion according to the relative pose, three-dimensional coordinates of the matched feature points, and the pixel depth value of each pixel to generate a sparse three-dimensional point cloud; construct a voxel grid for the sparse three-dimensional point cloud, each voxel in the voxel grid corresponding to a pixel; calculate a signed distance of each voxel under multiple views based on the depth value of each pixel, the signed distance representing a distance of the current voxel to a boundary of the sparse three-dimensional point cloud; remove a target signed distance from multiple signed distances corresponding to each voxel, and obtain an effective distance corresponding to each voxel according to the remaining signed distances; the number of difference values exceeding a preset value in a difference value between the target signed distance and other signed distances corresponding to the voxel is greater than a preset number; and generate a three-dimensional mesh model of a surface of the underwater concrete scene based on the effective distance corresponding to each voxel. It can be understood that the three-dimensional reconstruction module 1300 is further configured to perform steps S801-S805.
[0188] Based on the same inventive concept disclosed above, the embodiments of the present application also provide a block schematic diagram of an electronic device 2000 for performing the above method. Please refer to Figure 15 , Figure 15 The block schematic diagram of the electronic device 2000 provided by the embodiments of the present application includes a processor 2100, a memory 2200, a bus 2300, and a communication interface 2400. The processor 2100 and the memory 2200 are connected through the bus 2300, and the processor 2100 communicates with external devices through the communication interface 2400.
[0189] The processor 2100 can be an integrated circuit chip with a processing capability of signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 2100 or the instruction in the form of software. The processor 2100 described above can be a general processor 2100, including a central processing unit 2100 (CPU), a network processor 2100 (NP), etc.; can also be a digital signal processor 2100 (DSP), an application specific integrated circuit (ASIC), a ready programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
[0190] The memory 2200 is used for storing a computer program, for example, the underwater concrete scene three-dimensional reconstruction system 1000 in the embodiment of the application, including at least one software function module stored in the memory 2200 in the form of software or firmware, and the processor 2100 executes the program to realize the underwater concrete scene three-dimensional reconstruction method in the embodiment of the application after receiving an execution instruction.
[0191] The memory 2200 can include a high-speed random access memory 2200 (RAM), and can also include a non-volatile memory 2200. Optionally, the memory 2200 can be a storage device built in the processor 2100, or can be a storage device independent of the processor 2100.
[0192] The bus 2300 can be an ISA bus 2300, a PCI bus 2300 or an EISA bus 2300, etc. Figure 15 Only one bidirectional arrow is used for representation, but it does not mean that there is only one bus 2300 or only one type of bus 2300.
[0193] The electronic device 2000 can be a mobile phone, a tablet computer, a notebook computer, a desktop computer or the like.
[0194] Based on the same inventive concept, the embodiment of the application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by the processor 2100 to realize the underwater concrete scene three-dimensional reconstruction method as described above. The computer readable storage medium can include a U disk, a mobile hard disk, a read-only memory 2200 (ROM), a random access memory 2200 (RAM), a magnetic disk or an optical disk and various storage program codes.
[0195] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. The present application can have various changes and modifications for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for three-dimensional reconstruction of underwater concrete scene, characterized in that, The method includes: Acquire video data of underwater concrete surfaces; Obtaining multiple keyframe images from the video data includes: Select multiple video images from the video data; Perform background light stripping on each of the video images to obtain a background light stripped image; By performing castability compensation on the background light stripped image, a preliminary clear image is obtained; Adaptive color correction is performed on the initially clear image to obtain the enhanced keyframe image; Feature points are extracted from each of the keyframe images, and the feature points represent conspicuous parts of the underwater concrete surface. Perform epipolar geometry comparison on feature points in adjacent keyframe images to obtain feature matching point pairs in adjacent keyframe images. The feature matching point pairs contain the pixel coordinates of the same matched feature point in adjacent keyframe images. The relative pose of adjacent keyframe images is calculated based on the pixel coordinates of the same matched feature point in adjacent keyframe images. The relative pose represents the rotation angle and translation vector of the camera corresponding to the adjacent keyframe images. The three-dimensional coordinates of the matched feature point are calculated based on the pixel coordinates of the same matched feature point in adjacent keyframe images and the relative pose. Calculate the pixel depth value of each pixel in the keyframe image based on the relative pose; A three-dimensional mesh model of the underwater concrete scene is generated by multi-view fusion based on the relative pose, the three-dimensional coordinates of the matched feature points, and the pixel depth value of each pixel.
2. The method of claim 1, wherein, The step of performing epipolar geometric comparison on feature points in adjacent keyframe images to obtain feature matching point pairs in adjacent keyframe images includes: Obtain the pixel coordinates of feature points in adjacent keyframe images; Generate a binary descriptor based on the pixel coordinates of the feature points; Initial feature matching point pairs are determined based on the binary descriptor; The initial feature matching point pairs are subjected to epipolar geometry alignment to obtain feature matching point pairs in adjacent keyframe images.
3. The method according to claim 1, characterized in that, The step of calculating the relative pose of adjacent keyframe images based on the pixel coordinates of the same matched feature point in adjacent keyframe images includes: Based on the pixel coordinates of the same matched feature point in adjacent keyframe images and the preset parameters of the camera, an epipolar geometric constraint equation is established and the essential matrix is solved. The relative poses of adjacent keyframe images are obtained by decomposing the essential matrix.
4. The method according to claim 1, characterized in that, The step of calculating the pixel depth value of each pixel in the keyframe image based on the relative pose includes: Calculate the depth values of feature points in the keyframe image based on the relative pose; Each pixel in the keyframe image is assigned a preset number of initial depth values, and the initial depth values are close to the depth values of the feature points closest to the pixel. The preset number of initial depth values are respectively compared with the depth value of each target feature point corresponding to the pixel to calculate photometric consistency, so as to obtain the score of different initial depth values relative to the depth value of each target feature point. The initial depth value with the highest score is selected from all the scores of the initial depth values corresponding to each target feature point as the target depth value. The target depth value corresponds one-to-one with the target feature point. The target feature point is the feature point in the keyframe image whose distance from the pixel is within a preset range. The pixel depth value of the pixel is determined based on the target depth values corresponding to all target feature points.
5. The method according to claim 1, characterized in that, The step of generating a 3D mesh model of the underwater concrete scene by performing multi-view fusion based on the relative pose, the 3D coordinates of the matched feature points, and the pixel depth value of each pixel includes: Multi-view fusion is performed based on the relative pose, the three-dimensional coordinates of the feature points, and the pixel depth value of each pixel to generate a sparse three-dimensional point cloud. A voxel mesh is constructed for the sparse 3D point cloud, wherein each voxel in the voxel mesh corresponds to a pixel. Based on the depth value of each pixel, the symbolic distance of each voxel in multiple viewpoints is calculated. The symbolic distance represents the distance of the current voxel from the boundary of the sparse 3D point cloud. The target symbolic distance is removed from the multiple symbolic distances corresponding to each voxel, and the effective distance corresponding to each voxel is obtained based on the remaining symbolic distances; among the differences between the target symbolic distance and the other symbolic distances corresponding to the voxel, the number of differences exceeding a preset value is greater than a preset number; A three-dimensional mesh model of the underwater concrete scene surface is generated based on the effective distance corresponding to each voxel.
6. The method according to claim 1, characterized in that, The step of performing cast rate compensation on the background light stripped image to obtain a preliminary clear image includes: Transmittance compensation is performed on the background light stripping image to obtain a compensated image; The compensated image is then regularized to obtain a preliminary clear image.
7. A three-dimensional reconstruction system for underwater concrete scenes, characterized in that, The system includes: The data acquisition module is used to acquire video data of the underwater concrete surface, and to acquire multiple keyframe images from the video data. Specifically, the data acquisition module is used to select multiple video images from the video data; to perform background light stripping on each video image to obtain a background light stripped image; to perform cast rate compensation on the background light stripped image to obtain a preliminary clear image; and to perform adaptive color correction on the preliminary clear image to obtain an enhanced keyframe image. A 3D reconstruction module is used to extract feature points from each keyframe image, where the feature points represent conspicuous parts of the underwater concrete surface; perform epipolar geometry comparison on the feature points in adjacent keyframe images to obtain feature matching point pairs in adjacent keyframe images, where each feature matching point pair contains the pixel coordinates of the same feature point in the adjacent keyframe images; calculate the relative pose of the adjacent keyframe images based on the pixel coordinates of the same feature point in the adjacent keyframe images, where the relative pose represents the rotation angle and translation vector of the camera corresponding to the adjacent keyframe images; calculate the 3D coordinates of the feature point based on the pixel coordinates of the same feature point in the adjacent keyframe images and the relative pose; calculate the pixel depth value of each pixel in the keyframe image based on the relative pose; and perform multi-view fusion based on the relative pose, the 3D coordinates of the feature point, and the pixel depth value of each pixel to generate a 3D mesh model of the underwater concrete scene.
8. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a computer program that can be executed by the processor to implement the three-dimensional reconstruction method for underwater concrete scenes according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the three-dimensional reconstruction method for underwater concrete scenes as described in any one of claims 1-6.
Citation Information
Patent Citations
Non-cooperative target three-dimensional space pose estimation method and system, equipment and medium
CN118196188A
End-to-end three-dimensional reconstruction method and device based on unmanned aerial vehicle image
CN118781280A