Concrete crack three-dimensional reconstruction method and system based on multi-modal fusion and medium

Through the multimodal fusion method, the three-dimensional reconstruction of cracks is combined with images and lidar data, which solves the problem of insufficient three-dimensional detection accuracy in the prior art, and realizes the simultaneous detection of crack width, position and direction information with high accuracy.

CN120259573AActive Publication Date: 2025-07-04CENT SOUTH UNIV

Patent Information

Application Number
CN202510751650.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

In the prior art, the concrete crack skeleton is relied on only single frame image data to reconstruct the concrete crack skeleton and measure the width, making it difficult to realize three-dimensional crack detection of the structure. In addition, the existing three-dimensional reconstruction methods lack the density, accuracy and robustness in the local detail reconstruction of point clouds, and cannot meet the detection needs of microstructures.

Method used

The multimodal fusion method is adopted, combining structural image frame sequences and structural lidar point cloud frame sequences, segmentation processing is performed through the crack semantic segmentation model, and a three-dimensional point cloud map is constructed by combining spherical interpolation and radar real-time estimation of poses to obtain visible point cloud coloring and semantic segmentation point cloud maps, and finally fusion obtains the overall three-dimensional point cloud marking map to realize the measurement of crack three-dimensional geometric information.

Benefits of technology

High-precision three-dimensional crack detection is realized, and the width information, position and direction information of the crack can be obtained at the same time, adapting to the detection needs of various three-dimensional structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259573A_ABST
    Figure CN120259573A_ABST
Patent Text Reader

Abstract

The invention discloses a concrete crack three-dimensional reconstruction method and system based on multi-modal fusion and a medium, and relates to the technical field of structural engineering detection, and the method comprises the following steps: obtaining a structural image frame sequence and a structural laser radar point cloud frame sequence of a target structure; obtaining a crack mask image frame after crack segmentation; an integral structure de-noised point cloud picture is obtained; obtaining a visible point cloud coloring point cloud map and a visible point cloud semantic segmentation point cloud map; obtaining an overall three-dimensional coloring point cloud map and an overall three-dimensional semantic segmentation point cloud map; obtaining an overall three-dimensional point cloud marking map; and performing three-dimensional attribute measurement on the crack based on the three-dimensional point cloud marking map to obtain three-dimensional geometric information of the crack. According to the method provided by the invention, a multi-frame and multi-modal fused crack structure reconstruction framework is designed, the method can adapt to crack detection of various three-dimensional structures, and simultaneous detection of crack width information, crack position and crack trend information can be realized based on an overall three-dimensional point cloud marking map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of structural engineering detection, and particularly relates to a three-dimensional reconstruction method, system and medium for concrete cracks based on multi-modal fusion. Background Art

[0002] With the extension of the service life of concrete infrastructures such as bridges and tunnels, the problem of surface cracks caused by structural aging has become increasingly prominent; as a key indicator for evaluating the stress state, safety and durability of structures, the accurate detection and three-dimensional quantitative analysis of cracks have become the core requirements of health monitoring.

[0003] In the prior art, cracks are mainly detected based on manual inspection and computer vision technology; among them, manual inspection relies on visual observation, which has defects such as low efficiency, strong subjectivity, and high risks in high-altitude operations, and manual inspection cannot obtain the three-dimensional geometric information of cracks, making it difficult to meet the comprehensive and accurate requirements of modern engineering for structural health assessment; computer vision technology mainly performs crack segmentation and detection based on two-dimensional images, and three-dimensional information such as crack size, position, and topology cannot be directly obtained from two-dimensional crack information.

[0004] Currently, for three-dimensional crack detection, some researchers have reconstructed the crack skeleton and measured the crack width using image data in a single frame. However, using only a single image and its related processing techniques is not sufficient to capture the complete three-dimensional spatial details of cracks, making it difficult to achieve three-dimensional crack detection of structures. Further, in order to generate comprehensive three-dimensional crack information, researchers have extended the detection from a single local image to the entire structure, and three-dimensional reconstruction methods are generally divided into two categories: methods based on Structure from Motion (SfM) and methods based on Simultaneous Localization and Mapping (SLAM). The current three-dimensional reconstruction methods are still insufficient in the reconstruction density, accuracy, and robustness of local details of the point cloud, and cannot meet the detection requirements of such tiny structures as cracks. Moreover, the existing methods usually perform crack width measurement and crack localization separately, where crack width measurement is carried out on two-dimensional images, while crack localization is carried out on three-dimensional point clouds, and it is impossible to achieve synchronous automatic quantification of the 3D position and width of cracks.

[0005] In view of this, it is necessary to propose a three-dimensional reconstruction method, system and medium for concrete cracks based on multi-modal fusion to solve or at least alleviate some of the above defects. Summary of the Invention

[0006] The main object of the present invention is to provide a three-dimensional reconstruction method, system and medium for concrete cracks based on multi-modal fusion, aiming to solve the technical problem in the prior art that only relying on the image data in a single frame to reconstruct the crack skeleton and measure the crack width makes it difficult to achieve three-dimensional crack detection of the structure.

[0007] To achieve the above object, the present invention provides a three-dimensional reconstruction method for concrete cracks based on multi-modal fusion, including the following steps: S10, acquiring a sequence of structural image frames and a sequence of structural lidar point cloud frames of the target structure according to a preset acquisition trajectory, and acquiring the pose of the acquisition device corresponding to each acquisition moment; S20, using a crack semantic segmentation model to segment each frame of the structural image frame to obtain a crack mask map frame after crack segmentation; S30, performing three-dimensional point cloud mapping based on the structural lidar point cloud frame and the corresponding radar measurement pose and processing to obtain an overall structure denoised point cloud map, and the radar measurement pose is obtained based on the external parameters of the device and the pose of the acquisition device corresponding to the corresponding acquisition moment; S40, obtaining the radar real-time estimated pose at any moment by spherical interpolation calculation based on at least two radar measurement poses, and obtaining the corresponding key shooting poses at the key shooting moments of each camera according to the radar real-time estimated pose and the external parameters of the device, where the key shooting moment of the camera is the shooting moment corresponding to the key frame of the structural image, and the key frame of the structural image is determined from the sequence of structural image frames; S50, obtaining the visible three-dimensional point cloud corresponding to the key shooting moment, and projecting the three-dimensional points of the visible three-dimensional point cloud onto the corresponding key frame of the structural image and the crack mask map frame according to the key shooting pose for point cloud coloring and semantic segmentation to obtain a visible point cloud colored point cloud map and a visible point cloud semantic segmentation point cloud map; S60, performing point cloud coloring and semantic segmentation processing on all the key frames of the structural image and the crack mask map frames to obtain an overall three-dimensional colored point cloud map and an overall three-dimensional semantic segmentation point cloud map; S70, fusing the overall three-dimensional colored point cloud map and the overall three-dimensional semantic segmentation point cloud map to obtain an overall three-dimensional point cloud labeled map with crack information and coloring; S80, performing three-dimensional attribute measurement of the cracks based on the three-dimensional point cloud labeled map to obtain three-dimensional geometric information of the cracks, where the three-dimensional geometric information of the cracks includes the width information, the position of the cracks, and the crack trend information.

[0008] Further, step S20 specifically includes: Using a crack semantic segmentation model to segment each frame of the structural image frame to obtain an initial mask map after crack segmentation, and the initial mask map has black and white markings and semantic information; The initial mask image is cut into several feature images, and the feature images are processed to extract the crack skeleton image, and the crack skeleton image has pixel colors representing the distance values of each pixel point from the nearest edge; The crack skeleton image is converted into a set of pixel coordinate point sets, and pixel key points are extracted after downsampling; The pixel key point list is sorted and downsampled, and the feature images are marked based on a quantity threshold; According to the markings of the feature images as prompt points for the large model, the corresponding structural image frames of the feature images are imported into the large model for crack mask optimization, and a crack mask image frame is obtained by screening using a width optimization threshold.

[0009] Further, step S30 specifically includes: Based on the structured lidar point cloud frame and the corresponding radar measurement pose, 3D point cloud mapping is performed to obtain an overall 3D point cloud map; The overall 3D point cloud map is subjected to a cropping process; The cropped overall 3D point cloud map is denoised and smoothed to obtain an overall structure denoised point cloud map.

[0010] Further, in step S40: Based on at least two radar measurement poses, the rotation part is calculated by spherical interpolation, and the position part is interpolated linearly to obtain the radar real-time estimated pose at any moment; The camera pose is calibrated at each moment according to the radar real-time estimated pose and the radar-camera extrinsic calibration parameters, and the key shooting pose at the key shooting moment of the camera is obtained through coordinate transformation and time synchronization calculation; Further, step S50 specifically includes: Visibility filtering is performed on the target moment point cloud frame based on the current viewing angle of the key shooting pose to obtain the key moment visible point cloud frames corresponding to each camera key shooting moment, and the key moment visible point cloud frames are obtained based on the overall structure denoised point cloud map; The point clouds in the key moment visible point cloud frames are projected onto the image pixel points on the corresponding camera's image plane for matching, and point cloud coloring is performed to obtain a visible point cloud colored point cloud map; The point clouds in the key moment visible point cloud frames are projected onto the pixel points of the crack mask map on the corresponding camera's image plane for matching, and point cloud semantic segmentation is performed to obtain a visible point cloud semantic segmentation point cloud map.

[0011] Further, in step S50: Time synchronization is used to determine the radar real-time estimated pose corresponding to the key shooting moment, and visibility filtering is performed on the overall structure denoised point cloud map by the HRP operator according to the radar real-time estimated pose corresponding to the key shooting moment to obtain the visible 3D point cloud corresponding to the key shooting moment.

[0012] Further, step S60 specifically includes: Perform point cloud coloring processing on all key frames of the structural images to obtain an overall three-dimensional colored point cloud map; Perform point cloud semantic segmentation processing on the crack mask map frames corresponding to all key frames of the structural images to obtain an overall three-dimensional semantic segmentation point cloud map; Further, the step of "performing point cloud coloring processing on all key frames of the structural images to obtain an overall three-dimensional colored point cloud map" specifically includes: obtaining the associated structural image frames corresponding to each three-dimensional point in all visible point cloud frames at critical moments, where the associated structural image frames have associated pixel points matching the three-dimensional points; among them, determine the associated weight score corresponding to the three-dimensional point based on the weighted scoring method of camera angle and distance, and determine the top N frames with the highest associated weight score as the associated structural image frames of the three-dimensional point; use the normalized weights to fuse the RGB colors of the associated pixel points in all the associated structural image frames to determine the color value of the three-dimensional point; The step of "performing point cloud semantic segmentation processing on the crack mask map frames corresponding to all key frames of the structural images to obtain an overall three-dimensional semantic segmentation point cloud map" specifically includes: obtaining the associated mask map frames corresponding to each three-dimensional point in all visible point cloud frames at critical moments, where the associated mask map frames have associated pixel points matching the three-dimensional points; among them, determine the associated weight score corresponding to the three-dimensional point based on the weighted scoring method of camera angle and distance, and determine the top N frames with the highest associated weight score as the associated mask map frames of the three-dimensional point; use the normalized weights to fuse the colors of the associated pixel points in all the associated mask map frames to determine the semantic information of the three-dimensional point.

[0013] Further, step S70 specifically includes: Determine the crack recognition points on the two-dimensional skeleton image according to the crack mask map frames, and calculate the gradient within the domain of measuring the crack position to determine the crack direction vector; On the two-dimensional skeleton map, locate the left and right edge points relative to the skeleton points by following the direction and its perpendicular direction, and iteratively track to determine the two crack edge points of the measured crack position; Project the point cloud onto the image, select the three-dimensional point matching the crack center point on the image based on the minimum error, then calculate the plane equation at this point in the point cloud map based on this point, generate a coordinate grid within a specified radius based on this plane, and calculate the corresponding coordinates using the plane equation, which will obtain a set of sampled three-dimensional points distributed on the plane; Project the sampled three-dimensional points onto the two-dimensional image plane according to the internal parameters of the camera, and select the three-dimensional point best matching the crack edge point based on the minimum projection error to obtain the specific coordinate values of the two crack edge points.

[0014] Further, before step S20 includes: Construct a DeepLab-v3+ network model with ResNet as the backbone, determine the initial hyperparameters of the network model, and obtain a pre-trained model after pre-training; Use the labeled training set and test set to train the pre-trained model to obtain a crack semantic segmentation model.

[0015] The present invention also provides a three-dimensional reconstruction system for concrete cracks based on multimodal fusion, including a processing device, which is used to implement the steps of the above-mentioned three-dimensional reconstruction method for concrete cracks based on multimodal fusion.

[0016] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-mentioned three-dimensional reconstruction method for concrete cracks based on multimodal fusion.

[0017] The three-dimensional reconstruction method of concrete cracks based on multimodal fusion provided by the present invention first collects the structural image frame sequence and the structural lidar point cloud frame sequence of the target structure, and obtains the poses of the acquisition devices corresponding to each acquisition moment; and based on the structural lidar point cloud frame and the corresponding radar measurement pose, three-dimensional point cloud mapping is performed and processed to obtain the overall structural denoised point cloud map, and a crack semantic segmentation model is used to segment each frame of the structural image frame to obtain the crack mask map frame after crack segmentation; then, the structural image frame is determined from the key frame sequence of the structural image and the corresponding key shooting moment is obtained, and based on at least two radar measurement poses, the radar real-time estimated pose at any moment is obtained through spherical interpolation calculation. According to the radar real-time estimated pose and the external parameters of the device, the corresponding key shooting pose at the key shooting moment is obtained, and the accurate positioning of the associated shooting pose is realized when the shooting frame rates of the radar and the camera are not synchronized based on the radar real-time estimated pose; then, after obtaining the visible three-dimensional point cloud corresponding to the key shooting moment, the three-dimensional points of the visible three-dimensional point cloud are projected onto the corresponding structural image key frame and the crack mask map frame respectively according to the key shooting pose for point cloud coloring and semantic segmentation, and the visible point cloud colored point cloud map and the visible point cloud semantic segmentation point cloud map are obtained; then, all the structural image key frames and the crack mask map frames are subjected to point cloud coloring and semantic segmentation processing to obtain the overall three-dimensional colored point cloud map and the overall three-dimensional semantic segmentation point cloud map. Finally, the overall three-dimensional colored point cloud map and the overall three-dimensional semantic segmentation point cloud map are fused to obtain the overall three-dimensional point cloud labeled map with crack information and coloring, and the three-dimensional geometric information of the crack is obtained based on the three-dimensional point cloud labeled map. The three-dimensional geometric information of the crack includes the width information, the crack position and the crack orientation information, so as to realize the three-dimensional crack detection of the structure; the three-dimensional reconstruction method of concrete cracks based on multimodal fusion provided by the present invention designs a crack structure reconstruction framework that combines multiple frames and multimodal fusion when performing three-dimensional crack detection, combines lidar point cloud with structural plane images to generate a dense and high-precision overall three-dimensional point cloud labeled map with crack semantics and real-world dimensions, can adapt to the crack detection of various three-dimensional structures, and can simultaneously detect the width information, the crack position and the crack orientation information of the crack based on the overall three-dimensional point cloud labeled map. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings. Among them, the dashed boxes in the figures represent the processing procedures.

[0019] Figure 1Schematic flowchart of a three-dimensional reconstruction method for concrete cracks based on multimodal fusion in an embodiment of the present invention; Figure 2 Schematic flowchart of obtaining a crack mask map frame in step S20 in an embodiment of the present invention; Figure 3 Schematic three-dimensional structure diagram of the overall three-dimensional point cloud map after cropping processing in step S30 in an embodiment of the present invention; wherein, a is one of the structural diagrams of the overall three-dimensional point cloud map, and b is the other structural diagram of the overall three-dimensional point cloud map; Figure 4 Overall structural denoised point cloud map obtained by performing denoising and smoothing processing in step S30 in an embodiment of the present invention; wherein, a is one of the schematic diagrams of the overall structural point cloud, and b is the other schematic diagram of the overall structural point cloud; Figure 5 Schematic flowchart of obtaining an overall three-dimensional point cloud labeled map in an embodiment of the present invention; Figure 6 Overall three-dimensional colored point cloud map in step S60 in an embodiment of the present invention; Figure 7 Overall three-dimensional semantic segmentation point cloud map (three-dimensional point cloud map with crack information marked) in step S60 in an embodiment of the present invention; Figure 8 Overall three-dimensional point cloud labeled map (colored, three-dimensional point cloud map with crack information marked) in step S70 in an embodiment of the present invention; Figure 9 Schematic flowchart of three-dimensional attribute measurement of cracks in step S80 in an embodiment of the present invention.

[0020] The realization of the object, functional features and advantages of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners

[0021] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] It should be noted that all the directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative positional relationship, movement conditions, etc. between components in a specific posture (as shown in the attached drawings). If this specific posture changes, the directional indications will also change accordingly.

[0024] In addition, the descriptions involving "first", "second", etc. in the present invention are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0025] Please refer to the attached Figure 1 , the present invention provides a three-dimensional reconstruction method for concrete cracks based on multimodal fusion, including the following steps: S10, acquiring a sequence of structural image frames and a sequence of structural lidar point cloud frames of the target structure according to a preset acquisition trajectory, and acquiring the pose of the acquisition device corresponding to each acquisition moment; S20, using a crack semantic segmentation model to perform segmentation processing on each frame of the structural image frame to obtain a crack mask map frame after crack segmentation; S30, performing three-dimensional point cloud mapping based on the structural lidar point cloud frame and the corresponding radar measurement pose and processing to obtain an overall structural denoised point cloud map, and the radar measurement pose is obtained based on the external parameters of the device and the pose of the acquisition device corresponding to the corresponding acquisition moment; S40, obtaining the radar real-time estimated pose at any moment through spherical interpolation calculation based on at least two radar measurement poses, and obtaining the corresponding key shooting poses at each key shooting moment of the camera according to the radar real-time estimated pose and the external parameters of the device. The key shooting moment of the camera is the shooting moment corresponding to the key frame of the structural image, and the key frame of the structural image is determined from the structural image frames; S50, obtaining the visible three-dimensional point cloud corresponding to the key shooting moment, projecting the three-dimensional points of the visible three-dimensional point cloud onto the corresponding key frame of the structural image and the crack mask map frame according to the key shooting pose for point cloud coloring and semantic segmentation to obtain a visible point cloud colored point cloud map and a visible point cloud semantic segmentation point cloud map; S60, performing point cloud coloring and semantic segmentation processing on all the key frames of the structural image and the crack mask map frames to obtain an overall three-dimensional colored point cloud map and an overall three-dimensional semantic segmentation point cloud map; S70. Integrate the overall three-dimensional colored point cloud map and the overall three-dimensional semantic segmentation point cloud map to obtain an overall three-dimensional point cloud labeled map with crack information and coloring. S80. Perform three-dimensional attribute measurement of cracks based on the three-dimensional point cloud labeled map to obtain three-dimensional geometric information of cracks. The three-dimensional geometric information of cracks includes crack width information, crack position, and crack orientation information.

[0026] The three-dimensional concrete crack reconstruction method based on multi-modal fusion provided by the present invention first collects a sequence of structural image frames and a sequence of structural lidar point cloud frames of a target structure, and obtains the pose of the acquisition device corresponding to each acquisition moment. And based on the structural lidar point cloud frames and the corresponding radar measurement poses, three-dimensional point cloud mapping is performed and processed to obtain an overall structure denoised point cloud map, and a crack semantic segmentation model is used to segment each frame of the structural image frame to obtain a crack mask map frame after crack segmentation. Then, determine the structural image frames from the sequence of structural image key frames and obtain their corresponding key shooting moments, and obtain the radar real-time estimated pose at any moment through spherical interpolation calculation based on at least two radar measurement poses. According to the radar real-time estimated pose and the external parameters of the device, obtain the corresponding key shooting pose at the key shooting moment. Based on the radar real-time estimated pose, accurate positioning of the associated shooting pose is achieved when the shooting frame rates of the radar and the camera are not synchronized. Then, after obtaining the visible three-dimensional point cloud corresponding to the key shooting moment, project the three-dimensional points of the visible three-dimensional point cloud onto the corresponding structural image key frame and crack mask map frame respectively according to the key shooting pose for point cloud coloring and semantic segmentation, and obtain a visible point cloud colored point cloud map and a visible point cloud semantic segmentation point cloud map. Then, perform point cloud coloring and semantic segmentation processing on all the structural image key frames and crack mask map frames to obtain an overall three-dimensional colored point cloud map and an overall three-dimensional semantic segmentation point cloud map. Finally, integrate the overall three-dimensional colored point cloud map and the overall three-dimensional semantic segmentation point cloud map to obtain an overall three-dimensional point cloud labeled map with crack information and coloring, and perform three-dimensional attribute measurement of cracks based on the three-dimensional point cloud labeled map to obtain three-dimensional geometric information of cracks. The three-dimensional geometric information of cracks includes crack width information, crack position, and crack orientation information, realizing three-dimensional crack detection of the structure. The three-dimensional concrete crack reconstruction method based on multi-modal fusion provided by the present invention, when performing three-dimensional crack detection, designs a crack structure reconstruction framework for multi-frame and multi-modal fusion, combines lidar point cloud and structural plane images to generate a dense and high-precision overall three-dimensional point cloud labeled map with crack semantics and real-world dimensions, can adapt to crack detection of various three-dimensional structures, and can simultaneously detect the crack width information, crack position, and crack orientation information based on the overall three-dimensional point cloud labeled map.

[0027] Understandably, a structural image key frame refers to a representative image frame selected during data processing. A complete set of selected structural key frames is a core image set that can completely represent the crack information on the surface of the target structure. Understandably, the method of the present invention is mainly used for the detection of three-dimensional cracks such as curved surface cracks.

[0028] Further, step S20 specifically includes: using a crack semantic segmentation model to perform segmentation processing on each frame of the structural image frame to obtain an initial mask map after crack segmentation. The initial mask map has black and white markings and semantic information (the structural image frame has color information, and the initial mask map is a black and white map with black and white markings and has crack semantic information); cutting the initial mask map into several feature pictures, processing the feature pictures to extract a crack skeleton map, and the crack skeleton map has pixel colors representing the distance values of each pixel point to the nearest edge; converting the crack skeleton map into a set of pixel coordinate point sets, performing downsampling and then extracting pixel key points; sorting and downsampling the list of pixel key points, and marking the feature pictures based on a quantity threshold; using the markings of the feature pictures as prompt points for the large model, importing the structural image frame corresponding to the feature picture into the large model for crack mask optimization, and screening using a width optimization threshold to obtain a crack mask map frame.

[0029] Understandably, in the solution of the present invention, the collected sequence of structural image frames is subjected to crack image segmentation using the deep learning network DeepLab-v3+, and the large model SAM is used to optimize the segmentation result. As Figure 2As shown, specifically, first, image, lidar point cloud, and pose data of the cracked structure are collected. Then, a sequence of structural image frames is extracted from the obtained rosbag data. The structural image frames are segmented for crack images using the trained DeepLab-v3+ deep learning network model to obtain an initial mask image (Mask) after crack segmentation. Among them, the initial mask image is a black-and-white image, where black pixels on the initial mask image represent pixel points that are not cracks, and white pixels represent pixel points that are cracks. The initial mask image is cut into several small pictures, and the EDT (Euclidean Distance Transform) is used to extract the crack skeleton image from the cut initial mask image. The crack skeleton image reflects the distance value of each pixel point to the nearest edge, and the pixel color of the crack skeleton image represents the distance change. The pixel points of the crack skeleton image are converted into a set of coordinate point sets, downsampled for a given threshold, hint points are selected, and then contour detection is performed. Each contour is traversed, and pixel key points are extracted from it. The list of pixel key points is sorted and downsampled. After selecting a given number of feature points, marks are made at the corresponding positions of the original image crop. The large model Segment Anything Model is called to optimize the crack mask, a width optimization result evaluation program is set, and the optimized results are automatically screened to remove the optimization results with unreasonable widths. The screened results are used as the crack mask image frames.

[0030] Furthermore, step S30 specifically includes: performing 3D point cloud mapping based on the structural lidar point cloud frames and the corresponding lidar measurement poses to obtain an overall 3D point cloud map, and performing a cropping process on the overall 3D point cloud map; performing denoising and smoothing processes on the cropped overall 3D point cloud map to obtain an overall structural denoised point cloud map.

[0031] It can be understood that lidar point cloud and pose data are extracted from the collection, and 3D point cloud mapping is performed using the FAST-LIO2 technology to obtain a 3D point cloud map and lidar pose data. Then, the 3D point cloud map is denoised to obtain a clean 3D point cloud map; Specifically, structural lidar point cloud frames and the corresponding lidar measurement poses are extracted from the measured data, and 3D point cloud mapping is performed using the FAST-LIO2 technology to obtain an overall 3D point cloud map and lidar pose data. Then, the overall 3D point cloud map is denoised to obtain an overall structural denoised point cloud map. Please refer to the figure. Specifically, structural lidar point cloud frames and the corresponding lidar measurement poses are extracted from the measured data, and using the FAST-LIO2 technology, the point cloud data obtained by the lidar is stitched together to obtain the stitched overall 3D point cloud map and the lidar measurement pose when collecting data; the stitched overall 3D point cloud map is appropriately cropped according to the position and size of the target structure (in this solution, the concrete structure is taken as an example). The result after cropping is as Figure 3As shown; by using the Statistical Outlier Removal filter in the PCL library to remove the measured noise values (outliers) from the point cloud dataset; by using the Moving Least Square (MLS) algorithm to smooth the point cloud map, an overall structure denoised point cloud map is obtained, as Figure 4 shown.

[0032] Furthermore, it can be understood that the lidar acquisition frequency and the camera acquisition frequency may not be consistent. For example, if the structural image frame acquired by the camera at 1.5 seconds is the key frame of the structural image, at this time, the lidar only acquires data at 1.4 seconds and 1.6 seconds, and the camera pose can only be calibrated by relying on the lidar acquisition pose at the same acquisition moment. Therefore, the camera pose is calibrated at the moment by combining linear interpolation and the lidar-camera extrinsic calibration parameters, and the key shooting pose corresponding to the key shooting moment is calibrated.

[0033] Specifically, in step S40, based on at least two lidar measurement poses, the rotation part is calculated by spherical interpolation, and the position part is combined with linear interpolation to obtain the real-time estimated lidar pose at any moment; Step S50 specifically includes: calibrating the camera pose at the moment according to the real-time estimated lidar pose and the lidar-camera extrinsic calibration parameters, and calculating and obtaining the key shooting pose at the key shooting moment of the camera through coordinate transformation and time synchronization; performing visibility filtering on the point cloud map frame at the target moment based on the current viewing angle of the key shooting pose to obtain the key moment visible point cloud frames corresponding to each camera key shooting moment, and the key moment visible point cloud frames are obtained based on the overall structure denoised point cloud map; projecting the points in the key moment visible point cloud frames onto the image pixel points on the image plane of the corresponding camera for matching, and performing point cloud coloring to obtain the visible point cloud colored point cloud map; projecting the points in the key moment visible point cloud frames onto the pixel points of the crack mask map on the image plane of the corresponding camera for matching, and performing point cloud semantic segmentation to obtain the visible point cloud semantic segmentation point cloud map.

[0034] Furthermore, in step S50, time synchronization is used to determine the real-time estimated lidar pose corresponding to the key shooting moment, and visibility filtering is performed on the overall structure denoised point cloud map by using the Hidden Point Removal (HPR) operator according to the real-time estimated lidar pose corresponding to the key shooting moment to obtain the visible three-dimensional point cloud corresponding to the key shooting moment.

[0035] Optionally, the key shooting pose at the key shooting moment of the camera is calculated by SE(3) spherical interpolation using the lidar pose and the radar-camera extrinsic calibration parameters. The lidar point cloud is transferred to the camera pose coordinate system, and then the points of the point cloud are projected onto the image and the mask respectively for point cloud coloring, obtaining a colored 3D point cloud map of the concrete structure and a point cloud map with crack segmentation information. Optionally, the lidar pose and the camera pose are calibrated based on the timestamp technology. Then, the visibility filtering of the frame is performed on the overall structure denoised point cloud map according to the lidar pose at a certain moment, filtering out the invisible point cloud at this pose, and obtaining the visible 3D point cloud at this moment. Then, the visible 3D point cloud is projected onto the image plane for crack coloring.

[0036] Further, step S60 specifically includes: performing point cloud coloring processing on all key frames of the structural images to obtain an overall 3D colored point cloud map; performing point cloud semantic segmentation processing on the crack mask map frames corresponding to all key frames of the structural images to obtain an overall 3D semantic segmentation point cloud map; Further, the step of "performing point cloud coloring processing on all key frames of the structural images to obtain an overall 3D colored point cloud map" specifically includes: obtaining the associated structural image frames corresponding to each 3D point in all visible point cloud frames at the key moments, where the associated structural image frames have associated pixel points matching the 3D points; among them, the associated weight score corresponding to the 3D point is determined by a weighted scoring method based on the camera angle and distance, and the top N frames with the highest associated weight scores are determined as the associated structural image frames of the 3D point; the RGB colors of the associated pixel points in all the associated structural image frames are fused using the normalized weights to determine the color value of the 3D point; the step of "performing point cloud semantic segmentation processing on the crack mask map frames corresponding to all key frames of the structural images to obtain an overall 3D semantic segmentation point cloud map" specifically includes: obtaining the associated mask map frames corresponding to each 3D point in all visible point cloud frames at the key moments, where the associated mask map frames have associated pixel points matching the 3D points; among them, the associated weight score corresponding to the 3D point is determined by a weighted scoring method based on the camera angle and distance, and the top N frames with the highest associated weight scores are determined as the associated mask map frames of the 3D point; the colors of the associated pixel points in all the associated mask map frames are fused using the normalized weights to determine the semantic information of the 3D point.

[0037] Please refer to Figure 5The flow chart of lidar and camera data fusion coloring is shown as follows. In specific implementation: from the respective lidar pose data during the acquisition of the structured lidar point cloud frame sequence, according to the lidar-camera extrinsic calibration parameters between the lidar and the camera, spherical interpolation is performed through the Euclidean transformation SE(3) to calculate the pose data when the camera acquires data, so as to realize the conversion of the point cloud coordinates to the camera coordinate system; after obtaining the camera pose, before projecting the lidar points onto the image plane of the camera, it is necessary to determine which points are visible from the current camera perspective, which is achieved by using the HPR operator. This operator can efficiently filter out the occluded points in the camera's field of view; using the internal parameters of the camera, project the filtered lidar points determined to be visible by the HPR operator onto the image plane of the camera, and find the corresponding image pixel points for each point in the point cloud based on the minimum error, that is, the coloring of the point cloud in one image is completed; as this process iterates over all image frames, each point in the point cloud map will accumulate color information from multiple perspectives. For each point, color data from multiple frames will be collected; to select the optimal color information for each point, a weighted scoring method based on the camera angle and distance is introduced. For each three-dimensional point, the first N frames with the highest weights are selected as the best perspectives; the RGB colors of these frames are fused using the normalized weights to determine the final color of the point, so as to ensure that the final color assigned to each three-dimensional point is derived from the most reliable and accurately captured perspective, thereby generating a high-quality, colored point cloud map, as Figure 6 shown; operate by replacing the image with a mask map according to the above similar method, and fuse to obtain a three-dimensional point cloud map with crack information marked, as Figure 7 shown. Optionally, N can be 5, or other values such as 8, 10, etc.

[0038] Furthermore, in step S70, the coloring result and the segmentation result are fused together to obtain an overall three-dimensional point cloud labeled map with coloring and crack information marked, as Figure 8 shown.

[0039] Further, step S80 specifically includes: determining crack recognition points on the two-dimensional skeleton image according to the crack mask map frame, calculating the gradient within the domain for measuring the crack position to determine the crack direction vector; locating the left and right edge points relative to the skeleton points by following the direction and its perpendicular direction on the two-dimensional skeleton map, and iteratively tracking to determine the two crack edge points of the measured crack position; projecting the point cloud onto the image, selecting the three-dimensional points that match the crack center point on the image based on the minimum error, then calculating the plane equation at this point in the point cloud map based on this point, generating a coordinate grid within a specified radius based on this plane, and calculating the corresponding coordinates using the plane equation, which will obtain a set of sampled three-dimensional points distributed on the plane; projecting the sampled three-dimensional points onto the two-dimensional image plane according to the internal parameters of the camera, and selecting the three-dimensional points that best match the crack edge points based on the minimum projection error to obtain the specific coordinate values of the two crack edge points.

[0040] Please refer to Figure 9 , in the specific implementation, it mainly includes: calculating the crack skeleton direction: to estimate the crack direction of a given point in the two-dimensional skeleton image, first extract the local neighborhood around this point on the mask map and apply Gaussian smoothing to reduce noise, then use the Sobel operator to calculate the gradient within this neighborhood to determine the direction vector. The average gradient provides the direction vector, and then normalization processing is performed; finding edge points: use the calculated direction vector to trace the edges of the crack. By following the direction and its perpendicular direction, locate the left and right edge points relative to the skeleton points, and iteratively execute this tracing until the edge is detected; sampling 3D points on the plane: to accurately perform the three-dimensional reconstruction of the crack edge, project the overall three-dimensional semantic segmentation point cloud map onto the image through the Inverse Perspective Mapping (IPM) technology, select the 3D points that match the crack center point on the image based on the minimum error, then calculate the plane equation at this point in the point cloud map based on this point, generate a coordinate grid within a specified radius based on this plane, and calculate the corresponding coordinates using the plane equation, which will obtain a set of sampled 3D points distributed on the plane; finding 3D edge points: project the sampled 3D points onto the 2D image plane using the internal parameters of the camera, identify the 3D points corresponding to the actual crack edge by comparing the projected points with the edge points observed in the image, and select the best-matching 3D points based on the minimum projection error to obtain the 3D coordinate values of the two crack edge points; calculating the crack width: the crack width is calculated as the Euclidean distance between the identified left and right edge points in the three-dimensional space.

[0041] Further, before step S20, it includes: constructing a DeepLab-v3+ network model with ResNet as the backbone, determining the initial hyperparameters of the network model, and obtaining a pre-trained model after pre-training; using the labeled training set and test set to train the pre-trained model to obtain a crack semantic segmentation model.

[0042] Specifically, a DeepLab-v3+ network model with ResNet as the backbone is constructed. It uses the PyTorch deep learning framework and takes ResNet as the backbone to achieve crack semantic segmentation. After the network model is configured, the network training starts. The network training includes two parts: pre-training and formal training. First, the network is pre-trained. The crack image dataset for pre-training is divided into a training set and a validation set in an appropriate proportion, and the initial hyperparameters such as batchsize, epoch, and learning rate are determined. Then the pre-training starts. After the training is completed, a pre-trained model is obtained. On the basis of the pre-trained model, the crack image dataset for formal training is divided into a training set and a validation set in an appropriate proportion, and the initial hyperparameters such as batchsize, epoch, start epoch, and learning rate are determined. Then the formal training starts. After the training is completed, a trained model is obtained for subsequent image segmentation.

[0043] The present invention also provides a three-dimensional reconstruction system for concrete cracks based on multi-modal fusion, including a processing device, which is used to implement the steps of the above-mentioned three-dimensional reconstruction method for concrete cracks based on multi-modal fusion.

[0044] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-mentioned three-dimensional reconstruction method for concrete cracks based on multi-modal fusion.

[0045] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A three-dimensional reconstruction method for concrete cracks based on multimodal fusion, characterized in that, It includes the following steps: S10. Collect a sequence of structural image frames and a sequence of structural lidar point cloud frames of the target structure according to a preset acquisition trajectory, and obtain the pose of the acquisition device corresponding to each acquisition moment; S20. Use a crack semantic segmentation model to segment each frame of the structural image frame to obtain a crack mask map frame after crack segmentation; S30. Perform 3D point cloud mapping based on the structural lidar point cloud frame and the corresponding radar measurement pose and process it to obtain an overall structure denoised point cloud map, where the radar measurement pose is obtained based on the external parameters of the device and the pose of the acquisition device at the corresponding acquisition moment; S40. Obtain the radar real-time estimated pose at any moment through spherical interpolation calculation based on at least two of the radar measurement poses, and obtain the corresponding key shooting poses at each camera key shooting moment according to the radar real-time estimated pose and the external parameters of the device. The camera key shooting moment is the shooting moment corresponding to the key frame of the structural image, and the key frame of the structural image is determined from the structural image frames; S50. Obtain the visible 3D point cloud corresponding to the key shooting moment, project the 3D points of the visible 3D point cloud onto the corresponding key frame of the structural image and the crack mask map frame according to the key shooting pose for point cloud coloring and semantic segmentation to obtain a visible point cloud colored point cloud map and a visible point cloud semantic segmentation point cloud map; S60. Perform point cloud coloring and semantic segmentation processing on all the key frames of the structural image and the crack mask map frames to obtain an overall 3D colored point cloud map and an overall 3D semantic segmentation point cloud map; S70. Fuse the overall 3D colored point cloud map and the overall 3D semantic segmentation point cloud map to obtain an overall 3D point cloud labeled map with crack information and coloring; S80. Perform crack three-dimensional attribute measurement based on the three-dimensional point cloud labeled map to obtain crack three-dimensional geometric information, where the crack three-dimensional geometric information includes crack width information, crack position, and crack orientation information.

2. The three-dimensional reconstruction method of concrete cracks based on multi-modal fusion according to claim 1, characterized in that Step S20 specifically includes: Use a crack semantic segmentation model to segment each frame of the structural image frame to obtain an initial mask map after crack segmentation. The initial mask map has black and white markings and semantic information; Cut the initial mask map into several feature pictures, process the feature pictures to extract a crack skeleton map, and the crack skeleton map has pixel colors representing the distance values of each pixel point from the nearest edge; Convert the crack skeleton map into a set of pixel coordinate point sets, perform downsampling, and extract pixel key points; Sort and downsample the pixel key point list, and mark the feature pictures based on a quantity threshold; Use the markings of the feature pictures as prompt points for the large model, import the structural image frame corresponding to the feature picture into the large model for crack mask optimization, and perform screening using a width optimization threshold to obtain the crack mask map frame.

3. The three-dimensional reconstruction method of concrete cracks based on multimodal fusion according to claim 1, characterized in that, Step S30 specifically includes: Perform 3D point cloud mapping based on the structural lidar point cloud frame and the corresponding radar measurement pose to obtain an overall 3D point cloud map; Perform cropping processing on the overall 3D point cloud map; Denoise and smooth the overall three-dimensional point cloud map after cropping to obtain an overall structure denoised point cloud map.

4. The three-dimensional reconstruction method of concrete cracks based on multimodal fusion according to claim 1, characterized in that In step S40: Based on at least two of the radar measurement poses, calculate the rotation part through spherical interpolation, and combine linear interpolation for the position part to obtain the real-time estimated pose of the radar at any moment; Calibrate the camera pose at each moment according to the real-time estimated pose of the radar and the radar-camera extrinsic calibration parameters, and calculate and obtain the key shooting pose at the key shooting moment of the camera through coordinate transformation and time synchronization; Specifically included in step S50: Perform visibility filtering on the point cloud map frames at the target moment based on the current perspective of the key shooting pose to obtain the key moment visible point cloud frames corresponding to each key shooting moment of the camera, and the key moment visible point cloud frames are obtained based on the overall structure denoised point cloud map; Project the point cloud in the key moment visible point cloud frame to the image pixel points on the image plane of the corresponding camera for matching, and perform point cloud coloring to obtain a visible point cloud colored point cloud map; Project the point cloud in the key moment visible point cloud frame to the pixel points of the crack mask map on the image plane of the corresponding camera for matching, and perform point cloud semantic segmentation to obtain a visible point cloud semantic segmentation point cloud map; Specifically included in step S60: Perform point cloud coloring processing on all key frames of the structure images to obtain an overall three-dimensional colored point cloud map; Perform point cloud semantic segmentation processing on the crack mask map frames corresponding to all key frames of the structure images to obtain an overall three-dimensional semantic segmentation point cloud map.

5. The three-dimensional reconstruction method of concrete cracks based on multi-modal fusion according to claim 4, characterized in that The step "Perform point cloud coloring processing on all key frames of the structure images to obtain an overall three-dimensional colored point cloud map" specifically includes: obtaining the associated structure image frames corresponding to each three-dimensional point in all key moment visible point cloud frames, and the associated structure image frames have associated pixel points that match the three-dimensional points; wherein, determine the associated weight score corresponding to the three-dimensional points using a weighted scoring method based on camera angle and distance, and determine the top N frames with the highest associated weight score as the associated structure image frames of the three-dimensional points; use the normalized weights to fuse the RGB colors of the associated pixel points in all the associated structure image frames to determine the color value of the three-dimensional points. The step "Perform point cloud semantic segmentation processing on the crack mask map frames corresponding to all key frames of the structure images to obtain an overall three-dimensional semantic segmentation point cloud map" specifically includes: obtaining the associated mask map frames corresponding to each three-dimensional point in all key moment visible point cloud frames, and the associated mask map frames have associated pixel points that match the three-dimensional points; wherein, determine the associated weight score corresponding to the three-dimensional points using a weighted scoring method based on camera angle and distance, and determine the top N frames with the highest associated weight score as the associated mask map frames of the three-dimensional points; use the normalized weights to fuse the colors of the associated pixel points in all the associated mask map frames to determine the semantic information of the three-dimensional points.

6. The three-dimensional reconstruction method of concrete cracks based on multimodal fusion according to claim 1, characterized in that, In step S50, time synchronization is used to determine the real-time estimated pose of the radar corresponding to the key shooting moment. Based on the real-time estimated pose of the radar corresponding to the key shooting moment, visibility filtering is performed on the denoised point cloud map of the overall structure through the HRP operator to obtain the visible three-dimensional point cloud corresponding to the key shooting moment.

7. The three-dimensional reconstruction method of concrete cracks based on multimodal fusion according to claim 1, characterized in that, Step S80 specifically includes: Determine the crack recognition points on the two-dimensional skeleton image according to the crack mask map frame, and calculate the gradient within the domain of the measured crack position to determine the crack direction vector; On the two-dimensional skeleton map, locate the left and right edge points relative to the skeleton points by following the direction and its perpendicular direction, and iteratively track to determine the two crack edge points of the measured crack position; Project the point cloud onto the image, select the three-dimensional points that match the crack center point on the image based on the minimum error, then calculate the plane equation at this point based on the point cloud map, and generate a coordinate grid within the specified radius based on the plane equation. Calculate the corresponding coordinates using the plane equation, which will obtain a set of sampled three-dimensional points distributed on the plane; Project the sampled three-dimensional points onto the two-dimensional image plane according to the internal parameters of the camera, and select the three-dimensional points that best match the crack edge points based on the minimum projection error to obtain the specific coordinate values of the two crack edge points.

8. The three-dimensional reconstruction method of concrete cracks based on multimodal fusion according to claim 1, characterized in that, Before step S20, it includes: Construct a network model of DeepLab-v3+ with ResNet as the backbone, determine the initial hyperparameters of the network model, and obtain a pre-trained model after pre-training; Use the labeled training set and test set to train the pre-trained model to obtain a crack semantic segmentation model.

9. A three-dimensional reconstruction system for concrete cracks based on multimodal fusion, characterized in that, It includes a processing device, and the processing device is used to implement the steps of the three-dimensional reconstruction method of concrete cracks based on multi-modal fusion according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the three-dimensional reconstruction method of concrete cracks based on multi-modal fusion according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Laser and visual information fused robust positioning and mapping method and system

    CN112985416A

  • Three-dimensional model measurement method for concrete structure crack and related device

    CN116152697A

  • Building three-dimensional reconstruction and crack detection method and system, terminal and medium

    CN118570411A

  • Tunnel crack detection method and tunnel crack detection system

    CN119510293A

  • Civil engineering structure defect detection method and system based on image processing technology

    CN119810319A

Cited By

  • Aircraft engine part fatigue crack propagation identification method based on machine vision

    CN120509325A

  • Fatigue crack growth identification method for aero-engine components based on machine vision

    CN120509325B

  • Crack feature extraction and quantification method and device based on point cloud big language model

    CN121259617A

  • Crack feature extraction and quantification method and device based on point cloud large language model

    CN121259617B

  • Modal fusion bridge crack 4D point cloud microscopic motion enhancement method and system

    CN121708213A