A three-dimensional reconstruction method, system and medium for concrete cracks based on multimodal fusion

Through the multimodal fusion method, combined with structural images and lidar point cloud, the problem of difficulty in realizing three-dimensional crack detection in the prior art is solved, and high-precision simultaneous detection of crack width, position and direction information is achieved.

CN120259573BActive Publication Date: 2025-08-29CENT SOUTH UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510751650.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-29
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

In the prior art, it is difficult to realize the three-dimensional crack detection of the structure by reconstructing the concrete crack skeleton based on single-frame image data and measuring the width. In addition, the existing three-dimensional reconstruction methods lack the density, accuracy and robustness of local details of point clouds, and cannot meet the detection needs of microstructures.

Method used

The multimodal fusion method is adopted, combining structural image frame sequences and structural lidar point cloud frame sequences, and the overall structural denoising point cloud map is obtained through the crack semantic segmentation model and three-dimensional point cloud map construction, and the key shooting postures are calculated in combination with spherical interpolation and time synchronization to realize the coloring and semantic segmentation of visible three-dimensional point clouds. Finally, the overall three-dimensional point cloud marking map with crack information is obtained to measure the crack three-dimensional geometric information.

Benefits of technology

High-precision three-dimensional crack detection is realized, and the width information, position and direction information of the crack can be obtained at the same time, adapted to crack detection of various three-dimensional structures, and generated a dense and high-precision overall three-dimensional point cloud marking map.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259573B_ABST
    Figure CN120259573B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, and medium for three-dimensional reconstruction of concrete cracks based on multimodal fusion, which relates to the field of structural engineering detection technology and includes the following steps: obtaining a structural image frame sequence and a structural lidar point cloud frame sequence of a target structure; obtaining a crack mask frame after crack segmentation; obtaining a denoised point cloud map of the overall structure; obtaining a colored point cloud map of a visible point cloud and a semantically segmented point cloud map of a visible point cloud; obtaining an overall three-dimensional colored point cloud map and an overall three-dimensional semantically segmented point cloud map; obtaining an overall three-dimensional point cloud marker map; and performing three-dimensional crack attribute measurement based on the three-dimensional point cloud marker map to obtain three-dimensional crack geometric information. The method provided by the present invention designs a multi-frame and multimodal fusion crack structure reconstruction framework that can adapt to crack detection in various three-dimensional structures and can simultaneously detect crack width information, crack location, and crack direction information based on the overall three-dimensional point cloud marker map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of structural engineering detection technology, and in particular to a method, system, and medium for three-dimensional reconstruction of concrete cracks based on multimodal fusion. Background Art

[0002] As the service life of concrete infrastructure such as bridges and tunnels increases, the problem of surface cracks caused by structural aging is becoming increasingly prominent. As cracks are key indicators for assessing structural stress state, safety, and durability, their precise detection and three-dimensional quantitative analysis have become core requirements for health monitoring.

[0003] In existing technologies, crack detection is mainly based on manual inspection and computer vision technology. Among them, manual inspection relies on visual observation, which has defects such as low efficiency, strong subjectivity, and high risks of high-altitude operations. In addition, manual inspection cannot obtain three-dimensional geometric information of cracks, which makes it difficult to meet the comprehensiveness and accuracy requirements of modern engineering for structural health assessment. Computer vision technology is mainly based on two-dimensional images for crack segmentation and detection, and cannot directly obtain three-dimensional information such as crack size, location, and topology from two-dimensional crack information.

[0004] Currently, for 3D crack detection, some researchers have used single-frame image data to reconstruct the crack skeleton and measure crack width. However, using only a single image and its associated processing techniques is insufficient to capture the complete 3D spatial details of the crack, making 3D crack detection difficult. Furthermore, to generate comprehensive 3D crack information, researchers have expanded detection from a single local image to the entire structure. 3D reconstruction methods are generally divided into two categories: Structure from Motion (SfM)-based and Simultaneous Localization and Mapping (SLAM)-based methods. Current 3D reconstruction methods still lack the density, accuracy, and robustness to reconstruct local details in point clouds, failing to meet the detection requirements for small crack structures. Furthermore, existing methods typically separate crack width measurement from crack localization: crack width measurement is performed on a 2D image, while crack localization is performed on a 3D point cloud. This makes it impossible to simultaneously and automatically quantify the 3D position and width of cracks.

[0005] In view of this, it is necessary to propose a three-dimensional reconstruction method, system and medium for concrete cracks based on multimodal fusion to solve or at least alleviate some of the above defects. Summary of the Invention

[0006] The main purpose of the present invention is to provide a method, system and medium for three-dimensional reconstruction of concrete cracks based on multimodal fusion, aiming to solve the technical problem in the existing technology that the crack skeleton is reconstructed and the crack width is measured by relying solely on image data under a single frame, which makes it difficult to achieve three-dimensional crack detection of the structure.

[0007] To achieve the above objectives, the present invention provides a method for 3D reconstruction of concrete cracks based on multimodal fusion, comprising the following steps:

[0008] S10, acquiring a structural image frame sequence and a structural lidar point cloud frame sequence of the target structure according to a preset acquisition trajectory, and obtaining the acquisition device posture corresponding to each acquisition moment;

[0009] S20, using a crack semantic segmentation model to segment each structural image frame to obtain a crack mask image frame after crack segmentation;

[0010] S30, constructing a three-dimensional point cloud map based on the structural lidar point cloud frame and the corresponding radar measurement pose and processing it to obtain an overall structural denoised point cloud map, wherein the radar measurement pose is obtained based on the device external parameters and the acquisition device pose at the corresponding acquisition time;

[0011] S40, obtaining a real-time estimated radar pose at any time by spherical interpolation calculation based on at least two radar measured poses, and obtaining a key shooting pose corresponding to each camera at a key shooting moment based on the real-time estimated radar pose and device external parameters, where the key shooting moment of the camera is the shooting moment corresponding to a key frame of the structural image, and the key frame of the structural image is determined from the structural image frame;

[0012] S50, obtaining a visible 3D point cloud corresponding to the key shooting moment, and projecting the 3D points of the visible 3D point cloud onto the corresponding structural image key frame and crack mask frame according to the key shooting posture, performing point cloud coloring and semantic segmentation to obtain a visible point cloud coloring point cloud map and a visible point cloud semantic segmentation point cloud map;

[0013] S60, performing point cloud coloring and semantic segmentation processing on all structural image key frames and crack mask frames to obtain an overall three-dimensional colored point cloud map and an overall three-dimensional semantically segmented point cloud map;

[0014] S70, fusing the overall three-dimensional colored point cloud map and the overall three-dimensional semantic segmentation point cloud map to obtain an overall three-dimensional point cloud labeled map with crack information and coloring;

[0015] S80: Measure the three-dimensional properties of the crack based on the three-dimensional point cloud marker map to obtain three-dimensional geometric information of the crack, where the three-dimensional geometric information of the crack includes information on the width, position, and direction of the crack.

[0016] Furthermore, step S20 specifically includes:

[0017] The crack semantic segmentation model is used to segment each structural image frame to obtain an initial mask image after crack segmentation. The initial mask image has black and white marks and semantic information.

[0018] The initial mask image is cut into several feature images, and the feature images are processed to extract a crack skeleton image, where the crack skeleton image has a pixel color representing the distance value of each pixel point to the nearest edge;

[0019] Convert the crack skeleton map into a set of pixel coordinate points, perform downsampling and extract pixel key points;

[0020] Sort and downsample the pixel keypoint list, and mark the feature image based on the quantity threshold;

[0021] According to the mark of the feature image as the prompt point of the large model, the structural image frame corresponding to the feature image is imported into the large model for crack mask optimization, and the crack mask frame is obtained by screening using the width optimization threshold.

[0022] Furthermore, step S30 specifically includes:

[0023] Based on the structured lidar point cloud frame and the corresponding lidar measurement pose, a 3D point cloud map is constructed to obtain an overall 3D point cloud map.

[0024] Crop the entire 3D point cloud map;

[0025] The overall three-dimensional point cloud map after cropping is denoised and smoothed to obtain the overall structure denoised point cloud map.

[0026] Furthermore, in step S40:

[0027] Based on the poses measured by at least two radars, the rotation part is calculated by spherical interpolation, and the position part is combined with linear interpolation to obtain the real-time estimated pose of the radar at any time;

[0028] The camera pose is calibrated at all times based on the real-time estimated pose of the radar and the radar-camera extrinsic calibration parameters. The key shooting pose of the camera at the key shooting moment is obtained through coordinate transformation and time synchronization calculation.

[0029] Furthermore, step S50 specifically includes:

[0030] Based on the current perspective of the key shooting pose, the visibility of the target moment point cloud frame is filtered to obtain the key moment visible point cloud frame corresponding to the key shooting moment of each camera. The key moment visible point cloud frame is obtained based on the overall structure denoised point cloud map;

[0031] Project the point cloud in the visible point cloud frame at the key moment onto the image pixel points on the image plane of the corresponding camera for matching, and perform point cloud coloring to obtain a colored point cloud map of the visible point cloud;

[0032] The point cloud in the visible point cloud frame at the key moment is projected onto the pixel points of the crack mask on the image plane of the corresponding camera for matching, and point cloud semantic segmentation is performed to obtain a semantic segmentation point cloud map of the visible point cloud.

[0033] Furthermore, in step S50:

[0034] Time synchronization is used to determine the real-time estimated radar pose corresponding to the key shooting moment. According to the real-time estimated radar pose corresponding to the key shooting moment, visibility filtering is performed on the overall structure denoised point cloud image through the HRP operator to obtain the visible 3D point cloud corresponding to the key shooting moment.

[0035] Furthermore, step S60 specifically includes:

[0036] Perform point cloud coloring on all key frames of the structural image to obtain the overall three-dimensional colored point cloud map;

[0037] Perform point cloud semantic segmentation processing on the crack mask frames corresponding to all structural image key frames to obtain the overall three-dimensional semantic segmentation point cloud map;

[0038] Furthermore, the step of "performing point cloud coloring on all structural image key frames to obtain an overall three-dimensional colored point cloud map" specifically includes: obtaining associated structural image frames corresponding to each three-dimensional point in all visible point cloud frames at key moments, the associated structural image frames having associated pixel points matching the three-dimensional points; wherein, a weighted scoring method based on camera angle and distance is used to determine an associated weight score corresponding to the three-dimensional point, and the top N frames with the highest associated weight scores are determined as associated structural image frames of the three-dimensional point; and using normalized weights to fuse the RGB colors of associated pixel points in all associated structural image frames to determine the color value of the three-dimensional point.

[0039] The step of "performing point cloud semantic segmentation processing on the crack mask frames corresponding to all structural image key frames to obtain an overall three-dimensional semantic segmentation point cloud map" specifically includes: obtaining associated mask frames corresponding to each three-dimensional point in all visible point cloud frames at key moments, the associated mask frames having associated pixel points matching the three-dimensional points; wherein, a weighted scoring method based on camera angle and distance is used to determine the associated weight score corresponding to the three-dimensional point, and the top N frames with the highest associated weight score are determined as the associated mask frames of the three-dimensional point; and the colors of the associated pixel points in all associated mask frames are fused using normalized weights to determine the semantic information of the three-dimensional point.

[0040] Furthermore, step S70 specifically includes:

[0041] Determine crack identification points on the two-dimensional skeleton image according to the crack mask frame, calculate the gradient in the area where the crack position is measured and determine the crack direction vector;

[0042] On the two-dimensional skeleton graph, the left and right edge points are located relative to the skeleton point by following the direction and its perpendicular direction, and the two crack edge points of the measured crack position are determined by iterative tracking;

[0043] Project the point cloud onto the image, select a 3D point that matches the crack center on the image based on the minimum error, then calculate the plane equation at that point on the point cloud. Based on this plane, generate a coordinate grid within a specified radius, and use the plane equation to calculate the corresponding coordinates. This will result in a set of sampled 3D points distributed on the plane.

[0044] The sampled 3D points are projected onto the 2D image plane according to the intrinsic parameters of the camera. The 3D points that best match the crack edge points are selected based on the minimum projection error to obtain the specific coordinate values ​​of the two crack edge points.

[0045] Furthermore, before step S20, the method includes:

[0046] Build a DeepLab-v3+ network model with ResNet as the backbone, determine the initial hyperparameters of the network model, and perform pre-training to obtain a pre-trained model;

[0047] The pre-trained model is trained using the labeled training set and test set to obtain the crack semantic segmentation model.

[0048] The present invention also provides a three-dimensional reconstruction system for concrete cracks based on multimodal fusion, including a processing device, which is used to implement the steps of the above-mentioned three-dimensional reconstruction method for concrete cracks based on multimodal fusion.

[0049] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the above-mentioned method for three-dimensional reconstruction of concrete cracks based on multimodal fusion are implemented.

[0050] The present invention provides a three-dimensional reconstruction method for concrete cracks based on multimodal fusion. First, a structural image frame sequence and a structural lidar point cloud frame sequence of a target structure are collected, and the acquisition device posture corresponding to each acquisition moment is obtained. A three-dimensional point cloud is constructed and processed based on the structural lidar point cloud frame and the corresponding radar measurement posture to obtain a denoised point cloud map of the overall structure, and a crack semantic segmentation model is used to segment each structural image frame to obtain a crack mask frame after crack segmentation. Then, a structural image frame is determined from a structural image key frame sequence and its corresponding key shooting moment is obtained. A real-time estimated radar posture at any moment is obtained by spherical interpolation calculation based on at least two radar measurement postures. The key shooting posture corresponding to the key shooting moment is obtained according to the real-time estimated radar posture and the external parameters of the device. Based on the real-time estimated radar posture, the key shooting posture corresponding to the key shooting moment is accurately positioned when the shooting frame rate of the radar and the camera are not synchronized. Then, after obtaining the visible three-dimensional point cloud corresponding to the key shooting moment, the three-dimensional points of the visible three-dimensional point cloud are projected onto the corresponding structural image key frame and crack mask frame according to the key shooting posture for point cloud projection. Color and semantic segmentation are performed to obtain a colored point cloud map of visible point clouds and a semantically segmented point cloud map of visible point clouds; then, point cloud coloring and semantic segmentation are performed on all structural image key frames and crack mask frames to obtain an overall three-dimensional colored point cloud map and an overall three-dimensional semantically segmented point cloud map; finally, the overall three-dimensional colored point cloud map and the overall three-dimensional semantically segmented point cloud map are fused to obtain an overall three-dimensional point cloud labeled map with crack information and color; three-dimensional attribute measurement of cracks is performed based on the three-dimensional point cloud labeled map to obtain three-dimensional geometric information of cracks, where the three-dimensional geometric information of cracks includes crack width information, crack position and crack direction information, thereby realizing three-dimensional crack detection of the structure; the three-dimensional reconstruction method of concrete cracks based on multimodal fusion provided by the present invention designs a multi-frame and multi-modal fusion crack structure reconstruction framework when performing three-dimensional crack detection, combines the lidar point cloud and the structural plane image to generate a dense and high-precision overall three-dimensional point cloud labeled map with crack semantics and real-world size, which can adapt to crack detection of various three-dimensional structures, and can realize simultaneous detection of crack width information, crack position and crack direction information based on the overall three-dimensional point cloud labeled map. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on the structures shown in these drawings without inventive effort. The dotted boxes in the drawings represent the processing process.

[0052] Figure 1Schematic diagram of a flow chart of a method for 3D reconstruction of concrete cracks based on multimodal fusion in one embodiment of the present invention;

[0053] Figure 2 Schematic diagram of the process of obtaining a crack mask frame in step S20 in one embodiment of the present invention;

[0054] Figure 3 Schematic diagram of the three-dimensional structure of the entire three-dimensional point cloud map after the cropping process in step S30 in one embodiment of the present invention; wherein a is one schematic diagram of the structure of the entire three-dimensional point cloud map, and b is a second schematic diagram of the structure of the entire three-dimensional point cloud map;

[0055] Figure 4 The denoising and smoothing process is performed in step S30 in one embodiment of the present invention to obtain a denoised point cloud image of the overall structure; wherein a is one schematic diagram of the overall structure point cloud, and b is a second schematic diagram of the overall structure point cloud;

[0056] Figure 5 A schematic diagram of a process for obtaining an overall three-dimensional point cloud marked map in one embodiment of the present invention;

[0057] Figure 6 The overall three-dimensional colored point cloud map in step S60 in one embodiment of the present invention;

[0058] Figure 7 The overall three-dimensional semantic segmentation point cloud map (the three-dimensional point cloud map with crack information marks) in step S60 in one embodiment of the present invention;

[0059] Figure 8 The entire three-dimensional point cloud marked map in step S70 in one embodiment of the present invention (the colored three-dimensional point cloud map with crack information marks);

[0060] Figure 9 FIG. 1 is a schematic diagram of a process for measuring three-dimensional properties of cracks in step S80 in one embodiment of the present invention.

[0061] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0062] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0064] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0065] In addition, the descriptions of "first", "second", etc. in the present invention are for descriptive purposes only and should not be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" or "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the fact that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0066] Please refer to the attached Figure 1 The present invention provides a three-dimensional reconstruction method for concrete cracks based on multimodal fusion, comprising the following steps:

[0067] S10, acquiring a structural image frame sequence and a structural lidar point cloud frame sequence of the target structure according to a preset acquisition trajectory, and obtaining the acquisition device posture corresponding to each acquisition moment;

[0068] S20, using a crack semantic segmentation model to segment each structural image frame to obtain a crack mask image frame after crack segmentation;

[0069] S30, constructing a three-dimensional point cloud map based on the structural lidar point cloud frame and the corresponding radar measurement pose and processing it to obtain an overall structural denoised point cloud map, wherein the radar measurement pose is obtained based on the device external parameters and the acquisition device pose at the corresponding acquisition time;

[0070] S40, obtaining a real-time estimated radar pose at any time by spherical interpolation calculation based on at least two radar measured poses, and obtaining a key shooting pose corresponding to each camera at a key shooting moment based on the real-time estimated radar pose and device external parameters, where the key shooting moment of the camera is the shooting moment corresponding to a key frame of the structural image, and the key frame of the structural image is determined from the structural image frame;

[0071] S50, obtaining a visible 3D point cloud corresponding to the key shooting moment, and projecting the 3D points of the visible 3D point cloud onto the corresponding structural image key frame and crack mask frame according to the key shooting posture, performing point cloud coloring and semantic segmentation to obtain a visible point cloud coloring point cloud map and a visible point cloud semantic segmentation point cloud map;

[0072] S60, performing point cloud coloring and semantic segmentation processing on all structural image key frames and crack mask frames to obtain an overall three-dimensional colored point cloud map and an overall three-dimensional semantically segmented point cloud map;

[0073] S70, fusing the overall three-dimensional colored point cloud map and the overall three-dimensional semantic segmentation point cloud map to obtain an overall three-dimensional point cloud labeled map with crack information and coloring;

[0074] S80: Measure the three-dimensional properties of the crack based on the three-dimensional point cloud marker map to obtain three-dimensional geometric information of the crack, where the three-dimensional geometric information of the crack includes information on the width, position, and direction of the crack.

[0075] The present invention provides a three-dimensional reconstruction method for concrete cracks based on multimodal fusion. First, a structural image frame sequence and a structural lidar point cloud frame sequence of a target structure are collected, and the acquisition device posture corresponding to each acquisition moment is obtained. A three-dimensional point cloud is constructed and processed based on the structural lidar point cloud frame and the corresponding radar measurement posture to obtain a denoised point cloud map of the overall structure, and a crack semantic segmentation model is used to segment each structural image frame to obtain a crack mask frame after crack segmentation. Then, a structural image frame is determined from a structural image key frame sequence and its corresponding key shooting moment is obtained. A real-time estimated radar posture at any moment is obtained by spherical interpolation calculation based on at least two radar measurement postures. The key shooting posture corresponding to the key shooting moment is obtained according to the real-time estimated radar posture and the external parameters of the device. Based on the real-time estimated radar posture, the key shooting posture corresponding to the key shooting moment is accurately positioned when the shooting frame rate of the radar and the camera are not synchronized. Then, after obtaining the visible three-dimensional point cloud corresponding to the key shooting moment, the three-dimensional points of the visible three-dimensional point cloud are projected onto the corresponding structural image key frame and crack mask frame according to the key shooting posture for point cloud projection. Color and semantic segmentation are performed to obtain a colored point cloud map of visible point clouds and a semantically segmented point cloud map of visible point clouds; then, point cloud coloring and semantic segmentation are performed on all structural image key frames and crack mask frames to obtain an overall three-dimensional colored point cloud map and an overall three-dimensional semantically segmented point cloud map; finally, the overall three-dimensional colored point cloud map and the overall three-dimensional semantically segmented point cloud map are fused to obtain an overall three-dimensional point cloud labeled map with crack information and coloring, and three-dimensional attribute measurement of cracks is performed based on the three-dimensional point cloud labeled map to obtain three-dimensional geometric information of cracks, where the three-dimensional geometric information of cracks includes crack width information, crack position and crack direction information, thereby realizing three-dimensional crack detection of the structure; the three-dimensional reconstruction method of concrete cracks based on multimodal fusion provided by the present invention designs a multi-frame and multimodal fusion crack structure reconstruction framework when performing three-dimensional crack detection, combines the lidar point cloud and the structural plane image to generate a dense and high-precision overall three-dimensional point cloud labeled map with crack semantics and real-world size, which can adapt to crack detection of various three-dimensional structures, and can realize simultaneous detection of crack width information, crack position and crack direction information based on the overall three-dimensional point cloud labeled map.

[0076] As can be understood, structural image keyframes refer to representative image frames selected during the data processing process. A complete set of selected structural keyframes constitutes a core image set that fully characterizes surface crack information for the target structure. As can be understood, the method of the present invention is primarily used for detecting three-dimensional cracks, such as curved surface cracks.

[0077] Furthermore, step S20 specifically includes: using a crack semantic segmentation model to segment each frame of the structural image frame to obtain an initial mask map after crack segmentation, the initial mask map having black and white marks and semantic information (the structural image frame has color information, and the initial mask map is a black and white image with black and white marks, and has crack semantic information); cutting the initial mask map into several feature images, processing the feature images to extract a crack skeleton map, the crack skeleton map having pixel colors that represent the distance value of each pixel point from the nearest edge; converting the crack skeleton map into a set of pixel coordinate points, extracting pixel key points after downsampling; sorting and downsampling the pixel key point list, and marking the feature image based on a quantity threshold; using the mark of the feature image as a prompt point of the large model, importing the structural image frame corresponding to the feature image into the large model for crack mask optimization, and using the width optimization threshold for screening to obtain the crack mask map frame.

[0078] It can be understood that in the solution of the present invention, the captured structural image frame sequence is segmented using the deep learning network DeepLab-v3+, and the segmentation results are optimized using the large model SAM. Figure 2 As shown, specifically, first, the image, lidar point cloud and posture data of the cracked structure are collected, and then the structural image frame sequence is extracted from the obtained rosbag data; the structural image frame is segmented by the trained DeepLab-v3+ deep learning network model to obtain the initial mask map (Mask) after crack segmentation, where the initial mask map is a black and white image, and the black pixels on the initial mask map represent pixels that are not cracks, and the white pixels represent pixels that are cracks; the initial mask map is cut into several small images, and the crack skeleton map is extracted from the cropped initial mask map through EDT (Euclidean Distance Transform). The crack skeleton map reflects the distance value of each pixel point to the nearest edge, and the pixel color of the crack skeleton map represents the distance change; the pixel points of the crack skeleton map are converted into a set of coordinate points, downsampled for a given threshold, selected prompt points, and then contour detection is performed, each contour is traversed, and pixel key points are extracted from it; the pixel key point list is sorted and downsampled, and after selecting a given number of feature points, the corresponding positions of the original image crop are marked; the large model Segment is called Anything Model performs crack mask optimization, sets a width optimization result evaluation program, and automatically screens the optimized results to remove optimization results with unreasonable widths. The screened results are used as crack mask frames.

[0079] Furthermore, step S30 specifically includes: performing three-dimensional point cloud mapping based on the structural lidar point cloud frame and the corresponding radar measurement posture to obtain an overall three-dimensional point cloud map, and cropping the overall three-dimensional point cloud map; denoising and smoothing the cropped overall three-dimensional point cloud map to obtain an overall structural denoised point cloud map.

[0080] Understandably, the LiDAR point cloud and pose data are extracted from the acquisition, and the 3D point cloud is mapped using FAST-LIO2 technology to obtain a 3D point cloud map and LiDAR pose data. The 3D point cloud map is then subjected to noise reduction processing to obtain a clean 3D point cloud map.

[0081] Specifically, the structural lidar point cloud frames and the corresponding radar measurement poses are extracted from the measured data, and the FAST-LIO2 technology is used to construct the 3D point cloud map to obtain the overall 3D point cloud map and lidar pose data. The overall 3D point cloud map is then subjected to denoising to obtain the overall structural denoised point cloud map. Please refer to the figure, specifically, the structural lidar point cloud frames and the corresponding radar measurement poses are extracted from the measured data, and the FAST-LIO2 technology is used to stitch the point cloud data obtained by the lidar together to obtain the stitched overall 3D point cloud map and the radar measurement pose of the lidar when collecting data; the stitched overall 3D point cloud map is appropriately cropped according to the position and size of the target structure (this solution takes the concrete structure as an example), and the cropped result is as follows: Figure 3 As shown in the figure, the measurement noise values ​​(outliers) are removed from the point cloud dataset by using the Statistical OutlierRemoval filter in the PCL library; the point cloud image is smoothed by using the Moving Least Square (MLS) algorithm to obtain the overall structure denoised point cloud image, as shown in the figure. Figure 4 shown.

[0082] Furthermore, it is understandable that the lidar acquisition frequency may be inconsistent with the camera acquisition frequency. For example, if the structural image frame acquired by the camera at 1.5 seconds is the structural image key frame, the lidar only acquires data from 1.4 seconds and 1.6 seconds, and the camera shooting posture can only be calibrated based on the radar acquisition posture at the same acquisition moment. Therefore, the camera posture is calibrated at all times by combining linear interpolation and radar-camera extrinsic calibration parameters to set the key shooting posture corresponding to the key shooting moment for calibration.

[0083] Specifically, in step S40, based on the poses measured by at least two radars, the rotation part is calculated by spherical interpolation, and the position part is combined with linear interpolation to obtain the real-time estimated pose of the radar at any time;

[0084] Step S50 specifically includes: calibrating the camera pose at each moment according to the real-time estimated pose of the radar and the radar-camera extrinsic calibration parameters, and obtaining the key shooting pose of the camera at the key shooting moment through coordinate transformation and time synchronization calculation; performing visibility filtering on the point cloud frame at the target moment based on the current viewing angle of the key shooting pose to obtain the key moment visible point cloud frame corresponding to the key shooting moment of each camera, and obtaining the key moment visible point cloud frame based on the overall structure denoised point cloud map; projecting the point cloud in the key moment visible point cloud frame onto the image pixel points on the image plane of the corresponding camera for matching, performing point cloud coloring to obtain a visible point cloud colored point cloud map; projecting the point cloud in the key moment visible point cloud frame onto the pixel points of the crack mask map on the image plane of the corresponding camera for matching, and performing point cloud semantic segmentation to obtain a visible point cloud semantic segmentation point cloud map.

[0085] Furthermore, in step S50, time synchronization is used to determine the real-time estimated radar pose corresponding to the key shooting moment. Based on the real-time estimated radar pose corresponding to the key shooting moment, visibility filtering is performed on the overall structure denoised point cloud image using a hidden point removal (HPR) operator to obtain a visible 3D point cloud corresponding to the key shooting moment.

[0086] Optionally, the key camera pose at the key shooting moment is calculated by combining the lidar pose with the radar-camera extrinsic calibration parameters through SE (3) spherical interpolation, and the lidar point cloud is transferred to the camera pose coordinate system. The points of the point cloud are then projected onto the image and mask respectively for point cloud coloring, thereby obtaining a colored concrete structure 3D point cloud map and a point cloud map with crack segmentation information. Optionally, the radar pose and camera pose are calibrated based on timestamp technology. Then, according to the radar pose at a certain moment, the visibility of the frame is filtered in the overall structure denoised point cloud map, and the invisible point cloud at that pose is filtered out to obtain the visible 3D point cloud at that moment. The visible 3D point cloud is then projected onto the image plane for crack coloring.

[0087] Furthermore, step S60 specifically includes: performing point cloud coloring processing on all structural image key frames to obtain an overall three-dimensional colored point cloud map; performing point cloud semantic segmentation processing on the crack mask image frames corresponding to all structural image key frames to obtain an overall three-dimensional semantic segmentation point cloud map;

[0088] Furthermore, the step of "performing point cloud coloring on all structural image key frames to obtain an overall 3D colored point cloud map" specifically includes: obtaining associated structural image frames corresponding to each 3D point in all visible point cloud frames at key moments, wherein the associated structural image frames have associated pixel points that match the 3D points; wherein, a weighted scoring method based on camera angle and distance is used to determine the associated weight scores corresponding to the 3D points, and the top N frames with the highest associated weight scores are determined as the associated structural image frames of the 3D points; using normalized weights to fuse the RGB colors of the associated pixel points in all associated structural image frames to determine the color value of the 3D points; step The step "performing point cloud semantic segmentation processing on the crack mask frames corresponding to all structural image key frames to obtain an overall three-dimensional semantic segmentation point cloud map" specifically includes: obtaining the associated mask frames corresponding to each three-dimensional point in the visible point cloud frames at all key moments, the associated mask frames having associated pixel points matching the three-dimensional points; wherein, a weighted scoring method based on camera angle and distance is used to determine the associated weight score corresponding to the three-dimensional point, and the top N frames with the highest associated weight score are determined as the associated mask frames of the three-dimensional point; and the colors of the associated pixel points in all the associated mask frames are fused using normalized weights to determine the semantic information of the three-dimensional point.

[0089] Please refer to Figure 5 The flowchart of the laser radar and camera data fusion colorization shown in the figure is implemented as follows: the radar pose data of the structured laser radar point cloud frame sequence when collecting data are calculated based on the radar-camera extrinsic calibration parameters between the laser radar and the camera, and the pose data of the camera when collecting data are calculated by spherical interpolation through Euclidean transformation SE (3), so as to realize the conversion of the point cloud coordinates to the camera coordinate system; after obtaining the camera pose, before projecting the laser radar points to the image plane of the camera, it is necessary to determine which points are visible from the current camera perspective, which is achieved by using the HPR operator, which can efficiently filter out the obscured points in the camera field of view; using the internal parameters of the camera, those filtered laser radar points determined to be visible by the HPR operator are projected to the image plane of the camera, and the image pixel points corresponding to each point in the point cloud are found based on the minimum error, thus completing the point cloud colorization of an image; as this process is iterated on all image frames, each point in the point cloud will accumulate color information from multiple perspectives. For each point, color data from multiple frames is collected; in order to select the optimal color information for each point, a weighted scoring method based on camera angle and distance is introduced. For each 3D point, the top N frames with the highest weight are selected as the best perspective; the RGB colors of these frames are fused using normalized weights to determine the final color of the point, thereby ensuring that the final color assigned to each 3D point comes from the most reliable and accurately captured perspective, thereby generating a high-quality, colored point cloud map, such as Figure 6As shown; replace the image with the mask map according to the similar method above, and fuse it to obtain a three-dimensional point cloud map with crack information marks, as shown Figure 7 Optionally, N can be 5, or other values ​​such as 8 or 10.

[0090] Furthermore, in step S70, the coloring result and the segmentation result are fused together to obtain a colored overall three-dimensional point cloud map with crack information marks, such as Figure 8 shown.

[0091] Furthermore, step S80 specifically includes: determining the crack identification point on the two-dimensional skeleton image according to the crack mask frame, calculating the gradient in the area of ​​measuring the crack position to determine the crack direction vector; locating the left and right edge points relative to the skeleton point by following the direction and its perpendicular direction on the two-dimensional skeleton image, and iteratively tracking to determine the two crack edge points of the measured crack position; projecting the point cloud onto the image, selecting the three-dimensional point that matches the crack center point on the image based on the minimum error, and then calculating the plane equation at the point on the point cloud image based on this point, based on the plane, generating a coordinate grid within a specified radius, and using the plane equation to calculate the corresponding coordinates, which will obtain a set of sampled three-dimensional points distributed on the plane; projecting the sampled three-dimensional points onto the two-dimensional image plane according to the intrinsic parameters of the camera, selecting the three-dimensional point that best matches the crack edge point based on the minimum projection error, and obtaining the specific coordinate values ​​of the two crack edge points.

[0092] Please refer to Figure 9In the specific implementation, it mainly includes: calculating the crack skeleton direction: in order to estimate the crack direction of a given point in the two-dimensional skeleton image, first extract the local neighborhood around the point on the mask map, and apply Gaussian smoothing to reduce noise, then use the Sobel operator to calculate the gradient in this neighborhood to determine the direction vector, the average gradient provides the direction vector, and then normalize it; finding edge points: use the calculated direction vector to track the edge of the crack. By following the direction and its perpendicular direction, the left and right edge points are located relative to the skeleton points. This tracking is performed iteratively until the edge is detected; sampling 3D points on the plane: In order to accurately reconstruct the three-dimensional crack edge, the overall three-dimensional semantic segmentation point cloud map is projected onto the image using the Inverse Perspective Mapping (IPM) technology. The 3D point that matches the center point of the crack on the image is selected based on the minimum error. Then, based on this point, the plane equation at that point is calculated on the point cloud map through plane calculation. Based on this plane, a coordinate grid is generated within the specified radius, and the corresponding coordinates are calculated using the plane equation. This will result in a set of sampled 3D points distributed on the plane; finding 3D edge points: The sampled 3D points are projected onto the 2D image plane using the intrinsic parameters of the camera. By comparing the projected points with the edge points observed in the image, the 3D points corresponding to the actual crack edges are identified. The best matching 3D points are selected based on the minimum projection error to obtain the 3D coordinate values ​​of the two crack edge points; calculating the crack width: The crack width is calculated as the Euclidean distance between the left and right edge points identified in three-dimensional space.

[0093] Furthermore, before step S20, the method includes: constructing a DeepLab-v3+ network model with ResNet as the backbone, determining the initial hyperparameters of the network model, and performing pre-training to obtain a pre-trained model; and training the pre-trained model using a labeled training set and a test set to obtain a crack semantic segmentation model.

[0094] Specifically, a DeepLab-v3+ network model with ResNet as the backbone was constructed. It adopted the PyTorch deep learning framework and used ResNet as the backbone to realize semantic segmentation of cracks. After the network model configuration was completed, the network training began. The network training included two parts: pre-training and formal training. First, the network was pre-trained. The crack image dataset used for pre-training was divided into a training set and a validation set in an appropriate ratio. The initial hyperparameters such as batchsize, epoch, and learning rate were determined, and pre-training began. After the training, a pre-trained model was obtained. Based on the pre-trained model, the crack image dataset used for formal training was divided into a training set and a validation set in an appropriate ratio. The initial hyperparameters such as batchsize, epoch, start epoch, and learning rate were determined, and formal training began. After the training, a trained model was obtained for subsequent image segmentation.

[0095] The present invention also provides a three-dimensional reconstruction system for concrete cracks based on multimodal fusion, including a processing device, which is used to implement the steps of the above-mentioned three-dimensional reconstruction method for concrete cracks based on multimodal fusion.

[0096] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the above-mentioned method for three-dimensional reconstruction of concrete cracks based on multimodal fusion are implemented.

[0097] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for 3D reconstruction of concrete cracks based on multimodal fusion, characterized in that: The steps include: S10, acquiring a structural image frame sequence and a structural lidar point cloud frame sequence of the target structure according to a preset acquisition trajectory, and obtaining the acquisition device posture corresponding to each acquisition moment; S20, using a crack semantic segmentation model to segment each of the structural image frames to obtain a crack mask image frame after crack segmentation; Step S20 specifically includes: using a crack semantic segmentation model to segment each frame of the structural image frame to obtain an initial mask image after crack segmentation, wherein the initial mask image has black and white marks and semantic information; cutting the initial mask image into a number of feature images, processing the feature images to extract a crack skeleton image, wherein the crack skeleton image has pixel colors representing the distance value of each pixel point from the nearest edge; converting the crack skeleton image into a set of pixel coordinate points, extracting pixel key points after downsampling; sorting and downsampling the pixel key point list, and marking the feature image based on a quantity threshold; using the mark of the feature image as a prompt point of the large model, importing the structural image frame corresponding to the feature image into the large model for crack mask optimization, and using the width optimization threshold for screening to obtain a crack mask image frame; S30, constructing a three-dimensional point cloud map based on the structural lidar point cloud frame and the corresponding radar measurement pose, and processing the three-dimensional point cloud to obtain a denoised point cloud map of the overall structure, wherein the radar measurement pose is obtained based on device external parameters and the pose of the acquisition device at the corresponding acquisition time; S40, obtaining a real-time estimated radar pose at any time by spherical interpolation calculation based on at least two of the radar measured poses, and obtaining a key shooting pose corresponding to each camera at a key shooting moment according to the real-time estimated radar pose and device external parameters, wherein the key shooting moment of the camera is a shooting moment corresponding to a key frame of the structural image, and the key frame of the structural image is determined from the structural image frame; S50, obtaining a visible 3D point cloud corresponding to the key shooting moment, and projecting the 3D points of the visible 3D point cloud onto corresponding structural image key frames and crack mask frames according to the key shooting pose, performing point cloud coloring and semantic segmentation to obtain a visible point cloud coloring point cloud map and a visible point cloud semantic segmentation point cloud map; S60, performing point cloud coloring and semantic segmentation processing on all structural image key frames and crack mask frames to obtain an overall three-dimensional colored point cloud map and an overall three-dimensional semantically segmented point cloud map; S70, fusing the overall three-dimensional colored point cloud map and the overall three-dimensional semantic segmentation point cloud map to obtain an overall three-dimensional point cloud labeled map with crack information and coloring; S80 , performing three-dimensional attribute measurement of the crack based on the three-dimensional point cloud marker map to obtain three-dimensional geometric information of the crack, wherein the three-dimensional geometric information of the crack includes information on crack width, crack position, and crack direction.

2. The method for 3D reconstruction of concrete cracks based on multimodal fusion according to claim 1, characterized in that: Step S30 specifically includes: Based on the structured lidar point cloud frame and the corresponding lidar measurement pose, a 3D point cloud map is constructed to obtain an overall 3D point cloud map; Performing cropping on the overall three-dimensional point cloud map; The overall three-dimensional point cloud map after cropping is denoised and smoothed to obtain the overall structure denoised point cloud map.

3. The method for 3D reconstruction of concrete cracks based on multimodal fusion according to claim 1, characterized in that: In step S40: Based on the measured poses of at least two of the radars, a rotation portion is calculated by spherical interpolation, and a position portion is obtained by linear interpolation to obtain a real-time estimated pose of the radar at any time; The camera pose is calibrated at a moment in time according to the real-time estimated pose of the radar and the radar-camera extrinsic calibration parameters, and the key shooting pose of the camera at the key shooting moment is obtained through coordinate transformation and time synchronization calculation; Step S50 specifically includes: Performing visibility filtering on the target moment point cloud image frame based on the current viewing angle of the key shooting posture, filtering out the invisible point cloud in the posture, and obtaining the key moment visible point cloud frames corresponding to the key shooting moments of each camera, wherein the key moment visible point cloud frames are obtained based on the overall structure denoised point cloud image; Project the point cloud in the visible point cloud frame at the key moment onto the image pixel points on the image plane of the corresponding camera for matching, and perform point cloud coloring to obtain a colored point cloud map of the visible point cloud; Project the point cloud in the visible point cloud frame at the key moment onto the pixel points of the crack mask on the image plane of the corresponding camera for matching, and perform point cloud semantic segmentation to obtain a semantic segmentation point cloud map of the visible point cloud; Step S60 specifically includes: Perform point cloud coloring on all key frames of the structural image to obtain the overall three-dimensional colored point cloud map; Perform point cloud semantic segmentation processing on the crack mask frames corresponding to all structural image key frames to obtain the overall three-dimensional semantic segmentation point cloud map.

4. The method for 3D reconstruction of concrete cracks based on multimodal fusion according to claim 3, characterized in that: The step of "performing point cloud coloring on all structural image key frames to obtain an overall three-dimensional colored point cloud map" specifically includes: obtaining associated structural image frames corresponding to each three-dimensional point in all visible point cloud frames at key moments, wherein the associated structural image frames have associated pixel points that match the three-dimensional points; wherein, an associated weight score corresponding to the three-dimensional point is determined based on a weighted scoring method of camera angle and distance, and the top N frames with the highest associated weight scores are determined as the associated structural image frames of the three-dimensional point; and using normalized weights to fuse the RGB colors of the associated pixel points in all the associated structural image frames to determine the color value of the three-dimensional point. The step of "performing point cloud semantic segmentation processing on the crack mask frames corresponding to all structural image key frames to obtain an overall three-dimensional semantic segmentation point cloud map" specifically includes: obtaining associated mask frames corresponding to each three-dimensional point in all visible point cloud frames at key moments, the associated mask frames having associated pixel points matching the three-dimensional points; wherein, the associated weight scores corresponding to the three-dimensional points are determined based on a weighted scoring method based on camera angle and distance, and the top N frames with the highest associated weight scores are determined as the associated mask frames for the three-dimensional points; and the colors of the associated pixel points in all associated mask frames are fused using normalized weights to determine the semantic information of the three-dimensional points.

5. The method for 3D reconstruction of concrete cracks based on multimodal fusion according to claim 1, characterized in that: In step S50, time synchronization is used to determine the real-time estimated radar pose corresponding to the key shooting moment, and visibility filtering is performed on the overall structure denoised point cloud image using the HRP operator based on the real-time estimated radar pose corresponding to the key shooting moment. The invisible point cloud under the pose is filtered out to obtain the visible three-dimensional point cloud corresponding to the key shooting moment.

6. The method for 3D reconstruction of concrete cracks based on multimodal fusion according to claim 1, characterized in that: Step S80 specifically includes: Determine crack identification points on the two-dimensional skeleton image according to the crack mask frame, and calculate the gradient within the area of ​​measuring the crack position to determine the crack direction vector; On the two-dimensional skeleton graph, the left and right edge points are located relative to the skeleton point by following the direction and its perpendicular direction, and the two crack edge points of the measured crack position are determined by iterative tracking; Project the point cloud onto the image, select the 3D point that matches the crack center on the image based on the minimum error, then calculate the plane equation at that point based on the point in the point cloud. Based on the plane equation, generate a coordinate grid within the specified radius, and use the plane equation to calculate the corresponding coordinates. This will result in a set of sampled 3D points distributed on the plane. The sampled 3D points are projected onto the 2D image plane according to the intrinsic parameters of the camera. The 3D points that best match the crack edge points are selected based on the minimum projection error to obtain the specific coordinate values ​​of the two crack edge points.

7. The method for 3D reconstruction of concrete cracks based on multimodal fusion according to claim 1, characterized in that: Before step S20, the following steps are included: Build a DeepLab-v3+ network model with ResNet as the backbone, determine the initial hyperparameters of the network model, and perform pre-training to obtain a pre-trained model; The pre-trained model is trained using the labeled training set and test set to obtain the crack semantic segmentation model.

8. A concrete crack 3D reconstruction system based on multimodal fusion, characterized in that: It comprises a processing device, which is used to implement the steps of the three-dimensional reconstruction method of concrete cracks based on multimodal fusion as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for three-dimensional reconstruction of concrete cracks based on multimodal fusion according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Three-dimensional model measurement method for concrete structure crack and related device

    CN116152697A