Real-time tunnel surface crack detection method based on image recognition

By employing multi-sensor fusion and dynamic drift compensation technology, the problems of image blurring and positioning errors caused by vibration and track irregularities in tunnel crack detection have been solved, achieving sub-centimeter-level crack detection and structural health monitoring, and improving the accuracy and automation level of detection.

CN122016824AInactive Publication Date: 2026-05-12CHINA RAILWAY FIRST BUREAU GRP RAILWAY CONSTR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA RAILWAY FIRST BUREAU GRP RAILWAY CONSTR CO LTD
Filing Date
2026-04-13
Publication Date
2026-05-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing tunnel crack detection technologies are susceptible to track irregularities and vehicle vibrations in high-speed moving environments, resulting in blurred images and geometric distortions, with positioning errors reaching the meter level, making it difficult to achieve sub-centimeter-level pixel alignment and crack evolution trend monitoring.

Method used

A multi-sensor fusion system is adopted, combining high-resolution imaging, inertial measurement and laser ranging units to perform visual synchronous positioning and map building, implement dynamic drift compensation and image distortion correction, and utilize deep semantic segmentation and feature extraction logic to achieve sub-centimeter-level global positioning and historical evolution monitoring.

Benefits of technology

It achieves high-precision detection of tunnel cracks in dynamic environments, reduces positioning errors to the sub-centimeter level, supports long-term structural health monitoring, and improves the automation level and decision-making efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122016824A_ABST
    Figure CN122016824A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image data processing and structure detection, and particularly relates to a tunnel surface crack real-time detection method based on image recognition. According to the method, tunnel surface images and motion information are collected through multi-sensor fusion, pose parameters are estimated in real time by utilizing a visual synchronous positioning and map construction algorithm, a reverse motion compensation model is constructed according to the pose parameters to eliminate motion blur and perspective distortion generated by jitter of a detection platform, and a high-definition standardized image is generated; utilizing a deep semantic segmentation network to automatically extract the geometric contour and pixel mask features of the crack; and projecting the crack coordinates to a global coordinate system, and performing spatial alignment with historical inspection data to realize increment comparison of the crack degradation trend. According to the invention, the problems of image blurring and low positioning precision in a dynamic environment are solved, sub-centimeter-level global positioning and cross-cycle millimeter-level change monitoring are realized, reliable data support is provided for long-term safety assessment of a tunnel structure, and the automation degree and decision-making efficiency of inspection operation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image data processing and structure detection technology, specifically relating to a real-time detection method for tunnel surface cracks based on image recognition. Background Technology

[0002] With the continuous expansion of transportation infrastructure, long-term health monitoring of tunnel structures has become an important part of ensuring traffic safety. Cracks on the tunnel lining surface are a key indicator for assessing structural stability and durability, and the accuracy and timeliness of their detection are crucial to the scientific basis of subsequent maintenance decisions. In recent years, automated inspections using high-speed inspection vehicles equipped with high-resolution imaging equipment have gradually replaced traditional manual visual inspection methods, becoming the mainstream technology for large-scale, high-efficiency tunnel defect detection.

[0003] Computer vision-based crack identification technology acquires continuous images of the tunnel walls and uses deep learning or feature extraction algorithms to automatically locate and quantify the geometric parameters of minute cracks. In practical operation scenarios, to complete inspections without disrupting traffic, the detection system typically needs to be mounted on a high-speed moving platform. By integrating image acquisition and positioning modules, it achieves full-coverage scanning of the entire tunnel surface. This process requires the system to maintain sampling frequency and positioning accuracy during dynamic movement to ensure that every defect feature is clearly captured and assigned accurate spatial attributes.

[0004] Existing technologies face engineering challenges in high-speed detection. Traditional image acquisition methods are highly susceptible to track irregularities or vehicle vibrations, resulting in motion blur and geometric distortion in the raw images captured by the camera. Current systems rely primarily on single odometers or simple positioning devices for localization. In long-distance, enclosed tunnel environments, accumulated errors can cause physical coordinate shifts in cracks to the order of meters, making it difficult to establish high-precision global spatial correlations. Existing detection methods lack dynamic perception and compensation mechanisms for minute camera pose changes, failing to perform sub-centimeter pixel alignment of crack images from different detection cycles, thus reducing their reliability in monitoring crack evolution trends and structural degradation. Summary of the Invention

[0005] The purpose of this invention is to provide a real-time detection method for surface cracks in tunnels based on image recognition, which can solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is: a real-time detection method for tunnel surface cracks based on image recognition, comprising the following specific steps: Step 1: Construct a dynamic sensing system with multi-sensor fusion. Utilize the high-resolution imaging unit, inertial measurement unit, and laser ranging unit installed on the detection platform to synchronously trigger data acquisition commands during the movement of the detection platform, and acquire continuous original sequence images of the tunnel lining surface, instantaneous angular velocity information, instantaneous acceleration information, and predetermined distance data of the detection platform relative to the tunnel inner wall. Step 2: Execute the front-end mileage calculation method based on visual synchronous localization and map construction. By extracting feature points from the original sequence images and performing inter-frame matching, and combining the motion increment output by the inertial measurement unit, the six-degree-of-freedom pose parameters of the detection platform in the local spatial coordinate system of the tunnel are estimated in real time. The pose parameters include a three-dimensional rotation matrix and a three-dimensional translation vector. Step 3: Implement dynamic drift compensation and image distortion correction. Calculate the minute vibration displacement of the detection platform at the moment of exposure using the pose parameters estimated in Step 2, construct a reverse motion compensation model, and perform pixel-level resampling and geometric correction on the original sequence images to eliminate motion blur and perspective distortion caused by the shaking of the detection platform, generating a high-definition standardized tunnel surface image. Step 4: Perform deep semantic segmentation and feature extraction logic. Input the standardized tunnel surface image into the preset crack detection neural network model to automatically identify and extract the geometric contour features, width features and length features of the crack, and generate a crack mask image with pixel-level labels. Step 5: Achieve sub-centimeter-level global positioning and historical evolution monitoring. Utilize the sparse point cloud map of the tunnel generated by synchronous positioning and map construction, map the pixel coordinates in the crack mask image to the global world coordinate system through projection transformation, calculate the absolute physical coordinates of the crack, and spatially align and incrementally compare the current detection results with the historical inspection data stored in the preset database to determine the deterioration trend of the crack.

[0007] Preferably, in step 1, the high-resolution imaging unit includes multiple sets of industrial cameras arranged circumferentially along the detection platform. Each set of cameras maintains a predetermined overlapping field of view to ensure full-coverage imaging of the tunnel lining surface without blind spots. The trigger frequency of the imaging unit forms a closed-loop feedback adjustment with the travel speed of the detection platform; as the travel speed increases, the trigger frequency increases accordingly to maintain a constant longitudinal spatial resolution. The inertial measurement unit employs a high-frequency sampling mode to record minute disturbances of the detection platform in the three-axis directions.

[0008] Preferably, in step 2, the front-end odometry of the visual synchronous localization and map construction performs grayscale conversion and histogram equalization processing on the image to enhance image contrast in the dark tunnel environment. Feature point extraction uses feature descriptors with rotation invariance and scale invariance, and stable features are searched at different resolution levels by establishing a feature point pyramid structure. The inter-frame matching process uses bidirectional optical flow tracking technology combined with a random sampling consensus algorithm to eliminate mismatched points and ensure the robustness of pose estimation.

[0009] Preferably, the estimation process of the pose parameters further includes a back-end optimization step. By constructing a factor graph model, the visual reprojection error, the pre-integration error of the inertial measurement unit, and the constraint terms provided by the laser ranging unit are used as the optimization objective function. The motion trajectory of the detection platform is smoothed using the Levenberg Marquardt algorithm to eliminate the cumulative drift error generated by the sensor.

[0010] Preferably, in step 3, the construction process of the reverse motion compensation model involves transforming the estimated rotation matrix and translation vector into transformation operators for the image coordinate system. For translational blur caused by vehicle vibration, a degradation function is constructed in the frequency domain for inverse filtering by calculating the displacement difference between adjacent time points. For perspective distortion caused by camera tilt, the homography matrix is ​​used to project the tilted imaging plane onto a reference plane parallel to the tunnel lining surface, ensuring that the geometric proportions of the cracks are not distorted.

[0011] Preferably, in step 4, the preset crack detection neural network model adopts a symmetrical encoder-decoder structure. The encoder part gradually extracts deep semantic information of the image through multi-layer convolution and pooling operations to capture the subtle texture features of the crack. The decoder part restores the spatial resolution of the image through upsampling operations and skip connection structures to achieve accurate localization of crack edges. At the end of the network structure, an attention mechanism module is introduced to enhance the model's ability to perceive slender, low-contrast cracks and reduce the false detection rate caused by background interference such as water seepage and oil stains in the tunnel.

[0012] Preferably, the geometric contour feature extraction of the crack includes calculating the maximum width, average width, and total length of the crack. By performing skeletonization extraction on the crack mask image, the central axis of the crack is obtained, and edge pixels are searched along the normal direction of the central axis to accurately calculate the physical width of the crack. All measured values ​​are converted to units based on the camera's intrinsic parameter matrix and real-time feedback object distance data to ensure accuracy in converting from pixel scale to physical scale.

[0013] Preferably, in step 5, the establishment of the global world coordinate system is based on a preset reference point at the tunnel entrance. The synchronous positioning and mapping algorithm continuously updates and maintains a sparse point cloud map during operation, which records the coordinates of spatial points with geometric features within the tunnel. When the detection vehicle passes through the same area again, descriptor matching is used to associate the current frame with map points, triggering a loop closure detection mechanism. Loop closure constraints are then used to correct the global trajectory, achieving sub-centimeter-level positioning accuracy.

[0014] Preferably, the historical evolution monitoring process includes automated alignment logic. The system retrieves historical images and historical crack features within the same physical coordinate range from the database and performs point cloud registration using an improved iterative nearest-point algorithm. By calculating the area growth rate, width expansion, and new branching of the same crack under different periods, a tunnel structure health diagnosis report is automatically generated. If the crack change exceeds a preset threshold, the system will automatically trigger an alarm command.

[0015] Preferably, the method further includes ambient lighting compensation processing. During image acquisition, a high-power constant current lighting system is simultaneously activated, and a preset diffuser ensures that the light is evenly distributed on the tunnel wall, eliminating the influence of light spots and shadows caused by artificial lighting on image recognition.

[0016] Preferably, the method also involves a multi-threaded parallel processing architecture. Image acquisition, pose estimation, crack identification, and data storage tasks are allocated to different computing cores, and high-speed data exchange is achieved through a memory-sharing mechanism. Even when the detection platform is operating at high speed, the system can complete the entire process analysis of a single frame image within a predetermined processing cycle, ensuring real-time detection.

[0017] Preferably, the loop closure detection mechanism in the simultaneous localization and mapping (SMR) algorithm employs a bag-of-words model for image retrieval. By calculating the similarity between the feature vector of the current scene and the feature vector of historical scenes, the system identifies whether the detection platform has traversed previously visited locations. Once a loop closure is detected, the system initiates global pose graph optimization, adjusting the poses of all historical nodes to minimize the global error and resolve the coordinate drift problem in long-distance detection.

[0018] Preferably, the inertial measurement unit pre-integration technique is used to address the mismatch between the sampling frequency of the visual sensor and the inertial sensor. By integrating high-frequency acceleration and angular velocity data between two adjacent image frames, the relative motion increment of the detection platform during that time period is obtained. This processing method can offset the severe shaking effect on the pose trajectory caused by the instantaneous impact force generated by the vehicle at uneven track sections.

[0019] Preferably, the crack identification process also includes an online update strategy for the deep learning model. When the system encounters new types of defects or special background textures, the manual intervention unit confirms the suspected areas and adds the marked new samples to the training set. The model weights in the on-board computing unit are updated periodically through an incremental learning algorithm, enabling the system to have continuously evolving identification capabilities.

[0020] Preferably, the laser ranging unit is used not only to constrain the pose but also to monitor the lateral distance changes between the detection platform and the tunnel sidewall in real time. This distance information is fed back to the image calibration module in real time to dynamically correct the depth factor in the projection transformation matrix, ensuring that the extracted crack physical dimensions remain highly consistent even when the detection platform undergoes lateral shifts or serpentine movements.

[0021] Preferably, the database management system adopts a spatial index structure. The detection data is stored in segments based on tunnel mileage and ring segment number. Users can retrieve all historical image sequences for a specific location through a graphical interface, enabling a visual reproduction of the entire disease evolution process.

[0022] Preferably, the method further includes simultaneous monitoring of ambient temperature and humidity. While acquiring crack images, current environmental parameters are recorded and stored as additional attributes in the crack information archive. In subsequent data analysis, the system will perform correlation analysis between physical changes and environmental fluctuations, eliminating periodic crack closure errors caused by thermal expansion and contraction.

[0023] Preferably, the point cloud map generated by the synchronous positioning and map construction is simplified using voxelized raster filtering. While ensuring the integrity of the tunnel contour features, redundant smoothed area point clouds are eliminated, reducing memory usage and improving the computational efficiency of global registration and loop closure retrieval.

[0024] Preferably, the process of generating the standardized tunnel surface image also includes correction of lens distortion. Using pre-calibrated radial and tangential distortion coefficients, coordinate offset correction is performed on each pixel position to eliminate stretching or compression at image edges, ensuring that the geometric measurement accuracy of the crack remains constant at any location in the image.

[0025] Preferably, the method is applied to an automated tunnel inspection robot. The inspection robot has the ability to autonomously plan its path, and when it detects a suspected crack area, it will automatically reduce its travel speed or execute a stop-and-shoot command to obtain detailed images with a higher signal-to-noise ratio for in-depth verification.

[0026] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention deeply couples synchronous positioning and map building technology into the tunnel crack detection process, overcoming the problem of image blurring caused by platform vibration in dynamic environments under traditional detection methods, and achieving accurate reverse compensation for motion blur.

[0027] 2. By utilizing the six-degree-of-freedom pose parameters provided by the multi-sensor fusion algorithm, the system can perceive the minute motion trajectory of the detection platform in real time, improving the positioning accuracy of cracks in the tunnel global coordinate system and reducing the positioning error from the meter level to the sub-centimeter level.

[0028] 3. The historical evolution monitoring mechanism based on point cloud registration established in this invention enables millimeter-level monitoring of the same crack across cycles and long spans, providing scientific and reliable data support for the long-term safety assessment of tunnel structures and improving the automation level and decision-making efficiency of inspection operations. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of the overall technical solution architecture of the present invention; Figure 2 This is a schematic diagram of the core principle framework of the present invention, which performs reverse motion compensation and image distortion correction based on pose parameters generated by synchronous positioning and map construction. Figure 3 This is a flowchart illustrating the main stages of the process for deep semantic segmentation and crack geometric feature extraction of standardized tunnel surface images in this invention. Figure 4 This is a schematic diagram of the multi-level interaction relationship and data flow between multi-source sensing data, real-time pose transformation operator and crack detection results in this invention; Figure 5 This is a flowchart illustrating the logical process framework for monitoring crack deterioration trends based on tunnel global coordinate system mapping and spatial registration of historical inspection data in this invention. Detailed Implementation

[0030] Example 1: Please refer to the appendix Figure 1 To be continued Figure 5 To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments.

[0031] In the image recognition-based real-time detection method for tunnel surface cracks, step 1 is performed: a multi-sensor fusion dynamic perception system is constructed. The high-resolution imaging unit, inertial measurement unit, and laser ranging unit installed on the detection platform are used to synchronously trigger data acquisition commands during the movement of the detection platform to acquire continuous original sequence images of the tunnel lining surface, instantaneous angular velocity information, instantaneous acceleration information, and predetermined distance data of the detection platform relative to the tunnel inner wall.

[0032] The high-resolution imaging unit in step 1 consists of multiple sets of industrial cameras arranged in a fan-shaped or ring-shaped pattern along the circumference of the inspection platform. To ensure full coverage and no blind spots in imaging the tunnel lining surface, an overlap area is pre-defined between adjacent camera sets. The width of this overlap area is typically configured to be 10% to 20% of the width of a single image to facilitate subsequent image stitching and feature continuity analysis. The triggering mechanism of the imaging unit is not a fixed frequency, but rather forms a closed-loop feedback adjustment system with the travel speed of the inspection platform. The system monitors the travel speed in real time through high-line-count incremental encoders mounted on the wheels of the inspection platform and feeds the speed signal back to the central controller. When the travel speed increases, the central controller automatically calculates and increases the triggering frequency to ensure that the displacement increment of adjacent images in the longitudinal space remains constant, maintaining a constant longitudinal spatial resolution.

[0033] During data acquisition, the inertial measurement unit is configured in high-frequency sampling mode, with a sampling frequency typically set above 200 Hz, capable of completely recording minute angular velocity disturbances and instantaneous linear acceleration changes of the detection platform in three axes. The laser ranging unit calculates the lateral distance between the center or edge point of the detection platform and the tunnel wall in real time by emitting high-frequency pulsed lasers and receiving reflected signals. To eliminate the influence of uneven artificial lighting distribution, light spots, shadows, and weak ambient light on image recognition within the tunnel, the method also includes ambient light compensation processing. During image acquisition, the system simultaneously activates multiple sets of high-power constant-current lighting fixtures installed around the imaging unit. These fixtures are equipped with customized diffusers at their front ends, enabling the output light to be evenly distributed on the tunnel wall, providing a high color rendering and uniform illumination environment.

[0034] In step 1, the system also simultaneously monitors ambient temperature and humidity parameters. Using a digital temperature and humidity sensor integrated into the detection platform, while acquiring crack images, the current ambient temperature and relative humidity values ​​are stored as metadata attributes in the corresponding image information archive. This provides data support for subsequently eliminating periodic crack closure errors caused by thermal expansion and contraction, ensuring the physical consistency of crack width measurements.

[0035] Step 2: Execute the front-end mileage calculation method based on visual synchronous positioning and map construction. By extracting feature points from the original sequence images and performing inter-frame matching, and combining the motion increment output by the inertial measurement unit, the six-degree-of-freedom pose parameters of the detection platform in the local spatial coordinate system of the tunnel are estimated in real time. The pose parameters include a three-dimensional rotation matrix and a three-dimensional translation vector.

[0036] In step 2, the front-end odometry of the visual simultaneous localization and mapping (VLS) system performs preprocessing operations when processing each frame of image. This preprocessing includes converting the color image to a single-channel grayscale image and performing finite contrast adaptive histogram equalization to enhance feature saliency in dark tunnel environments and low-contrast conditions. The feature point extraction logic employs feature descriptors that are rotation-invariant, scale-invariant, and robust to affine transformations. To capture stable features at different distances and resolutions, the algorithm establishes a feature point pyramid structure, searching for extreme points in parallel across multiple levels from the original resolution to multi-level downsampling resolutions.

[0037] In the inter-frame matching stage, the system employs bidirectional optical flow tracking technology. It not only calculates the pixel motion vector of the current frame relative to the previous frame but also calculates the displacement of the previous frame relative to the current frame. Only when the bidirectional error is less than a predetermined threshold is the feature point retained. To further eliminate mismatched points, the system introduces a random sample consensus algorithm. This algorithm iteratively selects the minimum sample set to construct a homography matrix or essential matrix, eliminating outliers that do not conform to geometric constraints and ensuring the robustness of pose estimation.

[0038] The estimation process of the pose parameters also deeply integrates data from the inertial measurement unit (IMU). IMU pre-integration technology is employed to address the significant difference in sampling frequency between the visual sensor and the inertial sensor. Within the time span between two adjacent visual image frames, continuous time integration is performed on high-frequency acceleration and angular velocity data to obtain the relative velocity increment, displacement increment, and angle change of the detection platform during that time period. This processing method can offset the severe shaking caused by the instantaneous impact force generated by the vehicle at uneven track sections, providing a smoother and more reliable motion trajectory prediction than relying solely on visual matching.

[0039] Following step 2, a global optimization process is also included. A factor graph model is constructed, using the visual reprojection error, the inertial measurement unit pre-integration error, and the lateral distance constraint provided by the laser ranging unit as the objective function. The Levenburg-Marquardt algorithm is then used to iteratively smooth the motion trajectory of the detection platform within a sliding window. The goal of this optimization process is to minimize the sum of squared residuals between all observed data and the motion model, eliminate the cumulative drift errors generated by various sensors, and ensure the accuracy of the six-DOF pose parameters under long-term operation.

[0040] Step 3: Implement dynamic drift compensation and image distortion correction. Calculate the minute vibration displacement of the detection platform at the moment of exposure using the pose parameters estimated in Step 2, construct a reverse motion compensation model, and perform pixel-level resampling and geometric correction on the original sequence images to eliminate motion blur and perspective distortion caused by the shaking of the detection platform, generating a high-definition standardized tunnel surface image.

[0041] In step 3, the construction of the inverse motion compensation model involves converting the six-DOF pose parameters obtained in step 2 into a spatial transformation operator for the image coordinate system. For instantaneous translational blur caused by vehicle vibration, the system calculates the displacement difference between poses at adjacent time points. If the pixel displacement corresponding to this displacement difference exceeds a predetermined pixel threshold within the exposure time, motion blur is identified. The system constructs a specific degradation function in the frequency domain, which characterizes the kernel function of the motion blur, and uses Wiener filtering or inverse filtering algorithms to deconvolve the image, restoring the blurred edge details.

[0042] To address perspective distortion caused by the camera's optical axis not being perpendicular to the tunnel wall, the system uses a homography matrix to project the current tilted imaging plane onto a pre-defined reference plane that is ideally parallel to the tunnel lining surface. This projection process involves remapping the coordinates of each pixel in the image to ensure that the geometric proportions of the cracks in the image are not distorted due to perspective. The normalized image generation process also includes depth correction for intrinsic lens distortion. Using pre-calibrated radial and tangential distortion coefficients using a high-precision calibration plate, a coordinate offset correction mapping table covering the entire pixel area is established. Sub-pixel level resampling is performed at each pixel location to eliminate stretching or compression phenomena in image edge regions.

[0043] During the geometric correction process, the system also dynamically adjusts the depth factor in the projection transformation matrix based on the object distance data fed back in real time by the laser ranging unit. This step ensures that even if the imaging distance changes due to lateral shift or serpentine motion of the detection platform, the final standardized image has a uniform physical size correspondence at every point, meaning that the actual physical length represented by each pixel remains constant.

[0044] Step 4: Perform deep semantic segmentation and feature extraction logic. Input the standardized tunnel surface image into the preset crack detection neural network model, automatically identify and extract the geometric contour features, width features and length features of the crack, and generate a crack mask image with pixel-level labels.

[0045] In step 4, the pre-defined crack detection neural network model employs a highly optimized symmetrical encoder-decoder structure. The encoder, through cascaded multi-layer convolutional operations, normalization operations, and non-linear activation operations, progressively compresses the spatial dimension of the feature map and increases the feature channel depth, capturing the deep semantic information and subtle textures of cracks in the image. The decoder, on the other hand, gradually restores the spatial resolution of the image through transposed convolution or upsampling operations, and utilizes a skip connection structure to fuse low-level high-resolution features from the encoder with high-level semantic features from the decoder, achieving pixel-level precise localization of crack edges.

[0046] To enhance the model's ability to perceive low-contrast, elongated cracks against complex backgrounds, a spatial and channel attention mechanism module is introduced at the end of the network structure. This module automatically learns the importance weights of different regions in the image, allowing the model to focus more on pixel sets with crack features during processing, while suppressing background interference caused by water seepage, oil stains, construction joints, or cables within the tunnel. After obtaining the pixel-level crack mask, the system further performs geometric feature extraction logic. This includes morphological skeletonization of the crack mask image, obtaining the crack's central axis through a continuous peeling algorithm. Edge pixels are searched along the normal direction of this central axis, and the Euclidean distance between two pairs of edge pixels is calculated to obtain the crack width at that location.

[0047] During the extraction process, the system calculates the maximum width, average width, and total length accumulated from the central axis of the entire crack. All pixel-scale measurements are ultimately combined with the camera's intrinsic parameter matrix data, distortion correction coefficients, and real-time object distance data to transform them into physically meaningful millimeter-level values. This ensures high measurement accuracy in crack identification results. The system also has online learning and updating capabilities. When encountering previously unseen types of damage or unique background textures during inspections, the manual intervention unit manually confirms and labels suspected areas, adding the new sample to the training dataset in real-time or periodically. Incremental learning algorithms update the model weights in the onboard computing unit, allowing the robustness of the identification system to continuously evolve with increasing inspection mileage.

[0048] Step 5: Achieve sub-centimeter-level global positioning and historical evolution monitoring. Utilize the sparse point cloud map of the tunnel generated by synchronous positioning and map construction, map the pixel coordinates in the crack mask image to the global world coordinate system through projection transformation, calculate the absolute physical coordinates of the crack, and spatially align and incrementally compare the current detection results with the historical inspection data stored in the preset database to determine the deterioration trend of the crack.

[0049] In step 5, the establishment of the global world coordinate system begins with a preset reference point at the tunnel entrance, or with absolute geographic coordinates obtained using a high-precision global positioning system. During operation, the simultaneous localization and mapping (SMR) algorithm continuously maintains and dynamically updates the sparse point cloud map. This map not only records the coordinates of spatial points with geometric features but also includes feature vector descriptors for these feature points. When the detection vehicle re-enters the same tunnel area, the system quickly matches the feature descriptors extracted from the current frame with historical descriptors in the sparse point cloud map, thus establishing a connection between the current frame and the map.

[0050] Once a match is successful, the system triggers a loop closure detection mechanism. This mechanism uses a bag-of-words model to transform image features into visual word vectors. By calculating the similarity between the current vector and historical image vectors, it identifies whether the detection platform has returned to a specific spatial location it has previously visited. When a loop closure is detected, the system initiates a global pose graph optimization algorithm. This algorithm uses the poses of all historical nodes as variables to be optimized and loop closure constraints as error terms. By adjusting the global trajectory, it minimizes the cumulative coordinate drift error generated during long-distance travel.

[0051] During historical evolution monitoring, the system retrieves historical inspection data consistent with the current physical coordinates from a database based on a spatial index structure. The database manages data in segments according to tunnel mileage and ring segment numbers to ensure efficient retrieval. The system employs an improved iterative nearest-point algorithm to perform precise rigid or non-rigid registration between the current crack mask point cloud and historical crack mask point clouds. By calculating the percentage increase in crack area, width expansion, length extension, and the presence of new branches at the same physical location under different periods, the system automatically generates a tunnel structural health diagnosis report.

[0052] If the change in a certain indicator calculated by the system exceeds a preset safety threshold, the central controller will automatically trigger a tiered alarm command and mark the area as a key inspection target. When the method is applied to an automated tunnel inspection robot, the robot has the ability to autonomously plan its path. When it receives a suspected crack command from the detection system, the inspection robot will automatically reduce its travel speed and may even stop and take pictures at predetermined coordinates to obtain images with higher signal-to-noise ratios using long exposure or multi-angle shooting, for in-depth verification by the backend.

[0053] At the data storage and processing level, this invention relates to a multi-threaded parallel processing architecture. Image acquisition, pose estimation based on visual inertial odometry, crack identification based on deep learning, and global localization based on point cloud maps are allocated to different cores of the computing unit. High-speed data exchange is achieved through a high-speed memory sharing mechanism and a double-buffer mechanism. This architecture ensures that even when the detection platform is traveling at high speed, the system can complete the entire complex analysis process within a predetermined cycle of single-frame image acquisition, guaranteeing real-time detection. The generated sparse point cloud map also undergoes voxelized raster filtering to remove redundant point clouds while preserving the core geometric features of the tunnel, reducing memory consumption and improving the efficiency of global registration and large-scale map retrieval.

[0054] In the specific application scenario of the method, the detection platform is mounted on top of a track maintenance vehicle. As the maintenance vehicle moves at a constant speed, the high-resolution imaging unit acquires images at a frequency of 30 frames per second, while the inertial measurement unit records the vibration of the maintenance vehicle at a frequency of 500 Hz. When the maintenance vehicle experiences severe up-and-down jolting at a track joint, the inertial measurement unit senses a vertical acceleration pulse. The odometer in step 2 immediately captures this attitude change and transmits it to the reverse motion compensation module in step 3. This module quickly constructs a vertical translation compensation matrix, performing pixel-level correction on the image that would otherwise have produced vertical blurring due to jolting, ensuring that the crack edges are clearly discernible.

[0055] The identified crack was projected onto a global map based on mileage coordinates. The system found that this crack overlapped with crack number 1024 in last year's inspection record. Comparison revealed that the average width of this crack had increased from 0.25 mm to 0.28 mm. Although the increase was minor, the system, leveraging the sub-centimeter-level positioning capabilities provided by this invention, confirmed the authenticity of this change and automatically recorded this subtle evolution process in the health diagnosis report, achieving millimeter-level refined management of tunnel structural health.

[0056] Example 2: In another preferred embodiment, the image recognition-based real-time detection method for tunnel surface cracks is configured to have higher intelligent redundancy and fault tolerance mechanisms to cope with the extreme and harsh environment that may occur inside the tunnel.

[0057] Regarding the data acquisition stage in step 1, this embodiment introduces real-time imaging quality assessment logic based on multi-sensor fusion. After each frame of image acquisition is completed, the system immediately calculates the average gradient and information entropy of the image in the hardware buffer. If local overexposure due to abnormally strong light or insufficient illumination due to temporary malfunction of the supplementary lighting is detected, the system will send a power adjustment signal to the actuator to dynamically adjust the exposure parameters of the next frame and mark the current frame as sub-optimal quality data. In the subsequent processing in step 3, the system will prioritize using adjacent high-quality image sequences to perform spatiotemporal interpolation repair on the sub-optimal data, ensuring the continuity of the crack detection logic and preventing missed detections due to abnormal data in a single frame.

[0058] In the pose estimation process of step 2, to further improve the stability of the algorithm in areas lacking feature points (such as newly repaired flat lining surfaces), this embodiment incorporates an edge constraint term into the visual odometry. In addition to extracting point features, the algorithm also utilizes the Canney edge detection operator to extract the longitudinal and circumferential joint lines between lining blocks. By establishing line feature descriptors and performing line matching, the reprojection error of the line features is added to the factor graph optimization model. Since the joints of the tunnel lining have extremely strong geometric regularity, this point-line fusion SLAM algorithm can enhance the system's positioning and orientation accuracy in texture-scarce scenarios, preventing pose loss due to insufficient feature points.

[0059] Regarding the image distortion correction process in step 3, this embodiment further considers the geometric deformation caused by the camera's rolling shutter effect. Because industrial cameras read data line by line during high-speed movement, there are slight time deviations between different lines of the image. The system uses high-frequency angular velocity data provided by the inertial measurement unit to calculate the camera's rotation angle increment within each scan line. By compensating for each scan line of the image, the tilt or stretching deformation originally caused by the rolling shutter is corrected. This refined correction process improves the accuracy of feature extraction by the crack recognition model at high driving speeds, exceeding 60 kilometers per hour.

[0060] In the identification stage of step 4, this embodiment introduces a multi-scale feature pyramid attention mechanism. By calculating spatial attention weights at different feature levels of the neural network, the model can not only identify large through cracks but also capture tiny mesh cracks. In the post-processing stage of mask generation, a physics-based classifier is added. This classifier automatically distinguishes identified targets into structural stress cracks, surface shrinkage cracks, water seepage interference, and construction scratches by utilizing the crack's geometric aspect ratio, directional distribution features, and relative positional relationship with the lining joints. This deep semantic understanding provides higher-dimensional classification data references for the subsequent degradation trend analysis in step 5.

[0061] For the historical evolution monitoring in step 5, this embodiment implements a crack evolution model based on spatiotemporal topology at the database level. The system not only records the physical coordinates of cracks but also establishes the interrelationships between cracks. By analyzing the aggregation, convergence, and growth rate of multiple cracks within a certain area, the system can assess the stress field distribution changes in the tunnel lining at that location. During incremental comparisons, in addition to calculating absolute value changes, the system also uses heatmaps to display active areas of crack development. When the inspection vehicle passes through again, using the closed-loop detection provided in this embodiment, the system can seamlessly integrate with data from several years ago. Even when ambient light and shadow change, the registration accuracy can still be guaranteed to be within 10 millimeters through registration technology based on point cloud structure descriptors.

[0062] This embodiment also relates to a distributed computing architecture. The raw big data stream collected by the detection platform undergoes primary processing through vehicle-mounted high-performance edge computing nodes, including completing all real-time tasks from steps 1 to 4. Step 5, involving large-scale point cloud map maintenance, global trajectory optimization, and deep correlation analysis of historical inspection data, involves the vehicle-mounted nodes extracting feature summaries, which are then synchronized to a remote cloud server or station-level management center via a high-speed wireless communication network deployed within the tunnel. This cloud-edge collaborative processing method ensures real-time response speed at the inspection site while leveraging the powerful computing and storage resources of the backend, enabling unified management of massive amounts of tunnel health data across lines and years.

[0063] In practice, when the inspection robot discovers that the width growth slope of a crack increases progressively over three consecutive inspections, the intelligent decision-making module in this embodiment automatically upgrades the crack's risk level from general to major risk. The system immediately retrieves all standardized image sequences of that location over the past five years and uses the distortion-free images generated by this invention to automatically stitch together a panoramic view of the temporal evolution of the affected area. This panoramic view displays the entire process of the crack from its initial initiation to its expansion and the development of secondary damage in a high-definition, proportional manner. This provides tunnel maintenance departments with extremely detailed and irreplaceable scientific decision-making basis for formulating reinforcement and repair plans.

[0064] Example 3: In another embodiment of the present invention, the method particularly enhances the ability to handle heterogeneity of multi-source data and the limit control of cumulative coordinate error in ultra-long-distance tunnel inspection.

[0065] In step 1, the high-resolution imaging unit added a multispectral imaging mode. In addition to the visible light camera, a near-infrared sensor was integrated. Near-infrared light can penetrate some of the water vapor and light dust on the tunnel surface, further enhancing the contrast of cracks in the image. The inertial measurement unit uses high-precision devices at the fiber optic gyroscope level, with an extremely low random walk coefficient, enabling it to maintain a high-precision motion trajectory solely through inertial calculation even in the event of a short-term loss of visual signal. The system achieves active object distance tracking through a laser ranging unit, and the imaging unit's lens has an autofocus function. Based on the real-time measured object distance, the system uses a voice coil motor to drive the lens compensation group for millisecond-level focusing, ensuring that the imaging plane remains within the optimal depth of focus range even when the tunnel wall is uneven or the platform oscillates.

[0066] In the pose estimation logic of step 2, to address the scale uncertainty issue in monocular visual SLAM, the system introduces laser odometry constraints. The laser ranging unit not only measures distance in a single direction but also employs a multi-line rotating lidar to acquire the local geometry of the surrounding environment through real-time scanning of the tunnel segment contours. The system couples and optimizes feature line matching of the laser point cloud with visual feature point matching. In the backend optimization, the absolute scale information provided by the lidar is injected into the factor map as a strong constraint term, resolving the visual scale drift caused by vehicle speed fluctuations and ensuring the accuracy of the absolute physical dimensions of the three-dimensional translation vector in the six-degree-of-freedom pose parameters.

[0067] For distortion correction in step 3, this embodiment employs an image reprojection technique based on a non-parametric model. The system pre-establishes a dense lookup table containing all complex nonlinear distortions within the camera using a large field-of-view calibration field. When processing the pose increment output in step 2, the system not only performs overall translation and rotation compensation but also adjusts each image sub-region locally using affine transformations based on the camera imaging model. This refined reconstruction method perfectly corrects the severe barrel distortion caused by the wide-angle edge of the lens, resulting in a geometrically perfect continuity of the stitched tunnel roll image, without any seams or ghosting.

[0068] In the recognition logic of step 4, this embodiment introduces knowledge distillation technology to transfer the crack discrimination knowledge learned by the complex and parameter-intensive expert model to a lightweight student model that can run at ultra-high speed on an in-vehicle embedded computing core. This allows the system to maintain a very high recognition rate while reducing the processing time of a single frame image. The system utilizes a multi-task learning framework to simultaneously output a crack depth estimation map while outputting the crack mask, assisting in determining whether the crack has penetrated the interior of the lining.

[0069] In the global monitoring section of step 5, this embodiment implements an adaptive global map update strategy. Considering that the appearance of the tunnel lining surface may change due to water seepage, scale buildup, equipment replacement, or dust removal operations, the synchronous positioning and mapping algorithm periodically assesses the lifespan decay of old feature points in the point cloud map. Only feature points that are consistently observed across multiple inspections are retained as the foundation of the pyramid, while changed regional features are dynamically overwritten by newly acquired data. When comparing with historical data, the system employs a nonlinear registration logic based on manifold learning, which automatically identifies and isolates non-pathological changes caused by external construction such as support installation and pipeline laying, focusing instead on the trajectory of changes in the lining cracks themselves.

[0070] In the detection of extremely long tunnels, the system utilizes loop closure detection and external control points (such as measuring markers buried at predetermined intervals within the tunnel) for dual verification. Once a control point is detected, the system immediately uses its high-precision absolute coordinates as the true value and performs a global readjustment of the pose maps of all preceding road segments. In this way, the present invention can control the cumulative coordinate error of tunnel detection over tens of kilometers to within the centimeter range.

[0071] In the final data output stage, the system constructs a three-dimensional, visualized digital twin tunnel model. Users simply click on any location on the model, and the system automatically retrieves historical crack images, physical dimension parameters, environmental temperature and humidity records, and predicted evolution curves of crack development for that location. This highly integrated information display method enhances the intuitiveness and accuracy of tunnel safety assessments.

[0072] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A method for real-time detection of surface cracks in tunnels based on image recognition, characterized in that, The method includes the following steps: Step 1: Construct a multi-sensor fusion dynamic perception system. Utilize the high-resolution imaging unit, inertial measurement unit, and laser ranging unit installed on the detection platform to synchronously trigger data acquisition commands during the movement of the detection platform, and acquire continuous original sequence images of the tunnel lining surface, instantaneous angular velocity information, instantaneous acceleration information, and predetermined distance data of the detection platform relative to the tunnel inner wall. Step 2: Execute the front-end mileage calculation method based on visual synchronous localization and map construction. By extracting feature points from the original sequence images and performing inter-frame matching, and combining the motion increment output by the inertial measurement unit, the six-degree-of-freedom pose parameters of the detection platform in the local spatial coordinate system of the tunnel are estimated in real time. The six-degree-of-freedom pose parameters include a three-dimensional rotation matrix and a three-dimensional translation vector. Step 3: Implement dynamic drift compensation and image distortion correction. Calculate the minute vibration displacement of the detection platform at the moment of exposure using the pose parameters estimated in Step 2, construct a reverse motion compensation model, and perform pixel-level resampling and geometric correction on the original sequence images to generate high-definition standardized tunnel surface images. Step 4: Perform deep semantic segmentation and feature extraction logic. Input the standardized tunnel surface image into the preset crack detection neural network model to automatically identify and extract the geometric contour features, width features and length features of the crack, and generate a crack mask image with pixel-level labels. Step 5: Achieve sub-centimeter-level global positioning and historical evolution monitoring. Utilize the sparse point cloud map of the tunnel generated by synchronous positioning and map construction, map the pixel coordinates in the crack mask image to the global world coordinate system through projection transformation, calculate the absolute physical coordinates of the crack, and spatially align and incrementally compare the current detection results with the historical inspection data stored in the preset database to determine the deterioration trend of the crack.

2. The real-time detection method for tunnel surface cracks based on image recognition according to claim 1, characterized in that, In step 1, the high-resolution imaging unit includes multiple sets of industrial cameras arranged in a fan shape along the circumference of the detection platform. Adjacent sets of cameras maintain a predetermined field-of-view overlap area, the width of which is 10% to 20% of the width of a single image. The trigger frequency of the high-resolution imaging unit forms a closed-loop feedback adjustment with the travel speed of the detection platform. The travel speed is monitored in real time by a high-line-count incremental encoder installed on the wheels of the detection platform, and the speed signal is fed back to the central controller. When the travel speed increases, the central controller automatically increases the trigger frequency so that the displacement increment of adjacent images in the longitudinal space remains constant. The sampling frequency of the inertial measurement unit is set above 200 Hz to record the minute angular velocity disturbances and instantaneous linear acceleration changes of the detection platform in the three-axis directions. The method also includes ambient light compensation processing. During the image acquisition process, multiple sets of high-power constant current lighting fixtures installed around the imaging unit are turned on simultaneously. The front end of the multiple sets of high-power constant current lighting fixtures is equipped with diffuser plates so that the output light is evenly distributed on the tunnel wall. The multi-sensor fusion dynamic sensing system also utilizes digital temperature and humidity sensors integrated on the detection platform to simultaneously monitor environmental temperature and humidity parameters. While acquiring crack images, it stores the current environmental temperature and relative humidity values ​​as metadata attributes in the corresponding image information archive.

3. The real-time detection method for tunnel surface cracks based on image recognition according to claim 2, characterized in that, In step 2, when processing images, the front-end odometer of visual synchronous positioning and map construction first converts the color image into a single-channel grayscale image and performs limited contrast adaptive histogram equalization processing. Feature point extraction employs feature descriptors with rotation invariance and scale invariance. By establishing a feature point pyramid structure, extreme points are searched in parallel at multiple levels from the original resolution to multi-level downsampling resolution. The inter-frame matching process uses bidirectional optical flow tracking technology combined with random sampling consensus algorithm to calculate the pixel motion vector of the current frame relative to the previous frame, and calculate the displacement of the previous frame relative to the current frame in reverse. When the bidirectional error is less than a predetermined threshold, feature points are retained, and the homography matrix is ​​constructed by iteratively selecting the minimum sample set to remove outliers that do not conform to geometric constraints. The estimation process of the six-degree-of-freedom pose parameters also includes a back-end optimization stage. By constructing a factor graph model, the visual reprojection error, the pre-integration error of the inertial measurement unit, and the lateral distance constraint term provided by the laser ranging unit are used as the optimization objective function. The Levenberg Marquardt algorithm is used to iteratively smooth the motion trajectory of the detection platform within a sliding window, minimizing the sum of squared residuals between all observation data and the motion model, and eliminating the cumulative drift error generated by the sensor.

4. The real-time detection method for tunnel surface cracks based on image recognition according to claim 3, characterized in that, In step 2, the inertial measurement unit pre-integration technology performs continuous time integration on high-frequency acceleration and angular velocity data within the time span between two adjacent visual image frames to obtain the relative velocity increment, displacement increment, and angle change of the detection platform during that time period, thereby offsetting the impact of the instantaneous impact force generated by the vehicle at the uneven track on the pose trajectory. In the pose estimation process, the method also introduces an edge constraint term, uses the Canney edge detection operator to extract the longitudinal and circumferential joint lines between the lining blocks, and adds the reprojection error of the line features to the factor graph optimization model by establishing line feature descriptors and performing line matching. To address the rolling shutter effect of the camera, the high-frequency angular velocity data provided by the inertial measurement unit is used to calculate the camera's rotation angle increment during each scan line, and the image is compensated line by line to correct the deformation caused by the rolling shutter.

5. The real-time detection method for tunnel surface cracks based on image recognition according to claim 4, characterized in that, In step 3, the process of constructing the reverse motion compensation model involves transforming the estimated rotation matrix and translation vector into a transformation operator of the image coordinate system. To address the instantaneous translational blur caused by vehicle vibration, the displacement difference between poses at adjacent time points is calculated. When the pixel displacement corresponding to the displacement difference exceeds a predetermined pixel threshold within the exposure time, a degradation function representing the motion blur kernel function is constructed in the frequency domain, and the image is deconvolved using an inverse filtering algorithm. To address the perspective distortion caused by the camera's optical axis not being perpendicular to the tunnel wall, the homography matrix is ​​used to project the current tilted imaging plane onto a reference plane parallel to the tunnel lining surface, and the coordinates of each pixel in the image are remapped. During geometric correction, the system also dynamically corrects the depth factor in the projection transformation matrix based on the object distance data fed back in real time by the laser ranging unit, ensuring that the actual physical length represented by each pixel in the final standardized image remains constant when the detection platform shifts laterally.

6. The real-time detection method for tunnel surface cracks based on image recognition according to claim 5, characterized in that, The process of generating the standardized tunnel surface image also includes the correction of intrinsic lens distortion. Using pre-calibrated radial and tangential distortion coefficients, a coordinate offset correction mapping table covering the entire pixel area is established, and sub-pixel level resampling is performed on each pixel position to eliminate the stretching phenomenon generated in the image edge area. Step 3 also employs an image reprojection technique based on a non-parametric model. A dense lookup table containing nonlinear distortions within the camera is pre-established using a large field-of-view calibration field. When processing pose increments, local affine transformations are performed on each image sub-region according to the camera imaging model. At the data processing level, the method involves a multi-threaded parallel processing architecture, which allocates image acquisition tasks, pose estimation tasks, crack recognition tasks, and global localization tasks to different cores of the computing unit, and realizes data flow exchange through a high-speed memory sharing mechanism and a double buffer mechanism.

7. The real-time detection method for tunnel surface cracks based on image recognition according to claim 6, characterized in that, In step 4, the preset crack detection neural network model adopts a symmetrical structure of encoder and decoder. The encoder part compresses the spatial dimension of the feature map and increases the feature channel depth layer by layer through cascaded multi-layer convolution operation, normalization operation and nonlinear activation operation to extract the deep semantic information of cracks in the image. The decoder part gradually restores the spatial resolution of the image through upsampling operations and uses a skip connection structure to fuse the low-level high-resolution features in the encoder with the high-level semantic features in the decoder. At the end of the network structure, a spatial and channel attention mechanism module is introduced to automatically learn the importance weights of each region of the image, focusing on the pixel set with crack features, and suppressing false detections caused by background interference from water seepage, oil stains, and construction joints in the tunnel. The crack detection neural network model also incorporates knowledge distillation technology, which transfers the crack discrimination knowledge learned by the expert model to the lightweight student model, making the processing time of a single frame image less than 20 milliseconds.

8. The real-time detection method for tunnel surface cracks based on image recognition according to claim 7, characterized in that, The extraction of the geometric contour features of the crack includes morphological skeletonization of the crack mask image, obtaining the central axis of the crack through a continuous peeling algorithm, searching for edge pixels on both sides along the normal direction of the central axis, calculating the Euclidean distance between two pairs of edge pixels to obtain the crack width, and calculating the maximum width, average width and total length of the crack accumulated from the central axis. All pixel-scale measurements are combined with the camera's intrinsic parameter matrix data, distortion correction coefficients, and real-time feedback object distance data to be converted into physically meaningful millimeter-level values. The method also includes an online update strategy for deep learning models. When the system encounters new diseases or special background textures, the manual intervention unit will confirm the suspected areas and add the marked new samples to the training set, and update the model weights in the vehicle computing unit through an incremental learning algorithm. In the post-processing stage of mask generation, a physical prior-based classifier is added. By utilizing the geometric aspect ratio, directional distribution characteristics, and relative positional relationship of the cracks with the lining joints, the identified targets are distinguished into structural stress cracks, surface shrinkage cracks, water seepage interference, and construction scratches.

9. The real-time detection method for tunnel surface cracks based on image recognition according to claim 8, characterized in that, In step 5, the establishment of the global world coordinate system is based on the preset reference point at the tunnel entrance. The synchronous positioning and map building algorithm continuously maintains and dynamically updates a sparse point cloud map that records the coordinates of spatial points with significant geometric features and feature vector descriptors during operation. When the detection vehicle passes through the same area again, it uses descriptor matching to associate the current frame with map points, triggering a loop closure detection mechanism. The loop closure detection mechanism uses a bag-of-words model to convert image features into visual word vectors, and calculates the similarity between the current vector and historical image vectors to identify whether the detection platform has passed through previously visited locations. Once a loop closure is detected, the system starts a global pose graph optimization algorithm, using the poses of all historical nodes as variables to be optimized and the loop closure constraint as an error term to adjust the global trajectory and minimize the global error. The sparse point cloud map is simplified using voxelized raster filtering, and the old feature points in the point cloud map are evaluated for lifetime decay, retaining the feature points that are stably observed in multiple inspections.

10. The method for real-time detection of tunnel surface cracks based on image recognition according to claim 9, characterized in that, The historical evolution monitoring process includes automated alignment logic. The system retrieves historical images and historical crack features within the same physical coordinate range from a database based on a spatial index structure. The database stores the detection data in segments according to the tunnel mileage and ring segment number. An improved iterative nearest point algorithm is used to register the current crack mask point cloud with the historical crack mask point cloud. By calculating the percentage increase in area, the amount of width expansion, the amount of length extension, and the number of new branches of the same crack under different periods, a tunnel structure health diagnosis report is automatically generated. If the change in crack size exceeds the preset safety threshold, the system will automatically trigger a graded alarm command. The method is applied to an automated tunnel inspection robot, which has the ability to autonomously plan its path. When a suspected serious crack area is detected, it automatically reduces its travel speed or executes a stop and shooting command at a predetermined coordinate position to obtain a high signal-to-noise ratio image using long exposure. The system ultimately constructs a three-dimensional visualized digital twin tunnel model, displaying crack images from historical years, physical size parameters, environmental temperature and humidity records, and predicted evolution curves of crack development.