Deep learning driven adaptive inspection method and system for tunnel crack identification

By combining the SegFormer semantic segmentation network with adaptive speed control, high-precision and high-efficiency detection of tunnel cracks is achieved, solving the problems of low detection efficiency and insufficient accuracy in existing technologies. It is applicable to the automated detection of various structures, reduces hardware dependence, and improves security.

CN121482473APending Publication Date: 2026-02-06SHANDONG UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511663953.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing technologies for tunnel crack identification suffer from low detection efficiency and insufficient accuracy, and are difficult to achieve efficient and accurate adaptive control in complex environments, resulting in poor detection performance.

Method used

The SegFormer semantic segmentation network architecture, combined with adaptive speed control, generates crack segmentation masks, skeleton lines, and confidence maps through real-time image acquisition, preprocessing, and deep learning model recognition. It also calculates image clarity, recognition uncertainty, and crack continuity indices, and dynamically adjusts the inspection speed.

Benefits of technology

It achieves high-precision and high-efficiency detection of tunnel cracks, improves detection accuracy and efficiency, has good cross-scenario applicability and engineering promotion value, supports the detection of multiple types of structures, reduces hardware dependence and improves safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482473A_ABST
    Figure CN121482473A_ABST
Patent Text Reader

Abstract

The invention provides a deep learning driven adaptive inspection method and system for tunnel crack recognition, and belongs to the technical field of tunnel crack recognition, and the method comprises the steps: collecting an image in a tunnel in real time, and carrying out the preprocessing of the collected image; inputting the preprocessed image into a deep learning crack recognition model to obtain a recognition result; wherein the deep learning crack identification model comprises an encoder, a decoder and an output end; the encoder encodes the input preprocessed image and outputs four feature layers with different resolutions, and the details, the trend, the overall form and the global context of the crack are reserved respectively; the decoder performs up-sampling and fusion on the four feature layers, recovers high-resolution features, and generates a crack segmentation mask; calculating an image definition score, identification uncertainty and a crack continuity index; and controlling the speed of inspecting the image in the tunnel based on the calculated image definition score, the recognition uncertainty and the crack continuity index self-adaptive speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of tunnel crack identification technology, and particularly relates to a deep learning-driven adaptive inspection method and system for tunnel crack identification. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] The causes of cracks in concrete structures include material shrinkage, load, environmental erosion, and construction defects. Cracks not only affect the appearance of the structure, but can also become the starting point for water seepage, steel corrosion, and load-bearing capacity degradation, posing a serious threat to structural safety.

[0004] Methods for inspecting cracks in concrete structures include traditional inspection methods and crack recognition based on convolutional neural networks. Traditional inspection methods primarily rely on manual labor, typically using the naked eye or handheld devices. This approach is inefficient, subjective, and risky, and has significant limitations in inspecting long tunnels and high-pier bridges. In recent years, automatic recognition methods based on image processing and deep learning have gradually emerged, but they still mainly rely on inspection vehicles operating at low, constant speeds to ensure image clarity and recognition stability, resulting in low efficiency.

[0005] While crack detection methods based on convolutional neural networks (CNNs) have achieved some success, with lightweight models such as LR-ASPP and BiSeNet V2 demonstrating good real-time performance in certain scenarios, their computational cost remains high, making it difficult to balance accuracy, real-time performance, and robustness. In complex backgrounds or lighting conditions, CNN models are prone to misclassification or missed detection. Furthermore, existing real-time recognition methods using deep learning in conjunction with drones or inspection vehicles generally employ fixed speeds or preset sampling strategies, lacking adaptive control mechanisms based on recognition results. This makes it difficult to resolve the contradiction between detection accuracy and efficiency, and also prevents dynamic optimization of detection strategies in complex environments, thus affecting the effectiveness of practical engineering applications. Summary of the Invention

[0006] To overcome the shortcomings of the prior art, this invention provides a deep learning-driven adaptive inspection method and system for tunnel crack identification, achieving a balance between high accuracy and high efficiency in crack detection.

[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: Firstly, a deep learning-driven adaptive inspection method for tunnel crack identification is disclosed, including: Real-time acquisition of images inside the tunnel and preprocessing of the acquired images; The preprocessed image is input into a deep learning crack recognition model to obtain the recognition result; The deep learning crack recognition model is built on the SegFormer semantic segmentation network architecture, which includes a Transformer encoder, a lightweight MLP decoder, and an output. SegFormer, as a general semantic segmentation framework, consists of a hierarchical Transformer encoder and a lightweight MLP decoder. The Transformer encoder encodes the preprocessed input image and outputs four feature layers with different resolutions, which respectively preserve the details, direction, overall shape and global context of the crack. The lightweight MLP decoder upsamples and fuses four feature layers to recover high-resolution features and generate a crack segmentation mask. The output terminal is used to output the crack segmentation mask, skeleton lines, and confidence map; Based on the data output from the output terminal, calculate the image sharpness score, identification uncertainty, and crack continuity index; The speed of inspecting images inside the tunnel is adaptively controlled based on the calculated image sharpness score, identification uncertainty, and crack continuity index.

[0008] Secondly, a deep learning-driven adaptive inspection system for tunnel crack identification is disclosed, including: The image acquisition module is configured to: acquire images inside the tunnel in real time and preprocess the acquired images; The recognition module is configured to input the preprocessed image into a deep learning crack recognition model to obtain the recognition result. The deep learning crack recognition model is built on the SegFormer semantic segmentation network architecture, which includes a Transformer encoder, a lightweight MLP decoder, and an output. SegFormer, as a general semantic segmentation framework, consists of a hierarchical Transformer encoder and a lightweight MLP decoder. The Transformer encoder encodes the preprocessed input image and outputs four feature layers with different resolutions, which respectively preserve the details, direction, overall shape and global context of the crack. The lightweight MLP decoder upsamples and fuses four feature layers to recover high-resolution features and generate a crack segmentation mask. The output terminal is used to output the crack segmentation mask, skeleton lines, and confidence map; The evaluation module is configured to: calculate the image sharpness score, recognition uncertainty, and crack continuity index based on the data output by the output terminal, and comprehensively evaluate the reliability of the recognition result; The adaptive speed control module is configured to adaptively control the speed of images inspected inside the tunnel based on the calculated image sharpness score, identification uncertainty, and crack continuity index.

[0009] The above one or more technical solutions have the following beneficial effects: This invention integrates the SegFormer semantic segmentation network with adaptive speed control to achieve real-time crack identification, millimeter-level parameter measurement, and long-term evolution analysis. This method not only improves detection accuracy and efficiency but also possesses good cross-scenario applicability and engineering application value.

[0010] This invention introduces the SegFormer semantic segmentation network, whose multi-scale feature fusion and global modeling capabilities surpass those of traditional CNN architectures. Combined with skeleton constraint branches, it enhances crack connectivity analysis. Experiments show that this method can effectively identify minute cracks as narrow as 0.3 mm, avoiding the misjudgments of noise, structural seams, and reflective areas found in traditional methods. It achieves millimeter-level precision in crack geometric parameter measurement, significantly improving detection accuracy.

[0011] Significantly improved detection efficiency: By introducing adaptive speed control based on image quality, uncertainty, and crack continuity, this invention overcomes the technical bottleneck of "high speed leading to inaccuracy, and slow speed leading to inefficiency." Vehicles or drones can automatically adjust their operating speed according to the detection results, maintaining efficient inspection while ensuring accuracy, thereby improving the detection efficiency of long-distance tunnels and long-span bridges by more than 20%.

[0012] Enhanced robustness and adaptability: In complex environments, such as uneven lighting, surface water seepage, crack obstruction, or rough backgrounds, this invention dynamically adjusts the detection strategy through joint evaluation of the mass fraction Qt and uncertainty Ut, ensuring the stability and reliability of the model output. This robustness far surpasses traditional detection methods that rely on fixed speeds and single thresholds.

[0013] Cross-platform and multi-scenario applicability: This invention is not only applicable to tunnel crack detection, but can also be extended to various structures such as bridges, subways, high piers, and underground utility tunnels. Its detection platform can be flexibly deployed on ground inspection vehicles, drones, or robots, supporting a "vehicle-mounted + drone" joint detection mode, and has strong engineering scalability.

[0014] Risk Assessment and Lifecycle Management: This invention enables multi-period comparative analysis of cracks, quantifies crack propagation rates, classifies risks, and generates long-term health monitoring curves. This function provides a scientific basis for the lifecycle management of infrastructure, supporting maintenance departments in shifting from a "reactive maintenance" to a "proactive prevention" management model.

[0015] Low hardware dependency and flexible deployment: The SegFormer model has a lightweight structure with small computational load and parameter size, enabling it to run in real time on embedded platforms and edge computing modules with low to medium computing power. This invention achieves online detection without the need for high-performance servers, reducing engineering application costs and facilitating rapid deployment on conventional inspection equipment.

[0016] Enhanced Safety and Automation: Traditional manual inspections require personnel to enter tunnels or climb bridge structures, posing significant safety risks. This invention employs automated inspection methods, replacing manual labor with inspection vehicles and drones, reducing the time personnel are exposed to high-risk environments and significantly improving the safety and automation of inspection operations.

[0017] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0018] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0019] Figure 1 This is an overall flowchart of the method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the SegFormer encoder and decoder structure according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the layout of inspection vehicle equipment according to an embodiment of the present invention. Detailed Implementation

[0020] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0021] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0022] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0023] In addition, the Transformer architecture, which has emerged in recent years, has demonstrated powerful global modeling capabilities in the field of computer vision. SegFormer, as a semantic segmentation model that combines a lightweight Transformer with a multilayer perceptron (MLP) structure, has performed well on datasets such as ADE20K and Cityscapes.

[0024] Studies have shown that SegFormer-based crack detection models significantly outperform CNNs in terms of accuracy, real-time performance, and robustness. This provides a new approach for engineering inspection, but currently lacks a comprehensive solution that incorporates an adaptive inspection control mechanism.

[0025] Therefore, how to fully leverage the advantages of the SegFormer model in image detection and combine it with adaptive speed adjustment of vehicles / UAVs to achieve both high accuracy and high efficiency in crack detection is a key issue that urgently needs to be addressed.

[0026] Example 1 This embodiment discloses a deep learning-driven adaptive inspection method for tunnel crack identification, including: Step 1: Data Acquisition and Preprocessing: Acquire images using vehicle-mounted cameras, drones, or auxiliary sensors (IMU, odometer); perform illumination normalization, deblurring, and geometric correction.

[0027] In practice, vehicle-mounted cameras are typically positioned at the front or top of the vehicle, kept perpendicular or nearly perpendicular to the detection surface (angle ≤15°) to reduce perspective distortion.

[0028] Drones: Depending on the location of the bridge or tunnel wall, keep the camera's optical axis as directly as possible facing the inspection surface; for arches or high piers, tilt shooting within the range of 30°–60° can be used to cover the entire surface of the structure.

[0029] Step 2: Deep learning-based crack identification based on SegFormer: 2-1) The encoder part adopts a hybrid Transformer (Mix), which outputs four feature layers {F1, F2, F3, F4} with different resolutions through overlapping patch embedding and hierarchical Transformer blocks, respectively preserving the details, direction, overall shape and global context of the crack.

[0030] The input image first undergoes overlapping patch embedding, which divides the image into sub-blocks with overlapping regions and maps them to a high-dimensional feature space to preserve the details of the crack edges. Then, it is processed layer by layer by a hierarchical Transformer block, which extracts high-resolution detail features at the lower layers and obtains low-resolution global semantic information at the higher layers. Finally, it outputs four feature layers {F1, F2, F3, F4} with different resolutions, which respectively represent the details, direction, overall shape and global context of the crack.

[0031] 2-2) The decoder section employs a lightweight full MLP architecture, performing unified channel processing and layer-by-layer upsampling on the four different resolution feature layers {F1, F2, F3, F4} output by the encoder. The low-resolution feature layers are upsampled to a spatial size matching the high-resolution feature layers via interpolation, and then feature stitching and fusion are performed using a multilayer perceptron, preserving both the detailed information and global context of the cracks. The fused high-resolution features are mapped to a crack segmentation mask by the prediction head, representing the location and morphology of crack regions in the input image.

[0032] 2-3) A skeleton constraint branch is added to the decoder. The center line of the crack is extracted through orientation-aware convolution to enhance connectivity. See Appendix for the specific structure. Figure 2 As shown.

[0033] 2-4) The output includes a crack mask, skeleton lines, and uncertainty map, where the confidence is obtained by MC Dropout or a deep ensemble method.

[0034] The aforementioned skeleton line is generated by the high-resolution feature map output by the decoder through a skeleton constraint branch. This branch extracts the crack centerline features through orientation-aware convolution and refines the crack region by combining it with a crack segmentation mask, thereby obtaining a crack skeleton line map, which is used to characterize the crack's orientation and connectivity.

[0035] The confidence map described above is generated by an uncertainty estimation mechanism based on the model's predicted output. During the inference phase, MC Dropout or deep ensemble methods are used to perform multiple inferences on the crack segmentation results, calculating the variance of the predicted results for each pixel and mapping it to the same resolution as the input image, thus obtaining the uncertainty distribution map. This confidence map is used to evaluate the reliability of the crack identification results and provides input for adaptive speed control.

[0036] Step 3: Quality and Uncertainty Assessment: Calculate the image sharpness score (Q), recognition uncertainty (U), and crack continuity index (C) to comprehensively evaluate the reliability of the recognition results.

[0037] The above quality and uncertainty assessment includes the following calculation process: Image sharpness score Q is calculated by obtaining the image sharpness index through blur kernel estimation or Laplacian variance method, and then combining the uniformity of brightness histogram and the sharpness of crack segmentation edges for weighted calculation to obtain the normalized image quality score Q. Calculation of uncertainty U: During the model inference stage, MC Dropout or deep ensemble method is used to perform multiple forward inferences, calculate the variance or entropy value of the predicted probability of each pixel, and perform a weighted average on the whole image to obtain the uncertainty index U; Calculation of crack continuity index C: The connectivity of the crack centerline output by the skeleton branch is analyzed, and the number of fracture segments or the proportion of fracture length of the skeleton line are counted to quantitatively reflect the completeness of the crack prediction results and obtain the continuity index C.

[0038] Step 4: Adaptive Speed ​​Control: Control Law:

[0039] in, This is the quality threshold. If Qt is low or Ut is high, automatic deceleration occurs; if Qt is high and Ut is low, acceleration occurs; in suspected crack areas, short-term deceleration is triggered to acquire multiple frames.

[0040] In one implementation example, for adaptive speed control, the image sharpness score Qt is related to a preset quality threshold. Compare, if Qt is lower The system determines the image is unclear and triggers deceleration; it identifies the uncertainty Ut and compares it with a preset uncertainty threshold. If the comparison is made, and Ut is higher than... If the system determines that the model prediction is unreliable, it will trigger deceleration; when Qt is higher than... And Ut is lower than When the system determines that the image is clear and the prediction is reliable, the vehicle or drone will automatically accelerate; when the crack continuity index Ct is large and there is a suspected crack area, the system will trigger a short-term deceleration and acquire multiple frames of images.

[0041] In one implementation example, The quality threshold is a parameter set for the system. It can be defined as the minimum acceptable sharpness required for the recognition task, and can be set to 0.7 in engineering. The uncertainty threshold is a parameter set for the system to evaluate the reliability of crack identification results. In engineering, it can be set to 0.85.

[0042] Step 5: Precise Measurement of Crack Parameters: Using camera calibration and vehicle pose information, the crack pixel size is converted into millimeter-level parameters such as length and width. Inspection data from multiple time periods are registered to analyze the crack propagation rate and risk level.

[0043] In this embodiment, the accurate measurement of crack parameters includes the following steps: obtaining intrinsic and extrinsic parameters through camera calibration, and establishing a mapping relationship between pixel coordinates and actual physical dimensions; Using vehicle pose information, spatial localization is performed on the acquired images, and the crack pixels are converted into a point set in the world coordinate system; Pixel-level measurements are performed on the crack segmentation mask to calculate geometric features such as crack length and width, and the results are converted into millimeter-level true parameters based on the calibration results. Register the inspection data from multiple time periods to unify the crack measurement results from different periods into the same coordinate system; By comparing the geometric parameters of cracks at different time periods, the crack propagation rate is calculated, and the crack risk level is determined and classified in combination with a preset threshold.

[0044] In step one, to achieve real-time crack identification and speed-adaptive inspection based on deep learning, the ground inspection vehicle is equipped with a global shutter industrial camera (1920×1080, 120 fps), an LED strobe light source (100µs pulse width), an IMU, and an odometer during image acquisition. Figure 3 As shown, the vehicle cruised at 6 m / s in the tunnel. Images were input into the SegFormer variant model for crack segmentation. If the identification uncertainty exceeded 0.3, the controller reduced the vehicle speed to 1.5 m / s and continuously acquired 5 frames for super-resolution reconstruction. Experimental results showed a 15% improvement in crack detection rate, while maintaining an average speed of 5 m / s across the entire route.

[0045] The Qt image quality score is derived from preprocessing and the output of the SegFormer model. It is calculated in the "Quality and Uncertainty Evaluation" step and includes image sharpness (blur kernel estimation), brightness uniformity, and edge sharpness of the crack segmentation results. It is used to reflect whether the current frame image is sharp enough. The larger the value, the higher the image quality.

[0046] The quality threshold is a system-defined parameter used to identify the minimum acceptable sharpness required for a task, for example... This serves as a baseline for judging whether the image quality meets the requirements for crack recognition.

[0047] Ut (Identification Uncertainty) is derived from the confidence output of the SegFormer model. It is obtained in the "Deep Learning Crack Identification" step through MC Dropout or deep ensemble and is used to represent the confidence reliability of crack segmentation prediction. The larger the value, the more uncertain the model prediction is.

[0048] Ct (crack continuity index) is derived from the skeleton branch output. It is calculated in the "skeleton extraction" step by detecting the skeleton connectivity and degree of fracture, such as by statistically analyzing the skeleton length ratio or the number of fracture segments. The larger the value, the less consistent the crack prediction results are, and the vehicle needs to slow down for verification.

[0049] k1, k2, and k3 (weighting coefficients) are control strategy parameters set in the "Adaptive Speed ​​Control" step. They are used to balance the effects of image sharpness, uncertainty, and crack continuity on speed control. The sensitivity of speed adjustment can be controlled by adjusting the weighting coefficients. For example, k1=0.5, k2=0.7, and k3=0.4.

[0050] v t With v t+1 (Vehicle speed) is derived from the vehicle operation control module, where v t v is the current speed. t+1 To achieve the target speed adjusted according to the control law, the clip function is used to constrain the speed range, keeping it within a certain range. The range, such as [0.5, 3.0] m / s, is used to ensure the safety and stability of the inspection process.

[0051] For adaptive velocity control under complex lighting and water seepage environments, in tunnel sections with strong reflections and water seepage, the image sharpness score Q drops below 0.5. The deep learning model outputs a crack mask that is broken. Based on the aforementioned control law, the controller automatically reduces the velocity from 6.0 m / s to 2.0 m / s, increases the stroboscopic brightness, and re-acquires images. After fusion processing, the crack measurement error is controlled within ±0.05 mm.

[0052] Crack evolution trend analysis: By registering inspection data from two separate inspections of the same tunnel, six months apart, it was found that the width of a longitudinal crack increased from 0.32 mm to 0.58 mm at a certain location. The system calculated the propagation rate and identified the area as high-risk, issuing an early warning. This demonstrates that the method can support long-term crack evolution analysis and risk classification.

[0053] Joint UAV-vehicle inspection: UAVs are used to conduct extensive inspections of the bridge's underside to initially identify suspected cracks, followed by the deployment of vehicle-mounted inspection vehicles to measure millimeter-level parameters. Experiments have shown that this joint approach significantly improves inspection efficiency and accuracy.

[0054] The aforementioned drone-vehicle joint inspection includes the following steps: Drones conducted extensive inspections of the underside of the bridge, collecting panoramic images and identifying suspected crack areas based on lightweight models. The spatial location information of the suspected crack area is obtained through GPS / RTK or SLAM technology and transmitted to the ground control center; The control center maps the drone detection results to a unified coordinate system and generates inspection task instructions for the vehicle-mounted inspection vehicle. The vehicle-mounted inspection vehicle navigates to the corresponding area according to the instructions and decelerates within the target area. The vehicle-mounted inspection vehicle uses a high-resolution camera and adaptive speed control to measure millimeter-level parameters of the cracks and uploads the results to the monitoring system.

[0055] Example 2 The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the above-described method.

[0056] Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.

[0057] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the above method.

[0058] Example 4 The purpose of this embodiment is to provide a deep learning-driven adaptive inspection system for tunnel crack identification, including: The image acquisition module is configured to: acquire images inside the tunnel in real time and preprocess the acquired images; The recognition module is configured to input the preprocessed image into a deep learning crack recognition model to obtain the recognition result. The deep learning crack recognition model includes an encoder, a decoder, and an output. The encoder encodes the preprocessed input image and outputs four feature layers with different resolutions, which respectively preserve the details, direction, overall shape and global context of the crack. The decoder upsamples and fuses the four feature layers to restore high-resolution features and generate a crack segmentation mask. The output terminal is used to output the crack segmentation mask, skeleton lines, and confidence map; The evaluation module is configured to: calculate the image sharpness score, recognition uncertainty, and crack continuity index based on the data output by the output terminal, and comprehensively evaluate the reliability of the recognition result; The adaptive speed control module is configured to adaptively control the speed of images inspected inside the tunnel based on the calculated image sharpness score, identification uncertainty, and crack continuity index.

[0059] Example 5 The purpose of this embodiment is to provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods and functions involved in any of the above embodiments. The steps and methods involved in the apparatus of the above embodiments correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0060] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0061] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A deep learning-driven adaptive inspection method for tunnel crack identification, characterized by: include: Real-time acquisition of images inside the tunnel and preprocessing of the acquired images; The preprocessed image is input into a deep learning crack recognition model to obtain the recognition result; The deep learning crack recognition model is built on the SegFormer semantic segmentation network architecture, which includes a Transformer encoder, a lightweight MLP decoder, and an output. SegFormer, as a general semantic segmentation framework, consists of a hierarchical Transformer encoder and a lightweight MLP decoder. The Transformer encoder encodes the preprocessed input image and outputs four feature layers with different resolutions, which respectively preserve the details, direction, overall shape and global context of the crack. The lightweight MLP decoder upsamples and fuses four feature layers to recover high-resolution features and generate a crack segmentation mask. The output terminal is used to output the crack segmentation mask, skeleton lines, and confidence map; Based on the data output from the output terminal, calculate the image sharpness score, identification uncertainty, and crack continuity index; The speed of inspecting images inside the tunnel is adaptively controlled based on the calculated image sharpness score, identification uncertainty, and crack continuity index.

2. The deep learning-driven adaptive inspection method for tunnel crack identification as described in claim 1, characterized in that, The Transformer encoder encodes the preprocessed input image and outputs four feature layers at different resolutions, specifically including: The input image is first processed by overlapping block embedding, which divides the image into sub-blocks with overlapping regions and maps them to a high-dimensional feature space to preserve the details of the crack edges. Then, it is processed layer by layer by the hierarchical Transformer module, which extracts high-resolution detail features at the lower layers and obtains low-resolution global semantic information at the higher layers. Finally, four feature layers {F1, F2, F3, F4} with different resolutions are output, which respectively represent the details, direction, overall shape and global context of the crack.

3. The deep learning-driven adaptive inspection method for tunnel crack identification as described in claim 2, characterized in that, The lightweight MLP decoder upsamples and fuses four feature layers to recover high-resolution features, specifically including: The low-resolution feature layers from the four different resolution feature layers are sampled to match the spatial dimensions of the high-resolution feature layers through interpolation. These features are then stitched and fused using a multilayer perceptron to simultaneously preserve both the detailed information and global contextual information of the cracks. The fused high-resolution features are mapped by a prediction head to a crack segmentation mask, used to characterize the location and morphology of crack regions in the input image.

4. The deep learning-driven adaptive inspection method for tunnel crack identification as described in claim 1, characterized in that, When the adaptive speed control measures the speed of image inspection within the tunnel, the image sharpness score Qt is compared with a preset quality threshold. Compare, if Qt is lower The system determines the image is unclear and triggers deceleration; it identifies the uncertainty Ut and compares it with a preset uncertainty threshold. If the comparison is made, and Ut is higher than... If the system determines that the model prediction result is unreliable, it will trigger deceleration. When Qt is higher And Ut is lower than When the system determines that the image is clear and the prediction is reliable, the vehicle or drone will automatically accelerate; when the crack continuity index Ct is large and there is a suspected crack area, the system will trigger a short-term deceleration and acquire multiple frames of images.

5. The deep learning-driven adaptive inspection method for tunnel crack identification as described in claim 1, characterized in that, The confidence map is generated by the uncertainty estimation mechanism of the model prediction output. In the inference stage, MC Dropout or deep ensemble method is used to perform multiple inferences on the crack segmentation results, calculate the variance of the prediction results of each pixel and map it to the same resolution as the input image, thereby obtaining the uncertainty distribution map. This confidence map is used to evaluate the reliability of the crack identification results.

6. The deep learning-driven adaptive inspection method for tunnel crack identification as described in claim 1, characterized in that, When calculating image sharpness score, recognition uncertainty, and crack continuity index, the following are included: Image sharpness score Q: The image sharpness index is obtained by means of blur kernel estimation or Laplacian variance method, and the normalized image quality score Q is obtained by weighting the uniformity of brightness histogram and the sharpness of crack segmentation edge. Uncertainty U is identified by performing multiple forward inferences using MC Dropout or deep ensemble methods during the model inference stage, calculating the variance or entropy of the predicted probability of each pixel, and then performing a weighted average across the entire image to obtain the uncertainty index U. Crack continuity index C: The connectivity of the crack centerline output by the skeleton branch is analyzed, and the number of fracture segments or the proportion of fracture length of the skeleton line are counted to quantitatively reflect the completeness of the crack prediction results, and the continuity index C is obtained.

7. A deep learning-driven adaptive inspection system for tunnel crack identification, characterized by: include: The image acquisition module is configured to: acquire images inside the tunnel in real time and preprocess the acquired images; The recognition module is configured to input the preprocessed image into a deep learning crack recognition model to obtain the recognition result. The deep learning crack recognition model includes an encoder, a decoder, and an output. The encoder encodes the preprocessed input image and outputs four feature layers with different resolutions, which respectively preserve the details, direction, overall shape and global context of the crack. The decoder upsamples and fuses the four feature layers to restore high-resolution features and generate a crack segmentation mask. The output terminal is used to output the crack segmentation mask, skeleton lines, and confidence map; The evaluation module is configured to: calculate the image sharpness score, recognition uncertainty, and crack continuity index based on the data output by the output terminal, and comprehensively evaluate the reliability of the recognition result; The adaptive speed control module is configured to adaptively control the speed of images inspected inside the tunnel based on the calculated image sharpness score, identification uncertainty, and crack continuity index.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method described in any one of claims 1-6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it performs the steps of the method described in any one of claims 1-6 above.

Citation Information

Cited By

  • Intelligent tunnel disease inspection method based on multi-model fusion and consistency verification

    CN121921660A