Mountain area tunnel lining crack automatic identification method combining unmanned aerial vehicle inspection and target detection model, storage medium and equipment

By using an unmanned aerial vehicle (UAV) platform and an improved YOLOv10 model, the problems of low efficiency and high safety risks in tunnel lining crack detection have been solved, enabling efficient and safe crack identification and quantitative analysis, and improving the reliability of tunnel health monitoring.

CN121811234APending Publication Date: 2026-04-07CHINA THREE GORGES PROJECTS DEV CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing methods for detecting cracks in tunnel linings are inefficient and pose high safety risks. Furthermore, existing target detection models are difficult to deploy efficiently in resource-constrained scenarios, resulting in missed or false detections.

Method used

An unmanned aerial vehicle (UAV) platform combined with an improved YOLOv10 model was used for automatic identification of tunnel lining cracks. Through an efficient multi-scale attention module, an AutoAugment data augmentation strategy, and focus loss, high-precision location and identification of cracks were achieved.

Benefits of technology

It enables high-precision image data acquisition that is available 24/7, in all weather conditions, and without contact, improving the efficiency and safety of tunnel maintenance work and providing a stable basis for quantitative analysis of cracks and risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811234A_ABST
    Figure CN121811234A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of mountainous area tunnel lining defect detection, and particularly provides a mountainous area tunnel lining crack automatic identification method combining unmanned aerial vehicle inspection and a target detection model, and the method comprises the steps: obtaining the section size and shape of a mountainous area tunnel and camera sensor parameters; the unmanned aerial vehicle is controlled to fly at a constant speed in the axial direction of the tunnel, the binocular camera collects a lining image sequence, and a structured original data set is constructed; performing image screening, optical preprocessing, stereo matching and cylindrical surface expansion on the original data set to generate a two-dimensional expansion view of the tunnel lining; performing bounding box labeling on cracks in the two-dimensional expansion drawing, and dividing the cracks into a training set, a verification set and a test set according to a preset proportion; carrying out crack detection training by adopting an improved YOLOv10 model; and inputting the two-dimensional expansion graph into the trained improved YOLOv10 model, and automatically outputting position information of the crack. The unmanned aerial vehicle is used for automatic inspection, the crack area is recognized, and the tunnel maintenance work efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of mountainous tunnel lining defect detection, and particularly relates to a mountainous tunnel lining crack automatic identification method combining unmanned aerial vehicle inspection and target detection model, a storage medium and equipment. BACKGROUND

[0002] With the continuous extension of transportation, water conservancy and other infrastructure networks to complex geological conditions in the mountains, the number and scale of mountainous tunnel projects continue to grow, and the long-term health and safety of their operation period are facing severe challenges. The tunnel lining structure, as the final barrier against rock pressure, prevents rock weathering and collapse, and groundwater seepage, its integrity directly determines the safe operation of the entire tunnel and even the entire line. Once the lining has defects such as cracks and is not discovered and treated in time, it may lead to damage to facilities inside the tunnel, deterioration of the traffic environment, and even catastrophic accidents such as lining instability, surrounding rock loosening and even overall collapse, which not only causes huge direct economic losses and long-term traffic disruption, but also seriously threatens the safety of the people. Therefore, implementing systematic and precise tunnel lining crack detection is a key prerequisite for overall structural state assessment, early detection, early warning and treatment of potential safety hazards.

[0003] Currently, tunnel lining crack detection mainly faces two outstanding problems: On the one hand, the existing inspection methods are inefficient and have high safety risks. Traditional detection methods mainly rely on manual visual inspection, with technical personnel riding high-altitude work platforms for close-range investigation, relying on naked-eye observation and manual recording of cracks and other diseases. This method, although simple to operate and low in equipment cost, has obvious drawbacks: the detection results are highly subjective, low in precision and efficiency, and it is difficult to obtain quantitative data, and the high-altitude work environment is complex and has high safety risks. Even if a digital camera is used to take lining images and then manually interpreted in the office, although the images are recorded, a large amount of manpower is still required to process the data, which is heavily dependent on personnel experience and lacks real-time feedback capability, resulting in a long detection cycle and high cost.

[0004] On the other hand, with the rapid development of deep learning technology, significant breakthroughs have been made in the field of target detection, especially in the balance between detection accuracy and inference speed. However, existing methods still generally face problems such as post-processing redundancy, low utilization of model parameters, and difficulty in balancing delay and performance. Many detectors rely heavily on post-processing operations such as non-maximum suppression (NMS) to eliminate duplicate boxes, which not only introduces additional computational overhead, but also may lead to missed or false detections. In addition, the coupling design between feature extraction and target positioning in traditional architectures limits the overall efficiency of the model, making it difficult to achieve efficient deployment in resource-constrained actual scenarios. SUMMARY

[0005] The technical problem solved by the present application is to provide a mountain tunnel lining crack automatic identification method combined with unmanned aerial vehicle inspection and target detection model, a storage medium and equipment, which uses unmanned aerial vehicle automatic inspection to identify crack areas and effectively improves tunnel maintenance work efficiency.

[0006] To solve the above technical problems, the technical solution adopted by the present application is: a mountain tunnel lining crack automatic identification method combined with unmanned aerial vehicle inspection and target detection model, comprising the following steps: Step 1: preliminary data acquisition and flight planning: acquire the cross-sectional size, shape and camera sensor parameters of the mountain tunnel, and calculate and determine the flight height, flight speed and image heading overlap rate of the unmanned aerial vehicle; Step 2: unmanned aerial vehicle inspection and image acquisition: control the unmanned aerial vehicle to fly at a uniform speed along the tunnel axis, acquire a sequence of lining images through the mounted binocular camera, monitor the image quality in real time and mark the invalid frames for reflight, and construct a structured original data set; Step 3: data preprocessing: image screening, optical preprocessing, stereo matching and cylindrical unwrapping are performed on the original data set to generate a two-dimensional unwrapping diagram of the tunnel lining; Step 4: data set making: the boundaries of the cracks in the two-dimensional unwrapping diagram are labeled, the unwrapping diagram is cropped to the model input size, and the unwrapping diagram is divided into a training set, a validation set and a test set according to a preset ratio; Step 5: improved YOLOv10 model construction and training: an improved YOLOv10 model is used for crack detection training, and the improvements of the improved YOLOv10 model include: replacing the multi-head self-attention module in the original pyramid attention network with an efficient multi-scale attention module; using an AutoAugment automatic data enhancement strategy to replace Mosaic enhancement; using a focal loss to replace a binary cross-entropy loss; Step 6: crack automatic identification and result output: input the two-dimensional unwrapping diagram of the tunnel lining into the trained improved YOLOv10 model, and automatically output the position information of the cracks for tunnel maintenance and risk assessment.

[0007] In the preferred scheme, the image heading overlap rate in step 1 is ≥80%, and the flight parameters are determined by a geometric imaging model to ensure that the image ground sampling distance can completely cover the tunnel vault, sidewall and inverted arch area.

[0008] In the preferred scheme, in step 2, the unmanned aerial vehicle is mounted with a binocular camera, an RTK / UWB positioning module and an IMU unit to acquire RAW format images.

[0009] In the preferred solution, in step 2, the unmanned aerial vehicle inspection adopts an automatic exposure bracketing and synchronous triggering strategy, and the binocular camera synchronously collects images; during the flight, the ground station monitors the image definition, binocular synchronization and stereoscopic coverage range in real time, and invalid frames with blurring, overexposure or large synchronization deviation are removed.

[0010] In the preferred solution, in step 3, the image screening adopts a Laplacian variance algorithm and average gray value judgment: the blurred frames are removed by calculating the response variance of the image Laplacian operator, and the blurred frames are determined if the corresponding variance is lower than the set threshold; the overexposed or underexposed frames are removed by calculating the average gray value of the image, and the exposure abnormality is determined if the gray value is outside the set range.

[0011] In the preferred solution, the calculation formula of the Laplacian variance algorithm is: ; Wherein, is the response variance; x i is the pixel value of the image after the Laplacian operator processing; μ is the arithmetic mean of the pixel value; n is the total number of pixels.

[0012] In the preferred solution, the calculation formula of the average gray value is: ; Wherein, I(i,j) is the gray value of the pixel point (i,j); M represents the number of pixels in the vertical direction of the image, that is, the height of the image; N represents the number of pixels in the horizontal direction of the image, that is, the width of the image.

[0013] In the preferred solution, in step 3, the optical preprocessing adopts Zhang Zhengyou calibration method, and the radial distortion and tangential distortion of the image are eliminated by using the preset camera intrinsic matrix and lens distortion coefficient.

[0014] In the preferred solution, in step 3, the specific process of stereomatching and cylindrical development is: the half-global matching algorithm is used to calculate the disparity map of the corrected binocular image pair, the disparity is converted into depth information according to the disparity map, the camera intrinsic parameter and the baseline distance, each pixel point in the depth map is back projected to the three-dimensional space, and the three-dimensional point cloud is generated; an ideal cylindrical surface aligned with the tunnel center axis is defined, the three-dimensional point cloud is projected to the ideal cylindrical surface, and after being cut along the generatrix, it is developed into a two-dimensional tunnel lining development map, and the image texture information is synchronously mapped.

[0015] In a preferred solution, in step 4, the crack annotation uses an external rectangular frame to define each independent crack instance, and the bounding box tightly wraps the crack body and distinguishes the crack from other similar features; the image cropping uses an overlapping sampling method to crop the expanded graph into a standard input size of 640x640 pixels.

[0016] In a preferred solution, in step 4, the data set is divided into training set, validation set and test set in the ratio of 8:1:1, and the division follows the scene-independent principle, and the subgraphs of the same original expanded graph are not distributed across the sets.

[0017] In a preferred solution, in step 5, the specific implementation of the efficient multi-scale attention module EMA is: the input feature map is divided into G sub-features in the channel dimension, and is processed in parallel through a 1x1 convolution branch and a 3x3 convolution branch, the 1x1 convolution branch combines a one-dimensional global average pooling to aggregate spatial information, and the 3x3 convolution branch expands the receptive field to capture multi-scale structures; the 2D global average pooling encodes the global spatial information output by the 1x1 branch to generate a spatial attention map, and outputs after weighted fusion of the features of the two branches.

[0018] In a preferred solution, in step 5, the AutoAugment automatic data augmentation strategy automatically generates an optimal enhancement operation sequence containing rotation, translation, and color enhancement through data-driven strategy search, which is applied to the training set during training, and the validation set and test set are not subjected to data augmentation.

[0019] In a preferred solution, in step 5, the mathematical formula of the focal loss FocalLoss is: ; Wherein, y∈{0,1} is the true label, the crack is 1, and the background is 0; p is the probability of the model predicting positive examples, α is the positive and negative sample balance parameter, and γ is the focusing parameter, which is used to control the intensity of the modulation factor.

[0020] In a preferred solution, in step 5, the improved YOLOv10 model adopts a backbone-neck-head architecture: the backbone network is composed of multiple C2f modules stacked, and the expression of the C2f module is: ; Wherein: X is the input feature map; is the i-th bottleneck block; is the channel dimension splicing operation.

[0021] Each Block is a depth separable convolution; the neck is a PAN structure feature pyramid, which realizes multi-scale feature fusion through top-down and bottom-up paths; the head is a double-branch design, which shares regression parameters and independently learns classifiers to realize NMS-free inference.

[0022] In a preferred solution, in step 6, the crack position information is output in the form of a bounding box coordinate, a crack category and a confidence.

[0023] The application further provides a computer readable storage medium storing computer instructions for causing the computer to execute the mountain tunnel lining crack automatic identification method of a joint unmanned aerial vehicle inspection and target detection model.

[0024] An electronic device, characterized by comprising a memory and a processor, which are in communication connection with each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the mountain tunnel lining crack automatic identification method of a joint unmanned aerial vehicle inspection and target detection model.

[0025] The application provides a mountain tunnel lining crack automatic identification method of a joint unmanned aerial vehicle inspection and target detection model, a storage medium and an equipment, which has the following beneficial effects: 1. The application first uses an unmanned aerial vehicle platform to realize all-weather, non-contact and high-precision image data acquisition of the appearance of the mountain tunnel lining; on this basis, an optimized YOLOv10 model is used to realize high-speed and high-precision positioning of the crack target, and the excellent detection performance thereof provides a stable and reliable data basis for quantitative analysis (such as estimating the maximum length and width based on the size of the bounding box) and risk assessment of the crack, and provides reliable technical support for the structural health monitoring and early treatment of the disease of the mountain tunnel.

[0026] The visible light camera and various sensor modules carried on the unmanned aerial vehicle are used to acquire the lining images and flight data of the tunnel in real time, and the images are finally obtained through splicing to obtain a two-dimensional unfolded image of the tunnel; the data set is expanded through cropping and data enhancement algorithms, and is used to train the YOLOv10 model proposed in the application, and finally the whole tunnel two-dimensional unfolded image is imported into the model, so that the cracks on the surface are efficiently and accurately identified.

[0027] 2. The unmanned aerial vehicle can completely get rid of the restrictions of ground obstacles, accumulated water and uneven road surfaces due to its unique air maneuvering advantage, and can independently and efficiently complete large-scale tunnel scanning. Compared with artificial close-range exploration and ground movement of unmanned ground vehicles, the unmanned aerial vehicle not only greatly improves the operation efficiency and safety, but also can stably obtain comprehensive and uniform high-definition images, which provides a solid foundation for subsequent crack identification and quantitative analysis based on deep learning.

[0028] 3. This invention proposes an improved YOLOv10 model, enhancing its detection performance in complex scenarios through three core innovations: First, it replaces the multi-head self-attention module (MHSA) in the original Pyramid Attention Network (PSA) with an efficient multi-scale attention module (EMA), enhancing the model's ability to fuse multi-scale features and significantly reducing computational complexity; second, it introduces an Auto Augment automated data augmentation strategy to replace the preset Mosaic augmentation logic, improving the diversity and generalization of training data through data-driven strategy search; finally, it replaces the binary cross-entropy loss (BCE loss) with a focal loss, effectively alleviating the problem of imbalanced positive and negative samples. Attached Figure Description

[0029] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is a schematic diagram of the tunnel lining image acquired by the UAV according to the present invention; Figure 2 This is a flowchart of the overall process for automatic identification of cracks in the lining of mountain tunnels using a drone inspection and target detection model. Figure 3 This is a structural diagram of the improved YOLOv10 model; Figure 4 This is a structural diagram of the EMA module. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention.

[0031] Example 1: An automatic identification method for lining cracks in mountain tunnels using a combined UAV inspection and target detection model, such as Figure 2 As shown, it includes the following steps: Step 1: Preliminary data acquisition and flight planning: Obtain the cross-sectional dimensions, shape, and camera sensor parameters of the mountain tunnel, and calculate and determine the UAV's flight altitude, flight speed, and image heading overlap rate.

[0032] Image quality is fundamental to the effectiveness of deep learning model training. High-quality training images provide clear and accurate feature information, directly determining the model's ability to learn target features and its generalization performance. For tunnel lining crack detection, factors such as illumination uniformity, detail resolution, and motion blur in the images are particularly critical. Uneven or insufficient illumination can obscure crack textures, low-resolution images cannot provide sufficient pixel-level information, and motion blur can distort edge features. These quality issues can significantly reduce the model's recognition accuracy and robustness, and even lead to false positives and false negatives. Therefore, obtaining a high-quality tunnel lining image dataset is an indispensable prerequisite for subsequent model training and applications.

[0033] To obtain high-quality images that meet the requirements of model training, this embodiment employs a systematic approach to determine the UAV flight parameters. First, based on the tunnel cross-sectional dimensions (diameter / width, height, shape) and camera sensor parameters, theoretical flight parameters are planned using a geometric imaging model. The specific steps are as follows: 1. Create a baseline model First, a backpack laser scanner or total station is used to quickly scan the tunnel, obtaining a preliminary, reasonably accurate 3D point cloud model. The acquired point cloud data is then imported into specialized software to generate a triangular network model of the tunnel, which forms the digital benchmark for all subsequent planning.

[0034] 2. Planned flight parameters Based on the geometric imaging model obtained in the previous step, determine the position of the UAV relative to the tunnel sidewalls and arch; and set the camera orientation and angle to ensure that the entire target area is captured; strictly set the heading overlap rate (usually ≥80%). This is the lifeline for subsequent high-precision 3D modeling.

[0035] Based on this, the maximum permissible flight speed is precisely calculated according to the camera frame rate, field of view, and the overlap rate required for image stitching (typically ≥80% forward overlap). This ensures that while covering the entire tunnel surface, high-quality, high-overlap image sequences suitable for 3D reconstruction and deep learning model training are obtained. For tunnels with a cross-sectional diameter of 8-12 meters, the flight height (from the arch) is typically 3-5 meters, and the speed needs to be... 2m / s.

[0036] Step 2: UAV Inspection and Image Acquisition: Control the UAV to fly at a constant speed along the tunnel axis, acquire lining image sequences through the onboard binocular camera, monitor image quality in real time, mark invalid frames for re-flying, and construct a structured original dataset.

[0037] Based on the completed flight plan, this phase involves a systematic UAV image acquisition operation, aiming to obtain a high-quality image sequence that fully covers the tunnel lining surface, providing a foundation for subsequent data processing and analysis.

[0038] First, based on the preset flight path and acquisition parameters, the drone is controlled to fly at a constant speed along the tunnel axis, maintaining a stable flight attitude and shooting distance relative to the tunnel. By adjusting the camera exposure parameters and trigger interval, it is ensured that the acquired image sequence has good illumination consistency and sufficient directional overlap.

[0039] Secondly, image quality is monitored in real time during flight, and invalid frames caused by motion blur, defocus, or abnormal lighting are marked. If local areas are found to have missing coverage or substandard quality, a re-flight procedure is promptly executed to ensure that the lining surface is completely covered.

[0040] After acquisition, all images are systematically screened, and all invalid frames are removed. Qualified images are then organized and archived according to tunnel mileage and spatial location, forming a structured raw dataset. This provides complete and reliable input data for subsequent image preprocessing, panoramic stitching, and crack identification tasks.

[0041] This phase utilizes a high-resolution, global shutter binocular camera as the core sensor, reducing motion blur while generating depth maps and providing stereo vision. The UAV integrates an RTK / UWB positioning module and an IMU unit to enhance positioning and attitude stability in environments with weak satellite signals. Before flight, the tunnel environment is surveyed to identify obstacles and areas with abnormal lighting. The binocular camera is calibrated, and the baseline and field of view configuration are adjusted to ensure stereo imaging quality and coverage integrity. (See schematic diagram below.) Figure 1 As shown.

[0042] Based on the principle of binocular vision, the left and right cameras are synchronously controlled to acquire images in RAW format. Automatic exposure bracketing (AEB) and a synchronous triggering strategy are set to avoid exposure differences between binocular images. By automatically adjusting ISO and shutter speed to adapt to changes in lighting within the tunnel, highly consistent image pairs suitable for stereo matching are obtained while ensuring image quality. The system maintains a constant and stable flight speed during flight to ensure that the image sequence has a high overlap rate and stable baseline constraints.

[0043] The binocular image acquisition stream is monitored in real time via ground station to check imaging sharpness, binocular synchronization, and stereo coverage. Online dense matching verification is performed on the acquired images to evaluate the depth map generation effect. If areas with severe occlusion, matching failure, or uneven lighting are found, the mission can be interrupted in real time and local re-flight can be performed to ensure the availability of stereo data and the needs of model training.

[0044] After acquisition, the binocular image pairs are filtered to remove invalid data that is blurry, overexposed, or has excessive synchronization deviation. High-quality image pairs are retained, organized according to tunnel mileage and structural regions, and depth maps and point cloud data are generated simultaneously. A structured dataset including the original images, calibration parameters, and depth information is constructed to provide a foundation for crack detection.

[0045] Step 3: Data preprocessing: Image filtering, optical preprocessing, stereo matching and cylindrical unfolding are performed on the original dataset to generate a two-dimensional unfolded diagram of the tunnel lining.

[0046] In view of the characteristics of tunnel lining crack identification, this step aims to synthesize the acquired binocular image sequence into a detailed and geometrically accurate tunnel lining unfolded map through a stable and reliable image processing workflow, so as to provide a high-quality image foundation for subsequent crack identification.

[0047] First, image screening and optical preprocessing are performed. Based on indicators such as image sharpness, illumination uniformity, and binocular synchronization, invalid frames that are blurry, overexposed, or have excessive synchronization deviations are automatically removed. Furthermore, lens distortion is eliminated to achieve row alignment of the left and right views, establishing an accurate geometric foundation for subsequent stereo matching and panoramic stitching.

[0048] In the image registration and stitching stage, a 3D reconstruction-based method is used to generate geometrically accurate unfolded images, rather than traditional 2D stitching. Considering the cylindrical geometry of the tunnel walls, cylindrical projection or perspective transformation models are used for image mapping, and a multi-band fusion algorithm is employed to achieve a natural transition at the joints, effectively eliminating ghosting and stitching marks.

[0049] Finally, the generated panoramic unfolded image undergoes quality verification, with a focus on checking the continuity and clarity of the crack areas to ensure there are no registration misalignments, blurred details, or artifacts. The output is a complete and accurate unfolded image of the tunnel lining surface, providing reliable input data for subsequent deep learning-based crack detection and quantitative analysis.

[0050] To obtain high-quality tunnel lining images, preprocessing includes the following steps: (1) Image filtering: The Laplacian variance algorithm is used to calculate the response variance of the Laplacian operator for each image. Images with variance below a set threshold are considered blurry frames and automatically discarded. Simultaneously, the average grayscale value of the images is calculated; images with excessively high (overexposed) or excessively low (underexposed) grayscale values ​​are discarded. The formula is as follows: ; in: This represents the pixel values ​​of the image after processing with the Laplacian operator. pixel value The arithmetic mean of the variance. Images below a set threshold are considered blurry and are removed.

[0051] ; in: This represents the grayscale value of pixel (i,j); M represents the number of pixels in the vertical direction of the image, i.e., the height of the image; N represents the number of pixels in the horizontal direction of the image, i.e., the width of the image. When If the exposure exceeds the reasonable range, it is considered an exposure anomaly.

[0052] (2) Optical pretreatment: Using the camera intrinsic parameter matrix and lens distortion coefficients obtained in advance through Zhang Zhengyou's calibration method, functions in the OpenCV library are called to correct each image and eliminate radial and tangential distortion.

[0053] (3) Image registration: For stereo-corrected image pairs, a semi-global matching algorithm is used to calculate disparity maps. Then, based on the disparity maps, camera intrinsics, and baseline distance, the disparity is converted into depth information using triangulation principles. Specifically: The conversion relationship between parallax and depth is determined by a core formula: ; Where: Z is the distance from the target to the camera; f is the pixel unit; B is the distance between the left and right cameras, with a typical baseline distance of 0.2-0.5 meters; d is the difference in position of the same target pixel in the left and right images.

[0054] The depth information of each pixel is then obtained, and a 3D point cloud can be generated through back projection. For each pixel (u, v) and its depth Z in the depth map, the formula for calculating its 3D spatial coordinates (x, y, z) is: ; ; ; in: These are the coordinates of the camera's principal point; It is the focal length in the x and y directions.

[0055] (4) Cylindrical development: An ideal cylindrical surface aligned with the tunnel's central axis is defined. Each 3D point on the reconstructed triangular mesh surface model is vertically projected onto this ideal cylindrical surface. Finally, the cylindrical surface is "cut" along a generatrix and unfolded onto a 2D plane to generate the final unfolded diagram of the tunnel lining. During this process, the texture (i.e., image color information) on the surface is also simultaneously mapped onto the unfolded diagram.

[0056] Step 4: Dataset creation: Label the cracks in the 2D unfolded image with bounding boxes, crop the unfolded image to the model input size, and divide it into training set, validation set and test set according to the preset ratio.

[0057] The dataset creation process comprises three core steps. First, crack targets in the tunnel lining images are standardized and labeled. By precisely defining the detection range of each independent crack instance, a high-quality labeling benchmark is established. Next, the original images undergo adaptive processing and enhancement, transforming them into a standardized size suitable for model training. This process effectively expands the sample size while meeting input requirements and enhancing the model's ability to perceive local features. Finally, a scientific partitioning strategy is employed to divide the processed dataset into three independent sets: training, validation, and testing. Ensuring the independence of content between different sets prevents overfitting, thereby guaranteeing the reliability and generalization ability of subsequent model training and evaluation. These three progressively build upon each other, constructing a high-quality data foundation for the crack recognition task.

[0058] Specifically, the dataset creation mainly includes three aspects: ① Crack sample bounding box annotation: A bounding box-based instance annotation method is adopted, using professional annotation tools to draw accurate bounding rectangles for each independent crack instance. During the annotation process, the bounding boxes are required to tightly enclose the crack body, and the continuously distributed cracks are reasonably divided into instances to ensure that each detection target has a clear boundary definition. All crack instances are uniformly labeled as "crack", and special attention is paid to distinguishing the differences between cracks and similar features such as lining joints and water stains to avoid mislabeling and omissions. ② Data cropping: The high-resolution original tunnel panoramic image is cropped into smaller images suitable for the input size of the target detection network (e.g., 640×640). This process not only adapts to the input requirements of model training, but more importantly, it serves as an effective data augmentation method. By performing gridded overlapping sampling on the original image, the number of training samples is significantly increased, and the model is forced to learn to identify crack features in the local field of view, improving its sensitivity to small-scale cracks and local features. ③ Dataset partitioning: The cropped image samples are divided into training, validation, and test sets according to a preset ratio (8:1:1). The partitioning process follows the "scene independence" principle, ensuring that sub-images from the same original panoramic image do not appear in different sets. This effectively prevents overfitting caused by the model memorizing specific background features, guaranteeing the accuracy of performance evaluation and generalization ability. After partitioning, the training set will be used for subsequent Auto Augment strategy search and model training, the validation set will be used for hyperparameter tuning and early stopping, and the test set will serve as the benchmark for evaluating the final model performance.

[0059] Step 5: Improved YOLOv10 Model Construction and Training: An improved YOLOv10 model is used for crack detection training, with an initial learning rate of 0.001, a batch size of 16, and 200 iterations. Improvements to the YOLOv10 model include: replacing the multi-head self-attention module in the original pyramid attention network with an efficient multi-scale attention module; replacing Mosaic augmentation with AutoAugment automated data augmentation; and replacing binary cross-entropy loss with focusing loss.

[0060] Over the past few years, the YOLO family of algorithms has become the most advanced object detection framework due to its effective balance between computational cost and detection performance. Its core idea is to transform the detection task into an end-to-end regression problem. Researchers have explored YOLO's architecture design, optimization objectives, and data augmentation strategies, achieving significant progress.

[0061] This invention is based on the latest architecture of YOLOv10 and proposes three key improvements: First, EMA replaces MHSA in the original PSA. This improvement aims to address the issue that the computational complexity of MHSA increases quadratically with the sequence length. EMA uses parallel multi-scale (1x1 and 3x3 convolutional branches) convolutional paths. The 1x1 branch handles the encoding of two spatial directions, while the 3x3 branch captures multi-scale spatial structure information. This significantly reduces computational overhead while maintaining the ability to fuse multi-scale features, enabling the model to capture local details and global context information more efficiently.

[0062] Secondly, an Auto Augment strategy is introduced to replace the preset Mosaic augmentation. Cracks typically appear as long, continuous lines or mesh-like structures; they may be long, but their pixel width is narrow. Mosaic augmentation, through scaling, can make already narrow cracks extremely blurry or even broken, thus increasing the difficulty of recognition. Auto Augment, through data-driven strategy search, automatically generates the most suitable augmentation combination for the target dataset, effectively improving the diversity of training samples and the model's generalization ability.

[0063] Finally, the binary cross-entropy loss is replaced with a focal loss to address the class imbalance problem in object detection.

[0064] In its implementation, the improved model centers on an EMA-enhanced C2f module. It reduces information loss through spatial channel decoupling downsampling and combines Auto Augment data augmentation with Focal Loss, forming a systematic improvement scheme encompassing network structure, training strategies, and optimization objectives. These three innovations complement each other, jointly enhancing the model's detection accuracy and robustness in complex scenarios.

[0065] YOLOv10's model structure follows the classic "backbone-neck-head" design paradigm. Its backbone network extracts features by stacking multiple C2f modules, with the core C2f module represented as follows: ; Where: X is the input feature map; This is the i-th bottleneck block; This is a channel-dimensional concatenation operation. Each block consists of depthwise separable convolutions to improve efficiency. The feature pyramid neck network uses a PAN structure, through... Top-down path and A bottom-up approach enables multi-scale feature fusion. The final detection head employs a dual-branch design, sharing regression parameters but independently learning a classifier; its output can be formalized as: ; Each branch contains classification predictions. and coordinate regression This structure maintains training stability while enabling NMS-free inference.

[0066] The model used in this invention is an improved YOLOv10 model, and a schematic diagram of the improved YOLOv10 model is shown below. Figure 3 As shown, based on the original model, the main change is that the EMA module replaces the MHSA in the original PSA module. Its main design goal is to enhance the feature representation capability while retaining the complete information of each channel and reducing the computational overhead.

[0067] Specifically, the EMA module first takes the input feature map It is divided into G sub-features along the channel dimension, that is Each of them This grouping operation ensures that spatial semantic features are evenly distributed within each feature group. Subsequently, the module reshapes some channels to the batch dimension, which helps avoid channel dimensionality reduction in subsequent convolutional operations, thus preserving more information. In addition, a parallel sub-network design is employed, feeding the grouped features into two parallel branches for processing: ① A 1x1 convolutional branch: This branch primarily utilizes 1x1 convolutions, supplemented by two one-dimensional global average pooling operations, aggregating spatial information along the height and width directions, respectively. This branch ensures cross-channel interaction and, by avoiding channel dimensionality reduction, can more effectively recalibrate channel weights. ② A 3x3 convolutional branch: This branch captures multi-scale spatial feature information through a 3x3 convolutional kernel, expanding the receptive field and helping the model capture more complex spatial structures. The output features of the two parallel branches are aggregated through cross-dimensional interaction. Specifically, the module uses 2D global average pooling to encode global spatial information in the output of the 1x1 branch and generates a spatial attention map through matrix dot product operations. This mechanism captures pixel-level pairwise relationships and generates better pixel-level attention weights for high-level feature maps. Ultimately, the weighted features highlight the global contextual information of all pixels, enhancing the model's ability to model long-range dependencies. The detailed EMA model structure is as follows: Figure 4 As shown.

[0068] With the development of training strategies, especially in the later stages of YOLOv5 / v6 / v7 and YOLOv8, many have found that Mosaic can lead to domain shift problems (excessive discrepancies between the training and real data distributions) towards the end of training. Therefore, a common practice is to disable Mosaic in the last N epochs of training, or replace it with a more refined automated augmentation strategy. This invention uses an Auto Augment data augmentation strategy instead of Mosaic, employing a series of predefined, optimal augmentation operations, including various strategies such as rotation, translation, and color enhancement. It is important to note that none of these augmentations are used during model validation, as this more accurately reflects the model's performance on real data.

[0069] In addition to changes to the model, this invention also improves the evaluation metric by replacing the binary cross-entropy loss with Focal Loss. The standard binary cross-entropy loss function measures the difference between the model's predicted probability and the true label. For a single sample, its formula can be expressed as: ; in, is the true label (e.g., crack is 1, background is 0), and p is the probability that the model predicts the target to be a positive example (crack).

[0070] However, in tunnel lining crack detection, the vast majority of the image area is background, with crack pixels accounting for a very small percentage. This leads to negative samples (background) dominating the loss contribution. The model optimization process is thus dominated by a large number of "simple" negative samples, making it difficult to effectively learn how to identify challenging "hard" cracks. Focused loss, by automatically reducing the weight of easily classified samples through an adjustment factor, allows the model training to focus more on learning hard examples, thereby improving the detection performance for difficult targets. Specifically, a dynamic modulation factor is introduced to reduce the weight of easily classified samples, thus making the model training more focused on hard examples. The formula is as follows: ; in, and The modulation factor is: when a sample is misclassified or difficult to classify, the modulation factor is close to 1, and the loss is basically unaffected; when a sample is correctly classified and the confidence level is high (for positive samples, p is close to 1), the modulation factor approaches 0, and the contribution of the sample to the total loss is significantly reduced. It is an adjustable focusing parameter used to control the intensity of the modulation factor; This is a balancing parameter used to adjust the importance weights of positive and negative samples. In this invention, the balancing parameter and the focusing parameter are 0.25 and 2.0, respectively.

[0071] Step 6: Automatic crack identification and result output: Input the two-dimensional unfolded diagram of the tunnel lining into the trained improved YOLOv10 model, and automatically output the location information of the cracks for tunnel maintenance and risk assessment.

[0072] In the automatic crack identification and output stage, the complete two-dimensional unfolded diagram of the tunnel lining is input into the trained improved YOLOv10 model. The model automatically performs inference and outputs the precise location information of the cracks (usually represented in bounding box coordinates, along with crack category and confidence score). This output is the final deliverable of the entire automated process and can directly serve tunnel maintenance and risk assessment: maintenance personnel can locate specific crack areas for review and treatment based on the identification results, while management departments can conduct structural safety assessments and maintenance decisions based on the distribution, quantity, and size information of the cracks. This upgrades traditional general inspections to targeted fine-tuning, significantly improving the efficiency, accuracy, and scientific nature of tunnel maintenance work.

[0073] Example 2: The present invention also provides a computer-readable storage medium storing computer instructions for causing the computer to execute the above-described method for automatic identification of cracks in the lining of mountain tunnels using a combined UAV inspection and target detection model.

[0074] Example 3: The present invention also provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the above-described method for automatic identification of cracks in the lining of mountain tunnels using a combined UAV inspection and target detection model.

[0075] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for automatic identification of cracks in the lining of mountain tunnels using a combined UAV inspection and target detection model, characterized in that, Includes the following steps: Step 1: Preliminary data acquisition and flight planning: Obtain the cross-sectional dimensions, shape, and camera sensor parameters of the mountain tunnel, and calculate and determine the UAV's flight altitude, flight speed, and image heading overlap rate; Step 2: UAV Inspection and Image Acquisition: Control the UAV to fly at a constant speed along the tunnel axis, acquire lining image sequences through the onboard binocular camera, monitor image quality in real time, mark invalid frames for re-flying, and construct a structured original dataset; Step 3: Data preprocessing: Image filtering, optical preprocessing, stereo matching and cylindrical unfolding are performed on the original dataset to generate a two-dimensional unfolded diagram of the tunnel lining; Step 4: Dataset creation: Label the cracks in the 2D unfolded image with bounding boxes, crop the unfolded image to the model input size, and divide it into training set, validation set and test set according to the preset ratio; Step 5: Improved YOLOv10 Model Construction and Training: The improved YOLOv10 model is used for crack detection training. The improvements to the YOLOv10 model include: replacing the multi-head self-attention module in the original pyramid attention network with an efficient multi-scale attention module; replacing Mosaic augmentation with the AutoAugment automated data augmentation strategy; and replacing the binary cross-entropy loss with a focusing loss. Step 6: Automatic crack identification and result output: Input the two-dimensional unfolded diagram of the tunnel lining into the trained improved YOLOv10 model, and automatically output the location information of the cracks for tunnel maintenance and risk assessment.

2. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 1, the image heading overlap rate is ≥80%, and the flight parameters are determined by the geometric imaging model to ensure that the ground sampling distance of the image can completely cover the tunnel arch, sidewalls and inverted arch area.

3. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 2, the UAV is equipped with a binocular camera, an RTK / UWB positioning module, and an IMU unit to acquire RAW format images.

4. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 2, the UAV inspection adopts an automatic exposure bracketing and synchronous triggering strategy to control the binocular camera to acquire images synchronously. During the flight, the image clarity, binocular synchronization and stereo coverage are monitored in real time by the ground station, and invalid frames that are blurry, overexposed or have large synchronization deviations are removed.

5. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 3, the image screening uses the Laplacian variance algorithm and the average gray value to determine: blurry frames are eliminated by calculating the response variance of the Laplacian operator of the image, and if the response variance is lower than the set threshold, it is determined to be blurry; overexposed or underexposed frames are eliminated by calculating the average gray value of the image, and if the gray value exceeds the set range, it is determined to be an exposure abnormality.

6. The method for automatic identification of cracks in mountain tunnel lining based on a combined UAV inspection and target detection model according to claim 5, characterized in that, The formula for calculating the Laplace variance algorithm is as follows: ; in, For response variance; x i The pixel values ​​of the image after processing with the Laplacian operator; μ This is the arithmetic mean of the pixel values; n This represents the total number of pixels.

7. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 5, characterized in that, The formula for calculating the average gray value is: ; Where I(i,j) is the gray value of pixel (i,j); M represents the number of pixels in the vertical direction of the image, i.e., the height of the image; and N represents the number of pixels in the horizontal direction of the image, i.e., the width of the image.

8. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 3, the optical preprocessing adopts the Zhang Zhengyou calibration method, which uses the preset camera intrinsic parameter matrix and lens distortion coefficient to eliminate radial and tangential distortion of the image.

9. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 3, the specific process of stereo matching and cylindrical unfolding is as follows: a semi-global matching algorithm is used to calculate the disparity map of the corrected binocular image pair. Based on the disparity map, camera intrinsic parameters and baseline distance, the disparity is converted into depth information through the principle of triangulation. Each pixel in the depth map is back-projected into three-dimensional space to generate a three-dimensional point cloud. An ideal cylindrical surface aligned with the central axis of the tunnel is defined. The three-dimensional point cloud is projected onto the ideal cylindrical surface, cut along the generatrix and unfolded into a two-dimensional tunnel lining unfolded diagram, synchronously mapping image texture information.

10. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 4, crack annotation uses an outer rectangle to define each independent crack instance. The bounding box tightly wraps the crack body and distinguishes the crack from other similar features. Image cropping uses an overlay sampling method to crop the unfolded image to a standard input size of 640×640 pixels.

11. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 4, the dataset is divided into training set, validation set and test set in a ratio of 8:1:

1. The division follows the principle of scene independence, and subgraphs of the same original unfolded graph are not distributed across sets.

12. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 5, the specific implementation of the efficient multi-scale attention module (EMA) is as follows: the input feature map is divided into G sub-features in the channel dimension, and processed in parallel through a 1x1 convolutional branch and a 3x3 convolutional branch. The 1x1 convolutional branch combines one-dimensional global average pooling to aggregate spatial information, and the 3x3 convolutional branch expands the receptive field to capture multi-scale structures. The global spatial information output by the 1x1 branch is encoded by 2D global average pooling to generate a spatial attention map, and the features of the two branches are weighted and fused before being output.

13. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 5, the AutoAugment automated data augmentation strategy automatically generates the optimal augmentation operation sequence, which includes rotation, translation, and color enhancement, through data-driven strategy search. This sequence is applied to the training set during training, but no data augmentation is performed on the validation and test sets.

14. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 5, the mathematical formula for the focusing loss (FocalLoss) is: ; Where y∈{0,1} is the true label, crack is 1, background is 0; p is the probability of the model predicting a positive example, α is the positive and negative sample balance parameter, and γ is the focusing parameter used to control the intensity of the modulation factor.

15. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 5, the improved YOLOv10 model adopts a backbone-neck-head architecture: the backbone network is composed of multiple stacked C2f modules, and the expression of the C2f module is: ; Where: X is the input feature map; This is the i-th bottleneck block; This is a channel-level splicing operation; Each block is a depthwise separable convolution; the neck is a PAN-structured feature pyramid that achieves multi-scale feature fusion through top-down and bottom-up paths; the head is a dual-branch design that shares regression parameters and learns classifiers independently, enabling NMS-free inference.

16. The method for automatic identification of cracks in mountain tunnel lining using a combined UAV inspection and target detection model according to claim 1, characterized in that, In step 6, the crack location information is output in the form of bounding box coordinates, crack category, and confidence level.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute a method for automatic identification of cracks in the lining of mountain tunnels according to any one of claims 1 to 16, which is based on a combined UAV inspection and target detection model.

18. An electronic device, characterized in that, include: The system includes a memory and a processor, which are interconnected. The memory stores computer instructions, and the processor executes these computer instructions to perform an automatic identification method for cracks in the lining of mountain tunnels using a combined UAV inspection and target detection model, as described in any one of claims 1 to 16.