Tunnel lining crack intelligent detection method based on improved instance segmentation algorithm

By improving the YOLOv8 instance segmentation algorithm, combined with the C2f_EMSC module, CBAM attention mechanism and EIoU loss function, the problem of high-precision pixel-level crack detection in tunnel environment is solved, and lightweight and efficient tunnel lining crack detection is achieved.

CN120388008APending Publication Date: 2025-07-29SOUTHEAST UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510528077.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high-precision, pixel-level crack detection in a tunnel environment, and the existing models are usually large in size and have large calculations, which are not suitable for lightweight equipment deployment.

Method used

Using the improved YOLOv8 instance segmentation algorithm, some modules of the backbone and neck network are replaced by the lighter C2f_EMSC module, and the CBAM attention mechanism is inserted after each C2f_EMSC module, and the CIoU loss function is replaced by the EIoU loss function, optimizing the model structure and loss function.

Benefits of technology

It realizes pixel-level high-precision detection of tunnel lining cracks, reduces the amount of model parameters and calculations, is suitable for lightweight equipment deployment, and improves detection accuracy and ability to adapt to complex tunnel environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388008A_ABST
    Figure CN120388008A_ABST
Patent Text Reader

Abstract

The invention discloses a tunnel lining crack intelligent detection method based on an improved instance segmentation algorithm. The method comprises the following steps: firstly, obtaining a tunnel lining picture; the method comprises the following steps: selecting a tunnel lining picture containing cracks, performing data enhancement, performing pixel-level labeling on the cracks by using a Label data labeling method, and establishing a crack data set; a part of C2f modules in a backbone network and a neck network of the YOLOv8 algorithm are replaced with lighter C2fEMSC modules, a CBAM attention mechanism is inserted behind each C2fEMSC module, an original CIoU loss function of the YOLOv8 algorithm is replaced with an EIoU loss function, and improvement of the YOLOv8 instance segmentation algorithm is completed. Training based on the crack data set and obtaining a tunnel lining crack pixel-level detection model; and inputting a tunnel lining picture, detecting the tunnel lining picture by using the detection model, and outputting a tunnel lining crack detection result. Compared with the prior art, the method has the advantages that pixel-level and high-precision detection of tunnel lining cracks is realized by using a lightweight model, and the method can be deployed on lightweight equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of tunnel engineering, and in particular relates to an intelligent detection method for tunnel lining cracks based on an improved instance segmentation algorithm. Background Art

[0002] Transportation tunnel construction has become a crucial component of infrastructure development. Currently, tunnels are transitioning from a construction-focused approach to a phase that prioritizes both construction and maintenance. In the future, they will face the dual pressures of rapidly increasing operating mileage and increasing tunnel performance degradation. Tunnel linings are constantly impacted by factors such as the geological environment and transportation loads, making them susceptible to structural defects such as cracks. These defects can affect the tunnel's structural safety and service life. Therefore, crack detection in tunnel linings is crucial for the safe operation and sustainable development of tunnels.

[0003] Initially, the status of highway tunnels in my country was primarily determined by manual observation. However, this method suffers from low detection efficiency, strong subjectivity, and potential safety hazards. In recent years, while image processing and machine learning algorithms have been increasingly applied to tunnel lining crack detection, these methods require high image quality and are generally only suitable for environments with minimal interference around the cracks, making them difficult to apply to the complex environments of operational highway tunnels.

[0004] In recent years, a large number of scholars have continuously improved the accuracy, efficiency and generalization ability of crack detection models based on deep learning algorithms. However, there are still some shortcomings, which are specifically manifested in the following aspects:

[0005] 1. Most methods only achieve the target identification of cracks, and the identification results cannot meet the subsequent crack quantitative analysis requirements.

[0006] 2. Most existing methods are aimed at ground structures such as concrete, road pavements, and bridges. However, tunnels have characteristics such as poor lighting conditions and many environmental interferences. It is difficult to directly apply crack detection models for ground structures, and further research is needed.

[0007] 3. The existing crack detection models have room for improvement in both detection accuracy and efficiency. Models with higher detection accuracy are usually larger and require more storage space, which is not conducive to deployment on lightweight equipment.

[0008] Therefore, it is necessary to study a new pixel-level intelligent detection method for lightweight lining cracks suitable for tunnel environments. Summary of the Invention

[0009] To solve the above problems, the present invention discloses an intelligent detection method for tunnel lining cracks based on an improved instance segmentation algorithm. The improved YOLOv8 instance segmentation algorithm is used to detect tunnel lining images, realizing pixel-level, high-precision detection of tunnel lining cracks, and can be deployed on lightweight equipment.

[0010] To achieve the above object, the technical solution of the present invention is as follows:

[0011] An intelligent detection method for tunnel lining cracks based on an improved instance segmentation algorithm, comprising the following steps:

[0012] Obtain tunnel lining pictures through tunnel structure detection equipment;

[0013] Input the tunnel lining pictures, and use the crack pixel-level detection model to detect the tunnel lining pictures, and output the tunnel lining crack detection results;

[0014] Among them, the training process of the crack pixel-level detection model includes:

[0015] Select the tunnel lining pictures containing cracks obtained by the tunnel structure detection equipment and perform data augmentation, use the Labelme data annotation method to perform pixel-level annotation on the cracks therein, and establish a crack data set;

[0016] Based on the crack data set, use the improved YOLOv8 instance segmentation algorithm to train and obtain the tunnel lining crack pixel-level detection model.

[0017] Furthermore, the establishment process of the improved YOLOv8 algorithm includes:

[0018] Modify the backbone network for extracting crack image features and the neck network for feature fusion in the YOLOv8 instance segmentation algorithm, and replace the 3rd and 4th C2f modules in the backbone network and the 1st, 3rd, and 4th C2f modules in the neck network with the more lightweight C2f_EMSC module, so as to greatly reduce the number of parameters and computational complexity of the model without changing the detection accuracy of the model, and achieve the lightweight of the model;

[0019] Insert a CBAM attention mechanism behind each C2f_EMSC module in the backbone network of the YOLOv8 algorithm, with a total of two insertions. The attention mechanism can enable the detection model to pay more attention to the important crack image features extracted by the backbone network, thereby improving the detection accuracy of the detection model;

[0020] Replace the original CIoU loss function of the YOLOv8 algorithm with the EIoU loss function to improve the regression accuracy of the model, thereby improving the detection accuracy of the model. Calculate the model bounding box regression loss through the EIoU loss function, which not only considers the overlapping area between the predicted box and the ground truth box, but also calculates and compares the length difference and width difference between the predicted box and the ground truth box respectively, making the loss function more sensitive to the accuracy of the position and size of the predicted box. In addition, EIoU combines the Focal Loss strategy to further optimize the sample imbalance problem in the model bounding box regression.

[0021] Further, the C2f_EMSC module first uses a 1×1 convolution to change the number of channels of the input feature map sequence. Then, it uses the Split method to evenly divide the feature map sequence into two groups along the channel dimension. The first part of the feature map sequence undergoes no operation and directly serves as the input feature map sequence of the Concat layer. The other part of the feature map sequence successively passes through n (n is 2 for the first C2f_EMSC module of the backbone network, and n is 1 for the remaining 4 C2f_EMSC modules) Bottleneck_EMSC layers. Each Bottleneck_EMSC layer contains a 3×3 convolution and a more lightweight EMSConv convolution. Every time the feature map sequence passes through 1 Bottleneck_EMSC layer, the output feature map sequence is divided into two groups. One group serves as the input feature map sequence of the Concat layer, and the other group continues to enter the next Bottleneck_EMSC layer. Finally, the Concat layer combines and connects all the input feature map sequences and outputs them to the last 1×1 convolution to adjust the number of channels of the feature map sequence for subsequent operations.

[0022] Further, the EMSConv convolution first performs a Split operation on the input feature map sequence, evenly dividing it into two groups. One group serves as the input feature map sequence of the Concat layer, and the other group is further evenly divided into two subgroups. One subgroup performs a convolution operation with a 3×3 convolution kernel to extract image features, and the other subgroup performs a convolution operation with a 5×5 convolution kernel to extract image features. Then, the Concat method is used to combine and connect the feature map sequences output by the three branches to obtain a complete feature map sequence. Finally, a 1×1 convolution is used to exchange channel information to complete image feature extraction and adjustment of the number of channels of the feature map sequence.

[0023] Further, the CBAM module is composed of two parts in series: a channel attention module (Channel Attention Module, CAM) and a spatial attention module (Spatial Attention Module, SAM). Among them, the channel attention module enhances the channel attention of the input feature map sequence and improves the information interaction between different channels; the spatial attention module enhances the spatial attention of the input feature map sequence and improves the information correlation between different positions of the same feature map.

[0024] Further, the channel attention module mainly includes the following running steps:

[0025] Feature Compression: In the spatial dimensions H×W, global average pooling and global max pooling operations are simultaneously performed on the input feature map sequence X (with a shape of C×H×W, where C represents the number of input image channels, H represents the height of the input image, and W represents the width of the input image), outputting two groups of feature map sequences with a size of C×1×1, that is, keeping the number of input channels unchanged, and each channel is compressed into a single value, representing the maximum value and the average value of that channel respectively;

[0026] Feature Transformation: The two groups of feature map sequences obtained from the above pooling are input into a shared multi-layer perceptron (MLP), and the pooled feature map sequences are fused and transformed to extract the important features of each channel, generating an attention feature map;

[0027] Feature Fusion and Activation: The two groups of feature map sequences processed by the MLP are added together, and the attention weights of each channel are obtained through the Sigmoid activation function, reflecting the importance of each channel for the current task, and finally the channel attention weights are mapped to each channel of the input feature map sequence.

[0028] Furthermore, the multi-layer perceptron (MLP) usually consists of 2 fully connected layers. The first fully connected layer is used for channel dimensionality reduction to reduce the number of parameters during feature map fusion transformation, and the second fully connected layer is used for channel dimensionality increase to restore to the original number of channels for subsequent operations. The ReLU activation function is used between the two fully connected layers.

[0029] Furthermore, the spatial attention module mainly includes the following running steps:

[0030] Feature Compression: Average pooling and max pooling operations are performed on the feature map sequence X' processed by the channel attention module in the channel dimension C, outputting two feature maps with a size of 1×H×W, that is, each pixel position of the input feature map sequence is compressed into a single value, representing the maximum value and the average value of that position respectively;

[0031] Feature Concatenation and Convolution: The above two pooled feature maps are concatenated in the channel dimension to obtain a group of feature map sequences with a size of 2×H×W. Then, a convolutional layer with a kernel size of 7×7 is used to reduce the number of channels to 1 to capture local and global spatial information;

[0032] Feature Activation: The feature map output by the convolutional layer is input into the Sigmoid activation function to obtain the attention weights of each spatial position, reflecting the importance of each pixel position for the current task, and finally the spatial attention weights are mapped to each channel of the input feature map sequence.

[0033] Furthermore, the calculation process of the EIoU loss function is as follows:

[0034] Calculate IoU, and the formula is as follows:

[0035]

[0036] Among them, the intersection area is the area of the overlapping part of the predicted box and the ground truth box, and the union area is the sum of all regions covered by the predicted box and the ground truth box;

[0037] Calculate the distance ρ(b, b gt ) between the centers of the predicted box and the ground truth box, and the formula is as follows:

[0038]

[0039] Among them, the center coordinates of the predicted box are (x, y), and the center coordinates of the ground truth box are (x gt , y gt );

[0040] Calculate the difference ρ(w, w gt ) in width and the difference ρ(h, h gt ) in height between the predicted box and the ground truth box;

[0041] Calculate the width c w and height c h of the smallest bounding box that can completely cover the two bounding boxes. The calculation process is as follows:

[0042] Calculate the maximum coverage range of the predicted box and the ground truth box in the x-axis direction:

[0043] Left boundary:

[0044] Right boundary:

[0045] Then the width c w of the smallest bounding box = x max - x min ;

[0046] Calculate the maximum coverage range of the predicted box and the ground truth box in the y-axis direction:

[0047] Upper boundary:

[0048] Lower boundary:

[0049] Then the height c h of the smallest bounding box = y max - y min ;

[0050] Substitute the above results into the EIoU loss calculation formula to obtain the EIoU loss value, and the formula is as follows:

[0051]

[0052] The present invention has the following beneficial effects:

[0053] 1. Aiming at the problems that most current crack detection methods only achieve the target recognition of cracks, the recognition results cannot meet the requirements of subsequent crack quantification analysis, and the crack detection model for ground structures is difficult to be directly applied to the tunnel environment, etc., the present invention is based on an improved YOLOv8 instance segmentation algorithm, uses tunnel lining crack pictures for model training, establishes a pixel-level detection model for tunnel lining cracks, and realizes the pixel-level detection of tunnel lining cracks.

[0054] 2. Aiming at the problem that models with higher detection accuracy are usually larger, the present invention proposes a more lightweight EMSConv convolution, and based on this convolution, a new C2f_EMSC module is proposed to optimize the C2f module of the YOLOv8 instance segmentation algorithm. While maintaining the detection accuracy of the model, the number of parameters and the amount of calculation of the model are greatly reduced, realizing model lightweighting.

[0055] 3. Aiming at the problem that there is still room for improvement in the detection accuracy of existing crack detection models, the present invention introduces the CBAM attention mechanism and replaces the original CIoU loss function of the YOLOv8 instance segmentation algorithm with the EIoU loss function. Among them, the CBAM attention mechanism enables the detection model to pay more attention to the important crack image features extracted by the backbone network, and the EIoU loss function can improve the accuracy of model regression, thereby improving the detection accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is the overall flow schematic diagram of the present invention;

[0057] Figure 2 is the structural diagram of the improved YOLOv8 instance segmentation algorithm provided by the present invention;

[0058] Figure 3 is the structural decomposition diagram of the C2f_EMSC module provided by the present invention;

[0059] Figure 4 is the structural schematic diagram of the CBAM attention mechanism provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0061] Example 1:

[0062] Reference Figure 1As shown in the figure, this embodiment provides an intelligent detection method for tunnel lining cracks based on an improved instance segmentation algorithm, including the following steps:

[0063] Step 1: Obtain tunnel lining pictures;

[0064] In Step 1, in this embodiment, 24,360 tunnel lining scanned pictures were obtained at the site of a highway tunnel project using a tunnel structure detection device.

[0065] Step 2: Establish a crack dataset for training a pixel-level detection model for tunnel lining cracks;

[0066] Not all of the tunnel lining pictures obtained in Step 1 contain crack targets. Therefore, through manual screening, tunnel lining pictures containing cracks are selected to form a tunnel lining crack picture set, totaling 1,200 pictures;

[0067] The tunnel lining crack picture set contains crack pictures under different lighting conditions, crack angles, crack widths, crack positions, and complex environments. The complex environments include stains (linear spider webs), manual markings, lamp post obstructions, expansion joints, and lighting fixtures, etc. A tunnel lining crack picture set with high richness can provide more diverse crack image features for the model during model training, improving the detection accuracy and generalization ability of the pixel-level crack detection model;

[0068] Since the 1,200 tunnel lining crack pictures collected still cannot well meet the requirement of picture richness for model training, based on the Transforms data augmentation algorithm, the present invention designs a program to randomly perform operations such as rotation, cropping, scaling, deformation, etc. on the 1,200 crack pictures and adjust parameters such as brightness, saturation, contrast, and blur. Among them, the geometric transformation of the pictures can improve the richness and quantity of the crack picture set, the brightness adjustment of the pictures can simulate the low-brightness environment inside the tunnel, and the adjustments of saturation, contrast, blur, etc. of the pictures can simulate the adverse conditions during the collection of tunnel lining pictures;

[0069] The picture set expanded according to the above method contains a total of 6,000 pictures. Then, the Labelme data annotation method is used to perform pixel-level annotation on the cracks in the pictures one by one. The crack areas in the pictures are marked out by geometric polygons to form annotation frames and annotation masks, and the "Crack" label is attached to establish a tunnel lining crack dataset;

[0070] In the tunnel lining crack dataset, the crack pictures are divided into a training set, a validation set, and a test set according to the ratio of 7:2:1. The training set contains a total of 4,200 pictures, the validation set contains a total of 1,200 pictures, and the test set contains a total of 600 pictures, completing the establishment of the crack dataset for training a pixel-level detection model for tunnel lining cracks.

[0071] Step 3: Construct an improved YOLOv8 instance segmentation algorithm;

[0072] Refer to Figure 2 As shown, in this embodiment, the present invention uses a more lightweight C2f_EMSC module to replace some C2f modules in the backbone network and neck network of the YOLOv8 instance segmentation algorithm, inserts a CBAM attention mechanism behind each C2f_EMSC module, and uses the EIoU loss function to replace the original CIoU loss function of the YOLOv8 instance segmentation algorithm. The specific construction process is as follows:

[0073] Step1: Use a more lightweight C2f_EMSC module to replace some C2f modules in the backbone network and neck network of the YOLOv8 instance segmentation algorithm;

[0074] In the YOLOv8 instance segmentation algorithm, the backbone network is responsible for extracting the image features of the tunnel lining cracks, while the neck network is responsible for fusing the image features of different depths extracted by the backbone network to improve the detection accuracy of the tunnel lining crack pixel-level detection model for cracks of different sizes;

[0075] The backbone network of the YOLOv8 instance segmentation algorithm is composed of 5 convolutional modules, 4 C2f modules, and 1 SPPF pooling module connected in series. The first convolutional module is responsible for receiving the input image, and the feature map sequence after SPPF pooling will be transmitted to the neck network. Although the YOLOv8 instance segmentation algorithm already has very strong image feature extraction capabilities, the parameter quantity and computational complexity of this algorithm are relatively large. Therefore, the crack detection model established based on this algorithm not only consumes more computational resources during model prediction, but also has a relatively large model size, requiring more storage space, which is not conducive to deployment on lightweight devices.

[0076] Therefore, refer to Figure 3 As shown, in this embodiment, the present invention first proposes a more lightweight EMSConv convolution. The EMSConv convolution first performs a Split operation on the input feature map sequence, dividing it into 2 groups on average. One group serves as the input feature map sequence of the Concat layer, and the other group is further divided into 2 small groups on average. One small group performs a convolution operation with a convolution kernel size of 3×3 to extract image features, and the other small group performs a convolution operation with a convolution kernel size of 5×5 to extract image features. Then, the Concat method is used to combine and connect the feature map sequences output by the 3 branches to obtain a complete feature map sequence. Finally, a 1×1 convolution is used to exchange channel information to complete image feature extraction and adjustment of the channel number of the feature map sequence;

[0077] Furthermore, refer to Figure 3As shown, in this embodiment, the present invention constructs a C2f_EMSC module by referring to the C2f module in the YOLOv8 algorithm. The C2f_EMSC module first uses a 1×1 convolution to change the number of channels of the input feature map sequence, and then uses the Split method to evenly divide the feature map sequence into 2 groups along the channel dimension. The first part of the feature map sequence does not perform any

[0078] operations and directly serves as the input feature map sequence of the Concat layer; the other part of the feature map sequence sequentially passes through n (n is 2 for the first C2f_EMSC module of the backbone network, and n is 1 for the remaining 4 C2f_EMSC modules) Bottleneck_EMSC layers. Each Bottleneck_EMSC layer contains a 3×3 convolution and a more lightweight EMSConv convolution. Every time the feature map sequence passes through 1 Bottleneck_EMSC layer, the output feature map sequence is divided into 2 groups. One group serves as the input feature map sequence of the Concat layer, and the other group continues to enter the next Bottleneck_EMSC layer. Finally, the Concat layer combines and connects all the input feature map sequences and outputs them to the last 1×1 convolution to adjust the number of channels of the feature map sequence for subsequent operations;

[0079] Finally, referring to Figure 3 As shown, in this embodiment, the present invention uses the C2f_EMSC module to replace the 3rd and 4th C2f modules in the backbone network and the 1st, 3rd, and 4th C2f modules in the neck network, which can greatly reduce the number of parameters and computational amount of the model without changing the detection accuracy of the model, thus realizing the lightweight of the model.

[0080] Step2: Insert the CBAM attention mechanism behind each C2f_EMSC module;

[0081] Currently, most crack detection methods only achieve the target recognition of cracks, and the recognition results cannot meet the requirements of subsequent crack quantification analysis. Therefore, the present invention provides a pixel-level intelligent crack detection method. Pixel-level crack detection methods require the detection model to accurately identify each pixel point of the crack, which has a high requirement for the accuracy of the detection model. At the same time, the tunnel environment usually has lower brightness and more environmental interferences, and most of the current crack detection models for ground structures are not applicable to the more complex tunnel environment;

[0082] The attention mechanism can enable the detection model to pay more attention to the important crack image features extracted by the backbone network, thereby improving the detection ability of the detection model for each crack pixel point and reducing the influence of tunnel environment interferences at the same time;

[0083] Therefore, in this embodiment, the present invention proposes a CBAM attention mechanism to improve the detection performance of the detection model. The CBAM module is composed of two parts in series: a channel attention module (Channel Attention Module, CAM) and a spatial attention module (Spatial Attention Module, SAM). Among them, the channel attention module enhances the channel attention of the input feature map sequence and improves the information interaction between different channels; the spatial attention module enhances the spatial attention of the input feature map sequence and improves the information association between different positions of the same feature map;

[0084] Refer to Figure 4 As shown, the channel attention module mainly includes the following running steps:

[0085] Feature compression: On the spatial dimension H×W, global average pooling and global max pooling operations are simultaneously performed on the input feature map sequence X (with a shape of C×H×W, where C represents the number of input image channels, H represents the height of the input image, and W represents the width of the input image), and two groups of feature map sequences with a size of C×1×1 are output, that is, while keeping the number of input channels unchanged, each channel is compressed into a single value, representing the maximum value and the average value of that channel respectively;

[0086] Feature transformation: The two groups of feature map sequences obtained by the above pooling are input into a shared multi-layer perceptron (MLP), and the pooled feature map sequences are fused and transformed to extract the important features of each channel, generating an attention feature map;

[0087] Feature fusion and activation: The two groups of feature map sequences processed by the MLP are added together, and the attention weight of each channel is obtained through the Sigmoid activation function, reflecting the importance of each channel for the current task, and finally mapping the channel attention weight to each channel of the input feature map sequence;

[0088] Refer to Figure 4 As shown, the spatial attention module mainly includes the following running steps:

[0089] Feature compression: Average pooling and max pooling operations are performed on the feature map sequence X' processed by the channel attention module in the channel dimension C, and two feature maps with a size of 1×H×W are output, that is, each pixel position of the input feature map sequence is compressed into a single value, representing the maximum value and the average value of that position respectively;

[0090] Feature concatenation and convolution: The above two pooled feature maps are concatenated in the channel dimension to obtain a group of feature map sequences with a size of 2×H×W. Then, a convolutional layer with a convolution kernel size of 7×7 is used to reduce the number of channels to 1 to capture local and global spatial information;

[0091] Feature activation: The feature map output by the convolutional layer is input into the Sigmoid activation function to obtain the attention weights at each spatial position, reflecting the importance of each pixel position for the current task, and finally mapping the spatial attention weights to each channel of the input feature map sequence;

[0092] Insert a CBAM attention mechanism behind each C2f_EMSC module to improve the attention of the YOLOv8 instance segmentation algorithm to the characteristics of cracks themselves, thereby improving the ability of the tunnel lining crack pixel-level detection model trained based on the YOLOv8 instance segmentation algorithm to accurately detect crack pixels and reduce the influence of environmental interference.

[0093] Step3: Replace the original CIoU loss function of the YOLOv8 instance segmentation algorithm with the EIoU loss function;

[0094] The loss function originally used by the YOLOv8 instance segmentation algorithm is the CIoU loss function. Although this loss function measures the model loss through three elements: the overlapping area, the center offset, and the width-height ratio, its "width-height ratio" measurement element uses a single parameter v to describe the overall ratio difference and fails to independently measure the specific deviation of the width and height. This coupled measurement method is prone to causing the optimization direction to deviate from the actual requirements. Especially when the center of the prediction box coincides with the ground truth box but there are significant size differences, it is difficult to accurately adjust the length and width parameters.

[0095] In contrast, the EIoU loss function decomposes the geometric difference into two independent dimensions: horizontal and vertical. By constructing explicit measurement elements for the width difference and height difference, the YOLOv8 instance segmentation algorithm can optimize the width and height dimension deviations specifically during the model training process. At the same time, the EIoU loss function introduces the Focal Loss dynamic weighting strategy, which effectively balances the optimization weights of easy and difficult samples during training by reducing the gradient contribution of low-quality prediction samples, enhancing the localization generalization ability of the model in complex environments while improving the model convergence efficiency.

[0096] In this embodiment, the calculation process of the EIoU loss function is as follows:

[0097] Calculate IoU, and the formula is as follows:

[0098]

[0099] Among them, the intersection area is the area of the overlapping part of the prediction box and the ground truth box, and the union area is the sum of all areas covered by the prediction box and the ground truth box;

[0100] Calculate the distance ρ(b, b gt ) between the centers of the prediction box and the ground truth box, and the formula is as follows:

[0101]

[0102] Among them, the center point coordinates of the prediction box are (x, y), and the center point coordinates of the ground truth box are (x gt , y gt );

[0103] Calculate the difference ρ(w, w gt ) in width and the difference ρ(h, h gt ) in height between the prediction box and the ground truth box;

[0104] Calculate the width c w and height c h of the smallest bounding box that can completely cover the two bounding boxes. The calculation process is as follows:

[0105] Calculate the maximum coverage range of the prediction box and the ground truth box in the x-axis direction:

[0106] Left boundary:

[0107] Right boundary:

[0108] Then the width c w of the smallest bounding box = x max - x min ;

[0109] Calculate the maximum coverage range of the prediction box and the ground truth box in the y-axis direction:

[0110] Upper boundary:

[0111] Lower boundary:

[0112] Then the height c h of the smallest bounding box = y max - y min ;

[0113] Substitute the above results into the EIoU loss calculation formula to obtain the EIoU loss value. The formula is as follows:

[0114]

[0115] Step 4: Based on the crack dataset established in Step 2, use the improved YOLOv8 instance segmentation algorithm in Step 3 to train and obtain a tunnel lining crack pixel-level detection model, and finally evaluate the detection performance and size of the model;

[0116] The specific steps are as follows:

[0117] Step1: Based on the improved YOLOv8 instance segmentation algorithm, train a tunnel lining crack pixel-level detection model;

[0118] In this embodiment, Python language and Pytorch deep learning framework are used for algorithm construction during training. The Python version is 3.8.16, and the Pytorch version is 1.12.1. The training computer is equipped with a 12th Gen Intel(R) Core(TM) i7-12700 processor, an NVIDIA GeForce RTX 3060 graphics card, 32G of memory, and 12G of video memory. During the model training process, GPU acceleration is used for computing, and the CUDA version is 11.4. In terms of training hyperparameter settings, the input image size is 640×640, the number of training iterations Epochs is 500 times, the initial learning rate is 0.01, and the cosine annealing learning rate is used. The rest all adopt the default settings of the YOLOv8 instance segmentation algorithm;

[0119] Step2: Evaluate the detection performance of the tunnel lining crack pixel-level detection model;

[0120] In this embodiment, the detection accuracy of the model is evaluated through three indicators: precision (P), recall (R), and mean average precision (mAP). The detection efficiency of the model is evaluated through the number of images output per second (FPS), the number of floating-point operations (FLOPs), the number of model parameters, and the model size. The present invention compares the detection performance of the models trained by the YOLOv8 instance segmentation algorithm before and after improvement, and the evaluation results are shown in Table 1;

[0121] Table 1 Evaluation Results of Model Detection Performance

[0122]

[0123] As can be seen from the table, in this embodiment, the crack pixel-level detection model trained by using the improved YOLOv8 instance segmentation algorithm proposed by the present invention is superior to the model trained by using the original YOLOv8 algorithm in all evaluation indicators. Among them, the mAP, an indicator for evaluating the crack pixel-level detection ability of the model mask has increased by 3.5%, the FPS, an indicator for evaluating the detection efficiency and lightweight level of the model, has increased by 8.8%, the FLOPs has decreased by 4.2%, the number of parameters has decreased by 6.3%, and the model size has decreased by 7.4%. The above results prove that the tunnel lining crack pixel-level detection model trained by using the improved YOLOv8 instance segmentation algorithm proposed by the present invention has better detection performance and is more suitable for deployment on lightweight devices.

[0124] Step Five: Input the tunnel lining image, and use the tunnel lining crack pixel-level detection model to detect the tunnel lining image, and output the tunnel lining crack detection result.

[0125] It should be noted that the above content only illustrates the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.

Claims

1. An intelligent detection method for tunnel lining cracks based on an improved instance segmentation algorithm, characterized in that, The steps include: Obtain tunnel lining pictures through a tunnel structure detection device; Input the tunnel lining pictures and use a crack pixel-level detection model to detect the tunnel lining pictures, and output the tunnel lining crack detection results; Among them, the training process of the crack pixel-level detection model includes: Select tunnel lining pictures containing cracks obtained by the tunnel structure detection device and perform data augmentation. Use the Labelme data annotation method to perform pixel-level annotation on the cracks therein to establish a crack dataset; Based on the crack dataset, use the improved YOLOv8 instance segmentation algorithm to train and obtain a tunnel lining crack pixel-level detection model.

2. The intelligent detection method for tunnel lining cracks based on an improved instance segmentation algorithm according to claim 1, wherein The establishment process of the improved YOLOv8 instance segmentation algorithm includes: Modify the backbone network for extracting crack image features and the neck network for feature fusion in the YOLOv8 instance segmentation algorithm. Replace the 3rd and 4th C2f modules in the backbone network and the 1st, 3rd, and 4th C2f modules in the neck network with more lightweight C2f_EMSC modules; Insert a CBAM attention mechanism behind each C2f_EMSC module in the YOLOv8 algorithm backbone network, with a total of two insertions; Replace the original CIoU loss function of the YOLOv8 algorithm with the EIoU loss function. Calculate the model bounding box regression loss through the EIoU loss function, which not only considers the overlapping area between the predicted box and the ground truth box, but also calculates and compares the length difference and width difference between the predicted box and the ground truth box respectively. In addition, EIoU combines the Focal Loss strategy to further optimize the sample imbalance problem in the model bounding box regression.

3. The intelligent detection method for tunnel lining cracks based on an improved instance segmentation algorithm according to claim 2, characterized in that, The C2f_EMSC module first uses a 1×1 convolution to change the number of channels of the input feature map sequence, and then uses the Split method to evenly divide the feature map sequence into 2 groups along the channel dimension. The first part of the feature map sequence is directly used as the input feature map sequence of the Concat layer; the other part of the feature map sequence passes through n Bottleneck_EMSC layers in sequence, and each Bottleneck_EMSC layer contains a 3×3 convolution and a more lightweight EMSConv convolution; every time the feature map sequence passes through 1 Bottleneck_EMSC layer, the output feature map sequence is divided into 2 groups, one group is used as the input feature map sequence of the Concat layer, and the other group continues to enter the next Bottleneck_EMSC layer; finally, the Concat layer combines and connects all the input feature map sequences and outputs them to the last 1×1 convolution to adjust the number of channels of the feature map sequence for subsequent operations.

4. The intelligent detection method for tunnel lining cracks based on an improved instance segmentation algorithm according to claim 3, wherein The EMSConv convolution first performs a Split operation on the input feature map sequence, dividing it into two groups on average. One group serves as the input feature map sequence of the Concat layer, and the other group is further divided into two smaller groups on average. One of the smaller groups performs a convolution operation with a kernel size of 3×3 to extract image features, and the other smaller group performs a convolution operation with a kernel size of 5×5 to extract image features. Then, the Concat method is used to combine and connect the feature map sequences output by the three branches to obtain a complete feature map sequence. Finally, a 1×1 convolution is used to exchange channel information to complete image feature extraction and adjust the number of channels of the feature map sequence.

5. The intelligent detection method for tunnel lining cracks based on an improved instance segmentation algorithm according to claim 2, characterized in that, The CBAM module is composed of a channel attention module and a spatial attention module in series. Among them, the channel attention module enhances the channel attention of the input feature map sequence and improves the information interaction between different channels. The spatial attention module enhances the spatial attention of the input feature map sequence and improves the information correlation between different positions of the same feature map.

6. The intelligent tunnel lining crack detection method based on the improved instance segmentation algorithm according to claim 5, characterized in that, The channel attention module includes the following running steps: Feature compression: In the spatial dimension H×W, for the input feature map sequence X with a shape of C×H×W, where C represents the number of input image channels, H represents the height of the input image, and W represents the width of the input image, global average pooling and global max pooling operations are performed simultaneously, and two groups of feature map sequences with a size of C×1×1 are output, that is, the number of input channels remains unchanged, and each channel is compressed into a single value, representing the maximum value and the average value of that channel respectively. Feature transformation: The two groups of feature map sequences obtained by the above pooling are input into a shared multi-layer perceptron to perform a fusion transformation on the pooled feature map sequences, extract the important features of each channel, and generate an attention feature map. Feature fusion and activation: The two groups of feature map sequences processed by the MLP are added together, and the attention weight of each channel is obtained through the Sigmoid activation function, reflecting the importance of each channel for the current task, and finally the channel attention weight is mapped to each channel of the input feature map sequence.

7. The intelligent detection method for tunnel lining cracks based on an improved instance segmentation algorithm according to claim 6, characterized in that, The multi-layer perceptron is composed of two fully connected layers. The first fully connected layer is used for channel dimensionality reduction to reduce the number of parameters during feature map fusion transformation, and the second fully connected layer is used for channel dimensionality increase to restore to the original number of channels for subsequent operations. The ReLU activation function is used between the two fully connected layers.

8. The intelligent detection method for tunnel lining cracks based on an improved instance segmentation algorithm according to claim 5, characterized in that, The spatial attention module includes the following running steps: Feature compression: Average pooling and max pooling operations are performed on the feature map sequence X' processed by the channel attention module in the channel dimension C, and two feature maps with a size of 1×H×W are output, that is, each pixel position of the input feature map sequence is compressed into a single value, representing the maximum value and the average value of that position respectively. Feature concatenation and convolution: The above two pooled feature maps are concatenated in the channel dimension to obtain a group of feature map sequences with a size of 2×H×W. Then, a convolution layer with a kernel size of 7×7 is used to reduce the number of channels to 1 to capture local and global spatial information. Feature activation: The feature map output by the convolutional layer is input into the Sigmoid activation function to obtain the attention weights at each spatial position, reflecting the importance of each pixel position for the current task, and finally mapping the spatial attention weights to each channel of the input feature map sequence.

9. The intelligent detection method for tunnel lining cracks based on an improved instance segmentation algorithm according to claim 2, characterized in that, The calculation process of the EIoU loss function is as follows: Calculate IoU, and the formula is as follows: Among them, the intersection area is the area of the overlapping part of the predicted box and the ground truth box, and the union area is the sum of all regions covered by the predicted box and the ground truth box; Calculate the distance ρ(b, b gt ), and the formula is as follows: Among them, the center point coordinates of the prediction box are (x, y), and the center point coordinates of the ground truth box are (x gt , y gt ); Calculate the difference ρ(w, w gt ) between the widths of the predicted bounding box and the ground truth bounding box, and the difference ρ(h, h gt ) between their heights; Calculate the width c of the smallest bounding box that can completely cover two bounding boxes w and the height c h , and the calculation process is as follows: Calculate the maximum coverage range of the predicted box and the ground truth box in the x-axis direction: Left boundary: Right border: Then the width c of the minimum bounding box w = x max - x min ; Calculate the maximum coverage range of the predicted box and the ground truth box in the y-axis direction: Upper boundary: Lower boundary: Then the height c of the minimum bounding box h = y max - y min ; Substitute the above results into the EIoU loss calculation formula to obtain the EIoU loss value, and the formula is as follows:

Citation Information

Cited By

  • Low-altitude weak and small target detection and tracking method based on deep learning

    CN121353329A

  • Lightweight lithium mineral microscopic image real-time detection and instance segmentation method

    CN121437530A

  • Tunnel lining crack detection method based on improved RT-DETR

    CN121685461A

  • Tunnel lining crack detection method based on improved RT-DETR

    CN121685461B