Tunnel crack detection method and system
By introducing a multi-head detection structure and CBAM attention mechanism in tunnel crack detection, the problems of low accuracy, low real-time and insufficient adaptability in the prior art are solved, and the crack detection effect of high-precision, rapid and adaptable to complex environments is achieved.
Patent Information
- Application Number
- CN202410446502.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-04-15
AI Technical Summary
The prior art has problems of low accuracy, low real-time and insufficient adaptability in tunnel lining crack detection, especially when detecting small cracks and adapting to complex environments.
A multi-head detection structure and CBAM attention mechanism are introduced. By connecting the characteristics of multiple levels of backbone networks with different segmentation heads, combining channel and spatial attention mechanisms, the feature fusion process is optimized and the model's attention to the crack area is improved.
It significantly improves the accuracy and accuracy of tunnel crack detection, reduces missed and missed detection, improves detection quality, and improves computing speed and resource utilization efficiency through lightweight design, which is suitable for real-time detection.
Smart Images

Figure CN118230060B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a tunnel crack detection method and system. Background Art
[0002] Tunnels play a vital role in urban infrastructure. As an integral part of the transportation system, their integrity and safety are extremely demanding. As a hidden project in the underground geotechnical medium, various types of defects (such as cracks, leakage, and exposed steel bars) are inevitably observed in the tunnel lining under the action of complex environments. Among them, cracks are the most common defects, mainly caused by eccentric loads, groundwater, long-term differential settlement, temperature changes, internal vehicle loads, and adjacent excavations. Once a crack is generated and not treated in time, it can serve as a leakage channel, leading to steel corrosion and concrete carbonization. Therefore, it is of great practical value to detect and treat lining cracks for high-quality tunnel operation.
[0003] In traditional methods, the detection of cracks in tunnel linings usually relies on manual inspections and visual observations. This method has the problems of strong subjectivity, low efficiency, and difficulty in accurately determining the extent and scope of cracks. In particular, for small cracks or cracks located high in the tunnel, it is difficult for the human eye to accurately identify and locate them, which may lead to inaccurate detection results and crack omissions. In recent years, with the development of computer vision and deep learning technology, automated crack detection technology has become a hot topic of research. The emergence of deep learning methods heralds a breakthrough in image processing technology, and more and more researchers are applying deep learning methods to the field of crack detection. Deep learning-based crack detection methods can be divided into three categories: image-level classification, patch-level object detection, and pixel-level semantic segmentation. The first two methods can locate the position of cracks in an image, but their results are rough and cannot determine the morphology and quantification of cracks. Semantic segmentation assigns a label to each pixel in the image, thereby achieving accurate positioning of crack pixels. Therefore, it is naturally suitable for crack detection tasks.
[0004] Most of the semantic segmentation algorithms used in existing research on tunnel lining crack recognition are based on convolutional neural networks (CNNs). Due to the locality of CNNs, such algorithms have good local feature extraction capabilities, but cannot fully utilize contextual semantic information, resulting in difficulty in capturing global features. Therefore, it has the problem of insufficient robustness and bottleneck in improving crack detection accuracy. In recent years, a visual Transformer based on a self-attention mechanism has been proposed, which is not restricted by local interactions and can mine long-range feature dependencies to learn the most appropriate inductive bias. However, such a network lacks the inherent inductive bias of CNN, including translation invariance and local sensitivity, which leads to the shortcomings of strong data dependence and easy loss of local features in tunnel lining crack segmentation. Recently, SwinTransformer and CNN are integrated into the encoding and decoding framework of DeepLabv3+ to propose a hybrid tunnel crack semantic segmentation algorithm SCDeepLab. However, such a network has the problems of large computational complexity and low real-time performance.
[0005] In general, the current deep learning-based tunnel lining crack detection technology faces several major problems. First, for the problem of small crack detection, the traditional convolutional neural network cannot make full use of contextual semantic information, and has insufficient robustness and bottlenecks in improving crack detection accuracy. Although some Transformer-based neural networks perform well in capturing global information, they have not yet achieved satisfactory results in the problem of small crack detection due to some technical difficulties. Secondly, the real-time problem is one of the challenges that current technology urgently needs to solve. Faced with a large number of tunnel linings, efficient, fast and accurate crack detection is required, and the existing technology has not yet fully met the real-time requirements. Adaptability issues in multiple environments: The surface of the tunnel lining has complex textures, lighting changes and noise, and more robust low-detection algorithms are needed to meet these challenges.
[0006] The above challenges limit the application and effectiveness of existing technologies in the field of tunnel lining crack detection, so further technological innovation and improvement are needed to improve detection accuracy, real-time performance and adaptability.
[0007] Current tunnel lining crack detection technology faces several major problems:
[0008] 1. Accuracy issues: Mainly for the problem of small crack detection, traditional convolutional neural networks cannot fully utilize contextual semantic information, and have insufficient robustness and bottlenecks in improving crack detection accuracy. Although some Transformer-based neural networks perform well in capturing global information, they have not yet achieved satisfactory results in the problem of small crack detection due to some technical difficulties.
[0009] 2. Real-time problem: Although a large number of methods have achieved good results in crack segmentation, it is still challenging to perform efficient, fast and accurate crack detection on a large number of tunnel lining surfaces due to the large amount of computation required by ordinary deep learning neural networks and high hardware requirements.
[0010] 3. Adaptability issues in different environments: The tunnel lining surface has complex textures, lighting changes and noise, and more robust detection algorithms are needed to meet these challenges. Summary of the invention
[0011] In order to overcome the deficiencies of the prior art, the present invention provides a tunnel crack detection method and system to solve the problems of low accuracy and the like in the prior art.
[0012] The technical solution adopted by the present invention to solve the above problems is:
[0013] A tunnel crack detection method introduces a multi-head detection structure in a semantic segmentation model, connecting the features of multiple-level backbone networks with different segmentation heads; wherein the multi-head detection structure can establish multiple branches in the backbone network, each branch is connected to a different level of the backbone.
[0014] As a preferred technical solution, for the feature maps of multiple segmentation heads, a channel attention mechanism and a spatial attention mechanism are added to the outputs of different segmentation heads.
[0015] As a preferred technical solution, the following steps are included:
[0016] S1, image acquisition: collecting tunnel lining surface images, annotating the original images, and dividing the original images into different pixel categories; the pixel categories include background and cracks;
[0017] S2, training: construct a tunnel crack detection model and use the original image data to train the tunnel crack detection model.
[0018] As a preferred technical solution, step S2 includes the following steps:
[0019] S21, environment preparation: prepare the operating environment of the tunnel crack detection model;
[0020] S22, downloading the pre-trained model weight file: downloading the pre-trained model weight file, where the pre-trained model refers to a model that has been pre-trained using a data set;
[0021] S23, data set preparation: prepare crack data set for tunnel crack detection model training;
[0022] S24, optimizer and learning rate setting: set the optimizer and learning rate;
[0023] S25, loss function setting: setting the loss function;
[0024] S26, training setting: setting training parameters;
[0025] S27, feature extraction and image segmentation: using the pre-trained weight file and the original image data to train the tunnel crack detection model, and perform feature extraction and image segmentation of the original image;
[0026] S28, crack identification and positioning: Based on the different feature maps output by different fusion layers of the backbone network of the tunnel crack detection model, the cracks are identified and positioned.
[0027] As a preferred technical solution, in step S24, the learning rate is set to a value in the range of (0.0001, 0.001).
[0028] As a preferred technical solution, in step S24, the momentum coefficient of the optimizer is set to a value in the range of (0.9, 0.99), and the weight attenuation of the optimizer is set to a value in the range of (0.001, 0.1).
[0029] As a preferred technical solution, in step S25, the calculation formula of the loss function is as follows:
[0030]
[0031] Where i represents the number of pixels, N represents the total number of pixels, c represents the number of pixel categories, C represents the number of pixel categories, and y represents the number of pixel categories. i,c represents the true label, y i,c The value of p is 0 or 1. i,c Indicates the probability that the model predicts the pixel category numbered i.
[0032] As a preferred technical solution, in step S28, different segmentation heads are added to different feature maps, the output results are added, and then the channel attention mechanism and the spatial attention mechanism are used to add the added results to the channel attention module and the spatial attention module, so that the tunnel crack detection model pays more attention to the crack area and realizes the identification and positioning of the cracks.
[0033] As a preferred technical solution, in step S2, the structure of the constructed tunnel crack detection model includes a backbone network, a CBAM attention head, an addition operation module, and n segmentation heads; wherein n represents the total number of segmentation heads, n≥2 and n is an integer;
[0034] The backbone network includes a downsampling layer, and the output end of the downsampling layer is connected to two parallel branches: respectively denoted as the first branch and the second branch; the first branch includes n lightweight feature extraction blocks, and the second branch includes n fusion modules;
[0035] The segmentation head, lightweight feature extraction block, and fusion module are all numbered from 1 to n according to the distance from the downsampling layer.
[0036] Each fusion module includes conv+BN layer A, conv+BN layer B, sigmod activation layer, Up layer, and convolution operation module; conv+BN layer B, sigmod activation layer, and Up layer are connected in series in sequence, conv+BN layer A and Up layer are respectively connected to the input end of the convolution operation module, the conv+BN layer A of the first fusion module is connected to the output end of the sampling layer, and the conv+BN layer B of the i-th fusion module is connected to the output end of the i-th lightweight feature extraction block; wherein i represents the number of the segmentation head, lightweight feature extraction block, or fusion module, 1≤i≤n, and l is an integer;
[0037] The output end of the convolution operation module of the nth fusion module is connected to the input end of the nth segmentation head;
[0038] The output ends of the n segmentation heads are respectively connected to the input ends of the sum operation module, and the output end of the sum operation module is connected to the input end of the CBAM attention head.
[0039] A tunnel crack detection system, used to implement the tunnel crack detection method, comprises the following modules connected in sequence:
[0040] Image acquisition module: used to acquire tunnel lining surface images, annotate original images, and divide original images into different pixel categories; wherein pixel categories include background and cracks;
[0041] Training module: used to build a tunnel crack detection model and train the tunnel crack detection model using original image data.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] (1) Improved detection precision and accuracy: By introducing a multi-head lightweight semantic segmentation model and CBAM attention mechanism, the present invention significantly improves the precision and accuracy in tunnel crack detection; especially for the identification of small cracks, it reduces missed detection and false detection, thereby improving the overall detection quality;
[0044] (2) Lightweight model: Compared with traditional models, the present invention designs a lightweight structure, reduces the computational complexity and resource consumption of the model, and is suitable for mobile terminal deployment. This feature not only improves the computing speed, but also saves computing resources and energy consumption. The present invention can be deployed on edge devices for real-time segmentation.
[0045] (3) Effective attention mechanism improves recognition ability: The introduced CBAM attention mechanism optimizes the feature fusion process and more effectively focuses on and extracts key information in channels and spaces, thereby improving the ability to locate and identify cracks;
[0046] (4) Friendly to raw materials and the environment: By improving detection efficiency and accuracy, the present invention can reduce missed detection and mispositioning when maintaining and repairing tunnel structures, thereby saving maintenance materials and resources. This helps to reduce maintenance costs while reducing negative impacts on the environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a structural schematic diagram of a lightweight tunnel crack detection model with a multi-head structure and attention fusion according to the present invention;
[0048] Figure 2 It is a schematic diagram of the structure of the channel attention mechanism and spatial attention mechanism model;
[0049] Figure 3 Schematic diagram of the original crack image and the real label image randomly selected from the experimental dataset;
[0050] Figure 4 The original image and the effect diagram of crack segmentation implementation. DETAILED DESCRIPTION
[0051] The present invention will be further described in detail below in conjunction with embodiments and drawings, but the embodiments of the present invention are not limited thereto.
[0052] Example 1
[0053] like Figures 1 to 4 As shown in the figure, this technology uses deep learning technologies such as multi-head detection mechanism and different attention mechanisms to automatically identify and locate tunnel lining cracks. While ensuring accuracy and real-time performance, it improves the degree of automation of detection and has higher accuracy and efficiency than traditional methods.
[0054] In response to problem 1 faced by the current tunnel lining crack detection technology pointed out in the background technology, the present invention will be based on the lightweight backbone network Seaformer (Squeeze-enhanced Axial Transformer), extract different features of the input image through multi-level fusion, and then use the attention mechanism to make the neural network pay more attention to the crack information, and finally achieve better results.
[0055] In response to the problem 2 pointed out by the background technology that the current tunnel lining crack detection technology faces, we selected a lightweight backbone network Seaformer, which is an attention module with compressed axial and detail enhancement, which can be better applied on mobile terminals.
[0056] In view of the problem 3 of the current tunnel lining crack detection technology pointed out in the background technology, the present invention mainly solves the adaptability problem of the algorithm in segmentation under different environments by selecting crack surfaces in a variety of different environments as data sets, including different scenes, different lighting conditions, different textures, etc. At the same time, this method solves the robustness problem of the algorithm through data enhancement methods such as Resize, RandomCrop, RandmoFlip, PhotoMetricDistorion, etc.
[0057] The technical implementation plan mainly includes the following steps.
[0058] 1. Image acquisition and preprocessing:
[0059] Use a professional camera or industrial camera to collect tunnel lining surface images, and use the labeling tool (AnyLabeling) to label the original images, and divide the original images into two categories: background and cracks. The specific operation will be described in the specific implementation plan in Section 6. Preprocess the images, including cropping, denoising, brightness equalization, and image enhancement, to improve the accuracy and stability of subsequent processing.
[0060] Use image acquisition devices with different resolutions.
[0061] 2. Training:
[0062] Environment preparation: This method is implemented based on the existing segmentation framework MMSegmentation. You need to prepare mmcv-full provided by MMSegmentation, version 1.3.14. You also need to install pytorch version greater than 1.5.
[0063] Download the pre-trained model weight file: Next, you need to download the pre-trained model weight file. A pre-trained model refers to a model that has been pre-trained using a large-scale data set. The training phase of this model usually includes learning the structure, grammar, semantics, and rich contextual information of the language, and has a good effect on transfer learning and feature extraction. This method requires downloading the pre-trained weight file SeaFormer.pth.
[0064] Dataset preparation: This method is experimented on multiple crack datasets. It is recommended to symbolically link the dataset root directory to the data directory under the MMSegmentation framework. If your directory structure is different, you may need to change the corresponding path in the configuration file.
[0065] For optimizer and learning rate settings, we used the AdamW optimizer and set the learning rate to 0.0005. The momentum coefficient of the AdamW optimizer was set to (0.9, 0.99). The weight decay was set to 0.01 to regularize the model parameters. Special configurations were made for different model parameters, such as no weight decay for the position embedding layer (pos_emb), and different learning rate multipliers for certain layers (head, norm). The learning rate adopted a polynomial decay strategy, and a linear warm-up method was used to gradually increase the learning rate in the initial stage of training to help the model adapt to the training data more effectively. The learning rate was warmed up within 1500 iterations, and the warm-up learning rate was 1e of the initial learning rate. -6 The exponent of the polynomial decay is 1.0. The decay and warm-up of the learning rate are based on the number of iterations rather than epochs.
[0066] Loss function setting, this method uses cross entropy loss to predict the category of each pixel. In semantic segmentation tasks, pixel-level cross entropy loss is usually used. For each pixel, the model outputs a probability distribution, indicating the probability that the pixel belongs to each category. The calculation formula of the loss function is as follows:
[0067]
[0068] Where N is the total number of pixels, C is the number of categories, and y i,c is the true label 0 or 1, p i,c is the probability of the model predicting the pixel category. This loss function evaluates the classification results of the model at each pixel position, and optimizes the model by minimizing this loss to make it better suited to the segmentation task.
[0069] Training settings,During training, we set the model to be saved every 4000 iterations.,Evaluation is performed every 4000 iterations, using mIoU as the evaluation metric.,Logs are output every 50 iterations, using TextLogger and TensorboardLogger.,Use IterBasedRunner as the trainer, with a maximum number of iterations of 200,000.
[0070] Feature extraction and image segmentation: The backbone network and pre-trained weight files of the lightweight model SeaFormer are used for training. This backbone network design adds axis compression to enhance attention of position information. On the one hand, the Q, K, and V features are compressed and then enhanced. On the other hand, the Q, K, and V features are enhanced with a convolutional network. Finally, the two are fused to output enhanced features. We use the output results of different feature layers of the backbone network as input to different segmentation heads. In this way, the floating-point operations of our model on Crack11k are reduced by 2.7 times compared to the traditional method SegFormer (b0). This multi-level feature fusion method is used for feature extraction, and three different fusion layers are extracted to facilitate the subsequent accurate identification of cracks in the tunnel lining.
[0071] Crack identification and positioning: Different lightweight segmentation heads are added to different feature maps output by different fusion layers of the backbone network, and the output results are added (Concat). Then, the channel attention mechanism and spatial attention are used to add the added results to the channel attention module and the spatial attention module, so that the tunnel crack detection model pays more attention to the crack area, and finally realizes the identification, positioning and marking of the cracks.
[0072] 3. Testing:
[0073] Test setup, the data processing flow defined during the test includes operations such as image loading, multi-scale flipping, and normalization. Specifically, MultiScaleFlipAug is used to perform multi-scale test data enhancement. At the same time, multiple indicators such as MIoU, FScore, and MAcc are used to evaluate the segmentation ability of the model during the test. Finally, the segmentation results of the model on the test data are visualized. The entire test process mainly involves model construction, data loading, inference, and result preservation and evaluation.
[0074] The differences between this method and the prior art are described below:
[0075] This study proposes an improved model based on a lightweight and efficient Backbone structure for crack semantic segmentation tasks. The model introduces a multi-head detection structure, which connects the features of multiple levels of Backbone with different segmentation heads to capture cracks on the tunnel surface, especially subtle cracks. In addition, we apply an attention mechanism to the outputs of different segmentation heads, combining channel and spatial attention to enhance the model's accurate attention to crack areas.
[0076] Multi-head detection structure,Traditional semantic segmentation models usually use a single-head structure for feature extraction, but the model proposed in this study introduces a multi-head detection structure. This structure can establish multiple branches in the network, each branch is connected to a different level of the Backbone, and make full use of multi-scale feature information, thereby enhancing the model's perception of different levels of fracture surface features.
[0077] For the feature maps of multiple segmentation heads, we introduced the CBAM (Convolutional Block Attention Module) attention mechanism to enhance the model's attention to the crack area. This mechanism combines channel and spatial attention, adjusts the feature map in a specific spatial range and channel dimension, and helps the model to learn and express the salient features of the crack area more centrally.
[0078] The multi-head detection structure and CBAM attention mechanism proposed in this paper have achieved significant improvement in the crack semantic segmentation task. These innovative contributions provide new ideas and solutions for the field of semantic segmentation and are expected to be promoted and applied in a wider range of scenarios. Compared with traditional methods, it has higher accuracy, faster detection speed and better adaptability in complex tunnel environments.
[0079] Same features as prior art:
[0080] The steps of image acquisition, preprocessing and crack identification are similar to general image processing technology, but they differ in deep learning model detection method, crack detection accuracy and real-time performance.
[0081] About the dataset: The researchers of this invention conducted corresponding experiments on three different datasets, including two public datasets, CrackSeg9k and Crack Segmentation Dataset, and a self-made tunnel crack dataset, Tunnel Crack. The two public datasets contain seven identical sub-datasets. However, there are still some differences. CrackSeg9k is refined to address the presence of noisy annotations, while the latter consists of raw images without preprocessing. Tunnel Crack was made by ourselves, including real lining surfaces in tunnels, as well as cracks on relatively smooth surfaces selected in public datasets, which are very similar to cracks on the lining surfaces in tunnels. In terms of quantity, CrackSeg9k contains 9,255 images (400×400 resolution), Crack Segmentation Dataset contains 11298 images (448×448 resolution), and Tunnel Crack contains 6583 images (448×448 resolution).
[0082] Compared with the prior art, the present invention has the following advantages, including but not limited to:
[0083] 1. Improve detection precision and accuracy:
[0084] By introducing a multi-head lightweight semantic segmentation model and a CBAM attention mechanism, the present invention significantly improves the precision and accuracy in tunnel crack detection. In particular, for the identification of small cracks, missed detections and false detections are reduced, thereby improving the overall detection quality. Our evaluation indicators mainly include mIoU and mAcc, which are used to calculate the average intersection over union (IoU) and average accuracy of the target area, respectively. On Crack9k, our method improves mIoU, an important evaluation indicator for segmentation, by 0.31 percentage points compared to existing methods, and improves mAc by 2.1 percentage points. On Crack Segmentation Dataset, our method improves mIoU by 1.9 percentage points and mAc by 2 percentage points. On these two datasets, our method outperforms some currently known crack detection methods. On Tunnel Crack, our method also achieves good results in mIoU and mAc.
[0085] 2. Model lightweight:
[0086] Compared with the traditional model, the present invention designs a lightweight structure, reduces the computational complexity and resource consumption of the model, and is suitable for mobile deployment. This feature not only improves the computing speed, but also saves computing resources and energy consumption. It performs well in the two parameters of evaluating model size (Params) and computational complexity (GFLOPS). When the input model image is 256*256, the GFLOPS is only 1.61, and the Params is 14.01M. The number of parameters is only half of that of the traditional model PSPNet (ResNet18), and the GFLOPS is only one-fortieth. These two data show that our method can be deployed on edge devices for real-time segmentation.
[0087] 3. Effective attention mechanism improves recognition ability:
[0088] The introduced CBAM attention mechanism optimizes the feature fusion process and focuses on and extracts key information more effectively in channels and spaces, thereby improving the ability to locate and identify cracks.
[0089] 4. Friendly to raw materials and environment:
[0090] By improving detection efficiency and accuracy, the present invention can reduce missed detection and wrong positioning when maintaining and repairing tunnel structures, thereby saving maintenance materials and resources. This helps to reduce maintenance costs while reducing negative impacts on the environment.
[0091] In general, the present invention has significant advantages and beneficial effects over the prior art, especially in the field of crack detection, in terms of improvements in accuracy, efficiency and resource conservation.
[0092] In the present invention, the technical points that play an important role include:
[0093] 1. In order to solve the problem of difficulty in segmenting small cracks in crack segmentation, we proposed a multi-level feature fusion mechanism. Through the multi-head structure, the model can achieve better results than existing models for crack types that are difficult to detect, especially subtle cracks.
[0094] 2. To address the problem that cracks are difficult to locate due to their small size relative to the overall image, a channel attention mechanism and a spatial attention mechanism are implemented behind the multi-head structure, which enables the model to better locate and mark cracks for further segmentation.
[0095] To address the problem of low real-time performance of traditional segmentation methods, the present invention selects a lightweight and efficient backbone network Seaformer and designs different lightweight segmentation heads to reduce the number of floating-point operations, thereby improving the speed of model reasoning.
[0096] Example 2
[0097] like Figures 1 to 4 As shown, as a further optimization of Example 1, based on Example 1, this embodiment also includes the following technical features:
[0098] Figure 1 The relevant English words and English abbreviations are translated as follows;
[0099] conv+BN-convolution+batch normalization;
[0100] Up-up sampling;
[0101] Resize: resize the image to the specified size;
[0102] RandomCrop: Randomly select an area in the image and crop it to the specified size;
[0103] RandomFilp-Random Flip: Randomly flip the image horizontally or vertically;
[0104] PhotoMetricDistortion: Randomly distort or transform the luminosity of an image to increase the diversity of the data.
[0105] Normalize: Normalize the image and scale the pixel values of the image to a specified range;
[0106] Pad: Fill the image with a specified pixel value or color around it to make it reach the specified dimensions.
[0107] The specific steps of the present invention are as follows:
[0108] Step 1: Data collection: Sampling cracks on the inner surface of a real tunnel to obtain crack images. The original images are cropped to 448×448 resolution images. The cropped images are then manually labeled and labeled for semantic segmentation. Figure 1 It is generally difficult to label, so we use AnyLabeling for labeling. This software can help us with the original automatic labeling. In the end, only the labeler needs to make fine adjustments to get the final label map. The specific operations are as follows. First, install AnyLableing. You can install it directly through Anaconda. After the installation is complete, choose to download the model. The download model is slow, so you can download it in advance and put it in the C drive user. Then use labelme to open the labeled folder. Make fine adjustments. After the labeling is completed, there will be a Json file with the same name as each picture in the file, which is the label of each picture.
[0109] At the same time, since there are fewer crack images in the tunnel scene, in order to improve our final detection effect, we also selected images similar to cracks on the tunnel lining surface from the public dataset. Finally, the prepared dataset is divided into a training set and a test set for the next step of training.
[0110] Step 2: Input the crack images in the training set into the model for training. The specific method is as follows:
[0111] Step 1: Data enhancement, including the following processes: Resize, RandomCrop, RandomFilp, PhotoMetricDistortion, Normalize, Pad. It is worth noting that these enhancement methods are not the only ones, and other enhancement methods can be used, but the final effect may not be good.
[0112] Step 2: Perform a 3×3 convolution on the preprocessed image and input the result into a 3×3 MobileNetV2 (MobileNet stands for mobile network module, second version, abbreviated as MV) block with a step size of 1;
[0113] Step 3: Input the result of the previous step into a 3×3 MobileNetV2 block with a step size of 2, and then input the result of this step into a 3×3 MobileNetV2 block with a step size of 1;
[0114] Step 4: Input the result of step 3 into a 5×5 MobileNetV2 block with a step size of 2, and then input the result of this step into a 5×5 MobileNetV2 block with a step size of 1;
[0115] Step 5: Feed the result of step 4 into a 3×3 MobileNetV2 block with a stride of 2, then feed the result of this step into a 3×3 MobileNetV2 block with a stride of 1, and then feed the result into the SeaFormer layer;
[0116] Step 6: Input the result of step 5 to a 5×5 MobileNetV2 block with a stride of 2, and then input the result to the SeaFormer layer;
[0117] Step 7: Input the result of step 6 into a 3×3 MobileNetV2 block with a stride of 2, and then input the result into the SeaFormer layer;
[0118] The above steps are to realize the backbone network of lightweight and effective feature extraction through SeaFormer layer and MobileNetV2 module. These two modules are the key to the lightweight and efficient nature of this method. The key point is that the SeaFormer layer design adds axis compression of position information to enhance attention. On the one hand, the Q, K, and V features are compressed on the axis and then the attention is enhanced. On the other hand, the Q, K, and V features are used to enhance local information using a convolutional network. Finally, the two are fused to output enhanced features.
[0119] In this embodiment, the lightweight feature extraction block adopts the MobileNetV2+SeaFormer block:
[0120] Each MobileNetV2+SeaFormer block includes a MobileNetV2 block and a SeaFormer block in series; if 1≤i≤n-1, the output of the convolution operation module of the i-th fusion module is connected to the input of the i-th segmentation head and the conv+BN layer A of the i+1-th fusion module respectively;
[0121] If 1≤i≤n-1, the SeaFormer block of the i-th MobileNetV2+SeaFormer block is connected to the conv+BN layer B of the i-th fusion module and the MobileNetV2 block of the i+1-th MobileNetV2+SeaFormer block respectively.
[0122] It is worth noting that the lightweight feature extraction block can also be implemented using other structures.
[0123] Step 8: Take the output of step 3 and the output of step 5 as input and input them into the fusion block. The specific operation of the fusion block is to input the result of step 3 into the convolution layer and batchnorm layer to get result a, and input the result of step 5 into the convolution layer, batchnorm layer, sigmoid activation layer and an upsampling layer in turn to get result b, and finally multiply the two results a and b to get the fusion output;
[0124] Step 9: Take the results of step 8 and step 6 as input and input them into the fusion module in the same way as step 8;
[0125] Step 10: Take the results of step 9 and step 7 as input and input them into the fusion module in the same way as step 8;
[0126] Step 11: Take the results of steps 8, 9, and 10 as inputs to different segmentation heads. The specific operation of the head is to input the above results to the convolution layer, batchnorm layer, and relu activation layer in sequence. The output results are recorded as r8, r9, and r10;
[0127] Steps 8, 9, 10, and 11 mainly input the multi-level features output by the backbone network into different lightweight segmentation heads designed by us. In this way, the problem that common crack detection has poor detection effect on small cracks can be solved.
[0128] Step 12: Perform a concat operation on the output results of step 11 and input them into the channel attention module and spatial attention module of CBAM. Through the attention mechanism in this step, the model focuses more on the cracks and finds the key points in the output results of multiple segmentation heads, thereby improving the segmentation accuracy;
[0129] Step 13: Input the output of step 12 into the 1×1 convolution layer, and the model output is completed;
[0130] Step 14: Perform a bilinear interpolation to expand the output of step 13 to the label Figure 1 The size of the image is then changed to the same as the label, and then a cross entropy loss is performed. Finally, the model parameters are updated through back propagation and the optimizer. Figure 4 In the rendering, the lines are where the cracks are.
[0131] Step 3: After the training of step 2 is completed, the crack images in the test set are tested using the trained segmentation model to obtain the test results.
[0132] As described above, the present invention can be preferably implemented.
[0133] All features disclosed in all embodiments in this specification, or steps in all methods or processes implicitly disclosed, except for mutually exclusive features and / or steps, can be combined and / or expanded or replaced in any manner.
[0134] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. According to the technical essence of the present invention, within the spirit and principles of the present invention, any simple modification, equivalent replacement and improvement made to the above embodiment still falls within the protection scope of the technical solution of the present invention.
Claims
1. A tunnel crack detection method, characterized in that: A multi-head detection structure is introduced into the semantic segmentation model to connect the features of multiple levels of backbone networks with different segmentation heads. The multi-head detection structure can establish multiple branches in the backbone network, and each branch is connected to a different level of the backbone. For the feature maps of multiple segmentation heads, add channel attention mechanism and spatial attention mechanism to the output of different segmentation heads; The following steps are involved: S1, image acquisition: collecting tunnel lining surface images, annotating the original images, and dividing the original images into different pixel categories; the pixel categories include background and cracks; S2, training: constructing a tunnel crack detection model and using the original image data to train the tunnel crack detection model; Step S2 includes the following steps: S21, environment preparation: prepare the operating environment of the tunnel crack detection model; S22, downloading the pre-trained model weight file: downloading the pre-trained model weight file, where the pre-trained model refers to a model that has been pre-trained using a data set; S23, data set preparation: prepare crack data set for tunnel crack detection model training; S24, optimizer and learning rate setting: set the optimizer and learning rate; S25, loss function setting: setting the loss function; S26, training setting: setting training parameters; S27, feature extraction and image segmentation: using the pre-trained weight file and the original image data to train the tunnel crack detection model, and perform feature extraction and image segmentation of the original image; S28, crack identification and location: Based on the different feature maps output by different fusion layers of the backbone network of the tunnel crack detection model, the cracks are identified and located; In step S2, the structure of the constructed tunnel crack detection model includes a backbone network, a CBAM attention head, an addition operation module, and n segmentation heads; wherein n represents the total number of segmentation heads, n≥2 and n is an integer; The backbone network includes a downsampling layer, and the output end of the downsampling layer is connected to two parallel branches: respectively denoted as the first branch and the second branch; the first branch includes n lightweight feature extraction blocks, and the second branch includes n fusion modules; The segmentation head, lightweight feature extraction block, and fusion module are all numbered from 1 to n according to the distance from the downsampling layer. Each fusion module includes conv+BN layer A, conv+BN layer B, sigmod activation layer, Up layer, and convolution operation module; conv+BN layer B, sigmod activation layer, and Up layer are connected in series in sequence, conv+BN layer A and Up layer are respectively connected to the input end of the convolution operation module, the conv+BN layer A of the first fusion module is connected to the output end of the sampling layer, and the conv+BN layer B of the i-th fusion module is connected to the output end of the i-th lightweight feature extraction block; wherein i represents the number of the segmentation head, lightweight feature extraction block, or fusion module, 1≤i≤n and 1 is an integer; The output end of the convolution operation module of the nth fusion module is connected to the input end of the nth segmentation head; The output ends of the n segmentation heads are respectively connected to the input ends of the sum operation module, and the output end of the sum operation module is connected to the input end of the CBAM attention head.
2. A tunnel crack detection method according to claim 1, characterized in that: In step S24, the learning rate is set to a value in the range of (0.0001, 0.001).
3. A tunnel crack detection method according to claim 1, characterized in that: In step S24, the momentum coefficient of the optimizer is set to a value in the range of (0.9, 0.99), and the weight decay of the optimizer is set to a value in the range of (0.001, 0.1).
4. A tunnel crack detection method according to claim 1, characterized in that: In step S25, the calculation formula of the loss function is as follows: Where i represents the number of pixels, N represents the total number of pixels, c represents the number of pixel categories, C represents the number of pixel categories, and y represents the number of pixel categories. i,c represents the true label, y i,c The value of p is 0 or 1. i,c Indicates the probability that the model predicts the pixel category numbered i.
5. A tunnel crack detection method according to claim 1, characterized in that: In step S28, different segmentation heads are added to different feature maps, the output results are added, and then the channel attention mechanism and the spatial attention mechanism are used to add the added results to the channel attention module and the spatial attention module, so that the tunnel crack detection model pays more attention to the crack area and realizes the recognition and positioning of the cracks.
6. A tunnel crack detection system, characterized in that: A tunnel crack detection method for implementing any one of claims 1 to 5, comprising the following modules connected in sequence: Image acquisition module: used to acquire tunnel lining surface images, annotate original images, and divide original images into different pixel categories; wherein pixel categories include background and cracks; Training module: used to build a tunnel crack detection model and train the tunnel crack detection model using original image data.
Citation Information
Patent Citations
Attention mechanism fused lightweight bridge surface crack segmentation method and equipment
CN115511787A