A Multi-Label Defect Recognition Method for Concrete Based on Deep Learning
The deep learning-based concrete defect recognition method using CA attention and SPP layers in the EfficientNetV2 network addresses the challenge of accurately classifying complex concrete defects, achieving a 6.7% improvement in accuracy and reducing computational costs.
Patent Information
- Application Number
- CN202310760285.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-06-25
AI Technical Summary
The existing concrete defect identification methods are low in accuracy and efficiency in non-contact detection, making it difficult to quickly identify multiple defect types, and are greatly affected by concrete surface intrusion factors.
Using a deep learning-based method, the SE attention module in the MBConv module of the EfficientNetV2 network is replaced by the CA attention module, and the adaptive pooling layer of the classification head is replaced by the SPP layer, combined with Grad-CAM for visual verification, and the training is performed using cosine annealing learning rate scheduling and a progressive learning strategy to retain the aspect ratio information of the image.
The accuracy of concrete multi-label defect recognition has been significantly improved, from 82.8% to 89.5%, effectively reducing the impact of image aspect ratio distortion, reducing calculation costs, and improving the generalization ability and recognition efficiency of the model.
Smart Images

Figure CN116740037B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of concrete defect identification, and specifically relates to a method for multi-label defect identification of concrete based on deep learning. Background Art
[0002] Bridge engineering is an important lifeline project related to the national economy and people's livelihood. Due to the lack of modern detection methods and intelligent means, many bridges around the world are facing the common problem of being in disrepair for years but still working with hidden damages. The quality of concrete plays an important role in ensuring the safety and beauty of bridges. As a composite material, the surface reflectivity, roughness, color, and surface coating of concrete have changed significantly. Bridge concrete is prone to degradation due to long-term exposure to adverse environmental conditions, and it is often more serious than other building concretes. Timely maintenance and repair are often needed to exacerbate this degradation. In order to evaluate the structural integrity and determine the degree of degradation, accurate identification is required. Specific defects are often very small and often overlap with other defect categories. We found that it is difficult to detect or distinguish the types of defects during the maintenance and repair process. Traditional contact methods for identifying concrete structure defects have poor safety, productivity, and reliability.
[0003] In recent years, researchers have turned to non-contact methods, especially image-based methods, to address these challenges. However, numerous intrusion factors on the concrete surface, such as graffiti, holes, paint, and stains, increase the complexity of the problem. Computer vision based on deep learning has shown great potential in dealing with such complex situations.
[0004] In summary, with the continuous aging of various concrete bridges, more and more bridges have defects. Therefore, there is an urgent need to develop new and effective identification methods. The above problems need to be solved urgently. For this reason, a method for multi-label defect identification of concrete based on deep learning is proposed. Summary of the Invention
[0005] The technical problem to be solved by the present invention is: how to quickly and accurately identify the type of defects, while being non-contact, capable of quickly and accurately identifying, realizing automatic classification and recognition through a machine, reducing labor costs, and providing a method for multi-label defect identification of concrete based on deep learning.
[0006] The present invention solves the above technical problems through the following technical solutions. The present invention includes the following steps:
[0007] S1: Sample preprocessing
[0008] Preprocess the multi-label defect image samples of concrete.
[0009] S2: Network construction
[0010] Replace the SE attention module in the MBConv module of the original EfficientNetV2 network with a CA attention module, and replace the adaptive pooling layer in the classification head of the original EfficientNetV2 network with an SPP layer, thereby obtaining a concrete multi-label defect recognition network;
[0011] S3: Network training
[0012] Use the training set to train the concrete multi-label defect recognition network to obtain a concrete multi-label defect recognition model;
[0013] S4: Defect recognition
[0014] Use the trained concrete multi-label defect recognition model to identify the image to be recognized, output the defect classification recognition result, and use Grad-CAM to visually verify the result.
[0015] Furthermore, in the step S1, the specific process is as follows:
[0016] S11: Obtain image samples from the concrete multi-label defect image database, and divide the obtained image samples into a training set, a validation set, and a test set at a set ratio;
[0017] S12: Sort the images in the training set according to the aspect ratio, and adjust the image size according to the average aspect ratio of each batch of images so that the area of the image remains at a set value.
[0018] Furthermore, in the step S11, the concrete defects include 6 categories, namely no defect, crack, spalling, weathering, steel bar exposure, and corrosion.
[0019] Furthermore, in the step S2, the CA attention module divides the channel attention into two vectors in different directions; then splice these two vectors and input them into a non-linear transformation. After that, divide the obtained vector into two along the spatial dimension, and use two 1×1 convolutions to make the number of channels equal to the input; finally, expand these two vectors so that they are exactly the same size as the input, and multiply their product by the original image as the attention weight.
[0020] Furthermore, in the step S2, the replaced MBConv module is an MBConv-CA module, which includes a depthwise separable convolutional layer, a CA attention layer, and a linear bottleneck layer. The depthwise separable convolutional layer, the CA attention layer, and the linear bottleneck layer are connected in sequence, and the CA attention layer is the CA attention module.
[0021] Furthermore, in the step S2, in the SPP layer, the feature map is divided into a group of sub-regions with different scales, each sub-region is pooled, and the features of all scales are connected to a feature vector of a fixed size.
[0022] Furthermore, in the step S2, the replaced classification head includes a convolutional layer, an SPP layer, and a linear classification layer, and the convolutional layer, the SPP layer, and the linear classification layer are connected in sequence.
[0023] Furthermore, in the step S3, the training process is specifically as follows:
[0024] S31: Use the training set to train the concrete multi-label defect recognition network, and use the cosine annealing learning rate scheduling strategy during the training process;
[0025] S32: Use a progressive learning method, train for a total of 300 rounds, and increase the image size and regularization strength every 100 rounds.
[0026] Furthermore, in the step S4, the specific process is as follows:
[0027] S41: The image to be recognized first undergoes a two-fold downsampling process through a 3×3 convolutional layer; then passes through 10 Fused-MBConv modules, where in the 3rd and 7th Fused-MBConv modules, a two-fold downsampling process is performed on it, and the input and output sizes of the remaining Fused-MBConv modules remain unchanged. Then it passes through thirty MBConv-CA modules, where in the 1st and 16th MBConv-CA modules, a two-fold downsampling process is performed on it; then passes through a 1×1 convolutional layer and an SPP layer, and finally judges its defect type through a linear classification layer;
[0028] S42: Use the Grad-CAM technology to perform a visualization operation on the last convolutional layer in the classification head where the defect is recognized to obtain its heat map.
[0029] The present invention has the following advantages compared with the prior art:
[0030] (1) Replace the SE attention module in the MBConv module of the original EfficientNetV2 network with a CA attention module to improve the feature extraction ability and thus enhance the model's performance. Experiments have shown that a network using the CA attention mechanism can achieve better results in various computer vision tasks. First, the CA attention mechanism can learn spatial coordinates as attention maps, enabling the model to focus more on important spatial positions and suppress irrelevant regions. Second, the CA attention mechanism can be easily combined with various deep neural network structures and has strong pluggability. Finally, the CA attention mechanism can be implemented using global pooling, convolution, and fully connected layers in different directions, with a relatively low computational cost. Therefore, it can enhance the model performance with almost no additional computational burden.
[0031] (2) Inputting the image while preserving the aspect ratio information can minimize the impact caused by an overly large aspect ratio. Since there are a large number of images with a large aspect ratio in the dataset of the present invention, and they are often adjusted to a square shape when input into the network, resulting in serious image distortion and a significant impact on defect recognition. Therefore, the present invention adopts an SPP layer that allows the network to input images with different aspect ratios for training while preserving the aspect ratio information. Experiments have proven that the improved model effectively enhances the recognition level on the dataset of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a schematic flowchart of the method for multi-label defect recognition of concrete based on deep learning in an embodiment of the present invention;
[0033] FIG. 2(a) is a schematic structural diagram of the CA attention module in an embodiment of the present invention;
[0034] FIG. 2(b) is a schematic structural diagram of the SE attention module in an embodiment of the present invention;
[0035] Figure 3 is a schematic structural diagram of the spatial pyramid pooling (SPP) layer in an embodiment of the present invention;
[0036] Figure 4 is a schematic structural diagram of the improved EfficientNetV2 network in an embodiment of the present invention;
[0037] FIG. 5(a) is a schematic structural diagram of the classification head (Head) of the original EfficientNetV2 network in an embodiment of the present invention;
[0038] FIG. 5(b) is a schematic structural diagram of the classification head (Head) of the improved EfficientNetV2 network in an embodiment of the present invention;
[0039] FIG. 6(a) is an example diagram of no defect in an embodiment of the present invention;
[0040] Figure 6(b) is an example diagram in the embodiment of the present invention where the defect types include weathering and cracks;
[0041] Figure 6(c) is an example diagram in the embodiment of the present invention where the defect types include exposed steel bars, corrosion, and spalling;
[0042] Figure 7(a) is an example diagram in the embodiment of the present invention where the defect types include exposed steel bars and corrosion;
[0043] Figure 7(b) is a heat map of the exposed steel bar defect generated by Grad-CAM in the embodiment of the present invention;
[0044] Figure 7(c) is a heat map of the corrosion defect generated by Grad-CAM in the embodiment of the present invention. Detailed implementation manners
[0045] The embodiments of the present invention will be described in detail below. These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.
[0046] This embodiment provides a technical solution: a method for multi-label defect recognition of concrete based on deep learning, including the following steps:
[0047] S1: Preprocess the multi-label defect dataset of concrete.
[0048] S2: In the MBConv module of the original EfficientNetV2 network, replace the SE attention module with the CA attention module to improve the feature extraction ability of the model, make the model pay more attention to important spatial positions, and suppress irrelevant regions.
[0049] S3: Modify the classification head of the original EfficientNetV2 network, and replace the adaptive pooling layer of the original classification head with the SPP layer.
[0050] S4: Test the results on the test set and use Grad-CAM for visual verification of the results.
[0051] In this embodiment, the content of step S1 includes the following:
[0052] S11: The dataset used in this invention is the CODEBRIM dataset released by Martin Mundt et al., which contains 7,860 concrete defect images. In addition, at least one label is provided for each image. This invention focuses on the classification and recognition of multi-labels of concrete, including Background (no defect), Crack (crack), Spallation (spalling), Efflorescence (efflorescence), Exposed Bar (exposed steel bars), and Corrosion Stain (corrosion). In this work, about 150 images are assigned to each category label for validation and about 150 images for testing. And it is ensured that the data used for the validation set and the test set do not exist in the training set and do not participate in any model training process.
[0053] S12: Sort the images in the training set according to the aspect ratio, and Resize the images according to the average aspect ratio of each batch of pictures. The area of the pictures is maintained at 300×300. The H and W of the picture size are calculated by the following formula:
[0054] H = 300 * sqrt(hw)
[0055] W = 300 / sqrt(hw)
[0056] where hw is the aspect ratio.
[0057] Furthermore, in step S2, the SE attention is replaced by CA attention in the MBConv module to improve the model's feature extraction ability and make the model pay more attention to the specific content of important spatial positions as follows:
[0058] S21: Coordinate Attention (CA) is an attention mechanism designed to more effectively obtain feature space information. Different from the traditional self-attention mechanism that calculates attention weights based on global feature vectors, CA divides the channel attention into two vectors along different directions. Then these two vectors are concatenated and input into a non-linear transformation. After that, the obtained vector is divided into two along the spatial dimension, and two 1×1 convolutions are used to make the number of channels equal to the input. Finally, these two vectors are expanded to make them exactly the same size as the input, and their product is multiplied by the original image as the attention weight.
[0059] S22: The MBConv module is an improvement based on MobileNetV2. It uses techniques such as depthwise separable convolution and channel expansion to reduce the amount of computation and the number of parameters. Specifically, the MBConv module contains a depthwise separable convolution layer, an SE attention layer, and a linear bottleneck layer. These layers are used to improve the training speed and stability of the model through techniques such as residual connection and batch normalization.
[0060] S23: Although the SE attention mechanism has been widely used in the past two years, it only considers modeling the channel relationship to re - measure the importance of each channel, ignoring the position information, which is very important for generating a spatial selection attention map. The CA attention mechanism can learn the spatial coordinates as the attention map, enabling the model to pay more attention to important spatial positions and suppressing irrelevant regions. In the present invention, the SE in EfficientNetV2 is replaced by CA, effectively improving the performance of the model. In addition, the CA attention module is integrated into the original EfficientNetV2 network and only involves a few convolutional layers, which is convenient for modification. In this embodiment, the MBConv module using the CA attention mechanism is called the MBConv - CA module.
[0061] In step S3, the spatial pyramid pooling (SPP) layer is a feature pooling strategy that can adaptively generate a fixed - size one - dimensional vector from a feature map of any size. However, the pooling layer has limited performance when dealing with objects of different sizes and lacks the ability to capture fine - grained spatial information. To solve these problems, the SPP layer is proposed. In the SPP layer, the feature map is divided into a set of sub - regions of different scales, each sub - region is pooled, and the features of all scales are concatenated into a fixed - size feature vector. This SPP layer enables the neural network to capture global and local information of different scales and is suitable for object detection and recognition tasks. The performance of the SPP layer has been verified in various datasets and benchmark tests, demonstrating its effectiveness in improving the accuracy and robustness of deep - learning models. Due to the large aspect - ratio differences in the dataset, the SPP layer adopted in the present invention can effectively reduce the influence of the aspect ratio compared with the adaptive pooling layer.
[0062] In step S3, the original classification head (Head) of the EfficientNetV2 network includes a convolutional layer, an adaptive pooling layer, and a linear classification layer. Now, the adaptive pooling layer is replaced by the SPP layer, and the output is still a vector of fixed length, which can be directly connected to the final linear classification layer for classification.
[0063] In this embodiment, in multi - label classification, the standard method is to use a sigmoid function in the output layer to determine the presence or absence of each label. However, another method is to directly connect multiple binary classification heads to the backbone. In the present invention, the performance of the two methods is compared, and it is found that directly connecting six binary classification heads to the backbone achieves better results. We connect the improved classification heads in parallel and then access them to the network, and each classification head is responsible for distinguishing one type of defect.
[0064] In this embodiment, step S3 includes the following content:
[0065] S31: Cosine Annealing Learning Rate is a learning rate scheduling strategy used to optimize neural network models. When training a neural network model, the learning rate is an important hyperparameter that determines the step size for adjusting the model's parameters in each iteration. The cosine annealing learning rate scheduling strategy gradually decreases the learning rate during training, not linearly but along a cosine curve. Specifically, the learning rate undergoes cosine annealing between the initial value and the minimum value, forming a cosine cycle. The cycle length can be updated after each epoch, each batch, or every certain number of iterations. The advantage of the cosine annealing learning rate is that, compared to traditional learning rate decay methods (such as exponential decay, polynomial decay, etc.), it can make more full use of the number of iterations during the training process, adapt to different training scenarios, reduce the oscillations and fluctuations of the learning rate decline, and thus improve the training effect and generalization ability of the model.
[0066] S32: Progressive Learning Compared with EffilentNetV1, EffilentNetv2 introduces progressive learning technology, dividing the training process into multiple stages with increasing complexity. In the first stage, a small network is trained on a dataset with low resolution and simple augmentations. In subsequent stages, the network gradually trains on more complex augmented datasets with higher resolutions. In each stage, the network is initialized with the weights learned in the previous stage and fine-tuned to learn more abstract and complex features. In this paper, we increase the resolution and regularize the hyperparameters every 100 training epochs. By using this progressive learning method, compared with previous models, EfficientNetV2 achieves more advanced performance on various image classification tasks and significantly reduces the training time.
[0067] In this embodiment, step S4 specifically includes the following content:
[0068] S41: The image to be recognized first undergoes downsampling by a factor of two through a 3×3 convolutional layer; then passes through 10 Fused-MBConv modules, where downsampling by a factor of two is performed on the 3rd and 7th Fused-MBConv modules, and the input and output sizes of the remaining Fused-MBConv modules remain unchanged. Then it passes through thirty MBConv-CA modules, where downsampling by a factor of two is performed on the 1st and 16th MBConv-CA modules; then through a 1×1 convolutional layer and an SPP layer, and finally the defect type is determined through a linear classification layer.
[0069] S42: Use the Grad-CAM technique to perform a visualization operation on the last convolutional layer in the classification head where the defect is identified, and obtain its heatmap.
[0070] Table 1 The number of various defects and the dataset division
[0071] Category Quantity Training Validation Testing No defect 2486 2186 150 150 Crack 3027 2728 150 149 Spalling 3060 2770 150 140 Weathering 2375 2085 149 141 Exposed steel bars 2693 2401 150 142 Corrosion 2683 2387 150 146
[0072] Table 1 shows the number of each category in the dataset and the division of the dataset.
[0073] Table 2 Comparison of experimental results of different attention mechanisms and different classification heads
[0074]
[0075]
[0076] Table 2 shows the comparison of experimental results of different attention mechanisms and the use of different classification heads. +C means replacing the SE attention module with the CA attention module. Through comparative experiments, it is found that the CA attention module is more suitable for the network and the dataset used in the present invention, mainly because the CA attention mechanism pays more attention to the spatial position information. +S means replacing the adaptive pooling layer in the classification head with the spatial pyramid pooling layer. Since the spatial pyramid pooling layer retains the original aspect ratio information of the picture, the picture can be trained with the original aspect ratio, reducing the picture distortion caused by adjusting the aspect ratio of the picture, and effectively improving the experimental results of the network on this dataset. +P means using the progressive learning strategy during the training process, which can not only enable the network to learn more accurate features, but also significantly reduce the training time required.
[0077] Table 3 Results of different networks on the concrete multi-label dataset
[0078] Model Accuracy Number of parameters VGG 76.3 128.8 ResNet50 78.8 23.5 DenseNet169 77.6 14.1 ShuffleNetv2 81.3 36.5 EfficientNetv2-s 82.8 21.8 MobileViT 77.2 5.6 Ours 89.5 23.2
[0079] Table 3 shows the experimental results of different networks on the concrete multi-label defect dataset. Other networks for comparison include the classic ResNet network, VGG network, DenseNet network, lightweight ShuffleNet network, MobileViT network, and the original EfficientNetV2-s network. After thoroughly comparing all the results, it can be seen that the improved EfficientNetV2 has an accuracy 6.7% higher than the original EfficientNetV2 network.
[0080] In summary, the proposed deep learning-based multi-label defect recognition method for concrete presents an improved EfficientNetV2 network. By modifying the attention module in the network, the ability of the network to extract important features is enhanced, while unnecessary features are suppressed. For example, in the present invention, the two categories of weathering and corrosion are easily confused, so an optimization scheme is proposed to significantly reduce these confounding features by modifying the attention module. In the invention, the classification head of the EfficientNetV2 network is also improved, changing from the adaptive pooling layer to the spatial pyramid pooling layer, enabling the network to train images with different original aspect ratios to reduce the negative impact brought by modifying the aspect ratio of the images. Correspondingly, a batch aspect ratio normalization method is proposed, which not only ensures the normal training of the BN layer but also retains the original aspect ratio of the images to the greatest extent. Finally, various comparative experiments also prove the effectiveness of our improvement, and the accuracy is successfully increased from 82.8% to 89.5%. The model of the present invention is also compared with other methods, demonstrating its best accuracy. At the same time, using the Grad-CAM method to visualize the heatmaps of the output feature maps, it can be seen that the extracted features almost overlap with the defects, indicating the effectiveness and accuracy of the present invention.
[0081] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for multi-label defect recognition of concrete based on deep learning, characterized in that, It includes the following steps: S1: Sample preprocessing Preprocess the concrete multi-label defect image samples; S2: Network construction Replace the SE attention module in the MBConv module of the original EfficientNetV2 network with the CA attention module, and replace the adaptive pooling layer in the classification head of the original EfficientNetV2 network with the SPP layer, thereby obtaining a concrete multi-label defect recognition network; In the step S2, the CA attention module divides the channel attention into two vectors along different directions; then splice these two vectors and input them into a non-linear transformation. After that, divide the obtained vector into two along the spatial dimension, and use two 1×1 convolutions to make the number of channels equal to the input; finally, expand these two vectors so that they are exactly the same size as the input, and multiply their product by the original image as the attention weight; S3: Network training Use the training set to train the concrete multi-label defect recognition network to obtain a concrete multi-label defect recognition model; S4: Defect recognition Use the trained concrete multi-label defect recognition model to recognize the image to be recognized, output the defect classification recognition result, and use Grad-CAM to visually verify the result.
2. The method for multi-label defect recognition of concrete based on deep learning according to claim 1, characterized in that: In the step S1, the specific process is as follows: S11: Obtain image samples from the concrete multi-label defect image database, and divide the obtained image samples into a training set, a validation set, and a test set at a set ratio; S12: Sort the images in the training set according to the aspect ratio, and adjust the image size according to the average aspect ratio of each batch of images so that the area of the image remains at a set value.
3. The method for identifying multi-label defects of concrete based on deep learning according to claim 2, wherein: In the step S11, the concrete defects include 6 categories, namely no defect, crack, spalling, weathering, steel bar exposure, and corrosion.
4. A method for identifying multi-label defects of concrete based on deep learning according to claim 3, characterized in that: In the step S2, the replaced MBConv module is the MBConv-CA module, which includes a depthwise separable convolution layer, a CA attention layer, and a linear bottleneck layer. The depthwise separable convolution layer, the CA attention layer, and the linear bottleneck layer are connected in sequence, and the CA attention layer is the CA attention module.
5. The method for identifying multi-label defects of concrete based on deep learning according to claim 4, characterized in that: In the step S2, in the SPP layer, the feature map is divided into a group of sub-regions with different scales, each sub-region is pooled, and the features of all scales are connected to a feature vector with a fixed size.
6. The method for identifying multiple - label defects of concrete based on deep learning according to claim 5, wherein: In the step S2, the replaced classification head includes a convolution layer, an SPP layer, and a linear classification layer. The convolution layer, the SPP layer, and the linear classification layer are connected in sequence.
7. A method for identifying multi-label defects of concrete based on deep learning according to claim 6, characterized in that: In the step S3, the training process is as follows: S31: Use the training set to train the concrete multi-label defect recognition network, and use the cosine annealing learning rate scheduling strategy during the training process; S32: Use a progressive learning method, and train for a total of 300 rounds, and increase the image size and regularization strength every 100 rounds.
8. A method for identifying multi-label defects of concrete based on deep learning according to claim 7, characterized in that: In the step S4, the specific process is as follows: S41: The image to be recognized first undergoes downsampling by a factor of two through a 3×3 convolutional layer; then passes through 10 Fused-MBConv modules, where downsampling by a factor of two is performed on the 3rd and 7th Fused-MBConv modules, and the input and output sizes of the remaining Fused-MBConv modules remain unchanged. Then it passes through thirty MBConv-CA modules, where downsampling by a factor of two is performed on the 1st and 16th MBConv-CA modules; then through a 1×1 convolutional layer and an SPP layer, and finally the defect type is determined through a linear classification layer; S42: Using the Grad-CAM technique, a visualization operation is performed on the last convolutional layer in the classification head where the defect is recognized to obtain its heatmap.
Citation Information
Patent Citations
Brain tumor classification method based on OfficientNetV2 network
CN115019098A
Fabric defect detection method based on improved OfficientDet model
CN115601610A