A crack detection method and system based on deep learning

By introducing the SA-Net attention module and the Focal loss and Dice loss fusion loss function with dynamic compensation weights in the deep learning model, the crack detection algorithm is improved, and the problems of low efficiency and low accuracy of traditional methods are solved, achieving more efficient and accurate crack detection.

CN116416244BActive Publication Date: 2025-07-29SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310492746.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2025-07-29
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

Traditional crack detection methods are inefficient, low accuracy and high limitations, and have high data costs, making it difficult to achieve efficient and timely crack detection.

Method used

Using deep learning-based crack detection method, the SA-Deeplab V3+ model is used to combine the SA-Net attention module and the fusion loss function of the Focal loss and Dice loss with dynamic compensation weights, and the DWF+SA-Deeplab V3+ algorithm is constructed to improve the attention and detection accuracy of crack pixels.

Benefits of technology

It improves the accuracy and efficiency of crack detection, enhances the generalization ability of the algorithm, can better identify cracks and generate more complete and detailed predictive images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116416244B_ABST
    Figure CN116416244B_ABST
Patent Text Reader

Abstract

The present invention provides a crack detection method and system based on deep learning, which relates to the technical field of crack detection. It includes obtaining the original crack image; using the Deeplab V3+ model including an encoder module and a decoder module as the basic model, where the encoder module includes a backbone feature extraction network and a pyramid part. The pyramid part of the encoder module is fused with the SA-Net attention module, and at the same time, the convolutional layer after fusing the shallow features and deep features of the decoder is replaced with a depthwise separable convolution to build the SA-Deeplab V3+ model; the original crack image is input into the SA-Deeplab V3+ model for feature extraction and a crack prediction image is obtained. The present invention has a higher detection accuracy, higher efficiency and stronger generalization ability for cracks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of crack detection, and particularly relates to a crack detection method and system based on deep learning. Background Technique

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Cracks are one of the common disasters, existing not only in buildings but also in highways and airport roads. In buildings, general cracks have no impact on the use of the building, but if the width of the crack exceeds a certain limit, it will become a harmful crack, and the existence of harmful cracks will seriously affect the service life of the building. Similarly, on highways and airport roads, small cracks do not affect the use of highways and airport roads, but if the road surface cracks are not repaired in time, it is very likely that the road surface damage will be further aggravated due to repeated loads and adverse weather conditions, and even structural damage may occur, further leading to accidents. Therefore, it is very necessary to detect cracks more accurately, timely and efficiently.

[0004] The inventors found that although traditional detection methods can detect cracks, they have low efficiency, large limitations, high data costs, and low accuracy. Summary of the Invention

[0005] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a crack detection method and system based on deep learning, which makes the algorithm more focused on crack pixels, and has higher detection accuracy, higher efficiency, and stronger generalization ability for cracks.

[0006] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0007] The first aspect of the present invention provides a crack detection method based on deep learning.

[0008] A crack detection method based on deep learning includes the following steps:

[0009] Obtain the original crack image;

[0010] Based on the Deeplab V3+ model that includes an encoder module and a decoder module, where the encoder module consists of a backbone feature extraction network and a pyramid part, the pyramid part of the encoder module is fused with the SA-Net attention module, and at the same time, the convolutional layer after fusing the shallow features and deep features in the decoder network is replaced with a depthwise separable convolution to build the SA-Deeplab V3+ model. Here, the shallow features are the features extracted after a relatively small number of convolutions in the network, and the deep features are the features extracted after a relatively large number of convolutions in the network;

[0011] Input the original crack image into the SA-Deeplab V3+ model for feature extraction and obtain the crack prediction image.

[0012] Furthermore, input the original crack image into the SA-Deeplab V3+ model for feature extraction and obtain the crack prediction image, specifically:

[0013] Input the original crack image into the backbone feature extraction network of the encoder, extract shallow features after 3 - 5 convolutions in the backbone feature extraction network respectively, and extract deep features after more than 14 convolutions in the backbone feature extraction network;

[0014] Input the deep features into the pyramid part, perform parallel sampling on the deep features using atrous convolutions with different sampling rates to capture the context of the image at multiple scales and obtain the feature map after parallel sampling;

[0015] Use the SA-Net attention module to assign attention weights to the feature map after parallel sampling, and weight the attention weights with the corresponding feature maps to obtain the deep features after feature weighting;

[0016] Upsample the deep features after feature weighting, jointly input the upsampled result and the shallow features into the decoder for stacking, and perform depthwise separable convolution on the stacked features to obtain the effective feature map;

[0017] Upsample the effective feature map to obtain the crack prediction image.

[0018] Furthermore, use the SA-Net attention module to assign attention weights to the feature map after parallel sampling, and weight the attention weights with the corresponding feature maps to obtain the deep features after feature weighting, specifically:

[0019] SA-Net first divides the feature map after parallel sampling into G groups to obtain G sub-features. Each sub-feature is divided into two branches along the channel dimension. One branch is used to generate the spatial attention map, and the other branch is used to generate the channel attention map. Each sub-feature is captured during the training process, and the SA-Net attention module generates the corresponding weight coefficients for each sub-feature;

[0020] Each sub-feature is weighted with the corresponding weight coefficient to obtain the weighted sub-feature;

[0021] Using the shuffle mechanism, each weighted sub-feature flows in the channel dimension, and finally all weighted sub-features are integrated in the channel dimension to obtain the processed overall feature, that is, the deep feature.

[0022] Furthermore, the depthwise separable convolution consists of a depthwise convolution and a pointwise convolution:

[0023] In the depthwise convolution, one convolution kernel is responsible for one channel height, and one channel is only convolved by one convolution kernel. The number of channels of the feature map generated by this process is exactly the same as the number of input channels;

[0024] The pointwise convolution only performs weighted combination in the channel direction, and the number of feature maps generated for the effective feature map is determined by the number of convolution kernels.

[0025] Furthermore, the loss function of the SA-Deeplab V3+ model adopts Focal loss and Dice loss. The calculation formula of Focal loss is:

[0026]

[0027] y and y' respectively refer to the label value and the predicted value of the image; a is the balance factor; γ is the adjustment factor; the calculation formula of Dice loss is:

[0028]

[0029] y i and y i ′ respectively refer to the label value and the predicted value of the image; N refers to the total number of pixels in the image.

[0030] Furthermore, dynamic compensation weights are fused into Focal loss:

[0031]

[0032] Among them, y and y' respectively refer to the label value and the predicted value of the image, and β1 and β2 are dynamic compensation weight coefficients.

[0033] Further, β1 and β2 are calculated by the following formula:

[0034]

[0035]

[0036] where F p is the false positive, F n is the false negative, and P is the total number of crack pixel points in the image; S n is the total number of pixels in the image, is the percentage of crack pixels in the entire image, is the percentage of non-crack pixels in the entire image.

[0037] The second aspect of the present invention provides a crack detection system based on deep learning.

[0038] A crack detection system based on deep learning includes:

[0039] An image acquisition module configured to acquire an original crack image;

[0040] A model construction module configured to use the Deeplab V3+ model including an encoder module and a decoder module as the basic model, where the encoder module includes a backbone feature extraction network and a pyramid part, fuse the pyramid part of the encoder module with the SA-Net attention module, and at the same time replace the convolutional layer after fusing the shallow features and deep features in the decoder network with a depthwise separable convolution to construct the SA-Deeplab V3+ model, where the shallow features are the features extracted by fewer convolutional operations in the network, and the deep features are the features extracted by more convolutional operations in the network;

[0041] A crack detection module configured to input the original crack image into the SA-Deeplab V3+ model for feature extraction and obtain a crack prediction image.

[0042] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the crack detection method based on deep learning as described in the first aspect of the present invention.

[0043] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the crack detection method based on deep learning as described in the first aspect of the present invention.

[0044] The above one or more technical solutions have the following beneficial effects:

[0045] Based on SA-Deeplab V3+, the present invention uses a combined loss function that fuses Focal loss and Dice loss after integrating dynamic compensation weights, constructs the overall algorithm of DWF+SA-Deeplab V3+, and conducts data training and testing based on this algorithm, achieving good crack extraction results.

[0046] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Brief Description of the Drawings

[0047] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0048] Figure 1 It is the overall architecture diagram of SA-Deeplab V3+ for the first embodiment.

[0049] Figure 2 It is the structural diagram of the SA-Net attention module for the first embodiment.

[0050] Figure 3 It is the schematic diagram of the ordinary convolution structure for the first embodiment.

[0051] Figure 4 It is the schematic diagram of the per-channel convolution structure for the first embodiment.

[0052] Figure 5 It is the schematic diagram of the pointwise convolution structure for the first embodiment.

[0053] Figure 6(a) is the original crack picture for the first embodiment.

[0054] Figure 6(b) is the crack label map for the first embodiment.

[0055] Figure 6(c) is the detection result of Deeplab V3+ for the first embodiment.

[0056] Figure 6(d) is the detection of DWF+SA-Deeplab V3+ for the first embodiment.

[0057] Figure 7 It is the system structure diagram for the second embodiment. Detailed Description of the Invention

[0058] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0059] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0060] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0061] The core structure of the original Deeplab V3+ neural network consists of two modules, namely the spatial pyramid pooling module and the encoder-decoder module.

[0062] The specific feature extraction process is as follows:

[0063] Input the image into the backbone feature extraction network, obtain shallow features and deep features through the backbone feature extraction network, input the deep features into the atrous spatial pyramid pooling module (ASPP), perform convolution and pooling with four atrous convolutional layers and one pooling layer respectively to obtain five feature layers, then splice these feature maps and send them into a 1x1 convolutional layer for operation, and then obtain a new feature layer A after upsampling.

[0064] The shallow features are directly reduced in dimension through a 1x1 convolution to obtain the reduced-dimension feature B, then spliced with A, and finally output the prediction result with the same size as the original image through a 3x3 convolutional layer and upsampling.

[0065] In the present invention, shallow features are extracted after fewer convolutions of the backbone feature extraction network, and deep features are extracted after convolutions of the entire backbone feature network.

[0066] The overall concept of the present invention:

[0067] Although traditional detection methods can perform crack detection, they have low efficiency, large limitations, high data costs, and low accuracy. With the development of computer computing power, the application of deep learning methods in crack detection has gradually increased. At the same time, they have higher accuracy, higher efficiency, and strong generalization ability. The present invention proposes a crack detection algorithm based on deep learning according to this situation.

[0068] The present invention mainly proposes a new and more accurate crack detection algorithm based on the existing deep learning network framework. The specific improvement contents can be summarized as the following aspects:

[0069] 1. Add the Sa-Net attention mechanism to the pyramid module of the original Deeplab V3+ neural network.

[0070] 2. Use a loss function that combines Dice loss and Focal loss to reduce the interference of loss calculation caused by the small proportion of crack foreground in the image and improve the network's attention to cracks.

[0071] 3. A dynamic compensation weight for a new loss function is proposed, which can be used in combination with loss functions such as the focal loss function and the cross-entropy loss function to compensate for the problem that the proportion of crack pixels in the whole image is too small during the crack detection process. The dynamic compensation weight is fused with the Focal loss and used in combination with the Dice loss to replace the original cross-entropy loss function, and together with the SA-Deeplab V3+ to form the DWF+SA-Deeplab V3+ algorithm model.

[0072] The algorithm improvement of the present invention is mainly to make the algorithm more focused on crack pixels, and good results have been achieved in actual experiments after training on data.

[0073] Embodiment 1

[0074] This embodiment discloses a crack detection method based on deep learning.

[0075] As Figure 1 shown, a crack detection method based on deep learning includes the following steps:

[0076] Obtain the original crack image;

[0077] Based on the Deeplab V3+ model including an encoder module and a decoder module, where the encoder module includes a backbone feature extraction network and a pyramid part, fuse the SA-Net attention module into the pyramid part of the encoder module, and at the same time replace the convolutional layer after fusing the shallow features and deep features in the decoder network with a depthwise separable convolution to build the SA-Deeplab V3+ model, where the shallow features are the features extracted after 3-5 times of convolution in the network, and the deep features are the features extracted after more than 14 times of convolution in the network;

[0078] Input the original crack image into the SA-Deeplab V3+ model for feature extraction and obtain the crack prediction image.

[0079] Further, inputting the original crack image into the SA-Deeplab V3+ model for feature extraction and obtaining the crack prediction image specifically includes:

[0080] Input the original crack image into the backbone feature extraction network of the encoder, extract shallow features after passing through the backbone feature extraction network with fewer times of convolution, and extract deep features after passing through the backbone feature extraction network with more times of convolution;

[0081] Input the deep features into the pyramid part, and perform parallel sampling on the deep features with dilated convolutions at different sampling rates to capture the context of the image at multiple scales and obtain the feature map after parallel sampling;

[0082] The SA-Net attention module is used to assign attention weights to the feature maps after parallel sampling, and the attention weights are weighted with the corresponding feature maps to obtain the deep features after feature weighting.

[0083] The deep features after feature weighting are upsampled, and the upsampled result and the shallow features are jointly input into the decoder for stacking, and the stacked features are subjected to depthwise separable convolution to obtain the effective feature maps.

[0084] The effective feature maps are upsampled to obtain the crack prediction image.

[0085] In this embodiment, the shallow features are the features extracted by 3-5 times of convolution in the network, and the deep features are the features extracted by more than 14 times of convolution in the network.

[0086] Specifically:

[0087] (1) The overall network framework SA-Deeplab V3+

[0088] In order to reduce the information loss caused by upsampling, the 3x3 convolution layer after fusing the shallow features and the deep features in the decoder part of the Deeplab V3+ model is replaced with depthwise separable convolution. At the same time, in order to make the algorithm pay more attention to the crack pixels, we fuse the SA-Net attention module in the pyramid part (ASPP) of the model. The specific network structure is as Figure 1 shown.

[0089] The core structure of SA-Deeplab V3+ consists of two parts: an encoder module and a decoder module.

[0090] Among them, the backbone network of the encoder part of the model is MobileNetV2. Different from the common ResNet residual structure, MobileNetV2 first performs dimensionality increase on the input feature matrix through 1x1 convolution to increase the size of the channel, then performs convolution processing through a 3x3 depth convolution kernel (DW convolution), and finally performs dimensionality reduction through a 1x1 convolution kernel. At the same time, MobileNetV2 uses the Rectified Linear Unit 6 (ReLU6) activation function. When the input value is less than 0, it is default to zero, and when the input value is between [0, 6], it remains unchanged, but when the input value is greater than 6, the output value will be set to 6.

[0091] Meanwhile, for the spatial feature pyramid part in the encoder module, we integrated SA-Net. For a given feature map, SA-Net first divides the given feature map into G groups, and then captures each sub-feature during the training process, generating corresponding weight coefficients for each sub-feature through the attention module. Specifically, each sub-feature is divided into two branches along the channel dimension. One branch is used to generate the spatial attention map, and the other branch is used to generate the channel attention map, enabling the model to focus more on meaningful parts. For the implementation of the channel attention mechanism, SA-Net does not use the traditional Squeeze Excitation (SE). Instead, it first performs global pooling, then moves and scales the channel vector through a pair of parameters, and finally activates the value with sigmoid. SA-Net also adds a shuffling mechanism, allowing each group of information to flow in the channel dimension and finally integrating the grouped information in the channel dimension to obtain the processed overall feature. The structure of the SA-Net attention module is as shown in Figure 2 shown.

[0092] For the decoder part, we replaced ordinary convolutions with depthwise separable convolutions. Different from the ordinary convolution method, the depthwise separable convolution consists of a depthwise convolution and a pointwise convolution. In the depthwise convolution, one convolution kernel is responsible for one channel height, and one channel is only convolved by one convolution kernel. The number of channels of the feature map generated in this process is exactly the same as the number of input channels. The pointwise convolution only performs weighted combination in the channel direction to generate a new feature map, and the number of generated feature maps is determined by the number of convolution kernels. Therefore, compared with ordinary convolutions, the depthwise separable convolution has fewer parameters and lower computational costs, and less loss during the feature extraction process. The schematic diagrams of the depthwise convolution and the pointwise convolution structures in ordinary convolution and depthwise separable convolution are respectively as shown in Figure 3 , 4 , and 5.

[0093] (2) Improvement of the loss function

[0094] Through the observation and analysis of a large number of crack images, we found that when using the semantic segmentation method to detect cracks, different from traditional semantic segmentation, cracks often only account for a very small part of the image. Therefore, for loss calculation, if the cross-entropy loss function is used, training interference is often caused due to the excessive proportion of the background.

[0095] For this situation, the focal loss function can, to a certain extent, solve the problem of foreground-background imbalance, increasing the weight of a small number of target categories and the weight of misclassified samples. The focal loss function is a modification based on the cross-entropy loss function. The calculation formula of the binary cross-entropy loss function is as follows:

[0096]

[0097]

[0098] y and y' respectively refer to the label value and the predicted value of the image; the label value of the image represents the actual situation of whether the image is a crack. If it is a crack, the value is 1, and if it is not a crack, the value is 0; the predicted value of the image represents the value output by the detection algorithm proposed in this paper.

[0099] Focal loss is a modification based on the cross-entropy loss function. A factor γ is added on the original basis. γ>0 reduces the loss of easy-to-classify samples, making the model pay more attention to difficult and misclassified samples. At the same time, changing the size of γ can also adjust the speed at which the weights of simple samples are reduced. When γ increases, the influence of the adjustment factor increases. In addition, a balancing factor a is added to balance the imbalance in the proportion of positive and negative samples themselves. The calculation formula of Focal loss is as follows:

[0100]

[0101] y and y' respectively refer to the label value and the predicted value of the image.

[0102] At the same time, the present invention also introduces Dice loss, which is applicable to binary segmentation of images and can, to a certain extent, alleviate the problem of imbalance in the number of positive and negative samples. For predictions with higher confidence, a lower Dice coefficient will be obtained, resulting in a smaller Dice loss. For predictions with lower confidence, a higher Dice coefficient will be obtained, resulting in a larger Dice loss. The calculation formula of Dice loss is as follows:

[0103]

[0104] y i and y i ' respectively refer to the label value and the predicted value of the image, and N refers to the total number of pixels in the image.

[0105] Using Dice loss alone will have a problem of loss saturation. Therefore, this patent uses Focal loss and Dice loss in combination.

[0106] In addition, the present invention proposes a dynamic compensation weight based on the F1 score. Mainly based on the ratios of false positives, false negatives to true positives during the model training process as weight coefficients, so as to perform loss compensation when the model prediction results are poor. The calculation formula of Focal loss after fusing this dynamic compensation weight is as follows:

[0107]

[0108] Similarly, y and y' represent the label value and the predicted value of the image respectively, and β1 and β2 are dynamic compensation weight coefficients.

[0109] β1 and β2 can be calculated by the following formula:

[0110]

[0111]

[0112] where F p refers to false positives, F n refers to false negatives, and P refers to the total number of crack pixels in the image. When there are more predicted false positives, the corresponding coefficient increases, making the prediction loss for the corresponding real crack pixels increase. When there are more predicted false negatives, that is, when the number of pixels misjudged as non-crack pixels during the prediction process is large, the loss of the corresponding background pixels increases. When the final prediction result gradually improves, β1 gradually approaches 0. S n is the total number of pixels in the image, is the percentage of crack pixels in the whole image. is the percentage of non-crack pixels in the whole image. Since the percentage of non-crack pixels in the image is large, this coefficient can compensate for the problem that the prediction loss accounts for a small proportion due to the small proportion of crack pixels in the whole image.

[0113] (III) Overall algorithm structure DWF+SA-Deeplab V3+

[0114] This patent uses a combined loss function that fuses Focal loss and Dice loss after integrating dynamic compensation weights on the basis of SA-Deeplab V3+ to construct the overall algorithm of DWF+SA-Deeplab V3+. Based on this algorithm, data training and testing are carried out, and good crack extraction results are obtained. The original images, label maps, detection results of Deeplab V3+, and detection results of the DWF+SA-Deeplab V3+ algorithm on the public dataset CrackForest are shown in Figures 6(a), 6(b), 6(c), and 6(d).

[0115] The proposed DWF+SA-Deeplab V3+ algorithm of the present invention can effectively achieve crack detection. We trained, validated, and compared the original algorithm and the DWF+SA-Deeplab V3+ algorithm on the entire CrackForest dataset, and concluded that the crack detection effect of the algorithm proposed in the present invention is better, the crack detection results are more complete, more detailed, closer to the label image, the F1 score has increased by 0.057, and the mean intersection over union (MIOU) has increased by 0.07. Good results have been achieved in actual crack detection, and it can well handle the crack detection task. The quantitative results comparison of the two algorithms is shown in Table 1.

[0116] Table 1 Comparison of detection results of two algorithms

[0117]

[0118] Example 2

[0119] This embodiment discloses a crack detection system based on deep learning.

[0120] As Figure 7 shown, a crack detection system based on deep learning includes:

[0121] An image acquisition module, configured to: acquire an original crack image;

[0122] A model construction module, configured to: use the Deeplab V3+ model including an encoder module and a decoder module as the basic model, where the encoder module includes a backbone feature extraction network and a pyramid part, fuse the SA-Net attention module to the pyramid part of the encoder module, and at the same time replace the convolutional layer after fusing the shallow features and deep features in the decoder network with a depthwise separable convolution to construct the SA-Deeplab V3+ model, where the shallow features are the features extracted by fewer convolutional operations in the network, and the deep features are the features extracted by more convolutional operations in the network;

[0123] A crack detection module, configured to: input the original crack image into the SA-Deeplab V3+ model, perform feature extraction, and obtain a crack prediction image.

[0124] Example 3

[0125] The purpose of this embodiment is to provide a computer-readable storage medium.

[0126] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the deep learning-based crack detection method as described in Embodiment 1 of the present disclosure.

[0127] Example 4

[0128] The objective of this example is to provide an electronic device.

[0129] An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the steps in the deep learning-based crack detection method as described in Embodiment 1 of the present disclosure.

[0130] The steps involved in the devices of the above Embodiments 2, 3, and 4 correspond to those in Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0131] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device for execution by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0132] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.

Claims

1. A crack detection method based on deep learning, characterized in that, Including the following steps: Obtain the original crack image; Based on the Deeplab V3+ model including an encoder module and a decoder module, where the encoder module includes a backbone feature extraction network and a pyramid part, fuse the SA-Net attention module into the pyramid part of the encoder module, and at the same time replace the convolutional layer after fusing the shallow features and deep features in the decoder network with depthwise separable convolution to build the SA-Deeplab V3+ model. Here, the shallow features are the features extracted after fewer convolutional operations in the network, and the deep features are the features extracted after more convolutional operations in the network; Input the original crack image into the SA-Deeplab V3+ model for feature extraction and obtain the crack prediction image; Input the original crack image into the SA-Deeplab V3+ model for feature extraction and obtain the crack prediction image, specifically: Input the original crack image into the backbone feature extraction network of the encoder, extract shallow features after fewer convolutional operations of the backbone feature extraction network, and extract deep features after more convolutional operations of the backbone feature extraction network; Input the deep features into the pyramid part, perform parallel sampling on the deep features using dilated convolutions with different sampling rates to capture the context of the image at multiple scales and obtain the feature map after parallel sampling; Use the SA-Net attention module to assign attention weights to the feature map after parallel sampling, and weight the attention weights with the corresponding feature map to obtain the deep features after feature weighting; Upsample the deep features after feature weighting, jointly input the upsampled result and the shallow features into the decoder for stacking, and perform depthwise separable convolution on the stacked features to obtain the effective feature map; Upsample the effective feature map to obtain the crack prediction image; Use the SA-Net attention module to assign attention weights to the feature map after parallel sampling, and weight the attention weights with the corresponding feature map to obtain the deep features after feature weighting, specifically: SA-Net first divides the feature map after parallel sampling into G groups to obtain G sub-features. Each sub-feature is divided into two branches along the channel dimension. One branch is used to generate the spatial attention map, and the other branch is used to generate the channel attention map. Each sub-feature is captured during the training process, and the SA-Net attention module generates the corresponding weight coefficients for each sub-feature; Weight each sub-feature with the corresponding weight coefficient to obtain the weighted sub-feature; Use the shuffle mechanism to make each weighted sub-feature flow in the channel dimension and finally integrate all the weighted sub-features in the channel dimension to obtain the processed overall feature, that is, the deep feature.

2. The deep learning-based crack detection method according to claim 1, wherein The depthwise separable convolution consists of a depthwise convolution and a pointwise convolution: In per-channel convolution, one convolutional kernel is responsible for one channel height, and one channel is only convolved by one convolutional kernel. The number of channels of the feature map generated in this process is exactly the same as the number of channels of the input; Pointwise convolution only performs weighted combination in the channel direction to generate an effective feature map, and the number of generated effective feature maps is determined by the number of convolutional kernels.

3. The crack detection method based on deep learning according to claim 1, characterized in that, The loss function of the SA-DeeplabV3+ model adopts Focal loss and Dice loss, where the calculation formula of Focal loss is: ; and respectively refer to the tag value and predicted value of the image; a is the balance factor; is the adjustment factor; The calculation formula of Dice loss is: ; and respectively refer to the label value and the predicted value of the image; N refers to the total number of pixels in the image.

4. The crack detection method based on deep learning according to claim 3, characterized in that Fuse the dynamic compensation weight for Focal loss: ; Among them, and respectively refer to the label value and the predicted value of the image, and are dynamic compensation weight coefficients.

5. The crack detection method based on deep learning according to claim 4, wherein and is calculated by the following formula: ; ; Among them, is a false positive, is a false negative, and P is the total number of crack pixel points in the image; is the total number of pixels in the image, is the percentage of crack pixels in the entire image, is the percentage of non-crack pixels in the entire image.

6. A crack detection system based on deep learning, characterized in that: Including: An image acquisition module, configured to: acquire the original crack image; A model construction module, configured to: based on the Deeplab V3+ model including an encoder module and a decoder module as the basic model, where the encoder module includes a backbone feature extraction network and a pyramid part, fuse the SA-Net attention module into the pyramid part of the encoder module, and at the same time replace the convolutional layer after fusing the shallow features and deep features in the decoder network with a depthwise separable convolution to construct the SA-Deeplab V3+ model, where the shallow features are the features extracted after a small number of convolutions in the network, and the deep features are the features extracted after a large number of convolutions in the network; A crack detection module, configured to: input the original crack image into the SA-Deeplab V3+ model for feature extraction and obtain a crack prediction image; Input the original crack image into the SA-Deeplab V3+ model for feature extraction and obtain a crack prediction image, specifically: Input the original crack image into the backbone feature extraction network of the encoder, extract shallow features after a small number of convolutions in the backbone feature extraction network respectively, and extract deep features after a large number of convolutions in the backbone feature extraction network; Input the deep features into the pyramid part, perform parallel sampling on the deep features with dilated convolutions at different sampling rates to capture the context of the image at multiple scales and obtain the feature map after parallel sampling; Use the SA-Net attention module to assign attention weights to the feature map after parallel sampling, and weight the attention weights with the corresponding feature map to obtain the deep features after feature weighting; Upsample the deep features after feature weighting, jointly input the upsampled result and the shallow features into the decoder for stacking, and perform depthwise separable convolution on the stacked features to obtain an effective feature map; Upsample the effective feature map to obtain a crack prediction image; Use the SA-Net attention module to assign attention weights to the feature map after parallel sampling, and weight the attention weights with the corresponding feature map to obtain the deep features after feature weighting, specifically: SA-Net first divides the feature map after parallel sampling into G groups to obtain G sub-features. Each sub-feature is divided into two branches along the channel dimension. One branch is used to generate a spatial attention map, and the other branch is used to generate a channel attention map. Each sub-feature is captured during the training process, and the SA-Net attention module generates corresponding weight coefficients for each sub-feature; Each sub-feature is weighted with the corresponding weight coefficient to obtain a weighted sub-feature; Using the shuffle mechanism, each weighted sub-feature flows in the channel dimension, and finally all weighted sub-features are integrated in the channel dimension to obtain the processed overall feature, that is, the deep feature.

7. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the deep learning-based crack detection method according to any one of claims 1-5.

8. An electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the deep learning-based crack detection method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Improved urban streetscape image segmentation method based on deep learning

    CN115035299A

  • Nematode image segmentation method and system based on deep learning and iterative feature fusion

    CN115830316A