Road crack detection method based on deep learning

By adopting an improved U-Net network in road crack detection, using edge refinement modules and multi-scale fusion modules based on attention mechanisms, the problem of missed detection and missed detection in the prior art is solved, and higher detection accuracy and continuity are achieved.

CN115035065BActive Publication Date: 2025-05-02CHANGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210660658.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-13
Publication Date
2025-05-02
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

The prior art has missed detection and misdetection in road crack detection, especially when the crack distribution is messy and irregular, the shape and size are not fixed, the topological structure is complex, and there are many small cracks.

Method used

Using an improved U-Net network, the accuracy and continuity of crack detection is improved by replacing the traditional bilayer convolutional structure with edge refinement modules in the encoding section and designing multi-scale fusion modules and fusion optimization modules based on attention mechanisms in the decoding section.

Benefits of technology

It effectively reduces the mis-checking phenomenon of road cracks, improves the ability to extract crack details information, and enhances the integrity and continuity of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035065B_ABST
    Figure CN115035065B_ABST
Patent Text Reader

Abstract

The invention discloses a road crack detection method based on deep learning, comprising: obtaining a plurality of road crack images, dividing the plurality of road crack images into a training set, a verification set and a test set; building a U-Net network, wherein the U-Net network has an encoding part and a decoding part, each of which has 5 layers; replacing the traditional double-layer convolution structure in the encoding part with an edge refinement module, each layer including 3 edge refinement modules, designing a multi-scale fusion module based on an attention mechanism at the bottom of the U-Net network, and designing fusion optimization modules at the 2nd, 3rd and 4th layers of the decoding part respectively, to obtain an improved U-Net network; loading the training set and the verification set into the improved U-Net network for training and verification, and saving the model with the best effect; using the model with the best effect to test the road crack images in the test set to obtain a test result. It can reduce the phenomenon of missed detection and false detection of road cracks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a road defect detection method, and in particular to a road crack detection method based on deep learning. Background Art

[0002] Cracks are the most common and most harmful type of pavement disease. While affecting the appearance of the pavement, they can also cause traffic accidents and shorten the service life of the road. Therefore, it is crucial to detect and repair road cracks in a timely manner. Manual crack detection methods rely entirely on the experience of the inspectors, and have the disadvantages of low efficiency, subjective evaluation results, high cost, and high risk. Automated crack detection can reduce costs, improve detection efficiency, and reduce the rate of missed reports.

[0003] The current crack detection methods mainly include crack detection methods based on digital image processing and crack detection methods based on deep learning. Traditional crack detection methods include threshold segmentation, Gabor filter, histogram, random structure forest, etc. Although these methods improve the detection efficiency compared with manual detection, they have high requirements for the quality of the data set, are easily affected by external environments such as light and water stains, and perform poorly on data sets with a lot of noise. In recent years, with the development of artificial intelligence, deep learning methods have begun to be applied to the task of image crack detection.

[0004] In the crack detection task, although the deep learning method improves the accuracy of detection compared with the traditional method, the integrity and continuity of the cracks need to be further improved. On the one hand, the cracks are distributed irregularly, with irregular shapes and sizes, and it is difficult for the current crack detection methods to ensure the continuity of the cracks. On the other hand, the cracks have complex topological structures and many small cracks. Many small cracks are easily affected by noise, resulting in missed detection. Summary of the invention

[0005] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a road crack detection method based on deep learning that can reduce the phenomenon of missed detection and false detection of road cracks.

[0006] In order to solve the above technical problems, the technical solution of the present invention is: a road crack detection method based on deep learning, comprising:

[0007] Acquire a plurality of road crack images, and divide the plurality of road crack images into a training set, a validation set, and a test set;

[0008] Building a U-Net network, wherein the U-Net network has an encoding part and a decoding part, and the encoding part and the decoding part each have 5 layers;

[0009] The traditional double-layer convolution structure in the encoding part is replaced by an edge refinement module, each layer includes three edge refinement modules, a multi-scale fusion module based on an attention mechanism is designed at the bottom of the U-Net network, and fusion optimization modules are designed at the 2nd, 3rd and 4th layers of the decoding part respectively, to obtain an improved U-Net network;

[0010] Loading the training set and the validation set into the improved U-Net network for training and validation, and saving the model with the best effect;

[0011] The road crack images in the test set are tested using the model with the best effect to obtain test results. Further, before dividing the plurality of road crack images into a training set, a validation set and a test set, the method further includes:

[0012] The road crack images are cropped into a uniform size.

[0013] Furthermore, the working method of each edge refinement module includes:

[0014] Step A1: Input the feature x∈R of the edge refinement module H×W×C After 1×1 convolution, it is evenly divided into n feature subsets x i , where the i-th subset x i ∈{1,2,...,n}, each subset x i The number of channels is C / n;

[0015] x i ∈{2,3,...,n} after the corresponding 3×3 convolution, the output is y i ∈{1,2,...,n}:

[0016]

[0017] Where C refers to the number of channels of the features input to the edge refinement module, and Conv(·) represents a convolution operation with a convolution kernel of 3×3.

[0018] Step A2: i ∈{1,2,...,n} is combined and restored to the original number of channels through 1X1 convolution, and the feature y∈R is output H×W×C ;

[0019] Step A3: Output features y∈R H×W×C After the channel attention CAM module, the output feature y∈R H×W×C The following processing is performed in the channel attention CAM module:

[0020] First, global features are aggregated through global average pooling;

[0021]

[0022] Then the convolution operation adjusts the channel weights;

[0023] W=σ(Con'(y avg )) (3)

[0024] Finally, the weight W is combined with the feature y∈R of the input channel attention CAM module H×W×C multiply;

[0025] Among them, y i,j ∈R C is the full channel feature, Con'(·) represents the one-dimensional convolution of size K, and σ represents the Sigmoid activation function;

[0026] Step A4: Connect the output features of the channel attention CAM module to the original input features x∈R of the edge refinement module through residual connection H×W×C To perform the fusion:

[0027] x=W·y+x (4).

[0028] Furthermore, the working method of the multi-scale fusion module based on the attention mechanism includes:

[0029] Step B1: The feature maps output by the first two coding layers of the coding part are respectively subjected to 1×1 convolution operation to transform the channels and then pooled to obtain feature maps with the same scale and number of channels, and the two feature maps with the same scale and number of channels are fused to obtain a fused feature map;

[0030] f1'=w(f(f1)) (5)

[0031] f2'=w(f(f2)) (6)

[0032] f 12 =Cat(f1',f2') (7)

[0033] Among them, f1 and f2 represent the outputs of the first two encoding layers respectively, f(·) represents the convolution operation with a 1×1 convolution kernel, w(·) represents the pooling operation, and Cat(·) represents the superposition of features in the channel dimension;

[0034] Step B2: Fuse the fused feature map with the feature map output by the last encoding layer of the encoder, and finally output the multi-scale fused feature map f∈R H×W×C :

[0035] f=Cat(f 12 ,f5) (8)

[0036] Among them, f5 represents the output of the last encoding layer;

[0037] Step B3: The output multi-scale fusion feature map is subjected to three convolution operations to obtain f φ 、f γ , whose dimensions are all R H×W×C , then f φ 、f γ Perform reshape operations separately:

[0038]

[0039] f φ = flat(W φ (f)) (10)

[0040] f γ = flat(W γ (f)) (11)

[0041] in, W φ , W γ There are three convolution operations, flat(·) means reshaping the image features;

[0042] Step SB4: After transposition, φ Multiply them to get a matrix, and perform a softmax operation on each point of the matrix to get the spatial attention feature S∈R N×N :

[0043]

[0044] Where σ represents the Softmax activation function, N = H × W;

[0045] Step SB5: Spatial attention features S and f γ Reshape into R after multiplication C×H×W , and the multi-scale fusion feature map f∈R H ×W×C Fusion is performed to obtain the final decoded input feature map f z :

[0046] f Z =σ(flat(f γ ·S))+f (13).

[0047] Furthermore, the working method of each of the fusion optimization modules includes:

[0048] Step SC1: After passing through the channel attention module CAM, feature F1 is channel-joined with feature F2 that has passed through pixel-shuffle upsampling, dilation convolution with a dilation rate of 2, and position attention module PAM in sequence to obtain the fused feature:

[0049]

[0050] Among them, the feature F1∈R H×W×C is low-level semantic information; feature F2 is high-level semantic information, is a dilated convolution with a kernel size of 3 and a dilation rate of 2. P(·) indicates that the feature is operated through the position attention module PAM, E(·) indicates that the feature is operated through the channel attention module CAM, pix(·) indicates pixel-shuffle upsampling, and Cat(·) indicates the superposition of features in the channel dimension.

[0051] Step SC2: The fusion feature is subjected to a dilation rate of 2 to enlarge the receptive field and then a convolution operation is performed to output F Z :

[0052]

[0053] Among them, Conv(·) represents the convolution operation with a convolution kernel size of 3.

[0054] After adopting the above technical solution, the present invention has the following beneficial effects:

[0055] 1. The present invention uses an edge refinement module in the encoding part to replace the traditional double-layer convolution, which improves the ability of the improved U-Net network to extract crack detail information, thereby solving the problem of missed detection of small cracks; the present invention designs a multi-scale fusion module based on the attention mechanism at the bottom of the U-Net network, and designs multiple fusion optimization modules in the decoding part, which solves the problem that crack detection is easy to break. The present invention effectively reduces the phenomenon of missed detection and false detection of road cracks;

[0056] 2. The edge refinement module of the present invention is designed by using the residual network and the channel attention mechanism, which can capture more crack detail feature information, suppress information irrelevant to the crack detection task, and thus enhance the ability to effectively extract features;

[0057] 3. In the encoding stage, image information is extracted through convolution and pooling operations. The extracted feature information can be divided into low-level semantic information and high-level semantic information. The low-level semantic information contains low-level information such as image contour and texture, and the high-level semantic information contains more abstract and advanced features. However, since the pooling operation is used multiple times in the feature extraction process, the resolution of the feature map is reduced and the receptive field is increased, so that a lot of image detail information and spatial information are lost, and some small cracks are easily missed. The multi-scale fusion module of the present invention can fuse feature information of different scales, that is, fuse the low-level semantic information with the high-level semantic information, so that the fused information contains richer crack feature information;

[0058] 4. The fusion optimization module of the present invention uses the attention mechanism to retain the crack detail information while adopting the dilated convolution to expand the receptive field, taking into account both the detection of small cracks and the continuity of crack detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a flow chart of an embodiment of a road crack detection method based on deep learning of the present invention;

[0060] Figure 2 A structural diagram of an edge refinement module of an embodiment of a road crack detection method based on deep learning of the present invention;

[0061] Figure 3 It is a structural diagram of a multi-scale fusion module based on an attention mechanism in one embodiment of a road crack detection method based on deep learning of the present invention;

[0062] Figure 4 It is a structural diagram of a fusion optimization module of an embodiment of a road crack detection method based on deep learning of the present invention;

[0063] Figure 5 This is a diagram of the overall network architecture of an embodiment of a road crack detection method based on deep learning of the present invention. DETAILED DESCRIPTION

[0064] In order to make the contents of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments in conjunction with the accompanying drawings.

[0065] The present invention first provides a road crack detection method based on deep learning, and its flow chart is as follows: Figure 1 shown.

[0066] Step S1: obtaining multiple road crack images, cropping them into a uniform size, and dividing the multiple road crack images into a training set, a validation set, and a test set;

[0067] In this embodiment, the road crack images are uniformly cropped to a size of 320×320.

[0068] Step S2: Building a U-Net network, wherein the U-Net network has an encoding part and a decoding part, wherein the encoding part is used to extract crack features, and the decoding part is used to restore the image and output a final feature map, and each of the encoding part and the decoding part has 5 layers;

[0069] Step S3:

[0070] Step S31: To solve the problem of missed detection of small cracks, the traditional double-layer convolution structure in the encoding part is replaced by an edge refinement module, and each layer includes three edge refinement modules;

[0071] The topological structure of the crack edge is complex and there are many small cracks. In the feature extraction stage, the traditional double-layer convolution layer structure in the convolution module of the encoding part extracts limited features, and as the network deepens, multiple convolution and pooling operations lead to the loss of image detail information in the process of extracting image features. In order to improve the network's ability to extract crack detail information, this embodiment designs an edge refinement module, namely ER.

[0072] Among them, the three edge refinement modules of the same layer are connected in series in sequence, and the original image with a size of 2H×2W×C0 is converted into H×W×C features after pooling, and then input into the first edge refinement module of the first coding layer of the encoding part. For the first four coding layers, the features output by the last edge refinement module of each layer are pooled and input into the first edge refinement module of the next layer.

[0073] like Figure 2 As shown, the processing process of each edge refinement module is as follows:

[0074] Step SA1: Input the feature x∈R of the edge refinement module H×W×C After 1×1 convolution in the edge refinement module, it is evenly divided into n feature subsets x i , where the i-th subset x i ∈{1,2,...,n}, in this embodiment, i is 4, and each subset x i The number of channels is C / n, x i ∈{2,3,...,n} undergoes the corresponding 3×3 convolution, and the output of the convolution is y i ∈{1,2,...,n}:

[0075]

[0076] Where C refers to the number of channels of the features input to the edge refinement module, and Conv(·) represents a convolution operation with a convolution kernel of 3×3.

[0077] Step SA2: Set y i ∈{1,2,...,n} is combined and restored to the original number of channels through 1×1 convolution, and the feature y∈R is output H×W×C ;

[0078] Step SA3: Output feature y∈R H×W×C After the channel attention CAM module, the feature y∈R H×W×C The following processing is performed in the channel attention CAM module:

[0079] First, global features are aggregated through global average pooling;

[0080]

[0081] Then the convolution operation adjusts the channel weights;

[0082] W=σ(Con'(y avg )) (3)

[0083] Finally, the weight W is combined with the feature y∈R of the input attention CAM module H×W×C multiply;

[0084] Among them, y i,j ∈R C It is a full channel feature. The convolution operation is a one-dimensional convolution of size k under the condition of the same dimension, where the size of the convolution kernel is k, which represents the coverage of local cross-channel interaction. In this embodiment, k is 3, which determines the coverage of the interaction. After convolution, the activation value is calculated using the Sigmoid function to obtain the weight W∈R 1×1×C Indicates the relevance and importance of each channel. In the above formula, Con'(·) represents a one-dimensional convolution of size K, σ represents the Sigmoid activation function, and the weight W is multiplied by the input feature y to complete the recoding of each channel feature, thereby assigning a larger weight to important features and assigning a smaller weight to non-task information to suppress it;

[0085] Step SA4: Connect the output features of the channel attention CAM module to the original input features x∈R of the edge refinement module through residual connection H×W×C To perform the fusion:

[0086] x=W·y+x (4)

[0087] Step S32: Aiming at the problem of easy fracture in crack detection, a multi-scale fusion module based on attention mechanism is designed;

[0088] The multi-scale fusion module fuses features of different scales and aggregates the features of each position. The feature information of the last layer of the encoding part loses a lot of crack detail information after multiple convolution pooling, and lacks the ability to solve the loss of crack edge information to a certain extent. Therefore, this embodiment proposes a pyramid structure for multi-layer output feature fusion, such as Figure 3 As shown in the figure, by using the output features of the specified coding layer for fusion and learning the positional relationship between feature points through the attention module, the image features of each layer can be fully utilized, which can not only reduce the loss of crack edge information but also ensure the continuity of crack information.

[0089] The encoding part is divided into 5 layers, represented by E1-E5, and the feature map scale of the i-th layer output is 1 / 2 of the original image size i , the low-level feature information includes the contour and edge information of the crack, and the high-level feature information includes the spatial information of the image. This embodiment fuses the low-level semantic information containing a lot of details output by the first two coding layers with the high-level global semantic information output by the last coding layer. Since the scale and number of channels of the feature maps are different, they cannot be directly fused. Therefore, Figure 3 As shown, the processing process of the multi-scale fusion module in this embodiment, namely AMFF, is as follows:

[0090] Step S321: After the feature map features output by the first two coding layers of the coding part are transformed through a 1×1 convolution operation channel, they are pooled to obtain feature maps of the same scale, and the two feature maps of the same scale are fused to obtain a fused feature map;

[0091] f1'=w(f(f1)) (5)

[0092] f2'=w(f(f2)) (6)

[0093] f 12 =Cat(f1',f2') (7)

[0094] Where f1 and f2 represent the outputs of the encoding layers E1 and E2, respectively; f(·) represents the convolution operation with a 1×1 convolution kernel; w(·) represents the pooling operation; and Cat(·) represents the superposition of features in the channel dimension.

[0095] Step S322: The fused feature map is fused with the feature map output by the last layer encoder E5, and finally a multi-scale fused feature map f∈R is output H×W×C :

[0096] f=Cat(f 12 ,f5) (8)

[0097] Among them, f5 represents the output of the encoding layer E5. Although the fusion of high-level semantic information and low-level semantic information can ensure the integrity of crack detection, it lacks the correlation between crack pixels. Therefore, it is difficult to maintain the coherence of crack segmentation, resulting in discontinuity. Therefore, a position attention PAM module is added after the output feature map to learn the spatial correlation of features through the position attention PAM module;

[0098] Step S323: The output multi-scale fusion feature map is subjected to three convolution operations to obtain f φ 、f γ , whose dimensions are all R H×W×C , then f φ 、f γ Perform reshape operations separately:

[0099]

[0100] f φ = flat(W φ (f)) (10)

[0101] f γ = flat(W γ (f)) (11)

[0102] in, W φ , W γ There are three convolution operations, flat(·) means reshaping the image features into N = H × W;

[0103] Step S324: After transposition, φ Multiply them to get a matrix, and perform softmax on each point of the matrix to get the spatial attention feature S∈R N×N :

[0104]

[0105] Among them, σ represents the Softmax activation function;

[0106] Step S:325: Spatial attention features S and f γ Reshape into R after multiplication C×H×W , and the multi-scale fusion feature map f∈R H×W×C Fusion is performed to obtain the final decoded input feature map f z :

[0107] f Z =σ(flat(f γ·S))+f (13)

[0108] Among them, step S323, step S324 and step S325 are performed in the position attention PAM module.

[0109] The multi-scale fusion module of this embodiment aggregates context information of different regions, extracts crack information from a global perspective, and extracts the correlation between each feature pixel point to enhance the integrity and continuity of pavement crack detection.

[0110] Step S33: In view of the problem that crack detection is prone to breakage, multiple fusion optimization modules are further designed in the decoding part; this embodiment uses the connection ideas of pixel-shuffle, hole convolution and attention mechanism to design a fusion optimization module. The existing network mainly uses zero padding or bilinear interpolation methods for upsampling. Since crack segmentation is a pixel-level classification task, the use of traditional upsampling methods makes the feature pixels easily interfered by surrounding pixels, affecting the final detection results. The main function of pixel-suffle is to obtain a high-resolution feature map through convolution and multi-channel recombination of low-resolution feature maps. Pixel-shuffle is an upsampling method commonly used in the study of super-resolution reconstruction problems. Compared with conventional upsampling methods, it can reduce information loss and have higher detection accuracy. In this embodiment, the pixel-shuffle convolution layer is mainly used to replace the commonly used transposed convolution operation; the feature map is convolved with pixel-shuffle and then uses hole convolution to achieve the growth of the receptive field without reducing the resolution of the feature map, so that the output of each convolution contains a wider range of information; the feature map after hole convolution captures more crack position relationships through the position attention module to prevent crack breakage. A CAM module is added in the jump connection process to filter information and highlight more crack details.

[0111] like Figure 4 As shown in FIG. 1 , the structure of the fusion optimization module, namely, FO, is as follows: Figure 5 As shown, there are three fusion optimization modules, namely FO1, FO2, and FO3 from top to bottom. The processing process is as follows:

[0112] Step S331: set i=4;

[0113] Step S332: Features and feature F1 i In the fusion optimization module FO i-1 Perform the operation in formula (14-1):

[0114]

[0115] Step S333: F' in the fusion optimization module FOi-1 Perform the operation in formula (15-1):

[0116]

[0117] Step S334: i=i-1;

[0118] Step S335: Determine whether i is 1. If so, the processing ends. As the final output of the decoding part, if not, return to step S332;

[0119] Among them, F1 i ∈R H×W×C (i=1, ..., 4) is the low-level semantic information, which is output by the first four encoding layers E1-E4 respectively. is output from the fusion optimization module (FO) or from the multi-scale fusion module, where By the i-th fusion optimization module FO i (i=1,2,3) output, Output from the multi-scale fusion module, is a dilated convolution with a kernel size of 3 and a dilation rate of 2. P(·) indicates that the feature is operated through position attention PAM, E(·) indicates that the feature is operated through channel attention CAM, pix(·) indicates pixel-shuffle upsampling, and Cat(·) indicates the superposition of features in the channel dimension; Conv(·) represents a convolution operation with a kernel size of 3.

[0120] In this embodiment, the feature Perform pixel-shuffle upsampling to make its resolution consistent with F1 i The same; the dilation rate is 2, the hole convolution is to increase the receptive field; the position attention module PAM is to extract the correlation between feature pixels; the feature F1 output by the encoding part is passed through the channel attention module CAM to extract more crack detail information.

[0121] In this embodiment, if Figure 5 As shown in the figure, the feature map output by the decoding part is upsampled by pixel-shuffle to restore the original image size, and then convolved by 1X1 as the output of the overall network.

[0122] Step S4: Loading the training set and the validation set into the improved U-Net network for training and validation, and saving the model with the best effect;

[0123] Step S5: Use the model with the best effect to test the road crack images in the test set, obtain test results, and complete road crack detection.

[0124] The specific embodiments described above further illustrate the technical problems, technical solutions and beneficial effects solved by the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A road crack detection method based on deep learning, characterized in that: include: Acquire a plurality of road crack images, and divide the plurality of road crack images into a training set, a validation set, and a test set; Building a U-Net network, wherein the U-Net network has an encoding part and a decoding part, and the encoding part and the decoding part each have 5 layers; The traditional double-layer convolution structure in the encoding part is replaced by an edge refinement module, each layer includes three edge refinement modules, a multi-scale fusion module based on an attention mechanism is designed at the bottom of the U-Net network, and fusion optimization modules are designed at the 2nd, 3rd and 4th layers of the decoding part respectively, to obtain an improved U-Net network; Loading the training set and the validation set into the improved U-Net network for training and validation, and saving the model with the best effect; Using the model with the best effect to test the road crack images in the test set, to obtain a test result; The working method of each edge refinement module includes: Step A1: Input the feature x∈R of the edge refinement module H×W×C After 1×1 convolution, it is evenly divided into n feature subsets x i , where the i-th subset x i ∈{1,2,...,n}, each subset x i The number of channels is C / n; x i ∈{2,3,...,n} after the corresponding 3×3 convolution, the output is y i ∈{1,2,...,n}: Where C refers to the number of channels of the features input to the edge refinement module, and Conv(·) represents a convolution operation with a convolution kernel of 3×3. Step A2: i ∈{1,2,...,n} is combined and restored to the original number of channels through 1×1 convolution, and the feature y∈R is output H ×W×C ; Step A3: Output features y∈R H×W×C After the channel attention CAM module, the output feature y∈R H×W×C The following processing is performed in the channel attention CAM module: First, global features are aggregated through global average pooling; Then the convolution operation adjusts the channel weights; W=σ(With'(and avg )) (3) Finally, the weight W is combined with the feature y∈R of the input channel attention CAM module H×W×C multiply; Among them, y i,j ∈R C is the full channel feature, Con'(·) represents the one-dimensional convolution of size K, and σ represents the Sigmoid activation function; Step A4: Connect the output features of the channel attention CAM module to the original input features x∈R of the edge refinement module through residual connection H×W×C To perform the fusion: x=W·y+x (4); The working method of the multi-scale fusion module based on the attention mechanism includes: Step B1: The feature maps output by the first two coding layers of the coding part are respectively subjected to 1×1 convolution operation to transform the channels and then pooled to obtain feature maps with the same scale and number of channels, and the two feature maps with the same scale and number of channels are fused to obtain a fused feature map; f1'=w(f(f1)) (5) f2'=w(f(f2)) (6) f 12 =Cat(f1',f2') (7) Among them, f1 and f2 represent the outputs of the first two encoding layers respectively, f(·) represents the convolution operation with a 1×1 convolution kernel, w(·) represents the pooling operation, and Cat(·) represents the superposition of features in the channel dimension; Step B2: Fuse the fused feature map with the feature map output by the last encoding layer of the encoder, and finally output the multi-scale fused feature map f∈R H×W×C : f=Cat(f 12 ,f5) (8) Among them, f5 represents the output of the last encoding layer; Step B3: The output multi-scale fusion feature map is subjected to three convolution operations to obtain f φ 、f γ , whose dimensions are all R H×W×C , then f φ 、f γ Perform reshape operations separately: f φ =flat(W φ (f)) (10) f γ =flat(W γ (f)) (11) in, W φ , W γ There are three convolution operations, flat(·) means reshaping the image features; Step SB4: After transposition, φ Multiply them to get a matrix, and perform a softmax operation on each point of the matrix to get the spatial attention feature S∈R N×N : Where σ represents the Softmax activation function, N = H × W; Step SB5: Spatial attention features S and f γ After multiplication, reshape into R C×H×W , and the multi-scale fusion feature map f∈R H×W×C Fusion is performed to obtain the final decoded input feature map f z : f Z =σ(flat(f γ ·S))+f (13); The working method of each fusion optimization module includes: Step SC1: After passing through the channel attention module CAM, feature F1 is channel-joined with feature F2 that has passed through pixel-shuffle upsampling, dilation convolution with a dilation rate of 2, and position attention module PAM in sequence to obtain the fused feature: Among them, the feature F1∈R H×W×C is low-level semantic information; feature F2 is high-level semantic information, is a dilated convolution with a kernel size of 3 and a dilation rate of 2. P(·) indicates that the feature is operated through the position attention module PAM, E(·) indicates that the feature is operated through the channel attention module CAM, pix(·) indicates pixel-shuffle upsampling, and Cat(·) indicates the superposition of features in the channel dimension. Step SC2: The fusion feature is subjected to a dilation rate of 2 to enlarge the receptive field and then a convolution operation is performed to output F Z : Among them, Conv(·) represents the convolution operation with a convolution kernel size of 3.

2. The road crack detection method based on deep learning according to claim 1, characterized in that: Before dividing the plurality of road crack images into a training set, a validation set and a test set, the method further includes: The road crack images are cropped into a uniform size.

Citation Information

Patent Citations

  • High-precision crack detection method

    CN111222580A

  • Mountain crack detection method based on improved self-attention mechanism and transfer learning

    CN114022770A