Pavement crack detection method based on multi-scale receptive field and mixed attention mechanism
The UNet network with MRFE and DMA modules addresses crack detection issues by enhancing receptive fields and attention mechanisms, ensuring accurate and complete segmentation of cracks of varying sizes and reducing noise interference.
Patent Information
- Application Number
- CN202510424242.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-15
AI Technical Summary
The existing road surface crack detection methods are prone to missed and missed when facing cracks of extreme size, and are easily disturbed by similar textures, and have poor anti-interference ability, resulting in inaccurate detection results.
The road surface crack detection method based on multi-scale receptive field and hybrid attention mechanism is adopted. By introducing a multi-scale receptive field expansion module and a dual hybrid attention module in the UNet network, the range of the receptive field is expanded, the network's ability to segment cracks of different sizes is enhanced, and background noise interference is reduced through the attention mechanism.
Continuous and complete segmentation of cracks of different sizes is achieved, missed and missed detection is reduced, and the accuracy and robustness of detection is improved, especially in complex backgrounds and noise environments, which can clearly segment fractures.
Smart Images

Figure CN120318494A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image analysis models, and specifically to a pavement crack detection method based on multi-scale receptive fields and a hybrid attention mechanism. Background Art
[0002] Cracks are potential threats to road safety. Regular pavement crack detection is considered a key link in highway maintenance, which can timely detect pavement cracks, estimate the severity of damage, and then make corresponding rapid and reliable repair decisions according to various damage conditions to keep the pavement in good condition.
[0003] Crack detection can be achieved through two methods: manual detection and automatic detection. Manual detection requires a large amount of labor, which is neither economical nor time-saving. Currently, automatic crack detection methods are generally used for crack detection. Automatic crack detection includes two stages: data collection stage and crack detection stage. In the data collection stage, devices such as road scanners, vehicle-mounted cameras, drones, and mobile phones are used to collect crack images. The crack images are an important prerequisite for subsequent detection of cracks using various image processing techniques or deep learning algorithms. After the crack images are collected, it enters the crack detection stage. This stage is to analyze and process the collected images and is the most important stage for realizing automatic crack detection. Researchers have proposed a large number of crack detection methods.
[0004] At first, traditional image processing methods were used to detect cracks. This method uses various algorithms with specific rules for images, such as edge detection, mathematical morphology, minimum path, etc. for quantization processing. The crack detection method based on traditional image processing is fast and simple, but it is easily interfered by background noise and has low accuracy. In addition, the feature extractor of this method usually uses manual design, which is complex and inefficient, and it is designed for specific tasks or image sets. Once the environmental conditions or the image features to be segmented change, the detection effect will be greatly reduced.
[0005] With the progress and maturity of machine learning technology, researchers began to use machine learning methods to detect pavement cracks. The crack detection method based on machine learning mainly extracts features manually first, and then classifies them using a machine learning-based classifier. However, in a complex environment, it is difficult to extract effective features using this method, which easily leads to poor reliability of the detection model. That is, the crack detection method based on machine learning can only be used within a limited range and has poor robustness and generalization ability.
[0006] Subsequently, with the rapid development of computer algorithms and high-performance computing devices, various deep learning algorithms have been widely applied to crack detection tasks. According to different task objectives, there are mainly three categories of methods: crack detection methods based on image classification, crack detection methods based on object detection, and crack detection methods based on image segmentation. Different from image classification and object detection methods, image segmentation methods can obtain detailed information about cracks, thereby obtaining the geometric features corresponding to the cracks. Currently, the mainstream method in crack detection tasks is the crack detection method based on image semantic segmentation, and researchers use various classic semantic segmentation methods based on the encoder-decoder structure or improved methods to implement crack detection tasks.
[0007] Although the current deep learning-based crack detection methods have far exceeded the traditional image processing-based crack detection methods in terms of detection effect, the existing technologies have the following disadvantages:
[0008] (1) Cracks usually appear as slender, irregular straight or curved shapes, with different sizes and complex topological structures. Some cracks occupy a very small proportion in the image, and sometimes there are only a few pixel widths, often resulting in missed detections, leading to unclear crack details or boundary segmentation results. Some cracks may be very large, even filling an entire image, and there are also images with cracks of various sizes. The receptive fields of existing crack detection methods are small and single, resulting in discontinuous and incomplete segmentation results for large-size cracks, and missed detections and misclassifications for small-size cracks.
[0009] (2) Cracks are often hidden in pavement textures that are extremely similar to themselves, with low contrast between the two, and crack images usually contain complex background noises, such as oil stains, lane markings, manhole covers, leaves, shadows, etc., which are likely to interfere with crack detection and lead to inaccurate detection results. Existing crack detection methods have poor anti-interference ability in these situations, do not fully utilize the feature information in crack images, and are extremely prone to misdetection. Summary of the Invention
[0010] In view of the problems in the prior art of being prone to missed detections and misclassifications for cracks with extreme sizes and being greatly interfered by similar textures, the present invention discloses a pavement crack detection method based on multi-scale receptive fields and a hybrid attention mechanism. The technical solution adopted is as follows:
[0011] Step 1: Input a pavement photo into the UNet network, and the UNet network also includes an encoder and a decoder, and the encoder is connected to the decoder;
[0012] Step 2: Perform encoding and decoding operations on the picture by the UNet network;
[0013] Step 3, highlighting the identified road surface gaps;
[0014] A multiscale receptive field expansion (MRFE) module is provided at the connection between the encoder and the decoder, and a dual mixed attention (DMA) module is provided at each hop connection between the encoder and the decoder in the UNet network.
[0015] Multi-scale receptive field: The receptive field of the network plays a vital role in convolutional neural networks. A receptive field that is too small will limit the network's ability to obtain complete target features, while a receptive field that is too large will lead to inaccurate positioning of target pixels and blurred segmentation results. The cracks in pavement crack images usually vary in size and length. Existing methods all use a fixed receptive field, but a fixed receptive field does not work well in crack segmentation. Multi-scale receptive fields can expand the receptive field of existing methods and provide receptive fields of different scales, which is conducive to extracting crack features of different sizes, allowing the network to more effectively segment pavement cracks.
[0016] Attention mechanism: When humans look at images, they selectively focus on specific targets or areas and ignore the rest of the objects or content in the image. Inspired by this way of visual processing, the field of computer vision introduced the attention mechanism, which learns and adjusts adaptive weights based on the input image features, and makes a differentiated non-equal distribution of the weights originally evenly distributed in the image, thereby shifting the network's attention to the useful parts of the image. Pavement crack images have the characteristics of low pavement background contrast, complex pavement background and a lot of noise. Adding an attention mechanism to the crack segmentation network helps the network focus more on the crack area in the image to be segmented, ignore the background and noise areas in the image as much as possible, and reduce the interference of background and noise.
[0017] As a preferred technical solution of the present invention, the multi-scale receptive field expansion module also includes multiple parallel branches, some of which are dilated convolutions with different expansion coefficients. The dilated convolutions with different expansion coefficients have receptive fields of different sizes, which can perform multi-scale feature extraction on the crack image, which is beneficial to simultaneously segment cracks of various sizes in the image; the other branches are parallel SoftPool pooling branches along different directions, which can obtain more complete and effective long-distance features of the cracks.
[0018] As a preferred technical solution of the present invention, the dilated convolution has three branches, and their expansion coefficients are 2, 4, and 8 respectively.
[0019] As a preferred technical solution of the present invention, there are two SoftPool pooling branches, along the horizontal and vertical directions respectively.
[0020] As a preferred technical solution of the present invention, the dual hybrid attention module further includes an improved channel attention branch and an improved spatial attention branch, and the two branches are parallel branches.
[0021] As a preferred technical solution of the present invention, the improved channel attention branch includes the following operation steps:
[0022] Step 11, compress the spatial dimension of the input feature map. Therefore, after the fusion of the input high-level and low-level feature maps, a SoftPool pooling operation is performed to aggregate the spatial information of the features.
[0023] Step 12, send the spatial information obtained in Step 11 into a multi-layer perceptron for non-linear learning.
[0024] Step 13, then use the Sigmoid activation function for calculation to obtain the weight coefficient of the channel attention.
[0025] Step 14, finally, multiply the weight coefficient with the input feature by matrix multiplication to obtain the channel attention feature.
[0026] As a preferred technical solution of the present invention, the channel attention calculation formula obtained in Step 14 is: where M c represents the improved channel attention, σ represents the Sigmoid activation function, MLP represents being sent into a multi-layer perceptron for operation, R represents the pooling kernel, and F i represents the activation value of pixel i in the SoftPool corresponding pooling kernel region on the input feature map F, and F j represents the activation value of the corresponding pixel in this pooling kernel.
[0027] As a preferred technical solution of the present invention, the improved spatial attention branch includes the following operation steps:
[0028] Step a, first perform a SoftPool pooling operation to compress the feature dimension.
[0029] Step b, send the result after pooling into a 7×7 convolution for calculation.
[0030] Step c, then pass through the Sigmoid activation function to calculate and obtain the weight coefficient of the spatial attention.
[0031] Step d: Multiply the spatial attention weight with the input features through matrix multiplication to obtain the spatial attention features.
[0032] As a preferred technical solution of the present invention, the spatial attention calculation formula obtained in step d is: where M S represents the spatial attention, Conv 7×7 represents the convolution operation with a convolution kernel size of 7×7, σ represents the Sigmoid activation function, R represents the pooling kernel, and F i represents the activation value of pixel i in the pooling kernel region corresponding to SoftPool on the input feature map F, and F j represents the activation value of the corresponding pixel in the pooling kernel.
[0033] Advantages of the present invention: By adding the multi-scale receptive field expansion (MRFE) module proposed by the present method at the connection between the encoder and decoder of the original UNet network, the receptive field of the network during feature extraction is expanded, enabling the network to segment large-scale long cracks continuously and completely. The multi-scale receptive field expansion MRFE module can provide receptive fields of different scales on five parallel branches, which is beneficial for the network to more effectively segment cracks of different sizes in the image, being able to segment large cracks well without missing small cracks. In addition, the multi-scale receptive field expansion MRFE module uses two long-strip SoftPool pooling branches along the horizontal and vertical directions respectively in the last two branches, which helps the network extract the long-range dependencies of the cracks to further obtain complete and continuous global features of the cracks.
[0034] Furthermore, by adding the dual mixed attention (DMA) module proposed by the present method to each layer of skip connection between the encoder and decoder of the original UNet, the information from low-level and high-level image features can be fully utilized and processed to achieve more effective fusion. The network can not only pay more attention to the abstract semantic information in high-level image features, reducing the mis-segmentation of background noise as cracks, but also pay more attention to the crack detail information in low-level image features, reducing the missed detection of small cracks by the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0036] Figure 1 It is a schematic diagram of the overall structure of the UNet network of the present invention;
[0037] Figure 2 This is the structural diagram of the MRFE module of the present invention;
[0038] Figure 3 This is the structural diagram of the DMA module of the present invention;
[0039] Figure 4 This is the schematic diagram of the test image segmentation result of the Crack500 dataset of the present invention;
[0040] Figure 5 This is the schematic diagram of the test image segmentation result of the DeepCrack dataset of the present invention;
[0041] Figure 6 This is the schematic diagram of the test image segmentation result of the GAPs384 dataset of the present invention;
[0042] Figure 7 This is the schematic diagram of the test image segmentation result of the CFD dataset of the present invention. Detailed implementation manners
[0043] Example 1
[0044] As Figures 1 to 3 shown, the present invention discloses a road surface crack detection method based on a multi-scale receptive field and a hybrid attention mechanism. The adopted technical solution includes the following steps:
[0045] Step 1, input a road surface photo into the UNet network. The UNet network also includes an encoder and a decoder. There is a multi-scale receptive field expansion module at the connection between the encoder and the decoder, and a dual hybrid attention DMA module is added to each skip connection layer of the encoder and the decoder, and this module is used to replace the direct simple fusion operation of the original UNet network. The specific designs of the MRFE module and the DMA module are as follows:
[0046] (1) Multi-scale receptive field expansion MRFE module
[0047] The specific structure of the MRFE module is as Figure 2 shown, which is composed of five parallel branches. The first three branches are dilated convolutions with dilation coefficients of 2, 4, and 8 respectively. Dilated convolutions with different dilation coefficients have different sizes of receptive fields, which can perform multi-scale feature extraction on crack images and are beneficial to segmenting cracks of various sizes in the image at the same time; the latter two branches are two parallel long-strip SoftPool pooling branches along the horizontal and vertical directions respectively, which can obtain more complete and effective long-distance features of cracks.
[0048] The different parallel branches of the MRFE module can provide receptive fields of different scales, which is beneficial to extracting crack features of different sizes and enabling the network to perform pavement crack segmentation more effectively. In addition, in the pavement crack segmentation task, cracks usually appear as strips and are slender. If the scope of action of all branches is square, it is not conducive to obtaining correct and complete long-strip crack feature information more effectively. Therefore, the MRFE module uses rectangular pooling kernels in the latter two branches. This way, a larger range of feature information in the network can be obtained, which is beneficial to extracting remote context information in the feature image and establishing long-range dependencies.
[0049] When existing methods use pooling kernels with a vertical dimension of H×1 and a horizontal dimension of 1×W to pool the input image, they usually perform average pooling or max pooling operations. However, the method adopted here is the SoftPool pooling operation that combines the advantages of both.
[0050] Compared with the two common pooling methods, average pooling and max pooling, SoftPool pooling can reduce the loss of detailed information in the pooling operation to a greater extent because the pooling processes of average pooling and max pooling are too simple. Average pooling calculates the average value of the activation values corresponding to all pixels in the pooling area, and max pooling obtains the maximum value among the activation values corresponding to all pixels in the pooling area. In contrast, SoftPool pooling performs calculations in the high-dimensional feature space using a more complex Softmax weighting method. SoftPool pooling first calculates the weights based on the natural exponential e, and the size of the weights depends on the size of the corresponding activation values. In the output of SoftPool pooling, pixels with higher activation values play a greater role, and SoftPool pooling uses the activation values of all pixels in the pooling kernel area, combining the advantages of max pooling and average pooling. Moreover, the SoftPool pooling operation can perform differential operations, which is beneficial to the training of the network and can help the network learn more effective crack features.
[0051] (2) Dual Mixed Attention DMA Module
[0052] In order to reduce the loss of semantic features during the downsampling operation in the encoder part of the original UNet network, skip connections are used layer by layer to fuse low-level and high-level image features. However, this fusion method uses a direct and simple fusion operation, which cannot fully utilize the local context feature information. The degree of detail recovery of the image features is not ideal, and it is easy to cause the network to miss small cracks. Therefore, this method proposes to replace the direct and simple fusion operation in the original UNet skip connection with a Dual Mixed Attention (DMA) module. This module contains two parallel branches, which perform correlation modeling from the two dimensions of channels and space respectively, increase the weight of crack features and enhance the dependence between channels, enabling the network to not only extract rich semantic information to effectively distinguish cracks from background noise, but also extract sufficient spatial detail information to segment cracks more precisely, which is also beneficial to the segmentation of small cracks.
[0053] The DMA module proposed by this method is as Figure 3 shown, divided into upper and lower branches. The upper branch is the improved channel attention part, and the lower branch is the improved spatial attention part.
[0054] ① Improved channel attention branch
[0055] The upper branch is used to obtain channel attention features. To obtain channel attention, first, it is necessary to compress the spatial dimension of the input feature map. Therefore, after the fused high-level and low-level feature maps are input, a SoftPool pooling operation is performed to aggregate the spatial information of the features. Then, it is sent into a Multilayer Perceptron (MLP) for non-linear learning. This MLP contains a hidden layer. Then, the Sigmoid activation function is used for calculation to obtain the weight coefficient of channel attention. Finally, the weight coefficient is multiplied by the input features in matrix form to obtain the channel attention features, which can help the network focus on more useful semantic information. The calculation formula for the improved channel attention Mc is as follows:
[0056] where σ represents the Sigmoid activation function, MLP represents being sent into the Multilayer Perceptron for operation, R represents the pooling kernel, and F i represents the activation value of pixel i in the pooling kernel region corresponding to SoftPool on the input feature map F, and F j also represents the activation value of the corresponding pixel in this pooling kernel.
[0057] ② Improved spatial attention branch
[0058] The lower branch is used to obtain spatial attention features. Similar to the upper branch, the input fused features first undergo a SoftPool pooling operation to compress the feature dimension. Then, the pooled result of this branch is fed into a 7×7 convolution for calculation, and then passed through a Sigmoid activation function to calculate the weight coefficients of spatial attention. Finally, the spatial attention weights are multiplied by the input features in matrix form to obtain spatial attention features, which encode the specific spatial positions that need to be attended to or suppressed, and can help the network focus on the areas with the richest detailed information in the feature map. The spatial attention M obtained by the lower branch S The corresponding specific calculation formula is as follows:
[0059] where Conv 7×7 represents a convolution operation with a 7×7 convolution kernel, and σ represents the Sigmoid activation function.
[0060] The low-level image features in a convolutional neural network usually contain rich spatial information, which can provide a large amount of crack details and is beneficial to the fine segmentation of cracks and the segmentation of small cracks; while the high-level image features usually contain rich semantic information, which is beneficial to the correct judgment and classification of crack pixels and improves the accuracy of crack segmentation. The method in this paper layer by layer feeds the low-level image features from the encoder and the high-level image features after upsampling from the decoder into the dual hybrid attention module. After the feature fusion operation, channel attention and spatial attention are calculated simultaneously to enhance the spatial and semantic information of the features extracted by the network. Finally, the new features obtained on the two branches are fused to output more complete and diverse features. The dual hybrid attention module can effectively improve the segmentation accuracy of the network in two aspects. One is to improve the network's fine segmentation ability for cracks and even small cracks through the improved spatial attention branch in the module, and the other is to reduce the influence of background noise on crack segmentation through the improved channel attention branch in the module and reduce the situation of mis-segmenting background noise as cracks. Comparative experiments
[0061] The public datasets Crack500, DeepCrack, GAPs384, and CFD are used as experimental datasets. The sizes of the images in the above four datasets are uniformly adjusted to 448×448 pixels. The training set and validation set provided in Crack500, the training set provided in DeepCrack, and 450 randomly selected images from GAPs384 are merged, and they are data-augmented, including horizontal flipping and vertical flipping operations, to increase the number of images to three times the original, obtaining the training dataset for the network proposed in the present invention. The remaining images in Crack500, DeepCrack, and GAPs384, as well as the entire CFD dataset, are used as the test dataset for the proposed network.
[0062] Then, this technical solution is trained on the training set of the experimental dataset with four representative segmentation networks, including the original UNet network, AttentionUNet, CENet, and U-HDN. Subsequently, the obtained network model is evaluated on the test set. Partial segmentation results of all methods on the test set are as Figures 4 to 7 shown, respectively from the Crack500 dataset, DeepCrack dataset, GAPs384 dataset, and CFD dataset.
[0063] From the segmentation result diagrams as Figures 4 to 7 shown, it can be seen that when facing Figure 4 such crack images with very complex background textures and a large amount of noise, this method can also clearly segment the cracks; when facing Figure 5 cracks with different sizes and thicknesses in Figure 6 , this method can not only accurately segment cracks of different sizes but also precisely segment small cracks to avoid missed detections; when facing Figure 7 interference from objects in the background that are very similar to the cracks, the segmentation result of this method is clear, the obtained crack shape is good, and almost no segmentation is performed on the interfering objects. The segmentation effect of the Attention UNet network is second only to this method, indicating that the attention mechanism has a certain inhibitory effect on background noise interference; when this method segments
[0064] long and thin, cross-shaped cracks with a relatively complex topological structure, it can also achieve continuous and complete segmentation with fewer missing and broken cases, having a high segmentation accuracy and being able to capture most of the details in the cracks.
[0064] The specific numerical results of the comparative experiments on the evaluation indicators are shown in the following four tables: Table 1 Comparison of different methods on the Crack500 dataset Table 2 Comparison of different methods on the DeepCrack dataset Table 3 Comparison of different methods on the GAPs384 dataset Table 4 Comparison of different methods on the CFD dataset
[0065] From the results of the comparative experiments, it can be seen that the performance of this method on the four public datasets is better than that of the other four comparative methods, and the results of this method are better than those of the other four methods in terms of the comprehensive index F1 score and the IoU index, which proves the effectiveness of the multi-scale receptive field expansion MRFE module and the dual mixed attention DMA module proposed in this method. Ablation experiment
[0066] An ablation experiment was carried out on this method on the Crack500 dataset. The basic network is UNet, and the multi-scale receptive field expansion MRFE module and the dual mixed attention DMA module proposed in this chapter are added to UNet in turn. The parameters of each network containing different modules during training are the same as those used in the network proposed in this chapter.
[0067] Table 5 shows the specific numerical results of each network on the four evaluation indicators during the ablation experiment. From the results of the ablation experiment, it is easy to see that adding the MRFE module and the DMA module is a correct and effective improvement for the original UNet network. Table 5 Ablation experiment on the Crack500 dataset
[0068] Although the specific embodiments of the present invention have been described in detail above, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the spirit of the present invention, and modifications or deformations that do not involve creative labor are still within the protection scope of the present invention.
Claims
1. A road surface crack detection method based on multi-scale receptive fields and hybrid attention mechanism, comprising the following steps: Step 1, input a road surface photo into the UNet network, the UNet network further includes an encoder and a decoder, and the encoder is connected to the decoder; Step 2, perform encoding and decoding operations on the picture by the UNet network; Step 3, highlight the identified road surface cracks; It is characterized in that: There is a multi-scale receptive field expansion module at the connection between the encoder and the decoder, and a dual hybrid attention module is provided at each layer skip connection between the encoder and the decoder of the UNet network.
2. The pavement crack detection method based on multi-scale receptive fields and hybrid attention mechanism according to claim 1, characterized in that: The multi-scale receptive field expansion module further includes a plurality of parallel branches, wherein a part of the branches are dilated convolutions with different dilation coefficients, and another part of the branches are parallel SoftPool pooling branches along different directions.
3. A pavement crack detection method based on multi-scale receptive fields and hybrid attention mechanism according to claim 2, characterized in that: There are three branches of the dilated convolution, and their dilation coefficients are 2, 4, and 8 respectively.
4. A pavement crack detection method based on multi-scale receptive fields and hybrid attention mechanism according to claim 2, characterized in that: There are two SoftPool pooling branches, along the horizontal and vertical directions respectively.
5. A pavement crack detection method based on multi-scale receptive fields and hybrid attention mechanism according to claim 1, characterized in that: The dual hybrid attention module further includes an improved channel attention branch and an improved spatial attention branch, and the two branches are parallel branches.
6. The pavement crack detection method based on multi-scale receptive fields and hybrid attention mechanism according to claim 5, characterized in that, The improved channel attention branch includes the following operation steps: Step 11, compress the spatial dimension of the input feature map, so that the input high-level and low-level feature maps are subjected to SoftPool pooling operation after fusion to aggregate the spatial information of the features; Step 12, send the spatial information obtained in Step 11 into a multi-layer perceptron for non-linear learning; Step 13, then use the Sigmoid activation function to calculate to obtain the weight coefficient of the channel attention; Step 14, finally multiply the weight coefficient with the input feature by matrix multiplication to obtain the channel attention feature.
7. A pavement crack detection method based on multi-scale receptive fields and hybrid attention mechanism according to claim 6, characterized in that The channel attention calculation formula obtained in Step 14 is: Among them, M c represents the improved channel attention, σ represents the Sigmoid activation function, MLP represents being fed into a multi-layer perceptron for operation, R represents the pooling kernel, F i represents the activation value of pixel i in the pooling kernel region corresponding to SoftPool on the input feature map F, F j represents the activation value of the corresponding pixel in this pooling kernel.
8. A pavement crack detection method based on multi-scale receptive fields and hybrid attention mechanism according to claim 5, characterized in that, The improved spatial attention branch includes the following operation steps: Step a, first perform SoftPool pooling operation to compress the feature dimension; Step b, send the pooled result into a 7×7 convolution for calculation; Step c, then pass through the Sigmoid activation function to calculate to obtain the weight coefficient of the spatial attention; Step d, multiply the spatial attention weight with the input feature by matrix multiplication to obtain the spatial attention feature.
9. A pavement crack detection method based on multi-scale receptive fields and hybrid attention mechanism according to claim 8, characterized in that The spatial attention calculation formula obtained in Step d is: Among them, M S represents spatial attention, Conv 7×7 represents a convolution operation with a convolution kernel size of 7×7, σ represents the Sigmoid activation function, R represents the pooling kernel, F i represents the activation value of pixel i in the pooling kernel region corresponding to SoftPool on the input feature map F, F j represents the activation value of the corresponding pixel in this pooling kernel.
Citation Information
Cited By
Crack detection model training method, crack detection method and device
CN121353230A