Mountain Fracture Detection Method Based on Improved Self-Attention Mechanism and Transfer Learning
By improving the self-attention mechanism and transfer learning method, combined with the Swin-Transformer and Patch-merging modules, the problems of low automation and insufficient accuracy of mountain crack detection are solved, and high-precision mountain crack detection is achieved.
Patent Information
- Application Number
- CN202111335474.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-11
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-11-11
AI Technical Summary
The existing mountain crack detection methods have low degree of automation, and deep learning-based methods have insufficient accuracy, high computational complexity, and lack of data, making it difficult to be practical, especially in mountain environments.
The improved self-attention mechanism and transfer learning method are adopted to extract the global dependence and multi-scale information of the picture through the Swin-Transformer module and the Patch-merging module, and optimize network training using Focal-Loss loss function, combined with transfer learning to apply it to mountain crack detection.
It improves the accuracy and generalization ability of mountain crack detection, reduces the computational complexity, saves manpower and material resources, and is suitable for high-precision detection of mountain cracks.
Smart Images

Figure CN114022770B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and computer vision, and more specifically, to a mountain crack detection method based on an improved self-attention mechanism and transfer learning. Background Art
[0002] Mountain cracks are a common natural hazard. Unwarned mountain cracks often evolve into natural disasters such as landslides, causing huge property losses and casualties. Therefore, it is very necessary to detect mountain cracks in time. Most traditional mountain crack detection methods use instrument monitoring plus manual identification. Although these methods have certain detection effects, they are easily restricted by the field environment, and the automation degree is low, which requires a large amount of manpower and material resources.
[0003] The crack detection method based on deep learning has a good development prospect due to its high automation degree, but there are still many problems that lead to its inapplicability in many scenarios due to insufficient accuracy. Specifically, the current mainstream deep learning-based detection methods mainly include three categories: methods based on upsampling and downsampling, methods based on encoder-decoder structures, and methods based on self-attention mechanisms. Among them, the information extracted by the network of the first category of methods is not rich enough and does not consider the uniqueness of different-scale features, resulting in low detection accuracy. For the second category of methods, since it is difficult for convolutional neural networks to pay attention to the global dependencies of pixels, the segmentation effect at fine granularity still needs to be improved. For the third category of methods, a self-attention mechanism is added to the network to extract the global dependencies of pixels, which improves the detection accuracy of the network to a certain extent. However, the computational complexity of calculating the self-attention weights of this type of method is too high, resulting in inefficient training and limited accuracy improvement, generally unable to meet the requirements. At the same time, there is also a problem of lack of data in mountain crack pictures, which affects network training. Summary of the Invention
[0004] The present invention provides a mountain crack detection method based on an improved self-attention mechanism and transfer learning to efficiently detect and warn mountain cracks.
[0005] In order to achieve the above technical effects, the technical solution of the present invention is as follows:
[0006] A mountain crack detection method based on an improved self-attention mechanism and transfer learning, comprising the following steps:
[0007] S1: Preprocess the dataset of the crack detection network to obtain a training set and a validation set;
[0008] S2: Use the training set obtained in step S1 to train the crack detection network;
[0009] S3: Use the validation set obtained in step S1 to select the best-performing model obtained by training the crack detection network;
[0010] S4: Test the performance and generalization of the model on different road crack datasets.
[0011] S5: Apply the model processed in step S4 to mountain crack detection using transfer learning method.
[0012] Furthermore, the process of step S1 is as follows:
[0013] Resize all training samples and label samples for easier loss calculation, and remove samples with excessive noise in the dataset. Then, divide the remaining qualified samples into a training set and a validation set, where the training set is used to train the network and the validation set is used to select the model with the best performance.
[0014] Furthermore, the process of step S2 is as follows:
[0015] S21: Use the Swin-Transformer backbone network part of the ImageNet-1k pre-trained crack detection network.
[0016] S22: Use the training set to train the entire network.
[0017] Furthermore, the Swin-Transformer backbone network in step S21 contains two basic network modules: the Swin-Transformer module and the Patch-merging module:
[0018] 1), Swin-Transformer module:
[0019] Swin-Transformer is a module that contains an improved self-attention mechanism. It extracts the global dependencies of images and at the same time improves the problem of slow training of the general self-attention module. There are several consecutive pairs of Swin-Transformer blocks in Swin-Transformer. The input first passes through a linear normalization layer and a window-based multi-head self-attention layer in the first block, calculates and adds the residual, then passes through a linear normalization layer and a multi-layer perceptron layer and adds the residual again. The obtained output is given to the next block, and the above process is still calculated. The only difference is that the window-based multi-head self-attention layer is transformed into a shifted window-based multi-head self-attention layer. Let the input be z l-1 , the above process can be formulated as:
[0020]
[0021]
[0022]
[0023]
[0024] The W-MSA divides the image into smaller patches, and then divides the image with a window. Each patch only performs self-attention calculation with the patches within the same window:
[0025]
[0026] Among them, Q, K, and V are vectors obtained by multiplying the embedding of each patch in the image by a transformation matrix after passing through a convolutional neural network. Suppose each image can be divided into H×W patches, and the lengths of the Q, K, and V vectors obtained for each patch are C. And each window range contains M×M patches. Then the computational time complexities of the self-attention mechanism MSA and W-MSA are respectively:
[0027] Ω(MSA) = 4HWC 2 + 2(HW) 2 C
[0028] Ω(W-MSA) = 4HWC 2 + 2M 2 HWC
[0029] The computational complexity of the former is the square of the product of the length and width of the image HW, while the complexity of the latter is linearly related to HW when M 2 << HW, so when the image size is large, W-MSA can significantly accelerate the calculation of self-attention weights and thus speed up the network training process;
[0030] Although the computational complexity of W-MSA is significantly reduced compared to MSA, it limits the self-attention calculation within each window. Therefore, it cannot extract the dependencies between patches in different windows. This problem needs to be solved by SW-MSA. By sliding the entire window, a new window division method is obtained, so as to calculate the attention weights between patches that were originally in different windows;
[0031] 2) Patch-merging module:
[0032] Patch-merging aims to approximately downsample the feature map by merging the patch vectors of the feature map, so as to extract the multi-scale information of the image;
[0033] The specific operation of the Patch-merging module for merging the patch vectors of the feature map is as follows:
[0034] Connect the vectors of length c within adjacent 2×2 ranges to obtain a vector of length 4c, and then input it into the fully connected layer to obtain a vector of final length 2C. Suppose the size of the input feature map is h×w×c, where h and w are the width and height of the image, and c is the number of channels of the image. Then, after Patch-merging, we will get feature maps;
[0035] In the Swin-Transformer module, to solve the problem caused by the inconsistent window partitioning methods of SW-MSA and W-MSA, SW-MSA approximates window sliding by sliding the image in the opposite direction of window sliding, and then uses masks to calculate the W-MSA of each window.
[0036] Furthermore, the specific process of step S22 is as follows:
[0037] The input image first passes through a Ublock to extract features, then through patch embbeding and is input into the Swin-Transformer backbone network. The Swin-Transformer captures the global dependencies of the image and uses patch-merging to output features of different scales. These features are upsampled and fused with the upper-layer features, and then pass through several Ublocks to finally output prediction results of the same size;
[0038] Among them, Ublock is a micro encoder-decoder structure composed of alternating residual blocks and upsampling / downsampling layers. If the size of the input feature map is h×w×c, after the image is input into Ublock and goes through a combination of three residual blocks with two downsamplings at intervals, the size becomes After that, the feature map is restored to its original size through a combination of two upsamplings and residual blocks;
[0039] The network finally outputs four prediction maps of the same size, and uses the deep supervision mechanism to add each prediction to the calculation of the loss and backpropagate the gradient. The deep supervision uses the Focal-Loss loss function, which adds a weight parameter on the basis of the cross-entropy loss function. By adjusting the weight parameter, the network's attention to crack pixels can be increased, and at the same time, the influence of background pixels on the model can be weakened. Let y be the actual type of the pixel, be the prediction result of the model. Then the formula for Focal-Loss is:
[0040]
[0041] Among them, α is the sample balance factor, which is used to balance the uneven ratio of positive and negative samples themselves, and γ makes the loss of easy-to-separate samples much smaller than that of difficult-to-separate samples.
[0042] Further, in step S22, when γ = 0, the Focal-Loss degenerates into the cross-entropy loss function.
[0043] Further, the process of step S3 is as follows:
[0044] S31: Input the road crack pictures sampled from the validation set into the network, and the result outputs 4 prediction maps with dimensions of H OUT ×W OUT ×C OUT where C OUT = 2, indicating that the number of channels of each prediction map is 2. The first channel represents the probability that the pixel is a background pixel, and the second channel represents the probability that the pixel is a crack pixel;
[0045] S32: Select the prediction result of the last layer as the final output, and select the channel according to the value of the label pixel to obtain the prediction result of the model
[0046] S33: Calculate the average value of the following evaluation metrics:
[0047]
[0048]
[0049]
[0050]
[0051]
[0052] where TP represents the number of pixels whose prediction result is a positive sample and the label is also a positive sample, TN represents the number of pixels whose prediction result is a negative sample and the label is also a negative sample, FP represents the number of pixels whose prediction result is a positive sample but the label is a negative sample, FN represents the number of pixels whose prediction result is a negative sample but the label is a positive sample. In each metric, P represents precision, R represents recall rate, F1 represents F1 score, A represents accuracy rate, and IoU represents intersection over union;
[0053] S34: Select the model with the best performance on the validation set by integrating each metric;
[0054] In step S31, when the label pixel value is 0, select the first channel, and when the label pixel is 1, select the second channel.
[0055] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0056] The present invention preprocesses the dataset of the crack detection network to obtain a training set and a validation set; uses the obtained training set to train the crack detection network; uses the obtained validation set to select the best-performing model obtained by training the crack detection network; tests the performance and generalization of the model on different road crack datasets; and applies the processed model to mountain crack detection using transfer learning. The present invention aims at high-precision mountain crack detection application scenarios, studies the extraction methods of different-scale features of pictures, proposes a new crack detection network structure, and elaborates on the new network in detail from the structural level and the formula level. The present invention uses different data experiments to illustrate how the network is applied to specific detection scenarios, and demonstrates the advantages of the new algorithm by comparing the performance with representative methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a structural diagram of the crack detection network;
[0058] Figure 2 It is a schematic diagram of the calculation of W-MSA and SW-MSA;
[0059] Figure 3 It is a detection result diagram of road cracks taken by a drone;
[0060] Figure 4 It is a detection result diagram of mountain cracks. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0062] For better illustration of this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the dimensions of the actual product;
[0063] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0064] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.
[0065] The present invention provides a mountain crack detection method based on an improved self-attention mechanism and transfer learning, mainly including a network training stage and a model testing stage. Specifically:
[0066] Network training stage:
[0067] S1: First, preprocess the dataset. Specifically, resize all training samples and label samples to facilitate loss calculation, and remove samples with excessive noise in the dataset. Then, divide the remaining qualified samples into a training set and a validation set, where the training set is used to train the network, and the validation set is used to select the best-performing model;
[0068] S2: Train the crack detection network. The network structure is as shown in the appendix Figure 1 as follows:
[0069] The specific process of step S2 is as follows:
[0070] S21: Use the Swin-Transformer backbone network part of the ImageNet-1k pre-trained network.
[0071] S22: Use the training set to train the entire network.
[0072] The Swin-Transformer in step S21 mainly includes two basic network modules: the Swin-Transformer block and the Patch-merging module. The following is an explanation of these two modules:
[0073] 1), Swin-Transformer
[0074] Swin-Transformer is a module that contains an improved self-attention mechanism. It can extract the global dependencies of images and at the same time improve the problem of slow training of general self-attention modules. There are several consecutive pairs of Swin-Transformer blocks in Swin-Transformer. The input first passes through a linear normalization layer (LinearNormalization) and a window-based multi-head self-attention layer (W-MSA) in the first block, calculates and adds the residual, then passes through a linear normalization layer and a multi-layer perceptron layer and adds the residual again. The obtained output is given to the next block, and the above process is still calculated. The only difference is that the window-based multi-head self-attention layer is transformed into a shifted window-based multi-head self-attention layer (SW-MSA). Let the input be z l-1 , and the above process can be formulated as follows:
[0075]
[0076]
[0077]
[0078]
[0079] W-MSA divides the image into smaller patches, and then divides the image into smaller patches using windows. Each patch only performs self-attention calculations with the patches in the same window:
[0080]
[0081] Among them, Q, K, and V are vectors obtained by multiplying each patch in the image by the transformation matrix after embedding through the convolutional neural network. Assume that each image can be divided into H×W patches, the length of the Q, K, and V vectors obtained for each patch is C, and each window contains M×M patches. The computational time complexity of the general self-attention mechanism (MSA) and W-MSA are:
[0082] Ω(MSA)=4HWC 2 +2(HW) 2 C
[0083] Ω(W-MSA)=4HWC 2 +2M 2 HWC
[0084] The computational complexity of the former is the square of the product of the length and width of the image HW, while the complexity of the latter is M 2 When <<HW, it is linearly related to HW. Therefore, when the image size is large, W-MSA can significantly accelerate the calculation of self-attention weights and thus speed up the network training process.
[0085] Although W-MSA has significantly reduced computational complexity compared to MSA, it limits the self-attention calculation to each window, so it is unable to extract the dependencies between patches in different windows. To solve this problem, SW-MSA was born. Figure 2 A new window division method can be obtained by sliding the image to the upper left in the middle of the image window, thereby calculating the attention weights between patches in different windows. Furthermore, in order to solve the code writing difficulties caused by the inconsistent window division methods of SW-MSA and W-MSA, SW-MSA improves the window sliding method, and realizes the approximation of window sliding by sliding the image in the opposite direction of the window sliding, and then uses the mask to calculate the W-MSA of each window;
[0086] 2) Patch-merging
[0087] Patch-merging aims to approximately downsample the feature map by merging the patch vectors of the feature map, so as to extract the multi-scale information of the picture. The specific operation is to concatenate the vectors of length c within the adjacent 2×2 range to obtain a vector of length 4c, and then input it into the fully connected layer to obtain a final vector of length 2C. Suppose the size of the input feature map is h×w×c, where h and w are the width and height of the image, and c is the number of channels of the image. After experiencing Patch-merging, we will get the feature map;
[0088] The specific process of step S22 is as follows: The input picture first passes through a Ublock to extract features, then passes through patchembbeding and is input into the Swin-Transformer backbone network. The Swin-Transformer captures the global dependencies of the image and uses patch-merging to output features of different scales. These features are upsampled and fused with the upper-layer features and then pass through several Ublocks, and finally output prediction results of the same size.
[0089] Among them, Ublock is a micro encoder-decoder structure composed of alternating residual blocks and upsampling / downsampling layers. If the size of the input feature map is h×w×c, after the image is input into Ublock and passes through the combination of three residual blocks with two downsamplings in between, the size becomes After that, the feature map is restored to its original size through the combination of two upsamplings and residual blocks.
[0090] The network finally outputs four prediction maps of the same size, and uses the deep supervision mechanism to add each prediction to the calculation of the loss and backpropagate the gradient. Deep supervision uses the Focal-Loss loss function, which adds a weight parameter to the cross-entropy loss function. By adjusting the weight parameter, the network's attention to crack pixels can be increased, and at the same time, the influence of background pixels on the model can be weakened. Let y be the actual type of the pixel, be the prediction result of the model, then the formula of Focal-Loss is:
[0091]
[0092] Among them, α is the sample balance factor, which is used to balance the uneven proportion of positive and negative samples themselves. γ makes the loss of easy-to-separate samples much smaller than that of difficult-to-separate samples. When γ = 0, Focal-Loss will degenerate into the cross-entropy loss function.
[0093] Model testing phase:
[0094] S3: Use the data validation set to select the model with the best performance obtained by training the network;
[0095] S4: Test the performance and generalization of the model on different road crack datasets;
[0096] S5: Apply the crack detection model to mountain crack detection using transfer learning methods;
[0097] The specific process of step S3 is as follows:
[0098] S31: Input the road crack images sampled from the validation set into the network, and the result outputs 4 prediction maps with dimensions of H OUT ×W OUT ×C OUT where C OUT = 2, indicating that the number of channels of each prediction map is 2. The first channel represents the probability that the pixel is a background pixel, and the second channel represents the probability that the pixel is a crack pixel;
[0099] S32: Select the prediction result of the last layer as the final output, and select the channel according to the value of the label pixel (select the first channel when the label pixel value is 0, and select the second channel when the label pixel is 1) to obtain the prediction result of the model
[0100] S33: Calculate the average value of the following evaluation metrics:
[0101]
[0102]
[0103]
[0104]
[0105]
[0106] Among them, TP (True Positive) represents the number of pixels where the prediction result is a positive sample and the label is also a positive sample, TN (True Negative) represents the number of pixels where the prediction result is a negative sample and the label is also a negative sample, FP (False Positive) represents the number of pixels where the prediction result is a positive sample but the label is a negative sample, and FN (False Negative) represents the number of pixels where the prediction result is a negative sample but the label is a positive sample. Among each metric, P represents Precision, R represents Recall, F1 represents F1-Score, A represents Accuracy, and IoU represents Intersection over Union;
[0107] S34: Select the model with the best performance on the validation set by integrating various metrics.
[0108] Implementation Modes and Performance Comparison
[0109] The present invention uses the Crack500 road crack public dataset as the training dataset to train the crack detection network. For the processing of the dataset, first, all training samples and label samples are resized to 512×512 in size, and samples with excessive noise are removed. Then, the remaining qualified samples are divided into a training set containing 1,840 pictures and a validation set of 1,124 pictures (including corresponding labels). In the test phase, the present invention uses the Crack500 validation dataset to test the model performance and select the model with the best performance.
[0110] Regarding the performance comparison, the present invention designs and implements detection experiments of the model on road cracks captured by drones to test the generalization of the model. The results are as shown in the appendix Figure 3 As shown, the model can well overcome the influence of factors such as shadows, background interference, and complex shapes, and has considerable generalization. Then, the present invention applies the model to mountain crack detection using the transfer learning method, and the results are as shown in the appendix Figure 4 As shown. In the performance comparison, the present invention designs a comparison experiment on the Crack500 test set data, and selects three methods, namely FPHBN, SegNet, and CrackUnet, as benchmarks. Among them, FPHBN belongs to the method based on upsampling and downsampling, while SegNet and CrackUnet belong to the semantic segmentation methods based on the encoder-decoder. The experimental results are shown in Table 1. It can be seen that the crack detection method based on the improved self-attention mechanism and transfer learning adopted by the present invention effectively improves the detection accuracy of cracks and greatly saves the manpower and material resources for mountain crack early warning. The technologies mainly involved in the present invention are wireless charging, multi-agent cooperation, game theory, multi-agent reinforcement learning algorithms, etc.
[0111] Table 1 Results on the Crack500 Test Set
[0112] IoU F1-Score Precision Recall Accuracy FPHBN 0.492 0.687 0.614 0.675 0.908 SegNet 0.507 0.662 0.670 0.728 0.926 CrackUnet 0.541 0.688 0.6963 0.719 0.967 The present invention 0.569 0.712 0.680 0.787 0.968
[0113] The same or similar reference numerals correspond to the same or similar components;
[0114] The descriptions of the positional relationships in the drawings are only for illustrative purposes and should not be construed as limitations on this patent;
[0115] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. A mountain crack detection method based on an improved self-attention mechanism and transfer learning, characterized in that, It includes the following steps: S1: Preprocess the dataset of the crack detection network to obtain a training set and a validation set; S2: Use the training set obtained in step S1 to train the crack detection network. The process is as follows: S21: Use the ImageNet-1k pre-trained Swin-Transformer backbone network part of the crack detection network; The Swin-Transformer backbone network contains two basic network modules: the Swin-Transformer module and the Patch-merging module: 1). The Swin-Transformer module: Swin-Transformer is a module that contains an improved self-attention mechanism. It extracts the global dependencies of images while improving the problem of slow training of the general self-attention module. There are several consecutive pairs of Swin-Transformer blocks in Swin-Transformer. The input first passes through a layer normalization layer and a window-based multi-head self-attention layer in the first block, and the residual is added. Then, it passes through a layer normalization layer and a multi-layer perceptron layer, and the residual is added again. The obtained output is given to the next block, and the above process is still calculated. The only difference is that the window-based multi-head self-attention layer is transformed into a shifted window-based multi-head self-attention layer. Let the input be , the above process can be formulated as follows: Among them, W-MSA represents the window-based multi-head self-attention layer; LN represents the layer normalization; MLP represents the multi-layer perceptron; SW-MSA represents the shifted window-based multi-head self-attention layer; is the input to the first block in the l-th Swin-Transformer block pair; is the output of the l-th window-based multi-head self-attention layer; is the output of the first block in the l-th Swin-Transformer block pair; is the output of the l-th shifted window-based multi-head self-attention layer; is the output of the next block in the l-th Swin-Transformer block pair; W-MSA divides the picture into smaller patches, and then divides the picture with a window. Each patch only performs self-attention calculation with the patches within the same window: Among them, Q, K, and V are vectors obtained by multiplying the embedding of each patch in the picture by a convolutional neural network and then multiplying by a transformation matrix. d is the dimension size of the vector K; B is a preset bias term. Suppose each picture can be divided into patches, and the lengths of the Q, K, and V vectors obtained for each patch are . Each window range contains patches. Then the computational time complexities of the self-attention mechanism MSA and W-MSA are respectively: 2). The Patch-merging module: Patch-merging aims to approximately downsample the feature map by merging the patch vectors of the feature map, so as to extract the multi-scale information of the picture; S22: Use the training set to train the entire network; The input picture first extracts features through a layer of Ublock, and then after patch embbeding, it is input into the Swin-Transformer backbone network. The Swin-Transformer captures the global dependencies of the image and uses patch-merging to output features of different scales. These features are upsampled and fused with the upper-layer features, and then pass through several Ublocks to finally output prediction results of the same size; Among them, Ublock is a micro encoder-decoder structure composed of alternating residual blocks and upsampling layers. If the size of the input feature map is , after the image is input into Ublock and undergoes a combination of three residual blocks with two downsamplings in between, the size becomes ; then, the feature map is restored to its original size through a combination of two upsamplings and residual blocks; The network finally outputs four prediction maps of the same size, and uses the deep supervision mechanism to add each prediction to the calculation of the loss and backpropagate the gradient. The deep supervision adopts the Focal-Loss loss function, which adds a weight parameter on the basis of the cross-entropy loss function. By adjusting the weight parameter, the attention of the network to crack pixels can be increased, and at the same time, the influence of background pixels on the model can be weakened. Let be the actual type of the pixel, be the prediction result of the model, then the formula of Focal-Loss is: Among them, is the sample balance factor, which is used to balance the uneven proportion of positive and negative samples themselves, making the loss of easy-to-separate samples much smaller than that of difficult-to-separate samples; S3: Use the validation set obtained in step S1 to select the model with the best performance obtained by training the crack detection network; S4: Test the performance and generalization of the model on different road crack datasets; S5: Use the transfer learning method to apply the model processed in step S4 to mountain crack detection.
2. The mountain crack detection method based on the improved self-attention mechanism and transfer learning according to claim 1, wherein The process of step S1 is: Resize all training samples and label samples to facilitate loss calculation, and remove samples with too much noise in the dataset. Then, divide the remaining qualified samples into a training set and a validation set. The training set is used to train the network, and the validation set is used to select the model with the best performance.
3. The method for detecting mountain cracks based on the improved self-attention mechanism and transfer learning according to claim 1, characterized in that The specific operation of the Patch-merging module to merge the patch vectors of the feature map is: Connect adjacent vectors with length within to obtain a vector with length , then input it into a fully connected layer to obtain a vector with a final length of . Assume the size of the input feature map is , where h and w are the width and height of the image, and c is the number of channels of the image. Then, after undergoing Patch-merging, a feature map of will be obtained.
4. The mountain crack detection method based on the improved self-attention mechanism and transfer learning according to claim 3, characterized in that, In the Swin-Transformer module, to solve the problem caused by the inconsistent window division methods of SW-MSA and W-MSA, SW-MSA approximates the window sliding by sliding the picture in the opposite direction of the window sliding, and then uses a mask to calculate the W-MSA of each window.
5. The method for detecting mountain cracks based on the improved self-attention mechanism and transfer learning according to claim 1, wherein In step S22, when occurs, the Focal-Loss degenerates into the cross-entropy loss function.
6. The method for detecting mountain cracks based on the improved self-attention mechanism and transfer learning according to claim 5, characterized in that The process of step S3 is: S31: Input the road crack pictures sampled from the validation set into the network, and the result outputs 4 prediction maps with the size of . Among them, indicates that the number of channels of each prediction map is 2. The first channel represents the probability that the pixel is a background pixel, and the second channel represents the probability that the pixel is a crack pixel; S32: Select the prediction result of the last layer as the final output, and select the channel according to the value of the labeled pixel to obtain the prediction result of the model ; S33: Calculate the average value of the following evaluation metrics: Among them, TP represents the number of pixels where the predicted result is a positive sample and the label is also a positive sample, TN represents the number of pixels where the predicted result is a negative sample and the label is also a negative sample, FP represents the number of pixels where the predicted result is a positive sample but the label is a negative sample, FN represents the number of pixels where the predicted result is a negative sample but the label is a positive sample. In each metric, P represents precision, R represents recall, represents the F1 score, A represents accuracy, and IoU represents the intersection over union; S34: Select the model with the best performance on the validation set by synthesizing each metric.
7. The method for detecting mountain cracks based on the improved self-attention mechanism and transfer learning according to claim 6, wherein In step S31, when the label pixel value is 0, select the first channel, and when the label pixel is 1, select the second channel.
Citation Information
Patent Citations
Road crack detection method based on multi-source Unet + Attention network migration
CN111986164A
Bridge crack detection method based on improved PSPNet network
CN112560895A