Efficient pavement crack detection system and method fusing Swin-Transformer and yolov8
By integrating Swin-Transformer and yolov8 detection systems, the problem of crack detection on rural roads under complex background is solved, and efficient and accurate crack detection is achieved, which is suitable for multi-scale crack detection tasks.
Patent Information
- Application Number
- CN202510086441.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to effectively detect pavement cracks on rural roads in complex contexts, especially in the presence of a large number of disturbances and multi-scale cracks, resulting in inaccuracy and inefficiency of detection.
An efficient road surface crack detection system integrating Swin-Transformer and yolov8 is adopted. This system introduces the Swin-Transformer network and SPPF_Avg module in the Backbone structure, the ECSA adaptive attention mechanism module is introduced into the Neck structure, and the DBB reparameterization module is introduced into the Head structure to improve the accuracy of crack feature extraction and detection and anti-interference ability.
It significantly improves the accuracy and efficiency of road surface crack detection, can effectively adapt to the detection tasks of complex backgrounds and multi-scale cracks, and is suitable for crack detection on rural roads.
Smart Images

Figure CN120163761A_ABST
Abstract
Description
Technical Field:
[0001] The present invention belongs to the technical field of automatic detection of road surface cracks, and particularly relates to an efficient road surface crack detection system and method integrating Swin-Transformer and yolov8. Background Art:
[0002] Cracks are the most common type of road surface diseases, posing a serious threat to the service life of roads and the driving safety of vehicles. Compared with high-grade urban roads, rural roads have a lower maintenance level, a wider distribution, and a more complex scenario. The surface of rural roads is often shaded by roadside branches or other objects, which increases the complexity of the image background and brings additional challenges to crack detection. In addition, there are a large number of small interfering objects similar to weeds, branches, and dirt on rural road surfaces that are similar to crack textures, which easily lead to misjudgment problems in the detection network. During the detection of rural road surface diseases, cracks generated at different time periods and with different scales may exist on the same image, further increasing the detection difficulty. Currently, commonly used traditional road surface crack detection algorithms have limited application scenarios, single tasks, and poor anti-interference ability, so they cannot complete the crack detection task of rural roads with complex road surface conditions.
[0003] The road surface crack detection method based on deep learning uses various advanced sensors and technologies to quickly and accurately detect the highway surface. By collecting a large amount of data and applying machine learning algorithms, feature extraction and pattern recognition of cracks are carried out to achieve automatic crack detection. The detection method based on machine learning has many advantages. This method has made remarkable progress in the field of image processing and plays an important role in road surface crack detection. They can process a large amount of image data and significantly improve the accuracy and efficiency of crack detection by learning features and patterns.
[0004] Patent 202410911824.8 discloses "A Highway Crack Detection Method Based on Improved YOLOv8". In this method, a lightweight convolutional ASDown module is added to the Backbone structure of the yolov8 network, and a small target detection layer including a C2f module is added to the Neck structure. Although it more comprehensively integrates feature maps at different levels and captures multi-scale context information, the ASDown module and C2f module need to perform feature integration and channel transformation, which will increase the computational complexity of the model, resulting in an increase in the time cost of model training and inference. Improper use of these modules will lead to overfitting, thus affecting the generalization ability of the model. In the case of large-scale changes in the target, the model is prone to target missed detection or misdetection.
[0005] Patent 202410987176.4 discloses "A Road Crack Detection Method, Medium and Product". This method replaces the feature extraction network backbone of the yolov8 model with the lightweight convolutional neural network MobileNetV3, and embeds the coordinate attention mechanism CA module in the lightweight convolutional neural network MobileNetV3; adds a small target detection layer and an SE module at the neck end; although MobileNetV3 makes the overall model lightweight, it cannot fully retain the performance of the original backbone. Moreover, embedding the CA module and the SE module will also increase the computational complexity and the number of parameters.
[0006] Therefore, it is particularly important to select a deep learning model suitable for rural road crack detection and develop a rural road surface crack detection algorithm with strong anti-interference ability and multi-scale detection ability.
[0007] The information disclosed in this background section is only intended to increase the understanding of the overall background of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes prior art known to those of ordinary skill in the art. Summary of the Invention:
[0008] The purpose of the present invention is to provide an efficient road surface crack detection system and method that integrates Swin-Transformer and yolov8, so as to overcome the defects in the above-mentioned prior art.
[0009] To achieve the above purpose, the present invention provides an efficient road surface crack detection system that integrates Swin-Transformer and yolov8, including a Backbone structure, a Neck structure, and a Head structure; among them, the Backbone structure is used to extract features from the input image; the Neck structure is used to further process and fuse the features extracted by the Backbone; the Head structure is responsible for the final object detection task and generates the prediction results of each object.
[0010] A working method of an efficient road surface crack detection system that integrates Swin-Transformer and yolov8, the steps of which are:
[0011] S01. Use the perspective of the vehicle-mounted device during driving to take road surface crack images, and use the labelimg annotation tool to annotate the types and positions of cracks in the collected road surface crack images to create a YOLO format dataset; at the same time, the YOLO format dataset is divided into a training set and a test set according to a ratio of 8:2.
[0012] S02. Use the Mosaic data enhancement operation on the YOLO format dataset created in step S01, and add the sorted dataset to the existing RDD2022 public dataset;
[0013] S03, build the ST-yolov8 model, introduce the Swin-transformer network and SPPF_Avg module into the Backbone structure of the ST-yolov8 model, introduce the ECSA adaptive attention mechanism module into the Neck structure of the ST-yolov8 model, and introduce the DBB reparameterization module into the Head structure of the ST-yolov8 model;
[0014] S04. Input the RDD2022 public data set in step S02 into the ST-yolov8 model, iteratively train the ST-yolov8 model and obtain the optimal weight file best.pt, use the optimal weight file best.pt to identify the pavement crack image to be detected, detect the location and category of the cracks, and output the detection results.
[0015] Preferably, in the technical solution, in step S01, the vehicle-mounted device is a high-speed industrial matrix camera, the maximum resolution of the high-speed industrial matrix camera is set to 1280×1024, and the exposure time is set to automatic.
[0016] Preferably, in the technical solution, in step S01, a pavement crack image is collected, the image size is padded to 1280×1280, and then compressed to 512×512, and the pavement crack image is labeled with type and position using the labelimg labeling tool; wherein the labeled content includes the bounding box position of the crack and the category to which it belongs.
[0017] Preferably, in the technical solution, in step S02, the Mosaic data enhancement operation includes image rotation, translation, brightness change, noise addition and cropping.
[0018] Preferably, in the technical solution, in step S03, a Swin-Transformer network is used as a crack feature extraction part in the Backbone structure, and a SPPF_Avg module is used to perform crack feature fusion; wherein the Swin-Transformer network includes an Input module, a Patch Partition module, a Liner Embedding module, a Swin-Transformerblock module, and a patch merging module; the specific process of crack feature extraction is:
[0019] 3.1. Output a 512×512×3 pavement crack image from the Input module to the Patch Partition module for chunking. That is, every 4×4 adjacent pixels form a Patch, which is flattened in the channel direction and used as the input to the Linear Embeding module.
[0020] 3.2. The Linear Embeding module performs a linear transformation on the channel data of each pixel, changing the number of channels to the corresponding number required by the Swin-Transformer network and serving as the input to the Swin-Transformer block module.
[0021] 3.3. The Swin-Transformer block module extracts features from the pavement crack image. The Swin-Transformer block module contains two sub-modules, namely the Window Multi-Head Self-Attention (WMSA) module and the Shifted Window Multi-Head Self-Attention (SWMSA) module. Each sub-module contains two normalization layers, an attention module, and an MLP layer, and uses residual connections. The Swin-Transformer block module captures local and global information in the pavement crack image through the self-attention mechanism within local windows and the shifted window mechanism. The operation process of the Swin-Transformer block module is as follows:
[0022]
[0023] Among them, WMSA() represents the Window Multi-Head Self-Attention mechanism, SWMSA() represents the Shifted Window Multi-Head Self-Attention mechanism, MLP() represents the non-linear feature transformation, and LN() represents the normalization operation. represents the output of WMSA(). represents the output of SWMSA(), y l represents the output of MLP() in the WMSA module, y (l+1) represents the output of MLP() in the SWMSA module, y (l-1) represents the initial input of the Swin-Transformer block module, and l represents the ordinal number of the Swin-Transformer block module, where l≥1.
[0024] The formula for multi-head self-attention is as follows:
[0025]
[0026] Among them, Attention() represents multi-head self-attention, Q represents the query value, K TLet \(K\) denote the key value, \(V\) denote the value matrix, \(SoftMax()\) denote the weighted sum of \(V\), \(d\) denote the dimension, and \(B\) denote the offset;
[0027] 3.4. The Patch Merging module performs downsampling operations on the crack features, reduces the resolution of the feature map, and adjusts the number of channels, thereby forming a hierarchical design.
[0028] Preferably, in the technical solution, in step S03, the SPPF_Avg module has two major branches. One is the Max_Pool branch, and the other is the Avg_Pool branch. The Max_Pool branch extracts global and local features in the pavement crack image through max pooling and realizes feature fusion. The Avg_Pool branch compensates for the feature loss caused by the Max_Pool branch through average pooling, and the crack feature maps of the two branches are concatenated to realize crack feature fusion. The specific calculation process of crack feature fusion is as follows:
[0029]
[0030] Among them, \(Conv()\) is the convolution operation, \(Concat()\) is the feature fusion operation, \(Y1\) and \(Y2\) respectively represent the outputs of the Max_Pool branch and the Avg_Pool branch, \(M\) (1~3) represents the output after 1 to 3 max poolings, and \(A\) (1~3) represents the output after 1 to 3 average poolings, and \(Y\) represents the output of the fusion of the two branches.
[0031] Preferably, in the technical solution, in step S03, the ECSA adaptive attention mechanism module adaptively increases the weight of crack information and reduces the weight of interference information in the pavement crack image according to the detected crack features. The operation process of the ECSA adaptive attention mechanism module is as follows:
[0032] F' = F × \(\hat{F}\) C × \(\tilde{F}\) S ,
[0033] Among them, \(F\) is the effective feature map extracted by the Swin-Transformer network, \(\hat{F}\) C is the feature map given channel weights, \(\tilde{F}\) S is the feature map given spatial weights, and \(F'\) is the feature map output by the ECSA adaptive attention mechanism module.
[0034] Preferably, in the technical solution, in step S03, a DBB reparameterization module is added to the Head structure to prevent misdetection of crack types. The DBB reparameterization module adopts a separated branch structure, which includes a single-branch structure and a multi-branch structure. The multi-branch structure is adopted in the training stage of the ST-yolov8 model to obtain better weight parameters, and the multi-branch structure is converted into a single-branch structure in the inference stage. The specific operation process is as follows:
[0035]
[0036] Among them, k and b represent the weight and bias after crack feature fusion, k1 and b1 represent the weight and bias of the 1×1 convolution, k k and b k represent the weight and bias of the k×k convolution.
[0037] Preferably, in the technical solution, in step S04, the iterative training process of the ST-yolov8 model includes setting the initial learning rate of the ST-yolov8 model to 0.01, automatically adjusting the learning rate during training using the learning rate adjustment strategy, setting the number of training rounds to 300, using Stochastic Gradient Descent (SGD) as the optimizer, setting the momentum to 0.937, and setting the weight decay to 0.0005, and continuously updating the weights and biases of the ST-yolov8 model.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] It solves the problem of insufficient extraction of crack features in complex backgrounds, realizes a significant improvement in the accuracy of the model in object detection, can well adapt to the detection of objects with diverse scales such as road surface diseases, and is suitable for crack detection on rural roads. Description of the Drawings:
[0040] Figure 1 is the flow chart of the efficient road surface crack detection method that combines Swin-Transformer and yolov8 of the present invention;
[0041] Figure 2 is the schematic diagram of the network structure of the ST-yolov8 model of the present invention;
[0042] Figure 3 is the schematic diagram of the network structure of the Swin-transformer in the Backbone structure of the ST-yolov8 model of the present invention;
[0043] Figure 4 is the schematic diagram of the structure of the SPPF_Avg module in the Backbone structure of the ST-yolov8 model of the present invention;
[0044] Figure 5It is a schematic diagram of the structure of the ECSA adaptive attention mechanism module in the Neck structure of the ST-yolov8 model of the present invention;
[0045] Figure 6 It is a schematic diagram of the DBB re-parameterization module structure in the Head structure of the ST-yolov8 model of the present invention. Specific implementation method:
[0046] The specific embodiments of the present invention are described in detail below, but it should be understood that the protection scope of the present invention is not limited by the specific embodiments.
[0047] Unless explicitly stated otherwise, throughout the specification and claims, the term “comprise” or variations such as “include” or “comprising”, etc., will be understood to include the stated elements or components but not to exclude other elements or components.
[0048] The present invention provides an efficient pavement crack detection system integrating Swin-Transformer and yolov8, comprising a Backbone structure, a Neck structure and a Head structure; wherein the Backbone structure is used to extract features from an input image; the Neck structure is used to further process and fuse the features extracted by the Backbone; the Head structure is responsible for the final target detection task and generates a prediction result for each target.
[0049] like Figure 1 As shown, a working method of an efficient pavement crack detection system integrating Swin-Transformer and yolov8, the steps are as follows:
[0050] S01. Use a high-speed industrial matrix camera to capture road crack images during driving, set the highest resolution of the high-speed industrial matrix camera to 1280×1024, and set the exposure time to automatic; fill the acquired road crack image size to 1280×1280, and then compress it to 512×512, use the labelimg annotation tool to annotate the acquired road crack image with the type and location of the cracks, where the annotation content includes the bounding box location and category of the cracks, and create a YOLO format data set; at the same time, the YOLO format data set is divided into a training set and a test set in a ratio of 8:2; in the YOLO format data set, the road crack images are divided into four categories: D00 longitudinal shape, D10 transverse shape, D20 crocodile shape, and D40 pit shape;
[0051] S02. Apply Mosaic data augmentation operations to the YOLO format dataset created in step S01. The Mosaic data augmentation operations include image rotation, translation, brightness change, adding noise, and cropping. Add the sorted dataset to the existing RDD2022 public dataset.
[0052] S03. Build the ST-yolov8 model. As Figure 2 shown, introduce the Swin-transformer network and the SPPF_Avg module into the Backbone structure of the ST-yolov8 model, introduce the ECSA adaptive attention mechanism module into the Neck structure of the ST-yolov8 model, and introduce the DBB reparameterization module into the Head structure of the ST-yolov8 model.
[0053] In the Backbone structure, use the Swin-Transformer network as the crack feature extraction part and use the SPPF_Avg module for crack feature fusion. The Swin-Transformer network includes the Input module, PatchPartition module, Liner Embedding module, Swin-Transformer block module, and patch merging module. The specific process of crack feature extraction is as follows:
[0054] 3.1. Output a 512×512×3 pavement crack image from the Input module to the Patch Partition module for block division, that is, every 4×4 adjacent pixels are a Patch, and flatten in the channel direction as the input of the Linear Embeding module.
[0055] 3.2. The Linear Embeding module performs a linear transformation on the channel data of each pixel and changes the number of channels to the corresponding number of channels required by the Swin-Transformer network as the input of the Swin-Transformer block module.
[0056] 3.3. The Swin-Transformer block module extracts features from the pavement crack image. The Swin-Transformer block module contains two sub-modules, as Figure 3As shown, they are the Window Multi-Head Self-Attention (WMSA) module and the Shifted Window Multi-Head Self-Attention (SWMSA) module respectively. Each sub-module contains two normalization layers, an attention module, and an MLP layer, and residual connections are used; the Swin-Transformer block module captures local and global information in the pavement crack image through the self-attention mechanism within local windows and the shifted window mechanism; the operation process of the Swin-Transformer block module is as follows:
[0057]
[0058]
[0059] Among them, WMSA() represents the Window Multi-Head Self-Attention mechanism, SWMSA() represents the Shifted Window Multi-Head Self-Attention mechanism, MLP() represents the non-linear feature transformation, and LN() represents the normalization operation. represents the output of WMSA(). represents the output of SWMSA(), and y l represents the output of MLP() in the WMSA module, and y (l+1) represents the output of MLP() in the SWMSA module; y (l-1) represents the initial input of the Swin-Transformer block module, that is, the result after the Linear Embeding module performs a linear transformation on the channel data of each pixel; l represents the ordinal number of the Swin-Transformer block module, and l ≥ 1.
[0060] The formula for multi-head self-attention is as follows:
[0061]
[0062] Among them, Attention() represents multi-head self-attention, Q represents the query value, and K T represents the key value, V represents the value matrix, SoftMax() represents the weighted sum of V, d represents the dimension, and B represents the offset.
[0063] 3.4. The Patch Merging module performs downsampling on the crack features, reduces the resolution of the feature map, and adjusts the number of channels, thereby forming a hierarchical design.
[0064] Such as Figure 4As shown, the SPPF_Avg module has two major branches. One is the Max_Pool branch, and the other is the Avg_Pool branch. The Max_Pool branch extracts global and local features in the pavement crack image through max pooling and realizes feature fusion. The Avg_Pool branch compensates for the feature loss caused by the Max_Pool branch through average pooling. The crack feature maps of the two branches are concatenated to achieve crack feature fusion. The specific calculation process of crack feature fusion is as follows:
[0065]
[0066] Among them, Conv() is the convolution operation, Concat() is the feature fusion operation, Y1 and Y2 respectively represent the outputs of the Max_Pool branch and the Avg_Pool branch, and M (1~3) represents the output after 1 to 3 max poolings, and A (1~3) represents the output after 1 to 3 average poolings, and Y represents the output of the fusion of the two branches;
[0067] As Figure 5 shown, the ECSA adaptive attention mechanism module adaptively increases the weight of crack information and reduces the weight of interference information in the pavement crack image according to the detected crack features. The operation process of the ECSA adaptive attention mechanism module is as follows:
[0068] F' = F × F C × F S ,
[0069] Among them, F is the effective feature map extracted by the Swin-Transformer network, F C is the feature map with channel weights assigned, F S is the feature map with spatial weights assigned, and F′ is the feature map output by the ECSA adaptive attention mechanism module;
[0070] As Figure 6 shown, in step S03, a DBB reparameterization module is added to the Head structure to prevent misdetection of crack types. The DBB reparameterization module adopts a split-branch structure, and the split-branch structure includes a single-branch structure and a multi-branch structure. During the training stage of the ST-yolov8 model, a four-branch structure is adopted to obtain better weight parameters, and during the inference stage, the multi-branch structure is converted into a single-branch structure. Its specific operation process is as follows:
[0071]
[0072] Among them, k and b represent the weights and biases after crack feature fusion, k1 and b1 represent the weights and biases of the 1×1 convolution, k k and bk Represent the weights and biases of a k×k convolution;
[0073] S04. Input the RDD2022 public dataset in step S02 into the ST-yolov8 model, perform iterative training on the ST-yolov8 model, and obtain the optimal weight file best.pt;
[0074] The iterative training process of the ST-yolov8 model includes setting the initial learning rate of the ST-yolov8 model to 0.01, automatically adjusting the learning rate during training using a learning rate adjustment strategy, setting the number of training rounds to 300, using Stochastic Gradient Descent (SGD) as the optimizer, setting the momentum to 0.937, and setting the weight decay to 0.0005, and continuously updating the weights and biases of the ST-yolov8 model;
[0075] The learning rate adjustment strategy adopts a linear warmup + cosine annealing strategy; in the linear warmup stage, the initial learning rate linearly increases from a very small value to the set maximum learning rate; in the cosine annealing stage, when the linear warmup stage ends, the learning rate gradually decays to the lowest value according to the cosine function;
[0076] The Stochastic Gradient Descent (SGD) optimizer continuously adjusts the parameters of the ST-yolov8 model to minimize the loss function, thereby improving the performance of the ST-yolov8 model in object detection tasks;
[0077] Use the optimal weight file best.pt to identify the pavement crack image to be detected, detect the position and category of the crack, and output the detection result.
[0078] The foregoing description of the specific exemplary embodiments of the present invention is for the purposes of illustration and exemplification. These descriptions are not intended to limit the present invention to the precise forms disclosed, and obviously, many changes and variations are possible in light of the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present invention and its practical applications, so that those skilled in the art can implement and utilize the various different exemplary embodiments of the present invention, as well as various different selections and changes. The scope of the present invention is intended to be defined by the claims and their equivalents.
Claims
1. A working method of an efficient pavement crack detection system integrating Swin-Transformer and yolov8, characterized in that: The system includes Backbone structure, Neck structure and Head structure. The Backbone structure is used to extract features from the input image. The Neck structure is used to further process and fuse the features extracted by Backbone. The Head structure is responsible for the final target detection task and generates the prediction results for each target. The steps are: S01. Use the vehicle-mounted equipment to take road crack images during driving, use the labelimg annotation tool to annotate the types and locations of the collected road crack images, and create a YOLO format dataset; at the same time, the YOLO format dataset is divided into a training set and a test set in a ratio of 8:2; S02. Use the Mosaic data enhancement operation on the YOLO format dataset created in step S01, and add the sorted dataset to the existing RDD2022 public dataset; S03, build the ST-yolov8 model, introduce the Swin-transformer network and SPPF_Avg module into the Backbone structure of the ST-yolov8 model, introduce the ECSA adaptive attention mechanism module into the Neck structure of the ST-yolov8 model, and introduce the DBB reparameterization module into the Head structure of the ST-yolov8 model; S04. Input the RDD2022 public data set in step S02 into the ST-yolov8 model, iteratively train the ST-yolov8 model and obtain the optimal weight file best.pt, use the optimal weight file best.pt to identify the pavement crack image to be detected, detect the location and category of the cracks, and output the detection results.
2. According to claim 1, the working method of the efficient pavement crack detection system integrating Swin-Transformer and yolov8 is characterized in that: In step S01, the vehicle-mounted device is a high-speed industrial matrix camera, the maximum resolution of the high-speed industrial matrix camera is set to 1280×1024, and the exposure time is set to automatic.
3. The working method of the efficient pavement crack detection system integrating Swin-Transformer and yolov8 according to claim 1 is characterized in that: In step S01, a pavement crack image is collected, the image size is padded to 1280×1280, and then compressed to 512×512, and the pavement crack image is labeled with type and position using the labelimg labeling tool; The annotation content includes the bounding box location of the crack and the category to which it belongs.
4. The working method of the efficient pavement crack detection system integrating Swin-Transformer and yolov8 according to claim 1 is characterized in that: In step S02, the mosaic data enhancement operation includes image rotation, translation, brightness change, noise addition and cropping.
5. The working method of the efficient pavement crack detection system integrating Swin-Transformer and yolov8 according to claim 1 is characterized in that: In step S03, the Swin-Transformer network is used as the crack feature extraction part in the Backbone structure, and the SPPF_Avg module is used to fuse the crack features; the Swin-Transformer network includes an Input module, a Patch Partition module, a Liner Embedding module, a Swin-Transformer block module, and a patch merging module; the specific process of crack feature extraction is as follows: 3.
1. Output the 512×512×3 road crack image from the Input module to the Patch Partition module for block division, that is, each 4×4 adjacent pixels is a patch, flattened in the channel direction, and used as the input of the Linear Embedding module; 3.
2. The Linear Embedding module performs a linear transformation on the channel data of each pixel, changing the number of channels to the corresponding number of channels required by the Swin-Transformer network; this serves as the input of the Swin-Transformer block module; 3.
3. The Swin-Transformer block module extracts features from the pavement crack image. The Swin-Transformer block module contains two submodules, namely the window multi-head self-attention mechanism WMSA module and the sliding window multi-head self-attention mechanism SWMSA module. Each submodule contains two normalization layers, an attention module and an MLP layer, and uses residual connections. The Swin-Transformer block module captures local and global information in the pavement crack image through the self-attention mechanism and shift window mechanism within the local window. The operation process of the Swin-Transformer block module is as follows: Among them, WMSA() represents the window multi-head self-attention mechanism, SWMSA() represents the sliding window multi-head self-attention mechanism, MLP() represents the nonlinear feature transformation, and LN() represents the normalization operation. represents the output of WMSA(), represents the output of SWMSA(), y l Represents the output of MLP() in the WMSA module, y (l+1) represents the output of MLP() in the SWMSA module, y (l-1) represents the initial input of the Swin-Transformer block module, l represents the ordinal number of the Swin-Transformer block module, l≥1; The multi-head self-attention formula is as follows: Among them, Attention() represents multi-head self-attention, Q represents the query value, K T represents the key value, V represents the value matrix, SoftMax() represents the weighted sum of V, d represents the dimension, and B represents the offset; 3.
4. The Patch Merging module downsamples the crack features, reduces the resolution of the feature map, and adjusts the number of channels to form a hierarchical design.
6. The working method of the efficient pavement crack detection system integrating Swin-Transformer and yolov8 according to claim 5 is characterized in that: In step S03, the SPPF_Avg module has two major branches, one is the Max_Pool branch, and the other is the Avg_Pool branch. The Max_Pool branch extracts the global features and local features in the pavement crack image through maximum pooling and realizes feature fusion. The Avg_Pool branch compensates for the feature loss caused by the Max_Pool branch through average pooling. The crack feature maps of the two branches are spliced to realize crack feature fusion. The specific calculation process of crack feature fusion is as follows: Among them, Conv() is a convolution operation, Concat() is a feature fusion operation, Y1 and Y2 represent the outputs of the Max_Pool branch and the Avg_Pool branch respectively, and M (1~3) Represents the output after 1 to 3 maximum poolings, A (1~3) represents the output after 1 to 3 average poolings, and Y represents the output of the fusion of the two branches.
7. The working method of the efficient pavement crack detection system integrating Swin-Transformer and yolov8 according to claim 1 is characterized in that: In step S03, the ECSA adaptive attention mechanism module adaptively increases the weight of crack information and reduces the weight of interference information in the pavement crack image according to the detected crack features. The operation process of the ECSA adaptive attention mechanism module is as follows: F'=F×F C ×F S , Among them, F is the effective feature map extracted by the Swin-Transformer network, F C is the feature map assigned channel weights, F S is the feature map assigned with spatial weights, and F′ is the feature map output by the ECSA adaptive attention mechanism module.
8. The working method of the efficient pavement crack detection system integrating Swin-Transformer and yolov8 according to claim 1 is characterized in that: In step S03, a DBB re-parameterization module is added to the Head structure to prevent the problem of false detection of crack types. The DBB re-parameterization module adopts a separate branch structure, which includes a single branch structure and a multi-branch structure. The multi-branch structure is used in the ST-yolov8 model training stage to obtain better weight parameters, and the multi-branch structure is converted into a single branch structure in the inference stage. The specific operation process is as follows: Among them, k and b represent the weight and bias after the crack feature fusion, k1 and b1 represent the weight and bias of 1×1 convolution, k k and b k Represents the weights and biases of the k×k convolution.
9. The working method of the efficient pavement crack detection system integrating Swin-Transformer and yolov8 according to claim 1 is characterized by: In step S04, the iterative training process of the ST-yolov8 model includes setting the initial learning rate of the ST-yolov8 model to 0.01, using the learning rate adjustment strategy to automatically adjust the learning rate during training, setting the training rounds to 300, using stochastic gradient descent SGD as the optimizer, setting the momentum to 0.937, and the weight decay to 0.0005, and continuously updating the ST-yolov8 model weights and biases.
Citation Information
Patent Citations
Road crack detection method based on improved YOLOv8
CN118762298A
Road crack detection method, medium and product
CN118941526A
Container weak and small serial number target detection and identification method based on deep learning
CN117253154A
Road surface defect detection method and system based on deep learning
CN117274688A
Novel pavement disease detection method
CN117726957A
Cited By
High-altitude parabolic object detection method based on Swinin-Transform and YOLOv8 fusion detection algorithm
CN120726453A