Bridge structure concrete crack semantic segmentation method based on Swin Transform and CNN

By combining Swing Transformer and CNN to develop a semantic segmentation method for concrete cracks in bridge structures, the problems of insufficient local feature dependence and low computational efficiency in concrete crack detection in bridge structures are solved, achieving high-precision and low-parameter automatic detection.

CN120953599APending Publication Date: 2025-11-14YUYU +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510668541.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing methods for detecting concrete cracks in bridge structures suffer from insufficient dependence on local features, poor stability under complex backgrounds, large number of model parameters, and low computational efficiency, making it difficult to meet the requirements for lightweight design.

Method used

A pixel-level semantic segmentation method for concrete cracks in bridge structures based on Swin Transformer and CNN is adopted. It combines the global attention calculation module of Swin Transformer and the local feature calculation module of CNN, and performs feature fusion through a multi-scale feature pyramid decoder to achieve automatic semantic segmentation.

Benefits of technology

It improves the accuracy and stability of crack detection, reduces the number of model parameters and computational complexity, and is suitable for deployment on embedded devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953599A_ABST
    Figure CN120953599A_ABST
Patent Text Reader

Abstract

The invention discloses a bridge structure concrete crack semantic segmentation method based on Swin Transformer and CNN, and the method comprises the steps: collecting image data used for representing a bridge structure concrete crack, and determining a pairing data set; establishing a pixel-level bridge structure concrete crack semantic segmentation model PSC based on Swin Transform and a CNN (Convolutional Neural Network); inputting the paired data set into a pixel-level bridge structure concrete crack semantic segmentation model PSC based on Swin Transform and CNN so as to perform automatic semantic segmentation on the bridge structure concrete crack, and determining a detection result; and based on the detection result, generating a concrete crack detection network prediction result comparison graph and a concrete crack detection network parameter quantity and calculation efficiency comparison graph. According to the invention, the automatic detection of the concrete crack is realized, and the labor cost is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more specifically, to a semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN. Background Technology

[0002] As a common type of engineering structure, bridge concrete structures are constantly exposed to environmental influences such as natural disasters, traffic loads, and human-caused damage, frequently resulting in various surface defects such as cracks and spalling. Therefore, regular inspection and damage identification of bridge concrete structures are crucial. In the past, manual visual inspection was commonly used, but this method not only requires significant human and material resources but also carries the risk of errors and safety hazards due to human factors.

[0003] In recent years, image segmentation algorithms based on convolutional neural networks (CNNs) have been widely used in the detection of cracks in bridge concrete. However, the over-reliance of traditional CNN models on local features leads to their inadequacy in understanding the global structure and poor stability under complex background interference.

[0004] Recently, Transformer-based computer vision models have shown significant advantages in processing global information, effectively overcoming the local feature dependency problem of traditional CNN models. However, directly applying Transformer for image recognition often results in enormous computational demands and high resource consumption.

[0005] Today, deep learning and computer vision technologies have become mainstream methods in the field of bridge structure inspection and have made significant progress. However, existing methods still face two main limitations, namely the following problems: (1) In noisy environments, especially for segmenting details and branches of cracks, the performance is still insufficient. Although Transformer and CNN-based models have achieved certain results, existing models have not fully integrated the advantages of both, resulting in certain limitations in crack segmentation tasks; (2) The model has a large number of parameters, which restricts computational efficiency and makes it difficult to meet the requirements of lightweight design. Excessive model parameters not only reduce computational speed but also easily lead to overfitting problems, especially posing significant difficulties in deployment on embedded devices.

[0006] Application content

[0007] This application aims to at least partially address one of the aforementioned technical problems in the prior art.

[0008] An exemplary embodiment of this application provides a semantic segmentation method for concrete cracks in bridge structures based on Swin Transformer and CNN, the semantic segmentation method comprising:

[0009] Image data was acquired to characterize concrete cracks in bridge structures, and paired datasets were determined.

[0010] A pixel-level semantic segmentation model for concrete cracks in bridge structures, PSC, based on Swing Transformer and CNN, was established.

[0011] The paired dataset is input into the pixel-level semantic segmentation model PSC for concrete cracks in bridge structures based on Swing Transformer and CNN to automatically perform semantic segmentation of concrete cracks in bridge structures and determine the detection results.

[0012] Based on the detection results, a comparison chart of the prediction results of the concrete crack detection network and a comparison chart of the number of parameters and computational efficiency of the concrete crack detection network are generated.

[0013] According to one embodiment of this application, the acquisition of image data for characterizing concrete cracks in a bridge structure and the determination of paired datasets include:

[0014] The image data is preprocessed to determine the paired dataset;

[0015] The preprocessing includes at least image cropping, data augmentation, label creation, and creation of classification labels for the image data.

[0016] According to one embodiment of this application, the image cropping is performed manually, and the image cropping specification is set to 224x224.

[0017] And / or,

[0018] The labels are created manually on the cracks.

[0019] According to one embodiment of this application, the pixel-level semantic segmentation model PSC for concrete cracks in bridge structures based on Swing Transformer and CNN includes:

[0020] A global attention calculation module based on the Swing Transformer sliding window is used to extract global contextual features to eliminate noise interference;

[0021] The local feature computation module of VanillaNet, which is based on CNN, is used to extract local detailed features at different scales through multi-layer convolution processing.

[0022] The multi-scale feature pyramid decoder module is used to fuse multi-scale features through the feature pyramid decoder and progressively upsample to the segmentation result of the original resolution image.

[0023] According to one embodiment of this application, the pixel-level bridge structure concrete crack semantic segmentation model PSC based on Swin Transformer and CNN processes the original image independently and directly through parallel downsampling in a parallel sequence, so as to extract features of different styles respectively.

[0024] According to one embodiment of this application, the implementation process of the global attention calculation module based on the Swing Transformer sliding window includes:

[0025] The crack image is divided into multiple blocks using the Patch Partition layer, and then dimensional transformation is performed using the Linear Embedding layer.

[0026] Pixel recombination and dimension transformation are performed using the Patch Merging layer;

[0027] Input four consecutive Stages to process feature maps of different sizes, and obtain four feature maps of different dimensions, Stage1 to Stage4 respectively.

[0028] According to one embodiment of this application, Stage 1 passes through a linear layer first, while Stages 2, 3, and 4 pass through a Patch Merging layer.

[0029] According to one embodiment of this application, the implementation process of the local feature calculation module of VanillaNet developed based on CNN includes:

[0030] In the Stem stage, 4×4 convolutions are used to modify the dimensions of each layer of the original network from 512, 1024, 2048, and 4096 to 96, 192, 384, and 768, and the last two non-linear layers used for classification are removed.

[0031] The Stage part uses MaxPool for downsampling, and then uses 1×1 convolution for processing;

[0032] The head part uses two non-linear layers for classification.

[0033] According to one embodiment of this application, the implementation process of the multi-scale feature pyramid decoder module includes:

[0034] The feature maps obtained from the two downsampling operations are stitched together along the channel direction.

[0035] Use transposed convolution for dimensionality transformation and image filling;

[0036] The ReLU activation function is applied, and the sampled twice to a size of (1, 224, 224) is performed using a 1×1 transposed convolutional integral.

[0037] According to one embodiment of this application, generating a comparison chart of prediction results from the concrete crack detection network and a comparison chart of the number of parameters and computational efficiency of the concrete crack detection network based on the detection results includes:

[0038] The detection results are compared with those of U-Net, DeepLabV3, PSPNet, and FPN models to generate comparison charts of concrete crack detection network prediction results and comparison charts of concrete crack detection network parameter quantity and computational efficiency.

[0039] The technical advantages of this application are as follows: a semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN. This semantic segmentation method includes: acquiring image data to characterize concrete cracks in bridge structures and determining paired datasets; establishing a pixel-level semantic segmentation model (PSC) for concrete cracks in bridge structures based on Swing Transformer and CNN; inputting the paired datasets into the pixel-level semantic segmentation model (PSC) for concrete cracks in bridge structures to automatically perform semantic segmentation of the concrete cracks and determine the detection results; and generating a comparison chart of prediction results and parameter quantity and computational efficiency of the concrete crack detection network based on the detection results. This application achieves automatic detection of concrete cracks, greatly reducing labor costs. Attached Figure Description

[0040] Figure 1 The flowchart shows the steps of a semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN.

[0041] Figure 2 This is a diagram of the PSC network structure of the present invention;

[0042] Figure 3 This is a schematic diagram of SW-SAM operation;

[0043] Figure 4 The VanillaNet network structure is shown in the diagram as a "cylindrical" network.

[0044] Figure 5 Remove the two non-linear structure graphs used for classification from the VanillaNet network structure graph during the Stem stage;

[0045] Figure 6 This is a comparison chart of the prediction results of the concrete crack detection network of the present invention;

[0046] Figure 7 This is a comparison chart of the number of parameters and computational efficiency of the concrete crack detection network of this invention. Detailed Implementation

[0047] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0048] like Figures 1 to 7 As shown, an exemplary embodiment of this application provides a semantic segmentation method for concrete cracks in bridge structures based on Swin Transformer and CNN, characterized in that the semantic segmentation method includes:

[0049] Step S100: Acquire image data to characterize concrete cracks in the bridge structure and determine the paired dataset.

[0050] Step S200: Establish a pixel-level semantic segmentation model (PSC) for concrete cracks in bridge structures based on Swin Transformer and CNN.

[0051] Step S300: Input the paired dataset into the pixel-level semantic segmentation model PSC for bridge structure concrete cracks based on Swin Transformer and CNN to automatically perform semantic segmentation of bridge structure concrete cracks and determine the detection results.

[0052] Step S400: Based on the detection results, generate a comparison chart of the prediction results of the concrete crack detection network and a comparison chart of the number of parameters and computational efficiency of the concrete crack detection network.

[0053] In step S100, image data for characterizing concrete cracks in the bridge structure is acquired, and paired datasets are determined. This includes preprocessing the image data to determine the paired datasets. The preprocessing includes at least image cropping, data augmentation, label creation, and creation of classification labels for the image data.

[0054] The image cropping process was performed manually, with the cropping dimensions set to 224x224. Additionally, the labeling was done manually, creating labels for the cracks.

[0055] Specifically, the following methods can be used in the process of acquiring raw image data and preprocessing the image data:

[0056] Image data acquisition: Images of concrete cracks were captured using a camera, and a large number of public datasets were collected;

[0057] Image cropping: Based on the design characteristics of the model, the original concrete image is cropped into a 224×224 image.

[0058] Label creation: Manually add pixel-level labels to images containing concrete cracks to create a concrete crack sample set.

[0059] Data preprocessing: The dataset was expanded using data augmentation and other methods, and it was divided into a training set of 1744 images, a validation set of 256 images, and a test set of 480 images.

[0060] In step S200, the semantic segmentation structure model based on the Swin Transformer and CNN architecture and constructed using PSC parallel downsampling includes: a global attention calculation module for the Swin Transformer sliding window, a local feature calculation module for VanillaNet developed based on CNN, and a multi-scale feature pyramid decoder module.

[0061] The specific structure is as follows:

[0062] Figure 3 The diagram illustrates the SW-SAM operation. Specifically, assuming the result of W-SAM in layer L (Figure (a) is the result of layer L), SW-SAM is used in layer L+1 (Figure (b)). The diagram clearly shows window displacement, performing self-attention calculations across windows. The SW-SAM window is offset two patches to the lower right corner from the W-MSA layer, forming nine blocks of varying sizes. Cyclic shift is then used to translate and stitch these nine blocks into four blocks of the same size as the W-MSA blocks (Figure (c)). Masked MSA is then used to perform template calculations on these four stitched blocks to extract information. Finally, reverse cyclic shift is used to translate the patches containing the information data back to their original positions.

[0063] Reference Figure 3As shown, the global attention calculation module of the Swin Transformer sliding window works as follows: the image is divided into multiple blocks through the PatchPartition layer, and then dimensional transformation is performed through the Linear Embedding layer; pixel recombination and dimensional transformation are performed through the Patch Merging layer; the four stages contain 2, 2, 6, and 2 SwinTransformer Blocks respectively. Each Block consists of a window multi-head self-attention module (W-MSA) and a sliding window multi-head self-attention module (SW-MSA). W-MSA performs self-attention calculation within a local window, while SW-SAM performs cross-window self-attention calculation by shifting the window.

[0064] Therefore, the Swin Transformer outputs a global feature map of the concrete crack image. The global attention calculation module of the Swin Transformer uses a sliding window approach, which reduces the computational complexity of global self-attention. The self-attention mechanism of a local window can extract global contextual features to remove noise interference.

[0065] Figure 4 and Figure 5 The diagram above shows the VanillaNet network structure. Figure 4 The VanillaNet network structure is as follows: (1) Stem part: 4×4 convolution is used for transformation. (2) Stage part: MaxPool is used for downsampling for the first time, and then 1×1 convolution is used for processing. (3) Head part: Two non-linear layers are used for classification. The figure below shows that 4×4 convolution is still used in the Stem stage, and the dimensions of each layer of the original network (512, 1024, 2048, 4096) are modified to 96, 192, 384, 768, and the two non-linear layers used for classification are removed.

[0066] like Figure 4 and Figure 5 As shown, the implementation process of the VanillaNet module based on CNN is as follows:

[0067] Input a 224×224 image into the upper half of Stem, with dimensions from 3 to 24 and image size from 224 to 112.

[0068] Execute the SIAF activation function:

[0069] A s (x h,w,c )=∑ i,j∈{-n,n} a i,j,c A(x i+h,j+w,c +b c );

[0070] ReLU is chosen as the base activation function. In the above expression, n is the number of activation functions, and a and b are the scale and bias of the activation functions, to avoid simple stacking.

[0071] Input the lower half of Stem, dimension 24→24, image size 112×112→112×112, output feature map (24, 112, 112).

[0072] Stages 1, 2, and 3 perform dimensionality transformation, pooling operations, and execute the SIFA activation function to output feature maps (48, 56, 56), (96, 28, 28), and (192, 14, 14).

[0073] Among them, the local feature calculation module of VanillaNet, which is based on CNN, can extract local detail features at different scales with only 4 layers of convolution, which helps to process the details of cracks.

[0074] The implementation process of the multi-scale feature pyramid decoder module is as follows:

[0075] The feature maps obtained from the two downsampling operations are stitched together along the channel direction.

[0076] Use transposed convolution for dimensionality transformation and image filling.

[0077] Execute the ReLU activation function:

[0078] α i =max(z, 0);

[0079] The image was upsampled twice to a size of (1, 224, 224) using a 1×1 transposed convolutional integral. Multi-scale features were then fused using a feature pyramid decoder, progressively upsampling the image to the original resolution.

[0080] In step S300, the paired dataset is input into the pixel-level bridge structure concrete crack semantic segmentation model PSC based on Swin Transformer and CNN to automatically semantically segment the bridge structure concrete cracks and determine the detection results. This process includes training the concrete crack model.

[0081] First, parameter initialization settings: During training, the image size is maintained at 224×224; the training epoch is set to 200, the initial learning rate is set to 1×10-4, and the batch size is set to 8; the AdamW optimizer is selected; and an adaptive learning rate dynamic adjustment optimizer is used; the encoder is initialized using pre-trained weights obtained from ImageNet.

[0082] Then, training begins by inputting the training set samples into the PSC model for training.

[0083] Then, the parameters are updated, and the loss function used is DiceLoss. The output of each round of calculation is compared with the label image, and backpropagation training is performed to update the weight parameters in PSC, thereby reducing the loss value. The specific formula of the loss function is as follows:

[0084]

[0085] Where y i and Let represent the true label and the predicted result of the i-th pixel, respectively, and N represent the total number of pixels in the sample.

[0086] Finally, terminate the training. This involves repeating the steps of starting training and updating parameters. Training is stopped when the set number of training iterations is reached or the validation set accuracy stops increasing. The test set is then fed into the trained model, and the segmentation results are output to determine the detection outcome.

[0087] In step S400, the test image is input into the trained model, and the segmentation result is checked, such as... Figure 6 As shown, the detection results are compared with those of U-Net, DeepLabV3, PSPNet, and FPN models to generate comparison charts of concrete crack detection network prediction results and comparison charts of concrete crack detection network parameter quantity and computational efficiency.

[0088] This example aims to compare the model with traditional CNN models on four commonly used evaluation metrics in semantic segmentation: Precision, Recall, F1-score, and IoU. Experimental results and comparisons are shown in Table 1.

[0089]

[0090] Table 1

[0091] Analysis of the above experimental results shows that PSC significantly improves the evaluation metrics compared to segmentation methods such as U-Net, U-Net++, DeepLabV3, PSPNet, and FPN, demonstrating the effectiveness of the pixel-level semantic segmentation model for concrete cracks in bridge structures based on Swin Transformer and CNN.

[0092] The pixel-level semantic segmentation model for concrete cracks in bridge structures, PSC (Parallel Swin-CNN), based on Swing Transformer and CNN, was further compared with traditional CNN models in terms of parameter count and computational efficiency. Experimental results are as follows: Figure 7As shown. Analysis Figure 7 It can be seen that, compared with traditional CNN models, PSC has significantly fewer parameters, FPS, and FLOPs except for PSPNet. However, since PSPNet performs poorly in actual segmentation, it proves that the PSC model effectively reduces the number of model parameters and improves computational efficiency.

[0093] This example proposes a method for detecting concrete cracks in bridge structures based on a PSC network, achieving automatic detection of concrete cracks in bridge structures and significantly reducing labor costs. Specifically, a Swing Transformer-based encoder module is designed to improve the model's robustness against interference by extracting global features. A VanillaNet network based on CNN is also designed to extract local detail features, effectively improving the recognition accuracy of crack edges and the overall crack structure. Furthermore, a multi-scale feature fusion pyramid decoder is employed to fuse global and local features at multiple scales, thereby enhancing low-level feature information while preserving high-level features.

[0094] like Figures 1 to 7 As shown, in some embodiments, the pixel-level bridge structure concrete crack semantic segmentation model PSC based on Swin Transformer and CNN processes the original image independently and directly through parallel downsampling in a parallel order, so as to extract features of different styles respectively.

[0095] like Figures 1 to 7 As shown, in some embodiments, the implementation process of the global attention calculation module based on the Swing Transformer sliding window includes:

[0096] The crack image is divided into multiple blocks using the Patch Partition layer, and then dimensional transformation is performed using the Linear Embedding layer.

[0097] Pixel recombination and dimension transformation are performed using the Patch Merging layer;

[0098] Four consecutive stages are input to process feature maps of different sizes, resulting in four feature maps of different dimensions: Stage 1 to Stage 4. Stage 1 first passes through a linear layer, while Stages 2, 3, and 4 pass through a Patch Merging layer.

[0099] like Figures 1 to 7 As shown, in some embodiments, the implementation process of the local feature calculation module of VanillaNet, which is based on CNN, includes:

[0100] In the Stem stage, 4×4 convolutions are used to modify the dimensions of each layer of the original network from 512, 1024, 2048, and 4096 to 96, 192, 384, and 768, and the last two non-linear layers used for classification are removed.

[0101] The Stage part uses MaxPool for downsampling, and then uses 1×1 convolution for processing;

[0102] The head part uses two non-linear layers for classification.

[0103] like Figures 1 to 7 As shown, in some embodiments, the implementation process of the multi-scale feature pyramid decoder module includes:

[0104] The feature maps obtained from the two downsampling operations are stitched together along the channel direction.

[0105] Use transposed convolution for dimensionality transformation and image filling;

[0106] The ReLU activation function is applied, and the sampled twice to a size of (1, 224, 224) is performed using a 1×1 transposed convolutional integral.

[0107] The above example proposes a semantic segmentation method for concrete cracks in bridge structures based on Swin Transformer and CNN, specifically a novel pixel-level crack segmentation model, PSC (Parallel Swin-CNN). PSC retains common upsampling and downsampling image processing methods, using a parallel downsampling sequence to independently process the original image through the Transformer and CNN modules, extracting features of different styles to enhance feature extraction capabilities. The model uses the global attention computation module of the Swin Transformer sliding window as the feature extraction network, then employs the local feature computation module of VanillaNet (based on CNN) to extract local detail features at different scales. Finally, the multi-scale feature pyramid decoder module progressively upsamples to the original resolution image segmentation result.

[0108] Among them, such as Figure 2As shown, the new pixel-level crack segmentation model PSC (ParallelSwin-CNN) can be roughly divided into three parts: the Swing Transformer part, whose sliding window method reduces the computational complexity caused by global self-attention, and can extract global contextual features to remove noise interference through the self-attention mechanism of local windows; the CNN part, the new VanillaNet structure developed by CNN can extract local detail features at different scales with only 4 layers of convolution, which helps to process the details of cracks; and the upsampling part, which fuses multi-scale features through feature pyramid decoder and gradually upsamples to the segmentation result of the original resolution image.

[0109] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN, characterized in that, The semantic segmentation method includes: Image data was acquired to characterize concrete cracks in bridge structures, and paired datasets were determined. A pixel-level semantic segmentation model for concrete cracks in bridge structures, PSC, based on Swing Transformer and CNN, was established. The paired dataset is input into the pixel-level semantic segmentation model PSC for concrete cracks in bridge structures based on Swing Transformer and CNN to automatically perform semantic segmentation of concrete cracks in bridge structures and determine the detection results. Based on the detection results, a comparison chart of the prediction results of the concrete crack detection network and a comparison chart of the number of parameters and computational efficiency of the concrete crack detection network are generated.

2. The semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN as described in claim 1, characterized in that, The acquisition of image data used to characterize concrete cracks in bridge structures, and the determination of paired datasets, includes: The image data is preprocessed to determine the paired dataset; The preprocessing includes at least image cropping, data augmentation, label creation, and creation of classification labels for the image data.

3. The semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN according to claim 2, characterized in that, The image cropping is performed manually, and the cropping size is set to 224x224. And / or, The labels are created manually on the cracks.

4. The semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN according to claim 1, characterized in that, The pixel-level semantic segmentation model for concrete cracks in bridge structures, based on Swing Transformer and CNN, includes: A global attention calculation module based on the Swing Transformer sliding window is used to extract global contextual features to eliminate noise interference; The local feature computation module of VanillaNet, which is based on CNN, is used to extract local detailed features at different scales through multi-layer convolution processing. The multi-scale feature pyramid decoder module is used to fuse multi-scale features through the feature pyramid decoder and progressively upsample to the segmentation result of the original resolution image.

5. The semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN according to claim 4, characterized in that, The pixel-level semantic segmentation model for concrete cracks in bridge structures, PSC, based on Swin Transformer and CNN, processes the original image independently using a parallel downsampling sequence, allowing for the extraction of features of different styles.

6. The semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN according to claim 4, characterized in that, The implementation process of the global attention calculation module based on the Swing Transformer sliding window includes: The crack image is divided into multiple blocks through the Patch Partition layer, and then the dimensionality is transformed through the Linear Embedding layer. Pixel recombination and dimension transformation are performed using the Patch Merging layer; Input four consecutive Stages to process feature maps of different sizes, and obtain four feature maps of different dimensions, Stage1 to Stage4 respectively.

7. The semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN according to claim 6, characterized in that, Stage 1 passes through the linear layer first, while Stages 2, 3, and 4 pass through the PatchMerging layer.

8. The semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN according to claim 4, characterized in that, The implementation process of the local feature calculation module of VanillaNet, which is based on CNN, includes: In the Stem stage, 4×4 convolutions are used to modify the dimensions of each layer of the original network from 512, 1024, 2048, and 4096 to 96, 192, 384, and 768, and the last two non-linear layers used for classification are removed. The Stage part uses MaxPool for downsampling, and then uses 1×1 convolution for processing; The head part uses two non-linear layers for classification.

9. The semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN according to claim 4, characterized in that, The implementation process of the multi-scale feature pyramid decoder module includes: The feature maps obtained from the two downsampling operations are stitched together along the channel direction. Use transposed convolution for dimensionality transformation and image filling; The ReLU activation function is applied, and the sampled twice to a size of (1, 224, 224) is performed using a 1×1 transposed convolutional integral.

10. The semantic segmentation method for concrete cracks in bridge structures based on Swing Transformer and CNN according to any one of claims 1 to 9, characterized in that, Based on the detection results, a comparison chart of the prediction results of the concrete crack detection network and a comparison chart of the number of parameters and computational efficiency of the concrete crack detection network are generated, including: The detection results are compared with those of U-Net, DeepLabV3, PSPNet, and FPN models to generate comparison charts of concrete crack detection network prediction results and comparison charts of concrete crack detection network parameter quantity and computational efficiency.