Concrete defect detection method combining CNN (Convolutional Neural Network) and Transformer

By combining the improved YOLOv8 model of CNN and Transformer, the problems of global information capture and long-distance dependency in concrete defect detection are solved, and defect detection with higher accuracy and robustness are achieved.

CN120472206APending Publication Date: 2025-08-12SOUTHWEAT UNIV OF SCI & TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510492326.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Traditional convolution operations are difficult to capture global information and establish long-distance dependencies between features in concrete defect detection, and the network does not fully consider the interactive information between different dimensions when processing concrete images, resulting in information loss and decreased detection accuracy.

Method used

Combining the concrete defect detection method of CNN and Transformer, by introducing Bottleneck Transformer to replace the Bottleneck module, introducing the triple attention mechanism CTAM, and using Focaler-CIoU loss function, the YOLOv8 model is improved and the ability to global feature extraction and interactive information capture is enhanced.

Benefits of technology

It improves the network's ability to extract defect global information, improves detection accuracy and anti-interference ability, and can more accurately capture the long-distance dependence between feature points, enhancing the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472206A_ABST
    Figure CN120472206A_ABST
Patent Text Reader

Abstract

The invention discloses a concrete defect detection method combining CNN and Transformer, and belongs to the technical field of concrete defect detection, and the method comprises the following steps: obtaining a concrete defect image; preprocessing the concrete defect image; constructing a concrete defect data set; a concrete defect detection model is obtained in combination with CNN and Transform; training the concrete defect detection model to obtain an optimal concrete defect detection model; and inputting the collected concrete image into the concrete defect detection optimal model, and outputting a concrete defect detection result. According to the method, the local detail information of the concrete defect can be extracted, the global feature information of the defect can also be extracted, the method is more targeted, the long-distance dependency relationship between the feature points can be effectively captured, the global feature extraction of the network is improved, and the defect detection result is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of concrete defect detection, and specifically relates to a concrete defect detection method combining CNN and Transformer. Background Art

[0002] Concrete is one of the most commonly used building materials in modern infrastructure and plays a vital role in its construction. However, due to uneven concrete settlement, coupled with fluctuations in temperature and humidity, and the influence of various loads and pressures, various types of defects can form in concrete structures, including cracks, rust, spalling, and air holes. These defects directly impact the safety and durability of concrete structures such as bridges, dams, and tunnels. Therefore, regular, systematic, and scientific inspection and maintenance of concrete defects are crucial to ensuring the long-term stable operation of these infrastructures.

[0003] Traditional image processing techniques have achieved initial success in concrete defect detection, significantly reducing the burden of manual inspections. However, as these technologies are increasingly applied in real-world concrete infrastructure scenarios, their limitations are becoming increasingly apparent. For example, existing technologies are easily affected by external factors such as lighting variations and occlusion caused by stains. Furthermore, traditional machine learning-based methods often involve tedious manual feature extraction and subjective threshold setting. Among single-stage object detection algorithms, the YOLO family of techniques is one of the most popular and widely used. However, when directly applied to defect detection in concrete scenarios, it still has certain limitations. First, the characteristic structure of concrete defects can interfere with feature extraction. This is especially true for elongated defects such as cracks, where traditional convolution operations struggle to capture global information and establish long-range dependencies between features. Second, when processing concrete images, networks often fail to fully consider the interaction between different dimensions during feature enhancement, potentially leading to information loss in some dimensions. Finally, the network can significantly lose defect information during feature extraction, and the distribution of samples with varying difficulty levels can affect bounding box regression. These issues all have a certain impact on the network's recognition accuracy. Summary of the Invention

[0004] The purpose of the present invention is to address the above-mentioned deficiencies in the prior art and provide a concrete defect detection method combining CNN and Transformer to solve the problem that traditional convolution operations are difficult to capture global information and establish long-distance dependencies between features in concrete defect identification.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is:

[0006] A concrete defect detection method combining CNN and Transformer, comprising the following steps:

[0007] S1. Obtain the original image of concrete, classify and filter the original image using transfer learning training, and obtain the concrete defect image;

[0008] S2, preprocessing the concrete defect image;

[0009] S3. Perform target detection and annotation on the preprocessed concrete defect image to obtain a concrete defect dataset;

[0010] S4. Improve the YOLOv8 model by combining CNN and Transformer to obtain a concrete defect detection model;

[0011] S5. Using the concrete defect dataset to train the concrete defect detection model to obtain the optimal concrete defect detection model;

[0012] S6. Input the collected concrete image into the optimal model for concrete defect detection, and output the concrete defect detection result.

[0013] Furthermore, the S3 includes: using target detection and annotation software LabelImg to perform target detection and annotation on the preprocessed concrete defect image to obtain a concrete defect dataset, and dividing the concrete defect dataset into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0014] Furthermore, in S4, the YOLOv8 model is improved by combining CNN and Transformer, including:

[0015] The Bottleneck Transformer is used to replace the original Bottleneck to obtain the BoT module, and the BoT module is embedded into the end of the backbone network of the YOLOv8 model;

[0016] Introducing the triple attention mechanism CTAM in the neck network of the YOLOv8 model;

[0017] In the bounding box regression loss of the YOLOv8 model, the Focaler-CIoU loss is used to replace the original CIoU loss to obtain the final bounding box loss function.

[0018] Furthermore, the MHSA multi-head self-attention module is used to replace the convolution module in the Bottleneck Transformer;

[0019] The output of the MHSA multi-head self-attention module is:

[0020] Z=Softmax(qk T +qr T )×v

[0021] Among them, Z is the output of the MHSA multi-head self-attention module; Softmax is the activation function; qr T is the attention part obtained by matrix multiplication of the position encoding and the query matrix, qk T is the other part of the attention obtained by matrix multiplication of the query matrix and the key matrix; v is the value vector.

[0022] Furthermore, the triple attention mechanism (CTAM) is embedded between the C2f module and the output detection head at different levels in the neck network.

[0023] In the first attention mechanism CTAM branch, the input feature tensor χ with a shape of C×H×W is rotated 90° counterclockwise along the H axis to obtain a feature tensor χ1 with a shape of W×H×C. The feature tensor χ1 is simplified to a feature tensor χ1 of 2×H×C through Z-pool transfer. * , the feature tensor χ1 * Through convolution and normalization operations, a 1×H×C feature tensor is obtained. The 1×H×C feature tensor is passed through the sigmoid function to obtain the attention weight ω1. After multiplying the attention weight ω1 with the feature tensor χ1, it is rotated 90° clockwise along the H axis to keep the shape unchanged.

[0024] In the second attention mechanism CTAM branch, the input feature tensor χ with a shape of C×H×W is rotated 90° counterclockwise along the W axis to obtain a feature tensor χ2 with a shape of W×H×C. The feature tensor χ2 is simplified to a feature tensor χ2 of 2×H×C through Z-pool transfer. * , the feature tensor χ2 * Through convolution and normalization operations, a 1×H×C feature tensor is obtained. The 1×H×C feature tensor is passed through the sigmoid function to obtain the attention weight ω2. After multiplying the attention weight ω2 by the feature tensor χ2, it is rotated 90° clockwise along the W axis to keep the shape unchanged.

[0025] In the third attention mechanism CTAM branch, the input feature tensor χ of shape C×H×W is simplified to a feature tensor χ of shape 2×H×C through Z-pool transfer * , the characteristic tensor χ * Through convolution and normalization operations, a 1×H×C feature tensor is obtained. The 1×H×C feature tensor is passed through the sigmoid function to obtain the attention weight ω3, and the attention weight ω3 is multiplied by the feature tensor χ;

[0026] The features output by the first attention mechanism CTAM branch, the second attention mechanism CTAM branch, and the third attention mechanism CTAM branch are averaged and aggregated to obtain the output of the triple attention mechanism CTAM:

[0027]

[0028] Among them, y is the triple attention feature.

[0029] Furthermore, the Z-pool reduces the zeroth dimension of the input feature tensor to 2 through the maximum pooling layer Maxpool and the average pooling layer Avgpool:

[0030] Z-pool(χ)=[MaxPool 0d (χ),AvgPool 0d (χ)]

[0031] Among them, Z-pool(χ) is the dimension pooling operation; MaxPool is the maximum pooling operation; AvgPool is the average pooling operation; χ is the feature tensor; 0d is the 0th dimension through the maximum pooling operation Maxpool and the average pooling operation Avgpool.

[0032] Furthermore, the final bounding box loss function is:

[0033] L Focaler-CIoU =L CIoU +IoU-IoU focaler

[0034] Among them, L Focaler-CIoU is the final bounding box loss function; L CIoU is the CIoU loss function; IoU is the intersection-over-union ratio of the real box and the predicted box; IoU focaler is the IoU under linear interval mapping.

[0035] Furthermore, the CIoU loss function is:

[0036] L CIoU =1-CIoU

[0037]

[0038] IoU focaler for:

[0039]

[0040] Among them, B is the actual prediction box, B GT is the real box; c is the diagonal length of the minimum circumscribed rectangle of the actual prediction box and the real box; d is the center point distance between the actual prediction box and the real box; wgt and h gt Represent the width and height of the actual real box, w and h represent the width and height of the predicted box respectively; α is the weight parameter; ρ is the parameter for measuring the consistency of the aspect ratio; u is the boundary value.

[0041] The concrete defect detection method provided by the present invention, which combines CNN and Transformer, has the following beneficial effects:

[0042] 1. The present invention effectively combines CNN and Transformer to improve the network's ability to extract global defect information. Compared with existing machine vision algorithms, the method of the present invention can extract not only local detail information of defects, but also global feature information of defects. It is more targeted and can effectively capture the long-distance dependency between feature points, thereby improving the network's extraction of global features and providing more accurate defect detection results.

[0043] 2. The present invention uses concrete infrastructure as the scene basis, collects field data through a variety of intelligent equipment, and augments the images through a variety of image preprocessing algorithms to obtain a professional multi-category concrete defect dataset for multiple concrete scenes including dams, bridges, corridors and tunnels.

[0044] 3. Compared with traditional image processing methods, the present invention can greatly improve the anti-interference and generalization capabilities of defect detection through iterative training of convolutional neural networks, and effectively improve the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is the structural diagram of the MHSA multi-head self-attention module of the present invention.

[0046] Figure 2 This is the Bottleneck Transformer structure diagram of the present invention.

[0047] Figure 3 This is a structural diagram of the BoT module of the present invention.

[0048] Figure 4 This is the overall structure diagram of the attention mechanism CTAM of the present invention.

[0049] Figure 5 This is a flow chart of the concrete defect detection method combining CNN and Transformer in the present invention.

[0050] Figure 6 This is the network structure diagram of the concrete defect detection model YOLOv8-CDD of the present invention.

[0051] Figure 7 This is the concrete defect detection result diagram of the present invention, where: Figure 7 (a) in the figure is the crack detection result; Figure 7 (b) is the rust detection result; Figure 7 (c) in the figure is the peeling test result; Figure 7 (d) in the figure is the stoma detection result. DETAILED DESCRIPTION

[0052] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0053] Example 1

[0054] This embodiment is a concrete defect detection method combining CNN and Transformer, refer to Figure 5 and Figure 6 This embodiment can extract defect feature information in a more detailed and comprehensive manner in the concrete scene, which specifically includes the following contents:

[0055] Step S1: obtaining an original image of concrete, and using transfer learning training to classify and screen the original image to obtain a concrete defect image;

[0056] This embodiment uses a mature pre-training model to perform transfer learning training on the original image of concrete. The pre-training model and transfer learning training in this embodiment are both existing technologies in machine learning, so the detailed process is not repeated here.

[0057] Step S2, preprocessing the concrete defect image;

[0058] This embodiment adopts a data preprocessing solution using technologies such as image enhancement, image augmentation, and image registration to ensure the image quality and sample diversity of concrete defect data. The purpose is to improve image quality and information richness, reduce the influence of external interference factors, and enhance the ability of the later model to capture key features.

[0059] Step S3: performing target detection and annotation on the preprocessed concrete defect image to obtain a concrete defect dataset;

[0060] Specifically, this embodiment uses the target detection and annotation software LabelImg to perform target detection and annotation on the preprocessed concrete defect images to construct a high-quality, multi-category, and diverse concrete defect dataset. The concrete defect dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1 for the subsequent training of the concrete defect detection model.

[0061] Step S4: Improve the YOLOv8 model by combining CNN and Transformer to obtain a concrete defect detection model. The concrete defect detection model YOLOv8-CDD (Concrete Defect Detection) of this embodiment adopts a classic one-stage end-to-end structure. The network is iteratively trained using training set and validation set data to obtain a concrete defect detection model. The concrete defect detection model can quickly and efficiently detect and identify concrete defects. The specific structural improvements are as follows:

[0062] Improvement 1:

[0063] The characteristic structure of concrete defects can interfere with feature extraction, especially for slender defects such as cracks, where the length is much greater than the width. Traditional convolution operations have difficulty perceiving the global information of the defects and establishing long-range dependencies between features.

[0064] To address this issue, the present invention introduces the Bottleneck-Transformer (BoT) into the backbone network, effectively combining CNN and Transformer. This approach not only leverages the global perception capabilities of the self-attention mechanism, but also captures local features through convolution, enabling the network to effectively focus on both global and local features of the defect.

[0065] Specifically, this embodiment combines the idea of BoTNet and applies it to the C2f architecture, using Bottleneck Transformer to replace the original Bottleneck to obtain a new BoT module. The BoT module is embedded into the end of the backbone network, that is, the output end of the SPPF module, which enhances the network's overall extraction of global feature information of defects, thereby effectively improving the detection accuracy of the model; the specific structure of Bottleneck Transformer is as follows: Figure 2 As shown; the specific structure of the BoT module is as follows Figure 3 shown.

[0066] This embodiment replaces the original convolution with the MHSA multi-head self-attention module in the Transformer, and combines it with the local feature extraction capability of CNN, which can achieve better feature capture effect than using CNN or Transformer alone.

[0067] refer to Figure 1 , MHSA multi-head self-attention module is the core structure of Transformer, where q, k, v and r represent query, key, value and position encoding respectively. X as input size is H×W×d, where H, W and d represent the height and width of the input feature and the dimension of a single token respectively.

[0068] The output of the MHSA multi-head self-attention module is:

[0069] Z=Softmax(qk T +qr T )×v

[0070] Among them, Z is the output of the MHSA multi-head self-attention module; Softmax is the activation function; qr T is the attention part obtained by matrix multiplication of the position encoding and the query matrix, qk T is the other part of the attention obtained by matrix multiplication of the query matrix and the key matrix; v is the value vector.

[0071] Improvement 2:

[0072] When processing concrete images, the network usually does not fully consider the interaction information between different dimensions in the feature enhancement part, which may lead to information loss in some dimensions.

[0073] To solve this problem, refer to Figure 4 This embodiment introduces a triple attention mechanism (Convolutional Triplet Attention Module, CTAM) into the neck network of the YOLOv8 model. By adopting a three-branch structure, the triple attention mechanism CTAM can effectively capture the interactive relationship between various dimensions, significantly improving the model's ability to characterize key defect features with almost no increase in computational burden, thereby effectively improving the model's recognition accuracy. The triple attention mechanism CTAM consists of three parallel branches, of which the first two branches are responsible for calculating the interactive information between the channel dimension C and the spatial dimension H or W. The last branch is used solely to construct spatial attention. Finally, the outputs of the three branches are aggregated using a simple average to obtain the output feature.

[0074] In this embodiment, an attention mechanism (CTAM) is added between the C2f and the output detection head at different levels in the neck network, which can effectively enhance the regional information at different scales, realize feature enhancement at multiple resolution scales, and improve detection accuracy.

[0075] The specific process is as follows: given an input feature tensor

[0076] In the first attention mechanism CTAM branch, the input feature tensor χ with a shape of C×H×W is rotated 90° counterclockwise along the H axis to obtain a feature tensor χ1 with a shape of W×H×C. The feature tensor χ1 is simplified to a feature tensor χ1 of 2×H×C through Z-pool transfer. * , the feature tensor χ1 * Through convolution and normalization operations, a 1×H×C feature tensor is obtained. The 1×H×C feature tensor is passed through the sigmoid function to obtain the attention weight ω1. After multiplying the attention weight ω1 with the feature tensor χ1, it is rotated 90° clockwise along the H axis to keep the shape unchanged.

[0077] In the second attention mechanism CTAM branch, the input feature tensor χ with a shape of C×H×W is rotated 90° counterclockwise along the W axis to obtain a feature tensor χ2 with a shape of W×H×C. The feature tensor χ2 is simplified to a feature tensor χ2 of 2×H×C through Z-pool transfer. * , the feature tensor χ2 * Through convolution and normalization operations, a 1×H×C feature tensor is obtained. The 1×H×C feature tensor is passed through the sigmoid function to obtain the attention weight ω2. After multiplying the attention weight ω2 by the feature tensor χ2, it is rotated 90° clockwise along the W axis to keep the shape unchanged.

[0078] In the third attention mechanism CTAM branch, the input feature tensor χ of shape C×H×W is simplified to a feature tensor χ of shape 2×H×C through Z-pool transfer * , the characteristic tensor χ * Through convolution and normalization operations, a 1×H×C feature tensor is obtained. The 1×H×C feature tensor is passed through the sigmoid function to obtain the attention weight ω3, and the attention weight ω3 is multiplied by the feature tensor χ;

[0079] The features output by the first attention mechanism CTAM branch, the second attention mechanism CTAM branch, and the third attention mechanism CTAM branch are averaged and aggregated to obtain the output of the triple attention mechanism CTAM:

[0080]

[0081] Among them, y is the triple attention feature.

[0082] In Z-pool, the zeroth dimension of the input feature tensor is reduced to 2 through the maximum pooling layer Maxpool and the average pooling layer Avgpool. This not only retains the rich information of the features, but also further reduces the computational complexity. The specific calculation formula is as follows:

[0083] Z-pool(χ)=[MaxPool 0d(χ),AvgPool 0d (χ)]

[0084] Among them, Z-pool(χ) is the dimension pooling operation; MaxPool is the maximum pooling operation; AvgPool is the average pooling operation; χ is the feature tensor; 0d is the 0th dimension through the maximum pooling operation Maxpool and the average pooling operation Avgpool.

[0085] Improvement three:

[0086] The YOLOv8 model network suffers from significant loss of defect information during feature extraction, and the distribution of samples of varying difficulty levels influences bounding box regression, impacting network recognition accuracy. While CIoU can effectively improve model convergence speed and detection accuracy to a certain extent, it ignores the impact of the distribution of samples of varying difficulty levels on bounding box regression.

[0087] Based on this, this embodiment uses Focaler-CIoU loss to replace the original CIoU loss in the bounding box regression loss of the YOLOv8 model, which can focus on regression samples of different difficulty levels and allow the loss function to be sensitive to IoU within a certain range, which helps the model better learn samples of different difficulty levels, optimizes the training process, and can further improve the detection accuracy of the model as a whole. The Focaler-CIoU loss is specifically:

[0088]

[0089]

[0090] L CIoU =1-CIoU

[0091]

[0092] Among them, B is the actual prediction box, B GT is the real box; c is the diagonal length of the minimum circumscribed rectangle of the actual prediction box and the real box; d is the center point distance between the actual prediction box and the real box; w gt and h gt Represent the width and height of the actual real box, w and h represent the width and height of the predicted box respectively; α is the weight parameter; ρ is the parameter for measuring the consistency of aspect ratio; u is the boundary value; [d,u]∈[0,1], by adjusting the values of d and u, IoU can be focaler Pay attention to different samples. Applying this idea to CIoU, the final bounding box loss function is:

[0093] L Focaler-CIoU =L CIoU +IoU-IoUfocaler

[0094] Among them, L Focaler-CIoU is the final bounding box loss function; L CIoU is the CIoU loss function; IoU is the intersection-over-union ratio of the real box and the predicted box; IoU focaler is the IoU under linear interval mapping.

[0095] Step S5: Iteratively train the concrete defect detection model using the training set and validation set in the concrete defect dataset to obtain the optimal model for concrete defect detection, and test the detection effect on the test set.

[0096] Step S6, reference Figure 7 , the collected concrete images are input into the optimal model of concrete defect detection, and the concrete defect detection results are output; the detection results are processed and analyzed, and combined with the camera imaging principle and the actual camera parameters, the pixel-level area quantification results of the defects are converted into actual physical quantification results; according to the actual physical quantification results of the defect detection, the concrete structure defect loss is graded and evaluated.

[0097] Although the specific embodiments of the invention are described in detail in conjunction with the accompanying drawings, this should not be construed as limiting the scope of protection of this patent. Within the scope described by the claims, various modifications and variations that can be made by those skilled in the art without creative work still fall within the scope of protection of this patent.

Claims

1. A concrete defect detection method combining CNN and Transformer, characterized in that: The following steps are involved: S1. Obtain the original image of concrete, classify and filter the original image using transfer learning training, and obtain the concrete defect image; S2, preprocessing the concrete defect image; S3. Perform target detection and annotation on the preprocessed concrete defect image to obtain a concrete defect dataset; S4. Improve the YOLOv8 model by combining CNN and Transformer to obtain a concrete defect detection model; S5. Using the concrete defect dataset to train the concrete defect detection model to obtain the optimal concrete defect detection model; S6. Input the collected concrete image into the optimal model for concrete defect detection, and output the concrete defect detection result.

2. The concrete defect detection method combining CNN and Transformer according to claim 1 is characterized in that: The S3 includes: using the target detection and annotation software LabelImg to perform target detection and annotation on the preprocessed concrete defect image to obtain a concrete defect dataset, and dividing the concrete defect dataset into a training set, a validation set, and a test set in a ratio of 8:1:

1.

3. The concrete defect detection method combining CNN and Transformer according to claim 1 is characterized in that: In S4, the YOLOv8 model is improved by combining CNN and Transformer, including: The Bottleneck Transformer is used to replace the original Bottleneck to obtain the BoT module, and the BoT module is embedded into the end of the backbone network of the YOLOv8 model; Introducing the triple attention mechanism CTAM in the neck network of the YOLOv8 model; In the bounding box regression loss of the YOLOv8 model, the Focaler-CIoU loss is used to replace the original CIoU loss to obtain the final bounding box loss function.

4. The concrete defect detection method combining CNN and Transformer according to claim 3 is characterized in that: The convolution module in Bottleneck Transformer is replaced by the MHSA multi-head self-attention module; The output of the MHSA multi-head self-attention module is: Z=Softmax(qk T +qr T )×v Among them, Z is the output of the MHSA multi-head self-attention module; Softmax is the activation function; qr T is the attention part obtained by matrix multiplication of the position encoding and the query matrix, qk T is the other part of the attention obtained by matrix multiplication of the query matrix and the key matrix; v is the value vector.

5. The concrete defect detection method combining CNN and Transformer according to claim 3 is characterized in that: The triple attention mechanism CTAM is embedded between the C2f module and the output detection head at different levels in the neck network; In the first attention mechanism CTAM branch, the input feature tensor χ with a shape of C×H×W is rotated 90° counterclockwise along the H axis to obtain a feature tensor χ1 with a shape of W×H×C. The feature tensor χ1 is simplified to a feature tensor χ1 of 2×H×C through Z-pool transfer. * , the feature tensor χ1 * Through convolution and normalization operations, a 1×H×C feature tensor is obtained. The 1×H×C feature tensor is passed through the sigmoid function to obtain the attention weight ω1. After multiplying the attention weight ω1 with the feature tensor χ1, it is rotated 90° clockwise along the H axis to keep the shape unchanged. In the second attention mechanism CTAM branch, the input feature tensor χ with a shape of C×H×W is rotated 90° counterclockwise along the W axis to obtain a feature tensor χ2 with a shape of W×H×C. The feature tensor χ2 is simplified to a feature tensor χ2 of 2×H×C through Z-pool transfer. * , the characteristic tensor χ2 * Through convolution and normalization operations, a 1×H×C feature tensor is obtained. The 1×H×C feature tensor is passed through the sigmoid function to obtain the attention weight ω2. After multiplying the attention weight ω2 by the feature tensor χ2, it is rotated 90° clockwise along the W axis to keep the shape unchanged. In the third attention mechanism CTAM branch, the input feature tensor χ of shape C×H×W is simplified to a feature tensor χ of shape 2×H×C through Z-pool transfer * , the characteristic tensor χ * Through convolution and normalization operations, a 1×H×C feature tensor is obtained. The 1×H×C feature tensor is passed through the sigmoid function to obtain the attention weight ω3, and the attention weight ω3 is multiplied by the feature tensor χ; The features output by the first attention mechanism CTAM branch, the second attention mechanism CTAM branch, and the third attention mechanism CTAM branch are averaged and aggregated to obtain the output of the triple attention mechanism CTAM: Among them, y is the triple attention feature.

6. The concrete defect detection method combining CNN and Transformer according to claim 5 is characterized in that: In the Z-pool, the zeroth dimension of the input feature tensor is reduced to 2 through the maximum pooling layer Maxpool and the average pooling layer Avgpool: Z-pool(x)=[MaxPool 0d (x),AvgPool 0d (x)] Among them, Z-pool(χ) is the dimension pooling operation; MaxPool is the maximum pooling operation; AvgPool is the average pooling operation; χ is the feature tensor; 0d is the 0th dimension through the maximum pooling operation Maxpool and the average pooling operation Avgpool.

7. The concrete defect detection method combining CNN and Transformer according to claim 3 is characterized in that: The final bounding box loss function is: L Focaler-CIoU =L CIoU +IoU-IoU focaler Among them, L Focaler-CIoU is the final bounding box loss function; L CIoU is the CIoU loss function; IoU is the intersection-over-union ratio of the real box and the predicted box; IoU focaler is the IoU under linear interval mapping.

8. The concrete defect detection method combining CNN and Transformer according to claim 7 is characterized in that: The CIoU loss function is: L CIoU =1-CIoU IoU focaler for: Among them, B is the actual prediction box, B GT is the real box; c is the diagonal length of the minimum circumscribed rectangle of the actual prediction box and the real box; d is the center point distance between the actual prediction box and the real box; w gt and h gt They represent the width and height of the actual real box, w and h represent the width and height of the predicted box respectively; α is the weight parameter; ρ is the parameter for measuring the consistency of the aspect ratio; u is the boundary value.