Bridge crack detection method, equipment, medium and product
Through the improved Swin-TransUNet model, combined with the bidirectional feature pyramid fusion module, frequency domain feature enhancement module and Dice Loss loss function, the problems of low detection efficiency and poor environmental adaptability in bridge crack detection are solved, and high-precision and robust bridge crack detection are achieved.
Patent Information
- Application Number
- CN202510689373.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-29
AI Technical Summary
The existing bridge crack detection technology relies on manual experience, has low detection efficiency, high cost and is susceptible to subjective factors, has poor environmental adaptability, and is difficult to achieve large-scale automated monitoring. In addition, deep learning-based methods cannot effectively distinguish crack characteristics of different sizes, resulting in poor detection of small cracks and loss of high-frequency details, affecting segmentation accuracy and identification effect.
The bridge crack detection method based on the Swin-TransUNet model is adopted, combined with the bidirectional feature pyramid fusion module, frequency domain feature enhancement module and Dice Loss loss function, through bidirectional span-scale connection and frequency domain feature enhancement, the sensitivity to fine cracks is improved, complex background noise interference is suppressed, and category imbalance problem is alleviated.
It significantly improves the accuracy and robustness of bridge crack detection in complex backgrounds, can effectively detect large and small cracks, reduce the false detection rate, and improve detection accuracy and anti-interference ability.
Smart Images

Figure CN120563459A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of bridge crack detection, and in particular to a bridge crack detection method, equipment, medium and product. Background Art
[0002] Bridge crack detection is a key task in bridge structure health monitoring. Existing bridge crack detection technologies are mainly divided into two categories: traditional methods and deep learning-based methods, but there are still problems. Traditional methods mainly include manual detection and image processing-based technologies (such as edge detection, threshold segmentation, morphological operations, etc.). These methods have the following limitations: (1) They rely on manual experience, have low detection efficiency and high cost, and are easily affected by subjective factors, making it difficult to achieve large-scale automated monitoring. (2) Poor environmental adaptability: Traditional image processing methods are sensitive to lighting changes, noise, and complex backgrounds (such as stains and textures on the bridge surface), which can easily lead to false detection or missed detection. (3) Weak small crack detection capabilities: Since cracks usually only occupy a very small number of pixels in the image, traditional methods find it difficult to effectively extract the features of tiny cracks. For detection methods based on deep learning
[0003] Existing technologies often fail to effectively distinguish the features of cracks of different sizes, resulting in poor detection of small cracks and loss of high-frequency details (such as fine cracks), which in turn affects segmentation accuracy and the recognition of crack areas due to class imbalance.
[0004] Based on the above problems, there is an urgent need to provide a new bridge crack detection method to improve the accuracy and robustness of bridge crack detection in complex backgrounds. Summary of the Invention
[0005] The purpose of this application is to provide a bridge crack detection method, equipment, medium and product that can improve the accuracy and robustness of bridge crack detection under complex backgrounds.
[0006] To achieve the above objectives, this application provides the following solutions:
[0007] In a first aspect, the present application provides a bridge crack detection method, the bridge crack detection method comprising:
[0008] Based on the Swin-TransUNet model, a bridge crack detection model is constructed. The bridge crack detection model combines the Swin-TransUNet model with a bidirectional feature pyramid fusion module, a frequency domain feature enhancement module, and a DiceLoss loss function. The bidirectional feature pyramid fusion module is used to perform bidirectional cross-scale connections on the multi-scale feature maps output by the encoder in the Swin-TransUNet model. The frequency domain feature enhancement module is used to transform the bridge crack image into the frequency domain, and concatenate the frequency domain features with the multi-scale feature maps after the bidirectional cross-scale connections along the channel dimension, and input them into the decoder in the Swin-TransUNet model.
[0009] Acquire an image of a bridge crack to be detected;
[0010] According to the bridge crack image to be detected, the bridge crack detection model is adopted to obtain the bridge crack detection result.
[0011] Optionally, the step of obtaining an image of a bridge crack to be detected further includes:
[0012] Preprocessing is performed on the bridge crack image to be detected; the preprocessing includes: data enhancement and data normalization.
[0013] Optionally, the bidirectional feature pyramid fusion module specifically includes the following formula:
[0014]
[0015] in, is the feature map of the current layer i, is the feature map input to the current layer i, is the cross-branch feature map of the current layer i, is the feature map of the adjacent layer i-1, Resize is the scaling operation, w′ j , w′ j+1 , w′ j+2 are all learning weights, ε is a factor to prevent division by zero error, and Conv is a convolution operation.
[0016] Optionally, the frequency domain feature enhancement module performs discrete Fourier transform on the bridge crack image to obtain frequency domain features.
[0017] Optionally, the frequency domain feature enhancement module specifically includes the following formula:
[0018]
[0019] Among them, F(u,v) is the frequency domain feature, M is the image width, N is the image height, M×N is the image size, f(x,y) is the image pixel value in the time domain, j is the complex rotation factor, u and v represent the frequency variables in the frequency domain, corresponding to the horizontal and vertical directions of the image respectively, x and y are the pixel coordinates in the time domain, x is the horizontal direction, and y is the vertical direction.
[0020] Optionally, the Dice Loss loss function specifically includes the following formula:
[0021] L Dice =1-Dice;
[0022] Among them, L Dice is the Dice Loss loss value, Dice is the Dice coefficient, A is the predicted mask and B is the true label.
[0023] In a second aspect, the present application provides a bridge crack detection device, comprising:
[0024] A model construction module is used to construct a bridge crack detection model based on the Swin-TransUNet model; the bridge crack detection model is based on the Swin-TransUNet model and combines a bidirectional feature pyramid fusion module, a frequency domain feature enhancement module, and a Dice Loss loss function; the bidirectional feature pyramid fusion module is used to perform bidirectional cross-scale connections on the multi-scale feature maps output by the encoder in the Swin-TransUNet model; the frequency domain feature enhancement module is used to transform the bridge crack image into the frequency domain, and splice the frequency domain features with the multi-scale feature maps after the bidirectional cross-scale connection along the channel dimension, and input them into the decoder in the Swin-TransUNet model;
[0025] An image acquisition module, used for acquiring an image of a bridge crack to be detected;
[0026] The detection model is used to obtain a bridge crack detection result based on the bridge crack image to be detected using the bridge crack detection model.
[0027] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the bridge crack detection method.
[0028] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the bridge crack detection method when executed by a processor.
[0029] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which implements the bridge crack detection method when executed by a processor.
[0030] According to the specific embodiments provided in this application, this application has the following technical effects:
[0031] The present application provides a bridge crack detection method, equipment, medium and product. Through bidirectional weighted fusion of the bidirectional feature pyramid fusion module (BiFPN), the contribution of feature maps of each scale is dynamically adjusted through learnable weights, so that the bridge crack detection model can effectively detect large cracks and small cracks at the same time. The high-frequency components are extracted through the frequency domain feature enhancement module, and combined with the spatial domain features, the sensitivity to fine cracks is significantly improved. The interference of low-frequency noise such as stains and shadows on the bridge surface is effectively suppressed through frequency domain features, the false detection rate is reduced, and the anti-interference ability is improved. The vanishing gradient problem is then alleviated through Dice Loss, accelerating the convergence of the model. The present application can improve the accuracy and robustness of bridge crack detection under complex backgrounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0033] Figure 1 This is a flow chart of a bridge crack detection method in one embodiment of the present application;
[0034] Figure 2 Schematic diagram of the Swin-TransUNet model structure;
[0035] Figure 3 Schematic diagram comparing the traditional unidirectional feature pyramid and the bidirectional feature pyramid fusion module of this application;
[0036] Figure 4 Schematic diagram of the frequency domain feature enhancement module structure. DETAILED DESCRIPTION
[0037] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0038] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0039] In an exemplary embodiment, Figure 1 As shown, a bridge crack detection method is provided, which includes the following S101 to S103.
[0040] S101, based on Figure 2 The Swin-TransUNet model shown in the figure is used to construct a bridge crack detection model. The bridge crack detection model is based on the Swin-TransUNet model and combines a bidirectional feature pyramid fusion module, a frequency domain feature enhancement module, and a Dice Loss function. The bidirectional feature pyramid fusion module is used to perform bidirectional cross-scale connections on the multi-scale feature maps output by the encoder in the Swin-TransUNet model. The frequency domain feature enhancement module is used to transform the bridge crack image into the frequency domain and concatenate the frequency domain features with the multi-scale feature maps after the bidirectional cross-scale connections along the channel dimension, and input them into the decoder in the Swin-TransUNet model.
[0041] The scale of bridge cracks varies greatly (from fine cracks to wide cracks) and is often mixed with complex background textures (such as concrete particles and rust). Figure 3 As shown in Figure 2, traditional unidirectional feature pyramids (such as FPN or PANet) have difficulty effectively distinguishing crack features at different scales. The Swin-TransUNet model divides bridge crack images into non-overlapping patches (e.g., 4×4 pixel blocks), linearly embeds these patches into the Swin Transformer block, and then gradually downsamples them through multiple layers of Transformer blocks and patch merging to generate multi-scale feature maps (P3-P7).
[0042] The bidirectional feature pyramid fusion module specifically includes the following formula:
[0043]
[0044] in, is the feature map of the current layer i, is the feature map input to the current layer i, is the cross-branch feature map of the current layer i, is the feature map of the adjacent layer i-1, Resize is the scaling operation, w′ j , w′ j+1 , w′ j+2 are all learning weights, ε is a factor to prevent division by zero error, and Conv is a convolution operation.
[0045] like Figure 4 As shown in the figure, multi-scale features are added to represent the image at multiple scales. By performing crack detection at different resolutions (such as 1 / 4, 1 / 8, 1 / 16, and 1 / 32 resolutions), crack features of different sizes can be captured. Frequency domain features are introduced. Cracks are high-frequency signals in the image field. Using discrete Fourier transform, the image is transformed into the frequency domain, and frequency domain information is extracted. This enhances the network model's ability to represent crack features and helps improve classifier performance.
[0046] The frequency domain feature enhancement module performs a discrete Fourier transform (DFT) on the bridge crack image to obtain frequency domain features. A high-pass filter is used to retain high-frequency information (cracks appear as high-frequency signals).
[0047] The frequency domain feature enhancement module specifically includes the following formula:
[0048]
[0049] Among them, F(u,v) is the frequency domain feature, M is the image width, N is the image height, M×N is the image size, f(x,y) is the image pixel value in the time domain, j is the complex rotation factor, u and v represent the frequency variables in the frequency domain, corresponding to the horizontal and vertical directions of the image respectively, x and y are the pixel coordinates in the time domain, x is the horizontal direction, and y is the vertical direction.
[0050] The decoder of the Swin-TransUNet model gradually upsamples the image through a patch expansion layer to restore image resolution, and combines it with the encoder's skip connection to supplement low-level details. The segmentation head uses 1×1 convolution to map the feature map into a binary mask (crack / background), thereby accurately outputting pixel-level crack segmentation results.
[0051] The Dice Loss loss function is used to alleviate the class imbalance problem and improve the model's sensitivity to sparse crack pixels. The Dice Loss loss function specifically includes the following formula:
[0052] L Dice =1-Dice;
[0053] Among them, L Dice is the Dice Loss loss value, Dice is the Dice coefficient, A is the predicted mask and B is the true label.
[0054] S102, obtaining an image of a bridge crack to be detected;
[0055] S102 and later also include:
[0056] Preprocessing is performed on the bridge crack image to be detected; the preprocessing includes: data enhancement and data normalization.
[0057] S103 , using a bridge crack detection model according to the bridge crack image to be detected to obtain a bridge crack detection result.
[0058] On the Bridge Crack dataset, the introduction of BiFPN improved crack pixel accuracy (PA(ck)) from 68.22% to 71.76%, and the overall accuracy (mPA / Accuracy) reached 93.96%. The frequency domain feature enhancement module further improved PA(ck) to 69.56%. Experiments showed that while the Dice Loss function alone only improved PA(ck) by a limited 0.57%, combining it with BiFPN and DFT enabled the final bridge crack detection model to achieve an mPA / Accuracy of 94.76%, a significant improvement over the baseline (90.42%).
[0059] Based on the same inventive concept, embodiments of the present application also provide a bridge crack detection device for implementing the aforementioned bridge crack detection method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more bridge crack detection device embodiments provided below can be found in the aforementioned limitations of the bridge crack detection method and will not be further elaborated here.
[0060] In an exemplary embodiment, a bridge crack detection device is provided, comprising:
[0061] A model construction module is used to construct a bridge crack detection model based on the Swin-TransUNet model; the bridge crack detection model is based on the Swin-TransUNet model and combines a bidirectional feature pyramid fusion module, a frequency domain feature enhancement module, and a Dice Loss loss function; the bidirectional feature pyramid fusion module is used to perform bidirectional cross-scale connections on the multi-scale feature maps output by the encoder in the Swin-TransUNet model; the frequency domain feature enhancement module is used to transform the bridge crack image into the frequency domain, and splice the frequency domain features with the multi-scale feature maps after the bidirectional cross-scale connection along the channel dimension, and input them into the decoder in the Swin-TransUNet model;
[0062] An image acquisition module, used for acquiring an image of a bridge crack to be detected;
[0063] The detection model is used to obtain a bridge crack detection result based on the bridge crack image to be detected using the bridge crack detection model.
[0064] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0065] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0066] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0067] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0068] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0069] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.
[0070] In this application, all actions to obtain signals, information or data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0071] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0072] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A bridge crack detection method, characterized in that: The bridge crack detection method comprises: Based on the Swin-TransUNet model, a bridge crack detection model is constructed. The bridge crack detection model combines the Swin-TransUNet model with a bidirectional feature pyramid fusion module, a frequency domain feature enhancement module, and a Dice Loss function. The bidirectional feature pyramid fusion module is used to perform bidirectional cross-scale connections on the multi-scale feature maps output by the encoder in the Swin-TransUNet model. The frequency domain feature enhancement module is used to transform the bridge crack image into the frequency domain, and concatenate the frequency domain features with the multi-scale feature maps after the bidirectional cross-scale connections along the channel dimension, and input them into the decoder in the Swin-TransUNet model. Acquire an image of a bridge crack to be detected; According to the bridge crack image to be detected, the bridge crack detection model is adopted to obtain the bridge crack detection result.
2. The bridge crack detection method according to claim 1, characterized in that: The step of obtaining the image of the bridge crack to be detected further includes: Preprocessing is performed on the bridge crack image to be detected; the preprocessing includes: data enhancement and data normalization.
3. The bridge crack detection method according to claim 1, characterized in that: The bidirectional feature pyramid fusion module specifically includes the following formula: in, is the feature map of the current layer i, is the feature map input to the current layer i, is the cross-branch feature map of the current layer i, is the feature map of the adjacent layer i-1, Resize is the scaling operation, w′ j , w′ j+1 , w′ j+2 are all learning weights, ε is a factor to prevent division by zero error, and Conv is a convolution operation.
4. The bridge crack detection method according to claim 1, characterized in that: The frequency domain feature enhancement module performs discrete Fourier transform on the bridge crack image to obtain frequency domain features.
5. The bridge crack detection method according to claim 4, characterized in that: The frequency domain feature enhancement module specifically includes the following formula: Among them, F(u,v) is the frequency domain feature, M is the image width, N is the image height, M×N is the image size, f(x,y) is the image pixel value in the time domain, j is the complex rotation factor, u and v represent the frequency variables in the frequency domain, corresponding to the horizontal and vertical directions of the image respectively, x and y are the pixel coordinates in the time domain, x is the horizontal direction, and y is the vertical direction.
6. The bridge crack detection method according to claim 1, characterized in that: The Dice Loss loss function specifically includes the following formula: THE Dice =1-Dice; Among them, L Dice is the Dice Loss loss value, Dice is the Dice coefficient, A is the predicted mask and B is the true label.
7. A bridge crack detection device, characterized in that: The bridge crack detection equipment comprises: A model construction module is used to construct a bridge crack detection model based on the Swin-TransUNet model; the bridge crack detection model is based on the Swin-TransUNet model and combines a bidirectional feature pyramid fusion module, a frequency domain feature enhancement module, and a Dice Loss loss function; the bidirectional feature pyramid fusion module is used to perform bidirectional cross-scale connections on the multi-scale feature maps output by the encoder in the Swin-TransUNet model; the frequency domain feature enhancement module is used to transform the bridge crack image into the frequency domain, and splice the frequency domain features with the multi-scale feature maps after the bidirectional cross-scale connection along the channel dimension, and input them into the decoder in the Swin-TransUNet model; An image acquisition module, used for acquiring an image of a bridge crack to be detected; The detection model is used to obtain a bridge crack detection result based on the bridge crack image to be detected using the bridge crack detection model.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the bridge crack detection method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the bridge crack detection method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the bridge crack detection method according to any one of claims 1 to 6 is implemented.