Bridge crack detection method based on deep learning

By employing a deep learning-based bridge crack detection method, utilizing an improved U-Net model and multi-scale feature extraction, combined with hybrid loss function optimization, accurate and automated detection of bridge pavement cracks is achieved. This solves the problems of low detection accuracy and low efficiency in existing technologies, thereby improving the accuracy and efficiency of detection.

CN122023397AActive Publication Date: 2026-05-12SICHUAN JIAOTONG UNIV ENG TESTING CONSULTING CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN JIAOTONG UNIV ENG TESTING CONSULTING CO LTD
Filing Date
2026-04-10
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing methods for detecting cracks in bridge pavements suffer from low accuracy, low efficiency, and susceptibility to changes in lighting and environmental interference. Furthermore, traditional deep learning models are not optimized for the characteristics of long, thin, and discontinuous cracks in bridge pavements, causing small crack features to be easily obscured by the background, making it difficult to meet the actual needs of engineering projects.

Method used

A deep learning-based bridge crack detection method is adopted. Images are collected through a preset grid path, an improved U-Net model is constructed, multi-scale feature extraction and adaptive fusion are combined, and a hybrid loss function is used to optimize the model to generate a trained U-Net model, thereby achieving accurate crack localization and automated crack detection.

Benefits of technology

It significantly reduces the cost and time of manual inspections, enhances the ability to perceive minute cracks and complex backgrounds, reduces the probability of missed and false detections, and enables rapid, accurate, and automated detection of bridge pavement cracks, providing reliable technical support for infrastructure maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122023397A_ABST
    Figure CN122023397A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bridge crack identification, and discloses a bridge crack detection method based on deep learning, which comprises the following steps of: presetting equal-interval grid paths along a bridge pavement, acquiring an original image of the bridge pavement of each grid, carrying out filtering and enhancement processing, carrying out pixel-level crack labeling, and carrying out depth learning on the original image of the bridge pavement; generating a binary label graph and a bridge pavement image after filtering and enhancement processing; an improved U-Net model is constructed; taking the bridge pavement image as the input of an improved U-Net model, taking the binary label graph as the output, training the improved U-Net model, optimizing the trained improved U-Net model by adopting a mixed loss function, and generating a trained improved U-Net model; an original image of a to-be-detected bridge pavement is obtained, after filtering and enhancement processing, the trained improved U-Net model is input for crack detection, and a crack recognition result is obtained; according to the invention, rapid, accurate and automatic detection of bridge pavement cracks is realized, and reliable technical support is provided for infrastructure maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bridge crack identification technology, and more specifically to a bridge crack detection method based on deep learning. Background Technology

[0002] Bridge pavement cracks are a significant manifestation of bridge structural damage, and timely and accurate crack detection is crucial for bridge maintenance and extending bridge service life. Currently, bridge pavement crack detection mainly employs two methods: traditional manual inspection and image processing-based automated detection.

[0003] Traditional manual inspection relies on visual observation by inspectors, which is inefficient, labor-intensive, and easily affected by subjective experience and ambient lighting, resulting in a high rate of missed detection for small cracks. Early automatic inspection methods based on image processing employed traditional algorithms such as threshold segmentation and edge detection, which are weakly resistant to interference from changes in lighting and road surface stains, leading to low crack segmentation accuracy. Existing deep learning-based methods often directly use general segmentation models such as the original U-Net, without optimization for the characteristics of bridge and pavement cracks—long, discontinuous, and with low contrast. Multi-scale feature extraction is insufficient, and small crack features are easily obscured by the background, resulting in poor extraction accuracy and failing to meet practical engineering needs. Summary of the Invention

[0004] To address the aforementioned shortcomings in existing technologies, this invention provides a deep learning-based bridge crack detection method to solve the problem of low detection accuracy in existing bridge pavement crack detection methods.

[0005] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: A deep learning-based method for bridge crack detection includes the following steps: S1. Along the bridge road surface, a pre-set grid path with equal spacing is used to collect the original images of the bridge road surface of each grid. After filtering and enhancement processing, pixel-level crack annotation is performed to generate a binary label image and a bridge road surface image after filtering and enhancement processing. S2. Construct an improved U-Net model, including a multi-scale feature extraction module, a multi-scale feature fusion module, and a decoder; S3. Use the bridge road surface image as input to the improved U-Net model and the binary label image as output to train the improved U-Net model. Use the hybrid loss function to optimize the trained improved U-Net model and generate the trained improved U-Net model. S4. Obtain the original image of the bridge pavement to be tested. After filtering and enhancement, input it into the trained improved U-Net model for crack detection. Determine whether cracks exist. If so, output a binary segmentation image of the crack. Otherwise, the bridge pavement to be tested does not have cracks.

[0006] The present invention has the following beneficial effects: The bridge crack detection method proposed in this invention systematically acquires images through a pre-defined grid path and combines it with an improved U-Net model to achieve automated crack identification, significantly reducing the cost and time of manual inspection. Simultaneously, through multi-scale feature extraction and adaptive fusion, the model's ability to perceive subtle cracks and complex backgrounds is enhanced, achieving accurate crack localization. Supervised training using binary label maps ensures that the model outputs clear crack-free determinations or crack segmentation results, reducing the probability of missed and false detections. Ultimately, this method achieves rapid, accurate, and automated detection of bridge pavement cracks, providing reliable technical support for infrastructure maintenance. Attached Figure Description

[0007] Figure 1 This is a schematic diagram of the bridge crack detection method based on deep learning proposed in this invention. Figure 2 This is a schematic diagram of the original image of the bridge surface in the embodiment; Figure 3 This is a schematic diagram with annotations of the original image of the bridge pavement in the embodiment; Figure 4 This is a schematic diagram of the structure of the improved U-Net model in the embodiment; Figure 5 This is a schematic diagram of the multi-scale feature extraction module in the embodiment; Figure 6 This is a schematic diagram of the decoder structure in the embodiment. Detailed Implementation

[0008] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0009] like Figure 1 As shown, the bridge crack detection method based on deep learning includes the following steps S1-S4: S1. Along the bridge road surface, a pre-set grid path with equal spacing is used to collect the original images of the bridge road surface of each grid. After filtering and enhancement processing, pixel-level crack annotation is performed to generate a binary label image and a filtered and enhanced bridge road surface image.

[0010] In this embodiment, a mobile inspection device equipped with a high-definition industrial camera can be used to acquire original images along a preset path on the bridge surface to obtain an image dataset. The high-definition industrial camera has a resolution of at least 5120×3840 to ensure sufficient spatial detail for pixel-level crack identification; the frame rate is set to 10-20fps, matching the travel speed of the mobile platform to ensure a 60%-80% overlap between adjacent frames; the vertical distance between the lens and the bridge surface during acquisition is 1.5-2.0m to ensure sufficient image resolution for subsequent fine annotation while maintaining the field of view coverage and acquisition efficiency for each image. The acquisition path uses an equally spaced grid, with a horizontal spacing of no more than 0.5m between adjacent acquisition points to achieve full coverage of the bridge surface. This spacing ensures reasonable overlap between adjacent images and avoids blind spots.

[0011] Secondly, the original bridge surface images undergo data preprocessing, including image filtering and enhancement operations. Gaussian filtering is used for noise removal to effectively suppress Gaussian noise introduced by the camera sensor. Data enhancement includes geometric transformation and pixel transformation: geometric transformation involves randomly flipping, rotating, and scaling the image to increase the diversity of the dataset; pixel transformation involves randomly adjusting brightness and contrast or performing gamma correction to simulate complex lighting conditions, optimize image grayscale distribution, and improve the model's generalization ability.

[0012] Then, using professional image annotation software, skilled technicians perform fine-grained pixel-level crack annotations on the filtered and enhanced bridge pavement images, thereby generating binary label images, which are used as supervisory labels for model training. Examples of the original bridge pavement images and corresponding annotations collected in this embodiment are shown below. Figures 2-3 As shown, Figure 2 The original image is the input. Figure 3 The corresponding labeled image has a value of 1 for the pixels in the crack area (visually displayed as white) and a value of 0 for the pixels in the background area (visually displayed as black), which is used to provide pixel-level supervision signals for model training.

[0013] S2. Construct an improved U-Net model.

[0014] In this embodiment, the structure and connection relationships of the improved U-Net model are as follows: Figure 4 As shown, it includes a multi-scale feature extraction module, a multi-scale feature fusion module, and a decoder.

[0015] The structure and connection relationships of the multi-scale feature extraction module are as follows: Figure 5The encoder structure shown includes a first convolutional block, a second convolutional block, a third convolutional block, a fourth convolutional block, a first downsampling module, a second downsampling module, a third downsampling module, a fourth downsampling module, a first gap saliency subnetwork, a second gap saliency subnetwork, a third gap saliency subnetwork, and a fourth gap saliency subnetwork; wherein, Figure 5 Medium parameters These represent the extracted first-scale feature map, second-scale feature map, third-scale feature map, and fourth-scale feature map, respectively, with parameters... These represent the crack response values ​​of the first-scale feature map, the second-scale feature map, the third-scale feature map, and the fourth-scale feature map, respectively.

[0016] Furthermore, the first, second, third, and fourth convolutional blocks each include two 3×3 convolutional layers, one batch normalization (BN) layer, and one ReLU activation function layer connected in sequence; and the first, second, third, and fourth downsampling modules are all 2×2 max pooling layers. Therefore, when the preprocessed image enters the encoder, it first performs feature extraction from shallow to deep through four convolutional blocks, and then performs downsampling using max pooling after each layer, finally outputting four feature maps of different scales, thereby covering the feature information of cracks from fine to coarse.

[0017] Furthermore, the first, second, third, and fourth crack saliency subnetworks are lightweight subnetwork structures, each including a sequentially connected 1×1 convolutional layer, a global average pooling layer (GAP), and a fully connected layer. Therefore, this invention adds a lightweight subnetwork next to each downsampling node of the encoder, using a simplified structure of 1×1 convolution, global average pooling, and fully connected layers to calculate the crack response values ​​(as scalar values) of feature maps at each scale in real time, thereby providing adaptive weights for multi-scale fusion. The working principle of the first crack saliency subnetwork is as follows: First, after the first convolutional block performs preliminary feature extraction on the input image, it outputs shallow features, which include low-level features such as crack edges and textures. Then, after max pooling by the first downsampling module, a first-scale feature map is obtained. ; Secondly, the first crack saliency subnetwork receives the first-scale feature map and performs channel compression on the first-scale feature map through a 1×1 convolutional layer. If the number of channels in the first-scale feature map is... Then the number of channels after compression is (in The number of channels is a positive integer (multiples of 4), generating a dimensionality-reduced feature map. This 1×1 convolution operation not only reduces computational cost while compressing channels, but also preserves and enhances semantic features related to cracks and filters out irrelevant background noise. Then, the global average pooling layer receives the dimensionality-reduced feature map output from the 1×1 convolutional layer, calculates the average value of all pixels for each channel, and if the height of the dimensionality-reduced feature map is... , width is Then, after global average pooling, the generated... The feature vector, i.e., the feature map has a height of 1, a width of 1, and a number of channels. Each element corresponds to the global average feature of a channel, thereby aggregating the crack semantic information of the entire first-scale feature map; Finally, The feature vectors are input to the fully connected layer for linear mapping, and the input dimension is... The fully connected layer maps it to a 1-dimensional scalar, and then performs non-linear activation using the sigmoid function, constraining its value between 0 and 1 for easy quantization and comparison, thus obtaining the crack response value of the first-scale feature map. ,Right now:

[0018] in, For fully connected operation, for eigenvectors.

[0019] The working principles of the second, third, and fourth crack saliency subnetworks are the same as those of the first crack saliency subnetwork.

[0020] The structure and connection relationships of the decoder are as follows: Figure 6 As shown, it includes a first deconvolution block, a fifth convolution block, a second deconvolution block, a sixth convolution block, a third deconvolution block, a seventh convolution block, and an output layer connected in sequence.

[0021] Furthermore, the first, second, and third deconvolutional blocks each include a 2×2 deconvolutional layer; the fifth, sixth, and seventh convolutional blocks each include two 3×3 convolutional layers, one batch normalization (BN) layer, and one ReLU activation function layer; and the output layer is a 1×1 convolutional layer. Therefore, the decoder upsamples through three deconvolutional blocks to restore the feature map size, and then concatenates it with the corresponding scale feature map of the encoder to supplement detailed information and improve segmentation accuracy. Finally, it outputs a binary segmentation map of the crack through a 1×1 convolution.

[0022] S3. Use the bridge road surface image as input to the improved U-Net model and the binary label image as output to train the improved U-Net model. Use a hybrid loss function to optimize the trained improved U-Net model and generate a trained improved U-Net model.

[0023] In this embodiment, after multi-scale feature extraction combined with adaptive feature fusion, refined decoding is performed, and a hybrid loss function is introduced for optimization. This allows the improved U-Net model to be trained, enabling the trained model to possess accurate and robust binary segmentation capabilities for bridge pavement cracks. This solves the problems of insufficient capture of small cracks and significant background interference in the traditional U-Net model. Furthermore, the hybrid loss function optimization alleviates the shortcomings of unstable training gradients and low segmentation accuracy for small targets, ultimately resulting in a crack segmentation model with strong generalization ability and high segmentation accuracy. The operation process is as follows: Specifically, step S3 includes S31-S35: S31. Using the bridge pavement image as input data and the binary label image as supervision label, the input data is fed into the multi-scale feature extraction module of the improved U-Net model for multi-scale feature extraction, generating multi-scale feature maps and their corresponding crack response values.

[0024] Specifically, the process of inputting the input data into the multi-scale feature extraction module of the improved U-Net model to perform multi-scale feature extraction and generate multi-scale feature maps and their corresponding crack response values ​​is as follows: The input data is fed into the first convolutional block for convolution, batch normalization, and activation operations. Then, it is fed into the first downsampling module for downsampling to generate a first-scale feature map. Simultaneously, the first-scale feature map is fed into the first crack saliency sub-network for convolution, global average pooling, and activation operations. By extracting crack semantic features, crack response values ​​of the first-scale feature map are generated.

[0025] The first-scale feature map is input into the second convolutional block for convolution, batch normalization, and activation operations. Then, it is input into the second downsampling module for downsampling to generate the second-scale feature map. At the same time, the second-scale feature map is input into the second crack saliency sub-network for convolution, global average pooling, and activation operations. By extracting crack semantic features, the crack response value of the second-scale feature map is generated.

[0026] The second-scale feature map is input into the third convolutional block for convolution, batch normalization, and activation operations. Then, it is input into the third downsampling module for downsampling to generate the third-scale feature map. At the same time, the third-scale feature map is input into the third crack saliency sub-network for convolution, global average pooling, and activation operations. By extracting crack semantic features, the crack response value of the third-scale feature map is generated.

[0027] The third-scale feature map is input into the fourth convolutional block for convolution, batch normalization, and activation operations. Then, it is input into the fourth downsampling module for downsampling to generate the fourth-scale feature map. At the same time, the fourth-scale feature map is input into the fourth crack saliency sub-network for convolution, global average pooling, and activation operations. By extracting crack semantic features, the crack response value of the fourth-scale feature map is generated.

[0028] In this embodiment, not only is full coverage of multi-scale crack features achieved, but the semantic response of cracks is also enhanced, providing high-quality and highly discriminative basic features for subsequent crack feature fusion. Specifically, by combining four-level downsampling and convolutional blocks, feature maps of four scales are generated sequentially to capture the feature information of cracks at different scales in the bridge pavement, avoiding the omission of crack or wide crack features at a single scale. Furthermore, each scale feature map corresponds to a crack saliency sub-network, which extracts crack semantic features and generates response values ​​through convolution, global average pooling, and activation operations, effectively amplifying the feature differences between cracks and the pavement background and suppressing interference from non-crack areas.

[0029] S32. Input the multi-scale feature map and its corresponding crack response value into the multi-scale feature fusion module for feature fusion to generate a fused feature map, specifically: The crack response values ​​of the multi-scale feature maps are normalized using the softmax function to obtain the adaptive weights for each scale feature map, i.e.:

[0030] in, Indicates the first Attention weights for scale feature maps , They represent the first , No. Crack response values ​​from scale feature maps This represents an exponential function.

[0031] In this embodiment, the crack response values ​​of the four scales are normalized by the softmax function to obtain adaptive weights, which improves the effectiveness of dynamically judging the features of each scale. That is, scales with high crack response values, such as small scales containing fine cracks and large scales containing complete cracks, are assigned higher weights, while invalid scales, such as background scales without cracks, are given lower weights, thus avoiding the submergence of effective features caused by traditional fixed weight fusion.

[0032] The multi-scale feature maps and their corresponding adaptive weights are weighted and summed to generate a fused feature map, i.e.:

[0033] in, Represents the fused feature map. Indicates the first Scale feature map.

[0034] In this embodiment, the fusion feature map generated by weighted summation integrates the most effective information for crack segmentation at each scale, preserving both the detailed features of fine cracks and the overall morphological features of cracks, thus enabling the feature map to have a dual discriminative capability that combines fine-grained details with coarse-grained structure.

[0035] S33. The fused feature map is input into the decoder for deconvolution, and then skip connections and convolution operations are performed with the feature maps of the corresponding scale from the multi-scale feature extraction module to progressively upsample the data. Finally, a binary segmentation map of the crack is output, specifically: S331. After inputting the fused feature map into the first deconvolution block for deconvolution operation, it is concatenated with the third-scale feature map to obtain the first concatenated feature. The first concatenated feature is then input into the fifth convolution block for convolution, batch normalization, and activation operations to obtain the first feature map.

[0036] S332. After inputting the first feature map into the second deconvolution block for deconvolution operation, it is concatenated with the second scale feature map to obtain the second concatenated feature. The second concatenated feature is then input into the sixth convolution block for convolution, batch normalization, and activation operations to obtain the second feature map.

[0037] S333. After inputting the second feature map into the third deconvolution block for deconvolution, it is concatenated with the first scale feature map to obtain the third concatenated feature. The third concatenated feature is then input into the seventh convolution block for convolution, batch normalization, and activation operations to obtain the third feature map.

[0038] S334. The third feature map is input to the output layer and convolutional operation is performed to obtain the binary segmentation map of the crack.

[0039] In this embodiment, a three-stage deconvolution is used to progressively upsample the fused feature map, restoring it from a low resolution to a resolution matching the input image. This avoids the loss of detailed information such as crack edges and endpoints during the upsampling process. Simultaneously, the deconvolutioned feature map in the decoder is concatenated with the corresponding scale feature map in the encoder (multi-scale feature extraction module). This directly introduces the original detailed features at each scale retained by the encoder into the decoder, solving the problems of feature blurring and detail degradation during deconvolution. This results in a clearer edge and a more complete crack morphology in the final output binary crack segmentation map. Finally, the concatenated features are further fused with high- and low-dimensional features through convolutional blocks, batch normalization, and activation operations. This eliminates feature redundancy caused by concatenation and improves feature representation efficiency.

[0040] S34. Calculate the mixed loss function of Dice loss and cross-entropy loss, and use it as the mixed loss value of the binary segmentation map and the binary label map of the crack.

[0041] Specifically, the formula for calculating the hybrid loss function of Dice loss and cross-entropy loss is as follows:

[0042]

[0043]

[0044] in, Represents the mixed loss function. Represents the Dice loss function. Represents the cross-entropy loss function. Represents pixels The true value is given, and in the binary label image, the crack area is 1 and the background area is 0. Represents the logarithmic function. Represents pixels The predicted value, This represents a smoothing term with a value of 0.00001. Its purpose is to avoid numerical instability when both the numerator and denominator are 0, such as preventing division by zero when neither the actual label nor the prediction has any cracks.

[0045] In this embodiment, the Dice loss function, by calculating the overlap between the predicted segmentation map and the ground truth label map (i.e., the binary label map), is more sensitive to the segmentation error of small cracks (small targets), effectively improving the segmentation accuracy of micro-cracks and slits, and avoiding the problem of traditional loss functions ignoring small targets; the cross-entropy loss function, on the other hand, calculates the error using log probability, optimizing the gradient flow during training, making model training more stable, and reducing the risk of gradient explosion or vanishing; therefore, the Dice loss function is adopted... The weight allocation not only ensures the segmentation accuracy of small cracks, i.e., the core segmentation targets, but also stabilizes the overall training process through cross-entropy loss, achieving a dual optimization of accurate segmentation of small targets and stable global segmentation results.

[0046] S35. Update the network parameters of the improved U-Net model using the mixed loss value to obtain the trained improved U-Net model.

[0047] In this embodiment, the network parameters of the improved U-Net model are adaptively updated through backpropagation of the mixed loss values, so that the model parameters are continuously optimized in the direction of reducing crack segmentation error. Finally, a training model with optimal parameters and strong generalization ability is obtained, so that it can still accurately distinguish cracks from the background and output high-quality binary segmentation maps on unseen bridge pavement images.

[0048] S4. Obtain the original image of the bridge pavement to be tested. After filtering and enhancement, input it into the trained improved U-Net model for crack detection. Determine whether cracks exist. If so, output a binary segmentation image of the crack. Otherwise, the bridge pavement to be tested does not have cracks.

[0049] In this embodiment, leveraging the multi-scale feature extraction, adaptive fusion, and fine-grained decoding capabilities of the trained, improved U-Net model, the model can accurately identify cracks of different scales in the test image, including small micro-cracks, medium-length cracks, and clustered cracks. Simultaneously, it effectively distinguishes cracks from interfering factors such as road stains and normal textures, solving the problems of missed detection of small cracks and false detection of interfering textures in traditional detection methods, thus achieving high-precision automated crack detection. Furthermore, for the detected cracks, i.e., the crack binary segmentation map output by the model, the crack location and area can be determined based on this binary segmentation map, enabling timely crack repair.

[0050] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

[0051] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A bridge crack detection method based on deep learning, characterized in that, Includes the following steps: S1. Along the bridge road surface, a pre-set grid path with equal spacing is used to collect the original images of the bridge road surface of each grid. After filtering and enhancement processing, pixel-level crack annotation is performed to generate a binary label image and a bridge road surface image after filtering and enhancement processing. S2. Construct an improved U-Net model, including a multi-scale feature extraction module, a multi-scale feature fusion module, and a decoder; S3. Use the bridge road surface image as input to the improved U-Net model and the binary label image as output to train the improved U-Net model. Use the hybrid loss function to optimize the trained improved U-Net model and generate the trained improved U-Net model. S4. Obtain the original image of the bridge pavement to be tested. After filtering and enhancement, input it into the trained improved U-Net model for crack detection. Determine whether cracks exist. If so, output a binary segmentation image of the crack. Otherwise, the bridge pavement to be tested does not have cracks.

2. The bridge crack detection method based on deep learning according to claim 1, characterized in that, The multi-scale feature extraction module includes a first convolutional block, a second convolutional block, a third convolutional block, a fourth convolutional block, a first downsampling module, a second downsampling module, a third downsampling module, a fourth downsampling module, a first crack saliency subnetwork, a second crack saliency subnetwork, a third crack saliency subnetwork, and a fourth crack saliency subnetwork.

3. The bridge crack detection method based on deep learning according to claim 2, characterized in that, The decoder includes a first deconvolution block, a second deconvolution block, a third deconvolution block, a fifth convolution block, a sixth convolution block, a seventh convolution block, and an output layer.

4. The bridge crack detection method based on deep learning according to claim 3, characterized in that, Step S3 specifically includes: S31. Take the bridge pavement image as input data, the binary label image as supervision label, and input the input data into the multi-scale feature extraction module of the improved U-Net model to perform multi-scale feature extraction, and generate multi-scale feature maps and their corresponding crack response values. S32. Input the multi-scale feature map and its corresponding crack response value into the multi-scale feature fusion module to perform feature fusion and generate a fused feature map. S33. Input the fused feature map into the decoder for deconvolution, and perform skip connections and convolution operations with the feature map of the corresponding scale in the multi-scale feature extraction module to gradually upsample and finally output the binary segmentation map of the crack. S34. Calculate the mixed loss function of Dice loss and cross-entropy loss, and use it as the mixed loss value of the binary segmentation map and the binary label map of the crack; S35. Update the network parameters of the improved U-Net model using the mixed loss value to obtain the trained improved U-Net model.

5. The bridge crack detection method based on deep learning according to claim 4, characterized in that, The process of inputting the input data into the multi-scale feature extraction module of the improved U-Net model to extract multi-scale features and generate multi-scale feature maps and their corresponding crack response values ​​is as follows: After inputting the data into the first convolutional block for convolution, batch normalization, and activation, it is then input into the first downsampling module for downsampling to generate a first-scale feature map. Simultaneously, the first-scale feature map is input into the first crack saliency sub-network for convolution, global average pooling, and activation. By extracting crack semantic features, crack response values ​​of the first-scale feature map are generated. The first-scale feature map is input into the second convolutional block for convolution, batch normalization, and activation operations. Then, it is input into the second downsampling module for downsampling to generate the second-scale feature map. At the same time, the second-scale feature map is input into the second crack saliency sub-network for convolution, global average pooling, and activation operations. By extracting crack semantic features, the crack response value of the second-scale feature map is generated. The second-scale feature map is input into the third convolutional block for convolution, batch normalization, and activation operations. Then, it is input into the third downsampling module for downsampling to generate the third-scale feature map. At the same time, the third-scale feature map is input into the third crack saliency sub-network for convolution, global average pooling, and activation operations. By extracting crack semantic features, the crack response value of the third-scale feature map is generated. The third-scale feature map is input into the fourth convolutional block for convolution, batch normalization, and activation operations. Then, it is input into the fourth downsampling module for downsampling to generate the fourth-scale feature map. At the same time, the fourth-scale feature map is input into the fourth crack saliency sub-network for convolution, global average pooling, and activation operations. By extracting crack semantic features, the crack response value of the fourth-scale feature map is generated.

6. The bridge crack detection method based on deep learning according to claim 5, characterized in that, Step S32 specifically includes: The crack response values ​​of the multi-scale feature maps are normalized using the softmax function to obtain the adaptive weights for each scale feature map, i.e.: in, Indicates the first Attention weights for scale feature maps , They represent the first , No. Crack response values ​​from scale feature maps Represents an exponential function; The multi-scale feature maps and their corresponding adaptive weights are weighted and summed to generate a fused feature map, i.e.: in, Represents the fused feature map. Indicates the first Scale feature map.

7. The bridge crack detection method based on deep learning according to claim 6, characterized in that, Step S33 specifically includes: S331. After inputting the fused feature map into the first deconvolution block for deconvolution operation, it is concatenated with the third scale feature map to obtain the first concatenated feature. The first concatenated feature is then input into the fifth convolution block for convolution, batch normalization, and activation operations to obtain the first feature map. S332. After inputting the first feature map into the second deconvolution block for deconvolution operation, it is concatenated with the second scale feature map to obtain the second concatenated feature. The second concatenated feature is then input into the sixth convolution block for convolution, batch normalization, and activation operations to obtain the second feature map. S333. After inputting the second feature map into the third deconvolution block for deconvolution, it is concatenated with the first scale feature map to obtain the third concatenated feature. The third concatenated feature is then input into the seventh convolution block for convolution, batch normalization, and activation operations to obtain the third feature map. S334. The third feature map is input to the output layer and convolutional operation is performed to obtain the binary segmentation map of the crack.

8. The bridge crack detection method based on deep learning according to claim 7, characterized in that, The formula for calculating the hybrid loss function of Dice loss and cross-entropy loss is as follows: in, Represents the mixed loss function. Represents the Dice loss function. Represents the cross-entropy loss function. Represents pixels The true value, Represents the logarithmic function. Represents pixels The predicted value, This indicates the smoothing term.