Bridge defect detection method based on UGMB multi-scale feature extraction and fusion

By improving the YOLOv11n network, introducing FasterCGLU, FPSharedConv and DTAB modules, and designing the AMSF-Pyramid-YOLOv11n model, the multi-scale and complexity problems in bridge defect detection are solved, and high-precision and efficient bridge defect detection is achieved.

CN120612288APending Publication Date: 2025-09-09HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510635156.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In the existing technology, the YOLO series of detection algorithms in bridge defect detection have insufficient detection accuracy and efficiency due to the complexity and multi-scale problems of the data set, and cannot meet the needs of safety hazard investigation.

Method used

A bridge defect detection method based on UGMB multi-scale feature extraction and fusion was adopted. By improving the YOLOv11n network, introducing FasterCGLU, FPSharedConv and DTAB modules, and designing the AMSF-Pyramid-YOLOv11n model, the feature pyramid structure and feature fusion method were optimized to enhance the multi-scale target detection capability.

Benefits of technology

It significantly improves the accuracy and speed of bridge defect detection, enables more timely provision of bridge maintenance plans, and enhances the model's detection performance in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612288A_ABST
    Figure CN120612288A_ABST
Patent Text Reader

Abstract

The invention discloses a bridge defect detection method based on UGMB multi-scale feature extraction and fusion. The method comprises the following steps: designing a feature extraction module MANetFaster CGLU; according to the method, feature information of different scales is fused more efficiently, the detection capability of the model on targets of different sizes is enhanced, and an original SPPF module is replaced with a feature pyramid shared convolution module FeaturePyramidShared Conv; in the aspect of an attention mechanism, in order to improve the attention and characterization capability of the model on target features, a DTAB module which performs feature enhancement on space and channel dimensions through Dilatation G-CSA and Dilatation M-WSA is adopted to replace a C2PSA module; a brand-new multi-scale feature extraction and fusion pyramid module UGMB is designed, and the module further improves the detection performance of the model for a multi-scale target by optimizing the structure of a feature pyramid and a feature fusion mode, and provides a more stable and more efficient solution for bridge defect detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to a bridge defect detection method based on UGMB multi-scale feature extraction and fusion. Background Art

[0002] Bridges play an indispensable role in transportation networks. However, due to their long-term exposure to the elements and the impact of various factors, such as vehicle loads, wind loads, and temperature fluctuations, they are prone to various structural defects. These include, but are not limited to, corrosion, cracks, degraded concrete, concrete voids, moisture, spalling, and carbonization in concrete structures, and rust and fatigue cracks in steel structures. If these defects are not detected and addressed promptly, they will gradually develop and worsen, seriously affecting the bridge's load-bearing capacity and service life, and may even lead to major safety accidents such as bridge collapse, resulting in significant economic losses and casualties. Therefore, accurate and timely detection of bridge defects is of paramount importance.

[0003] Existing one-stage detection algorithms, such as the YOLO series, have improved the accuracy and efficiency of bridge defect detection. However, due to the varying angles and lighting conditions of their datasets, bridge defects appear small in images, with small distances between classes. Furthermore, the varying sizes and shapes of bridge defects, coupled with the presence of significant image noise, can lead to misjudgments during subsequent hazard detection, making them unable to meet the high demands of bridge inspections for safety hazard detection. Therefore, a new detection model is urgently needed that can further improve detection accuracy and develop more accurate hazard detection plans by addressing the limitations of complex defect morphologies, dense defect distribution, and multi-scale defect recognition. Summary of the Invention

[0004] Purpose of the invention: In order to solve the problems mentioned in the background technology, the present invention discloses a bridge defect detection method based on UGMB multi-scale feature extraction and fusion. By taking the YOLOv11n network as the benchmark and designing a detection model with a customized network architecture, high-precision detection of bridge defects of different scales in real-world collected data can be achieved.

[0005] Technical solution:

[0006] The present invention discloses a bridge defect detection method based on UGMB multi-scale feature extraction and fusion, the method comprising the following steps:

[0007] S1 builds a bridge defect detection dataset;

[0008] S2 uses YOLOv11n as the baseline model for improvement and builds the AMSF-Pyramid-YOLOv11n bridge defect detection model;

[0009] S2.1 introduces the FasterCGLU structure to replace the ConvNeck in the MANet structure to form the MANet_FasterCGLU structure;

[0010] S2.2 uses the FPSharedConv module to perform feature extraction using shared convolution kernels at different expansion rates instead of the traditional SPPF module;

[0011] S2.3 replaces the traditional C2PSA module with the DTAB module that combines channel attention and spatial attention mechanisms. It also designs a multi-scale feature extraction and fusion pyramid module UGMB by optimizing the feature pyramid structure and feature fusion method, replacing the baseline model neck structure.

[0012] S3 obtains the final model, sets training hyperparameters, and uses the training set to train and verify the improved model; S4 deploys the final model to implement bridge defect detection.

[0013] Furthermore, the dataset described in S1 includes corrosion, cracks, degraded concrete, concrete voids, wet bridge column surfaces and pavement surface defects, and the dataset is divided into a training set, a validation set and a test set in a ratio of 8:1:1.

[0014] Furthermore, the FasterCGLU structure described in S2.1 is implemented as follows:

[0015] FasterCGLU first determines whether the number of input channels inc is consistent with the number of output channels dim. If not, 1x1 convolution is used to adjust the number of channels. For input feature maps of size (N, C in ,H,W), where N is the batch size, C in is the number of input channels, H and W are the height and width of the feature map. After 1x1 convolution, the size of the output feature map is N×C out ×H×W, where Cout=dim, the mathematical expression of this layer is: adjusted =Conv 1×1 (X); spatial feature mixing is performed through partial convolution, and the input is X adjusted , output features Figure X mixed The size is still N×dim×H×W, and the mathematical expression of the partial convolution is: mixed =PartialConv(X adjusted ); By combining the multi-layer perceptron with gated linear units, the input is X mixed , output features Figure X mlp The size is still N×dim×H×W, and the mathematical expression of this layer is: mlp=GLU(X mixed ); finally, through the residual connection layer and the droppath layer; the residual connection layer will process the features Figure X mlp With the original input features Figure X adjusted Add, form residual connection, output features Figure X residual The size is still (N, dim, H, W), the droppath layer is used for path discarding to prevent overfitting, and the input is X residual , output features Figure X output The size is still (N, dim, H, W), and the mathematical expression of the droppath layer is X output =DropPath(X residual ).

[0016] Furthermore, the FPSharedConv module structure described in S2.2 is implemented as follows:

[0017] For an input feature map of size N×C in ×H×W, using 1x1 convolution to reduce the number of input channels to C in / / 2, output features Figure X The size of 1 is N×C in / / 2×H×W, the mathematical expression of 1x1 convolution is X1=Conv 1×1 (X); perform multiple dilated convolutions on X1, each using the same shared convolution kernel but a different dilation rate d. The mathematical expression of dilated convolution is:

[0018]

[0019] Among them, X d It indicates the result of a certain position in the output feature map after the dilation convolution operation with the dilation rate d; X i+d·k,j+d·l Represents the value of the element with coordinates (i+d·k,j+d·l) in the input feature map, where i and j are the base coordinates of the input feature map, d is the dilation rate, and k and l are the row and column indices in the convolution kernel; W k,l represents the weight value at coordinate (k, l) in the shared convolution kernel, where k and l are the number of rows and columns of the convolution kernel, respectively, i.e. the convolution kernel size is K×L; b represents the bias term;

[0020] Output features Figure X d The size is still N×C in ×H×W; all the feature maps after the expansion convolution are spliced ​​according to the channel dimension to obtain a size of N×C in / / 2×(1+len(dilations))×H×W features Figure X concat ; Use 1×1 convolution to adjust the number of channels of the concatenated feature map to the number of output channels C out , and obtain the final output features Figure X output , size N×C out ×H×W.

[0021] Furthermore, the DTAB module structure described in S2.3 is composed of three functional regions: DilatedG-CSA, DilatedFFN and DilatedM-WSA. The three functional regions are connected in series in sequence and perform feature transfer.

[0022] Furthermore, the DilatedG-CSA part: for the input features Figure X Perform normalization and get X norm , using three parallel dilated depthwise separable convolution operations on X norm After processing, three feature maps F1, F2, and F3 are obtained. The feature maps are spliced ​​in the channel dimension to obtain F concat , through the element-by-element multiplication operation and the input feature X, the feature interaction between channels is performed, and the enhanced feature map F is obtained through the residual connection gcsa :

[0023] The DilatedFFN part: Feature map F after DilatedG-CSA processing gcsa Enter this section and click F gcsa Perform normalization and get F norm , further extract the feature F through an expansion depth separable convolution operation conv1 , apply the GeLU activation function to get F gelu , perform dilated depth-separable convolution to obtain F conv2 ;

[0024] The DilatedM-WSA part: the characteristic graph F ffn Perform normalization and get F norm , the channel dimension is reduced by three 1×1 convolution operations, F conv1 ,F conv2 ,F conv3 The feature map after dimensionality reduction is divided into windows in the spatial dimension. The features in each window interact through the attention mechanism. The attention mechanism dynamically adjusts the feature weight by calculating the correlation between the query Q, key K, and value V to obtain the attention score A:

[0025]

[0026] Apply it to the value V to get the feature interaction result F within the window attn :

[0027] F attn =A·V(4)

[0028] Among them, dk is the dimension of the key. Finally, the information in the window is integrated through feature aggregation operation to complete the enhancement of spatial features.

[0029] Furthermore, the UGMB module structure described in S2.3 is implemented as follows:

[0030] In terms of feature extraction and upsampling, the WFU module is applied to upsample and enhance feature maps at different levels, convert low-resolution feature maps into high-resolution feature representations, and fuse multi-scale information at the same time; the MANet_FasterCGLU module is used to further extract and enhance the feature maps; the information flow is dynamically controlled through partial convolution and gated linear units GLU; through multiple Concat operations, feature maps from different sources and different scales are spliced ​​to integrate multi-scale information, and finally the fused feature maps are subjected to hierarchical attention fusion through the HAFB module to highlight important feature areas and suppress irrelevant feature information.

[0031] Furthermore, the WFU module structure is implemented as follows:

[0032] The WFU module receives two sets of input feature maps; Figure X big Size: H×W×C big and small features Figure X small Size is H′×W′×C small , through the HaarWavelet module for large features Figure X big Perform wavelet transform and decompose it into four sub-bands: approximate component A, horizontal detail component H, vertical detail component V and diagonal detail component D. After decomposition, add the horizontal, vertical and diagonal detail components to get the sum of detail components H+V+D, and then perform feature enhancement through residual block, which contains two 3×3 convolutional layers and a residual connection. Figure X small The enhanced detail component and the transformed approximate component are spliced ​​according to the channel dimension, and the feature map is reconstructed by inverse wavelet transform.

[0033] Furthermore, the HAFB module structure is implemented as follows:

[0034] The HAFB module receives two sets of input feature maps, denoted as X1: size H×W×C1 and X2: size H×W×C1, through two 1×1 convolutional layers W x1 and W x2 Adjust the number of channels of X1 and X2 respectively and convert them into feature maps W with the same number of channels x1 and W x2 , adjust the feature map W after the channel x1 and W x2 Add together to get the basic path feature bp, and further extract and enhance the features through a 3×3 convolution layer to adjust the feature map W after the channel x1 and W x2 Apply the local attention mechanism respectively to obtain the weighted local feature representation; x1 and W x2 The global attention mechanism is applied to concatenate the feature maps obtained by the local and global attention branches according to the channel dimension to obtain two sets of fused feature maps x1 and x2. Then, x1, x2 and the basic path feature bp are concatenated according to the channel dimension to form a feature map containing rich feature information. A 1×1 convolutional layer is used to perform feature compression and channel number adjustment. The RepConv module is applied to further extract and enhance the fused feature map. A 1×1 convolutional layer is used to perform the final feature transformation and channel number adjustment on the output of the RepConv module to obtain an output feature map of size H×W×C.

[0035] Beneficial effects:

[0036] 1. The present invention designs a module called FasterCGLU that combines CGLU to improve computational efficiency and feature expression capabilities, and a module (MANet) that enhances feature extraction capabilities through multi-scale feature fusion and attention mechanism to form MANet_FasterCGLU. While ensuring feature extraction accuracy, it effectively reduces the computational complexity of the model and significantly improves the speed of bridge defect detection, allowing for more timely responses and the provision of bridge maintenance solutions.

[0037] 2. In terms of attention mechanism, this paper replaces the C2PSA module with the DTAB (DualAttentionBlock) module. The DTAB module combines channel attention and spatial attention mechanisms, using DilatedG-CSA and DilatedM-WSA to enhance features in the spatial and channel dimensions, respectively. This improves the model's focus on and representation of target features, enabling it to address bridge defect detection scenarios in various environments and complex conditions.

[0038] 3. This paper designs a new multi-scale feature extraction and fusion pyramid module UGMB, which further improves the model's detection performance for multi-scale targets by optimizing the structure of the feature pyramid and the feature fusion method. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a flow chart of the overall improvement scheme of the present invention;

[0040] Figure 2 This is a schematic diagram of the improved bridge defect detection model;

[0041] Figure 3 Comparative structural diagram of MANet_FasterCGLU and MANet designed for the present invention;

[0042] Figure 4 This is the structural diagram of FasterCGLU;

[0043] Figure 5 This is the structural diagram of FPSharedConv;

[0044] Figure 6 is the structural diagram of DTAB;

[0045] Figure 7 This is the UGMB pyramid structure diagram;

[0046] Figure 8 This is the WFU structure diagram in the UGMB pyramid;

[0047] Figure 9 The structural diagram of HAFB in the UGMB pyramid

[0048] Figure 10 This is a diagram of the model experiment results of the present invention; DETAILED DESCRIPTION

[0049] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0050] like Figure 1 As shown, the present invention discloses a bridge defect detection method based on AMSF-Pyramid-YOLOv11n network, and the specific implementation steps are as follows:

[0051] S1 builds a bridge defect detection dataset;

[0052] This example extracts some images from the dacl10k-toolkit, CRACK500, GAPs384, cracktree200, and dacl1k datasets, and combines them with bridge defect images obtained from the network to synthesize a new bridge defect detection dataset. The dataset described in the present invention contains 6306 color images with a unified image resolution of 2048×1536, covering five typical defects on the surfaces of bridge columns and road surfaces, specifically including corrosion, cracks, degraded concrete, concrete voids, and moisture; the dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1. The training set contains 5046 images, the validation set contains 630 images, and the test set contains 630 images.

[0053] S2 Figure 2 As shown in the figure, YOLOv11n is used as the baseline model to improve the model network;

[0054] S2.1 introduces the Faster_CGLU structure to replace the ConvNeck in the MANet structure to form the MANet_FasterCGLU structure. The comparison structure diagram of MANet_FasterCGLU and MANet is shown in the figure below. Figure 3 shown.

[0055] The FasterCGLU module combines convolution operations and a gating mechanism for feature extraction and transformation. This allows it to perform nonlinear transformations on input features and extract complex relationships between them. The convolution operation, through partial convolution, effectively captures spatial information in feature maps while reducing computational effort, enabling the model to better understand the spatial structure of the image. The gating mechanism allows the model to adaptively select important feature channels and suppress unimportant features, thereby enhancing feature representation. Regarding channel alignment, if the number of input channels (inc) does not equal the module's dimension (dim), a 1×1 convolutional layer is used for channel alignment. This allows the module to flexibly adapt to feature maps with varying numbers of input channels, enhancing its versatility and compatibility. An adaptive layer scaling mechanism is introduced. The output of the Convolutional GLU (CGLU) is scaled using learnable parameters before being added to the residual connection. This mechanism helps adjust the output amplitude of each layer, thereby improving model training stability.

[0056] The structure of FasterCGLU is shown in the figure Figure 4 As shown:

[0057] The FasterCGLU module first determines whether the number of input channels inc is consistent with the number of output channels dim. If they are inconsistent, 1x1 convolution is used to adjust the number of channels. For input feature maps of size (N, Cin ,H,W), where N is the batch size, C in is the number of input channels, H and W are the height and width of the feature map. After 1x1 convolution, the size of the output feature map is (N,C out ,H,W), where Cout=dim. The mathematical expression of this layer is: X adjusted =Conv 1×1 (X); Secondly, spatial feature mixing is performed through partial convolution (PartialConv). The input is X adjusted , output features Figure X mixed The size of is still (N, dim, H, W). The mathematical expression of partial convolution is: mixed =PartialConv(X adjusted ); then pass through a multi-layer perceptron (MLP) combined with a gated linear unit (GLU). The input is X mixed , output features Figure X mlp The size is still (N, dim, H, W). The mathematical expression of this layer is: mlp =GLU(X mixed ); finally, through the residual connection layer and the droppath layer; the residual connection layer will process the features Figure X mlp With the original input features Figure X adjusted Add together to form a residual connection. Output features Figure X residual The size is still (N, dim, H, W). The mathematical expression of the residual connection layer is shown in formula (4); the droppath layer is used to drop the path to prevent overfitting. The input is X residual , output features Figure X output The size is still (N, dim, H, W). The mathematical expression of the droppath layer is X output =DropPath(X residual ).

[0058] S2.2 uses the FPSharedConv module to perform feature extraction using shared convolution kernels at different expansion rates instead of the traditional SPPF module;

[0059] The FPSharedConv (FeaturePyramidSharedConv) module is used to replace the traditional SPPF module. This module realizes multi-scale feature extraction of input features by using shared convolution kernels and convolution operations with different dilation rates. Compared with the traditional SPPF module, it reduces the number of parameters and computational complexity while maintaining a strong feature extraction capability, thus improving the performance and efficiency of the model. The core idea of ​​FPSharedConv is to capture multi-scale features through shared convolution and dilated convolution. The structural diagram of FPSharedConv is shown in the figure below. Figure 5 As shown:

[0060] The convolution layer is the core module for extracting features in the neural network. in ,H,W), first, use 1x1 convolution to reduce the number of input channels to C in / / 2, output features Figure X The size of 1 is (N,C in / / 2,H,W); the mathematical expression of 1x1 convolution is X1=Conv 1×1 (X). Secondly, multiple dilated convolutions are performed on X1, each time using the same shared convolution kernel but a different dilation rate d. The mathematical expression of dilated convolution is:

[0061]

[0062] Among them, X d It indicates the result of a certain position in the output feature map after the dilation convolution operation with the dilation rate d; X i+d·k,j+d·l Represents the value of the element with coordinates (i+d·k,j+d·l) in the input feature map, where i and j are the base coordinates of the input feature map, d is the dilation rate, and k and l are the row and column indices in the convolution kernel; W k,l represents the weight value at coordinate (k, l) in the shared convolution kernel, where k and l are the number of rows and columns of the convolution kernel, respectively, i.e. the convolution kernel size is K×L; b represents the bias term;

[0063] Output features Figure X d The size is still (N,C in ,H,W). All the feature maps after dilation convolution are spliced ​​according to the channel dimension to obtain a size of (N,C in / / 2×(1+len(dilations)),H,W) features Figure X concat Finally, a 1×1 convolution is used to adjust the number of channels of the concatenated feature map to the number of output channels C. out , and obtain the final output features Figure X output , size (N,Cout ,H,W), this design reduces the number of parameters and computational complexity by sharing convolution kernels and convolution operations with different expansion rates, while enhancing the model's ability to capture multi-scale features.

[0064] S2.3 replaces the traditional C2PSA module with the DTAB module that combines channel attention and spatial attention mechanisms. It also designs a multi-scale feature extraction and fusion pyramid module UGMB by optimizing the feature pyramid structure and feature fusion method, replacing the baseline model neck structure.

[0065] The DTAB module structure is as follows Figure 6 As shown in the figure, the module mainly consists of three functional areas: DilatedG-CSA (Grouped Channel-wise Spatial Attention), DilatedFFN (Feed-Forward Network) and DilatedM-WSA (Multi-Window Spatial Attention).

[0066] DilatedG-CSA (GroupedChannel-wiseSpatialAttention): This part first performs the input feature Figure X Perform normalization and get X norm Next, three parallel dilated depthwise separable convolution operations are performed on X norm After processing, three feature maps F1, F2, and F3 are obtained. These three feature maps are spliced ​​in the channel dimension to obtain F concat , and then perform feature interaction between channels with the input feature X through element-by-element multiplication operation, and finally obtain the enhanced feature map F through residual connection gcsa :

[0067] F gcsa =X+F concat ⊙X

[0068] where ⊙ represents element-wise multiplication.

[0069] DilatedFFN (Feed-Forward Network): Feature map F after G-CSA processing gcsa Enter this section. First, gcsa Perform normalization and get F norm . Then a dilated depthwise separable convolution operation is performed to further extract the feature F conv1 , then apply the GeLU activation function to get F gelu . Perform the dilated depth-wise separable convolution again to get F conv2This series of operations aims to improve the discriminability and robustness of features through nonlinear transformation and feature reorganization:

[0070] F ffn =F conv2 +F gcsa

[0071] For the feature map F ffn Perform normalization and get F norm , the channel dimension is reduced by three 1×1 convolution operations, F conv1 ,F conv2 ,F conv3 The feature map after dimensionality reduction is divided into windows in the spatial dimension. The features in each window interact through the attention mechanism. The attention mechanism dynamically adjusts the feature weight by calculating the correlation between the query Q, key K, and value V to obtain the attention score A:

[0072]

[0073] Apply it to the value V to get the feature interaction result F within the window attn :

[0074] F attn =A·V(4)

[0075] Among them, dk is the dimension of the key. Finally, the information in the window is integrated through feature aggregation operation to complete the enhancement of spatial features.

[0076] The feature pyramid module UGMB designed in the present invention is an innovative multi-scale feature extraction and fusion structure, which aims to improve the target detection model's ability to capture targets of different scales. Compared with the traditional YOLOv11 feature pyramid, in terms of feature extraction and upsampling, the feature maps of different levels are upsampled and enhanced by applying the WFU module multiple times, and the low-resolution feature maps are converted into high-resolution feature representations, while fusing multi-scale information. In terms of feature enhancement and conversion, the MANet_FasterCGLU module is used to further extract and enhance the feature maps, and the information flow is dynamically controlled by partial convolution and gated linear units (GLU) to improve the feature expression ability. In terms of feature fusion, feature maps from different sources and different scales are spliced ​​through multiple Concat operations to integrate multi-scale information and enrich feature representation. In terms of feature integration and output, the fused feature maps are finally subjected to hierarchical attention fusion through the HAFB module to highlight important feature areas, suppress irrelevant feature information, and improve the model's adaptability to complex scenes and the accuracy of target detection. More efficient and more refined feature extraction and fusion are achieved. The structural diagram of the multi-scale feature extraction and fusion feature pyramid module UGMB is shown in Figure 2. Figure 7 shown.

[0077] The WFU (Wavelet Feature Upsampling) module is a feature upsampling and fusion module based on wavelet transform. It aims to improve the model's detection ability for targets of different scales through multi-scale feature extraction and fusion. This module decomposes the large feature map into sub-bands of different frequencies through wavelet transform, combines the small feature map for feature enhancement, and finally reconstructs the feature map through inverse wavelet transform. The WFU structure diagram is shown in the figure. Figure 8 As shown:

[0078] The WFU module receives two sets of input feature maps; Figure X big Size: H×W×C big and small features Figure X small Size is H′×W′×C small First, the HaarWavelet module is used to Figure X big Perform wavelet transform and decompose it into four sub-bands: approximate component A, horizontal detail component H, vertical detail component V, and diagonal detail component D. These four sub-bands contain low-frequency information and high-frequency information in different directions, respectively, and can characterize the characteristics of the input feature map from multiple dimensions.

[0079] After decomposition, the horizontal, vertical and diagonal detail components (H, V, D) are added together to obtain the sum of the detail components H+V+D, and then feature enhancement is performed through the residual block (RB). The residual block contains two 3×3 convolutional layers and a residual connection, which can effectively improve the feature expression ability and learning efficiency. At the same time, the small features are Figure X small The sum of the approximate components A is concatenated according to the channel dimension, and then the channel transformation module performs feature fusion and channel number adjustment. The channel transformation module contains two 1×1 convolutional layers, which can flexibly adjust the number of channels and enhance the expressiveness of features.

[0080] Finally, the enhanced detail components and the transformed approximate components are spliced ​​according to the channel dimension and the inverse wavelet transform is used to transform the

[0081] Reconstruct the feature map. The inverse wavelet transform reassembles subbands of different frequencies into a complete feature map, restoring its original size and structure. The entire WFU module effectively integrates and enhances features of different scales through wavelet and inverse wavelet transforms, improving the model's ability to detect multi-scale objects.

[0082] The WFU module plays a key role in the new feature pyramid. By fusing multi-scale features, it improves the model's detection capabilities for objects of varying scales, enhances feature representation, optimizes computational efficiency, and strengthens model robustness. These features give the WFU module a significant advantage in high-precision object detection tasks, providing the model with powerful feature representation capabilities.

[0083] HAFB: HAFB (Hierarchical Attention Fusion Block) aims to achieve effective fusion and enhancement of feature maps from different sources through a hierarchical attention mechanism and a multi-branch structure. The HAFB module has significantly improved multi-scale feature fusion capabilities, feature expression capabilities, computational efficiency, and model robustness, providing powerful feature representation capabilities for high-precision target detection tasks. Compared with traditional feature fusion methods, the HAFB module can more finely control the feature fusion process, highlight important feature areas, and suppress irrelevant feature information, thereby improving the model's adaptability to complex scenes and the accuracy of target detection. The HAFB structure diagram is shown below. Figure 9 As shown:

[0084] The HAFB module receives two sets of input feature maps, denoted as X1 (size is H×W×C1) and X2 (size is H×W×C1). These two sets of feature maps usually come from feature extraction modules at different levels, with different number of channels and semantic information. First, two 1×1 convolutional layers (W x1 and W x2 ) Adjust the number of channels of X1 and X2 respectively and convert them into feature maps W with the same number of channels (hidc=ouc / / 2) x1 and W x2 The purpose of this step is to prepare for the subsequent feature fusion operation and ensure that the feature maps from different sources can be effectively spliced ​​and fused in the channel dimension. x1 and W x2 Add them together to get the basic path feature bp. Then, a 3×3 convolution layer is used to further extract and enhance the features, providing basic feature information for subsequent feature fusion. x1 and W x2The local attention mechanism (LocalGlobalAttention, parameter p=2) is applied respectively. This mechanism divides the feature map into local blocks of size p×p and calculates the attention score in each block to highlight the key features in the local area. Specifically, the feature map is first unfolded to divide it into multiple local blocks; then the features in each block are averaged and pooled to obtain block-level feature representation; then the attention score is calculated through the multi-layer perceptron (MLP) and softmax function; finally, the attention score is multiplied with the original block-level feature to obtain the weighted local feature representation. Similarly, W x1 and W x2 A global attention mechanism (LocalGlobalAttention, parameter p=4) is applied. Unlike the local attention mechanism, the global attention mechanism uses a larger block size (p=4), which allows it to capture a wider range of contextual information and global feature dependencies. The feature maps obtained by the local and global attention branches are then concatenated along the channel dimension to obtain two fused feature maps x1 and x2. This step integrates feature information of different scales and different semantics, enriching the feature representation. x1, x2, and the base path features bp are then concatenated along the channel dimension to form a feature map containing rich feature information. A 1×1 convolutional layer is then used to perform feature compression and adjust the number of channels, reducing computational effort while maintaining feature expressiveness. The RepConv module is then applied to the fused feature map for further feature extraction and enhancement. The RepConv module achieves multi-scale feature extraction and efficient parameter learning by combining 3×3 and 1×1 convolutions, along with optional batch normalization. During training, the RepConv module can learn more complex feature representations; during deployment, it can be fused into an equivalent convolutional layer to improve computational efficiency. Finally, a 1×1 convolutional layer performs the final feature transformation and channel number adjustment on the output of the RepConv module, resulting in an output feature map of size H×W×C.

[0085] S3 obtains the improved final model, sets training hyperparameters, and uses the training set to train and verify the improved model;

[0086] S4 deploys the final model to implement bridge defect detection.

[0087] In this example, training was performed on a laboratory host computer using Python 3.8, CUDA 11.3, an NVIDIA RTX 4090 GPU, and 24GB of memory. The training hyperparameters were set as follows: 400 training rounds and a batch size of -1. The validation data showed that the model achieved a 12.4% improvement in precision, a 7.9% improvement in recall, and a 10.5% improvement in mAP@0.5 compared to YOLOv11n.

[0088] The comparative experiments are shown in Table 1:

[0089] Table 1

[0090]

[0091]

[0092] Any experiment in deep learning is random. In order to improve the credibility of the experimental results, we will compare the improved model with other models and use evaluation indicators including Precision, Recall, mAP, Parameters, and Gflops. In this invention, mAP@0.5 and mAP@0.5:0.95 (the threshold range of IoU is from 0.5 to 0.95) are used as accuracy measurement indicators. mAP is the average accuracy of the detection results of each type. AP refers to the area of ​​the curve enclosed by the horizontal and vertical axes with precision and recall. The calculation formulas for Precision and Recall are as follows:

[0093]

[0094] Where TP is the number of correctly identified targets. Generally, when the IoU threshold is greater than or equal to 0.5, it is considered to be a correctly identified target; FP is the number of incorrectly identified targets; FN is the number of missed targets; Precision is the proportion of correct targets among the targets detected by the model; Recall is the proportion of targets correctly identified by the model among the total number of real targets.

[0095] The number of detected categories in this example is 5, so the mAP is as follows:

[0096]

[0097] Through comparative analysis of experimental data, the proposed model shows significant performance advantages for bridge defect detection datasets including corrosion, cracks, degraded concrete, concrete voids, moisture, etc. Figure 10As shown in the figure, its mAP50 index continues to outperform the yolov11n model throughout the entire training cycle, and converges to a higher level in the later stage (after 200 epochs), indicating that the model has better comprehensive detection capabilities for various bridge defects such as corrosion and cracks, and can accurately identify and locate different types of defects; Figure 10 The Precision metric of the proposed model is significantly higher than that of the original model during training, and it maintains a higher level of accuracy in the later stages, effectively reducing the false positive rate for defects such as corrosion and degraded concrete, and improving the accuracy of various defect determinations. In summary, when applied to bridge defect detection scenarios, the proposed model, through optimized design, surpasses the original model in core metrics such as mAP50 and Precision, achieving more accurate detection and identification, as well as more reliable classification and determination of various bridge defects such as corrosion, cracks, degraded concrete, concrete voids, and moisture, providing more efficient and stable technical support for bridge structural health monitoring and safety assessment.

[0098] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made in accordance with the spirit of the present invention shall be included in the scope of protection of the present invention.

Claims

1. A bridge defect detection method based on UGMB multi-scale feature extraction and fusion, characterized in that: The method comprises the following steps: S1 builds a bridge defect detection dataset; S2 uses YOLOv11n as a baseline model to improve and build a bridge defect detection model; S2.1 introduces the FasterCGLU structure to replace the ConvNeck in the MANet structure to form the MANet_FasterCGLU structure; S2.2 uses the FPSharedConv module to perform feature extraction using shared convolution kernels at different expansion rates instead of the traditional SPPF module; S2.3 replaces the traditional C2PSA module with the DTAB module that combines channel attention and spatial attention mechanisms. It also designs a multi-scale feature extraction and fusion pyramid module UGMB by optimizing the feature pyramid structure and feature fusion method, replacing the baseline model neck structure. S3 obtains the final model, sets training hyperparameters, and uses the training set to train and verify the improved model; S4 deploys the final model to implement bridge defect detection.

2. The bridge defect detection method based on UGMB multi-scale feature extraction and fusion according to claim 1 is characterized in that: The dataset described in S1 includes corrosion, cracks, degraded concrete, concrete voids, wet bridge column surfaces, and pavement surface defects. The dataset is divided into training set, validation set, and test set in a ratio of 8:1:

1.

3. The bridge defect detection method based on UGMB multi-scale feature extraction and fusion according to claim 2 is characterized in that: The FasterCGLU structure described in S2.1 is implemented as follows: FasterCGLU first determines whether the number of input channels inc is consistent with the number of output channels dim. If not, 1x1 convolution is used to adjust the number of channels. For input feature maps of size (N, C in ,H,W), where N is the batch size, C in is the number of input channels, H and W are the height and width of the feature map. After 1x1 convolution, the size of the output feature map is N×C out ×H×W, where Cout=dim, the mathematical expression of this layer is: adjusted =Conv 1×1 (X); spatial feature mixing is performed through partial convolution, and the input is X adjusted , output feature map X mixed The size is still N×dim×H×W, and the mathematical expression of the partial convolution is: mixed =PartialConv(X adjusted ); By combining the multi-layer perceptron with gated linear units, the input is X mixed , output feature map X mlp The size is still N×dim×H×W, and the mathematical expression of this layer is: mlp =GLU(X mixed ); Finally, through the residual connection layer and the droppath layer; the residual connection layer will process the feature map X mlp With the original input feature map X adjusted Add together to form a residual connection and output feature map X residual The size is still (N, dim, H, W), the droppath layer is used for path discarding to prevent overfitting, and the input is X residual , output feature map X output The size is still (N, dim, H, W), and the mathematical expression of the droppath layer is X output =DropPath(X residual ).

4. The bridge defect detection method based on UGMB multi-scale feature extraction and fusion according to claim 3 is characterized in that: The FPSharedConv module structure described in S2.2 is implemented as follows: For an input feature map of size N×C in ×H×W, using 1x1 convolution to reduce the number of input channels to C in / / 2, the size of the output feature map X1 is N×C in / / 2×H×W, the mathematical expression of 1x1 convolution is X1=Conv 1×1 (X); perform multiple dilated convolutions on X1, each using the same shared convolution kernel but a different dilation rate d. The mathematical expression of dilated convolution is: Among them, X d It indicates the result of a certain position in the output feature map after the dilation convolution operation with the dilation rate d; X i+d·k,j+d·l Represents the value of the element with coordinates (i+d·k,j+d·l) in the input feature map, where i and j are the base coordinates of the input feature map, d is the dilation rate, and k and l are the row and column indices in the convolution kernel; W k,l represents the weight value at coordinate (k, l) in the shared convolution kernel, where k and l are the number of rows and columns of the convolution kernel, respectively, i.e. the convolution kernel size is K×L; b represents the bias term; Output feature map X d The size is still N×C in ×H×W; all the feature maps after the expansion convolution are spliced ​​according to the channel dimension to obtain a size of N×C in / / 2×(1+len(dilations))×H×W feature map X concat ; Use 1×1 convolution to adjust the number of channels of the concatenated feature map to the number of output channels C out , and get the final output feature map X output , size N×C out ×H×W.

5. The bridge defect detection method based on UGMB multi-scale feature extraction and fusion according to claim 1 is characterized in that: The DTAB module structure described in S2.3 consists of three functional regions: DilatedG-CSA, DilatedFFN and DilatedM-WSA. The three functional regions are connected in series and perform feature transfer.

6. The bridge defect detection method based on UGMB multi-scale feature extraction and fusion according to claim 5 is characterized in that: The DilatedG-CSA part: normalizes the input feature map X to obtain X norm , using three parallel dilated depthwise separable convolution operations on X norm After processing, three feature maps F1, F2, and F3 are obtained. The feature maps are spliced ​​in the channel dimension to obtain F concat , through the element-by-element multiplication operation and the input feature X, the feature interaction between channels is performed, and the enhanced feature map F is obtained through the residual connection gcsa : The DilatedFFN part: Feature map F after DilatedG-CSA processing gcsa Enter this section and click F gcsa Perform normalization and get F norm , further extract the feature F through an expansion depth separable convolution operation conv1 , apply the GeLU activation function to get F gelu , perform dilated depth-separable convolution to obtain F conv2 ; The DilatedM-WSA part: the characteristic graph F ffn Perform normalization and get F norm , the channel dimension is reduced by three 1×1 convolution operations, F conv1 ,F conv2 ,F conv3 The feature map after dimensionality reduction is divided into windows in the spatial dimension. The features in each window interact through the attention mechanism. The attention mechanism dynamically adjusts the feature weight by calculating the correlation between the query Q, key K, and value V to obtain the attention score A: Apply it to the value V to get the feature interaction result F within the window attn : F attn =A·V(4) Among them, dk is the dimension of the key. Finally, the information in the window is integrated through feature aggregation operation to complete the enhancement of spatial features.

7. The bridge defect detection method based on UGMB multi-scale feature extraction and fusion according to claim 6 is characterized in that: The UGMB module structure described in S2.3 is implemented as follows: In terms of feature extraction and upsampling, the WFU module is applied to upsample and enhance feature maps at different levels, convert low-resolution feature maps into high-resolution feature representations, and fuse multi-scale information at the same time; the MANet_FasterCGLU module is used to further extract and enhance the feature maps; the information flow is dynamically controlled through partial convolution and gated linear units GLU; through multiple Concat operations, feature maps from different sources and different scales are spliced ​​to integrate multi-scale information, and finally the fused feature maps are subjected to hierarchical attention fusion through the HAFB module to highlight important feature areas and suppress irrelevant feature information.

8. The bridge defect detection method based on UGMB multi-scale feature extraction and fusion according to claim 7 is characterized in that: The WFU module structure is implemented as follows: The WFU module receives two sets of input feature maps; the large feature map X big Size: H×W×C big and small feature map X small Size is H ′ ×W ′ ×C small , through the HaarWavelet module to the large feature map X big Perform wavelet transform and decompose it into four sub-bands: approximate component A, horizontal detail component H, vertical detail component V and diagonal detail component D. After decomposition, add the horizontal, vertical and diagonal detail components to get the sum of detail components H+V+D, and then perform feature enhancement through residual block, which contains two 3×3 convolutional layers and a residual connection. At the same time, the small feature map X small The enhanced detail component and the transformed approximate component are spliced ​​according to the channel dimension, and the feature map is reconstructed by inverse wavelet transform.

9. The bridge defect detection method based on UGMB multi-scale feature extraction and fusion according to claim 7 is characterized in that: The HAFB module structure is implemented as follows: The HAFB module receives two sets of input feature maps, denoted as X1: size H×W×C1 and X2: size H×W×C1, through two 1×1 convolutional layers W x1 and W x2 Adjust the number of channels of X1 and X2 respectively and convert them into feature maps W with the same number of channels x1 and W x2 , adjust the feature map W after the channel x1 and W x2 Add together to get the basic path feature bp, and further extract and enhance the features through a 3×3 convolution layer to adjust the feature map W after the channel x1 and W x2 Apply the local attention mechanism respectively to obtain the weighted local feature representation; x1 and W x2 The global attention mechanism is applied to concatenate the feature maps obtained by the local and global attention branches according to the channel dimension to obtain two sets of fused feature maps x1 and x2. Then, x1, x2 and the basic path feature bp are concatenated according to the channel dimension to form a feature map containing rich feature information. A 1×1 convolutional layer is used to perform feature compression and channel number adjustment. The RepConv module is applied to further extract and enhance the fused feature map. A 1×1 convolutional layer is used to perform the final feature transformation and channel number adjustment on the output of the RepConv module to obtain an output feature map of size H×W×C.