A bridge inspection-oriented unmanned aerial vehicle image apparent disease intelligent identification method

CN122551213APending Publication Date: 2026-08-11NANNING UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-09
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,桥梁病害具有尺度差异大、形态复杂、背景干扰多、小目标占比高等特点,现根据YOLOv12的算法进行改进,解决了光照不足、图像模糊、病害尺度极小或纹理相似的情况下,识别性能显著下降的问题

Benefits of technology

[0043]本发明为一种面向桥梁巡检的无人机影像表观病害智能识别方法,通过StarNet主干特征提取网络与SegNext多尺度卷积注意力模块,提高捕捉桥梁病害的多尺度特征与上下文信息,同时,使用损失函数NWD模块,可以进行小目标检测,提高对于桥梁病害检测的准确性,也可以精准识别不同的桥梁病害类型。因此,本发明的方法能够综合利用图像增强、多尺度特征提取、注意力机制与小目标优化等技术,实现高效、准确、自动化的桥梁表现病害检测与分类,为桥梁维护与安全管理提供可靠的技术支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551213A_ABST
    Figure CN122551213A_ABST
Patent Text Reader

Abstract

The application discloses a kind of unmanned aerial vehicle image apparent disease intelligent identification methods for bridge inspection, and the high-definition image data of bridge disease is collected by unmanned aerial vehicle inspection system, and bridge disease dataset is constructed;The data set is preprocessed;Bridge disease identification model based on YOLOv12 architecture is constructed;Disease identification model is improved in combination with SegNext attention module and loss function NWD module;The improved model is trained using the preprocessed data set;The bridge inspection image to be identified is input into the trained disease identification model, and the class label and the confidence boundary box of the bridge disease image are output by the model, to realize the intelligent identification of bridge disease.The method of the application can realize efficient, accurate, automated bridge performance disease detection and classification, and provide reliable technical support for bridge maintenance and safety management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of UAV image appearance defect recognition technology, specifically a method for intelligent recognition of bridge defects. Background Technology

[0002] Drone technology, with its advantages of flexibility, efficiency, and low cost, is widely used in bridge inspection. Drones equipped with high-definition cameras capture images of bridge defects from multiple angles and close range, effectively compensating for the shortcomings of traditional inspections, improving the safety of bridge inspections, and facilitating timely repairs. However, many drone scans rely heavily on manual interpretation, which is time-consuming and labor-intensive. Furthermore, the images are easily affected by weather and other factors, leading to errors in interpretation by staff and making it difficult to guarantee the consistency and accuracy of the identification results.

[0003] With the advancement of deep learning technology, automated methods for identifying defects in basic images have gradually become a research hotspot. In particular, target detection algorithms have shown great potential in identifying defects such as cracks, decay, and spalling on bridges. However, bridge defects are characterized by large scale differences, complex shapes, numerous background interferences, and a high proportion of small targets. This paper presents an improvement on the YOLOv12 algorithm, which solves the problem of significantly reduced recognition performance under conditions of insufficient lighting, image blur, extremely small defect scale, or similar textures.

[0004] Furthermore, existing methods mostly focus on identifying single types of diseases, lacking a unified and lightweight identification framework for multiple types of diseases. Meanwhile, in model design, effectively integrating multi-scale features, enhancing small target detection capabilities, and optimizing loss functions to improve localization accuracy remain challenges in current technology.

[0005] Therefore, there is an urgent need for an intelligent UAV identification method for bridge defect inspection scenarios, which can comprehensively utilize technologies such as image enhancement, multi-scale feature extraction, attention mechanisms and small target optimization to achieve efficient, accurate and automated detection and classification of bridge defects, providing reliable technical support for bridge maintenance and safety management. Summary of the Invention

[0006] To achieve the above objectives, the present invention provides an intelligent identification method for apparent defects in bridge inspection using UAV images.

[0007] The following technical solution was adopted in this method:

[0008] S1. Collect high-definition image data of bridge defects and construct a bridge defect dataset. The bridge defect types include four types: corrosion, cracks, substandard concrete, and road surface deterioration.

[0009] S2. Preprocess the dataset by using Wiener filtering to perform inverse filtering on the original image data to remove blurring, and use histogram equalization to make the gray-level histogram of the original image data evenly distributed across the entire gray-level range to obtain the preprocessed image.

[0010] S3. Construct a bridge defect identification model based on the YOLOv12 architecture, improve the backbone network model, use the StarNet module as the backbone feature extraction network of the model, use the modified demo block to extract features, and perform element-wise multiplication, which includes depthwise separable convolution. At the same time, the channel factor of the network width is fixed at 4. This backbone model can significantly reduce the number of parameters and computation, and the unique feature expression method can effectively improve the ability to capture small cracks.

[0011] S4. Based on the improved backbone network model, combined with the SegNext attention module, a pyramid structure is adopted, and a new multi-scale convolutional attention module is formed by drawing on the ViT structure. This module is used to capture multi-scale contextual information of the image. The output is directly used as the attention weight to re-weight the input of the multi-scale convolutional attention module. Features are extracted through multiple convolutional kernels of different scales to effectively distinguish the bridge piers, water areas, and vegetation interference in the bridge background of the image after preprocessing in step S2.

[0012] S5. In the improved bridge defect identification model, the loss function NWD module is introduced for small target detection of bridge defects. The boundary of the constructed bridge defect prediction box is modeled as a two-dimensional Gaussian distribution. The Wasserstein distance between the prediction box and the real box is measured to optimize the localization of early cracks with slender shapes and a pixel ratio less than a set value.

[0013] S6. After labeling the preprocessed image, input it into the bridge defect identification model improved by steps S3, S4 and S5 for training to obtain the trained defect identification model.

[0014] S7. Input the bridge inspection image to be identified into the trained defect identification model. The model outputs the category label and confidence bounding box of the bridge defect image to achieve intelligent identification of bridge defects.

[0015] Furthermore, in step S1, a professional drone is used to take photos of bridge defects. For defects such as substandard concrete on bridge piers and towers and road surface cracking, the drone is positioned at a distance of 5-15 meters and tilted at an angle of 45°-75° to ensure that the defects are clearly displayed in the images. For localized defects such as corrosion and cracks, the drone is positioned at a distance of 3-8 meters and tilted at a small angle, i.e., 0°-30°, to take photos.

[0016] Furthermore, Wiener filtering can achieve a balance between deblurring and noise suppression, and its performance in the frequency domain is expressed by the following formula:

[0017]

[0018] in, It is the transfer function of the Wiener filter that needs to be solved; It is the Fourier transform of the degenerate function. yes The spectrum is equal to , yes The complex conjugate; It is the power spectrum of the noise; It is the power spectrum of the image data acquired in step S1; It is the noise-to-signal power ratio.

[0019] Furthermore, the gray-level histogram of the Wiener-filtered image data is calculated. The gray-level histogram is calculated based on the probability of each gray level appearing in the image, using the following formula:

[0020]

[0021] in, It is the first grayscale value; It is grayscale. The number of pixels appearing in the image; It is the total number of pixels in the image; It is grayscale. The probability of occurrence, i.e., the normalized histogram; Indicates the current grayscale index; It represents the total number of grayscale values.

[0022] Calculate the cumulative distribution function That is, from gray level 0 to The cumulative probability is given by the formula:

[0023]

[0024] in, It is grayscale. The cumulative distribution function, which represents the gray value less than or equal to The sum of the probabilities of all pixels, which is monotonically increasing in the interval [0,1]. Indicates the current grayscale index; Represents the variable to be summed; It represents the total number of grayscale values.

[0025] Based on the cumulative distribution function, it can be mapped to the entire grayscale range [0, L-1], and the transformation formula is:

[0026]

[0027] in, It is the original gray level. The new gray level obtained after equalization; It is a transformation function; It rounds to the nearest integer because grayscale levels must be integers. It is the maximum gray value. Multiplying by this value is to map the probability value [0, L-1] to the actual gray range.

[0028] Furthermore, in step S3, the StarNet module backbone feature extraction network adopts the star operation algorithm. In the first layer of the star operation, the "star operation" can be written as: .

[0029] This represents the input feature matrix, used to receive the feature map from the previous layer; This represents the weight matrix, which performs a linear transformation on the input values. This indicates that the bias vector adds a bias to the linear transformation.

[0030] Let the input features be... The number of channels is Through a single star operation, a dimension is implicitly expressed. A high-dimensional feature space to achieve The implicit dimension feature space representation shows that each term has a non-linear relationship with the input; the implicit dimension formula for a single-layer star operation is:

[0031]

[0032] in, This represents the star operation, i.e., element-wise multiplication; This represents the weight vector, which represents the weights on the input. Weights for linear transformation; This represents the input feature vector. For a specific input element, it represents the feature values ​​of all channels, and its dimension is also 1. ; This represents the number of input channels and is the actual, explicit width of the network. The summation index is used to iterate through each element of the vector; Represents a scalar value at a specific position in a vector.

[0033] When stacking multiple layers, if the number of channels of the input features in the first layer is... ,go through Layer star operations, implicitly obtaining belonging to The feature space, Indicates the first The dimension of the implicit feature space corresponding to the layer output.

[0034] The formula for the exponential growth of multiple stacked layers is:

[0035]

[0036] in, Indicates the index of the current layer; Indicates the first The output of the star operation is a vector. This indicates the output of the previous layer, which is also the input of the current layer; Indicates the first The layer has two weight matrices, which can handle multiple channel outputs.

[0037] Furthermore, in step S4, the new multi-scale convolutional attention module comprises three parts: depthwise convolution for aggregating local information; and multi-branch depthwise separable convolution for capturing multi-scale contextual information through convolutional kernels of different sizes. Convolution is used to establish relationships between different channels of the model. The output is directly used as attention weights to reweight the input of the multi-scale convolutional attention module. Stacking multiple such building blocks forms a convolutional encoder.

[0038] The new multi-scale convolutional attention module adopts a hierarchical structure, consisting of four stages, each with progressively decreasing spatial resolution. Each stage contains a sampling block, which consists of a convolutional layer with a stride of 2 and a kernel size of 3x3, and a batch normalization layer.

[0039] here and These represent the height and width of the input image, respectively. Each stage contains a sampling block and a stacking block. The sampling block has a convolution with a stride of 2 and a kernel size of 3x3.

[0040] Furthermore, in step S6, for the images obtained after preprocessing in step S2, different bridge defects are labeled with different colors using image feature annotation software, and the corresponding defect names are written on them for differentiation. Simultaneously, the image data is classified, and the images are grouped into training, testing, and validation sets using a 7:2:1 allocation ratio.

[0041] Furthermore, in step S6, the features of the images are labeled based on the YOLOv12 improvements. Simultaneously, the labeled, grouped image data are trained by group. During model training, incorrect or flawed labels are adjusted to improve recognition quality and efficiency.

[0042] Beneficial effects:

[0043] This invention presents an intelligent method for identifying apparent defects in bridge images from unmanned aerial vehicle (UAV) images. It utilizes a StarNet backbone feature extraction network and a SegNext multi-scale convolutional attention module to improve the capture of multi-scale features and contextual information of bridge defects. Simultaneously, the use of a loss function (NWD) module enables small object detection, enhancing the accuracy of bridge defect detection and allowing for precise identification of different defect types. Therefore, this method comprehensively leverages image enhancement, multi-scale feature extraction, attention mechanisms, and small object optimization techniques to achieve efficient, accurate, and automated detection and classification of bridge defects, providing reliable technical support for bridge maintenance and safety management. Attached Figure Description

[0044] Figure 1 This is a schematic diagram of the process of the present invention;

[0045] Figure 2 This is a structural diagram of the StarNet backbone network module of the present invention;

[0046] Figure 3 This is a result diagram of the attention SegNeXt module of the present invention;

[0047] Figure 4 This is a flowchart of the image processing grayscale histogram and Wiener filtering process of the present invention;

[0048] Figure 5 This is a flowchart of the algorithm for the StarNet backbone network module of this invention;

[0049] Figure 6 The flowchart of the attention SegNeXt module of this invention is shown below.

[0050] Figure 7 The flowchart of the NWD loss function module of this invention is shown below;

[0051] Figure 8 This is a comparison diagram of the model before and after the improvement of the present invention. Detailed Implementation

[0052] Example 1:

[0053] This invention discloses an intelligent method for identifying apparent defects in bridge inspection images using unmanned aerial vehicle (UAV) images, combining... Figure 1As shown, it includes the following steps:

[0054] S1. High-definition image data of bridge defects are collected from different angles and distances using a drone inspection system to construct a bridge defect dataset. The types of bridge defects include four types: corrosion, cracks, substandard concrete, and road surface deterioration.

[0055] S2. Perform overall processing and enhancement on the dataset. Use Wiener filtering to perform inverse filtering to deblur the image. Use histogram equalization to make the gray-level histogram of the original image uniformly distributed across the entire gray-level range, thus obtaining the preprocessed image.

[0056] S3. Construct a bridge defect identification model based on the YOLOv12 architecture, improve the backbone network model, and use the StarNet module based on star operation as the backbone feature extraction network of the model. This module uses the modified demo block module to extract features. This backbone network model can significantly reduce the number of parameters and computation, and its unique feature expression method can effectively improve the ability to capture small cracks. It includes depthwise separable convolution, and the channel factor of the network width is fixed at 4. This backbone model can significantly reduce the number of parameters and computation, and its unique feature expression method can effectively improve the ability to capture small cracks.

[0057] The StarNet module modifies the original demo block module by replacing the GELU set function with ReLU6 to improve inference speed; replacing Layer Normalization with Batch Normalization, which can be fused with convolutional layers to reduce computational overhead; and adding a depthwise convolution at the end of the original demo block to enhance the module's ability to model local spatial information.

[0058] S4. Based on the improved backbone network model, combined with the attention module SegNext, a pyramid structure is adopted. Drawing on the ViT structure, a new multi-scale convolutional attention module MSCA is formed to capture multi-scale contextual information of the image. The output is directly used as the attention weight to reweight the input of the multi-scale convolutional attention module MSCA. In this module, the convolutional attention module MSCA extracts features through multiple convolutional kernels of different scales. It effectively distinguishes between defects and background, reducing the false detection rate, and addresses interference such as bridge piers, water areas, and vegetation that are often included in the bridge background.

[0059] S5. Introduce the loss function module NWD in the bridge defect identification model for small target detection. Model the boundary as a two-dimensional Gaussian distribution and optimize the localization of early cracks with slender shapes and extremely small pixel ratio by measuring the Wasserstein distance between the predicted box and the real box.

[0060] S6. Input the bridge inspection image to be identified into the trained bridge defect identification model. The model outputs the category label and confidence bounding box of the bridge defect image to realize intelligent identification of bridge defects.

[0061] In step S1, a professional drone is used to take photos of bridge defects. For defects such as substandard concrete on bridge piers and towers and road surface cracking, the drone is positioned at a distance of 5-15 meters and tilted at an angle of 45°-75° to make the defects clear in the images. For local defects such as corrosion and cracks, the drone is positioned at a distance of 3-8 meters and tilted at a small angle, i.e., 0°-30°, to take photos.

[0062] In step S2, when preprocessing the dataset, Wiener filtering can achieve a balance between deblurring and noise suppression. Its formula in the frequency domain is:

[0063]

[0064] in, It is the transfer function of the Wiener filter that needs to be solved; It is the Fourier transform of the degenerate function. yes The spectrum is equal to , yes The complex conjugate; It is the power spectrum of the noise; It is the power spectrum of the original undegraded image, i.e., the image data acquired in step S1; It is the noise-to-signal power ratio.

[0065] like Figure 4 As shown, after Wiener filtering for noise reduction, the gray-level histogram of the image is calculated. The gray-level histogram is calculated based on the probability of each gray level appearing in the image, using the following formula:

[0066]

[0067] in, It is the first grayscale value; It is grayscale. The number of pixels appearing in the image; It is the total number of pixels in the image; It is grayscale. The probability of something appearing in an image, i.e., the normalized histogram; Indicates the current grayscale index; It represents the total number of grayscale values.

[0068] Calculate the cumulative distribution function, i.e., from gray level 0 to... The cumulative probability is given by the formula:

[0069]

[0070] in, It is grayscale. The cumulative distribution function, which represents the gray value less than or equal to The sum of the probabilities of all pixels, which is monotonically increasing in the interval [0,1].

[0071] Based on the cumulative distribution function, it can be mapped to the entire grayscale range [0, L-1], and the transformation formula is:

[0072]

[0073] in, It is the original gray level. The new gray level obtained after equalization; It is a transformation function; It rounds to the nearest integer because grayscale levels must be integers. It is the maximum gray value. Multiplying by this value is to map the probability value [0, L-1] to the actual gray range.

[0074] In step S3, such as Figure 2 and Figure 5 As shown, in the first layer of a single-layer "star operation" neural network, the "star operation" can be written as: ,

[0075] This represents the input feature matrix, used to receive the feature map from the previous layer; This represents the weight matrix, which performs a linear transformation on the input values. This indicates that the bias vector adds a bias to the linear transformation.

[0076] Let the input features be... The number of channels is Then, through a single star operation, a dimension of approximately [dimensionality missing] is implicitly expressed. A high-dimensional feature space to achieve , The implicit dimension feature space representation, where each term has a non-linear relationship with the input, is given by the formula for the implicit dimension of a single-layer star operation:

[0077]

[0078] in, This represents the star operation, i.e., element-wise multiplication; The weight vector represents the weights applied to the input feature vector. Weights for linear transformation; This represents the input feature vector. For a specific input element, it represents the feature values ​​of all channels, and its dimension is also 1. ; This represents the number of input channels and is the actual, explicit width of the network. The summation index is used to iterate through each element of the vector; Represents a scalar value at a specific position in a vector.

[0079] The multi-layer "star operation" occurs when multiple layers are stacked, and the number of channels in the first layer's input features is... ,go through The "star operation" of a layer can implicitly obtain the data belonging to the layer. The exponential growth formula for the feature space of multiple stacked layers is:

[0080]

[0081] in, Indicates the index of the current layer; Indicates the first The output of the star operation is a vector. This represents the output of the previous layer's star operation, and also the input of the current layer; Indicates the first The layer has two weight matrices, which can handle multiple channel outputs; Indicates the first The dimension of the implicit feature space corresponding to the layer output.

[0082] In step S4, such as Figure 3 and Figure 6 As shown, the SegNext attention module uses a pyramid structure, drawing inspiration from the ViT structure, to form a new multi-scale convolutional attention module, MSCA. The multi-scale convolutional attention module MSCA consists of three parts: depthwise convolutions for aggregating local information; multi-branch depthwise separable convolutions that capture multi-scale contextual information using convolutional kernels of different sizes; Convolution is used to establish relationships between different channels in the model. The output is directly used as attention weights to reweight the input of the Multi-Scale Convolutional Attention Module (MSCA). Stacking multiple such building blocks forms a convolutional encoder.

[0083] The Multi-Scale Convolutional Attention Module (MSCA) employs a common hierarchical structure, comprising four stages, each with progressively decreasing spatial resolution. Each stage contains a sampling block, which consists of a convolutional layer with a stride of 2 and a kernel size of 3x3, and a batch normalization layer.

[0084] here and These represent the height and width of the input image, respectively. Each stage contains a downsampling block and a stacking block. The downsampling block has a convolution with a stride of 2 and a kernel size of 3x3.

[0085] Furthermore, in step S5, this module first models the bounding boxes as two-dimensional Gaussian distributions to measure their distribution differences using the Wasserstein distance. Let the two-dimensional Gaussian distributions corresponding to the two bounding boxes be respectively... and , Let represent the mean vectors of the Gaussian distribution, respectively. Let represent the covariance matrices of the Gaussian distribution, respectively.

[0086] For horizontal bounding box parameters Model it as a two-dimensional Gaussian distribution The formula is:

[0087]

[0088] in, Represents the center point of the bounding box coordinate; Indicates the width and height of the bounding box; This represents the mean vector of a Gaussian distribution. Let represent the covariance matrix of the Gaussian distribution.

[0089] like Figure 7 As shown, according to the above modeling formula, the true bounding box of the model is... and the predicted bounding box are The predicted distribution parameters are derived from the predicted bounding box output by the model. The true distribution parameters are the true bounding boxes labeled in the dataset. The formulas are as follows:

[0090]

[0091]

[0092] Calculate two Gaussian distributions and The squared second-order Wasserstein distance between them is given by the formula:

[0093]

[0094] in, , Indicates a Gaussian distribution; This represents the coordinates of the center point of the actual bounding box. Indicates width, Indicates altitude; This indicates the coordinates of the center point of the predicted bounding box. Indicates width, Indicates altitude; express Norm; This represents the square of the second-order Wasserstein distance.

[0095] The Wasserstein distance is normalized using the following formula:

[0096]

[0097] in, Represents the normalized Wasserstein distance, with a range of . ; This represents the normalization constant, which is related to the dataset. This represents an exponential function, ensuring the output is between 0 and 1.

[0098] The formula for calculating the U-loss function is:

[0099]

[0100] in, express Loss function; The Gaussian distribution representing the true bounding box; This represents the Gaussian distribution of the predicted bounding boxes.

[0101] Example 2:

[0102] This implementation case utilizes image feature annotation software to label different diseases with different colors and write the corresponding disease names for differentiation. Simultaneously, the image data is categorized, grouping images according to a 7:2:1 allocation ratio, as shown in Table 1.

[0103] It mainly targets four types of defects in data images: corrosion, cracks, substandard concrete, and road surface deterioration.

[0104] Table 1 Statistics on the number of targets detected

[0105] category test set training set Validation set total Corroded 4625 15962 1931 22518 cracks 596 2207 282 3085 Deteriorated concrete 1399 4819 644 6862 pavement degradationpavement degradation 4 70 3 77 all 6624 23058 2860 32542

[0106] Multiple object detection models were compared, and different evaluation metrics were used to measure network performance. The metrics used included precision, recall, mean precision (mAP50), and mAP50:90. Precision and recall were used to evaluate the model's accuracy and false negative rate in bridge defect detection. mAP50 and mAP50:90 represent the model's average precision at intersection-over-union (IoU) thresholds of 0.5 and 0.5 to 0.95, respectively, and were used to evaluate the model's performance on small sample sizes in bridge defect detection.

[0107] like Figure 8 The figure shows a comparison of the model before and after the improvement. The evaluation metric for the YOLOv12 algorithm before the improvement is shown as the YOLOv12 curve in the figure. After the improvement, StarNet, SegNext, and NWD modules are introduced into the YOLOv12 algorithm, and the evaluation metric is shown as the SSN-YOLO curve in the figure. The comparison shows that the improvement increases the algorithm's running time and improves the recognition accuracy, resulting in enhanced accuracy. The same data augmentation method is used to expand the training data, improving generalization ability. Overall, the improved model improves detection performance, robustness, and efficiency.

[0108] The above-mentioned feature annotation is based on YOLOv12 improvements. Simultaneously, group-based training is performed, and during model training, incorrect or flawed annotations are adjusted to improve recognition quality and efficiency.

Claims

1. A method for intelligent recognition of apparent diseases in UAV images for bridge inspection, characterized in that, Includes the following steps: S1. Collect high-definition image data of bridge defects and construct a bridge defect dataset. The bridge defect types include four types: corrosion, cracks, substandard concrete, and road surface deterioration. S2. Preprocess the dataset by performing Wiener filtering on the acquired image data to remove blur, and use histogram equalization to make the gray-level histogram of the Wiener-filtered image data uniformly distributed across the entire gray-level range, thus obtaining the preprocessed image. S3. Construct a bridge defect identification model based on the YOLOv12 architecture, improve the backbone network model, use the StarNet module as the backbone feature extraction network of the model, use the modified demo block to extract features, and perform element-wise multiplication. S4. Based on the improved backbone network model, combined with the SegNext attention module, a pyramid structure is adopted, and a new multi-scale convolutional attention module is formed by drawing on the ViT structure. This module is used to capture multi-scale contextual information of the image. The output is directly used as the attention weight to re-weight the input of the multi-scale convolutional attention module. Features are extracted through multiple convolutional kernels of different scales to effectively distinguish the bridge piers, water areas, and vegetation interference in the bridge background of the image after preprocessing in step S2. S5. In the improved bridge defect identification model, the loss function NWD module is introduced for bridge defect target detection. The boundary of the constructed bridge defect prediction box is modeled as a two-dimensional Gaussian distribution. The Wasserstein distance between the prediction box and the real box is measured to optimize the localization of early cracks with slender shapes and a pixel ratio less than a set value. S6. After labeling the preprocessed image, input it into the bridge defect identification model improved by steps S3, S4 and S5 for training to obtain the trained defect identification model. S7. Input the bridge inspection image to be identified into the trained defect identification model, and output the category label and confidence bounding box of the bridge defect image to realize intelligent identification of bridge defects. 2.The bridge inspection-oriented unmanned aerial vehicle image apparent disease intelligent identification method according to claim 1, characterized in that, In step S1, a drone is used to photograph bridge defects. For defects such as substandard concrete on bridge piers and towers and road surface cracking, the drone is used to take pictures at a distance of 5-15 meters and an angle of 45°-75°. For defects such as corrosion and cracks, the drone is used to take pictures at a distance of 3-8 meters and an angle of 0°-30°. 3.The bridge inspection oriented unmanned aerial vehicle image apparent disease intelligent identification method of claim 1, wherein, In step S2, Wiener filtering achieves a balance between deblurring and noise suppression, and the formula performed in the frequency domain is: in, It is the transfer function of the Wiener filter that needs to be solved; It is the Fourier transform of the degenerate function. yes The spectrum is equal to , yes The complex conjugate; It is the power spectrum of the noise; It is the power spectrum of the image data acquired in step S1; It is the noise-to-signal power ratio.

4. The unmanned aerial vehicle image apparent disease intelligent identification method for bridge inspection according to claim 3, characterized in that, The grayscale histogram of the Wiener-filtered image data is calculated as follows: The gray-level histogram is calculated by the probability of each gray level appearing in an image, using the following formula: in, It is the first grayscale value; It is grayscale. The number of pixels appearing in the image; It is the total number of pixels in the image; It is grayscale. The probability of occurrence, i.e., the normalized histogram; Indicates the current grayscale index; It is the total number of grayscale values; Calculate the cumulative distribution function , wherein is a cumulative distribution function of a gray scale representing a sum of probabilities of all pixels having a gray value less than or equal to ; Based on the cumulative distribution function, mapping to the entire grayscale range [0, L-1], the transformation formula is: in, It is the original gray level. The new gray level obtained after equalization; It is a transformation function; It rounds to the nearest integer. 5.The bridge inspection oriented unmanned aerial vehicle image apparent disease intelligent identification method according to claim 1, characterized in that, In step S3, the StarNet module backbone feature extraction network adopts the star operation algorithm. In the first layer of the star operation, a single star operation implicitly expresses a dimension of... The high-dimensional feature space realizes the implicit dimensional feature space representation, where each item has a non-linear relationship with the input; When stacking multiple layers, if the number of channels of the input features in the first layer is... ,go through Layer star operations, implicitly obtaining belonging to The feature space, Indicates the first The dimension of the implicit feature space corresponding to the layer output. 6.The bridge inspection oriented unmanned aerial vehicle image apparent disease intelligent identification method according to claim 1, characterized in that, In step S4, the new multi-scale convolutional attention module consists of three parts: depthwise convolution, which aggregates local information, and multi-branch depthwise separable convolution, which captures multi-scale contextual information through convolutional kernels of different sizes. Convolution is used to establish relationships between different channels; multiple such building blocks are stacked to form a convolutional encoder by reweighting the input of a new multi-scale convolutional attention module.

7. The intelligent identification method for apparent defects in bridge inspection images using UAVs, as described in claim 6, is characterized in that... The new multi-scale convolutional attention module adopts a hierarchical structure, consisting of four stages, with the spatial resolution gradually decreasing in each stage. Each stage contains a sampling block, which consists of a convolutional layer with a stride of 2 and a kernel size of 3x3, and a batch normalization layer. 8.The bridge inspection oriented unmanned aerial vehicle image apparent disease intelligent identification method of claim 1, wherein, In step S6, different bridge defects in the images obtained after preprocessing in step S2 are labeled with different colors using image feature annotation software; at the same time, the image data is classified and grouped into training set, test set and validation set using a 7:2:1 allocation ratio.

9. The intelligent identification method for apparent defects in bridge inspection images using UAVs, as described in claim 8, is characterized in that... In step S6, the labeled and grouped image data are trained by group. During model training, the labels that are incorrect or flawed are adjusted.