Bridge disease automatic identification method based on unmanned aerial vehicle vision and deep learning

By combining multi-scale feature extraction, visual-inertial navigation tight coupling, and multi-task deep learning with LiDAR data, the precise three-dimensional localization and quantification of bridge defects were achieved. This solved the problems of inaccurate localization, insufficient recognition accuracy, and poor system fault tolerance in existing technologies, and generated an efficient and intelligent bridge defect detection report.

CN121834680APending Publication Date: 2026-04-10YUNNAN YUNLU ENG INSPECTION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing UAV-based bridge defect detection methods suffer from problems such as inaccurate positioning, blurry images, insufficient recognition accuracy, and poor system fault tolerance, making it difficult to achieve fully automated, intelligent, and high-precision bridge defect detection.

Method used

The system employs multi-scale feature extraction and quality assessment, tight coupling of vision and inertial navigation, and adaptive exposure control. It combines a visual bag-of-words model with geometric constraints for closed-loop detection, uses laser point cloud matching for verification, constructs a 3D point cloud model, and uses a multi-task deep learning model for disease identification and quantification. It also dynamically adjusts sensor fusion weights to achieve intelligent fault tolerance.

Benefits of technology

It enables precise three-dimensional localization and quantification of bridge defects, improves the objectivity, comprehensiveness and efficiency of inspection, ensures stable operation of the system under harsh conditions, and generates structured inspection reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834680A_ABST
    Figure CN121834680A_ABST
Patent Text Reader

Abstract

The invention discloses a bridge disease automatic identification method based on unmanned aerial vehicle vision and deep learning, and particularly relates to the technical field of bridge disease detection.The method comprises the steps that firstly, high-precision positioning and three-dimensional modeling of an unmanned aerial vehicle in a complex bridge environment are achieved by fusing vision, inertia and laser radar data; and sensor fault diagnosis and self-adaptive fusion capabilities are provided, so that the reliability of data acquisition is guaranteed. And then, performing bridge member semantic segmentation on the acquired high-definition image, and synchronously realizing accurate identification and classification of various diseases such as cracks and spalling by using a multi-task deep learning model based on HRNet. And finally, mapping a two-dimensional identification result to a three-dimensional model, completing geometric quantization and severity level evaluation of the disease in a real space, and automatically generating a structured detection report and a visual result. According to the invention, full-process automation from data acquisition to disease analysis is realized, and the efficiency, precision and intelligent level of bridge detection are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bridge disease detection, and more particularly to a bridge disease automatic identification method based on unmanned aerial vehicle vision and deep learning. BACKGROUND

[0002] As a key node of traffic infrastructure, the structural health of a bridge is directly related to public safety. Traditional bridge disease detection mainly relies on manual visual inspection or the use of bridge inspection vehicles and other equipment. This method is not only inefficient and costly, but also greatly influenced by personnel subjective experience, and has safety risks, making it difficult to achieve rapid, comprehensive and objective evaluation of large and medium-sized bridges.

[0003] In recent years, with the development of technology, visual detection methods based on unmanned aerial vehicle (UAV) platforms have emerged. By carrying a high-definition camera, this method can quickly obtain image data of the bridge surface, improving the detection efficiency and safety to a certain extent. However, the existing methods still have obvious limitations: first, the unmanned aerial vehicle is easily affected by wind disturbance and other factors during flight, resulting in inaccurate positioning and blurred images, and lacking accurate relative pose information with the bridge surface, making it difficult to accurately position and quantify the disease in three-dimensional space; second, traditional image processing algorithms have insufficient recognition accuracy and robustness for diseases such as cracks and spalling in complex backgrounds, and are easily disturbed by light, shadow and structural texture; in addition, existing methods rely on a single visual sensor and lack a backup plan in case of visual failure (such as strong light, smoke), and the overall fault tolerance and reliability of the system need to be improved.

[0004] Therefore, the existing technology has not effectively solved the problem of integrating autonomous and accurate positioning of unmanned aerial vehicles, reliable fusion of multiple sensors, and intelligent automatic identification and quantification of diseases. There is an urgent need for a comprehensive solution that can achieve full-process automation, intelligence and high precision from data acquisition, positioning and mapping to disease identification and quantification, in order to overcome the shortcomings of existing technology and meet the needs of modern bridge management for efficient, accurate and safe detection. SUMMARY

[0005] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present application provide a bridge disease automatic identification method based on unmanned aerial vehicle vision and deep learning.

[0006] To achieve the above-mentioned purpose, the present application provides the following technical solutions:

[0007] The bridge disease automatic identification method based on unmanned aerial vehicle vision and deep learning comprises the following steps:

[0008] S1. Localization and Mapping: Environmental perception is achieved through multi-scale feature extraction and quality assessment. Real-time pose calculation is optimized through tight coupling of visual and inertial navigation, and adaptive exposure control is used to ensure image quality, providing a data foundation for subsequent analysis.

[0009] S2. Loop Closure Correction and Optimization: Fast loop closure detection is performed using a visual bag-of-words model and geometric constraints, and then laser point cloud is used for matching verification. Finally, the loop closure constraints and inertial data are fused through pose graph optimization to construct a globally consistent 3D point cloud model of the bridge.

[0010] S3, Intelligent Fusion and Fault Tolerance: Calculate the health indicators of vision, IMU and LiDAR respectively, dynamically adjust the weight of multi-sensor fusion according to the indicators, and trigger replanning when the positioning uncertainty reaches the standard;

[0011] S4. Defect Identification and Quantification: First, the bridge image is segmented into components. Then, a multi-task deep learning model is used to detect and classify cracks and spalling defects. Finally, the two-dimensional identification results are mapped onto the three-dimensional model for cross-view fusion and quantification, and a structured inspection report and visualization results are generated.

[0012] Specifically, in S1:

[0013] Multi-scale feature extraction uses the ORB pyramid feature extraction method, with the number of pyramid layers set to 8 and the scale factor set to 1.2.

[0014] The number of features is dynamically adjusted based on the bridge surface material, including:

[0015] The number of target features on the concrete surface is 200-250;

[0016] The steel structure surface has 150-200 surfaces;

[0017] Minimum 50 features are retained in sparse texture regions.

[0018] Specifically, in S2:

[0019] The closed-loop detection adopts a multi-level screening: the first level uses a visual bag-of-words model for fast screening, with a dictionary size of 10k words and a similarity threshold set at 0.35;

[0020] The second layer performs geometric consistency verification, retaining candidate frames with an inlier ratio greater than 60%.

[0021] Perform laser point cloud matching on the selected candidate frames.

[0022] Specifically, in S3, sensor health monitoring includes:

[0023] Visual health is calculated based on a combination of feature quantity ratio, tracking success rate, and image blur.

[0024] IMU health is calculated based on the degree of deviation between acceleration measurements and expected values;

[0025] LiDAR health is calculated based on point cloud density and data packet loss rate;

[0026] The fusion weights of vision, inertial navigation, and lidar are dynamically adjusted based on health indicators.

[0027] Specifically, multiple fusion modes are set based on health indicators:

[0028] In normal mode, the weighting ratio of vision, inertial navigation, and lidar is 0.5:0.3:0.2;

[0029] In visual degradation mode, the weighting ratio of vision, inertial navigation, and lidar is 0.2:0.5:0.3;

[0030] In IMU abnormal mode, the weighting ratio of vision, inertial navigation, and lidar is adjusted to 0.6:0.1:0.3;

[0031] In severe degradation mode, a pure laser odometer is activated.

[0032] Specifically, in S4:

[0033] The component segmentation adopts the DeepLabV3+ semantic segmentation model, and the segmentation categories include bridge deck, beam body, pier or tower and bearing;

[0034] Disease identification employs a multi-task deep learning model based on HRNet, simultaneously performing target detection of blocky diseases and pixel-level segmentation of cracks.

[0035] Specifically, in S4, the quantification of diseases includes:

[0036] The two-dimensional recognition results are back-projected onto the three-dimensional point cloud model to calculate the actual length and average width of cracks and the actual area of ​​spalling or pits.

[0037] Based on the quantitative results and the bridge technical condition assessment standards, the severity level is classified, and a structured report containing the type, location, size, and level of the defects is output.

[0038] The technical effects and advantages of this invention are as follows:

[0039] This invention constructs a high-precision 3D point cloud model of a bridge by fusing visual, inertial, and lidar data, providing an accurate spatial coordinate reference for bridge defects. Utilizing a multi-task deep learning model, it automatically segments bridge components and classifies defects, fusing the 2D recognition results with the 3D model to achieve precise geometric quantification (such as length, width, and area) of defects like cracks and spalling, ultimately generating a structured inspection report. This completely changes the traditional, inefficient, manual inspection method, significantly improving the objectivity, comprehensiveness, and efficiency of inspections.

[0040] This invention designs a dynamic sensor fault diagnosis and adaptive weight adjustment mechanism, which can assess the health status of each sensor in real time and intelligently adjust the fusion strategy in cases of visual degradation or IMU anomalies, ensuring stable operation of the system under harsh conditions such as drastic changes in lighting and texture loss. Simultaneously, closed-loop detection enhanced by LiDAR effectively eliminates accumulated errors, guaranteeing global consistency between localization and the model. This integrated intelligent fault-tolerant design greatly enhances the practicality and reliability of the entire system in real-world bridge inspection scenarios. Attached Figure Description

[0041] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0043] like Figure 1 As shown, the steps of the automatic bridge defect identification method based on UAV vision and deep learning are as follows:

[0044] Step 1: Localization and mapping. Environmental perception is achieved through multi-scale feature extraction and quality assessment. Real-time pose calculation is optimized through tight coupling of visual and inertial navigation, and adaptive exposure control is used to ensure image quality, providing a stable and accurate data foundation for subsequent analysis.

[0045] The specific steps are as follows:

[0046] Multi-scale feature extraction and quality assessment:

[0047] Pyramid ORB feature extraction was used, with the number of pyramid layers set to 8 and the scale factor set to 1.2.

[0048] Define the feature quality scoring formula:

[0049]

[0050] in: This represents the overall quality score of the feature points. These are the corner response values ​​at the feature point locations, derived from the Harris corner detection algorithm. The image gradient magnitude at the location of the feature point. It is a measure of the consistency between a feature point and its surrounding pixels. , , These are the weight coefficients for the corresponding terms, and =0.4, =0.3, =0.3;

[0051] The number of features is dynamically adjusted based on the bridge surface material:

[0052] Concrete surface: 200-250 target features; Steel structure surface: 150-200 target features; Sparse texture areas: retain at least 50 high-quality features;

[0053] Visual-inertial navigation tight coupling optimization:

[0054] Establish a nonlinear optimization problem:

[0055]

[0056] in: This represents the minimization operation, which involves finding the state variable that minimizes the objective function value. Let x be the observation matrix for visual reprojection, and let x be the system's state variables, typically including position, attitude, velocity, IMU bias, etc. The velocity observation matrix, The observation matrix is ​​biased. This is due to visual reprojection error. For IMU speed error, This refers to the IMU bias error.

[0057] The LM optimization algorithm is used, with a maximum of 50 iterations.

[0058] Adaptive exposure control based on lighting conditions:

[0059] Real-time analysis of image histogram distribution:

[0060] If the overexposed pixel ratio is greater than 30%, reduce the exposure time by 20%.

[0061] If the underexposed pixel ratio is greater than 40%, increase the exposure time by 25%.

[0062] Keep the average gray level of the image within the range of 80-120.

[0063] Step 2: Closed-loop correction and optimization. Rapid loop closure detection is performed using a visual bag-of-words model and geometric constraints. Then, fine matching verification is performed using laser point clouds. Finally, the loop closure constraints and inertial data are fused through pose graph optimization to construct a globally consistent high-precision 3D point cloud model of the bridge.

[0064] Multi-level loopback candidate frame filtering:

[0065] First layer: Visual bag-of-words model for fast filtering: using the DBoW2 dictionary with a size of 10k words; similarity threshold set at 0.35;

[0066] Second layer: Geometric consistency verification: Calculate the epipolar geometric constraints of feature points between candidate frames and the current frame; retain candidate frames with an inlier ratio > 60%;

[0067] Fine-grained laser point cloud matching: For the selected candidate frames, perform the following point cloud processing procedure:

[0068] Voxel filtering downsampling, grid size 0.05m; normal vector estimation, search radius 0.1m; ICP fine registration, maximum corresponding distance set to 0.1m;

[0069] Calculate the registration score: ;in The score for laser point cloud registration. Root mean square error (RMSE) is the average distance error between corresponding points after ICP registration, expressed in meters. The in-point ratio is the proportion of successfully registered point pairs to the total number of point pairs, reflecting the degree of consistency in registration.

[0070] Optimized pose graph construction:

[0071] Node: Keyframe Pose ;in, Let be the pose of the i-th keyframe. Let be the rotation matrix for the i-th keyframe, and be a 3×3 orthogonal matrix. Let be the translation vector of the i-th keyframe, which is a 3×1 vector;

[0072] The edges include:

[0073] Visual inertial edges: connect consecutive keyframes; loop closure edges: connect loop closure keyframe pairs; prior edges: fix the initial pose;

[0074] Optimize the objective function:

[0075]

[0076] in: For visual inertial edge error, The information matrix weights for visual inertial edges, This is the lap edge error. The information matrix weights of the loop edges, For prior edge error, The weights are the information matrix weights of the prior edges.

[0077] Step 3: Intelligent fusion and fault tolerance. The health indicators of vision, IMU and LiDAR are calculated separately. The weights of multi-sensor fusion are dynamically adjusted according to these indicators. When the positioning uncertainty is too high, replanning is triggered to ensure the reliability and security of the data acquisition path.

[0078] Real-time sensor health monitoring:

[0079] Visual health index:

[0080]

[0081] in: For the health of the vision sensor, This is the weighting coefficient for the feature quantity term, with a value of 0.4. This represents the actual number of feature points extracted in the current frame. This represents the number of target feature points, with a value of 150. This is the weighting coefficient for the tracking rate term, with a value of 0.4. The feature point tracking success rate is the ratio of the number of successfully tracked feature points in the current frame to the number of feature points in the previous frame. This is the weighting coefficient for the ambiguity term, with a value of 0.2. The image blur rating is a numerical value between 0 and 1 obtained by methods such as calculating the Laplacian variance of the image.

[0082] IMU Health Indicators:

[0083]

[0084] in: For the health of the IMU sensor, The acceleration value actually measured by the IMU. For the desired acceleration value, The threshold value for acceleration error is set to... , This represents the absolute error between the measured acceleration value and the expected value.

[0085] LiDAR health status:

[0086]

[0087] in: For the health status of lidar, Point cloud density, The data packet loss rate is the proportion of point cloud data packets lost during transmission.

[0088] Adaptive fusion weight adjustment:

[0089] Normal mode: Vision:Inertial:Laser = 0.5:0.3:0.2;

[0090] Visual degradation mode ( <0.6): Visual weight decreases to 0.2, inertial navigation weight increases to 0.5, and laser weight increases to 0.3;

[0091] IMU abnormal mode ( <0.7): Inertial navigation weight is reduced to 0.1, visual weight is increased to 0.6, and laser weight is increased to 0.3;

[0092] Severe degradation mode: Enable pure laser odometer, weight 1.0.

[0093] Positioning uncertainty propagation and assessment:

[0094] Define location reliability: ,in: For the system's location reliability, Let P be the trace of the state covariance matrix. The threshold for the covariance trace is 0.1. To limit the amplitude of the normalized covariance trace and ensure that the result does not exceed 1; when When the value is less than 0.8, the path replanning mechanism is triggered.

[0095] Step 4: Defect identification and quantification. First, the bridge image is segmented into components. Then, a multi-task deep learning model is used to detect and classify defects such as cracks and spalling. Finally, the two-dimensional identification results are mapped onto the three-dimensional model for cross-view fusion and accurate quantification, and a structured inspection report and visualization results are generated.

[0096] Receive a sequence of high-resolution images of the bridge surface with precise pose information generated after processing in steps one and two, and automatically identify and classify the defects in the bridge surface.

[0097] First, the input image is segmented at the component level to enable refined disease identification based on the characteristics of different components.

[0098] Segmentation Network: The DeepLabV3+ semantic segmentation model is adopted. Its encoder-decoder structure can effectively fuse multi-scale information and accurately segment bridge components.

[0099] Segmentation categories: The model was trained to identify the following four categories of bridge components: bridge deck, beams, piers / towers, and bearings;

[0100] Post-processing: Use morphological operations (such as closing operations) to smooth the segmented component masks and eliminate small holes.

[0101] ROI extraction: Based on the mask obtained from segmentation, the image regions of each component are cropped from the original image and used as direct input for subsequent disease identification.

[0102] A multi-task deep learning model is used to simultaneously perform disease detection and classification:

[0103] Model backbone network: HRNet (High-Resolution Network) is selected as the feature extraction backbone, which can always maintain high-resolution feature representation and is particularly suitable for the identification of slender diseases such as cracks.

[0104] Task 1: Disease Target Detection (for patchy diseases):

[0105] Detection head: A RetinaNet detection head is connected to the feature map extracted by HRNet to identify blocky defects such as concrete spalling, potholes, and water seepage with obvious areas.

[0106] Output: For each detected disease, output its bounding box, category label, and confidence score;

[0107] Task 2: Pixel-level segmentation of diseases (for linear diseases):

[0108] Segmentation Header: A lightweight segmentation head is connected to another output of HRNet to perform pixel-level segmentation of the cracks.

[0109] Output: Generate a binary segmentation map with the same size as the input image, where white pixels represent cracks and black pixels represent the background.

[0110] Model Input and Preprocessing: The input is the extracted ROI image of the component, which is uniformly scaled to 640x640 pixels; normalization processing is performed to scale the pixel values ​​from [0,255] to [0,1].

[0111] Model inference: The preprocessed image is input into the multi-task model, and object detection and pixel segmentation are performed in parallel to obtain preliminary recognition results;

[0112] Cross-view defect fusion and quantification: Pose-based projection fusion: Using the keyframe poses and camera intrinsic parameters optimized in step two, the defect bounding boxes or crack pixels identified in the current frame are back-projected onto the bridge 3D point cloud map constructed in step two.

[0113] For cracks, the projections of the same crack in different frames in three-dimensional space are clustered and merged into a complete three-dimensional crack.

[0114] Quantitative analysis of defects: Cracks: Calculate their actual length (meters) and average width (millimeters) in the 3D model; Peeling / Pits: Calculate their actual area (square meters) in the 3D model.

[0115] Severity level assessment: Based on the quantitative results and in conjunction with the "Technical Condition Assessment Standard for Highway Bridges" (JTG / TH21), a preliminary severity level classification (e.g., mild, moderate, severe) is made for various types of defects.

[0116] Model training and data preparation:

[0117] The semantic segmentation model and the multi-task recognition model are trained specifically, as follows:

[0118] Dataset construction: Collect tens of thousands of bridge images with different environments, lighting conditions, and bridge types.

[0119] The images were finely annotated using tools such as Labelme and CVAT. The annotations included: polygons segmented from components, bounding boxes and categories of blocky defects, and pixel-level annotations of cracks.

[0120] Training strategy: Transfer learning: Initialize the model backbone network with weights pre-trained on large public datasets such as COCO and Cityscapes to accelerate convergence and improve performance;

[0121] Data augmentation: Online data augmentation techniques such as random rotation (±10°), brightness and contrast adjustment, Gaussian noise, and random cropping are used to improve model robustness;

[0122] Hyperparameters: The Adam optimizer was used, with an initial learning rate of 1e-4, a batch size of 8, and a total training duration of 300 epochs.

[0123] Recognition result output and visualization

[0124] Ultimately, structured disease reports and visualizations are generated;

[0125] Structured report: Output a report in JSON or XML format, including: bridge ID, inspection time;

[0126] The list of all identified defects includes, for each record: defect type, component, location in three-dimensional space, quantitative dimensions (length / width / area), and severity level;

[0127] Visualization results: Generate a 3D point cloud model of the bridge showing the distribution of defects, with different defects highlighted in different colors (e.g., red represents cracks, yellow represents spalling).

[0128] Detailed images are generated for each key disease, including the original image, an overlay of the recognition results, and quantitative data.

[0129] The above formulas are all normalized and calculated based on their numerical values. The formulas are derived from software simulations using a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0130] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, ATA hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state ATA hard disk.

[0131] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0132] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0133] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0134] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0135] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0136] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable ATA hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0137] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An automatic bridge defect identification method based on UAV vision and deep learning, characterized in that, Includes the following steps: S1. Localization and Mapping: Environmental perception is achieved through multi-scale feature extraction and quality assessment. Real-time pose calculation is optimized through tight coupling of visual and inertial navigation, and adaptive exposure control is used to ensure image quality, providing a data foundation for subsequent analysis. S2. Loop Closure Correction and Optimization: Fast loop closure detection is performed using a visual bag-of-words model and geometric constraints, and then laser point cloud is used for matching verification. Finally, the loop closure constraints and inertial data are fused through pose graph optimization to construct a globally consistent 3D point cloud model of the bridge. S3, Intelligent Fusion and Fault Tolerance: Calculate the health indicators of vision, IMU and LiDAR respectively, dynamically adjust the weight of multi-sensor fusion according to the indicators, and trigger replanning when the positioning uncertainty reaches the standard; S4. Defect Identification and Quantification: First, the bridge image is segmented into components. Then, a multi-task deep learning model is used to detect and classify cracks and spalling defects. Finally, the two-dimensional identification results are mapped onto the three-dimensional model for cross-view fusion and quantification, and a structured inspection report and visualization results are generated.

2. The automatic bridge defect identification method based on UAV vision and deep learning according to claim 1, characterized in that, In S1: Multi-scale feature extraction uses the ORB pyramid feature extraction method, with the number of pyramid layers set to 8 and the scale factor set to 1.

2. The number of features is dynamically adjusted based on the bridge surface material, including: The number of target features on the concrete surface is 200-250; The steel structure surface has 150-200 surfaces; Minimum 50 features are retained in sparse texture regions.

3. The automatic bridge defect identification method based on UAV vision and deep learning according to claim 1, characterized in that, In S2: The closed-loop detection adopts a multi-level screening: the first level uses a visual bag-of-words model for fast screening, with a dictionary size of 10k words and a similarity threshold set at 0.35; The second layer performs geometric consistency verification, retaining candidate frames with an inlier ratio greater than 60%. Perform laser point cloud matching on the selected candidate frames.

4. The automatic bridge defect identification method based on UAV vision and deep learning according to claim 1, characterized in that, In step S3, sensor health monitoring includes: Visual health is calculated based on a combination of feature quantity ratio, tracking success rate, and image blur. IMU health is calculated based on the degree of deviation between acceleration measurements and expected values; LiDAR health is calculated based on point cloud density and data packet loss rate; The fusion weights of vision, inertial navigation, and lidar are dynamically adjusted based on health indicators.

5. The automatic bridge defect identification method based on UAV vision and deep learning according to claim 4, characterized in that, Multiple fusion modes can be set based on health indicators: In normal mode, the weighting ratio of vision, inertial navigation, and lidar is 0.5:0.3:0.2; In visual degradation mode, the weighting ratio of vision, inertial navigation, and lidar is 0.2:0.5:0.3; In IMU abnormal mode, the weighting ratio of vision, inertial navigation, and lidar is adjusted to 0.6:0.1:0.3; In severe degradation mode, a pure laser odometer is activated.

6. The automatic bridge defect identification method based on UAV vision and deep learning according to claim 1, characterized in that, In S4: The component segmentation adopts the DeepLabV3+ semantic segmentation model, and the segmentation categories include bridge deck, beam body, pier or tower and bearing; Disease identification employs a multi-task deep learning model based on HRNet, simultaneously performing target detection of blocky diseases and pixel-level segmentation of cracks.

7. The automatic bridge defect identification method based on UAV vision and deep learning according to claim 6, characterized in that, In S4, the quantification of disease includes: The two-dimensional recognition results are back-projected onto the three-dimensional point cloud model to calculate the actual length and average width of cracks and the actual area of ​​spalling or pits. Based on the quantitative results and the bridge technical condition assessment standards, the severity level is classified, and a structured report containing the type, location, size, and level of the defects is output.