Pier damage identification method based on double-region flight and double-mainstem optimization network

By combining dual-region flight control and dual-backbone optimized networks, the problem of GNSS signal obstruction in UAV bridge pier damage identification was solved, achieving high-precision 3D reconstruction and damage identification, and improving detection efficiency and accuracy.

CN121661550BActive Publication Date: 2026-04-28CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
Filing Date
2026-02-06
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing UAV-based bridge pier damage identification technologies suffer from problems such as imaging blind spots caused by GNSS signal obstruction, parallax artifacts, insufficient image overlap, and inconsistent image quality, which affect the efficiency and accuracy of 3D reconstruction and damage identification.

Method used

A dual-region flight control strategy is adopted, combining RTK positioning and PID cascade control to achieve manual flight of the upper part of the bridge pier and automatic flight of the lower part. Image acquisition and recognition are performed by combining a dual-backbone optimization network, 3D reconstruction is performed by the bundle adjustment algorithm with RTK position constraints, and damage identification is performed by an improved YOLOv11 neural network.

Benefits of technology

It enables full-size 3D model acquisition under GNSS signal obstruction conditions, improves the accuracy of 3D reconstruction and damage identification, adapts to various pier cross-sectional shapes, reduces errors from manual control, and improves detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661550B_ABST
    Figure CN121661550B_ABST
Patent Text Reader

Abstract

The application discloses a pier damage identification method based on double-region flight and double-main optimization network, relates to the technical field of bridge vibration analysis, and comprises the following steps: firstly, a double-region flight control strategy combining PID cascade flight control and RTK positioning automatic flight control is adopted to realize stable image collection of the upper region of the pier and automatic and rapid image collection of the lower region of the pier in the absence of GNSS signals; secondly, a standard elliptical layered flight path is adopted to standardize flight paths of various pier section shapes; then, a double-region image set fusion three-dimensional reconstruction framework based on a motion recovery structure algorithm and RTK position constraint is proposed to generate a detailed model of the pier; finally, a YOLOv11 network is improved through a double-main and multi-layer detection head to improve the crack and defect detection precision of the neuron network in the visualized image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bridge pier damage identification technology, and in particular to a bridge pier damage identification method based on dual-region flight and dual-backbone optimization network. Background Technology

[0002] Bridge piers, as crucial load-bearing components of modern transportation infrastructure, are constantly subjected to complex coupled effects of multiple hazards, including cyclic vehicle loads, environmental erosion, and material aging. These factors inevitably lead to the development of surface defects such as cracks and spalling. Without timely and accurate identification, this damage can propagate, compromising structural integrity and causing serious safety consequences. Traditional inspection methods primarily rely on manual climbing or close-range visual inspection from vehicles beneath the bridge, which suffers from high operational risks, low efficiency, and strong subjectivity. Their qualitative approach hinders the quantitative, digital damage assessment required for modern structural health monitoring (SHM).

[0003] Machine vision technology has emerged as a promising new approach for structural health monitoring, gaining attention for its flexibility, efficiency, and ability to capture high-resolution images. This has spurred the development of specialized platforms, primarily wall-climbing robots and drones. For example, Jang et al. developed an automated ring robot for detecting cracks in bridge piers, while Ding et al. integrated a Feature Pyramid Transformer (FPT) network into a wall-climbing robot for fine-grained crack segmentation. Similarly, Du et al. proposed an intelligent detection framework using a PD-Mamba model for accurate defect detection. However, these robots often face the challenge of limited camera field of view, requiring repeated repositioning to achieve full coverage. This reduces detection efficiency and introduces challenges related to operational safety, endurance, and deployment in complex environments such as waterways or canyons.

[0004] Unmanned aerial vehicles (UAVs) offer a compelling solution due to their ease of operation, mobility, and cost-effectiveness, attracting significant research interest. Recent advances have focused on integrating UAVs with advanced algorithms. Ni et al. proposed a cross-dimensional collaborative YOLO (CDC-YOLO) model for multi-defect detection in UAV imagery. Path planning was optimized using an improved ant colony optimization algorithm and a multi-level representation learning strategy. Furthermore, high-precision 3D reconstruction and defect localization techniques have been developed, such as lightweight panoramic reconstruction and depth-guided 3D visualization crack segmentation. The integration of UAV photogrammetry, deep learning, and digital twin methods for automated assessment has also been explored. Despite these advances, key challenges remain in the inspection of vertical bridge pier structures. Autonomous flight near structures faces issues such as real-time collision avoidance, frequent GNSS signal obstruction (especially near the top of the pier), and a lack of intelligent path planning algorithms that dynamically adapt to complex structural geometries. Therefore, maintaining a balance between effectively executing autonomous flight paths and acquiring images suitable for high-precision 3D reconstruction and visual recognition remains a critical research gap.

[0005] Currently, there are two main methods for identifying bridge pier surface damage based on UAVs: (1) manual trajectory flight relying on pilot experience; and (2) automatic trajectory planning based on predefined models. Liu et al. achieved a resolution of up to 0.2 mm / pixel in single pier detection. However, in multi-pier groups, the occlusion of adjacent piers can lead to more than 35% of the imaging blind zone and produce parallax artifacts, which seriously affect the reconstruction quality. The cuboid-constrained ant colony optimization method proposed by Ni et al. improves the efficiency of aerial photography; however, its performance is limited by the accuracy and completeness of the initial model and cannot achieve high-resolution imaging at close range. The upper region of the bridge pier is close to the bottom of the beam, and the GNSS signal is blocked, which leads to the interruption of autonomous flight operations. In addition, long-term manual flight control often leads to imaging defects, including missed detections and insufficient image overlap.

[0006] The flight strategy of unmanned aerial vehicles (UAVs) is a decisive factor affecting the quality of 3D reconstruction and is the foundation for accurate damage assessment. Current methods typically involve manually guiding the UAVs to circle close to the bridge piers or rely on pre-planned paths based on simplified geometric assumptions. Compared with traditional methods, these strategies have limitations while reducing costs:

[0007] (1) Automatic flight mode cannot robustly adapt to obstruction and complex pier surfaces, often requiring high-precision prior models (>0.1 m) and facing the risk of incomplete data collection;

[0008] (2) Manual control leads to inconsistencies between image quality and georeferenced accuracy;

[0009] (3) The loss of GNSS signal on the upper section of the bridge pier seriously affects the quality of location data for image registration. These problems together affect the efficiency, model fidelity and measurement accuracy of UAV-based bridge pier inspection.

[0010] Parallel advancements in computer vision, such as the YOLO series and other deep learning-based object detection algorithms, have provided powerful tools for automatic damage identification. Lightweight architectures and enhanced multi-scale feature learning have improved the detection of small, low-contrast defects, enabling the identification of a wide variety of surface damage. However, achieving robust and accurate cross-scale defect identification in real-world environments with variable lighting, shading, and complex textures remains a significant challenge. Summary of the Invention

[0011] This invention provides a bridge pier damage identification method based on dual-region flight and dual-backbone optimization network. By fusing the substructure method with a dynamic explicit-implicit joint solution strategy, it achieves high-precision analysis of the dynamic response of key local areas of bridges.

[0012] To solve the above-mentioned technical problems, the technical solution proposed by this invention is as follows:

[0013] A bridge pier damage identification method based on dual-region flight and dual-backbone optimization network includes the following steps:

[0014] Step S1: Use drone close-range photogrammetry to capture the surface structure of the bridge piers, using a controlled vertical shooting distance;

[0015] Step S2: A dual-zone flight control strategy is adopted: the bridge pier is divided into a lower automatic flight zone and an upper manual control zone. PID cascade manual flight is used in the upper zone of the bridge pier, and RTK positioning automatic flight is used in the lower zone. Images are acquired along a standardized elliptical layered path.

[0016] Step S3, 3D Reconstruction: The dual-region images obtained in step S2 and their corresponding RTK positioning coordinate information are fused together, and the motion recovery structure 3D reconstruction is performed using the bundle adjustment algorithm with RTK position constraints to generate a 3D model of the bridge pier.

[0017] Step S4, Intelligent Surface Damage Recognition: Construct and train a dual-backbone enhanced YOLOv11 target detection neural network, input the bridge pier surface image collected in step S2 into the network, and identify and locate cracks and spalling defects in the image.

[0018] A further improvement to the above technical solution is as follows:

[0019] Preferably, in step S2, the automatic flight control of the lower region specifically involves: the UAV receiving the real-time position located by RTK as feedback, comparing it with the preset standardized elliptical layered path waypoint coordinates, and driving the UAV to automatically fly along the preset path and triggering the camera to take pictures through a cascaded PID controller.

[0020] Preferably, in step S2, the manual flight control of the upper region specifically includes:

[0021] Sensor fusion state estimation: Data from visual odometry, inertial measurement unit, ultrasonic or infrared ranging sensors are fused, and state estimation is performed through extended Kalman filter to output the position, velocity and attitude information of the UAV;

[0022] Cascaded PID control: It adopts a four-loop cascaded control structure of position loop, velocity loop, attitude loop and angular velocity loop. The desired control command is input, and the controller calculates and outputs the final control quantity based on the state estimation information, so as to control the UAV in the absence of satellite navigation signal.

[0023] Preferably, the standardized elliptical layered path is as follows: the lateral dimension of the pier determines the length of the straight line segment of the ellipse, the diameter of the semicircular segment depends on the longitudinal dimension of the pier and the object distance, and the flight track shooting points are planned through 4 elliptical control points, object distance d and overlap rate parameters.

[0024] Preferably, the three-dimensional reconstruction of the motion recovery structure in step S3 consists of four main stages:

[0025] Feature extraction and matching: Feature point detection and description are performed on all images acquired from dual regions, and the correspondence between images is established through feature matching algorithms;

[0026] Initial two-view reconstruction: Select two images with high overlap for initial two-view reconstruction, and calculate the fundamental matrix through epipolar geometry;

[0027] New incremental registration of views using PnP estimation combined with triangulation: Initial 3D sparse point cloud is generated by triangulation, and new camera pose parameters are estimated using the PnP method.

[0028] Bundle adjustment optimization: Construct a cost function with RTK constraints to optimize camera pose and 3D point coordinates.

[0029] Preferably, in step S3, the RTK position constraint is incorporated into the cost function, as shown below:

[0030]

[0031] in, It is the expected minimum total cost. and The covariance matrices of visual observation and RTK observation are respectively. These are RTK weight parameters. It is an indicator variable. These are the observed pixel coordinates of the j-th 3D point on the i-th image. Let K be the j-th 3D point, and K be the parameter matrix within the camera. Let be the rotation matrix in the external reference of the camera for the i-th image. Let be the three-dimensional coordinates of the projection center of the i-th image in the camera coordinate system in the traditional SfM algorithm.

[0032] Preferably, the dual-backbone enhanced YOLOv11 target detection neural network is a dual-backbone parallel feature extraction network, combined with a detection head structure optimized for small targets.

[0033] Preferably, the dual-backbone parallel feature extraction network is as follows: on the YOLOv11 baseline model, a parallel backbone network branch is added, the original backbone is used to extract deep semantic features, and the newly added branch adopts a shallow structure.

[0034] The bridge pier damage identification method based on dual-region flight and dual-backbone optimization network provided by this invention has the following advantages compared with the prior art:

[0035] (1) The bridge pier damage identification method based on dual-region flight and dual-backbone optimization network of the present invention adopts a dual-region flight strategy to realize the image acquisition of the full-size three-dimensional model of the bridge pier under the interference of GNSS information of the main beam, and uses the RTK image projection center position constraint enhancement SfM algorithm to realize the three-dimensional fusion reconstruction of images containing regional positioning information and non-positioning information images; in addition, the layered standard elliptical path adopted can meet the standardization of flight path of bridge piers with different structural shapes, which is crucial for obtaining accurate 3D modeling results.

[0036] (2) The pier damage identification method based on dual-region flight and dual-backbone optimization network of the present invention realizes three-dimensional reconstruction and standardized structural path planning, high-precision three-dimensional model reconstruction and neural network-based visual identification of piers with various cross-sectional sizes. Attached Figure Description

[0037] Figure 1 This is a schematic diagram illustrating the derivation process of close-up photography of bridge piers in this invention.

[0038] Figure 2 This is a schematic diagram of the dual-zone flight control principle of the present invention.

[0039] Figure 3 This is a schematic diagram of the standard elliptical flight path of the bridge pier of this invention.

[0040] Figure 4 This is a schematic diagram of the dual-backbone improved YOLOv11 neural network of the present invention.

[0041] Figure 5 This is a schematic diagram of the dual-region image fusion of pier #1 in the experimental verification.

[0042] Figure 6 This is a schematic diagram of the 3D point cloud fusion of pier #2 in the experimental verification.

[0043] Figure 7 Visual comparison of cracks in various improved YOLOv11 models during experimental verification.

[0044] Figure 8 This is a schematic diagram showing the dimensions and mass of the three-dimensional reconstruction model of pier #1 during the experimental verification.

[0045] Figure 9 The diagram shows the quality analysis of the connection point of pier #2, where (a) is the distribution diagram of the uncertainty of the connection point of pier #2, and (b) is the distribution diagram of the resolution of the connection point of pier #2. Detailed Implementation

[0046] The following provides a detailed description of specific embodiments of the present invention. It should be understood that the specific embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit the scope of the invention.

[0047] This invention presents a bridge pier damage identification method based on dual-region flight and a dual-backbone optimized network, primarily addressing flight path planning and surface damage identification for 3D reconstruction of bridge piers, a key challenge for unmanned aerial vehicles (UAVs) used in structural health monitoring. The method proposes an integrated automatic detection framework that combines an adaptive dual-region flight control strategy with a deep learning-based detection network. First, a dual-region flight control strategy combining PID cascade flight control and RTK positioning-based automatic flight control is employed to achieve stable image acquisition of the upper region of the bridge pier in the absence of GNSS signals and automated, rapid image acquisition of the lower region. Second, flight paths for various bridge pier cross-sectional shapes are standardized using a standard elliptical layered flight path. Then, a dual-region image set fusion 3D reconstruction framework based on the Structure from Motion (SfM) algorithm and RTK position constraints is proposed to generate detailed models of the bridge pier. Finally, the YOLOv11 network is improved through a dual-backbone and multi-layer detection head to enhance the accuracy of crack and defect detection in the neural network of the visualized image.

[0048] The bridge pier damage identification method based on dual-region flight and dual-backbone optimization network of the present invention includes the following steps:

[0049] Step S1: Close-range photogrammetry using a drone.

[0050] Close-range photogrammetry is used to capture the surface structure of bridge piers, replacing traditional aerial height measurement with controlled vertical shooting distance to ensure sub-millimeter image accuracy.

[0051] like Figure 1 As shown, the basic principle of photogrammetry is central perspective projection, which geometrically connects an object point (P*) in three-dimensional space, the optical center (O) of the camera lens, and the corresponding image point (P*) on the two-dimensional projection plane through a straight line relationship.

[0052] The collinearity condition equation for three collinear points is expressed as:

[0053] ;

[0054] (1)

[0055] Where (x, y) are the coordinates of the image point in the image plane coordinate system. The azimuth angle is the internal angle of the camera, with elements being the coordinates of the principal point of the image and the camera focal length. These are the coordinates of a point on the object in the world coordinate system. The coordinates (exterior orientation line elements) of the camera lens center in the world coordinate system. These are the elements that constitute the rotation matrix from the world coordinate system to the camera coordinate system.

[0056] In aerial photogrammetry, ground sampling distance (GSD) is a direct indicator of image granularity and an important metric for image resolution. In this aerial photograph, The relationship is represented as:

[0057] (2)

[0058] in, This indicates the image sampling resolution during image acquisition controlled by altitude and distance flight. Indicates the camera's focal length. Indicates the pixel size of the camera sensor; It refers to the flight altitude, i.e., the altitude at which aerial photography is conducted.

[0059] In close-up drone photography, the camera's optical axis is vertically aligned with the bridge pier surface. The physical relationship of the ground sampling distance formula changes from being controlled by aerial photography altitude to being controlled by shooting distance. Simultaneously, the sampling distance perpendicular to the object's surface determines the lower limit of damage identification capability. The relational expression is as follows:

[0060] (3)

[0061] in, This indicates the image sampling resolution during image acquisition controlled by the vertical distance of the object. D is the vertical distance from the center of the camera lens to the surface of the bridge pier, i.e., the shooting distance.

[0062] Step S2, dual-area flight strategy.

[0063] like Figure 2 As shown, this invention proposes a dual-region cooperative flight strategy for bridge piers with complex configurations. For the areas near the top of the bridge pier and the bottom of the beam that are shielded by GNSS, a multi-sensor fusion system combining visual odometry, inertial measurement, ultrasonic / infrared ranging, and barometric pressure data is implemented. The combined system achieves steady-state estimation through Kalman filtering and high-precision manual on-orbit flight through cascaded PID control, ensuring imaging stability and comprehensive coverage. In the area below the bridge pier where GNSS is enabled, dual-antenna RTK technology combined with PID flight control achieves centimeter-level positioning accuracy, which is beneficial for automatic elliptical path navigation and effectively enables full-coverage data acquisition.

[0064] S2-1, GNSS-shielded area.

[0065] Traditional single sensors are inherently limited in their capabilities. Visual odometry (VO) suffers from cumulative errors, while accelerometer readings from inertial measurement units (IMUs) are subject to drift and gravity distortion. Therefore, the integration of sensors is essential in this invention to compensate for their respective weaknesses. An extended Kalman filter (EKF) is used for this purpose. The Kalman filter operates through a cyclic prediction-update iterative process, predicting the current state of the UAV, including position, velocity, and attitude, by combining the optimal estimate from the previous time step with angular velocity and acceleration data measured by the IMU. The state transition equation in the Kalman filter prediction step is defined as follows:

[0066] (4)

[0067] in, It is the prior prediction value of the current state. It is the prior prediction value at time step (k-1). It is the state transition matrix. It is a control input matrix. It is the input vector of the IMU.

[0068] The prior prediction of the current state is corrected by the observations from the VO sensor, and the optimal estimate of the system state is obtained by fusing the prediction and observation information. for:

[0069] ;

[0070] (5)

[0071] in, It is Kalman gain. It is the observation matrix. It is the observation noise covariance matrix; The velocity is calculated from the sensor observations using VO. It is the prior estimation error covariance matrix. It is the transpose of the observation matrix. It is the symbol for the transpose of a matrix. It is the prior prediction value of the current state.

[0072] Through continuous filtering, an optimal, smooth, and unbiased state estimate is finally output, including position, velocity, and attitude angle.

[0073] This invention employs a cascaded PID controller for flight control of a UAV. A single PID controller performs well in simple, slow-response systems. However, for complex systems like UAVs that are fast, multivariable, and subject to internal and external disturbances, a single PID controller falls short. Therefore, this invention introduces a cascaded PID controller, layering the complex control problem by breaking down a single controller into multiple cascaded, clearly defined sub-controllers: the outer loop controls position, the middle loop controls velocity, and the inner loop controls attitude. The flight control system uses a cascaded PID control structure, from the outside in: position loop → velocity loop → attitude loop → angular velocity loop, decomposing the complex flight control problem into several simpler, faster-responding sub-problems. Specifically, this includes the following:

[0074] (1) The horizontal position control loop is the outermost loop, responsible for controlling the UAV in the horizontal plane ( , The absolute position on the target, input the desired position ( , ) and the actual location of the fusion estimate ( , The error between the two values ​​is used to determine the target horizontal speed that the drone needs to achieve. , The relationship between horizontal position and horizontal velocity is as follows:

[0075] ;

[0076] (6)

[0077] in, It is the positional error in the horizontal X-axis direction. It is the desired position along the X-axis of the horizontal plane. It is the actual position in the horizontal X-axis direction. It is the positional error in the Y-axis direction of the horizontal plane. It is the desired position in the Y-axis direction of the horizontal plane. It is the actual position in the Y-axis direction of the horizontal plane. It is the target velocity in the horizontal X-axis direction. It is the target velocity in the Y-axis direction of the horizontal plane. , , These are the PID parameters for the horizontal position loop, specifically the integral term. , Used to eliminate steady-state position error; differential term , Errors based on the rate of change of position are used to provide damping and prevent overshoot.

[0078] (2) The horizontal speed control loop controls the horizontal movement speed of the UAV. The desired horizontal speed is input. The actual horizontal velocity estimated by fusion The error between them determines the target attitude angle that the drone body needs to tilt. The relationship between the horizontal velocity and the target attitude angle is as follows:

[0079] ;

[0080] (7)

[0081] in, For the target pitch angle, For the target roll angle, , For the PID parameters of the horizontal speed loop, The velocity error is in the horizontal X-axis direction. Let be the desired velocity in the X-axis direction of the horizontal plane. This represents the actual velocity along the horizontal X-axis. The velocity error is in the Y-axis direction of the horizontal plane. Let be the desired velocity in the Y-axis direction of the horizontal plane. The actual velocity in the Y-axis direction of the horizontal plane. The proportional gain parameter for the velocity loop in the X-axis direction. The integral gain parameter for the velocity loop in the X-axis direction. The differential gain parameter of the velocity loop in the X-axis direction is... The proportional gain parameter for the velocity loop in the Y-axis direction. The integral gain parameter for the velocity loop in the Y-axis direction. The differential gain parameter of the velocity loop in the Y-axis direction.

[0082] (3) The attitude control loop is used to control the UAV's own orientation (attitude). The input is the error between the desired attitude angle and the actual attitude angle estimated by IMU fusion. The expected yaw angle is usually directly controlled by the pilot or given by the navigation system (e.g., waypoint flight). The output is the target angular velocity applied to the airframe. The relationship between the target attitude angle and the target angular velocity is:

[0083] ;

[0084] (8)

[0085] in, , , For the PID parameters of the attitude loop, , , These are the roll rate, pitch rate, and yaw rate of the fuselage target axis, respectively. The error between the target and the actual roll angle. This is the actual roll angle. The error between the target pitch angle and the actual pitch angle. This is the actual pitch angle. The error between the target and the actual yaw angle. For the expected yaw angle, This is the actual yaw angle. The proportional gain parameter for the roll angle attitude loop. The integral gain parameter of the roll angle attitude loop is... For the differential gain parameter of the roll angle attitude loop, The proportional gain parameter for the pitch rate attitude loop. For the integral gain parameter of the pitch angle attitude loop, For the differential gain parameter of the pitch angle attitude loop, This is the proportional gain parameter for the yaw angle attitude loop. Here are the integral gain parameters for the yaw angle attitude loop. The differential gain parameter of the yaw angle attitude loop.

[0086] (4) The angular velocity control loop is the innermost and fastest-responding loop, directly stabilizing the rotational motion of the UAV. It takes as input the error between the desired angular velocity and the actual angular velocity directly measured by the gyroscope, and outputs the control torque to be applied to the airframe. The relationship between the target angular velocity and the control torque is:

[0087] ;

[0088] (9)

[0089] in, The error between the expected and actual roll angular velocities. This is the actual roll angular velocity. The error between the expected and actual pitch angular velocity, This is the actual pitch angular velocity. The error between the expected and actual yaw rate, This is the actual yaw rate. For rolling torque, For pitching moment, For yaw moment, The proportional gain parameter for the roll velocity loop. The integral gain parameter for the roll velocity loop. The differential gain parameter of the roll velocity loop, The proportional gain parameter for the pitch angular velocity loop. The integral gain parameter for the pitch angular velocity loop. The differential gain parameter of the pitch angular velocity loop, The proportional gain parameter for the yaw rate loop. The integral gain parameter for the yaw rate loop. The differential gain parameter of the yaw rate loop.

[0090] (5) The height control loop is an independent but parallel cascade structure. Height control and horizontal control are parallel, using cascaded PID. The outer loop of the height control loop is the position loop that controls the height, and the inner loop is the velocity loop that controls the height. The input is the error between the desired height and the actual height, and the output is the target vertical velocity. The relationship between the height position and the target vertical velocity is:

[0091] ;

[0092] (10)

[0093] in, This refers to the height position error in the Z-axis direction. The desired height position in the Z-axis direction. This represents the actual height position along the Z-axis. For the target vertical velocity, The proportional gain parameter for the height position loop. The integral gain parameter for the height position loop. The differential gain parameter of the height position loop.

[0094] The input is the error between the target vertical velocity and the actual vertical velocity, and the output is the total thrust (T*) used to control the reference speed of all motors. The relationship between altitude speed and total thrust is:

[0095] ;

[0096] (11)

[0097] in, The error between the target velocity and the actual vertical velocity. For the target vertical velocity, This is the actual vertical velocity. The proportional gain parameter for the vertical velocity loop. The integral gain parameter for the vertical velocity loop. The differential gain parameter of the vertical velocity loop.

[0098] The control torque and total thrust are distributed to the four motors. The relationship between the speed of the four motors and the control torque and total thrust is as follows:

[0099] (12)

[0100] in, This is the rotational speed of the i-th motor, where i = 1, 2, 3, 4. It is the lever arm. It is the rolling torque. It is the pitching moment. It is the yaw moment. It is the anti-torque coefficient.

[0101] The dual-region cooperative flight strategy of this invention enables UAVs to maintain excellent flight stability and controllability even in GNSS-shielded environments such as behind bridge piers, indoors, and in forests, providing crucial technical support for manual precision flight layer-by-layer fine inspection of bridge pier surfaces.

[0102] S2-2, the area for GNSS signals.

[0103] The area beneath the bridge piers has good satellite signal coverage. The higher the pier, the less the satellite signal is blocked by the main beam. A flight strategy combining standard elliptical layered orbital flight with D-RTK real-time positioning technology is employed in the area beneath the piers, enabling high-precision automatic flight path and image acquisition. This saves significant manual flight time and reduces errors caused by human operation. The main sources of error in conventional GNSS positioning include satellite clock bias, receiver clock bias, ionospheric delay, tropospheric delay, and orbital errors. RTK positioning eliminates these errors through carrier phase observations and differential analysis.

[0104] The UAV's RTK system receives strong satellite signals in the area below the bridge pier. As the height of the bridge pier increases, the proportion of satellite signals blocked by the main beam decreases accordingly. The lower-area flight strategy of this invention employs a standard elliptical layered orbital flight mode combined with D-RTK real-time positioning technology to achieve automatic route navigation and high-precision image acquisition. This method significantly reduces manual flight time while minimizing errors related to manual control. The main error sources in conventional GNSS positioning include satellite clock error, receiver clock error, ionospheric delay, tropospheric delay, and orbital error. Using RTK technology, the main errors in GNSS positioning are eliminated through carrier phase observation and differential analysis. The carrier phase observation equation is as follows:

[0105] (13)

[0106] in, The measured carrier phase value, It is the geometric distance between the satellite and the receiver. It's the speed of light. For satellite clock deviation, It is the receiver's clock difference. It is the carrier wavelength. It is the ambiguity of an integer number of weeks. For ionospheric delay, For tropospheric delay, Error caused by tidal effects, For measuring noise.

[0107] RTK uses differential observations, while D-RTK uses a double-difference method to eliminate errors caused by satellite clock differences and receiver clock differences. The single difference between stations refers to the difference between the observations from the base station (r) and the rover station (m) of the same (s) satellite. The equation for the single difference between stations is: (14)

[0108] in, The difference in carrier phase values ​​between the base station and the rover station is measured for the (s) satellite. To measure the carrier phase value of the satellite for the mobile station, To measure the carrier phase value of the satellite at the reference station, To measure the geometric distance difference between the base station and the rover station and the satellite(s), The clock difference between the receivers of the base station and the rover station observation(s) satellite. The integer difference between the satellite ambiguities observed by the base station and the rover station. For the difference in ionospheric delay between the base station and the rover station observations (s) of the satellite, For the tropospheric delay difference between the base station and the rover station observations(s) satellite, The noise difference between the base station and the rover station observation(s) satellite measurements is used to measure the noise difference. The carrier wavelength.

[0109] The inter-satellite difference between satellite (s) and satellite (k) is calculated using the difference equation based on the single difference between stations. The inter-satellite double difference equation is as follows:

[0110] (15)

[0111] in, The difference between satellite (s) and satellite (k) is the double difference. The difference in carrier phase values ​​between the base station and the rover station is measured for the (k) satellite. The double difference of the geometric distance between satellite (s) and satellite (k) The integer double difference between the ambiguities between satellite (s) and satellite (k) The double difference in ionospheric delay between satellite (s) and satellite (k) The double difference in tropospheric delay between satellite (s) and satellite (k) The double difference in measurement noise between the (s) satellite and the (k) satellite.

[0112] In the two-difference equation, the main error terms (ionospheric and tropospheric delays) are highly correlated at short baselines (<20km). As the distance between the reference station and the rover decreases, the ionospheric and tropospheric delay errors decrease, and equation (15) simplifies to:

[0113] (16)

[0114] This simplified formula can eliminate errors caused by satellite clock differences and receiver clock differences, thereby improving positioning accuracy.

[0115] S2-3, Bridge Pier Route Planning

[0116] Bridge piers exhibit various cross-sectional geometries (such as rectangular, pointed, circular, single-column, or multi-column), with significant dimensional variations between their upper and lower regions. To obtain high-precision 3D models of the bridge piers, it is crucial to develop effective UAV flight trajectories that conform to the pier's geometry and regional specifications. To address the flight trajectory planning problem, this invention designs a layered, circular flight trajectory method based on standardized elliptical trajectories. The lateral dimension of the pier determines the length of the straight-line segment of the ellipse, while the diameter of the semi-circular segment depends on the longitudinal dimension of the pier and the object distance d. Figure 3 As shown, the flight path shooting points are planned using four elliptical control points, object distance d, and overlap rate parameters.

[0117] Step S3, 3D reconstruction.

[0118] To achieve both global accuracy and local detail in the 3D reconstruction of the bridge pier surface, this invention designs a dual-region strategy that integrates the high-precision absolute positioning of RTK with the flexibility and high resolution of manual flight. This strategy divides the reconstruction process into two regions: an "automatic flight zone with satellite signals" and a "manual flight zone with signal shielding." By fusing data within a unified optimization framework, the limitations of single methods in terms of flight efficiency, positioning signal, coverage, accuracy, and detail are effectively solved.

[0119] In this embodiment, a high-precision sparse point cloud and camera pose are obtained through initial image reconstruction. These poses are constrained by the RTK positioning area. Then, manually acquired images are registered onto the established model. For manually captured flight images, the three-dimensional points precisely measured in the RTK positioning area are used as fixed control points, and the initial pose of the camera is estimated by solving the PnP problem.

[0120] Specifically, the following steps are included:

[0121] S3-1, 3D reconstruction based on the SfM algorithm.

[0122] The 3D reconstruction process in Multi-View SfM integrates geometric constraints, statistical optimization, and incremental strategies. Its basic principle involves simultaneously solving for the camera's intrinsic and extrinsic parameters (motion) and the 3D coordinates of scene feature points. This is achieved by establishing correspondences between 2D images captured from multiple perspectives, while also utilizing the principles of multi-view geometry. The core algorithm consists of four main stages: (i) feature extraction and matching; (ii) initial two-view reconstruction; (iii) incremental registration of the new view using PnP estimation combined with triangulation; and (iv) bundle adjustment optimization.

[0123] S3-1-1, Feature Extraction and Matching.

[0124] Central perspective projection connects points in the three-dimensional world with points in a two-dimensional image. Let the homogeneous coordinates of a three-dimensional point be... The corresponding homogeneous coordinates of the image points are The central perspective projection equation is:

[0125] (17)

[0126] in, It is a scale factor. It is a camera projection matrix. R is the camera's intrinsic parameter matrix, and R and t are the rotation and translation matrices in the camera's external reference, respectively.

[0127] Including manufacturing errors in the camera and pixel components, the parameter matrix K within the camera is represented as:

[0128] (18)

[0129] in, and These are the focal lengths along the x-axis and y-axis, respectively. For the manufacturing error angle of the pixel component, It is the horizontal offset from the center point of the camera coordinate projection to the center point of the pixel coordinates. It is the vertical offset.

[0130] S3-1-2, Initial two-view reconstruction.

[0131] Initial dual-view reconstruction involves finding the relative relationships between two camera views, given a set of matching points, and points in one view. Corresponding point in another view It must lie on a straight line called the epipolar line. The relationship between epipolar geometry and the fundamental matrix is ​​as follows:

[0132] (19)

[0133] in, It is a fundamental matrix that encodes all the geometric information between the two cameras.

[0134] S3-1-3, Incremental registration of new views using PnP estimation combined with triangulation.

[0135] Based on the fundamental matrix F and the camera projection matrix M, the initial 3D point cloud is recovered through triangulation. Meanwhile, given the two-dimensional projection point set corresponding to the three-dimensional point set. The new camera pose parameters are estimated using the PnP (Perspective-n-Point) method. The formula for minimizing the reprojection error of PnP is:

[0136] (20)

[0137] in, These are the observed pixel coordinates of the j-th 3D point on the i-th image. For the j-th 3D point, This is the camera projection function.

[0138] S3-1-4, Beam adjustment optimization.

[0139] Bundle adjustment (BA) is a key global optimization step in the SfM algorithm, where all 3D point coordinates and camera pose parameters are jointly optimized to minimize the total reprojection error. The cost function relationship is as follows:

[0140] (twenty one)

[0141] Where m is the number of cameras and n is the number of 3D points. It is an indicator variable; if the point If it is visible in the i-th camera, the value is 1; if it is not visible, the value is 0. It is the Euclidean distance function.

[0142] The relationship for solving the target using bundle adjustment is as follows:

[0143] (twenty two)

[0144] in, It is the projection matrix of the i-th camera.

[0145] S3-2, 3D reconstruction integrating RTK image information.

[0146] Real-time kinematics (RTK) technology, based on real-time dynamic carrier phase difference measurement, provides high-precision absolute positioning. This is combined with traditional vision-based 3D reconstruction methods, overcoming the inherent limitations of purely vision-based approaches, including scale ambiguity, cumulative errors, and lack of absolute positioning. Through carrier phase difference processing, the camera center coordinates are determined with centimeter-level accuracy, denoted as... The scaling factor between pixel coordinates and world coordinates is:

[0147] ;

[0148] (twenty three)

[0149] in, It is the physical distance from the projection center point of the i-th image to the projection center point of the (i+1)-th image in the world coordinate system. It is the scaling factor between pixel coordinates and world coordinates. Let be the three-dimensional coordinates of the projection center point of the i-th image in the world coordinate system. Let be the three-dimensional coordinates of the projection center point of the (i+1)th image in the world coordinate system. Let be the three-dimensional coordinate position of the projection center of the (i+1)th image in the camera coordinate system in the traditional SfM algorithm. Let be the three-dimensional coordinates of the projection center of the i-th image in the camera coordinate system in the traditional SfM algorithm.

[0150] The camera projection center point set and RTK point set are decentralized. In the camera coordinate system, When transformed to the camera coordinate system, the optical center coordinate value is always 0 (because it is the origin of the coordinate system). Combining the formula (17) for transforming the optical center point of the UAV 3D world camera to the pixel coordinate system, the translation matrix relationship can be obtained as follows:

[0151] Discretize the camera projection center point set and RTK point set to establish their relative positions. In the camera coordinate system, when transformed to this coordinate system, the inherent analytical value of the optical center coordinates is zero because they represent the origin of the coordinate system. Combining formula (17) to transform the optical center point of the UAV 3D world camera to the pixel coordinate system, the translation matrix relationship is as follows:

[0152] (twenty four)

[0153] in, Let be the rotation matrix in the external reference of the camera for the i-th image. Let be the translation matrix in the external reference of the camera for the i-th image. Let be the three-dimensional coordinate position of the projection center of the i-th image in the camera coordinate system in the traditional SfM algorithm.

[0154] Substituting formula (24) into formula (20) eliminates the influence of the translation matrix, and formula (20) simplifies to:

[0155] (25)

[0156] Incorporating RTK position constraints into the BA cost function, the BA cost function equation containing ERK position constraints is as follows:

[0157] (26)

[0158] in, It is the expected minimum total cost. and The covariance matrices of visual observation and RTK observation are respectively. These are RTK weight parameters.

[0159] The dual-region strategy of 3D reconstruction, which combines manual and RTK methods, combines global accuracy with local detail, automation with flexibility, and is suitable for complex scenarios such as bridge pier inspection.

[0160] Step S4, based on the dual-backbone improved YOLOv 11.

[0161] This invention selects YOLOv11 as the basic model for bridge defect identification due to its strong real-time performance and wide adaptability in structural image analysis. However, the original YOLOv11 architecture has limited ability to perceive small-scale targets (such as well-defined cracks) in actual bridge defect detection, resulting in insufficient feature extraction and frequent omissions of small targets. To address these limitations, this invention designs an enhanced aerial target detection algorithm combining a dual-backbone network architecture, such as... Figure 4 As shown.

[0162] This invention presents an enhanced aerial target detection algorithm with a dual-backbone network architecture. First, it implements a parallel dual-path backbone to achieve complementary extraction of spatial details and deep semantic features, thereby enhancing the model's feature representation capability for defect regions. Second, it incorporates an enhancement layer specifically designed for small target detection into the feature fusion network, significantly improving the multi-scale perception accuracy of subtle cracks. Finally, it replaces the traditional detection head with a progressive feature fusion AFPNHead structure. Optimizing the multi-scale feature transmission path further improves the performance of small target defect reconstruction.

[0163] The YOLOv11 target detection neural network identifies two-dimensional damage information such as cracks and spalling defects. Based on the three-dimensional model reconstructed in step S3 and the camera pose, it is mapped to three-dimensional space to realize the visualization and integrated display of bridge pier surface damage on the three-dimensional model.

[0164] Experimental verification:

[0165] The bridge piers used in this experiment are the ramp piers of an overpass in City A and the mainline bridge piers in City B. The piers in City A are single-column piers, while those in City B are double-column piers. UAVs equipped with RTK modules were deployed for data collection. Using the WGS84 coordinate system, centimeter-level (≤1.5cm) positioning accuracy was achieved through multi-constellation GNSS (GPS, GLONASS, BeiDou, Galileo) fusion. Complete UAV platform specifications and camera parameters are detailed in Tables 1 and 2 to ensure repeatable data quality for high-precision structural assessment.

[0166] Table 1 Unmanned Aerial Vehicle Platform Parameters

[0167]

[0168] Table 2 Camera Parameters

[0169]

[0170] To ensure full coverage of the bridge pier modeling, the lower area was automatically flown by UAVs using a standard elliptical flight path, with a flight path overlap rate of 90% and a side overlap rate of 80%, and the flight direction was horizontal. The upper area was flown by a manual approximate standard elliptical flight path using PID cascade control. The modeling accuracy SSD was controlled by setting the close-range shooting distance D. Considering the site conditions, the object distance from the target to the bridge pier surface was set to 5m. Substituting the 5m D value and camera parameters into formula (3) yielded the theoretical SSD value of the bridge pier image, which was equal to 0.00068m / pixel. The height range of the manual PID-controlled flight area was generally controlled by the range of satellite signal obstruction, and the vertical flight height range of the PID in the background was 2~4m.

[0171] A 3D model of the No. 1 bridge pier in City A was reconstructed using a dual-region photogrammetry method. The model was generated based on 834 and 249 photographs captured in the lower and upper regions, respectively. The acquired photographic data was processed using Context Capture (CC) software, and a fused 3D model was established using the SfM algorithm with dual-region constraints. Figure 5 As shown, under the constraint of a reference image for RTK localization, the dual-region SfM strategy demonstrates the ability to integrate untagged images into a cohesive point cloud model.

[0172] Test pier #2 is a bridge pier in City B. The 3D model was created using 458 photographs taken of the lower area and 289 photographs taken of the upper area. The collected data and photographs were imported into the third-party software CC (Context Capture) to analyze the quality of the 3D model. Simultaneously, a point cloud model was created in RealityCapture software to verify the dimensional accuracy. Figure 6 It can be seen that even when the width of the main beam is large compared to the pier, the test bridge pier can still obtain good projection optical center point positioning solutions from UAV-captured images. The dual-region fusion flight strategy can effectively solve the problem of the difficulty in establishing full-size 3D models of bridge piers with large width-to-height ratios.

[0173] (1) Visual verification of cracks.

[0174] To evaluate the specific detection performance of the improved model, the proposed dual-backbone YOLO model is compared with the original object detection model. To more clearly and intuitively compare the detection performance of each model in different scenarios, three drone-captured crack images from the test dataset are randomly selected for detection and recognition. The test combines the detection of each model as follows: Figure 7As shown, in the same scene, the YOLOv11 improved with dual backbones and AFPNHead achieves higher accuracy in visualized crack images compared to the original YOLOv11 and other single-improvement YOLOv11 neural networks. Across all test set images, the YOLOv11 improved with dual backbones and AFPNHead exhibits fewer overlapping detection boxes due to multiple detections of the same target compared to other models.

[0175] The YOLOv11 model, enhanced and optimized with dual-backbone and multi-layer small target detection heads, showed some improvement in crack identification, with an accuracy increase of 0.5%, recall increase of 2.8%, mAP50 increase of 4.2%, and mAP50-95 increase of 0.7%. In the visual verification experiment of cracks, the method of the present invention significantly improved the confidence level of surface crack defect identification.

[0176] (2) Verification of the quality of the three-dimensional model.

[0177] The quality analysis of the 3D model is a crucial factor in evaluating the effectiveness of the dual-region flight strategy. The reprojection error of the 3D model of the tested pier #1 was 0.61 pixels. Based on RTK image localization information and minimum reprojection calculations, the average positioning error of the image projection center point of this model was found to be 0.02602m. Figure 8 As shown, markings were placed on-site using a 1m steel ruler, and the measured distance for modeling was 1.001m. The experimental results demonstrate that the 3D model established using the dual-region flight strategy has good dimensional accuracy, achieving millimeter-level dimensional accuracy in modeling.

[0178] The reprojection error of the 3D model of pier #2 tested was 0.41 pixels, indicating that the projection pixel error on the photographic dataset was small. Figure 9 As shown, the positioning accuracy of the connection points in the core area of ​​the model is between 0.00018 and 0.033 m, demonstrating the accurate three-dimensional spatial positioning of the point cloud data. The resolution of the connection points in the core modeling area is 0.00066 to 0.0033 m / pixel, verifying the model's high-resolution capability.

[0179] The median resolution of the UAV close-range photogrammetry 3D modeling using the dual-region flight strategy is 0.00068~0.0033 m / pixel, the projection point error is 0.41~0.61 pixels, and the median position uncertainty is 0.00018~0.033 m. The dimensional error of the reconstructed model is verified to be within the millimeter level by placing distance markers on two bridge pier 3D models. This effectively ensures the accuracy and feasibility of the UAV technology for reconstructing 3D bridge pier models.

[0180] The above embodiments are merely preferred examples of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Therefore, any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention should fall within the protection scope of the present invention.

Claims

1. A method for identifying bridge pier damage based on dual-region flight and dual-backbone optimization networks, characterized in that, Includes the following steps: Step S1: Use drone close-range photogrammetry to capture the surface structure of the bridge piers, using a controlled vertical shooting distance; Step S2: A dual-zone flight control strategy is adopted: the bridge pier is divided into a lower automatic flight zone and an upper manual control zone. PID cascade manual flight is used in the upper zone of the bridge pier, and RTK positioning automatic flight is used in the lower zone. Images are acquired along a standardized elliptical layered path. Step S3, 3D Reconstruction: The dual-region images obtained in step S2 and their corresponding RTK positioning coordinate information are fused together, and the motion recovery structure 3D reconstruction is performed using the bundle adjustment algorithm with RTK position constraints to generate a 3D model of the bridge pier. The motion recovery structure three-dimensional reconstruction in step S3 consists of four stages: Feature extraction and matching: Feature point detection and description are performed on all images acquired from dual regions, and the correspondence between images is established through feature matching algorithms; Initial two-view reconstruction: Select two images with high overlap for initial two-view reconstruction, and calculate the fundamental matrix through epipolar geometry; New incremental registration of views using PnP estimation combined with triangulation: Generate an initial 3D sparse point cloud through triangulation, and estimate new camera pose parameters using the PnP method. Bundle adjustment optimization: Construct a cost function with RTK constraints to optimize camera pose and 3D point coordinates; Step S4, Intelligent Surface Damage Recognition: Construct and train a dual-backbone enhanced YOLOv11 target detection neural network, input the bridge pier surface image collected in step S2 into the network, and identify and locate cracks and spalling defects in the image.

2. The pier damage identification method based on dual-region flight and dual-backbone optimization network according to claim 1, characterized in that, In step S2, the automatic flight control of the lower region specifically involves: the UAV receiving the real-time position located by RTK as feedback, comparing it with the preset standardized elliptical layered path waypoint coordinates, and driving the UAV to fly automatically along the preset path and triggering the camera to take pictures through a cascaded PID controller.

3. The pier damage identification method based on dual-region flight and dual-backbone optimization network according to claim 2, characterized in that, In step S2, the manual flight control of the upper region specifically includes: Sensor fusion state estimation: Data from visual odometry, inertial measurement unit, ultrasonic or infrared ranging sensors are fused, and state estimation is performed through extended Kalman filter to output the position, velocity and attitude information of the UAV; Cascade PID control: It adopts a four-loop cascade control structure of position loop, velocity loop, attitude loop and angular velocity loop. The desired control command is input, and the controller calculates and outputs the final control quantity based on the state estimation information, so as to control the UAV in the absence of satellite navigation signal.

4. The pier damage identification method based on dual-region flight and dual-backbone optimization network according to claim 1, characterized in that, The standardized elliptical layered path is as follows: the lateral dimension of the pier determines the length of the straight segment of the ellipse, the diameter of the semicircular segment depends on the longitudinal dimension of the pier and the object distance, and the flight path shooting points are planned through 4 elliptical control points, object distance d and overlap rate parameters.

5. The pier damage identification method based on dual-region flight and dual-backbone optimization network according to claim 1, characterized in that, In step S3, the RTK position constraints are incorporated into the cost function, as shown below: ; in, It is the expected minimum total cost. and The covariance matrices of visual observation and RTK observation are respectively. These are RTK weight parameters. It is an indicator variable. These are the observed pixel coordinates of the j-th 3D point on the i-th image. Let K be the j-th 3D point, and K be the parameter matrix within the camera. Let be the rotation matrix in the external reference of the camera for the i-th image. Let m be the 3D coordinate position of the projection center of the i-th image in the camera coordinate system in the traditional SfM algorithm, where m is the number of cameras and n is the number of 3D points. Let be the three-dimensional coordinates of the projection center point of the i-th image in the world coordinate system.

6. The pier damage identification method based on dual-region flight and dual-backbone optimization network according to claim 1, characterized in that, The dual-backbone enhanced YOLOv11 target detection neural network is a dual-backbone parallel feature extraction network combined with a detection head structure optimized for small targets.

7. The pier damage identification method based on dual-region flight and dual-backbone optimization network according to claim 6, characterized in that, The dual-backbone parallel feature extraction network is as follows: on the YOLOv11 baseline model, a parallel backbone network branch is added. The original backbone is used to extract deep semantic features, and the newly added branch adopts a shallow structure.

Citation Information

Patent Citations

  • Bridge pier structure surface crack detection and quantification method based on Gaussian sputtering

    CN121053564A

  • Automatic space coordinate registration method and device based on 3D Gaussian splash model

    CN121120727A