Bridge damage positioning and quantifying method based on unmanned aerial vehicle panoramic unfolding and TCFormer driving

By using a panoramic deployment of UAVs and a bridge damage localization and quantification method driven by TCFormer, the problem of identifying and quantitatively assessing multiple types of damage in bridge inspection has been solved, achieving efficient and accurate bridge damage detection and health assessment.

CN122072960APending Publication Date: 2026-05-22YANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YANGZHOU UNIV
Filing Date
2026-02-04
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing bridge inspection technologies suffer from high costs, low efficiency, insufficient accuracy, and difficulty in identifying and quantitatively assessing multiple types of damage. In particular, they lack accuracy under complex lighting and texture interference and lack standardized health assessment.

Method used

By employing a UAV panoramic deployment and TCFormer-driven approach, the entire process from image acquisition to damage quantification is automated through matrix-style flight path planning, panoramic image construction, multi-type damage segmentation, component-level localization, and health status assessment.

Benefits of technology

It significantly improves the efficiency and accuracy of bridge inspection, enables precise identification and sub-millimeter-level quantification of various types of damage, provides objective health assessment results, and is applicable to the inspection of different types of concrete bridges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122072960A_ABST
    Figure CN122072960A_ABST
Patent Text Reader

Abstract

The invention discloses a bridge damage positioning and quantifying method based on unmanned aerial vehicle panoramic unfolding and TCFormer driving, and the method comprises the steps: 1) unmanned aerial vehicle image collection: employing a matrix type flight path planning strategy, and dividing a bridge bottom into a plurality of independent collection regions; 2) panorama construction: performing single-component three-dimensional reconstruction on the acquired image through a motion recovery structure SfM and a multi-view stereoscopic vision MVS technology; 3) performing multi-class damage segmentation: performing semantic segmentation on the panoramic image cutting area based on a lightweight TCFormer model; 4) component-level positioning: establishing a standardized component coordinate system and an index system, and dividing the components into five types; according to the method, a health evaluation system including PMCI single component scoring and PCCI whole-span comprehensive scoring is constructed, full-process automation from image acquisition, damage identification, spatial positioning, geometric quantification to health evaluation is achieved, and the efficiency, precision and standardization level of bridge detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bridge structure inspection and health monitoring technology, and in particular to a bridge damage location and quantification method based on UAV panoramic unfolding and TCFormer driving. Background Technology

[0002] With the aging of transportation infrastructure and the increase in traffic load, the safety of bridge structures is receiving increasing attention. Early detection and quantitative assessment of damage are crucial to ensuring the long-term service of bridges. Traditional bridge inspection mainly relies on manual inspections and static monitoring equipment, which suffers from high costs, long cycles, and low efficiency. Furthermore, the inspection results are easily affected by subjective factors, making it difficult to meet the requirements of efficiency, objectivity, and digitalization in modern bridge operation and maintenance.

[0003] In recent years, the development of unmanned aerial vehicle (UAV) technology and computer vision technology has provided new pathways for bridge inspection. UAVs can flexibly acquire image information of areas that are difficult to reach by traditional methods, such as the bottom of bridges, significantly improving inspection coverage and safety. However, existing technologies still have many shortcomings: First, most studies focus on the identification of a single damage type, with limited generalization ability, making it difficult to cope with complex scenarios where multiple types of damage coexist, such as bridge cracks, spalling, and holes. Second, the complex texture of concrete surfaces, variable lighting, and shadow interference lead to insufficient recognition accuracy of existing segmentation models. Third, there is no stable GPS signal under the bridge, making it difficult for traditional positioning methods to achieve accurate spatial correspondence between damage and structural entities. Fourth, damage quantification relies on target calibration or manual ranging, which is easily affected by environmental factors and prone to errors, making it difficult to achieve sub-millimeter level accuracy measurement. Fifth, existing technologies lack effective integration of identification results with health assessment systems, making it difficult to achieve the transformation from "visual recognition" to "structural diagnosis."

[0004] Therefore, developing an integrated detection technology that combines the ability to identify multiple types of damage, high-precision positioning, sub-millimeter-level quantification accuracy, and standardized health assessment has become a pressing technical challenge in the field of intelligent bridge inspection. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a bridge damage localization and quantification method based on UAV panoramic deployment and TCFormer-driven approach. This method automates the entire process from image acquisition, damage identification, spatial localization, geometric quantization to health assessment, thereby improving the efficiency, accuracy, and standardization of bridge inspection.

[0006] The objective of this invention is achieved as follows: a bridge damage localization and quantification method based on UAV panoramic deployment and TCFormer driving, comprising the following steps:

[0007] 1) UAV image acquisition: A matrix-style flight path planning strategy is adopted to divide the bridge bottom into multiple independent acquisition areas, maintain a certain degree of overlap between adjacent fields of view, acquire multi-angle images of the bridge bottom, and control exposure parameters to ensure image consistency.

[0008] 2) Panoramic image construction: The acquired images are reconstructed into 3D single components using SfM (Structure from Motion) and MVS (Multi-view Stereo Vision) techniques. Orthogonal projection views of each component are extracted, and preprocessed using the multi-scale Retinex algorithm and color statistical normalization. Then, the images are stitched together according to the structural topology to form a panoramic unfolded image. Component naming rules and a pixel-to-physical scale mapping model are established.

[0009] 3) Multi-type damage segmentation: Based on the lightweight TCFormer model, semantic segmentation of the panoramic image cropping area is performed. Multi-scale features are extracted in parallel through a three-channel convolution module to achieve fine segmentation of six types of damage: cracks, peeling, holes, exposed reinforcement, water leakage and efflorescence.

[0010] 4) Component-level positioning: Establish a standardized component coordinate system and index system, divide components into five categories: cap beam, bottom plate, flange plate, web plate and wet joint, use letter-number combination to identify the partition, and realize the corresponding mapping of damage in panoramic image and physical component through image naming;

[0011] 5) Sub-millimeter quantization: Based on the component design dimensions and imaging parameters, determine the spatial resolution, and calculate the damage length, width, and area geometric parameters using the pixel-to-actual-size conversion formula;

[0012] 6) Health status assessment: Based on the regional bridge inspection standards, a health evaluation system is constructed that includes PMCI single component scoring and PCCI whole span comprehensive scoring. The deduction value is calculated by combining the weight of damage type, size index and location weight to determine the technical condition level of the bridge.

[0013] Furthermore, step 1) specifically includes:

[0014] Step 1-1) Data collection area division: Based on the structural characteristics of the bridge bottom, the data collection area is divided into four independent areas: web, flange, bottom plate and wet joint, and the spatial boundaries and structural range of each area are clearly defined.

[0015] Steps 1-2) Flight path planning: A matrix flight strategy of "partition scanning + overlap compensation" is adopted. Based on the sequence of waypoints in the three-dimensional coordinate system, a shortest path optimization model is established to balance flight efficiency and area coverage.

[0016] Flight path planning uses a sequence of waypoints in a three-dimensional coordinate system as the waypoint set, and achieves a balance between flight efficiency and coverage by solving the shortest path optimization model.

[0017]

[0018] Among them, C(P) i ) indicates that at position P i Image coverage index at location, θ v The camera's field of view;

[0019] During the acquisition process, fixed aperture and shutter speed parameters are used to control exposure consistency, ensuring the stitchability and geometric comparability of image data under varying lighting conditions;

[0020] Steps 1-3) Overlap control: Set the overlap between adjacent fields of view to 60%–80% to ensure the feature matching quality of multi-view images;

[0021] Steps 1-4) Introduction of flight constraints: Environmental point cloud constraints and obstacle boundary conditions are incorporated during the planning phase;

[0022] Steps 1-5) Unify the acquisition parameters: Fix the camera aperture and shutter speed parameters to control the consistency of exposure.

[0023] Furthermore, step 2) specifically includes:

[0024] Step 2-1) Generation of sparse point cloud based on motion-recovered structure SfM;

[0025] Feature point extraction and matching: The SIFT algorithm is used to extract feature points from multi-view images and generate descriptors. Euclidean distance is used to match the same feature points in adjacent images.

[0026] False match removal: The RANSAC algorithm is used to remove false match feature points;

[0027] Camera pose determination: Based on binocular vision constraints, derived from the essential matrix. To establish a pixel coordinate association, the expression is:

[0028]

[0029] Where x and x' are the normalized pixel coordinates of two adjacent images, respectively, and the essential matrix is... It consists of the camera's relative rotation R and translation vector t;

[0030] Sparse Point Cloud Generation and Optimization: Sparse point clouds are generated by combining triangulation. The triangulation formula is as follows:

[0031]

[0032] Where K is the intrinsic parameter matrix, λ is the scale factor, x is the normalized pixel coordinate, and X is the pixel coordinate; then the pose and point cloud coordinates are optimized by bundle adjustment.

[0033] Step 2-2) Dense 3D modeling based on MVS;

[0034] Dense Depth Map Generation: After obtaining the camera pose and sparse point cloud, the multi-view stereo vision MVS algorithm is used to solve for consistency constraints on pixels across multiple views; its core objective is to minimize the photometric error function and optimize the pixel depth d to ensure consistent brightness across different viewpoints.

[0035]

[0036] Where ρ(·) is the robust loss, I i Let I be the grayscale value of the i-th image, and I0 be the reference image selected in the multi-view reconstruction. This is the depth projection function, used to solve for pixel depth and generate a dense depth map;

[0037] 3D model construction: Dense point cloud is obtained by fusing depth maps, mesh model is generated by Poisson reconstruction, and realistic surface texture is given to the model through texture mapping to complete 3D reconstruction;

[0038] Steps 2-3) Extracting orthogonal projection views of components;

[0039] Local coordinate system establishment: Define the normal vector n for each component. i To determine the projection direction, a local coordinate system for the bridge is established;

[0040] Plane fitting: Perform plane fitting on the mesh model to obtain the plane equation. ;

[0041] Projection transformation: The 3D surface point cloud is mapped onto a 2D plane through projection transformation to obtain the frontal texture map of each component:

[0042]

[0043] Where u is the horizontal pixel coordinate of the image, v is the vertical pixel coordinate, and K is the camera intrinsic parameter matrix. This is the camera extrinsic parameter matrix;

[0044] Steps 2-4) Image preprocessing:

[0045] Size normalization: Resample all orthogonal views to a uniform resolution so that the pixel size corresponding to a unit length meets the standard scaling factor s, eliminating spatial resolution deviations caused by differences in camera viewing distance;

[0046] Illumination enhancement: All orthogonal views are resampled to a uniform resolution so that the pixel size per unit length meets the standard scaling factor s; secondly, the multi-scale retina enhancement algorithm MSR is used for illumination enhancement; the calculation process is expressed as follows:

[0047]

[0048] in, Let i be a Gaussian convolution kernel of size i, and * denote the convolution operation. It is the weight of the i-th reference scale; the weight is adaptively allocated according to the average brightness difference of the image, and the smaller the brightness difference, the larger the weight.

[0049] Color consistency matching: Let the mean and variance of the color channels of the current image be µ... c , σ c The mean and variance of the color channels of the reference image are respectively µ t , σ t The color matching result is:

[0050]

[0051] Among them, I c This represents the color channel values ​​of the matched image;

[0052] This linear transformation can adjust the brightness distribution and color saturation to minimize color differences in the splicing area;

[0053] Steps 2-5) Panoramic stitching and calibration:

[0054] View stitching: Match SIFT keypoints in the overlapping areas of orthogonal views of various components, and estimate the affine transformation matrix H. ij :

[0055]

[0056] After aligning multiple views step by step, the differences in stitching seams are eliminated by pyramid fusion; a local geometric compensation model is introduced to reproject and correct the texture of the projection boundary area.

[0057] Component Indexing and Dimensioning: Referencing bridge design drawings and on-site survey data, define component naming rules; based on the actual spatial length L of the component... real Pixel distance L in the unfolded image pixel and the Euclidean distance d in the 3D model 3D Construct a pixel-to-physical scale conversion model:

[0058]

[0059] The scaling factor S allows the two-dimensional detection results to be directly mapped to the real physical space, achieving a precise conversion from the image domain to the structural domain.

[0060] Furthermore, step 3) specifically includes:

[0061] Step 3-1) Design of TCFormer model architecture, wherein the TCFormer model includes a feature encoding layer, an inflated Transformer module layer, an efficient self-attention mechanism layer, and a decoding fusion layer;

[0062] The input image is processed by overlapping patch embedding through a feature encoding layer to extract initial features;

[0063] Each dilated Transformer module layer contains an efficient self-attention module and a three-channel convolutional module. The feature update process is as follows:

[0064]

[0065] Among them, EffAttn(·) represents the efficient self-attention module, and the three-channel convolution module completes multi-scale convolution fusion to realize the spatial enhancement of the Transformer structure;

[0066] Efficient self-attention mechanism layer: Before calculating attention weights, the input features are mapped to a low-dimensional space. The number of information channels is controlled by the dimensionality reduction mapping matrix. The expression for the attention mechanism is:

[0067]

[0068] Where Q, K, and V are the query, key, and value matrices, respectively, and d k For feature dimensions;

[0069] Through the dimension reduction mapping matrix W q W k W v Controlling the number of information channels reduces the computational complexity from O(n^2) to O(n^2). 2 The value is reduced to O(n);

[0070] The features at each stage are upsampled and fused by the decoding fusion layer to output the segmentation results of five types of damage;

[0071] Step 3-2) Design of the three-channel convolutional TCC module:

[0072] Parallel Convolution Branches: Three parallel convolution branches are set up with dilation rates of r1=1, r2=3, and r3=5, respectively. The dilated convolution expression is as follows:

[0073]

[0074] Where y[i] represents the i-th pixel value of the output feature map, and x is the input feature map. Here, K represents the kernel weights, K represents the kernel size, and r represents the dilation rate.

[0075] Activation and Normalization: After each branch is processed by the RPReLU activation function, they are concatenated along the channel dimension, and the module structure is shown in the following equation:

[0076]

[0077] Then, a normalization layer is used to normalize the features and balance the scale, and the output is:

[0078]

[0079] Step 3-3) Dataset Construction and Augmentation:

[0080] Data acquisition: High-resolution images are acquired from multiple angles and distances on various types of concrete structure surfaces using drone platforms and handheld cameras;

[0081] Annotation processing: Pixel-by-pixel semantic segmentation and annotation are performed using professional annotation tools, covering six types of damage: cracks, peeling, holes, exposed reinforcement, efflorescence, and water leakage;

[0082] Data augmentation: Random rotation, horizontal and vertical flipping, color jitter, and lighting variation simulation strategies were introduced to expand the image pool to 4032 images;

[0083] Steps 3-4) Model training configuration: Divide the dataset into training, validation and test sets in a 7:1:2 ratio; run the PyTorch framework in a specific environment; select the AdamW optimizer, set the momentum parameter, introduce L2 regularization to preset the weight decay coefficient and suppress overfitting; introduce an early stopping mechanism, set the initial learning rate to a preset value, and dynamically adjust it using a cosine annealing strategy;

[0084] Steps 3-5) employ a weighted combination of cross-entropy loss and Dice loss. The expression for the loss function is as follows:

[0085]

[0086] Wherein, the loss weight coefficients λ1, λ2, λ1+λ2=1, L CE This represents the cross-entropy loss; the loss function L balances the pixel imbalance problem between different categories.

[0087] Furthermore, step 4) specifically includes:

[0088] Step 4-1) Component Classification and Coordinate System Definition: Bridge components are divided into five categories: Cap, Bottom, Flange, Web, and Wet. Vertically, the components are labeled alphabetically from top to bottom, and horizontally, they are labeled with two-digit numbers from left to right, thus establishing a standardized component coordinate system.

[0089] Step 4-2) Image cropping and naming: Cropping the images of each component partition according to a fixed scale of a certain number of pixels, and using the naming rule of "component type - vertical partition - horizontal partition" to achieve unique identification of the damage location;

[0090] Step 4-3) Location mapping implementation: Through the naming and indexing system, the specific location in the panoramic view of the bridge can be quickly located directly by the image name, and a one-to-one correspondence mapping between "image damage" and "component area" can be established.

[0091] Furthermore, step 5) specifically includes:

[0092] Step 5-1) Determine spatial resolution: Based on the component design dimensions and camera imaging parameters, determine a spatial resolution of 0.78125mm / pixel as the benchmark for converting pixels to actual dimensions;

[0093] Step 5-2) Image size preprocessing: For the original image that is not a multiple of 640, a symmetrical zero-filling strategy is adopted to add an equal amount of black pixels to the image edges so that the width and height of the final image are both integer multiples of 640, while keeping the pixel spacing and relative geometric relationship unchanged.

[0094] Step 5-3) Quantization calculation: Using the pixel-to-actual-size conversion formula:

[0095]

[0096] The pixel-scale parameters of crack width, peeling area, and hole radius in the segmentation results are converted into actual physical quantities; by accurately calculating geometric parameters such as damage length, width, and area, sub-millimeter-level quantitative analysis is achieved.

[0097] Furthermore, step 6) specifically includes:

[0098] Step 6-1) Construct PMCI single component scoring;

[0099] Individual Deduction Value Calculation: The individual damage deduction value fully considers the type, severity, and component location effect of the damage. The calculation model is as follows:

[0100]

[0101] Among them, W type Type weights are set based on the degree of impact of damage on the load-bearing capacity and durability of the component; S metric α is the actual size of the damage, obtained by converting the image segmentation results and physical pixel size; loc Location weights are used to adjust the structural importance of the section where the damage occurs; This is a normalization factor used to convert the magnitude of physical dimensions into a reasonable deduction range.

[0102] Step 6-2) Component Comprehensive Score: The core calculation formula is based on 100 points, and is determined by cumulatively subtracting the deductions caused by all identified damages. The core calculation formula is as follows:

[0103]

[0104] Where N is the total number of damages detected on the component;

[0105] Step 6-3) PCCI overall score;

[0106] Worst-case component score determination: The PCCI is based on the weighted average of the PMCI values ​​of all components, and a "worst-case component correction mechanism" is introduced to reflect the impact of weak links on overall safety.

[0107]

[0108] in, PMCI is the average score for all main beams or components. min For the component with the lowest score, t is the component quantity correction factor;

[0109] The correction factor t is determined based on the number of components. When the number of components is ≤5, the correction factor t is 9.2. When the number of components is greater than 5 and less than 10, the correction factor t is 8.1.

[0110] Step 6-4) Technical Condition Classification: Five levels are defined based on the PCCI score:

[0111]

[0112] Compared with the existing technology, the beneficial effects of the present invention are as follows: 1) Significantly improved detection efficiency: Through the matrix path planning and panoramic deployment technology of UAVs, the bottom of the bridge can be quickly covered, greatly reducing manual intervention. The detection efficiency is more than 5 times higher than that of traditional manual inspection, effectively reducing detection costs and cycle.

[0113] 2) High accuracy in identifying multiple types of damage: The mIoU of the TCFormer model reaches 0.8822, which is better than mainstream models such as ConvNeXt and SegFormer. It can accurately identify six types of bridge damage, effectively solving the identification problem in complex lighting, texture interference and multiple damage coexisting scenarios, and significantly reducing the false detection rate and false detection rate.

[0114] 3) High positioning reliability: The innovative component-level index and coordinate system can achieve accurate correspondence between damage and physical components without relying on GPS signals. The positioning error is controlled within the centimeter level, solving the positioning problem in the absence of stable GPS under the bridge.

[0115] 4) High quantification accuracy: A pixel-to-actual size mapping relationship based on component geometry information is established. Combined with symmetrical zero-fill processing technology, sub-millimeter level damage quantification is achieved. The quantification error meets the actual assessment needs of engineering and provides accurate data support for judging the severity of damage.

[0116] 5) Standardization of health assessment: Construct a PMCI and PCCI scoring system based on regional standards, automatically generate technical condition levels, avoid the subjectivity and experience dependence of manual assessment, and ensure that the assessment results are objective and traceable, providing a scientific basis for maintenance decisions.

[0117] 6) Wide applicability: Applicable to various types of damage detection for different types of concrete bridges, it can be directly embedded into the bridge digital twin system to realize automatic inspection and time-series tracking of defects, and has broad application prospects in the field of inspection of transportation infrastructure such as highway and railway bridges. Attached Figure Description

[0118] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0119] Figure 1 This invention relates to the acquisition of frontal views of components and multi-scale Retinex image enhancement.

[0120] Figure 2 A schematic diagram illustrating the principle of panoramic unfolding of the bridge bottom in this invention.

[0121] Figure 3 The TCFormer network structure diagram of this invention.

[0122] Figure 4 Visual comparison of the various models in this invention.

[0123] Figure 5 The image indexing flowchart of this invention.

[0124] Figure 6 This invention includes data acquisition site and bridge component drawings.

[0125] Figure 7 A panoramic view of the bottom of the bridge in this invention.

[0126] Figure 8 The present invention provides the results of damage identification and quantification at the bottom of bridges.

[0127] Figure 9 This invention provides a panoramic view of the spatial location of damage at the bottom of a bridge.

[0128] Figure 10 The bridge position weight α of this invention loc Schematic diagram of value selection. Detailed Implementation

[0129] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0130] A bridge damage localization and quantification method based on UAV panoramic unfolding and TCFormer-driven approach includes the following steps:

[0131] 1) UAV image acquisition: A matrix-style flight path planning strategy is adopted to divide the bridge bottom into multiple independent acquisition areas, maintain a certain degree of overlap between adjacent fields of view, acquire multi-angle images of the bridge bottom, and control exposure parameters to ensure image consistency; as shown in Figure 6, the left side is the actual scene of data acquisition, and the right side is the bridge component drawings and UAV parameter settings;

[0132] Step 1-1) Data collection area division: Based on the structural characteristics of the bridge bottom, the data collection area is divided into four independent areas: web, flange, bottom plate and wet joint, and the spatial boundaries and structural range of each area are clearly defined.

[0133] In this embodiment, based on the structural design drawings of the No. 2 span of the Grand Canal Bridge on Tianjin Road in Huai'an City, the bridge bottom is divided into four independent data collection areas: web, flange, bottom plate, and wet joint. The spatial boundaries of each area are clearly defined (such as the range of web height and the range of wet joint length).

[0134] Drone parameters and configuration: A multi-rotor drone was selected, equipped with a high-resolution camera, with the following parameters: 82.1° field of view, 24mm equivalent focal length, f / 1.7 aperture, focus distance from 1m to infinity, 4K (3840×2160) video resolution at 30fps, and lens tilt adjustment range of -135° to 80° to ensure that all areas under the bridge can be effectively filmed;

[0135] Steps 1-2) Flight path planning: A matrix flight strategy of "partition scanning + overlap compensation" is adopted. Based on the sequence of track points in the three-dimensional coordinate system, a shortest path optimization model is established to balance flight efficiency and regional coverage. The planned track needs to cover the area under the North 2# span bridge and take into account the boundary conditions of obstacles such as the bridge support structure to avoid collision risks.

[0136] Flight path planning uses a sequence of waypoints in a three-dimensional coordinate system as the waypoint set, and achieves a balance between flight efficiency and coverage by solving the shortest path optimization model.

[0137]

[0138] Among them, C(P) i ) indicates that at position P i Image coverage index at location, θ v The camera's field of view;

[0139] During the acquisition process, fixed aperture and shutter speed parameters are used to control exposure consistency, ensuring the stitchability and geometric comparability of image data under varying lighting conditions;

[0140] Steps 1-3) Overlap control: Set the overlap between adjacent fields of view to 60%–80% to ensure the feature matching quality of multi-view images; in this embodiment, the overlap between adjacent fields of view is controlled at 70%.

[0141] Steps 1-4) Introduction of flight constraints: Environmental point cloud constraints and obstacle boundary conditions are incorporated during the planning phase;

[0142] Steps 1-5) Unified acquisition parameters: Fixed camera aperture and shutter parameters to control exposure consistency; flew along the planned flight path and acquired images, obtaining more than 1,200 high-resolution images of the bridge under the bridge from multiple angles, ensuring no blind spots in acquisition.

[0143] 2) Panoramic Image Construction: The acquired images are reconstructed into 3D for each component using Structure-of-Motion (SfM) and Multi-View Stereo Vision (MVS) techniques. Orthogonal projection views of each component are extracted, and preprocessed using the multi-scale Retinex algorithm and color statistical normalization. These are then stitched together according to structural topology to form a panoramic unfolded image. Component naming rules and a pixel-to-physical-scale mapping model are established, such as... Figure 1 As shown;

[0144] Step 2-1) Generation of sparse point cloud based on motion-recovered structure SfM;

[0145] Feature point extraction and matching: The SIFT algorithm is used to extract feature points from multi-view images and generate descriptors. Euclidean distance is used to match the same feature points in adjacent images. For more than 1,200 images, the SIFT algorithm is used to extract feature points from each image and generate descriptors. Euclidean distance is used to match the same feature points in adjacent images.

[0146] False match removal: The RANSAC algorithm is used to remove false match feature points. In this embodiment, the removal rate is about 15%, and valid matching point pairs are retained.

[0147] Camera pose determination: Based on binocular vision constraints, derived from the essential matrix. To establish a pixel coordinate association, the expression is:

[0148]

[0149] Where x and x' are the normalized pixel coordinates of two adjacent images, obtained by extracting feature points and generating descriptors from multi-view images captured by a drone using the SIFT algorithm, and matching them using Euclidean distance; the essential matrix It consists of the camera's relative rotation R and translation vector t;

[0150] Sparse point cloud generation and optimization: Sparse point clouds are generated by combining triangulation, and then the pose and point cloud coordinates are optimized by bundle adjustment to reduce the cumulative error (the point cloud coordinate error after optimization is ≤0.5cm).

[0151] The trigonometric formula is:

[0152]

[0153] Where K is the intrinsic parameter matrix, λ is the scale factor, x is the normalized pixel coordinate, and X is the pixel coordinate; then the pose and point cloud coordinates are optimized by bundle adjustment.

[0154] Step 2-2) Dense 3D modeling based on MVS;

[0155] Dense Depth Map Generation: After obtaining the camera pose and sparse point cloud, the multi-view stereo vision MVS algorithm is used to solve for consistency constraints on pixels across multiple views; its core objective is to minimize the photometric error function and optimize the pixel depth d to ensure consistent brightness across different viewpoints.

[0156]

[0157] Where ρ(·) is the robust loss, I i Let I be the grayscale value of the i-th image, and I0 be the reference image selected in multi-view reconstruction (usually an image with a better viewing angle and uniform illumination). This is the depth projection function, used to solve for pixel depth and generate a dense depth map;

[0158] 3D model construction: Dense point cloud is obtained by fusing depth maps, mesh model is generated by Poisson reconstruction, and realistic surface texture is given to the model through texture mapping to complete 3D reconstruction;

[0159] Steps 2-3) Extracting orthogonal projection views of components;

[0160] Local coordinate system establishment: Define the normal vector n for each component. i To determine the projection direction, a local coordinate system for the bridge is established;

[0161] Plane fitting: Perform plane fitting on the mesh model to obtain the plane equation. ;

[0162] Projection transformation: The 3D surface point cloud is mapped onto a 2D plane through projection transformation to obtain the frontal texture map of each component:

[0163]

[0164] Where u is the horizontal pixel coordinate of the image, v is the vertical pixel coordinate, and K is the camera intrinsic parameter matrix. This is the camera extrinsic parameter matrix;

[0165] Steps 2-4) Image preprocessing:

[0166] Size normalization: Resample all orthogonal views to a uniform resolution so that the pixel size corresponding to a unit length meets the standard scaling factor s, eliminating spatial resolution deviations caused by differences in camera viewing distance;

[0167] Illumination enhancement: All orthogonal views are resampled to a uniform resolution so that the pixel size per unit length meets the standard scaling factor s; secondly, the multi-scale retina enhancement algorithm MSR is used for illumination enhancement; the calculation process is expressed as follows:

[0168]

[0169] in, Let i be a Gaussian convolution kernel of size i, and * denote the convolution operation. It is the weight of the i-th reference scale; the weight is adaptively allocated according to the average brightness difference of the image, and the smaller the brightness difference, the larger the weight.

[0170] Color consistency matching: Let the mean and variance of the color channels of the current image be µ... c , σ c The mean and variance of the color channels of the reference image are respectively µ t , σ t The color matching result is:

[0171]

[0172] Among them, I c This represents the color channel values ​​of the matched image, where c corresponds to RGB and other color channels.

[0173] This linear transformation can adjust the brightness distribution and color saturation to minimize color differences in the splicing area;

[0174] Steps 2-5) Panoramic stitching and calibration: (e.g.) Figure 2 The diagram shown is a schematic diagram of the panoramic unfolding principle of the bridge bottom, which clearly shows the modeling, unfolding and splicing logic of each component at the bottom of the beam (West Flange, West Web, Bottom, East Web, Wet joints).

[0175] View stitching: Match SIFT keypoints in the overlapping areas of orthogonal views of various components, and estimate the affine transformation matrix H. ij :

[0176]

[0177] After aligning multiple views step by step, the differences in stitching seams are eliminated by pyramid fusion; a local geometric compensation model is introduced to reproject and correct the texture of the projection boundary area.

[0178] Component indexing and dimensional calibration: Referring to bridge design drawings and on-site survey data, define component naming rules (e.g., WF: West Flange, WW: West Web, B: Bottom, EW: East Web, Wet: Wet joints, etc.); based on the actual spatial length L of the component... real (e.g., the actual length of the bottom plate of beam #1 is 10m), pixel distance L in the unfolded diagram. pixel (Corresponding to 12800 pixels) and the Euclidean distance d in the 3D model 3D Construct a pixel-to-physical scale conversion model:

[0179]

[0180] The scaling factor S allows the two-dimensional detection results to be directly mapped to the real physical space, achieving a precise conversion from the image domain to the structural domain.

[0181] The scaling factor S = 0.78125 mm / pixel was calculated, completing the pixel-to-physical scale mapping. Finally, the images were stitched together to form a panoramic view of the south and north cap beams (2560×33920 pixels) and a panoramic view of the bottom unfolded (38400×57433 pixels).

[0182] 3) Multi-class damage segmentation: Semantic segmentation of the panoramic image cropping region is performed based on the lightweight TCFormer model. Multi-scale features are extracted in parallel through a three-channel convolutional module, achieving fine segmentation of six types of damage: cracks, spalling, holes, exposed reinforcement, water leakage, and efflorescence. Figure 3 As shown;

[0183] Step 3-1) Design of TCFormer model architecture, wherein the TCFormer model includes a feature encoding layer, an inflated Transformer module layer, an efficient self-attention mechanism layer, and a decoding fusion layer;

[0184] The input image is processed by overlapping patch embedding through a feature encoding layer to extract initial features;

[0185] Each dilated Transformer module layer contains an efficient self-attention module and a three-channel convolutional module. The feature update process is as follows:

[0186]

[0187] Among them, EffAttn(·) represents the efficient self-attention module, and the three-channel convolution module completes multi-scale convolution fusion to realize the spatial enhancement of the Transformer structure;

[0188] Efficient self-attention mechanism layer: Before calculating attention weights, the input features are mapped to a low-dimensional space. The number of information channels is controlled by the dimensionality reduction mapping matrix. The expression for the attention mechanism is:

[0189]

[0190] Where Q, K, and V are the query, key, and value matrices, respectively, and d k For feature dimensions;

[0191] Through the dimension reduction mapping matrix W q W k W v Controlling the number of information channels reduces the computational complexity from O(n^2) to O(n^2). 2 The value is reduced to O(n);

[0192] The features at each stage are upsampled and fused by the decoding fusion layer to output the segmentation results of five types of damage;

[0193] Step 3-2) Design of the three-channel convolutional TCC module:

[0194] Parallel Convolution Branches: Three parallel convolution branches are set up with dilation rates of r1=1, r2=3, and r3=5, respectively. The dilated convolution expression is as follows:

[0195]

[0196] Where y[i] represents the i-th pixel value of the output feature map, and x is the input feature map. Here, K represents the kernel weights, K represents the kernel size, and r represents the dilation rate.

[0197] Activation and Normalization: After each branch is processed by the RPReLU activation function, they are concatenated along the channel dimension, and the module structure is shown in the following equation:

[0198]

[0199] Then, a normalization layer is used to normalize the features and balance the scale, and the output is:

[0200]

[0201] Step 3-3) Dataset Construction and Augmentation:

[0202] Data acquisition: High-resolution images are acquired from multiple angles and distances on various types of concrete structure surfaces using drone platforms and handheld cameras;

[0203] Annotation processing: Pixel-by-pixel semantic segmentation and annotation are performed using professional annotation tools, covering six types of damage: cracks, peeling, holes, exposed reinforcement, efflorescence, and water leakage;

[0204] Data augmentation: Random rotation, horizontal and vertical flipping, color jitter, and lighting variation simulation strategies were introduced to expand the image pool to 4032 images;

[0205] 4032 valid images were selected from the collected images, and pixel-by-pixel semantic segmentation and annotation were performed using professional annotation tools, covering six types of damage: cracks (yellow), peeling (green), holes (gray), exposed tendons (blue), efflorescence (purple), and leakage (red). During the training phase, data augmentation strategies such as random rotation (0°, 90°, 180°, 270°), horizontal and vertical flipping, color jitter (brightness ±10%, saturation ±15%), and illumination change simulation were introduced. The images were divided into a training set (2822 images), a validation set (403 images), and a test set (807 images) in a 7:1:2 ratio.

[0206] Steps 3-4) Model Training Configuration: Divide the dataset into training, validation, and test sets in a 7:1:2 ratio; run on an NVIDIA RTX PRO 6000 GPU using the PyTorch framework; select the AdamW optimizer (momentum parameters β1=0.9, β2=0.999), introduce L2 regularization with preset weight decay coefficients (weight decay coefficient 0.0001) to suppress overfitting; introduce an early stopping mechanism (terminating training if there is no improvement in validation set performance after 15 consecutive rounds), set the initial learning rate to 0.001, and dynamically adjust it using a cosine annealing strategy; set the batch size to 8 and the number of training rounds to 200;

[0207] Steps 3-5) employ a weighted combination of cross-entropy loss and Dice loss (weight α=0.6) as the loss function expression:

[0208]

[0209] Where the loss weight coefficients λ1, λ2, λ1+λ2=1, L CE Cross-entropy loss is a commonly used basic loss in semantic segmentation, used to measure the difference in probability distribution between the predicted result and the true label; the loss function L balances the pixel imbalance problem between different categories.

[0210] After training, the model achieved an mIoU of 0.8822, a Precision of 0.9429, a Recall of 0.9279, and an F1-score of 0.9350 on the test set.

[0211] like Figure 4 The image shows a visualization comparison of the various models (ConvNeXt, UNet++, TransUNet, DeepLabv3+, SegFormer, TCFormer), including the original images, ground truth (GT), and segmentation outputs of each model. Detailed implementation:

[0212] The panoramic unfolded image was cropped into several image patches at a scale of 640×640 pixels, and input into the trained TCFormer model for semantic segmentation, outputting segmentation results for various types of damage. Comparative results show that the TCFormer model can accurately segment multiple types of damage with blurred edges and significant scale differences, effectively maintaining the integrity and connectivity of small targets (such as fine cracks), outperforming other mainstream models. The DeepLabv3+ model performed the worst, exhibiting problems such as crack breakage, voids in the peeling area, and misclassification of efflorescence, thus verifying the superiority of the TCFormer model. (See Table 1 for details.)

[0213]

[0214] 4) Component-level positioning: Establish a standardized component coordinate system and indexing system, classifying components into five categories: cap beams, bottom plates, flanges, webs, and wet joints. Use alphanumeric combinations to identify these zones, and achieve a mapping between damage in the panoramic view and the actual component through image naming. Figure 5 As shown;

[0215] Step 4-1) Component Classification and Coordinate System Definition: Bridge components are divided into five categories: Cap, Bottom, Flange, Web, and Wet. Vertically, the partitions are identified by alphabetical order (A, B, C, D...) (from top to bottom), and horizontally, the partitions are identified by two-digit serial numbers (01, 02, 03...) (from left to right). A standardized component coordinate system is established, and the image naming rules are shown in Table 2.

[0216]

[0217] Step 4-2) Image cropping and naming: Cropping the images of each component partition according to a fixed scale of a certain number of pixels, and using the naming rule of "component type - vertical partition - horizontal partition" to achieve a unique identification of the damage location; for example, the first partition on the leftmost side of the top layer of the cap beam is named Cap-A-01, and the third partition of area B of the bottom plate is named Bottom-B-03, resulting in more than 1,200 image blocks.

[0218] Step 4-3) Location Mapping Implementation: Through the naming and indexing system, the specific location in the bridge panoramic image can be quickly located directly by the image name, establishing a one-to-one correspondence between "image damage" and "component area". For example, "Web-C-04" corresponds to the 4th partition of the C area of ​​the web. Its position and size in the panoramic coordinates are fixed, realizing a one-to-one correspondence from "image damage" to "component area" without relying on GPS signals.

[0219] like Figure 7 The image shown is a panoramic view of the bridge's underside, displaying the bottom structure of beams 1 through 8 and the spatial distribution of each component; as shown... Figure 9The image shows a panoramic view of the spatial location of damage at the bottom of the bridge. Different colors indicate different damage types (yellow - holes, red - efflorescence, blue - exposed reinforcement, green - spalling, gray - holes, purple - leakage). Specific implementation: The segmentation results of all image blocks were mapped to the panoramic unfolded image according to the indexing system. Composite damage (cases where multiple types of damage exist in a single image) was marked using color gradients and transparent overlays, generating a panoramic view of the damage distribution across the entire bridge. The image clearly shows that hole-type damage is mainly concentrated at the junction of the bottom slab and web; spalling and exposed reinforcement damage mostly appear at the lower edge of the cap beam and the middle of the web; and efflorescence damage is mainly located at wet joints and the edge of the bottom slab, highly consistent with the results of manual inspection.

[0220] 5) Sub-millimeter quantization: Based on the component design dimensions and imaging parameters, the spatial resolution is determined, and the geometric parameters of the damage length, width, and area are calculated using the pixel-to-actual-size conversion formula; for example... Figure 8 The image shown is a result of damage identification and quantification at the bottom of the bridge. The left side is the original image, and the right side is the model output, which includes the damage category, pixel area, and the converted actual size (length, width, area).

[0221] Step 5-1) Determine spatial resolution: Based on the component design dimensions and camera imaging parameters, determine a spatial resolution of 0.78125mm / pixel as the benchmark for converting pixels to actual dimensions;

[0222] Step 5-2) Image size preprocessing: For the original image that is not a multiple of 640, a symmetrical zero-padding strategy is adopted to add an equal amount of black pixels to the image edges so that the width and height of the final image are both integer multiples of 640, while keeping the pixel spacing and relative geometric relationship unchanged;

[0223] Step 5-3) Quantization Calculation: Based on the established spatial resolution of 0.78125mm / pixel, the pixel-to-actual-size conversion formula is used:

[0224]

[0225] The pixel-scale parameters of crack width, peeling area, and hole radius in the segmentation results are converted into actual physical quantities. For example, if the pixel length of a crack in the image is 231px, the actual length after conversion is 231×0.78125mm=180.47mm. If the maximum pixel width is 5px, the actual width after conversion is 5×0.78125mm=3.90625mm (approximately 4mm). If the pixel area of ​​a peeling region is 37765px², the actual area after conversion is 37765×(0.78125mm)²=37765×0.61035mm²≈23050mm²=230.50cm². By accurately calculating the geometric parameters such as damage length, width, and area, sub-millimeter level quantitative analysis is achieved.

[0226] Quantitative result verification: 20 damage points (covering six types of damage) were randomly selected for on-site measurement. The error between the measured value and the model quantification value was ≤0.1mm (sub-millimeter level), which meets the actual engineering assessment requirements.

[0227] 6) Health Status Assessment: Based on regional bridge inspection standards, a health assessment system is constructed that includes PMCI single-component scoring and PCCI whole-span comprehensive scoring. Deductions are calculated by combining damage type weights, dimensional indicators, and location weights to determine the bridge's technical condition level; for example... Figure 10 The diagram shown illustrates the weighting of bridge locations, specifically at mid-span (α). loc =1.2) and beam end (α) loc The weight difference is 1.0, with the mid-span region (high bending moment zone) having a higher weight than the end region. Reference standards: A health evaluation system is constructed based on the "Technical Condition Assessment Standard for Highway Bridges" (JTG / T H21-2011) and the "Maintenance Specification for Highway Bridges and Culverts" (JTG 5120-2021).

[0228] Step 6-1) Construct PMCI single component scoring;

[0229] Individual Deduction Value Calculation: The individual damage deduction value fully considers the type, severity, and component location effect of the damage. The calculation model is as follows:

[0230]

[0231] Among them, W type The type weights are set according to the degree of impact of damage on the load-bearing capacity and durability of the component. In this embodiment, exposed reinforcement is 2.5, cracks are 2.0, spalling is 1.5, holes are 1.2, leakage is 1.0, and efflorescence is 0.25; S metric α is the actual size of the damage, obtained by converting the image segmentation results and physical pixel size; locThe positional weights are 1.2 for the mid-span region and 1.0 for the end region, used to adjust the structural importance of the section where the damage occurs; The normalization factor, set to 0.001, is used to convert physical size magnitudes into reasonable deduction ranges to control the overall rigor of the assessment. For example, a 180mm long crack in the mid-span region would have a single deduction value D. i =2.0×180×1.2×0.001=0.432.

[0232] Step 6-2) Component Comprehensive Score: The core calculation formula is based on 100 points, and is determined by cumulatively subtracting the deductions caused by all identified damages. The core calculation formula is as follows:

[0233]

[0234] Where N is the total number of damages detected on the component;

[0235] The calculations are shown in Table 3:

[0236] Step 6-3) PCCI overall score;

[0237] Worst-case component score determination: The PCCI is based on the weighted average of the PMCI values ​​of all components, and a "worst-case component correction mechanism" is introduced to reflect the impact of weak links on overall safety.

[0238]

[0239] in, PMCI is the average score for all main beams or members. min For the component with the lowest score, t is the component quantity correction factor;

[0240] Correction factor selection: The correction factor t is determined based on the number of components. When the number of beams is ≤ 5, the correction factor t is 9.2. When the number of beams is greater than 5 but less than 10, the correction factor t is 8.1.

[0241] Step 6-4) Technical Condition Classification: Five levels are defined based on the PCCI score:

[0242]

[0243] In this embodiment: PCCI overall score, average score calculation:

[0244]

[0245] Worst component score determined: PMCI min=94.34 (3# beam).

[0246] Correction factor selection: The number of components is 1, and the correction factor t = 8.1.

[0247] Overall score calculation:

[0248] PCCI=96.70−(96.70−94.34)×(8.1 / 10)=96.70−2.36×0.81≈95.89.

[0249] Grading: PCCI=95.89≥95, classified as Class 1 (intact), regular inspections recommended according to specifications.

[0250] The above description of the embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make several improvements and modifications to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A bridge damage localization and quantification method based on UAV panoramic unfolding and TCFormer driving, characterized in that, Includes the following steps: 1) UAV image acquisition: A matrix-style flight path planning strategy is adopted to divide the bridge bottom into multiple independent acquisition areas, maintain a certain degree of overlap between adjacent fields of view, acquire multi-angle images of the bridge bottom, and control exposure parameters to ensure image consistency. 2) Panoramic image construction: The acquired images are reconstructed into 3D single components using SfM (Structure from Motion) and MVS (Multi-view Stereo Vision) techniques. Orthogonal projection views of each component are extracted, and preprocessed using the multi-scale Retinex algorithm and color statistical normalization. Then, the images are stitched together according to the structural topology to form a panoramic unfolded image. Component naming rules and a pixel-to-physical scale mapping model are established. 3) Multi-type damage segmentation: Based on the lightweight TCFormer model, semantic segmentation of the panoramic image cropping area is performed. Multi-scale features are extracted in parallel through a three-channel convolution module to achieve fine segmentation of six types of damage: cracks, peeling, holes, exposed reinforcement, water leakage and efflorescence. 4) Component-level positioning: Establish a standardized component coordinate system and index system, divide components into five categories: cap beam, bottom plate, flange plate, web plate and wet joint, use letter-number combination to identify the partition, and realize the corresponding mapping of damage in panoramic image and physical component through image naming; 5) Sub-millimeter quantization: Based on the component design dimensions and imaging parameters, determine the spatial resolution, and calculate the damage length, width, and area geometric parameters using the pixel-to-actual-size conversion formula; 6) Health status assessment: Based on the regional bridge inspection standards, a health evaluation system is constructed that includes PMCI single component scoring and PCCI whole span comprehensive scoring. The deduction value is calculated by combining the weight of damage type, size index and location weight to determine the technical condition level of the bridge.

2. The bridge damage localization and quantification method based on UAV panoramic deployment and TCFormer driving according to claim 1, characterized in that, Step 1) specifically includes: Step 1-1) Data collection area division: Based on the structural characteristics of the bridge bottom, the data collection area is divided into four independent areas: web, flange, bottom plate and wet joint, and the spatial boundaries and structural range of each area are clearly defined. Steps 1-2) Flight path planning: A matrix flight strategy of "partition scanning + overlap compensation" is adopted. Based on the sequence of waypoints in the three-dimensional coordinate system, a shortest path optimization model is established to balance flight efficiency and area coverage. Flight path planning uses a sequence of waypoints in a three-dimensional coordinate system as the waypoint set, and achieves a balance between flight efficiency and coverage by solving the shortest path optimization model. ; Among them, C(P) i ) indicates that at position P i Image coverage index at location, θ v The camera's field of view; During the acquisition process, fixed aperture and shutter speed parameters are used to control exposure consistency, ensuring the stitchability and geometric comparability of image data under varying lighting conditions; Steps 1-3) Overlap control: Set the overlap between adjacent fields of view to 60%–80% to ensure the feature matching quality of multi-view images; Steps 1-4) Introduction of flight constraints: Environmental point cloud constraints and obstacle boundary conditions are incorporated during the planning phase; Steps 1-5) Unify the acquisition parameters: Fix the camera aperture and shutter speed parameters to control the consistency of exposure.

3. The bridge damage localization and quantification method based on UAV panoramic deployment and TCFormer driving according to claim 1, characterized in that, Step 2) specifically includes: Step 2-1) Generation of sparse point cloud based on motion-recovered structure SfM; Feature point extraction and matching: The SIFT algorithm is used to extract feature points from multi-view images and generate descriptors. Euclidean distance is used to match the same feature points in adjacent images. False match removal: The RANSAC algorithm is used to remove false match feature points; Camera pose determination: Based on binocular vision constraints, derived from the essential matrix. To establish a pixel coordinate association, the expression is: ; Where x and x' are the normalized pixel coordinates of two adjacent images, and the essential matrix is... It consists of the camera's relative rotation R and translation vector t; Sparse Point Cloud Generation and Optimization: Sparse point clouds are generated by combining triangulation. The triangulation formula is as follows: ; Where K is the intrinsic parameter matrix, λ is the scale factor, x is the normalized pixel coordinate, and X is the pixel coordinate; then the pose and point cloud coordinates are optimized by bundle adjustment. Step 2-2) Dense 3D modeling based on MVS; Dense Depth Map Generation: After obtaining the camera pose and sparse point cloud, the multi-view stereo vision MVS algorithm is used to solve for consistency constraints on pixels across multiple views; its core objective is to minimize the photometric error function and optimize the pixel depth d to ensure consistent brightness across different viewpoints. ; Where ρ(·) is the robust loss, I i Let I be the grayscale value of the i-th image, and I0 be the reference image selected in the multi-view reconstruction. This is the depth projection function, used to solve for pixel depth and generate a dense depth map; 3D model construction: Dense point cloud is obtained by fusing depth maps, mesh model is generated by Poisson reconstruction, and realistic surface texture is given to the model through texture mapping to complete 3D reconstruction; Steps 2-3) Extracting orthogonal projection views of components; Local coordinate system establishment: Define the normal vector n for each component. i To determine the projection direction, a local coordinate system for the bridge is established; Plane fitting: Perform plane fitting on the mesh model to obtain the plane equation. ; Projection transformation: The 3D surface point cloud is mapped onto a 2D plane through projection transformation to obtain the frontal texture map of each component: ; Where u is the horizontal pixel coordinate of the image, v is the vertical pixel coordinate, and K is the camera intrinsic parameter matrix. This is the camera extrinsic parameter matrix; Steps 2-4) Image preprocessing: Size normalization: Resample all orthogonal views to a uniform resolution so that the pixel size corresponding to a unit length meets the standard scaling factor s, eliminating spatial resolution deviations caused by differences in camera viewing distance; Illumination enhancement: All orthogonal views are resampled to a uniform resolution so that the pixel size per unit length meets the standard scaling factor s; secondly, the multi-scale retina enhancement algorithm MSR is used for illumination enhancement; the calculation process is expressed as follows: ; in, Let i be a Gaussian convolution kernel of size i, and * denote the convolution operation. It is the weight of the i-th reference scale; the weight is adaptively allocated according to the average brightness difference of the image, and the smaller the brightness difference, the larger the weight. Color consistency matching: Let the mean and variance of the color channels of the current image be µ... c , σ c The mean and variance of the color channels of the reference image are respectively µ t , σ t The color matching result is: ; Among them, I c This represents the color channel values ​​of the matched image; This linear transformation can adjust the brightness distribution and color saturation to minimize color differences in the splicing area; Steps 2-5) Panoramic stitching and calibration: View stitching: Match SIFT keypoints in the overlapping areas of orthogonal views of various components, and estimate the affine transformation matrix H. ij : ; After aligning multiple views step by step, the differences in stitching seams are eliminated by pyramid fusion; a local geometric compensation model is introduced to reproject and correct the texture of the projection boundary area. Component Indexing and Dimensioning: Referencing bridge design drawings and on-site survey data, define component naming rules; based on the actual spatial length L of the component... real Pixel distance L in the unfolded image pixel and the Euclidean distance d in the 3D model 3D Construct a pixel-to-physical scale conversion model: ; The scaling factor S allows the two-dimensional detection results to be directly mapped to the real physical space, achieving a precise conversion from the image domain to the structural domain.

4. The bridge damage localization and quantification method based on UAV panoramic deployment and TCFormer driving according to claim 1, characterized in that, Step 3) specifically includes: Step 3-1) Design of TCFormer model architecture, wherein the TCFormer model includes a feature encoding layer, an inflated Transformer module layer, an efficient self-attention mechanism layer, and a decoding fusion layer; The input image is processed by overlapping patch embedding through a feature encoding layer to extract initial features; Each dilated Transformer module layer contains an efficient self-attention module and a three-channel convolutional module. The feature update process is as follows: ; Among them, EffAttn(·) represents the efficient self-attention module, and the three-channel convolution module completes multi-scale convolution fusion to realize the spatial enhancement of the Transformer structure; Efficient self-attention mechanism layer: Before calculating attention weights, the input features are mapped to a low-dimensional space. The number of information channels is controlled by the dimensionality reduction mapping matrix. The expression for the attention mechanism is: ; Where Q, K, and V are the query, key, and value matrices, respectively, and d k For feature dimensions; Through the dimension reduction mapping matrix W q W k W v Controlling the number of information channels reduces the computational complexity from O(n^2) to O(n^2). 2 The value is reduced to O(n); The features at each stage are upsampled and fused by the decoding fusion layer to output the segmentation results of five types of damage; Step 3-2) Design of the three-channel convolutional TCC module: Parallel Convolution Branches: Three parallel convolution branches are set up with dilation rates of r1=1, r2=3, and r3=5, respectively. The dilated convolution expression is as follows: ; Where y[i] represents the i-th pixel value of the output feature map, and x is the input feature map. Here, K represents the kernel weights, K represents the kernel size, and r represents the dilation rate. Activation and Normalization: After each branch is processed by the RPReLU activation function, they are concatenated along the channel dimension, and the module structure is shown in the following equation: ; Then, a normalization layer is used to normalize the features and balance the scale, and the output is: ; Step 3-3) Dataset Construction and Augmentation: Data acquisition: High-resolution images are acquired from multiple angles and distances on various types of concrete structure surfaces using drone platforms and handheld cameras; Annotation processing: Pixel-by-pixel semantic segmentation and annotation are performed using professional annotation tools, covering six types of damage: cracks, peeling, holes, exposed reinforcement, efflorescence, and water leakage; Data augmentation: Random rotation, horizontal and vertical flipping, color jitter, and lighting variation simulation strategies were introduced to expand the image pool to 4032 images; Steps 3-4) Model training configuration: Divide the dataset into training, validation and test sets in a 7:1:2 ratio; run the PyTorch framework in a specific environment; select the AdamW optimizer, set the momentum parameter, introduce L2 regularization to preset the weight decay coefficient, and suppress overfitting; introduce an early stopping mechanism, set the initial learning rate to a preset value, and dynamically adjust it using a cosine annealing strategy; Steps 3-5) employ a weighted combination of cross-entropy loss and Dice loss. The expression for the loss function is as follows: ; Where the loss weight coefficients λ1 and λ 2, λ1+λ2=1, L CE This represents the cross-entropy loss; the loss function L balances the pixel imbalance problem between different categories.

5. The bridge damage localization and quantification method based on UAV panoramic unfolding and TCFormer driving according to claim 1, characterized in that, Step 4) specifically includes: Step 4-1) Component Classification and Coordinate System Definition: Bridge components are divided into five categories: Cap, Bottom, Flange, Web, and Wet. Vertically, the components are labeled alphabetically from top to bottom, and horizontally, they are labeled with two-digit numbers from left to right, thus establishing a standardized component coordinate system. Step 4-2) Image cropping and naming: Cropping the images of each component partition according to a fixed scale of a certain number of pixels, and using the naming rule of "component type - vertical partition - horizontal partition" to achieve a unique identifier of the damage location; Step 4-3) Location Mapping Implementation: Through the naming and indexing system, the specific location in the panoramic view of the bridge can be quickly located directly by the image name, establishing a one-to-one correspondence between "image damage" and "component area".

6. The bridge damage localization and quantification method based on UAV panoramic deployment and TCFormer driving according to claim 1, characterized in that, Step 5) specifically includes: Step 5-1) Determine spatial resolution: Based on the component design dimensions and camera imaging parameters, determine a spatial resolution of 0.78125mm / pixel as the benchmark for converting pixels to actual dimensions; Step 5-2) Image size preprocessing: For the original image that is not a multiple of 640, a symmetrical zero-filling strategy is adopted to add an equal amount of black pixels to the image edges so that the width and height of the final image are both integer multiples of 640, while keeping the pixel spacing and relative geometric relationship unchanged. Step 5-3) Quantization calculation: Using the pixel-to-actual-size conversion formula: ; The pixel-scale parameters of crack width, peeling area, and hole radius in the segmentation results are converted into actual physical quantities; by calculating the geometric parameters of damage length, width, and area, sub-millimeter-level quantitative analysis is achieved.

7. The bridge damage localization and quantification method based on UAV panoramic deployment and TCFormer driving according to claim 1, characterized in that, Step 6) specifically includes: Step 6-1) Construct PMCI single component scoring; Individual Deduction Value Calculation: The individual damage deduction value fully considers the type, severity, and component location effect of the damage. The calculation model is as follows: ; Among them, W type Type weights are set based on the degree of impact of damage on the load-bearing capacity and durability of the component; S metric α is the actual size of the damage, obtained by converting the image segmentation results and physical pixel size; loc Location weights are used to adjust the structural importance of the section where the damage occurs; This is a normalization factor used to convert the magnitude of physical dimensions into a reasonable deduction range. Step 6-2) Component Comprehensive Score: The core calculation formula is based on 100 points, and is determined by cumulatively subtracting the deductions caused by all identified damages. The core calculation formula is as follows: ; Where N is the total number of damages detected on the component; Step 6-3) PCCI overall score; Worst-case component score determination: The PCCI is based on the weighted average of the PMCI values ​​of all components, and a "worst-case component correction mechanism" is introduced to reflect the impact of weak links on overall safety. ; in, PMCI is the average score for all main beams or members. min For the component with the lowest score, t is the component quantity correction factor; Correction factor selection: The correction factor t is determined based on the number of components. When the number of components is ≤5, the correction factor t is 9.

2. When the number of components is greater than 5 and less than 10, the correction factor t is 8.

1. Step 6-4) Technical Condition Classification: Five levels are defined based on the PCCI score: 。