A Classification Method for Bridge Inspection Components Based on Unmanned Aerial Vehicles

By collecting 3D point cloud data of bridges using drones, performing data augmentation and preprocessing, and combining this with deep learning model training, the problems of scarce samples and high annotation costs in the automatic classification of bridge components have been solved. This has enabled accurate, robust, and efficient classification of various bridge components, adapting to different bridge types and data collection conditions.

CN122090137APending Publication Date: 2026-05-26XIAN INNO AVIATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN INNO AVIATION TECH CO LTD
Filing Date
2026-01-30
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, automatic classification methods for bridge components suffer from problems such as scarce samples, high annotation costs, insufficient model generalization ability, and low processing efficiency. In particular, when faced with various bridge types and complex data acquisition conditions, the classification performance drops significantly.

Method used

The bridge's 3D point cloud data was collected by drones, and data augmentation and preprocessing were performed. Combined with deep learning model training, a systematic data augmentation strategy was adopted to simulate real acquisition interference. Multimodal feature fusion and weighted loss functions were designed to build a robust point cloud semantic segmentation model. The model performance was improved through human-machine collaborative optimization.

Benefits of technology

It achieves accurate, robust and efficient classification of various bridge components, can adapt to different bridge types and data acquisition conditions, reduces annotation costs, improves model generalization ability and processing efficiency, and supports efficient automation of bridge inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090137A_ABST
    Figure CN122090137A_ABST
Patent Text Reader

Abstract

This invention relates to the fields of unmanned aerial vehicles (UAVs) and computer vision technology, and particularly to a UAV-based method for classifying bridge components. The technical problem is that existing point cloud-based automatic classification methods for bridge components suffer from issues such as scarce samples, high annotation costs, insufficient model generalization ability, and low processing efficiency. The technical solution is a UAV-based method for classifying bridge components, comprising a data acquisition step, a data augmentation step, a data preprocessing step, a model training step, and a component classification step. In the model training step, preprocessed 3D point cloud data and its corresponding component category labels are used to train a deep learning-based point cloud semantic segmentation model, enabling the model to learn the feature representations of different bridge components. This invention uses a systematic data augmentation strategy to simulate real-world acquisition interference, combined with an efficient point cloud deep learning model, to achieve accurate, robust, and efficient classification of various bridge components.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of unmanned aerial vehicles (UAVs) and computer vision technology, and in particular to a method for classifying bridge components based on UAVs. Background Technology

[0002] As an important transportation infrastructure, the safety of bridges needs to be ensured through regular inspections. Traditional bridge inspection mainly relies on bridge inspection vehicles, scaffolding, or manual climbing, which has problems such as low efficiency, high cost, traffic disruption, and safety risks. In recent years, drone technology has been introduced into the field of bridge inspection due to its flexibility, efficiency, and lack of terrain limitations. By being equipped with LiDAR or oblique photography equipment, drones can quickly acquire high-precision three-dimensional point cloud models of bridges.

[0003] However, the core step in transforming point cloud data collected by drones into structured information suitable for structural assessment lies in accurately classifying different components (such as main beams, bridge towers, cables, and hangers) within the point cloud. Currently, point cloud-based automatic bridge component classification methods are applicable when dealing with diverse bridge structures (beam bridges, arch bridges, cable-stayed bridges, suspension bridges, etc.), each with complex component compositions. Acquiring high-quality point cloud data covering multiple bridge types and operating conditions is inherently costly, and performing detailed point-by-point manual annotation requires significant expertise and time, resulting in a severe shortage of training samples suitable for supervised learning, especially for slender components (cables and hangers) in cable-stayed and suspension bridges. Furthermore, the actual point cloud data collected is subject to limitations imposed by drones. Factors such as human-machine flight trajectory, sensor installation posture, ambient lighting, and occlusion can lead to problems such as uneven density, noise, and viewpoint changes. Existing methods often overfit to specific datasets, and their classification performance deteriorates significantly when applied to new bridges or under different acquisition conditions. There is a lack of data preprocessing and augmentation methods that can effectively simulate these real-world scene variations and enhance model robustness. Moreover, bridge point cloud data is usually large-scale and unstructured. Traditional point cloud segmentation algorithms or early deep learning models have high computational costs when processing such data, making it difficult to meet the speed requirements of practical engineering. At the same time, the scale of bridge components varies greatly (from the macroscopic main tower to the small cables), requiring the model to have multi-scale feature learning capabilities.

[0004] Therefore, to address the above problems, we propose a method for classifying bridge components based on unmanned aerial vehicles (UAVs). This method uses a systematic data augmentation strategy to simulate real-world data acquisition interference and combines it with an efficient point cloud deep learning model to achieve accurate, robust, and efficient classification of various bridge components. Summary of the Invention

[0005] To overcome the problems of scarce samples, high annotation costs, insufficient model generalization ability, and low processing efficiency in existing point cloud-based automatic classification methods for bridge components.

[0006] The technical solution of this invention is: a method for classifying bridge components based on unmanned aerial vehicle (UAV) inspection, comprising the following steps:

[0007] S1: Data acquisition steps: Collect 3D point cloud data of the bridge using the lidar equipment carried by the drone;

[0008] S2: Data augmentation step: Augment the three-dimensional point cloud data to generate diverse training samples;

[0009] S3: Data preprocessing steps: Perform structured preprocessing on the enhanced 3D point cloud data to generate an intermediate data format that is easy for the model to read and calculate;

[0010] S4: Model training steps: Using the preprocessed 3D point cloud data and its corresponding component category labels, train a deep learning-based point cloud semantic segmentation model so that the model learns the feature representations of different bridge components.

[0011] S5: Component classification steps: Input the 3D point cloud data of the bridge to be classified into the trained point cloud semantic segmentation model, and output the bridge component category to which each point belongs.

[0012] Preferably, in the S2 data augmentation step, the augmentation process includes spatial transformation of the three-dimensional point cloud data, which is used to simulate the spatial attitude differences of the point cloud caused by sensor installation deviation or changes in the flight attitude of the UAV.

[0013] Preferably, the spatial transformation includes random rotation transformation around spatial coordinate axes, and the rotation angle around any coordinate axis is limited to a preset small angle range.

[0014] Preferably, in the S2 data augmentation step, the augmentation process further includes downsampling the three-dimensional point cloud data, the downsampling being used to simulate the point cloud density differences caused by changes in the UAV's flight speed or sampling distance.

[0015] Preferably, the sampling ratio of the downsampling is randomly selected within a predefined continuous interval.

[0016] Preferably, the S3 data preprocessing step specifically includes: performing downsampling on the three-dimensional point cloud data and constructing a spatial index structure for accelerating neighborhood queries.

[0017] Preferably, the spatial index structure is a KD tree.

[0018] Preferably, in the S4 model training step, the point cloud semantic segmentation model is a network model that uses a random sampling strategy for downsampling and includes a local feature aggregation module.

[0019] Preferably, in the S4 model training step, a weighted loss function is used to optimize the model, and the weighted loss function includes at least the cross-entropy loss function.

[0020] Preferably, before the S1 data acquisition step, the method further includes:

[0021] Labeling steps: Manually label the collected original bridge 3D point cloud data, assign corresponding bridge component category labels to the points in the point cloud, and form a training dataset; the bridge component categories include at least one or more of background, main beam, bridge column, and cable tower, and further include one or more of stay cable, main cable, suspension cable, anchor, and cable clamp, depending on the bridge type.

[0022] Preferably, in the S1 data acquisition step, a high-resolution orthophoto of the bridge and the three-dimensional point cloud data are acquired simultaneously; the method further includes:

[0023] S6: Multimodal Feature Fusion Step: Using the high-resolution orthophoto, extract the two-dimensional texture and contour features of the bridge components through an image semantic segmentation model. Then, perform cross-modal alignment and fusion of these features with the geometric and spatial structural features learned from the three-dimensional point cloud by the point cloud semantic segmentation model to generate enhanced component classification results.

[0024] Preferably, the method further includes a continuous model performance optimization loop, which includes:

[0025] S7: Uncertainty Quantification and Active Recommendation Step: After the component classification step in S5, the uncertainty of each point or each local region in the classification result is quantified by using the classification probability output by the point cloud semantic segmentation model.

[0026] S8: Key Sample Recommendation Step: Based on a preset uncertainty threshold or active learning strategy, automatically filter out the bridge component point cloud fragments that are most uncertain in model classification and most likely to contain classification errors or new features from the large amount of unlabeled point cloud data generated by actual inspections.

[0027] S9: Human-machine collaborative annotation and model iteration steps: Submit the recommended key samples to experts for rapid verification or fine annotation, and add the newly annotated data to the training dataset to trigger incremental learning or model fine-tuning in the model training step S4, thereby achieving closed-loop evolution of the classification model.

[0028] The beneficial effects of this invention are:

[0029] 1. This invention employs a small-angle rotation and random downsampling enhancement strategy to simulate sensor installation deviations and flight parameter variations. Without significantly increasing annotation costs, it expands the diversity and realism of training data, effectively improving the model's generalization ability and robustness under complex real-world acquisition conditions. By constructing a KD-tree index during the preprocessing stage, the costly neighborhood search computation is completed in advance. Combined with efficient random sampling and local feature aggregation models such as RandLANet, it enables rapid processing of large-scale bridge point cloud data, meeting the efficiency requirements of engineering applications. Through the design of a systematic classification system covering key components of major bridge types, and the optimization of model training using a weighted loss function, it ensures high-precision and balanced classification performance from large components such as main beams and bridge towers to slender components such as stay cables and suspension cables. This achieves accurate, robust, and efficient classification of various bridge components.

[0030] 2. By introducing multimodal feature fusion, this invention breaks through the information bottleneck of single laser point cloud data; by fusing the rich texture information of high-resolution images, the system obtains reliable supplementary discrimination features when facing sparse point cloud components (such as slings photographed from a distance) caused by excessive distance or occlusion, effectively solving the performance degradation problem of pure geometric methods in scenarios with insufficient feature extraction, making fine-grained perception of the surface state of components possible.

[0031] 3. The method of this invention establishes a continuous optimization cycle for model performance, which can proactively identify its own cognitive weaknesses and accurately guide human experts to intervene, forming an efficient human-machine collaborative knowledge growth closed loop. This not only exponentially reduces the annotation cost required for large-scale promotion in practice, but also enables the system to continuously learn and adapt to bridges of different regions and types during use, and even learn new structural features as the bridge ages, realizing the self-adaptation and lifelong learning capabilities of the classification model. Attached Figure Description

[0032] Figure 1 The diagram shows a detailed flowchart of the method for classifying bridge components based on unmanned aerial vehicles (UAVs) according to the present invention.

[0033] Figure 2 The diagram shown is a simplified flowchart of the method for classifying bridge components based on unmanned aerial vehicle (UAV) inspection according to the present invention.

[0034] Figure 3 The diagram shown illustrates the classification effect of the UAV-based bridge inspection component classification method for overpass beam bridge components according to the present invention.

[0035] Figure 4 The diagram shows the classification effect of the cable-stayed bridge components based on the UAV bridge inspection component classification method of the present invention. Detailed Implementation

[0036] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0037] Example 1

[0038] Please see Figure 1 , Figure 2 , Figure 3 and Figure 4 This invention provides an embodiment: a method for classifying bridge components based on unmanned aerial vehicles (UAVs), comprising the following steps:

[0039] S1: Data acquisition steps: Collect 3D point cloud data of the bridge using the lidar equipment carried by the drone;

[0040] S2: Data augmentation step: Augment the 3D point cloud data to generate diverse training samples;

[0041] S3: Data preprocessing steps: Perform structured preprocessing on the enhanced 3D point cloud data to generate an intermediate data format that is easy for the model to read and calculate;

[0042] S4: Model training steps: Using the preprocessed 3D point cloud data and its corresponding component category labels, train a deep learning-based point cloud semantic segmentation model so that the model learns the feature representations of different bridge components.

[0043] S5: Component classification steps: Input the 3D point cloud data of the bridge to be classified into the trained point cloud semantic segmentation model, and output the bridge component category to which each point belongs.

[0044] Furthermore, in the S2 data augmentation step, the augmentation process includes spatial transformation of the 3D point cloud data. This spatial transformation is used to simulate the spatial attitude differences of the point cloud caused by sensor installation deviations or changes in the UAV's flight attitude. Specifically, the spatial transformation is used to accurately simulate the small angular deviations caused by the installation of the lidar sensor on the UAV platform, or the high-frequency, small-amplitude attitude jitter caused by factors such as airflow during UAV flight. This transformation is achieved by applying a random rotation matrix to the point cloud in 3D space, thereby increasing the diversity of the training data in spatial attitude without changing the essential geometry of the components, and improving the robustness of the model to differences in acquisition perspective and sensor installation.

[0045] Furthermore, the spatial transformation includes random rotation transformations around spatial coordinate axes, and the rotation angle around any coordinate axis is limited to a preset small angle range; the spatial transformation is specifically a random rotation transformation around the X, Y, and Z axes of the spatial rectangular coordinate system; in order to conform to the physical reality and avoid generating unreasonable distorted data, the rotation angle around any coordinate axis is limited to a preset small angle range, such as [-2°, 2°]; the setting of this range is based on the analysis of the rigid connection structure and flight vibration characteristics of a typical UAV LiDAR system, ensuring that the augmented data introduces effective noise while maintaining the authenticity of the spatial relationship of bridge components.

[0046] Furthermore, in the S2 data augmentation step, the augmentation process also includes downsampling the 3D point cloud data. Downsampling is used to simulate the differences in point cloud density caused by changes in the UAV's flight speed or sampling distance. When the flight speed is faster or the distance is farther, the point cloud will be relatively sparse; conversely, it will be denser. By introducing this density variation in training, the model can be forced to learn features that are insensitive to point cloud density, thereby enhancing its generalization ability under various acquisition parameters.

[0047] Furthermore, the downsampling ratio is randomly selected within a predefined continuous interval, such as [0.7, 1.0]; a ratio of 1.0 represents maintaining the original density, while 0.7 represents randomly retaining 70% of the original points. This method of random sampling within a continuous interval can generate a variety of samples with smooth density variations, which can better cover the continuous density distribution in real-world scenes than fixed sampling rates.

[0048] Furthermore, the S3 data preprocessing steps specifically include: downsampling the 3D point cloud data to control the number of points input to the model and balance accuracy and computational burden; at the same time, constructing a spatial index structure to accelerate neighborhood queries; since the point cloud model needs to frequently query the neighboring points of each point to aggregate local features during training, constructing an efficient index in advance can decouple this computation from the expensive online training cycle and improve the overall training and inference speed.

[0049] Furthermore, the spatial index structure is a KD tree; a KD tree is a binary tree structure used to partition k-dimensional data space, which is particularly suitable for range search and nearest neighbor search; in point cloud processing, after constructing a KD tree, all neighboring points within a certain radius around a given point can be found quickly (average complexity O(log N)), providing key support for subsequent local feature extraction.

[0050] Furthermore, in the S4 model training step, the point cloud semantic segmentation model is a network model that uses a random sampling strategy for downsampling and includes a local feature aggregation module, such as the RandLANet network. The random sampling strategy of this model avoids complex computational overhead and can efficiently process large-scale point clouds. Its local feature aggregation module explicitly encodes the geometric and contextual information of each point and its local neighborhood through multilayer perceptron (MLP) and attention mechanisms, thereby effectively capturing discriminative features of components of different scales, from slender cables to large bridge towers.

[0051] Furthermore, in the S4 model training step, a weighted combination of loss functions is used to optimize the model. The weighted combination of loss functions includes at least the cross-entropy loss function. The cross-entropy loss function is the standard loss for classification tasks, used to measure the difference between the class probability distribution predicted by the model and the true label distribution. It can be further combined with loss functions based on set similarity, such as Lovasz-Softmax loss. The latter can directly optimize commonly used metrics for segmentation tasks, such as intersection-over-union (IoU), and is especially helpful in improving the classification performance of components that account for a small proportion in the point cloud (such as slings), thus alleviating the class imbalance problem.

[0052] Furthermore, prior to the S1 data acquisition step, the following steps are also included:

[0053] Labeling steps: The collected original 3D point cloud data of the bridge is manually labeled, and the points in the point cloud are assigned corresponding bridge component category labels to form a training dataset. The bridge component categories include at least one or more of the following: background, main beam, bridge column, and pylon. Depending on the bridge type, they may further include one or more of the following: stay cable, main cable, suspension cable, anchorage, and cable clamp. The systematic classification system ensures that the model output has clear engineering semantics and can directly serve subsequent health detection and evaluation.

[0054] Furthermore, in the S1 data acquisition step, high-resolution orthophotos of the bridge and the three-dimensional point cloud data are acquired simultaneously.

[0055] The method also includes a multimodal feature fusion step: using the high-resolution orthophoto, the two-dimensional texture and contour features of the bridge components are extracted through the image semantic segmentation model, and the features are cross-modal aligned and fused with the geometric and spatial structural features learned from the three-dimensional point cloud by the point cloud semantic segmentation model to generate enhanced component classification results;

[0056] Among them, laser point clouds excel at depicting precise geometric shapes, but have limitations in recognizing surface materials, colors, and subtle textures; while high-resolution images are rich in texture information; for example, the texture of the anti-corrosion coating on the main cable, water stains or crack patterns on the surface of the bridge tower are clearly distinguishable in the images; by accurately registering the 3D point cloud with the 2D image in space, and using a feature fusion network (such as through an attention mechanism), the texture features of the image can be "attached" to the corresponding 3D points; this allows the model to not only determine whether it is the main cable based on the "shape", but also to further distinguish cables of different materials or states by combining the "texture", providing additional and powerful discrimination criteria when facing components with sparse point clouds and indistinct geometric features, significantly improving the accuracy of classification.

[0057] Furthermore, the method also includes a continuous model performance optimization loop, which includes uncertainty quantification and proactive recommendation step S7, key sample recommendation step S8, and human-machine collaborative annotation and model iteration step S9. In step S7, the model not only outputs classification labels but also its prediction confidence (probability). The uncertainty of the prediction can be quantified by calculating entropy or coefficient of variation. In step S8, the system automatically scans massive amounts of daily inspection data to locate areas that cause the model to hesitate (such as new ancillary facilities, components around special damage areas, and connections between different components). Step S9 then pushes these issues to human experts, improving the input-output ratio of expert annotation and obtaining the data that best improves model performance with minimal human cost. After new data is injected, the model adapts quickly through incremental learning, enabling it to continuously learn new bridge types, new components, and even new degradation patterns in practical applications, thus achieving autonomous evolution of classification capabilities.

[0058] Through the above steps, this invention utilizes a small-angle rotation and random downsampling enhancement strategy to simulate sensor installation deviations and flight parameter changes. This expands the diversity and realism of training data without significantly increasing annotation costs, effectively improving the model's generalization ability and robustness under complex real-world acquisition conditions. By constructing a KD-tree index during the preprocessing stage, the costly neighborhood search computation is completed in advance. Combined with efficient random sampling and local feature aggregation models such as RandLANet, rapid processing of large-scale bridge point cloud data is achieved, meeting the efficiency requirements of engineering applications. By designing a systematic classification system covering key components of major bridge types and optimizing model training using a weighted loss function, high-precision and balanced classification performance is ensured, ranging from large components such as main beams and bridge towers to slender components such as stay cables and suspension cables. This achieves accurate, robust, and efficient classification of various bridge components. Furthermore, by introducing multi-mode... By fusing dynamic features, this invention breaks through the information bottleneck of single laser point cloud data. By fusing the rich texture information of high-resolution images, the system obtains reliable supplementary discriminative features when facing sparse point cloud components (such as slings photographed from a distance) caused by excessive distance or occlusion. This effectively solves the performance degradation problem of pure geometric methods in scenarios with insufficient feature extraction, making fine-grained perception of the surface state of components possible. The method of this invention establishes a continuous optimization loop for model performance, which can proactively discover its own cognitive weaknesses and accurately guide human experts to intervene, forming an efficient human-machine collaborative knowledge growth closed loop. This not only exponentially reduces the annotation cost required for large-scale promotion in practice, but also allows the system to continuously learn and adapt to bridges of different regions and types during use, and even learn new structural features as the bridge ages, realizing the adaptive and lifelong learning capabilities of the classification model.

[0059] Example 2

[0060] Optionally, this embodiment provides a basic implementation process for a bridge inspection component classification method based on unmanned aerial vehicles (UAVs).

[0061] First, perform the S1 data acquisition step; select a multi-rotor UAV platform equipped with a lightweight LiDAR scanner (such as Livox Avia); plan the UAV flight path to ensure that the path covers the entire main structure of the bridge to be inspected and maintains a relative distance of 10-50 meters from the bridge surface to obtain point cloud data with sufficient accuracy; collect the original three-dimensional point cloud data of the bridge through autonomous flight of the UAV. The data is usually stored in LAS or PCD format and includes the three-dimensional coordinates (X,Y,Z) and reflection intensity information of each point.

[0062] Next, the S2 data augmentation step is performed; the collected raw point cloud dataset (training set portion) is loaded into memory; in order to increase the diversity of training data, online augmentation processing is performed on each frame or each batch of training point clouds; in this embodiment, the augmentation processing includes, but is not limited to, basic geometric transformations such as random translation and small-scale scaling, in order to simulate the small uncertainties in the data acquisition process.

[0063] Next, the S3 data preprocessing step is performed; the enhanced point cloud data is normalized; for example, the point cloud coordinates are normalized so that they are distributed within a relatively standard spatial range; then, the processed point cloud and its corresponding ground truth labels (during the training phase) are converted into a data format that can be efficiently read by deep learning frameworks (such as PyTorch or TensorFlow), for example, by building a data loader for point cloud patches.

[0064] Then, perform the S4 model training step; construct a deep learning-based point cloud semantic segmentation model, such as a simplified PointNet++ network; input the preprocessed training data into the model for training; the model learns the mapping relationship from point cloud features to component category labels through multiple forward and backward propagation iterations; supervised learning is performed using the cross-entropy loss function during training.

[0065] Finally, the S5 component classification step is executed. For new, unlabeled bridge point cloud data (test set or actual inspection data), the S1 acquisition and S3 preprocessing (excluding enhancement) are performed similarly to the training phase. Then, the data is input into the final model trained in the S4 step. The model performs forward inference on each point in the input point cloud, outputs the probability distribution of each point belonging to each component category, and takes the category corresponding to the highest probability as the classification result of that point, thereby completing the component-level semantic segmentation of the entire bridge point cloud.

[0066] Example 3

[0067] Optionally, this embodiment further elaborates on the specific implementation of the data augmentation steps based on Embodiment 2.

[0068] In this embodiment, the S2 data augmentation step is specifically designed to simulate two main sources of variation in actual UAV inspections.

[0069] Simulated Spatial Attitude Variation Process: To simulate the point cloud angle deflection caused by minor installation deviations or high-frequency vibrations during flight of the lidar, random rotation transformations are applied to the point cloud around the Z-axis (vertical direction) and around the X and Y axes (horizontal directions); let the coordinates of a point in the original point cloud be... The applied rotation matrix is Then the coordinates of the transformed point Rotation matrix The result is obtained by multiplying the rotation angles α, β, and γ about the Z, Y, and X axes in sequence: ;in, Let represent the transformation matrix for rotation around the Z-axis by an angle α; where the rotation angles α, β, and γ are not arbitrary values, but are uniformly and randomly sampled from a strictly defined small angle range; in this embodiment, is set , , This parameter setting is based on the fact that the connection between the UAV and the lidar is usually rigid, and the attitude angle changes generated during actual flight are mainly high-frequency small-amplitude vibrations, and large-angle deflection does not conform to physical reality; this small angle transformation can effectively improve the robustness of the model to small changes in the acquisition perspective without distorting the main geometry of the bridge.

[0070] Simulating point cloud density variation: To simulate the differences in point cloud density caused by variations in flight speed and distance, the point cloud is randomly downsampled; a given downsampling ratio is used. Its value is randomly selected from a predefined continuous interval. In this embodiment, it is set as follows: For a frame containing The point cloud of points is downsampled randomly and uniformly, selecting points from it. Keep one point and discard the rest. When, the point cloud density remains unchanged; when At that time, the point cloud density dropped to 70% of its original value. This continuous ratio random sampling, compared to a few fixed ratios, can more continuously and comprehensively simulate various density situations that may occur in the real world, forcing the model to learn features that are insensitive to point cloud density.

[0071] During training, each frame of the point cloud input to the network will independently and randomly undergo a combination of the two transformations mentioned above, thereby virtually expanding the limited dataset several times and including real-world scene variations, thus improving the generalization ability of the final classification model.

[0072] Example 4

[0073] Optionally, this embodiment further elaborates on the specific implementation methods of efficient preprocessing and model training based on embodiment 2 or 3.

[0074] In this embodiment, the S3 data preprocessing step is optimized to improve the processing efficiency of large-scale point clouds; specifically, it includes:

[0075] Downsampling and Index Building Process: Before model training begins, the original training point cloud undergoes offline downsampling preprocessing, such as using voxelization downsampling to unify the point cloud to a scale of approximately 50,000 points to control computational load. Specifically, a KD-tree (k-dimensional tree) spatial index is built for each frame of downsampled point cloud. A KD-tree is a binary tree data structure used to organize points in k-dimensional space. Its construction process recursively divides the space along the dimension with the largest data variance, forming hyperrectangular regions, until the number of points within a sub-region falls below a threshold. After construction, a query for a point with a radius of... The average time complexity of operations on all points within a spherical neighborhood can be reduced from that of brute-force search. Reduce to During training, point cloud data and its pre-computed KD-tree index are loaded together. When the model needs to query a local neighborhood (such as when calculating local features), it directly calls the KD-tree for a fast range search, thereby moving a large amount of computation from the training loop to the preprocessing stage and accelerating the training and inference process.

[0076] In this embodiment, the S4 model training step employs a specific network architecture and optimization strategy:

[0077] Model selection: RandLANet was chosen as the point cloud semantic segmentation model; the innovation of this model lies in:

[0078] Random sampling: In each layer of the network, random sampling is used instead of complex farthest point sampling (FPS) or voxelization, which significantly reduces the computational overhead and enables it to efficiently process large-scale point clouds;

[0079] Local feature aggregation: This module explicitly encodes the geometric and contextual information of each point and its K nearest neighbors obtained by fast querying in the KD tree through a multilayer perceptron (MLP) and attention mechanism;

[0080] This enables the model to effectively capture the distinguishing features of components of different scales, from slender stay cables to massive bridge towers.

[0081] Loss function design: To optimize model performance, especially to improve the classification effect of small components (such as slings), this embodiment adopts a weighted combination loss function: ;

[0082] Representing cross-entropy loss, it is the standard multi-class classification loss; for a classification problem with C classes, the loss for a single point is calculated as follows: ,in It is an indicator function (1 if the true class of the point is equal to c, 0 otherwise). It is the probability that the point belongs to category c as predicted by the model;

[0083] Representing Lovasz-Softmax loss, it is a loss function based on submodulus optimization that can directly optimize the intersection-over-union ratio (IoU), a core metric for segmentation tasks, and is particularly good at handling class imbalance problems;

[0084] and These are the weighting coefficients that balance the two losses, and in this embodiment, they can be set to 1.0 and 0.5 respectively.

[0085] Training hyperparameters: The optimizer used is Adam, the initial learning rate is set to 0.01, and a cosine annealing strategy is used for decay; the batch size is set to 6, and the training epochs are 100.

[0086] Example 5

[0087] Optionally, this embodiment integrates the technical features of the foregoing embodiments and further illustrates the complete closed-loop process from data preparation to application, which is a comprehensive preferred implementation.

[0088] First, perform the annotation process; use professional point cloud processing software (such as CloudCompare) to perform 3D visualization and interactive annotation of the raw bridge point cloud collected by the drone; based on bridge structural engineering knowledge, establish a systematic component category labeling system; this system is hierarchical and scalable.

[0089] General-purpose components: applicable to most bridge types, including background (non-bridge structures, such as ground, vegetation), main beams, and bridge columns (or piers).

[0090] Bridge-specific components:

[0091] For cable-stayed bridges, add pylons and stay cables;

[0092] For suspension bridges, add towers, main cables, suspenders, anchorages, and cable clamps;

[0093] Ancillary facilities: Streetlights, guardrails, maintenance tracks, etc. can be further marked.

[0094] During annotation, the operator uses the polygon selection or brush tool in the software to manually select the point cloud areas belonging to the same component and assign corresponding category labels, ultimately generating a point cloud file with label information.

[0095] Then, steps S1 to S5 are executed sequentially; wherein, data augmentation in S2 adopts the small-angle rotation and random downsampling strategy of Example 3; preprocessing in S3 adopts the KD tree construction method of Example 4; and model training in S4 adopts the RandLANet network architecture, weighted combined loss function and hyperparameter settings of Example 4.

[0096] Finally, in the S5 component classification step, the trained model can automatically segment new bridge point clouds. For example, given a point cloud of a suspension bridge, the model can automatically output whether each point belongs to the background, main beam, tower, main cable, or suspension cable. The classification results can be saved as a color point cloud file, with different colored points representing different components for visualization. It can also generate structured reports to count the number of point clouds and spatial location of various components, providing accurate input data for subsequent component-level damage identification, deformation analysis, and inspection path planning.

[0097] Example 6

[0098] Optionally, based on any one of embodiments 2 to 5, this embodiment further elaborates on how to integrate the images and point cloud data collected simultaneously by the UAV to achieve more robust and refined component classification.

[0099] In this embodiment, the S1 data acquisition step is upgraded to multi-sensor synchronous acquisition; the UAV is equipped not only with a lidar but also with a high-resolution visible light camera; the two are synchronized by a hardware synchronization trigger or a high-precision time synchronization module (such as a PPS signal) to ensure data acquisition time alignment; flight planning must ensure sufficient overlap between the fields of view of the camera and the lidar; after acquisition, dense 3D point cloud (PCD format) of the bridge and highly overlapping sequence images (JPG / RAW format) are obtained respectively.

[0100] The specific implementation is as follows:

[0101] Data registration: First, the image and point cloud are spatially registered with high precision using indirect or direct methods; a preferred implementation method is:

[0102] a. Generate high-density 3D mesh models or sparse point clouds from sequential images using structure-of-motion (SfM) technology;

[0103] b. Register this mesh / sparse point cloud with the precise point cloud scanned by LiDAR using the ICP (Iterative Closest Point) algorithm to obtain the accurate transformation matrix. ;

[0104] Using this transformation matrix, the pixel coordinate system of each image can be associated with the world coordinate system of the 3D laser point cloud through the collinearity equation, that is, to find the corresponding pixel coordinates of each 3D point on at least one image.

[0105]

[0106] in, For pixel coordinates, For the camera intrinsic parameter matrix, The extrinsic parameter matrix from the lidar coordinate system to the camera coordinate system (obtained through registration). The coordinates of the laser point are 3D.

[0107] Dual-branch feature extraction:

[0108] Point cloud branch: Similar to Example 3, the RandLANet network is used to process the point cloud and extract the depth geometric features of each point. ;

[0109] Image Branch: A two-dimensional semantic segmentation network (such as DeepLabv3+ or HRNet) is used to process the registered orthophoto or the original image; for each laser point Based on the above coordinate transformation, its position in the image is found. The corresponding pixel position on Furthermore, two-dimensional texture features at that location are extracted from the feature map of the corresponding layer of the image segmentation network using bilinear interpolation. .

[0110] Feature alignment and fusion: aligning features from two modalities. and To achieve fusion; one specific implementation method is: firstly, Projected onto the interface through a fully connected layer The same feature dimensions are used; then, an attention-guided feature fusion module is employed; this module computes an attention weight map. ,in For the Sigmoid function, For splicing operations, Features of the projected image; final fused features ,in This indicates element-wise multiplication; this attention mechanism allows the network to dynamically decide whether to rely more on geometric or texture features for each point (or each local region).

[0111] Classification decision: merging the features The input is fed into a shared classifier (fully connected layer), and the output is the final component category probability distribution. This approach allows the model to determine the "main cable" based on both its cylindrical geometric features and its image features of "dark gray with parallel steel wire texture". It has a significant advantage in identifying components that are occluded by shadows, resulting in missing point clouds, or that have similar geometric shapes but different materials.

[0112] Example 7

[0113] Optionally, this embodiment further constructs an intelligent classification system capable of self-evaluation, proactive learning, and continuous evolution, based on any of the above embodiments.

[0114] This method establishes a long-running model optimization loop in addition to the conventional S1-S5 (or S1-S6) process, specifically including additional steps S7, S8, and S9:

[0115] S7: Uncertainty Quantification: In the component classification step of S5, the model (such as RandLANet) outputs a probability vector for each point belonging to each category. ,in The total number of categories; this step calculates the uncertainty of the classification prediction for each point; a widely used and effective metric is prediction entropy, the higher the entropy value, the more uncertain the model is about the classification of that point (the more uniform the probability distribution); in addition to the point level, the average entropy of each component instance or local block obtained by point cloud segmentation can also be calculated as the uncertainty score for that region.

[0116] S8: Key Sample Recommendation: The system maintains an "unlabeled pool" that stores a large amount of classified but unverified point cloud data and their uncertainty scores generated during daily inspections; an active learning strategy is used to select the most valuable samples from this pool; one implementation method is:

[0117] Uncertainty-based sampling: Periodically (e.g., weekly), select the top K point cloud fragments with the highest average uncertainty scores from the pool (e.g., each fragment is all points within a fixed-radius sphere extracted around a high-entropy point).

[0118] Diversity-based sampling: To ensure the diversity of recommended samples and avoid duplication, clustering methods (such as K-Means clustering of fragment features) can be used, and then the sample with the highest uncertainty can be selected from different clusters.

[0119] Finally, the system automatically generates a "list of tasks to be reviewed" that includes the 3D spatial location (such as center point coordinates) of these "difficult" fragments, snapshot images, and the current classification results of the model.

[0120] S9: Human-computer collaborative annotation and model iteration:

[0121] Bridge inspection experts can view recommended segments through a dedicated review interface; the interface displays the 3D point cloud of the segment (colored with model-predicted categories) and the corresponding real image side by side; experts can quickly confirm the correct classification or correct incorrect classification labels; this targeted review is far more efficient than labeling random massive amounts of data.

[0122] After the review is completed, the newly generated "fragment-truth value" paired data is added to the original training dataset; then, the system triggers an incremental training; the specific implementation can be as follows: load the pre-trained model weights as pre-training parameters, use the mixed data of the original dataset and the new dataset, and fine-tune it for a limited number of rounds (e.g., 20 rounds) with a low learning rate (e.g., 0.001); this process consumes far less computing resources than training from scratch, but can effectively allow the model to absorb new knowledge, correct original errors, and generalize to new similar scenarios.

[0123] When the system first inspects a new type of composite beam bridge, it may be uncertain about the classification of its unique component connection methods. After these uncertain segments are recommended to experts for annotation, the model can accurately identify such new structures in the next iteration. With long-term operation, the knowledge base and model performance accumulated by the system will continue to grow, eventually forming a highly specialized and highly adaptive intelligent agent for classifying bridge components.

[0124] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A method for classifying bridge components based on unmanned aerial vehicle (UAV) inspection, characterized in that: Includes the following steps: S1: Data acquisition steps: Collect 3D point cloud data of the bridge using the lidar equipment carried by the drone; S2: Data augmentation step: Augment the three-dimensional point cloud data to generate diverse training samples; S3: Data preprocessing steps: Perform structured preprocessing on the enhanced 3D point cloud data to generate an intermediate data format that is easy for the model to read and calculate; S4: Model training steps: Using the preprocessed 3D point cloud data and its corresponding component category labels, train a deep learning-based point cloud semantic segmentation model so that the model learns the feature representations of different bridge components. S5: Component classification steps: Input the 3D point cloud data of the bridge to be classified into the trained point cloud semantic segmentation model, and output the bridge component category to which each point belongs.

2. The method for classifying bridge components based on unmanned aerial vehicle (UAV) inspection according to claim 1, characterized in that: In the S2 data augmentation step, the augmentation process includes spatial transformation of the three-dimensional point cloud data. The spatial transformation is used to simulate the spatial attitude differences of the point cloud caused by sensor installation deviations or changes in the flight attitude of the UAV.

3. The method for classifying bridge components based on unmanned aerial vehicle (UAV) inspection according to claim 2, characterized in that: The spatial transformation includes random rotation transformations around spatial coordinate axes, and the rotation angle around any coordinate axis is limited to a preset small angle range.

4. The method for classifying bridge components based on unmanned aerial vehicle (UAV) inspection according to claim 2, characterized in that: In the S2 data augmentation step, the augmentation process further includes downsampling the three-dimensional point cloud data, which is used to simulate the point cloud density differences caused by changes in the UAV's flight speed or sampling distance.

5. The method for classifying bridge components based on unmanned aerial vehicle (UAV) inspection according to claim 4, characterized in that: The downsampling ratio is randomly selected within a predefined continuous interval.

6. The method for classifying bridge components based on unmanned aerial vehicle (UAV) inspection according to claim 1, characterized in that: The S3 data preprocessing step specifically includes: performing downsampling on the 3D point cloud data and constructing a spatial index structure for accelerating neighborhood queries; the spatial index structure is a KD tree.

7. The method for classifying bridge components based on unmanned aerial vehicle (UAV) inspection according to claim 1, characterized in that: In the S4 model training step, the point cloud semantic segmentation model is a network model that uses a random sampling strategy for downsampling and includes a local feature aggregation module; in the S4 model training step, a weighted combination loss function is used to optimize the model, and the weighted combination loss function includes at least the cross-entropy loss function.

8. The method for classifying bridge components based on unmanned aerial vehicle (UAV) inspection according to claim 1, characterized in that: Prior to the S1 data acquisition step, the following is also included: Labeling steps: Manually label the collected original bridge 3D point cloud data, assign corresponding bridge component category labels to the points in the point cloud, and form a training dataset; the bridge component categories include at least one or more of background, main beam, bridge column, and cable tower, and further include one or more of stay cable, main cable, suspension cable, anchor, and cable clamp, depending on the bridge type.

9. The method for classifying bridge components based on unmanned aerial vehicle (UAV) inspection according to claim 1, characterized in that: In the S1 data acquisition step, high-resolution orthophotos of the bridge and the three-dimensional point cloud data are acquired simultaneously; the method further includes: S6: Multimodal feature fusion step: Using the high-resolution orthophoto, extract the two-dimensional texture and contour features of the bridge components through the image semantic segmentation model, and perform cross-modal alignment and fusion with the geometric and spatial structural features learned from the three-dimensional point cloud by the point cloud semantic segmentation model to generate enhanced component classification results.

10. A method for classifying bridge components based on unmanned aerial vehicle (UAV) inspection, as described in claim 1 or 9, characterized in that: The method also includes a continuous model performance optimization loop, which includes: S7: Uncertainty Quantification and Active Recommendation Step: After the component classification step in S5, the uncertainty of each point or each local region in the classification result is quantified by using the classification probability output by the point cloud semantic segmentation model. S8: Key Sample Recommendation Step: Based on a preset uncertainty threshold or active learning strategy, automatically filter out the bridge component point cloud fragments that are most uncertain in model classification and most likely to contain classification errors or new features from the large amount of unlabeled point cloud data generated by actual inspections. S9: Human-machine collaborative annotation and model iteration steps: Submit the recommended key samples to experts for rapid verification or fine-tuning, and add the newly annotated data to the training dataset to trigger incremental learning or model fine-tuning in the model training step of S4.