Bridge tower apparent disease detection method based on unmanned aerial vehicle matrix sampling and panoramic stitching
By using a UAV matrix sampling and panoramic stitching method, the problems of low efficiency and poor accuracy in bridge tower surface image acquisition were solved, realizing high-resolution image acquisition, panoramic image stitching, and defect identification, thus meeting the precise positioning requirements for bridge structural health monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANTONG UNIV
- Filing Date
- 2026-03-31
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies suffer from low efficiency in acquiring images of bridge tower surfaces, incomplete coverage, poor panoramic stitching accuracy in weak texture scenes, insufficient accuracy in identifying minor defects, and a lack of spatial positioning capabilities for defects, making it difficult to achieve efficient acquisition, accurate identification, and precise positioning of surface defects on bridge towers.
A matrix sampling and panoramic stitching method using unmanned aerial vehicles (UAVs) is adopted. By constructing a matrix virtual acquisition grid, the UAV carrying a camera is controlled to fly at a deflection angle. Combined with a lightweight image matching network and a multi-scale context-aware attention focusing network, high-resolution image acquisition, panoramic image stitching, and disease identification are achieved.
It enables fully automatic, rapid, and comprehensive image acquisition of bridge tower surfaces, improves the accuracy of panoramic image stitching, enhances the accuracy of defect identification, and achieves precise spatial positioning of defects on bridge tower surfaces, meeting the actual needs of engineering inspection.
Smart Images

Figure CN122454444A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bridge engineering defect detection, specifically involving a method for detecting apparent defects in bridge towers based on UAV matrix sampling and panoramic stitching. Background Technology
[0002] As a key component of transportation infrastructure, the structural safety of large bridges directly affects public safety and transportation efficiency. Bridge towers, as core load-bearing components of large bridges such as suspension bridges and cable-stayed bridges, endure complex loads and are exposed to the natural environment for extended periods. Their surfaces are prone to surface defects such as cracks, holes, and spalling. Early identification and precise location of these defects are crucial for bridge structural health monitoring and preventative maintenance. Therefore, developing efficient, high-precision, and automated technologies for detecting surface defects in bridge towers has become a research hotspot in the field of bridge engineering.
[0003] In recent years, with the rapid development of UAV technology, using UAVs equipped with high-definition cameras for bridge surface image acquisition has gradually become a mainstream technical approach. Domestic and international scholars have conducted extensive research on UAV bridge inspection, achieving significant progress in flight path planning, image acquisition strategies, visual positioning, and navigation. However, existing research largely focuses on conventional components such as bridge main beams and piers, with relatively little research on the special structure of large bridge towers. Bridge towers typically feature large height, vertical facades, and complex surface textures. Existing UAV acquisition methods often employ a single camera flying at a fixed angle, resulting in low acquisition efficiency, insufficient image resolution, and incomplete coverage. Especially for towering structures like bridge towers, there is a lack of systematic solutions for achieving fully automated, rapid, and comprehensive surface image acquisition. A few studies have attempted to use multi-rotor UAVs for zoned acquisition, but the lack of quantitative modeling of acquisition parameters makes it difficult to ensure consistency in image resolution and overlap rate, affecting the quality of subsequent image stitching.
[0004] In image stitching, panoramic image stitching of bridge tower surfaces is fundamental to visualizing the defects of the entire tower. Existing image stitching methods are mainly divided into traditional methods based on feature points and matching methods based on deep learning. Feature point-based methods, such as SIFT and SURF, perform well in textured scenes, but bridge tower surfaces are mostly made of concrete with large smooth areas and a lack of texture information, resulting in insufficient feature point extraction, decreased matching accuracy, and defects such as misalignment and ghosting in the stitched images. Deep learning-based image matching methods have developed rapidly in recent years. Networks such as SuperPoint and LoFTR have shown strong robustness in complex scenes, but these networks usually have a large number of parameters and high computational complexity, making it difficult to meet the real-time requirements of engineering applications. In addition, for large-scale structural surfaces like bridge towers, existing stitching methods mostly use two-dimensional projection transformations, failing to fully utilize the geometric correspondences during the acquisition process, leading to serious error accumulation during long-distance stitching. Therefore, there is an urgent need for a bridge tower surface image stitching method that is lightweight, high-precision, and adaptable to weakly textured scenes.
[0005] In bridge defect identification, semantic segmentation technology based on deep learning has become the mainstream research direction. Scholars both domestically and internationally have successively proposed various improved convolutional neural network structures for the automatic identification of defects such as cracks and holes, such as U-Net, the DeepLab series, and SegFormer. These methods have achieved high recognition accuracy on standard datasets, but there are still the following shortcomings in dedicated research on bridge defects: First, bridge tower defects have multi-scale characteristics; crack widths can reach millimeter levels, and hole diameters can reach centimeter levels, making it difficult for existing networks to simultaneously capture the local features of minute cracks and the global context of larger defects. Second, concrete surfaces have complex backgrounds such as pores, stains, and light and shadow interference, and the ability of existing attention mechanisms to suppress background noise still needs improvement. Third, the segmentation accuracy for micro-cracks cannot yet meet the actual needs of engineering inspection, making it difficult to accurately extract millimeter-level cracks. Therefore, there is an urgent need for a high-precision defect identification network that can focus on multi-scale defect features and suppress background interference.
[0006] Regarding lesion localization, existing research largely focuses on image-level lesion annotation, only outputting the pixel positions of lesions in images, lacking the ability to locate lesions in the actual physical space of the bridge. A few studies use UAV positioning systems to obtain the geographic coordinates of the shooting points and combine them with camera attitude information to estimate lesion locations, but due to limitations in GPS positioning accuracy and attitude measurement errors, the localization results are difficult to meet the needs of precise maintenance. Furthermore, there is currently a lack of mature technical solutions for establishing a unified coordinate system through panoramic images to achieve systematic annotation and spatial mapping of lesion locations.
[0007] In summary, existing technologies have significant shortcomings in areas such as fully automated and rapid data acquisition for detecting surface defects in long-span bridge towers, high-precision panoramic stitching in low-texture scenes, accurate identification of multi-scale micro-defects, and spatial positioning of defects. A systematic solution is urgently needed to achieve efficient data acquisition, accurate identification, and precise positioning of surface defects in bridge towers. Summary of the Invention
[0008] The purpose of this invention is to provide a method and system for detecting surface defects of bridge towers based on UAV matrix sampling and panoramic stitching, so as to solve the technical problems of low efficiency of bridge tower surface image acquisition, incomplete coverage, poor panoramic stitching accuracy in weak texture scenes, insufficient accuracy of small defect identification, and lack of spatial positioning capability of defects in the prior art.
[0009] To achieve the above objectives, the present invention provides the following technical solution:
[0010] A method for detecting apparent defects in bridge towers based on UAV matrix sampling and panoramic stitching includes the following steps:
[0011] Step 1: Construct a matrix-style virtual acquisition grid facing the bridge tower facade, and control the camera on the UAV to carry out coverage acquisition according to the preset deflection angle and zigzag flight path to obtain a high-resolution ultra-high-definition image sequence with geometric correspondence.
[0012] Specifically, step 1 further includes:
[0013] Step 11: Use a wide-angle camera mounted on a drone to capture a wide-angle image of the bridge tower facade, and divide the wide-angle image into a matrix-style virtual acquisition grid according to a preset physical coverage area.
[0014] Step 12: Based on the calibration distance between the UAV and the bridge tower facade, the focal length of the telephoto camera, the actual resolution of the required image, the field of view of the wide-angle camera, the overlap rate of the telephoto image used for subsequent panoramic stitching, and the maximum deflection angle of the UAV gimbal, calculate and determine the spatial coordinates and gimbal attitude angle of each sampling point in the matrix virtual acquisition grid.
[0015] To achieve the desired true resolution image (e.g., an image resolution of 0.1 mm / pixel), the values of the parameters satisfy the following mathematical model:
[0016] Let the distance between the drone and the bridge tower facade be d, the focal length of the telephoto camera be f, and the pixel size be p, then the following conditions are met: This formula characterizes the constraint relationship between system resolution, acquisition distance, and camera parameters. That is, by adjusting the combination of UAV distance and camera focal length, it ensures that the physical size of a single pixel does not exceed 0.1mm, thereby achieving sub-millimeter-level high-resolution image acquisition.
[0017] Let the horizontal overlap rate of the telephoto image be... The number of pixels in the horizontal direction is Then the horizontal spacing Δx between adjacent sampling points satisfies This formula ensures that the images acquired by adjacent sampling points have sufficient overlap in the horizontal direction, providing ample matching features for subsequent panoramic stitching. The same applies to the vertical direction.
[0018] Let the horizontal field of view of the wide-angle camera be... If the horizontal coverage width of the bridge tower facade is W, then the number M of horizontal sampling points in the matrix virtual acquisition grid satisfies: This formula constrains the number of sampling points by using the field of view of a wide-angle camera, ensuring that the drone's flight path can completely cover the bridge tower facade.
[0019] In addition, the horizontal deflection angle of the grid edge sampling points relative to the center point Must meet: ,in This formula ensures that the deflection angle of the drone gimbal during edge sampling does not exceed the gimbal's maximum deflection capability, avoiding image distortion or degradation of imaging quality due to excessive angle. The maximum deflection angle of the UAV gimbal is given.
[0020] Step 13: Control the drone to fly to each sampling point in sequence, and drive the gimbal to deflect the telephoto camera according to the gimbal attitude angle, and collect a sequence of telephoto images covering the bridge tower facade in blocks.
[0021] During the data acquisition phase, a hovering-at-a-point acquisition scheme is adopted before triggering the shutter, i.e., arrival-hovering-shooting-loop. The specific process is as follows: when the drone arrives at the preset grid node... Then, the drone enters a hovering state, waiting for the vibration amplitude of each axis to fall below the threshold. Then, motion blur caused by platform movement is eliminated; in addition, active compensation by the gimbal ensures the camera's principal optical axis is maintained. Always perpendicular to the plane being measured That is, satisfying ( (where is the plane normal vector) to obtain the optimal geometric fidelity.
[0022] Through the above process, image sequences with uniform illumination, stable overlap, and high geometric fidelity can be obtained.
[0023] Step 2: A lightweight image matching network is used to extract and densely match the high-resolution ultra-high-definition image sequence to generate a depth feature descriptor for the bridge tower surface, so as to overcome the defect of insufficient texture on the bridge tower surface.
[0024] Specifically, step 2 further includes: constructing a lightweight image matching network to extract depth features of each image in the high-resolution ultra-high-definition image sequence, generating feature descriptors that combine geometric and semantic information to overcome matching errors caused by insufficient texture in the smooth areas of the bridge tower surface; performing dense matching on adjacent images based on the feature descriptors to obtain matching point pairs; calculating geometric transformation parameters between images based on the matching point pairs, performing geometric alignment and stitching processing, and generating a high-resolution panoramic image of the bridge tower surface.
[0025] The lightweight image matching network employs an encoder-decoder structure. The encoder extracts multi-scale features through depthwise separable convolutions, while the decoder fuses low-level geometric information with high-level semantic information through upsampling, generating feature descriptors that integrate local texture features and global structural information. During training, the network uses a contrastive loss function to enhance the discriminative power of key points in smooth regions. The lightweight image matching network includes a dense feature detection and scoring module and a high-discriminative descriptor generator. The dense feature detection and scoring module outputs reliability scores and sub-pixel-level local offsets for feature points, and uses kurtosis loss to constrain the score map distribution, forming sparse and sharp response peaks. The high-discriminative descriptor generator outputs a dense descriptor map, extracts descriptors from feature points through bilinear interpolation and performs L2 normalization, and is trained using triplet loss.
[0026] The multi-scale feature extraction backbone network acts as a shared encoder, responsible for extracting pyramid-shaped features with rich semantic information from the input image. Given an input image... The encoder samples at different depths to generate a set of multi-scale feature pyramids. ,in , This is the total downsampling step size. This represents the number of channels.
[0027] To integrate the strong semantic information of coarse-scale features with the precise localization capability of fine-scale features, enhanced multi-scale features are generated through upsampling and feature fusion operations. :
[0028]
[0029] in, This indicates a 2x upsampling operation, where `Concat` performs channel concatenation, and `Conv` is a convolutional layer used to reduce the number of channels and fuse information. This process proceeds from low to high resolution, ensuring that features at each scale are integrated into the global context, thus providing robust feature representations for subsequent detection and description tasks.
[0030] The dense feature detection and scoring module employs an efficient dense prediction mechanism, where excellent feature points correspond to local structures in the image with significant gradient changes. The detection module consists of a lightweight convolutional detection head. To achieve the final coordinates of feature points in the original image. It can be calculated using the following formula:
[0031]
[0032] Among them, reliability score , representing the confidence level that the point is a stable feature point. s is the total downsampling step size of the current feature map relative to the original image. Local offset. It is used to achieve sub-pixel level precise positioning.
[0033] To ensure the sparsity and saliency of the feature point distribution and avoid cumbersome non-maximum suppression post-processing, a distribution loss function is introduced during training. Specifically, kurtosis loss is used to constrain the distribution of the predicted score map S, making it tend towards a peak state, thus naturally forming sparse and sharp response peaks.
[0034]
[0035] in, It is a standard score regression loss based on the true value location (such as L1 loss). These are the weighting coefficients for balancing the two terms. Maximizing kurtosis is equivalent to minimizing... This drives the network to concentrate high scores on the most salient feature points.
[0036] The high-discriminacy descriptor generator acts as a decoder, generating a descriptor vector for each feature point that can be used for accurate matching. The descriptor generator is also a lightweight convolutional network. It outputs a dense descriptor graph. For any feature point with sub-pixel precision Its descriptor Extracted from D by bilinear interpolation:
[0037]
[0038] Subsequently, the descriptors are L2 normalized and projected onto a unit hypersphere to allow for similarity measurement using cosine or Euclidean distance:
[0039]
[0040] To train highly discriminative descriptors, the triplet loss function from metric learning is employed. Given an anchor descriptor... A positive sample descriptor A negative sample descriptor The loss function is defined as follows:
[0041]
[0042] This loss function enhances the discriminative power of the descriptor space by bringing positive sample pairs closer together and distancing negative sample pairs further apart. The final training loss of the model is a weighted sum of the detection loss and the descriptor loss:
[0043]
[0044] α and β are hyperparameters used to balance the importance of the two tasks.
[0045] Step 3: Based on the results of the dense matching, the high-resolution ultra-high-definition image sequence is geometrically aligned and stitched to generate a high-resolution panoramic image of the bridge tower surface.
[0046] Step 4: Construct a multi-scale context-aware attention-focusing network, and use the spatial pyramid pooling module and the channel shuffling attention module to identify and segment the defect features in the high-resolution panoramic image, and output the detection results of typical bridge defects, including defect identification, geometric parameter calculation and location.
[0047] Specifically, the multi-scale context-aware attention-focusing network includes a void space pyramid pooling module, which is placed at the end of the encoder. It uses multiple convolutional layers with different expansion rates to capture multi-scale features in parallel, enabling the network to simultaneously perceive the local details of fine cracks and the global context of larger pores, in order to solve the problem of scale changes in concrete surface defects caused by differences in imaging distance and the physical size of the defects themselves.
[0048] Furthermore, the multi-scale context-aware attention focusing network also includes a channel shuffling attention module, which is embedded in the bottleneck layer and the decoder skip connection. Through attention weighting of channel and spatial dimensions, the feature response is adaptively recalibrated to highlight the channels and spatial locations that are highly related to the disease and suppress the activation of background and texture noise.
[0049] To address the significant scale variations in the surface defects of concrete bridge towers (such as ranging from tiny cracks at the pixel level to large-area holes), this invention introduces a void spatial pyramid pooling module at the end of the U-Net network encoder. The multi-scale context-aware attention-focusing network uses U-Net as its basic framework.
[0050] Let the input deep feature map be... ,in For the number of channels, and These represent the height and width of the feature map, respectively.
[0051] feature map Input is sent to the first branch, which is configured to have The convolutional layer with convolutional kernels. This branch performs a linear combination across channels without changing the receptive field of the feature map space. Its purpose is to preserve high-frequency local details in the input feature map to the greatest extent, thereby ensuring the continuity of extremely fine cracks on the bridge tower surface during the feature extraction process.
[0052] feature map The inputs are fed into the second, third, and fourth branches, respectively. Each of the three branches is configured as a 3×3 dilated convolutional layer, with different void ratios set for each. The void ratios of the second, third, and fourth branches are respectively set as follows: , and Furthermore, the size of the output feature map is kept constant by setting the corresponding padding size. Output feature map In position The response can be calculated by the following formula:
[0053]
[0054] in, The convolution kernel weights are used. By configuring different porosity, the second branch focuses on capturing medium-scale diseases (such as localized network cracks), while the fourth branch focuses on covering large-scale areas (such as large-area pores or efflorescence areas), achieving robust coverage of multi-scale disease morphology.
[0055] feature map The input leads to the fifth branch. The fifth branch first uses adaptive global average pooling to compress the spatial dimensions, obtaining a dimension of... The global feature vector Z is calculated using the following formula:
[0056]
[0057] Subsequently, the number of channels was adjusted using 1×1 convolution, and the spatial dimension was upsampled to restore it to its original size using bilinear interpolation. .
[0058] The feature maps output from the five branches are concatenated along the channel dimension. Then, the concatenated high-dimensional feature map is input into the fusion layer, which is configured as follows: Convolutional layers are used for cross-channel information fusion and dimensionality reduction, outputting the final multi-scale disease feature representation. :
[0059]
[0060] in, For activation function, For batch normalization operations, The weights of the fusion layer.
[0061] Through the above steps, by utilizing a parallel multi-scale feature aggregation strategy, the detection model is able to accurately locate local minute cracks on complex bridge tower surfaces under the perspective of UAV inspection, and also has the robustness to eliminate environmental interference by utilizing the context of a large receptive field, which significantly improves the accuracy of defect detection.
[0062] Because images of concrete bridge towers and piers acquired by drone inspections often contain complex backgrounds (such as water stains, oil stains, aggregate spots, and local shadows), these background disturbances are easily confused with real microcracks or spalling holes in terms of local visual features. To maximize feature representation without significantly increasing the number of model parameters, a lightweight channel shuffling attention module is introduced. The channel shuffling attention module performs the following operations:
[0063] Let the feature map tensor input to the module be... ,in For batch size, For the number of channels, and This represents the spatial resolution of the feature map.
[0064] Feature grouping and channel-dimensional splitting: To achieve efficient allocation of computing resources, the module first splits the input feature map along the channel dimension. Divided into Each is an independent feature group. For the first... There are 1 feature groups, and their feature maps are represented as follows: Subsequently, within each group, the aforementioned The feature set is divided into two feature subsets proportionally along the channel dimension: the first feature subset. Second feature subset The mathematical expression for this splitting process is:
[0065]
[0066] in, The aforementioned Configured as input to the channel attention branch to extract semantic attributes, the It is configured to be input to the spatial attention branch to extract location information, thereby achieving decoupling of disease features in semantic and spatial dimensions.
[0067] Adaptive weighted channel attention based on global statistical information: for the first feature subset The channel attention branch obtains channel-level descriptors through global average pooling. Specifically, for The Each channel, its global statistics The calculation formula is:
[0068]
[0069] This yields the global description vector. Subsequently, learnable weight parameters are used. and bias parameters For the vector Perform an affine transformation, followed by the Sigmoid activation function. Generate a channel attention weight matrix. Finally, combine this weight matrix with the original features. Perform element-wise multiplication to obtain the channel-weighted features. :
[0070]
[0071] The significance of this step lies in dynamically evaluating the degree of response of each channel to the characteristics of the disease, adaptively assigning higher weights to channels containing the semantics of cracks or holes, while suppressing channels that respond to background noise such as water stains and oil stains.
[0072] Adaptive weighting of spatial attention based on group normalization: in parallel, targeting the second feature subset The spatial attention branch uses group normalization to obtain spatial statistical information of the feature map. First, the mean of the set of features is calculated. With variance :
[0073]
[0074]
[0075] Normalization was performed using the above statistics. (which is a very small constant), and is achieved through learnable weight parameters. and bias parameters Perform affine transformation and activation, ultimately combining with the original features Multiplying yields the spatially weighted features. :
[0076]
[0077] This step effectively eliminates spatial distribution differences caused by changes in the drone's shooting perspective or uneven local lighting, significantly enhancing the network's ability to accurately locate the geometric shape of the disease edge (such as the direction of cracks).
[0078] Multidimensional feature aggregation and channel shuffling mechanism: After completing the above-mentioned dual-branch parallel weighting, the... and Reassemble along the channel dimension to restore the original dimension of a single set of features, i.e.:
[0079]
[0080] To overcome the information isolation between channels caused by the aforementioned grouping operation, this invention further performs a channel shuffling operation on all aggregated group features. Let the tensor of the feature map after all groups are concatenated be... The channel shuffling process is essentially the transpose and reshaping of tensor dimension indices:
[0081]
[0082] in, This indicates a permutation operation, which physically swaps the "group dimension" (of size ). ) and "Intra-group Channel Dimension" (size is After being washed and flattened, it returned to its original state. The output feature map of the dimension enables deep cross-fusion of features across groups. Through this step, the final output disease features have both fine-grained spatial localization accuracy and global noise-resistant semantic expression, providing a feature foundation for subsequent high-precision disease segmentation.
[0083] Regarding the calculation of the geometric parameters of the disease, step 4 includes the following steps:
[0084] Step 41: Extract the skeleton lines of the segmented diseased areas to obtain the centerline path of the disease;
[0085] Step 42: Using an improved regional grayscale model, a subpixel edge detection algorithm is employed to perform subpixel interpolation on the pixel-level edges of the Canny operator after anisotropic diffusion filtering, thereby obtaining the precise subpixel-level coordinates of the disease edges.
[0086] Step 43: Based on the subpixel edge coordinates, calculate the distance between edge point pairs in the normal direction of the disease centerline to obtain geometric parameters of the disease, such as crack width.
[0087] Regarding the spatial location of the disease, step 4 includes the following steps:
[0088] Step 44: Based on the high-resolution panoramic image of the bridge tower surface generated by stitching, establish a panoramic image coordinate system with the preset reference point in the image as the origin and the preset direction as the coordinate axis;
[0089] Step 45: In the disease segmentation results output by the multi-scale context-aware attention focusing network, the center point of each disease region is taken as the position coordinate of the disease in the panoramic image coordinate system;
[0090] Step 46: Establish the coordinate transformation relationship between the panoramic image coordinate system and the actual physical space of the bridge tower, map the position coordinates of each defect to the actual physical space of the bridge tower, and obtain the actual spatial position of each defect on the surface of the bridge tower.
[0091] This invention also provides a bridge tower surface defect detection system based on UAV matrix sampling and panoramic stitching, used to perform the above-mentioned method. The system includes: a quadcopter UAV platform, a camera array, and an onboard computer. The camera array, mounted on the UAV platform, includes a wide-angle camera and a freely rotatable telephoto camera. The wide-angle camera is used to acquire wide-angle images of the bridge tower facade to assist in planning a matrix virtual acquisition grid. The telephoto camera is used to perform block-based shooting according to the matrix virtual acquisition grid to obtain a high-resolution ultra-high-definition image sequence. The onboard computer, mounted on the UAV platform and electrically connected to the camera array, is used to perform image stitching, defect identification, geometric parameter calculation, and spatial positioning processing of the high-resolution ultra-high-definition image sequence. The camera array also includes supplementary lighting. The quadcopter UAV platform has high-precision hovering and autonomous flight capabilities. The onboard computer is also used to control the UAV flight and gimbal deflection in real time and perform on-site image preprocessing.
[0092] Compared with the prior art, the present invention has the following beneficial effects:
[0093] 1. This invention establishes a quantitative mathematical model of acquisition parameters by combining a matrix-style virtual acquisition grid and a zigzag trajectory planning with a collaborative working mode of wide-angle and telephoto cameras. This enables fully automatic, rapid, and full-coverage acquisition of bridge tower facade images, ensuring consistency in image resolution and overlap rate, and providing a high-quality data foundation for subsequent stitching.
[0094] 2. This invention employs a lightweight image matching network, extracts deep feature descriptors that combine geometric and semantic information through an encoder-decoder structure, and enhances the discriminativeness of key points in smooth areas by combining a contrastive loss function. This effectively solves the matching error problem caused by insufficient surface texture of bridge towers and achieves high-fidelity panoramic image stitching.
[0095] 3. The multi-scale context-aware attention focusing network constructed in this invention captures multi-scale features through the hollow spatial pyramid pooling module, while simultaneously perceiving the contextual information of minute cracks and larger holes; through the channel shuffling attention module, it adaptively highlights disease-related features and suppresses background noise, significantly improving the accuracy and robustness of disease identification.
[0096] 4. This invention employs an improved regional grayscale model-based subpixel edge detection algorithm, combined with skeleton line extraction and normal direction calculation, to achieve subpixel-level accurate measurement of geometric parameters of defects such as crack width, meeting the actual needs of engineering inspection.
[0097] 5. By establishing a mapping relationship between a panoramic image coordinate system and the actual physical space of the bridge tower, this invention enables the calibration of the actual spatial location of each defect on the surface of the bridge tower, providing crucial spatial information support for bridge structural health monitoring and precise maintenance. Attached Figure Description
[0098] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of the present invention and are not intended to limit the present invention.
[0099] Figure 1 This is a flowchart of the present invention;
[0100] Figure 2 This is a schematic diagram of a UAV matrix sampling method according to a preferred embodiment of the present invention;
[0101] Figure 3 This is a schematic diagram of the panoramic stitching effect of a preferred embodiment of the present invention;
[0102] Figure 4 This is a schematic diagram of the apparent defect network of a bridge tower according to a preferred embodiment of the present invention;
[0103] Figure 5 This is a schematic diagram of the loss curve of a model according to a preferred embodiment of the present invention.
[0104] Figure 6 This is a diagram of a drone data acquisition and processing system according to a preferred embodiment of the present invention. Detailed Implementation
[0105] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention.
[0106] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.
[0107] Example 1:
[0108] like Figure 1-5As shown, this invention provides a method for detecting apparent defects in bridge towers based on UAV matrix sampling and panoramic stitching, comprising the following steps:
[0109] 1. UAV matrix sampling
[0110] (1) Matrix grid planning
[0111] First, for the elevation of the bridge tower to be tested, a model is established from the three-dimensional physical space to the two-dimensional projection plane. The mapping is as follows. To balance image acquisition accuracy (resolution) and coverage integrity, a virtual matrix acquisition grid based on geometric constraints is constructed. The specific process is as follows:
[0112] Based on the principles of imaging optics, let the focal length of the UAV camera be... The size of the photosensitive element is The preset ground resolution is The drone's shooting distance relative to the tower wall. satisfy:
[0113]
[0114] in This refers to the pixel size. Under this constraint, the actual coverage width of a single image on the tower wall. and height Defined as:
[0115]
[0116] To ensure the quality of the panoramic stitching after feature matching of the subsequent image sequences, an overlap rate constraint is set. Let the horizontal overlap rate be... Vertical overlap rate The horizontal sampling interval of the grid nodes Vertical step distance Defined as:
[0117]
[0118] (2) Aerial photography data collection by drones
[0119] The generation of drone aerial survey routes is based on the ground station's projection plane. Define a rectangular working area. Discretize this area to generate a series of waypoints. The spatial sequence formed.
[0120] To optimize operational efficiency and maintain the continuity of the movement trajectory, a top-down S-shaped scanning aerial photography route was adopted. Aerial photography route It can be represented as:
[0121]
[0122] This approach ensures that high-resolution image sequences have stable geometric correspondences and consistent neighborhoods in both the temporal and spatial dimensions.
[0123] During the fully automatic data acquisition phase, a method of triggering the shutter after hovering at a fixed point is adopted, i.e., arrival-hover-shoot-loop. The specific process is as follows:
[0124] When the drone arrives at the preset grid node Then, the drone enters a hovering state, waiting for the vibration amplitude of each axis to fall below the threshold. Then, motion blur caused by platform movement is eliminated; in addition, active compensation by the gimbal ensures the camera's principal optical axis is maintained. Always perpendicular to the plane being measured That is, satisfying ( (where is the plane normal vector) to obtain the optimal geometric fidelity.
[0125] Through the above process, image sequences with uniform illumination, stable overlap, and high geometric fidelity can be obtained.
[0126] 2. Image stitching based on deep learning Xfeat network
[0127] XFeat's overall architecture follows an encoder-decoder approach, jointly detecting highly repetitive feature points and their highly discriminative descriptors from a single image of acquired data. The model primarily consists of three core modules: a multi-scale feature extraction backbone network, a dense feature detection and scoring module, and a highly discriminative descriptor generator.
[0128] The multi-scale feature extraction backbone network acts as a shared encoder, responsible for extracting pyramid-shaped features with rich semantic information from the input image. Given an input image... The encoder samples at different depths to generate a set of multi-scale feature pyramids. ,in , This is the total downsampling step size. This represents the number of channels.
[0129] To integrate the strong semantic information of coarse-scale features with the precise localization capability of fine-scale features, enhanced multi-scale features are generated through upsampling and feature fusion operations. :
[0130]
[0131] in, This indicates a 2x upsampling operation, where `Concat` performs channel concatenation, and `Conv` is a convolutional layer used to reduce the number of channels and fuse information. This process proceeds from low to high resolution, ensuring that features at each scale are integrated into the global context, thus providing robust feature representations for subsequent detection and description tasks.
[0132] The dense feature detection and scoring module employs an efficient dense prediction mechanism, where excellent feature points correspond to local structures in the image with significant gradient changes. The detection module consists of a lightweight convolutional detection head. To achieve the final coordinates of feature points in the original image. It can be calculated using the following formula:
[0133]
[0134] Among them, reliability score , representing the confidence level that the point is a stable feature point. s is the total downsampling step size of the current feature map relative to the original image. Local offset. It is used to achieve sub-pixel level precise positioning.
[0135] To ensure the sparsity and saliency of the feature point distribution and avoid cumbersome non-maximum suppression post-processing, a distribution loss function is introduced during training. Specifically, kurtosis loss is used to constrain the distribution of the predicted score map S, making it tend towards a peak state, thus naturally forming sparse and sharp response peaks.
[0136]
[0137] in, It is a standard score regression loss based on the true value location (such as L1 loss). These are the weighting coefficients for balancing the two terms. Maximizing kurtosis is equivalent to minimizing... This drives the network to concentrate high scores on the most salient feature points.
[0138] The high-discriminacy descriptor generator acts as a decoder, generating a descriptor vector for each feature point that can be used for accurate matching. The descriptor generator is also a lightweight convolutional network. It outputs a dense descriptor graph. For any feature point with sub-pixel precision Its descriptor Extracted from D by bilinear interpolation:
[0139]
[0140] Subsequently, the descriptors are L2 normalized and projected onto a unit hypersphere to allow for similarity measurement using cosine or Euclidean distance:
[0141]
[0142] To train highly discriminative descriptors, the triplet loss function from metric learning is employed. Given an anchor descriptor... A positive sample descriptor A negative sample descriptor The loss function is defined as follows:
[0143]
[0144] This loss function makes the descriptor space highly discriminative by bringing positive sample pairs closer together and pushing negative sample pairs further apart.
[0145] The final training loss of the model is a weighted sum of the detection loss and the descriptor loss:
[0146]
[0147] α and β are hyperparameters used to balance the importance of the two tasks.
[0148] 3. Disease Detection Based on Lightweight Deep Learning UNet Network
[0149] An improved architecture based on U-Net is constructed, and performance is enhanced through the following two core strategies: First, a hollow spatial pyramid pooling module is introduced to enhance multi-scale context awareness. Placed at the end of the encoder, it utilizes convolutional layers with different dilation rates to capture multi-scale features in parallel, enabling the model to simultaneously understand the local details of fine cracks and the global context of larger pores, thus solving the recognition challenge caused by scale differences. Second, a channel shuffling attention module is introduced to achieve adaptive feature optimization. This lightweight module is embedded in the bottleneck layer and the skip connection between the decoder. Through channel and spatial dimension attention weighting, it automatically highlights disease-related features and suppresses irrelevant background noise, thereby improving the quality of feature representation and the robustness of the model.
[0150] (1) Hollow space pyramid pooling module
[0151] To address the significant scale variations in the surface defects of concrete bridge towers (such as ranging from pixel-level microcracks to large-area pores), this invention introduces a void spatial pyramid pooling module at the end of the U-Net network encoder. The multi-scale context-aware attention-focusing network uses U-Net as its basic framework.
[0152] Let the input deep feature map be... ,in For the number of channels, and These represent the height and width of the feature map, respectively.
[0153] feature map Input is sent to the first branch, which is configured to have The convolutional layer with convolutional kernels. This branch performs a linear combination across channels without changing the receptive field of the feature map space. Its purpose is to preserve high-frequency local details in the input feature map to the greatest extent, thereby ensuring the continuity of extremely fine cracks on the bridge tower surface during the feature extraction process.
[0154] feature map The inputs are fed into the second, third, and fourth branches, respectively. Each of the three branches is configured as a 3×3 dilated convolutional layer, with different void ratios set for each. The void ratios of the second, third, and fourth branches are respectively set as follows: , and Furthermore, the size of the output feature map is kept constant by setting the corresponding padding size. Output feature map In position The response can be calculated by the following formula:
[0155]
[0156] in, The convolution kernel weights are used. By configuring different porosity, the second branch focuses on capturing medium-scale diseases (such as localized network cracks), while the fourth branch focuses on covering large-scale areas (such as large-area pores or efflorescence areas), achieving robust coverage of multi-scale disease morphology.
[0157] feature map The input leads to the fifth branch. The fifth branch first uses adaptive global average pooling to compress the spatial dimensions, obtaining a dimension of... The global feature vector Z is calculated using the following formula:
[0158]
[0159] Subsequently, the number of channels was adjusted using 1×1 convolution, and the spatial dimension was upsampled to restore it to its original size using bilinear interpolation. .
[0160] The feature maps output from the five branches are concatenated along the channel dimension. Then, the concatenated high-dimensional feature map is input into a fusion layer, configured as a 1×1 convolutional layer, for cross-channel information fusion and dimensionality reduction, outputting the final multi-scale disease feature representation. :
[0161]
[0162] in, For activation function, For batch normalization operations, The weights of the fusion layer.
[0163] Through the above steps, by utilizing a parallel multi-scale feature aggregation strategy, the detection model is able to accurately locate local minute cracks on complex bridge tower surfaces under the perspective of UAV inspection, and also has the robustness to eliminate environmental interference by utilizing the context of a large receptive field, which significantly improves the accuracy of defect detection.
[0164] (2) Channel shuffling attention module
[0165] Because images of concrete bridge towers and piers acquired by drone inspections often contain complex backgrounds (such as water stains, oil stains, aggregate spots, and local shadows), these background disturbances are easily confused with real microcracks or spalling holes in terms of local visual features. To maximize feature representation without significantly increasing the number of model parameters, a lightweight channel shuffling attention module is introduced. The channel shuffling attention module performs the following operations:
[0166] Let the feature map tensor input to the SA module be... ,in For batch size, For the number of channels, and This represents the spatial resolution of the feature map.
[0167] Feature grouping and channel-dimensional splitting: To achieve efficient allocation of computing resources, the SA module first splits the input feature map along the channel dimension. Divided into Each is an independent feature group. For the first... There are 1 feature groups, and their feature maps are represented as follows: Subsequently, within each group, the aforementioned The feature set is divided into two feature subsets proportionally along the channel dimension: the first feature subset. Second feature subset The mathematical expression for this splitting process is:
[0168]
[0169] in, The aforementioned Configured as input to the channel attention branch to extract semantic attributes, the It is configured to be input to the spatial attention branch to extract location information, thereby achieving decoupling of disease features in semantic and spatial dimensions.
[0170] Adaptive weighted channel attention based on global statistical information: for the first feature subset The channel attention branch obtains channel-level descriptors through global average pooling. Specifically, for The Each channel, its global statistics The calculation formula is:
[0171]
[0172] This yields the global description vector. Subsequently, learnable weight parameters are used. and bias parameters For the vector Perform an affine transformation, followed by the Sigmoid activation function. Generate a channel attention weight matrix. Finally, combine this weight matrix with the original features. Perform element-wise multiplication ( ), to obtain channel weighted features :
[0173]
[0174] The significance of this step lies in dynamically evaluating the degree of response of each channel to the characteristics of the disease, adaptively assigning higher weights to channels containing the semantics of cracks or holes, while suppressing channels that respond to background noise such as water stains and oil stains.
[0175] Adaptive weighting of spatial attention based on group normalization: in parallel, targeting the second feature subset The spatial attention branch uses group normalization to obtain spatial statistical information of the feature map. First, the mean of the set of features is calculated. With variance :
[0176]
[0177]
[0178] Normalization was performed using the above statistics. (which is a very small constant), and is achieved through learnable weight parameters. and bias parameters Perform affine transformation and activation, ultimately combining with the original features Multiplying yields the spatially weighted features. :
[0179]
[0180] This step effectively eliminates spatial distribution differences caused by changes in the drone's shooting perspective or uneven local lighting, significantly enhancing the network's ability to accurately locate the geometric shape of the disease edge (such as the direction of cracks).
[0181] Multidimensional feature aggregation and channel shuffling mechanism: After completing the above-mentioned dual-branch parallel weighting, the... and Reassemble along the channel dimension to restore the original dimension of a single set of features, i.e.:
[0182]
[0183] To overcome the information isolation between channels caused by the aforementioned grouping operation, this invention further performs a channel shuffling operation on all aggregated group features. Let the tensor of the feature map after all groups are concatenated be... The channel shuffling process is essentially the transpose and reshaping of tensor dimension indices:
[0184]
[0185] in, This indicates a permutation operation, which physically swaps the "group dimension" (of size ). ) and "Intra-group Channel Dimension" (size is After being washed and flattened, it returned to its original state. The output feature map of the dimension enables deep cross-fusion of features across groups. Through this step, the final output disease features possess both fine-grained spatial localization accuracy and global noise-resistant semantic expression, providing a solid feature foundation for subsequent high-precision disease segmentation.
[0186] 4. Data Quantification
[0187] Obtain the homography matrix and transform the pixel coordinates of the diseased area to physical coordinates in the world coordinate system; based on the physical coordinates, calculate the actual physical dimensions of the disease features, including crack width, spalling area, and perimeter and area of cavities.
[0188] Regarding the calculation of geometric parameters of the disease, this embodiment includes the following steps:
[0189] Step 41: Extract the skeleton lines of the segmented diseased areas to obtain the centerline path of the disease;
[0190] Step 42: Using an improved regional grayscale model, a subpixel edge detection algorithm is employed to perform subpixel interpolation on the pixel-level edges of the Canny operator after anisotropic diffusion filtering, thereby obtaining the precise subpixel-level coordinates of the disease edges.
[0191] Step 43: Based on the subpixel edge coordinates, calculate the distance between edge point pairs in the normal direction of the disease centerline to obtain the geometric parameters of the disease.
[0192] Regarding the spatial location of the disease, this embodiment includes the following steps:
[0193] Step 44: Based on the high-resolution panoramic image of the bridge tower surface generated by stitching, establish a panoramic image coordinate system with the preset reference point in the image as the origin and the preset horizontal and vertical directions as coordinate axes;
[0194] Step 45: In the disease segmentation results output by the multi-scale context-aware attention focusing network, the center point of each disease region is taken as the position coordinate of the disease in the panoramic image coordinate system;
[0195] Step 46: Establish the coordinate transformation relationship between the panoramic image coordinate system and the actual physical space of the bridge tower, map the position coordinates of each defect to the actual physical space of the bridge tower, and obtain the actual spatial position of each defect on the surface of the bridge tower.
[0196] The above-described design method was used to conduct the verification.
[0197] 1. Data Collection
[0198] This invention uses a drone to systematically photograph the surface of a newly built bridge tower and precast beams, collecting a total of 2270 original images, each with a resolution of 5184×3888 pixels. Given that the network model input size is fixed at 512×512 pixels, directly downsampling the original images would result in significant loss of subtle defects such as cracks and holes, affecting recognition accuracy. Therefore, this study employs a sliding window method when constructing the training set, cutting the high-resolution original images into 512×512 pixel image blocks, and then performing meticulous manual annotation to mark cracks and holes pixel-by-pixel. After this processing, a total of 5000 image blocks containing valid defects were finally selected, including 3000 crack images and 2000 hole images, forming the complete dataset used in this study.
[0199] 2. Training and Validation
[0200] The designed network was trained on the aforementioned dataset and compared with the original U-Net. Experiments were conducted using the PyTorch framework on an NVIDIA RTX 4090 GPU platform. The dataset was randomly divided into training and validation sets in an 8:2 ratio. Training parameters were set as follows: a total of 200 training epochs, a batch size of 4, an initial learning rate of 0.0001, and dynamic adjustments using a cosine annealing strategy.
[0201] During model training, Dice Loss is used as the loss function for optimization. Average accuracy is used as the criterion for saving the optimal model, and the model performance is comprehensively evaluated by combining the mean Intersection over Union (mIoU) and the loss variation curve. The formula for calculating mIoU is:
[0202]
[0203] In this model, TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative. The loss curve recorded during training shows that the model converges well and its performance gradually stabilizes. Quantitative evaluation on the validation set shows that the improved model outperforms U-Net. Specifically, the improved model achieves an mIoU of 0.694, while U-Net achieves 0.688. Experimental results demonstrate that the improved model, incorporating the ASPP and ShuffleAttention modules, improves segmentation accuracy by approximately 2.3% compared to the unmodified U-Net, validating the effectiveness of the proposed improvement strategy.
[0204] Table 1. Results of Network Model Ablation Experiment
[0205]
[0206] Example 2:
[0207] The computer-readable storage medium of this embodiment stores a computer program that, when executed by a processor, implements the steps in the bridge tower appearance defect detection method based on UAV matrix sampling and panoramic stitching in Embodiment 1.
[0208] The computer-readable storage medium in this embodiment can be an internal storage unit of the terminal, such as the terminal's hard disk or memory; the computer-readable storage medium in this embodiment can also be an external storage device of the terminal, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc. equipped on the terminal; furthermore, the computer-readable storage medium can include both the terminal's internal storage unit and external storage devices.
[0209] The computer-readable storage medium of this embodiment is used to store computer programs and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0210] Example 3:
[0211] The computer device of this embodiment includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the bridge tower appearance defect detection method based on UAV matrix sampling and panoramic stitching in Embodiment 1.
[0212] In this embodiment, the processor can be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The memory can include read-only memory and random access memory, and provides instructions and data to the processor. A portion of the memory can also include non-volatile random access memory. For example, the memory can also store device type information.
[0213] Those skilled in the art will understand that the content disclosed in the embodiments can be provided as a method, system, or computer program product. Therefore, this solution can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this solution can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage) containing computer-usable program code.
[0214] This solution is described with reference to flowchart illustrations and / or block diagrams of methods and computer program products according to embodiments of this solution. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0215] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0216] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0217] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory, etc.
[0218] The examples described herein are merely preferred embodiments of the invention and are not intended to limit the concept and scope of the invention. Any modifications and improvements made by those skilled in the art to the technical solutions of the invention without departing from the design concept of the invention should fall within the protection scope of the invention.
[0219] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the specific embodiments described above. The specific embodiments and descriptions in the specification are merely for further illustrating the principles of the invention. Various changes and modifications can be made to the present invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the claims and their equivalents.
Claims
1. A method for detecting apparent defects in bridge towers based on UAV matrix sampling and panoramic stitching, characterized in that, Includes the following steps: Step 1: Construct a matrix-style virtual acquisition grid facing the bridge tower facade, and control the camera on the UAV to carry out coverage acquisition according to the preset deflection angle and "zigzag" trajectory to obtain a high-resolution ultra-high-definition image sequence with geometric correspondence. Step 2: A lightweight image matching network is used to extract and densely match the high-resolution ultra-high-definition image sequence to generate a depth feature descriptor for the bridge tower surface, so as to overcome the defect of insufficient texture on the bridge tower surface. Step 3: Based on the results of the dense matching, perform geometric alignment and stitching on the high-resolution ultra-high-definition image sequence to generate a high-resolution panoramic image of the bridge tower surface; Step 4: Construct a multi-scale context-aware attention-focusing network, and use the spatial pyramid pooling module and the channel shuffling attention module to identify and segment the defect features in the high-resolution panoramic image, and output the identification results, geometric parameters and spatial location information of typical bridge defects.
2. The method according to claim 1, characterized in that, Step 1 further includes: Step 11: Use a wide-angle camera mounted on a drone to capture a wide-angle image of the bridge tower facade, and divide the wide-angle image into a matrix-style virtual acquisition grid according to the preset physical coverage area; Step 12: Based on the calibration distance between the UAV and the bridge tower facade, the focal length of the telephoto camera, the actual resolution of the required image, the field of view of the wide-angle camera, the overlap rate of the telephoto image used for subsequent panoramic stitching, and the maximum deflection angle of the UAV gimbal, calculate and determine the spatial coordinates and gimbal attitude angle of each sampling point in the matrix virtual acquisition grid. Step 13: Control the drone to fly to each sampling point in sequence, and drive the gimbal to deflect the telephoto camera according to the gimbal attitude angle, and collect a sequence of telephoto images covering the bridge tower facade in blocks.
3. The method according to claim 2, characterized in that, In step 12, the values of each parameter satisfy the following mathematical model: Let the distance between the drone and the bridge tower facade be d, the focal length of the telephoto camera be f, and the pixel size be p, then the following condition is met: ; Let the horizontal overlap rate of the telephoto image be... The number of pixels in the horizontal direction is The horizontal spacing between adjacent sampling points satisfy Let the horizontal field of view of the wide-angle camera be... The horizontal coverage width of the bridge tower facade is The number of horizontal sampling points in the matrix-type virtual acquisition grid is... satisfy ; Horizontal deflection angle of grid edge sampling points relative to the center point satisfy ,in , The maximum deflection angle of the UAV gimbal is given.
4. The method according to claim 1, characterized in that, In step 2, the lightweight image matching network adopts an encoder-decoder structure, where the encoder extracts multi-scale features through depthwise separable convolution, and the decoder fuses low-level geometric information and high-level semantic information through upsampling. The lightweight image matching network includes a dense feature detection and scoring module and a high-discrimination descriptor generator. The dense feature detection and scoring module outputs the reliability score of feature points and sub-pixel-level local offset, and uses kurtosis loss to constrain the distribution of the score map. The high-discrimination descriptor generator outputs a dense descriptor map, extracts descriptors for feature points through bilinear interpolation and performs L2 normalization, and uses triplet loss for training.
5. The method according to claim 1, characterized in that, In step 4, the multi-scale context-aware attention focusing network uses U-Net as its basic framework, the hollow spatial pyramid pooling module is placed at the end of the encoder, and the channel shuffling attention module is embedded in the bottleneck layer and the decoder jump connection.
6. The method according to claim 5, characterized in that, The void space pyramid pooling module adopts a five-branch parallel structure: The first branch is characterized by Convolutional layers with convolutional kernels; The second, third, and fourth branches are all... The voided convolutional layers are set to void ratios of 100% and 100% respectively. , and ; The fifth branch is followed by global average pooling. Convolutional upsampling restores the size; the feature maps output from the five branches are concatenated along the channel dimension, and then... Convolutional layer fusion and dimensionality reduction.
7. The method according to claim 5, characterized in that, The channel shuffling attention module performs the following operations: The input feature map is divided along the channel dimension into Each group of independent features is divided into a first feature subset and a second feature subset proportionally along the channel dimension. The first feature subset is input into the channel attention branch, and channel-level descriptors are obtained through global average pooling. Channel attention weights are generated by Sigmoid activation and multiplied element-wise with the first feature subset to obtain channel-weighted features. The second feature subset is input into the spatial attention branch, and spatial statistics are obtained by group normalization. Spatial attention weights are generated by Sigmoid activation and multiplied element-wise with the second feature subset to obtain spatial weighted features. The channel weighted features and the spatial weighted features are concatenated along the channel dimension, and then channel shuffling is performed on all groups to exchange the group dimension and the channel dimension within the group.
8. The method according to claim 1, characterized in that, Step 4 also includes the following steps: Step 41: Extract the skeleton lines of the segmented diseased areas to obtain the centerline path of the disease; Step 42: Using an improved regional grayscale model, a subpixel edge detection algorithm is employed to perform subpixel interpolation on the pixel-level edges of the Canny operator after anisotropic diffusion filtering, obtaining the precise subpixel coordinates of the disease edges; Step 43: Based on the subpixel edge coordinates, the distance between edge point pairs is calculated in the normal direction of the disease centerline to obtain the geometric parameters of the disease; Step 44: Based on the high-resolution panoramic image of the bridge tower surface generated by stitching, a panoramic image coordinate system is established with a preset reference point in the image as the origin and preset horizontal and vertical directions as coordinate axes; Step 45: In the disease segmentation results output by the multi-scale context-aware attention focusing network, the center point of each disease region is used as the position coordinate of the disease in the panoramic image coordinate system; Step 46: A coordinate transformation relationship is established between the panoramic image coordinate system and the real physical space of the bridge tower, mapping the position coordinates of each disease to the real physical space of the bridge tower to obtain the real spatial position of each disease on the bridge tower surface.
9. A bridge tower surface defect detection system based on UAV matrix sampling and panoramic stitching for performing the method according to any one of claims 1 to 8, characterized in that, include: The system comprises a quadcopter drone platform, a camera array, and an onboard computer. The camera array, mounted on the drone platform, includes a wide-angle camera and a freely rotatable telephoto camera. The wide-angle camera is used to acquire wide-angle images of the bridge tower facade to assist in planning a matrix-style virtual acquisition grid. The telephoto camera is used to perform segmented shooting according to the matrix-style virtual acquisition grid to obtain a high-resolution ultra-high-definition image sequence. The onboard computer, mounted on the drone platform and electrically connected to the camera array, is used to perform image stitching, defect identification, geometric parameter calculation, and spatial positioning processing of the high-resolution ultra-high-definition image sequence.
10. The system according to claim 9, characterized in that, The camera array also includes supplementary lighting; the quadcopter UAV platform has high-precision hovering and autonomous flight capabilities; the onboard computer is also used to control the UAV flight and gimbal deflection in real time, and to perform on-site image preprocessing.