Bridge disease detection method and unmanned aerial vehicle
By voxelizing the three-dimensional bridge model and deploying a defect detection model on a drone, the problems of slow recognition speed, low accuracy, and high computing resource consumption in drone bridge defect detection were solved, achieving real-time and accurate defect detection.
Patent Information
- Application Number
- CN202510846856.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-23
AI Technical Summary
The existing bridge defect detection method based on drone technology has problems such as slow defect identification speed, insufficient identification accuracy, excessive consumption of computing resources and inability to perform real-time detection.
By voxelizing the three-dimensional model of the bridge, a refined drone flight path is generated, and a disease detection model is deployed on the drone. The backbone network is added with the SCSA mechanism, and the neck network introduces a simplified attention mechanism to achieve real-time image analysis.
It improves the recognition speed and efficiency of bridge defect detection, reduces the data processing pressure on the cloud or back-end server, enhances the accuracy and coverage of detection, and avoids the randomness and omission risks of manual operation.
Smart Images

Figure CN120685095A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of bridge health detection, and in particular to a bridge disease detection method and a drone. Background Art
[0002] With the large-scale construction of transportation infrastructure in my country, the maintenance and safety monitoring of critical structures such as bridges are becoming increasingly important. Traditional bridge inspection methods rely primarily on visual inspections by experienced engineers or require the deployment of a large number of sensor networks. The former is susceptible to subjective factors, while the latter faces challenges such as high deployment costs and complex maintenance. Both methods are inefficient and costly.
[0003] In recent years, drones have gradually emerged in the bridge inspection field due to their flexibility, high maneuverability, and low cost. However, current drone inspections still face many limitations, such as inefficient image acquisition due to manual remote control, which makes real-time image processing and analysis impossible. Furthermore, the massive amount of bridge image data places a significant strain on computer storage and computing resources. Meanwhile, edge computing, an emerging technology, can provide computing support close to the data source, significantly reducing data transmission latency and improving processing efficiency, offering new opportunities to address these issues.
[0004] It is worth noting that current drone-based bridge defect detection systems still have certain shortcomings. These include slow recognition speed, insufficient recognition accuracy, excessive computational resource consumption, and an inability to perform real-time detection. Therefore, developing an efficient, accurate, and automated real-time bridge defect detection method is crucial. Summary of the Invention
[0005] (1) Technical issues to be resolved
[0006] In view of the above-mentioned shortcomings and deficiencies of the existing technology, the present application provides a bridge defect detection method and drone, which solves the technical problems in the existing bridge defect detection method based on drone technology, such as slow defect recognition speed, insufficient recognition accuracy, excessive computing resource consumption and inability to perform real-time detection.
[0007] (2) Technical solution
[0008] In order to achieve the above objectives, the main technical solutions adopted in this application include:
[0009] In a first aspect, an embodiment of the present application provides a bridge defect detection method, comprising:
[0010] Determining a three-dimensional model of the area to be inspected of the bridge, and performing voxel processing on the three-dimensional model of the area to be inspected to obtain a voxelized model;
[0011] Performing UAV flight path planning based on the voxelized model to obtain a flight planning path;
[0012] Acquiring images captured by a drone, wherein the drone captures images of the bridge while flying along the planned flight path;
[0013] The captured image is input into a pre-trained disease detection model for detection to obtain a bridge disease image of the area to be detected, wherein the disease detection model is deployed on the drone, and the SCSA mechanism is added to the last layer of the backbone network of the disease detection model, and a simplified attention mechanism with slicing operation is added to the convolution module of the neck network.
[0014] Optionally, in some embodiments of the present application, performing UAV flight path planning based on the voxelized model includes:
[0015] Constructing a discrete grid map for the space where the voxelized model is located, and using the grid where the grid map contacts the voxelized model as a viewpoint;
[0016] An incomplete graph is constructed based on the viewpoints, and each viewpoint is connected on the incomplete graph based on a greedy algorithm to form the flight planning path.
[0017] Optionally, in some embodiments of the present application, constructing a discrete grid map for the space where the voxelized model is located includes:
[0018] respectively obtaining the extreme coordinates of the voxelized model along the x-axis, y-axis, and z-axis in a pre-established three-dimensional coordinate system;
[0019] A grid map that completely accommodates the voxelized model is established based on the extreme coordinates of the voxelized model along the x-axis, the y-axis, and the z-axis in a pre-established three-dimensional coordinate system, and the voxelized model is mapped to the grid map.
[0020] Optionally, in some embodiments of the present application, constructing an incomplete graph based on the viewpoint includes:
[0021] Obtaining a distance between each viewpoint and the remaining viewpoints based on the grid map, and obtaining a viewpoint direction unit vector corresponding to each viewpoint based on the distance between each viewpoint and the remaining viewpoints;
[0022] The incomplete graph is constructed according to the viewpoint direction unit vector corresponding to each viewpoint.
[0023] Optionally, in some embodiments of the present application, obtaining a viewpoint direction unit vector corresponding to each viewpoint based on the distance between each viewpoint and the remaining viewpoints includes:
[0024] According to the distance between any viewpoint and the remaining viewpoints, obtain the attraction vector of the remaining viewpoints to the viewpoint, wherein the magnitude of the attraction vector is a pre-configured empirical constant divided by the square of the distance between the two viewpoints, and the direction of the attraction vector is the direction from the viewpoint to the remaining viewpoints;
[0025] Add the attraction vectors of the remaining viewpoints to the viewpoint to obtain the three-dimensional force vector corresponding to each viewpoint;
[0026] The three-dimensional force vector corresponding to each viewpoint is normalized to obtain the viewpoint direction unit vector corresponding to each viewpoint.
[0027] Optionally, in some embodiments of the present application, the disease detection model is obtained by training with the YOLOv11 model as the benchmark network, wherein the last layer of the backbone network is the C2PSA layer. When the SCSA mechanism is added to the C2PSA layer, spatial attention extraction is performed on the original feature map output by the SPPF module along the width and height directions respectively to obtain a height attention map and a width attention map, and the height attention map, the width attention map and the original feature map are multiplied to obtain a corresponding spatial attention feature map, and corresponding attention weights are assigned to all channels in the spatial attention feature map based on the channel self-attention mechanism to obtain a channel attention weight matrix, and the spatial attention feature map and the channel attention weight matrix are multiplied to obtain a fused feature map, so that the neck network and the head network can mark the diseases in the captured image based on the fused feature map.
[0028] Optionally, in some embodiments of the present application, spatial attention extraction is performed on the original feature map output by the SPPF module along the width direction and the height direction respectively, including:
[0029] Decomposing the original feature map along the height direction and the width direction, and performing a global pooling operation to obtain a height feature map and a width feature map;
[0030] Decomposing the height feature map into at least one height sub-feature map, and decomposing the width feature map into at least one width sub-feature map;
[0031] The height sub-feature maps are respectively extracted by using a multi-receptive field shared one-dimensional depth convolution method, and are then spliced, grouped and normalized, and activated by a Sigmoid function to obtain the height attention map;
[0032] The width sub-feature maps are respectively extracted by using a multi-receptive field shared one-dimensional depth convolution method, and are then spliced, grouped and normalized, and activated with a Sigmoid function to obtain the width attention map.
[0033] Optionally, in some embodiments of the present application, corresponding attention weights are assigned to all channels in the spatial attention feature map based on a channel self-attention mechanism to obtain a channel attention weight matrix, including:
[0034] After performing average pooling and group normalization operations on the spatial attention feature map, a linear transformation is performed through a two-dimensional depthwise convolution to obtain the corresponding query vector, key vector, and value vector;
[0035] According to the query vector, key vector, and value vector, each channel in the spatial attention feature map is assigned a corresponding channel attention weight, and processed by a Sigmoid activation function to obtain the channel attention weight matrix.
[0036] Optionally, in some embodiments of the present application, the convolution module is a Conv-SWS structure, which is suitable for performing spatial slicing and / or channel slicing on the input feature map to obtain at least one slice map, and evaluate the importance of each slice map based on a preconfigured energy function, and perform feature enhancement on the feature map of the input Conv-SWS structure according to the importance of all slice maps.
[0037] In a second aspect, an embodiment of the present application provides a drone, including a memory and a processor;
[0038] The memory is used to store executable instructions of the processor, and the processor is configured to perform the above-mentioned bridge defect detection method by executing the executable instructions.
[0039] (3) Beneficial effects
[0040] The bridge defect detection method and drone provided in the embodiments of this application feature comprehensive technical optimizations from data acquisition to real-time processing. In terms of data acquisition, by voxelizing the three-dimensional bridge model, the bridge structure is decomposed into quantifiable spaces. This enables the generation of a fully covered, non-redundant drone flight path based on a refined spatial network. Compared to traditional manual remote control methods, this eliminates the need for manual remote control of the drone to capture each point, avoiding the randomness and risk of omissions associated with manual operation and improving detection accuracy and monitoring efficiency.
[0041] In terms of data processing, the disease detection model is directly deployed on the drone, realizing real-time analysis of the captured images without the need to transmit data to the cloud or back-end server, thereby improving the recognition speed and efficiency of disease detection and reducing the data processing pressure on the cloud or back-end server; at the model architecture level, the SCSA mechanism is added to the last layer of the backbone network of the disease detection model, and the accuracy of disease detection is improved by analyzing the attention weight of the input feature map; at the same time, a simplified attention mechanism for slicing operations is introduced in the convolution module of the neck network, which reduces background interference and improves the accuracy of disease recognition through local feature slicing and cross-layer interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 1 is a flow chart of a bridge defect detection method according to one embodiment of the present application;
[0043] Figure 2 This is a flowchart of drone flight path planning according to one embodiment of the present application;
[0044] Figure 3 A flowchart of constructing a discrete grid map according to one embodiment of the present application;
[0045] Figure 4 A flowchart of constructing an incomplete graph according to one embodiment of the present application;
[0046] Figure 5 This is a schematic diagram of a three-dimensional model of a bridge according to one embodiment of the present application;
[0047] Figure 6 This is a schematic diagram of the viewpoint, line of sight, and path planning process of a drone according to one embodiment of the present application;
[0048] Figure 7 Schematic diagram of the disease detection model structure according to one embodiment of the present application. DETAILED DESCRIPTION
[0049] In order to better explain the present application and facilitate understanding, the present application is described in detail below through specific implementation methods in conjunction with the accompanying drawings.
[0050] Bridge inspection methods in related technologies often rely on visual inspections by engineers or require the deployment of a large number of various sensor networks. This method is not only inefficient but also costly.
[0051] The drone-based bridge inspection method, which is highly flexible and low-cost, not only requires pilots to manually remotely shoot, but also fails to meet the needs of real-time image processing and analysis. In addition, the large amount of bridge image data also consumes a lot of computing and storage resources on remote servers.
[0052] To this end, the bridge defect detection method and drone provided in this application have undergone comprehensive technical optimization, from data acquisition to real-time processing. In terms of data acquisition, by voxelizing the three-dimensional bridge model, the bridge structure is decomposed into quantifiable spaces. This enables the generation of a fully covered, non-redundant drone flight path based on a refined spatial network. Compared to traditional manual remote control methods, this eliminates the need for manual remote control of the drone to capture each point, avoiding the randomness and risk of omissions associated with manual operation and improving detection accuracy and monitoring efficiency.
[0053] In terms of data processing, the disease detection model is directly deployed on the drone, realizing real-time analysis of the captured images without the need to transmit data to the cloud or back-end server, thereby improving the recognition speed and efficiency of disease detection and reducing the data processing pressure on the cloud or back-end server; at the model architecture level, the SCSA mechanism is added to the last layer of the backbone network of the disease detection model, and the accuracy of disease detection is improved by analyzing the attention weight of the input feature map; at the same time, a simplified attention mechanism for slicing operations is introduced in the convolution module of the neck network, which reduces background interference and improves the accuracy of disease recognition through local feature slicing and cross-layer interaction.
[0054] To better understand the above technical solutions, exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a clearer and more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0055] Figure 1 Schematic diagram of a flow chart of a bridge defect detection method according to an embodiment of the present application.
[0056] like Figure 1 As shown, the bridge disease detection method includes:
[0057] S1. Determine a three-dimensional model of a bridge area to be inspected, and perform voxelization processing on the three-dimensional model of the area to be inspected to obtain a voxelized model.
[0058] A voxel is the smallest unit in three-dimensional space, similar to a pixel in a two-dimensional image. Each voxel represents a small cubic area in three-dimensional space and contains data such as position, color, and density. Voxelized models decompose a three-dimensional model (such as a bridge structure) into a series of regularly arranged voxelized units, forming a discrete three-dimensional network representation. In this way, continuous geometric shapes can be converted into digital models that can be processed by computers.
[0059] In one possible implementation of this embodiment, the three-dimensional point cloud data of the bridge can be obtained through technologies such as laser radar and drone photography and measurement, and then a three-dimensional model can be constructed, or a three-dimensional model of the bridge structure can be generated through construction drawings; the three-dimensional model is divided into a regular network structure through a set voxel size, each voxel unit corresponds to a set of spatial points in the point cloud data, and each voxel unit is assigned geometric attributes (such as whether it belongs to the bridge structure), image attributes (such as image features of the corresponding area), etc.; finally, redundant voxels (such as background noise) are eliminated, and the grid boundaries are smoothed to ensure the accuracy of the voxelized model. Generally, the voxelization process is implemented by using an octree structure. This embodiment provides a technical basis for subsequent drone path planning and disease detection by converting the bridge structure into a computable three-dimensional model and performing voxelization processing. The advantage of voxelization processing is that it has strong spatial quantization capabilities, which improves the coverage, detection accuracy and detection efficiency in bridge detection.
[0060] Furthermore, in general, since the bridge deck structure and span structure (i.e., the upper structure of the bridge) are inspected by inspection vehicles, before voxelizing the 3D model, the 3D model can be divided into the upper structure and the lower structure (piers, abutments, and foundations) based on the bridge supports. The upper structure is discarded, and only the lower structure is voxelized. Only the lower structure that needs to be inspected by drones is voxelized, which reduces the amount of data processing and improves the data processing efficiency. The voxelization process is as follows: Figure 5 shown.
[0061] S2. Perform UAV flight path planning based on the voxelized model to obtain a planned flight path.
[0062] Among them, UAV path planning refers to generating an efficient, safe, and fully covered flight route for the UAV based on the detection target, environmental constraints, and mission requirements. In bridge detection, path planning directly affects detection efficiency, data integrity, and flight safety. In the current common UAV path planning, the UAV is equipped with a variety of sensors to perceive the environment in real time and dynamically adjust the path. This flight path planning method has high requirements for the real-time computing power of the processing device and is prone to omissions. However, this application proposes to decompose the bridge model into regular voxel units through the voxelized grid method, and the UAV scans line by line in the grid order, with strong coverage, which is suitable for complex geometric structures such as bridge disease detection. The path can be simulated in advance to avoid risks.
[0063] Optionally, in some embodiments of the present application, the UAV flight path planning is performed based on the voxelized model, such as Figure 2 As shown, including:
[0064] S21 . Construct a discrete grid map for the space where the voxelized model is located, and use the grid where the grid map contacts the voxelized model as a viewpoint.
[0065] S22: construct an incomplete graph based on the viewpoints, and connect the viewpoints on the incomplete graph based on a greedy algorithm to form the flight planning path.
[0066] It should be noted that the viewpoint refers to the position, posture, and viewing angle parameters of the drone when photographing the bridge, which directly affects image quality, defect identification accuracy, and detection coverage efficiency. In this embodiment, by discretizing the space where the voxelized model resides into a grid map, the continuous space problem is transformed into a discrete node problem, significantly reducing the amount of computation (the number of grids is usually much smaller than the number of voxels). Furthermore, only grids that touch the voxelized model are used as viewpoints, avoiding the ineffective modeling of empty grids (obstacle-free areas), further compressing the data size, and improving real-time performance.
[0067] A non-complete graph refers to a mathematical model that describes the "non-fully connected relationship" between the drone's viewpoint and objects (or areas) in the target scene.
[0068] The greedy algorithm is an electrostatic algorithm strategy for solving optimization problems. Its core idea is to take the best (or most favorable) decision under the current state in each step of selection to obtain the global optimal solution.
[0069] That is to say, in the embodiments of the present application, the incomplete graph only connects reachable and meaningful viewpoints, rather than fully connected, which greatly reduces the number of edges of the graph and reduces the complexity of the algorithm. Algorithmically, a feasible path is quickly generated through a greedy algorithm to adapt to dynamic scenarios. At each step, the current optimal connection is prioritized without backtracking or global traversal. The algorithm has low time complexity and is suitable for real-time path planning of drones.
[0070] Alternatively, in some embodiments of the present application, a discrete grid map is constructed for the space where the voxelized model is located, such as Figure 3 Shown, including:
[0071] S211, respectively obtaining the extreme coordinates of the voxelized model along the x-axis, y-axis, and z-axis in a pre-established three-dimensional coordinate system;
[0072] S212 : Establish a grid map that completely accommodates the voxelized model based on the extreme coordinates of the voxelized model along the x-axis, the y-axis, and the z-axis in a pre-established three-dimensional coordinate system, and map the voxelized model to the grid map.
[0073] Specifically, when establishing a viewpoint, a discrete grid map is constructed, and a cubic box with a side length of D is set so that it can accommodate a three-dimensional model and satisfy the following formula:
[0074] D≥max[(x max -x min ),(y max -y min ),(z max -z min )];
[0075] Where D is the side length of the cube box, that is, the grid map that can completely accommodate the voxelized model, x max is the maximum value of the voxelized model along the x-axis in the pre-established three-dimensional coordinate system, x min is the minimum value of the voxel model along the x-axis in the pre-established three-dimensional coordinate system, max is the maximum value of the voxelized model along the y-axis in the pre-established three-dimensional coordinate system, min is the minimum value of the voxelized model along the y-axis in the pre-established three-dimensional coordinate system, z max is the maximum value of the voxelized model along the z axis in the pre-established three-dimensional coordinate system, z min It is the minimum coordinate value of the voxelized model along the z-axis in the pre-established three-dimensional coordinate system.
[0076] The grid side length d is selected according to the detection requirements (e.g., d = 0.5m), where the index (i, j, k) of each grid corresponds to the spatial position, that is, (x, y, z) = (x min +i·d,y min +j·d,z min + k·d). The grid state is then marked. If the grid intersects with the bridge voxel model, it is an occupied grid and marked as 1. Otherwise, it is marked as a free grid and marked as 0. In this embodiment, an octree structure is used to accelerate the query and avoid voxel detection.
[0077] When mapping the voxelized model to the raster map, each voxel should correspond to one or more grid cells in the raster map to ensure that the model's geometric features (such as obstacle positions and hollow areas) are accurately expressed in the discrete space, and the raster map is fully aligned with the three-dimensional coordinates to facilitate coordinate conversion and spatial positioning for subsequent operations such as viewpoint selection and path calculation.
[0078] Furthermore, complex structures (such as bridge pier connections) can be locally refined to generate secondary grids to improve the full coverage of subsequent disease detection.
[0079] Furthermore, ensure that the cube box can fully accommodate the model in the x, y, and z directions to avoid the model partially exceeding the range of the grid map. When one dimension of the model is much smaller than the other directions, the bounding box can be adjusted to a non-cubic box to reduce redundant space, but the consistency of the grid division must be maintained.
[0080] Furthermore, the selection of grid resolution d should be reasonable. When the resolution is high (small d), the detection accuracy is high, but the computational complexity is large; when the resolution is low (large d), the computational efficiency is high, but small targets may be missed.
[0081] That is to say, in the embodiments of the present application, by establishing a minimum cubic area that completely accommodates the voxelized model, it is ensured that the grid map covers all parts of the model to avoid missing obstacles or passable areas; and the mechanism coordinates directly determine the size of the grid map, so that the map size matches the actual size of the model to avoid redundancy or insufficiency.
[0082] Alternatively, in some embodiments of the present application, an incomplete graph is constructed based on the viewpoint, such as Figure 4 Shown, including:
[0083] S221. Obtain the distance between each viewpoint and the remaining viewpoints based on the grid map, and obtain the viewpoint direction unit vector corresponding to each viewpoint based on the distance between each viewpoint and the remaining viewpoints.
[0084] S222: Construct the incomplete graph according to the viewpoint direction unit vector corresponding to each viewpoint.
[0085] Specifically, in some embodiments of the present application, for any viewpoint, the distance between each viewpoint and the remaining viewpoints is first calculated, and the attractive force is calculated using this distance. The magnitude of the attractive force is a constant (empirical coefficient) divided by the square of the distance. The direction of the attractive force is the unit vector pointing from the viewpoint to the remaining viewpoints. The attractive force vectors of all cubes are then added together to obtain a three-dimensional force vector, which is then normalized to obtain the viewpoint direction unit vector. The above process can be expressed by the following formula:
[0086]
[0087] f min <||p ti -p v || <f max ;
[0088] Among them, a is a constant, p v is the viewpoint position, p ti Indicates the position of the center of the i-th triangle, N is the number of triangle center points, [f min ,f max ] is the preset range.
[0089] That is, based on the distance between each viewpoint and the rest of the viewpoints, the viewpoint direction unit vector corresponding to each viewpoint is obtained, including:
[0090] According to the distance between any viewpoint and the remaining viewpoints, the attraction vector of the remaining viewpoints to the viewpoint is obtained, wherein the magnitude of the attraction vector is a pre-configured empirical constant divided by the square of the distance between the two viewpoints, and the direction of the attraction vector is the direction from the viewpoint to the remaining viewpoints; the attraction vectors of the remaining viewpoints to the viewpoint are added to obtain the three-dimensional force vector corresponding to each viewpoint; the three-dimensional force vector corresponding to each viewpoint is normalized to obtain the viewpoint direction unit vector corresponding to each viewpoint.
[0091] Specifically, in some embodiments of the present application, the path planning process is as follows: Figure 6 As shown, for each viewpoint, calculate its attraction vector to all free space grids:
[0092]
[0093] Then, the attraction vector is synthesized, that is, all the attraction forces of the viewpoint are added together to obtain the three-dimensional force vector of each viewpoint. The three-dimensional force vector of each viewpoint is normalized to obtain the unit direction vector of each viewpoint, that is, the direction of the camera on the drone. Among them, only the distance L∈[f min ,f max ] grid.
[0094] Define nodes and edges, where all viewpoints are considered nodes. Define edge connection rules, that is, the edges between two viewpoints must meet the following requirements: first, reachability, which means there are no obstacles in the straight path (verified by ray method); second, distance constraint, which means the distance between the two viewpoints must be less than the preset distance threshold; and finally, directional consistency, which means the steering angle must meet the following requirements:
[0095]
[0096] Based on the edge connection rules, edge weights are calculated. Specifically, the distance cost is calculated based on the distance constraint between the two viewpoints; the steering cost is calculated using the steering angle between the two viewpoints; and the coverage gain is calculated by calculating the proportion of newly covered undetected areas in the path between the two viewpoints. The corresponding edge weight is obtained by combining the distance cost, steering cost, and coverage enhancement.
[0097] Through the greedy algorithm, the viewpoint with the minimum weight and the largest coverage gain is selected as the corresponding flight path.
[0098] That is to say, in the embodiments of the present application, by combining multi-dimensional attraction vectors with graph theory optimization, and through multi-dimensional constraints and multiple target weights, efficient, smooth and full-coverage drone path planning is achieved.
[0099] Furthermore, before the drone begins capturing images based on its planned flight path, the calculated viewpoint coordinates must be converted to latitude, longitude, and relative altitude. Using the North-East-Down (NED) convention, the x, y, and z axes of the simulation model must be aligned with the real-world north, east, and down directions. This coordinate conversion yields the position of the drone's viewpoint in the real world. Furthermore, the calculated view direction must be adjusted to reflect the real world.
[0100] S3. Acquire images captured by a drone, wherein the drone captures images of the bridge while flying along the planned flight path;
[0101] The captured image is input into a pre-trained disease detection model for detection to obtain a bridge disease image of the area to be detected, wherein the disease detection model is deployed on the drone, and the SCSA mechanism is added to the last layer of the backbone network of the disease detection model, and a simplified attention mechanism with slicing operation is added to the convolution module of the neck network.
[0102] It should be noted that simplifying the attention mechanism requires defining the following energy function for each neuron in the network model:
[0103]
[0104] Among them, w t is the linear transformation weight of the target neuron, b t is the linear transformation bias of the target neuron, and is the target neuron t and other neurons x i The ideal output label.
[0105] Then, w t and b t Find the partial derivative and get the minimum energy:
[0106]
[0107] Therefore, the importance of a neuron can be obtained by the following formula:
[0108]
[0109] Finally, the feature map is enhanced according to the attention mechanism:
[0110]
[0111] The SCSA mechanism is an attention mechanism used in the field of computer vision. It consists of two parts: shared multi-semantic spatial attention (SMSA) and progressive channel self-attention (PCSA). SMSA: Multi-scale depth-wise 1D convolution is used to extract spatial information at different semantic levels from four independent sub-features, and GroupNorm is used to accelerate model convergence and avoid problems such as semantic information leakage. To address the limited receptive field problem caused by feature decomposition and one-dimensional convolution, lightweight shared convolution is used after the depth-wise 1D convolution for feature alignment. Finally, the semantic sub-features are aggregated, normalized by GroupNorm, and multiplied with the original features via the Sigmoid activation function to generate spatial attention. PCSA: Combining progressive compression and channel-specific self-attention mechanisms, it minimizes computational complexity while preserving the spatial prior within SMSA. The self-attention mechanism is used to further explore channel-level similarities, reduce the semantic differences between different sub-features, and generate an attention map.
[0112] In this embodiment, the SCSA mechanism is added to the last C2PSA layer of the backbone network, effectively combining the advantages of channel and spatial attention, fully utilizing multi-semantic information, and thus improving the performance of visual tasks. Furthermore, this embodiment improves the efficiency and practicality of the model by adding a simplified attention mechanism with slicing operations to the convolutional module of the neck network, reducing computational complexity and optimizing structural design, while maintaining or approaching the original attention effect.
[0113] This embodiment provides a bridge defect detection method and drone, featuring comprehensive technical optimizations from data acquisition to real-time processing. In terms of data acquisition, by voxelizing the three-dimensional bridge model, the bridge structure is decomposed into quantifiable spaces. This allows for the generation of a fully covered, non-redundant drone flight path based on a refined spatial network. Compared to traditional manual remote control methods, this eliminates the need for manual remote control of the drone to capture each point, avoiding the randomness and risk of omissions associated with manual operation and improving detection accuracy and monitoring efficiency. In terms of data processing, the defect detection model is deployed directly on the drone, enabling real-time analysis of captured images without transmitting data to the cloud or backend server. This improves the speed and efficiency of defect detection and reduces the data processing pressure on the cloud or backend server. At the model architecture level, an SCSA mechanism is added to the last layer of the backbone network of the defect detection model. This improves the accuracy of defect detection by analyzing the attention weights of the input feature maps. Furthermore, a simplified attention mechanism for slicing operations is introduced in the convolutional module of the neck network. This reduces background interference and improves the accuracy of defect recognition through local feature slicing and cross-layer interaction.
[0114] Optionally, in one embodiment of the present application, the disease detection model is obtained by training with the YOLOv11 model as the benchmark network, wherein the last layer of the backbone network is the C2PSA layer. When the SCSA mechanism is added to the C2PSA layer, the original feature map output by the SPPF module is subjected to spatial attention extraction along the width and height directions respectively to obtain a height attention map and a width attention map, and the height attention map, the width attention map and the original feature map are multiplied to obtain a corresponding spatial attention feature map, and corresponding attention weights are assigned to all channels in the spatial attention feature map based on the channel self-attention mechanism to obtain a channel attention weight matrix, and the spatial attention feature map and the channel attention weight matrix are multiplied to obtain a fused feature map, so that the neck network and the head network can mark the diseases in the captured image based on the fused feature map.
[0115] Specifically, the last C2PSA layer of the backbone network is added to the SCSA (Spatial and Channel Synergistic Attention) mechanism, aiming to effectively combine the advantages of channel and spatial attention, make full use of multi-semantic information, and thus improve the performance of visual tasks.
[0116] Optionally, in an embodiment of the present application, spatial attention extraction is performed on the original feature map output by the SPPF module along the width direction and the height direction respectively, including:
[0117] Decomposing the original feature map along the height direction and the width direction, and performing a global pooling operation to obtain a height feature map and a width feature map;
[0118] Decomposing the height feature map into at least one height sub-feature map, and decomposing the width feature map into at least one width sub-feature map;
[0119] The height sub-feature maps are respectively extracted by using a multi-receptive field shared one-dimensional depth convolution method, and are then spliced, grouped and normalized, and activated by a Sigmoid function to obtain the height attention map;
[0120] The width sub-feature maps are respectively extracted by using a multi-receptive field shared one-dimensional depth convolution method, and are then spliced, grouped and normalized, and activated with a Sigmoid function to obtain the width attention map.
[0121] Specifically, for the input feature map It is decomposed along the height H and width W directions, and then obtained through the global pooling operation process and Then it is further divided into K sub-features, and the number of channels for each sub-feature becomes C / K, and we get and Then to and Use multiple receptive fields to share one-dimensional depth convolution to operate on each sub-feature, and get and Then, the sub-features of height and width dimensions are concatenated respectively, and group normalization (GN) is used to reduce semantic interference. Finally, the sigmoid function is used to activate the spatial attention feature map. The specific expression can be expressed as:
[0122]
[0123] X SMSA =SMSA(X)=S H ×S H ×X;
[0124] Where σ(·) represents the Sigmoid activation function; and Indicates the application of K GN operations along the height and width directions; Concat(·) represents the stacking and splicing operation of the feature map; S H is the spatial attention map along the height direction; S W is the spatial attention map along the width direction; X SMSA This is the final feature map.
[0125] Optionally, in an embodiment of the present application, corresponding attention weights are assigned to all channels in the spatial attention feature map based on a channel self-attention mechanism to obtain a channel attention weight matrix, including:
[0126] After performing average pooling and group normalization operations on the spatial attention feature map, a linear transformation is performed through a two-dimensional depthwise convolution to obtain the corresponding query vector, key vector, and value vector;
[0127] According to the query vector, key vector, and value vector, each channel in the spatial attention feature map is assigned a corresponding channel attention weight, and processed by a Sigmoid activation function to obtain the channel attention weight matrix.
[0128] Specifically, the PCSA module first reduces the size of the feature map output by the previous SMSA module through an average pooling operation. A GN operation and a 1×1 two-dimensional depthwise convolution (DWConv) are then used to linearly transform the input feature map to generate a query vector, a key vector, and a value vector. A self-attention mechanism is then used to assign weights to each channel. Finally, the feature map passes through a Sigmoid activation function and is multiplied by the channel weights to generate a PCSA-enhanced feature map. The specific implementation process is shown in the following formula:
[0129] X p =Avg_pooling(X SMSA );
[0130]
[0131] Among them, X p is the feature map after pooling; Q, K and and N=H; is a 1×1 depthwise separable convolution, Represents the phenomenon projection calculation of Q, K and V.
[0132] The embodiments of the present application effectively combine the advantages of channel and spatial attention by adding the SCSA mechanism to the last C2PSA layer of the backbone network, making full use of multi-semantic information, thereby improving the performance of visual tasks.
[0133] Optionally, in an embodiment of the present application, the convolution of the backbone network and the neck network is replaced by wavelet convolution WTConv. WTConv decomposes the input into components of different frequencies by utilizing wavelet transform. Small kernel convolution is then performed, and each convolution focuses on a different frequency range. At this time, for any input X, the input feature map is decomposed into low-frequency and high-frequency components using Haar wavelet transform, and down-sampled. Small kernel depth-separable convolution is then performed on different frequency components. Finally, the convolution results of different frequency components are weighted and summed, and up-sampled to obtain the final output. Through the wavelet convolution WTConv structure, the feature extraction capability, computational efficiency and anti-interference capability of the model are improved.
[0134] Furthermore, the model structure of the disease detection model is as follows: Figure 7 As shown, it includes a backbone network, a neck network and a head network; wherein the backbone network includes: a first Conv layer, a second Conv layer, a first C3k2 layer, a third Conv layer, a second C3k2 layer, a fourth Conv layer, a third C3k2 layer, a fifth Conv layer, a fourth C3k2 layer, an SPPF layer and a C2PSA layer connected in sequence;
[0135] The neck network includes: the first Upsample layer, the first Concat layer, the fifth C3k2 layer, the second Upsample layer, the second Concat layer, the sixth C3k2 layer, the sixth Conv layer, the third Concat layer, the seventh C3k2 layer, the seventh Concat layer, the fourth Concat layer and the eighth C3k2 layer connected in sequence;
[0136] The head network includes three Detect layers, and the three Detect layers are connected to the sixth C3k2 layer, the seventh C3k2 layer, and the eighth C3k2 layer respectively;
[0137] Among them, the C2PSA layer is connected to the first Upsample layer;
[0138] The first Concat layer is connected to the third C3k2 layer;
[0139] The second Concat layer is connected to the second C3k2 layer;
[0140] The third Concat layer is connected to the fifth C3k2 layer;
[0141] The fourth Concat layer is connected to the C2PSA layer.
[0142] Furthermore, the first C3k2 layer, the second C3k2 layer, the third C3k2 layer, the fourth C3k2 layer, the fifth C3k2 layer, the sixth C3k2 layer, the seventh C3k2 layer and the eighth C3k2 layer are all wavelet convolution WTConv structures; the first Conv layer, the third Conv layer, the fourth Conv layer, the sixth Conv layer and the seventh Conv layer are all Conv-SWS structures.
[0143] This embodiment improves the feature extraction capability of the model by adding the SCSA mechanism to the last layer of the backbone network, replacing some C3k2 layers with WTConv structures, and adding a simplified attention mechanism to some Conv layers, making subsequent disease detection more accurate.
[0144] Furthermore, the defect detection model requires pre-training. This involves pre-collecting a dataset of bridge images containing three types of defects: cracks, efflorescence, and exposed rebar. The defect locations in the images are annotated, and the dataset is divided into training, validation, and test sets in a ratio of 7:2:1. When building a network model suitable for edge computing, the YOLOv11n model was used as the base model and improved upon.
[0145] After training is complete, the following evaluation metrics are used to evaluate the model. The trained model needs to be evaluated using these metrics to verify its accuracy. Typically, four metrics are used to verify the model: accuracy, recall, AP (average precision), and mAP (mean average precision). Additionally, the model size and computational overhead are important factors in the evaluation, measured by parameters and GFLOPs. The calculation formula is as follows:
[0146]
[0147] TP represents the number of correctly detected positive samples, FP represents the number of undetected positive samples, and FN represents the number of incorrectly detected negative samples. AP is the area enclosed by plotting Precision and Recall as the axes. mAP is the average mean average precision (AP) of all target categories and describes the model's detection performance for all target detection categories. A larger mAP value indicates better detection results, better recognition accuracy, and higher recognition precision.
[0148] In addition, an embodiment of the present application also proposes a bridge defect detection device, which includes an edge computing module and a remote server, and the remote server and the edge computing module are communicatively connected.
[0149] The remote server obtains a voxelized model by performing voxel processing on the three-dimensional model of the area to be detected of the predetermined bridge; and plans the flight path of the UAV based on the voxelized model to obtain a flight planning path.
[0150] The edge computing module controls the drone to photograph the bridge according to the flight planning path; and inputs the photographed image into a pre-trained disease detection model for detection to obtain the bridge disease image of the area to be detected.
[0151] Typically, the edge computing module only saves images of detected defects, optimizing storage space. After the drone completes the inspection, the defect images can be exported to a local computer for further analysis and quantification.
[0152] Finally, this embodiment also proposes a drone, comprising a memory and a processor;
[0153] The memory is used to store executable instructions of the processor, and the processor is configured to execute the bridge defect detection method described in the above embodiment by executing the executable instructions.
[0154] In summary, the bridge defect detection method and drone of this embodiment have undergone comprehensive technical optimization from data acquisition to real-time processing. In terms of data acquisition, by voxelizing the three-dimensional bridge model, the bridge structure is decomposed into quantifiable spaces. This enables the generation of a fully covered, non-redundant drone flight path based on a refined spatial network. Compared to traditional manual remote control methods, this eliminates the need for manual remote control of the drone to capture each point, avoiding the randomness and risk of omissions associated with manual operation and improving detection accuracy and monitoring efficiency.
[0155] In terms of data processing, the disease detection model is directly deployed on the drone, realizing real-time analysis of the captured images without the need to transmit data to the cloud or back-end server, thereby improving the recognition speed and efficiency of disease detection and reducing the data processing pressure on the cloud or back-end server; at the model architecture level, the SCSA mechanism is added to the last layer of the backbone network of the disease detection model, and the accuracy of disease detection is improved by analyzing the attention weight of the input feature map; at the same time, a simplified attention mechanism for slicing operations is introduced in the convolution module of the neck network, which reduces background interference and improves the accuracy of disease recognition through local feature slicing and cross-layer interaction.
[0156] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0157] In this application, unless otherwise specified or limited, the terms "mounted," "connected," "connect," "fixed," etc. should be understood broadly. For example, they can refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.
[0158] In this application, unless otherwise expressly specified or limited, when a first feature is “on” or “below” a second feature, it may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Moreover, when a first feature is “above”, “above”, or “above” a second feature, it may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is at a higher level than the second feature. When a first feature is “below”, “below”, or “below” a second feature, it may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is at a lower level than the second feature.
[0159] In the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.
[0160] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A bridge disease detection method, characterized in that: include: Determining a three-dimensional model of the area to be inspected of the bridge, and performing voxel processing on the three-dimensional model of the area to be inspected to obtain a voxelized model; Performing UAV flight path planning based on the voxelized model to obtain a flight planning path; Acquiring images captured by a drone, wherein the drone captures images of the bridge while flying along the planned flight path; The captured image is input into a pre-trained disease detection model for detection to obtain a bridge disease image of the area to be detected, wherein the disease detection model is deployed on the drone, and the SCSA mechanism is added to the last layer of the backbone network of the disease detection model, and a simplified attention mechanism with slicing operation is added to the convolution module of the neck network.
2. The bridge defect detection method according to claim 1, characterized in that: The UAV flight path planning is performed based on the voxelized model, including: Constructing a discrete grid map for the space where the voxelized model is located, and using the grid where the grid map contacts the voxelized model as a viewpoint; An incomplete graph is constructed based on the viewpoints, and each viewpoint is connected on the incomplete graph based on a greedy algorithm to form the flight planning path.
3. The bridge defect detection method according to claim 2, characterized in that: Constructing a discrete grid map of the space where the voxelized model is located, including: respectively obtaining the extreme coordinates of the voxelized model along the x-axis, y-axis, and z-axis in a pre-established three-dimensional coordinate system; A grid map that completely accommodates the voxelized model is established based on the extreme coordinates of the voxelized model along the x-axis, the y-axis, and the z-axis in a pre-established three-dimensional coordinate system, and the voxelized model is mapped to the grid map.
4. The bridge defect detection method according to claim 2, characterized in that: Constructing an incomplete graph based on the viewpoint, including: Obtaining a distance between each viewpoint and the remaining viewpoints based on the grid map, and obtaining a viewpoint direction unit vector corresponding to each viewpoint based on the distance between each viewpoint and the remaining viewpoints; The incomplete graph is constructed according to the viewpoint direction unit vector corresponding to each viewpoint.
5. The bridge defect detection method according to claim 4, characterized in that: Based on the distance between each viewpoint and the rest of the viewpoints, get the viewpoint direction unit vector corresponding to each viewpoint, including: According to the distance between any viewpoint and the remaining viewpoints, obtain the attraction vector of the remaining viewpoints to the viewpoint, wherein the magnitude of the attraction vector is a pre-configured empirical constant divided by the square of the distance between the two viewpoints, and the direction of the attraction vector is the direction from the viewpoint to the remaining viewpoints; Add the attraction vectors of the remaining viewpoints to the viewpoint to obtain the three-dimensional force vector corresponding to each viewpoint; The three-dimensional force vector corresponding to each viewpoint is normalized to obtain the viewpoint direction unit vector corresponding to each viewpoint.
6. The bridge defect detection method according to any one of claims 1 to 5, characterized in that: The disease detection model is obtained by training with the YOLOv11 model as the benchmark network, wherein the last layer of the backbone network is the C2PSA layer. When the SCSA mechanism is added to the C2PSA layer, spatial attention extraction is performed on the original feature map output by the SPPF module along the width and height directions respectively to obtain a height attention map and a width attention map, and the height attention map, the width attention map and the original feature map are multiplied to obtain a corresponding spatial attention feature map, and corresponding attention weights are assigned to all channels in the spatial attention feature map based on the channel self-attention mechanism to obtain a channel attention weight matrix, and the spatial attention feature map and the channel attention weight matrix are multiplied to obtain a fused feature map, so that the neck network and the head network can mark the diseases in the captured image based on the fused feature map.
7. The bridge defect detection method according to claim 6, characterized in that: The original feature map output by the SPPF module is subjected to spatial attention extraction along the width and height directions, including: Decomposing the original feature map along the height direction and the width direction, and performing a global pooling operation to obtain a height feature map and a width feature map; Decomposing the height feature map into at least one height sub-feature map, and decomposing the width feature map into at least one width sub-feature map; The height sub-feature maps are respectively extracted by using a multi-receptive field shared one-dimensional depth convolution method, and are then spliced, grouped and normalized, and activated by a Sigmoid function to obtain the height attention map; The width sub-feature maps are respectively extracted by using a multi-receptive field shared one-dimensional depth convolution method, and are then spliced, grouped and normalized, and activated with a Sigmoid function to obtain the width attention map.
8. The bridge defect detection method according to claim 6, characterized in that: Based on the channel self-attention mechanism, corresponding attention weights are assigned to all channels in the spatial attention feature map to obtain a channel attention weight matrix, including: After performing average pooling and group normalization operations on the spatial attention feature map, a linear transformation is performed through a two-dimensional depthwise convolution to obtain the corresponding query vector, key vector, and value vector; According to the query vector, key vector, and value vector, each channel in the spatial attention feature map is assigned a corresponding channel attention weight, and processed by a Sigmoid activation function to obtain the channel attention weight matrix.
9. The bridge defect detection method according to claim 1, characterized in that: The convolution module is a Conv-SWS structure, which is suitable for spatial slicing and / or channel slicing of the input feature map to obtain at least one slice map, and evaluate the importance of each slice map based on a preconfigured energy function, and perform feature enhancement on the feature map of the input Conv-SWS structure according to the importance of all slice maps.
10. A drone, characterized in that: including memory and processor; The memory is used to store executable instructions of the processor, and the processor is configured to execute the bridge defect detection method according to any one of claims 1 to 9 by executing the executable instructions.
Citation Information
Patent Citations
Bridge structure displacement detection method based on unmanned aerial vehicle system
CN116758149A
YOLOv6-based bridge disease inspection method and system
CN117151680A
Road disease monitoring method and system based on images acquired by unmanned aerial vehicle, and medium
CN118506217A
Cited By
Expressway disease size and area detection method based on binocular vision
CN121353381A