Unmanned aerial vehicle tower crane corrosion detection method based on improved YOLOv11
By improving the YOLOv11 model and combining with the drone platform, the tower crane corrosion detection system was built, which solved the problem of insufficient accuracy of tower crane corrosion detection in complex backgrounds, and achieved efficient and safe automated detection.
Patent Information
- Application Number
- CN202510423873.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-04
AI Technical Summary
The existing drone tower crane corrosion detection system has insufficient detection accuracy under complex backgrounds, which is prone to error and missed inspections, and manual inspections are dangerous and inefficient.
The improved YOLOv11 model is adopted, and the C3k2 module in the backbone network is replaced by the C3k2_DCNv2 module, the SPPF module is replaced by the AIFI module in RT-DETR, and the DyHead module is introduced in front of the head network to build a tower crane corrosion detection model and combine it with the drone platform for automated detection.
The accuracy and efficiency of tower crane corrosion detection are improved, the leakage detection rate and error detection rate are reduced, and efficient and safe tower crane corrosion detection is achieved.
Smart Images

Figure CN120259922A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and particularly to an unmanned aerial vehicle (UAV) tower crane corrosion detection method based on improved YOLOv11. Background Art
[0002] In the modern industrial and construction fields, tower cranes are indispensable key equipment. However, due to long-term exposure to harsh environmental conditions, the metal structures of tower cranes are prone to corrosion, which not only shortens their service life but may also lead to serious safety hazards. Therefore, regular corrosion detection and maintenance are crucial to ensure the safe operation of tower cranes.
[0003] Tower cranes are usually at high altitudes and have complex structures. Conducting corrosion inspections manually is a time-consuming and dangerous task. Inspectors need to climb to inaccessible heights and rely on visual inspection to identify corrosion areas, which is not only inefficient but may also miss problems or result in misjudgments. With the rapid development of UAV technology, the target detection systems equipped on UAVs are gradually being used in tower crane corrosion detection. UAVs can flexibly approach tower cranes, collect high-resolution images, and effectively improve the detection efficiency. However, corrosion features are complex, usually manifested as fine rust spots, irregular changes, and color differences, etc. This makes it easy for existing detection systems to have false detections and missed detections under the variable lighting conditions and complex backgrounds at construction sites.
[0004] To address these challenges, the present invention proposes an unmanned aerial vehicle tower crane corrosion detection system based on improved YOLOv11. By optimizing the YOLOv11 model, the detection accuracy and efficiency for corrosion areas under complex backgrounds are improved. This system combines the flexible mobility of UAVs with the powerful analysis ability of deep learning, providing an efficient, safe, and reliable solution for tower crane corrosion detection and maintenance. Summary of the Invention
[0005] The present invention aims to make up for the deficiencies in the detection accuracy of existing UAV tower crane corrosion detection in construction sites, and proposes an unmanned aerial vehicle tower crane corrosion detection method based on improved YOLOv11. This method has the characteristics of high detection efficiency, strong safety, and high-precision detection, providing an efficient solution for tower crane corrosion detection in construction sites.
[0006] An unmanned aerial vehicle tower crane corrosion detection method based on improved YOLOv11, the process is as Figure 1 shown, and includes the following steps:
[0007] Step 1: Start the tower crane corrosion detection task. The UAV loads the preset route file and takes off, and heads to the detection points of each tower crane to be inspected;
[0008] Step 2: During flight, the UAV carrying the gimbal camera continuously captures tower crane videos and transmits the tower crane video data to the ground control center;
[0009] Step 3: Input the tower crane video data into the constructed tower crane corrosion detection model. The model analyzes the tower crane structure, automatically identifies the corrosion areas and marks the relevant information;
[0010] Step 4: Visualize the detection results on the ground monitoring interface and upload the detected corrosion conditions to the cloud for long-term storage;
[0011] Step 5: After completing the corrosion detection task, the UAV returns safely and ends this detection operation.
[0012] For the flight path file proposed in Step 1, its construction steps are as follows:
[0013] Step 1.1: Enter the structural linear parameters and 3D model of the tower crane to be inspected into the flight control software;
[0014] Step 1.2: Set a patrol waypoint every 3 meters in the tower crane inspection area to ensure the coverage rate of corrosion inspection;
[0015] Step 1.3: Set the starting point and return point of the UAV, determine the video acquisition resolution as 1080p, and configure the video frame rate as 60fps at the same time to ensure the smoothness of the acquired images;
[0016] Step 1.4: Generate the UAV flight route through the flight control software, check the route, and save the flight path file in KML format.
[0017] Furthermore, the detection areas proposed in Step 1.3 include: the main chord of the boom, the main chord of the tower body section, the root of the tower cap, the top connection tie rod seat, the connection of the balance arm, and the connection of the slewing bearing seat.
[0018] For the tower crane corrosion detection model proposed in Step 2, its construction steps are as follows:
[0019] Step 2.1: Establish a tower crane corrosion data set, including collecting tower crane surface images covering various lighting conditions, different angles and distances, and different corrosion degrees, and using annotation tools to annotate the corrosion parts, generate annotation files corresponding to the images, and divide them into a training set, a validation set, and a test set according to the ratio of 8:1:1;
[0020] Step 2.2: Construct an improved YOLOv11 network, as Figure 2As shown in the figure, the network is based on YOLOv11, replacing the two C3k2 modules at the end of its backbone network with C3k2_DCNv2 modules as the feature extraction module, and using the AIFI module in RT-DETR to replace the SPPF module in the backbone network; a DyHead module is constructed in front of the head network;
[0021] Step 2.3: Use the improved YOLOv11 model to train the tower crane corrosion dataset to obtain the optimal tower crane corrosion detection model.
[0022] Furthermore, the C3k2_DCNv2 module proposed in Step 2.2, as Figure 3 shown, its construction structure is: the C3k2 module structure includes a CBS module, a Split module, n C3 modules, a Concat module, and a CBS module connected in sequence. Each C3 module stacks n Bottleneck modules, and replaces the second 3x3 convolution in the Bottleneck module with DCNv2 to form the C3k2_DCNv2 module.
[0023] DCNv2 enables the sampling points of the convolution kernel to be dynamically adjusted by introducing learnable offsets, so as to adaptively capture the shape and pose changes of the target. This adaptive sampling mechanism makes DCNv2 show stronger robustness when dealing with irregular and non-uniform corrosion, thus achieving higher accuracy and efficiency in detection and recognition tasks.
[0024] Furthermore, the AIFI module proposed in Step 2.2, as Figure 4 shown, its construction structure is: the AIFI module receives the high-level features extracted by the backbone network, combines them with the position encoding, and generates Key, Query, and Value through linear transformation; then uses multi-head attention to perform fine-grained interaction between features of the same scale to enhance the feature representation; then the output of the multi-head attention is first added to the input features through a residual connection and subjected to layer normalization processing. The layer-normalized features obtained are sent to the feed-forward network to learn more complex feature representations; finally, the output of the feed-forward network is added to the layer-normalized features again through a residual connection and layer normalization is performed, so as to output the features after deep fusion. Query, Key, and Value are calculated by the following formulas:
[0025] Q = XW Q , K = XW K , V = XW V
[0026] Where Q represents the Query query vector, K represents the Key key vector, V represents the Value value vector, WQ, WK, and WV are the input feature mapping matrices of Q, K, and V on each pixel respectively, and X represents the input image features.
[0027] The SPPF module extracts features through multi-scale feature fusion, while the AIFI module focuses on feature fusion of high-level features, which can effectively reduce the computational cost; the AIFI module enhances the network's focusing ability on key features through the multi-head attention mechanism, promotes richer feature fusion, and can reduce the missed detection rate and false detection rate.
[0028] Furthermore, the DyHead module proposed in the step 2.2, as Figure 5 and Figure 6 shown, its construction process is as follows: The DyHead module is constructed in front of the YOLOv11 head network. The DyHead module is composed of 4 stacked DyHead Block modules. Each DyHead Block module can improve the model's detection ability and feature expression ability in terms of scale perception, spatial perception, and task perception through the self-attention mechanism, which is achieved through the following formula:
[0029] W(F) = π C (π S (π L (F)·F)·F)·F
[0030] where F represents an input three-dimensional tensor L×S×C, and π L , π s and π c are three independent attention functions applied to the dimensions L, S, and C respectively.
[0031] In the scale perception attention mechanism, the input features first go through global average pooling to extract the global context information of the feature map; then, the information of different channels is integrated through a 1×1 convolutional layer to achieve cross-channel information interaction and enhance the feature expression ability; next, the ReLU activation function is applied to the output of the 1×1 convolutional layer to introduce non-linearity, enabling the model to learn more complex feature representations; finally, the hard-sigmoid function processes the output of ReLU, restricting the value to the range of 0 to 1 to generate the scale perception weight, which is achieved through the following formula:
[0032]
[0033] where, is the hard-sigmoid function, and f(·) is a linear function approximated by a 1×1 convolutional layer.
[0034] In the spatial perception attention mechanism, first, key spatial positions in the feature map are selected using indices; then, these positions are subjected to feature extraction through a 3×3 convolutional layer to capture local details; next, spatial offsets are calculated to adjust the sampling points, enabling the model to focus on more discriminative spatial regions; finally, the sigmoid function is applied to generate weights for these spatial positions, and this process is achieved through the following formula:
[0035]
[0036] where K is the number of sparse sampling positions, p k is the sampling position, Δp k is the spatial offset, p k +Δp k is the position adjusted by the self-learned spatial offset Δp k Δm k is the self-learned importance scalar at the sampling position p k .
[0037] In the task perception attention mechanism, the input features first undergo global average pooling to extract global information; then, the pooled features are processed through two fully connected layers, with the ReLU activation function in between to introduce non-linearity, mapping the global pooled features to the task-specific weight space; finally, normalization is performed through the activation function to generate weights adapted to different detection tasks, and this process is achieved through the following formula:
[0038] π C (F)·F = max(α 1 (F)·F C +β 1 (F), α 2 (F)·F C +β 2 (F))
[0039] where F C is the feature slice of the c-th channel, θ(·) = [α 1 , α 2 , β 1 , β 2 Τ is the hyperfunction for learning to control the activation threshold.
[0041] The visualization of the detection results proposed in step 3 includes the following steps:
[0042] Step 3.1: The ground monitoring system uses the OpenCV library to capture the drone video stream transmitted via the HTTP protocol;
[0043] Step 3.2: Extract each frame of the image from the UAV video stream and input it into the tower crane corrosion detection model for analysis;
[0044] After the tower crane corrosion detection model identifies the corrosion area in the image, draw a bounding box and label on the image, and display the processed image on the ground monitoring interface in real time for users to observe and evaluate. The beneficial effects of the present invention are as follows:
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] The present invention utilizes a customized flight trajectory and advanced target detection technology to realize the visualization of the detection process and digitally store the corrosion information, thereby realizing low-cost and efficient full-coverage corrosion detection of tower cranes.
[0047] The present invention improves the YOLOv11 network structure to make it suitable for the detection of tower crane corrosion images from the perspective of UAVs, improves the accuracy of tower crane corrosion detection, enhances the robustness of the network, and can effectively reduce interference factors such as irregular corrosion area shapes, complex backgrounds, image noise, and light reflection. Description of the Drawings
[0048] Figure 1 is a flowchart of the UAV tower crane corrosion detection method based on the improved YOLOv11;
[0049] Figure 2 is the structural diagram of the improved YOLOv11 network architecture;
[0050] Figure 3 is the structural diagram of the C3k2_DCNv2 module;
[0051] Figure 4 is the structural diagram of the AIFI module;
[0052] Figure 5 is the structural diagram of the DyHeadBlock module;
[0053] Figure 6 is a schematic diagram of the stacking method of the DyHeadBlock module. Detailed Embodiments
[0054] The specific implementation of the present invention is as follows:
[0055] A UAV tower crane corrosion detection method based on the improved YOLOv11, the process is as Figure 1 shown, including the following steps:
[0056] Step 1: Start the tower crane corrosion detection task, the UAV loads the preset route file and takes off, and goes to the detection points of each tower crane to be inspected;
[0057] Step 2: During flight, the UAV carrying the gimbal camera continuously captures tower crane videos and transmits the tower crane video data to the ground control center;
[0058] Step 3: Input the tower crane video data into the constructed tower crane corrosion detection model. The model analyzes the tower crane structure, automatically identifies the corrosion areas, and marks the relevant information;
[0059] Step 4: Visualize the detection results on the ground monitoring interface and upload the detected corrosion conditions to the cloud for long-term storage;
[0060] Step 5: After completing the corrosion detection task, the UAV returns safely and ends this detection operation.
[0061] The flight route file proposed in Step 1 is constructed as follows:
[0062] Step 1.1: Enter the structural linear parameters and 3D model of the tower crane to be inspected into the flight control software;
[0063] Step 1.2: Set an inspection waypoint every 3 meters in the tower crane inspection area to ensure the coverage rate of corrosion inspection;
[0064] Step 1.3: Set the starting point and return point of the UAV, determine the video acquisition resolution as 1080p, and configure the video frame rate as 60fps to ensure the smoothness of the acquired images;
[0065] Step 1.4: Generate the UAV flight route through the flight control software, check the route, and save the route file in KML format.
[0066] Furthermore, the detection areas proposed in Step 1.3 include: the main chord of the boom, the main chord of the tower body section, the root of the tower cap, the top connection pull rod seat, the connection of the balance arm, and the connection of the slewing bearing seat.
[0067] The tower crane corrosion detection model proposed in Step 2 is constructed as follows:
[0068] Step 2.1: Establish a tower crane corrosion data set, including collecting tower crane surface images covering various lighting conditions, different angles and distances, and different corrosion degrees, and using annotation tools to annotate the corrosion parts, generating annotation files corresponding to the images, and dividing them into a training set, a validation set, and a test set according to the ratio of 8:1:1;
[0069] Step 2.2: Construct an improved YOLOv11 network, as Figure 2As shown in the figure, the network is based on YOLOv11, replacing the two C3k2 modules at the end of its backbone network with C3k2_DCNv2 modules as the feature extraction module, and using the AIFI module in RT-DETR to replace the SPPF module in the backbone network; a DyHead module is constructed in front of the head network;
[0070] Step 2.3: Use the improved YOLOv11 model to train the tower crane corrosion dataset to obtain the optimal tower crane corrosion detection model.
[0071] Furthermore, the C3k2_DCNv2 module proposed in Step 2.2 is as Figure 3 shown, and its construction structure is: the C3k2 module structure includes a CBS module, a Split module, n C3 modules, a Concat module, and a CBS module connected in sequence. Each C3 module stacks n Bottleneck modules. Replace the second 3x3 convolution in the Bottleneck module with DCNv2 to form the C3k2_DCNv2 module.
[0072] By introducing learnable offsets, DCNv2 enables the sampling points of the convolution kernel to be dynamically adjusted, thereby adaptively capturing the shape and pose changes of the target. This adaptive sampling mechanism makes DCNv2 show stronger robustness when dealing with irregular and non-uniform corrosion, thus achieving higher accuracy and efficiency in detection and recognition tasks.
[0073] Furthermore, the AIFI module proposed in Step 2.2 is as Figure 4 shown, and its construction structure is: the AIFI module receives the high-level features extracted by the backbone network, combines them with the position encoding, and generates Key, Query, and Value through linear transformation; then uses multi-head attention to perform fine-grained interaction between features of the same scale to enhance the feature representation; then the output of the multi-head attention is first added to the input features through a residual connection and undergoes layer normalization processing. The layer-normalized features obtained are fed into the feed-forward network to learn more complex feature representations; finally, the output of the feed-forward network is added to the layer-normalized features again through a residual connection and layer normalization is performed, thereby outputting the features after deep fusion. Query, Key, and Value are calculated by the following formulas:
[0074] Q = XW Q , K = XW K , V = XW V
[0075] Where Q represents the Query query vector, K represents the Key key vector, V represents the Value value vector, WQ, WK, and WV are the input feature mapping matrices of Q, K, and V on each pixel respectively, and X represents the input image features.
[0076] The SPPF module extracts features through multi-scale feature fusion, while the AIFI module focuses on feature fusion of high-level features, which can effectively reduce the computational cost; the AIFI module enhances the network's focusing ability on key features through the multi-head attention mechanism, promotes richer feature fusion, and can reduce the missed detection rate and false detection rate.
[0077] Furthermore, the DyHead module proposed in step 2.2, as Figure 5 and Figure 6 shown, its construction process is as follows: The DyHead module is constructed before the YOLOv11 head network. The DyHead module is composed of 4 stacked DyHead Block modules. Each DyHead Block module can improve the model's detection ability and feature expression ability in terms of scale perception, spatial perception, and task perception through the self-attention mechanism, which is achieved through the following formula:
[0078] W(F) = π C (π S (π L (F)·F)·F)·F
[0079] where F represents an input three-dimensional tensor L×S×C, and π L , π s and π c are three independent attention functions applied to dimensions L, S, and C respectively.
[0080] In the scale perception attention mechanism, the input features first go through global average pooling to extract the global context information of the feature map; then, the 1×1 convolutional layer is used to integrate information from different channels, realizing cross-channel information interaction and enhancing the feature expression ability; next, the ReL activation function is applied to the output of the 1×1 convolutional layer to introduce non-linearity, enabling the model to learn more complex feature representations; finally, the hard-sigmoid function processes the output of ReLU, restricting the value to the range of 0 to 1 to generate the scale perception weight, and this process is achieved through the following formula:
[0081]
[0082] where, is the hard-sigmoid function, and f(·) is a linear function approximated by the 1×1 convolutional layer.
[0083] In the spatial perception attention mechanism, first, key spatial positions in the feature map are selected using indices; then, a 3×3 convolutional layer is used to extract features at these positions to capture local details; next, spatial offsets are calculated to adjust the sampling points, enabling the model to focus on more discriminative spatial regions; finally, a sigmoid function is applied to generate weights for these spatial positions, and this process is achieved through the following formula:
[0084]
[0085] where K is the number of sparse sampling positions, p k is the sampling position, Δp k is the spatial offset, p k +Δp k is the position adjusted by the self-learned spatial offset Δp k Δm k is the self-learned importance scalar at the sampling position p k .
[0086] In the task perception attention mechanism, the input features first undergo global average pooling to extract global information; then, the pooled features are processed through two fully connected layers, with a ReLU activation function in between to introduce non-linearity, mapping the global pooled features to a task-specific weight space; finally, normalization is performed through an activation function to generate weights adapted to different detection tasks, and this process is achieved through the following formula:
[0087] π C (F)·F = max(α 1 (F)·F C +β 1 (F), α 2 (F)·F C +β 2 (F))
[0088] where F C is the feature slice of the c-th channel, θ(·) = [α 1 , α 2 , β 1 , β 2 Τ is the hyperfunction for learning to control the activation threshold.
[0089] The visualization of the detection results proposed in step 3 includes the following steps:
[0090] Step 3.1: The ground monitoring system uses the OpenCV library to capture the drone video stream transmitted via the HTTP protocol;
[0091] Step 3.2: Extract each frame of the image from the UAV video stream and input it into the tower crane corrosion detection model for analysis;
[0092] Step 3.3: After the tower crane corrosion detection model identifies the corrosion area in the image, draw a bounding box and label on the image, and display the processed image on the ground monitoring interface in real time for users to observe and evaluate.
[0093] The above are the preferred embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, who makes changes or substitutions according to the technical solution of the present invention, should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. An unmanned aerial vehicle tower crane corrosion detection method based on improved YOLOv11, characterized in that, It includes the following steps: Step 1: Start the tower crane corrosion detection task. The drone loads the preset route file and takes off, and heads to the detection points of each tower crane to be inspected; Step 2: During the flight, the drone equipped with a gimbal camera continuously shoots tower crane videos and transmits the tower crane video data to the ground control center; Step 3: Input the tower crane video data into the constructed tower crane corrosion detection model. The model analyzes the tower crane structure, automatically identifies the corrosion areas and marks the relevant information; Step 4: Visualize the detection results on the ground monitoring interface and upload the detected corrosion conditions to the cloud for long-term storage; Step 5: After completing the corrosion detection task, the drone returns safely and ends this detection operation.
2. The method for detecting corrosion of tower cranes by drones based on improved YOLOv11 according to claim 1, characterized in that The route file proposed in Step 1 includes the following steps: Step 1: Enter the structural linear parameters and 3D model of the tower crane to be inspected in the flight control software; Step 2: Set a patrol waypoint every 3 meters in the tower crane inspection area to ensure the coverage rate of key parts; Step 3: Set the starting point and return point of the drone, determine the video acquisition resolution to be 1080p, and configure the video frame rate to be 60fps to ensure the smoothness of the acquired images; Step 4: Generate the drone flight route through the flight control software, check the route, and save the route file in KML format.
3. The route file according to claim 2, wherein The inspection area proposed in Step 2 includes: the main chord of the boom, the main chord of the tower body section, the root of the tower cap, the top connection pull rod seat, the connection of the balance arm, and the connection of the slewing bearing seat.
4. The method for detecting corrosion of tower cranes by drones based on improved YOLOv11 according to claim 1, wherein The tower crane corrosion detection model proposed in Step 3 includes the following steps: Step 1: Establish a tower crane corrosion data set, including collecting tower crane surface images covering various lighting conditions, different angles and distances, and different corrosion degrees, and using annotation tools to annotate the corrosion parts, generating annotation files corresponding to the images, and dividing them into a training set, a validation set, and a test set according to the ratio of 8:1:1; Step 2: Construct an improved YOLOv11 network. The network is based on YOLOv11, replaces the C3k2 module in its backbone network with the C3k2_DCNv2 module as the feature extraction module, and uses the AIFI module in RT-DETR to replace the SPPF module in the backbone network; construct a DyHead module before the head network; Step 3: Use the improved YOLOv11 network to train the tower crane corrosion data set to obtain the tower crane corrosion detection model.
5. The tower crane corrosion detection model according to claim 4, wherein In Step 2, there are four C3k2 modules in the backbone network of the original YOLOv11 network. Replace the two C3k2 modules at the end of the sequence with C3k2_DCNv2 modules to enhance the feature extraction ability and improve the recognition accuracy of tower crane corrosion features.
6. The C3k2_DCNv2 module according to claim 5, wherein Use DCNv2 to replace the second 3x3 convolution of the Bottleneck module in the C3 module, aiming to improve the model's ability to extract complex image features.
7. The tower crane corrosion detection model according to claim 4, wherein In Step 2, construct a DyHead module before the head network of the original YOLOv11 network to enhance the network's object detection ability and improve the recognition accuracy of tower crane corrosion features.
8. The DyHead module according to claim 7, characterized in that, This module is composed of four stacked DyHead Block modules, aiming to improve the detection accuracy and robustness through multi-layer feature fusion to better adapt to the tower crane corrosion detection task.
9. The method for detecting corrosion of the tower crane of the unmanned aerial vehicle based on the improved YOLOv11 according to claim 1, wherein The visualization of the detection results proposed in step 4 includes the following steps: Step 1: The ground monitoring system uses the OpenCV library to capture the drone video stream transmitted via the HTTP protocol; Step 2: Extract each frame of the image from the drone video stream and input it into the tower crane corrosion detection model for analysis; Step 3: After the tower crane corrosion detection model identifies the corrosion area in the image, draw bounding boxes and labels on the image, and display the processed image on the ground monitoring interface in real time for users to observe and evaluate.
Citation Information
Cited By
Hidden danger identification method and device for tower crane inspection, electronic equipment and storage medium
CN120580236A
A hidden danger identification method and device for tower crane inspection, an electronic device, and a storage medium
CN120580236B
Pipe inner surface defect detection method based on improved YOLOv11
CN122115433A