Unmanned aerial vehicle tower crane bolt state detection method based on multi-scale dynamic edge attention enhancement mechanism
By introducing a multi-scale dynamic edge enhancement attention mechanism and an adaptive anchor frame accuracy loss function in the YOLOv8s model, the status detection of the drone tower crane bolt is optimized, solving the problem of insufficient detection accuracy in complex backgrounds, and achieving more efficient tower crane bolt status recognition.
Patent Information
- Application Number
- CN202510357772.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing drone tower crane bolt state detection methods are insufficient in complex backgrounds to detect the accuracy, and the traditional algorithm model is large and the reasoning time is long, so it is not suitable for real-time applications.
The multi-scale dynamic edge enhancement attention mechanism (MDEA) and adaptive anchor box accuracy loss function (AAPL) based on the YOLOv8s model are adopted to optimize the model training process and improve the detection ability of tower crane bolt status through multi-scale pooling, edge enhancement and dynamic weight adjustment.
In a complex context, the model's detection accuracy and training speed of the tower crane bolt state are significantly improved, and the generalization ability and detection accuracy of the model are enhanced.
Smart Images

Figure CN120298772A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of tower crane bolt state detection, and specifically relates to a method for detecting the state of tower crane bolts by an unmanned aerial vehicle based on a multi-scale dynamic edge enhancement attention mechanism. Background Art
[0002] With the progress of technology, machine inspection technology has gradually become the focus of research, which significantly improves the efficiency and accuracy of detection through automated means. With the development of technology, visual inspection by unmanned aerial vehicles has achieved certain results in the fields of electric power and transportation. Unmanned aerial vehicle inspection can effectively prevent the problems of misdetection and missed detection existing in manual inspection, and the efficiency is greatly improved compared with manual inspection. Therefore, introducing unmanned aerial vehicle image recognition technology into the detection of the tightening degree of the connecting bolts of the tower crane standard section has potential application value and important practical significance.
[0003] In the field of object detection, traditional algorithms are mainly divided into two categories: two-stage detection and one-stage detection. Two-stage detection methods such as R-CNN, Fast R-CNN, and Faster R-CNN generate candidate regions through a Region Proposal Network (RPN), and then use a deep learning model for classification and bounding box regression to improve the accuracy of detection. However, the models of these methods are relatively large and the inference time is long, which is not suitable for application scenarios that require real-time response. In contrast, one-stage detection algorithms such as YOLO, SSD, and RetinaNet integrate the object detection process in a unified neural network, and can directly output the category and location information of the object without going through the region proposal stage, thus having an obvious advantage in inference speed and being more suitable for real-time application scenarios. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for detecting the state of tower crane bolts by an unmanned aerial vehicle based on a multi-scale dynamic edge enhancement attention mechanism. This method first collects picture samples of the connecting bolts of the tower crane standard section by an unmanned aerial vehicle to construct an image sample set of tower crane bolts under multiple perspectives. By analyzing the image features of the connecting bolts of the tower crane standard section, it is found that there is a large difference in color between them and the standard section, and their shape features are relatively obvious. Therefore, by introducing an attention mechanism on the basis of the YOLOv8s model and improving the anchor box loss function in the model training process, the model can enhance the accuracy of judging the state of the connecting bolts of the tower crane standard section in complex scenarios.
[0005] To achieve the above object, the technical solution of the present invention is: a method for detecting the state of tower crane bolts based on a multi-scale dynamic edge-enhanced attention mechanism. Image samples of the connecting bolts of the standard section of the tower crane are obtained by shooting with a drone, and a sample library under multiple scenarios is constructed; based on the YOLOv8 model, a multi-scale dynamic edge-enhanced attention mechanism MDEA is proposed, and an adaptive anchor box precision loss function AAPL is proposed as the loss function during the model training process; the trained model is used to judge the state and locate the image of the connecting bolts of the standard section of the tower crane to be detected.
[0006] Further, the method includes the following steps:
[0007] S1. Use a drone to shoot the states of the connecting bolts of the standard section of the tower crane at the construction site under different construction scenarios, and obtain the tower crane bolt images from different perspectives;
[0008] S2. Use the labelimg tool to annotate the obtained tower crane bolt image data, label the normal state as normal, the loose state as loose, and the missing state as lack, store the label name and the annotation position to construct a data set and divide it into a training set and a test set;
[0009] S3. Take the YOLOv8s network as the initial model, and add a multi-scale dynamic edge-enhanced attention mechanism MDEA (Multi-Scale Dynamic Edge-Enhanced Attention, MDEA) to the backbone part of the YOLOv8 network;
[0010] S4. Introduce an adaptive anchor box precision loss function AAPL (Adaptive Anchor Precision Loss, AAPL) as the loss function during the model training process;
[0011] S5. Input the constructed data set into the model proposed in steps S3 and S4 for training to obtain a model weight file;
[0012] S6. Use the trained model MDEA-YOLOv8s to judge the state and locate the image of the connecting bolts of the standard section of the tower crane to be detected.
[0013] Further, different perspectives include frontal view, upward view, downward view, relatively far distance, and relatively close distance views.
[0014] Further, the data set is randomly divided into a training set and a test set according to a ratio of 8:2.
[0015] Furthermore, MDEA includes three mechanisms: multi-scale spatial attention, edge enhancement module, and dynamic weight adjustment.
[0016] Furthermore, the specific implementation of MDEA is as follows:
[0017] The multi-scale spatial attention captures spatial features at different scales through multi-scale pooling operations, using max pooling and average pooling with kernel sizes of 3×3, 5×5, and 7×7 to enhance the model's ability to extract detailed information of small targets; specifically, for the input feature map X, the multi-scale pooling operation is first performed as shown in formulas (1) and (2):
[0018] M ki = MaxPool ki (X) (19)
[0019] A ki = AvgPool ki (X) (20)
[0020] where ki ∈ {3, 5, 7} is the size of the pooling kernel, M ki is the global max pooling feature of spatial attention corresponding to the pooling kernel size, A ki is the average pooling feature of spatial attention corresponding to the pooling kernel size, AvgPool(X) is average pooling, MaxPool(X) is global max pooling, and the multi-scale spatial attention is obtained by concatenating the features after multi-scale pooling as shown in formula (3):
[0021] F scale = Concat(M k1 , A k1 ,..., M kn , A kn ) (21)
[0022] Subsequently, an edge enhancement module is introduced. The Sobel operator is used to extract edge features, which are then concatenated with the multi-scale spatial attention. The edge features are extracted using the Sobel operator as shown in formula (4):
[0023] Edge = Sobel(X) (22)
[0024] where Edge is the extracted edge feature, and then the edge feature is concatenated with the multi-scale spatial attention as shown in formula (5):
[0025] F edge = Concat(F scale, , Edge) (23)
[0026] The normalized F is calculated through 1×1 convolution and the softmax functionedge To obtain the dynamic weights of image features as shown in Equation (6):
[0027] ω s = Softmax(Conv 1×1 (Norm(F edge ))) (24)
[0028] Where Norm is the normalization calculation of the maximum and minimum values.
[0029] The spatial attention map is generated through 1×1 convolution and the Sigmoid function, and combined with the dynamic weights to further enhance the model's detection ability for small targets; the expression of the spatial attention map is shown in Equation (7):
[0030] A s = σ(Conv 1×1 (F edge )) ⊙ ω s (25)
[0031] Where σ is the Sigmoid function, and ⊙ is the Hadamard product, that is, element-wise multiplication.
[0032] The channel attention map is generated through global max pooling and global average pooling, and then the channel attention weights are generated through a fully connected layer to enhance the feature representation of important channels and suppress the feature of unimportant channels. The global max pooling and global average pooling are shown in Equations (8) and (9):
[0033] Y avg = AvgPool(X) (26)
[0034] Y max = MaxPool(X) (27)
[0035] Where Y avg is the average pooling feature of channel attention, and Y max is the max pooling feature of channel attention.
[0036] Then the channel attention weights are generated through a fully connected layer as shown in Equation (10):
[0037] A c = σ(FC2(ReLU(FC1(Y avg + Y max )))) (28)
[0038] Among them, FC1 is the first fully connected layer, which maps the input feature vector to a lower-dimensional space. FC2 is the second fully connected layer used to map the output of FC1 back to the original channel dimension. ReLU (Rectified Linear Unit) is an activation function, and its calculation method is shown in Equation (11):
[0039] ReLU = max(0, x) (29)
[0040] Finally, the channel attention map and the spatial attention map are combined to weight the input feature map, and the final output feature map is obtained as shown in Equation (30):
[0041] F out = A c ⊙(A s ⊙X) (31).
[0042] Furthermore, AAPL optimizes the anchor box loss to better focus on defective bolt samples under the condition that there is an obvious imbalance in tower crane bolt samples in complex scenarios.
[0043] Furthermore, the anchor box loss function is implemented based on the center point distance between anchor boxes, and the calculation method of the anchor box loss function is as follows:[[]]
[0044] The performance of object detection and the accuracy of evaluation are highly related to the design of the loss function. The purpose of bounding box regression is to make the bounding box output by the detector approach the true bounding box by comparing it with the true bounding box; IoU (intersection over union) is used to evaluate the loss of the prediction box in the field of object detection, and its formula is shown in Equation (13):
[0045]
[0046] Where B pred is the predicted bounding box region; B gt is the labeled bounding box region;
[0047] The basic IoU assumes that all labels are good by default and does not consider the harmful gradients generated during the training process caused by possible low-quality labels; instead, it focuses on strengthening the fitting ability of the predicted bounding box region to the labeled bounding box region; at the same time, its ability to distinguish difficult-to-classify samples is weak, and its perception ability in complex environments is insufficient; in addition, the problem of inconsistent numbers of samples for different classes is not considered either; therefore, to address these shortcomings, AAPL is designed to reduce the impact of low-quality labeled bounding boxes on model training; the calculation of AAPL is shown in Equation (14):
[0048] L AAPL = ωc ·(1 - L IoU ) γ ·L WIoU (33)
[0049] In the formula, γ is the adjustment hyperparameter for difficult - to - classify samples, ω c is the sample weight of samples of class c, and its expression is shown in Equation (15):
[0050]
[0051] where N is the total number of samples, N c is the number of samples of the class, L IoU is the loss function for measuring the overlap degree between the predicted bounding box and the ground - truth bounding box, L WIoU is the loss function that introduces additional geometric features on the basis of L IoU The calculation formulas (16) of L IoU and L WIoU are shown in Equation and Equation (17):
[0052] L IoU = 1 - IoU (35)
[0053]
[0054] where (x pred , y pred ), (x gt , y gt ) represent the center - point coordinates of the predicted bounding box and the ground - truth bounding box respectively, w c and h c are the width and height of the smallest annotation bounding box, * means separating the smallest bounding box from the gradient calculation to reduce the adverse effects generated by model training; β represents the outlier degree of the predicted bounding box, and the outlier degree β of the predicted bounding box is shown in Equation (18):
[0055]
[0056] The closer the outlier value is to 1, the closer its annotation effect is to the average level. Assigning a smaller gradient gain to samples with smaller outlier values is beneficial to increasing the generalization ability of the model, and assigning a smaller gradient gain to samples with too large outlier values is also beneficial to reducing the large impact of low - quality annotation samples on the model accuracy.
[0057] The present invention also provides a UAV tower crane bolt state detection system based on a multi - scale dynamic edge - enhanced attention mechanism, including a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the method steps described above can be implemented.
[0058] The present invention also provides a computer-readable storage medium, on which computer program instructions capable of being run by a processor are stored. When the processor runs the computer program instructions, the method steps described in any one of the above can be implemented.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] 1. The present invention proposes a tower crane bolt detection model based on the optimized and improved YOLOv8s. By introducing the MDEA attention mechanism, the anchor box loss function (AAPL) in the model training process is optimized at the same time. The ability of the model to capture the feature information of the standard section connection bolts of the tower crane in a complex background is enhanced. At the same time, the prediction accuracy and the convergence speed of the training model are improved.
[0061] 2. The present invention introduces the MDEA attention mechanism in the Backbone part of the model to enhance the ability of the model to extract the features of the standard section connection bolts of the tower crane in a complex construction site scenario.
[0062] 3. The present invention optimizes the bounding box regression loss during the training process, adopts the AAPL loss function, reduces the harm of low-quality samples to the model, accelerates the convergence speed and regression accuracy of the model, and improves the generalization degree and detection accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a flowchart of the method of the present invention;
[0064] Figure 2 is a schematic structural diagram of the MDEA attention mechanism module;
[0065] Figure 3 is a schematic diagram of the features extracted from the images taken by the drone. DETAILED DESCRIPTION OF THE INVENTION
[0066] The technical solutions of the present invention will be specifically described below in conjunction with the drawings.
[0067] The present invention provides a method for detecting the state of tower crane bolts by drones based on a multi-scale dynamic edge enhancement attention mechanism. Image samples of the standard section connection bolts of the tower crane are obtained by taking pictures with drones, and a sample library under multiple scenarios is constructed; on the basis of the YOLOv8 model, a multi-scale dynamic edge enhancement attention mechanism MDEA is proposed, and an adaptive anchor box accuracy loss function AAPL is proposed as the loss function during the model training process; the trained model is used to judge the state and locate the state image of the standard section connection bolts of the tower crane to be detected.
[0068] The following is the specific implementation process of the present invention.
[0069] The present invention provides a method for detecting bolt status of a UAV tower crane based on a multi-scale dynamic edge enhancement attention mechanism, comprising the following steps:
[0070] 1) Use drones to shoot the status of the standard section connection bolts of tower cranes on construction sites in different construction scenarios, and obtain images of tower crane bolts from different perspectives: horizontal, upward, downward, long distance, and close distance.
[0071] 2) Use the labelimg tool to manually annotate the acquired tower crane bolt image data, annotate the normal state as normal, the loose state as loose, and the missing state as lack, store the label name and the annotation position to construct a data set, and then randomly divide the data set into a training set and a test set according to the ratio of 8:2.
[0072] 3) Using the basic YOLOv8s network as the initial model, the Multi-Scale Dynamic Edge-Enhanced Attention (MDEA) mechanism proposed in this patent is added to the backbone part of the YOLOv8 network. MDEA significantly improves the model's ability to capture small target features through multi-scale spatial attention, edge enhancement module and dynamic weight adjustment.
[0073] The above-mentioned MDEA attention mechanism is an attention mechanism that enhances the performance of convolutional neural networks, aiming to overcome the limitations of traditional convolutional neural networks in processing information of different scales, shapes and directions. MDEA significantly improves the model's ability to capture small target features through three mechanisms: multi-scale spatial attention, edge enhancement module, and dynamic weight adjustment, especially in the recognition of the state of the standard section connection bolts of tower cranes under complex backgrounds.
[0074] Multi-scale spatial attention uses multi-scale pooling operations, and adopts 3×3, 5×5, 7×7 maximum pooling and average pooling to capture spatial features of different scales, enhancing the model's ability to extract detail information of small objects. Specifically, for the input feature map X, a multi-scale pooling operation is first performed as shown in formulas (1) and (2):
[0075] M ki =MaxPool ki (X) (38)
[0076] A ki =AvgPool ki (X) (39)
[0077] where ki∈{3,5,7} is the size of the pooling kernel, M kiSpatial attention global maximum pooling feature corresponding to the pooling kernel size, A ki Spatial attention average pooling feature corresponding to the pooling kernel size, AvgPool(X) is average pooling, MaxPool(X) is global maximum pooling, and multi-scale spatial attention is obtained by concatenating the features after multi-scale pooling as shown in Equation (3):
[0078] F scale = Concat(M k1 , A k1 ,..., M kn , A kn ) (40)
[0079] Subsequently, an edge enhancement module is introduced. The Sobel operator is used to extract edge features, and they are concatenated with the multi-scale spatial attention. The edge features are extracted using the Sobel operator as shown in Equation (4):
[0080] Edge = Sobel(X) (41)
[0081] where Edge is the extracted edge feature, and then the edge feature is concatenated with the multi-scale spatial attention as shown in Equation (5):
[0082] F edge = Concat(F scale, Edge) (42)
[0083] Dynamically adjust the spatial attention weight according to the saliency of local features, enabling the model to adaptively focus on the key regions of small targets. Calculate the normalized F edge through a 1×1 convolution and the softmax function to obtain the dynamic weight of the image features as shown in Equation (6):
[0084] ω s = Softmax(Conv 1×1 (Norm(F edge ))) (43)
[0085] where Norm is the maximum-minimum normalization calculation.
[0086] The spatial attention map is generated through a 1×1 convolution and the Sigmoid function, and combined with the dynamic weight to further enhance the model's detection ability for small targets; the expression of the spatial attention map is shown in Equation (7):
[0087] A s = σ(Conv 1×1 (F edge )) ⊙ ω s (44)
[0088] where σ is the Sigmoid function, and ⊙ is the Hadamard product, i.e., element-wise multiplication.
[0089] The channel attention map is generated by global max pooling and global average pooling, and then the channel attention weights are generated through a fully connected layer, which are used to enhance the feature representation of important channels and suppress the feature of unimportant channels. The global max pooling and global average pooling are shown in Equations (8) and (9) as follows:
[0090] Y avg = AvgPool(X) (45)
[0091] Y max = MaxPool(X) (46)
[0092] where Y avg is the average pooling feature of the channel attention, and Y max is the max pooling feature of the channel attention.
[0093] Then, the channel attention weights are generated through a fully connected layer as shown in Equation (10):
[0094] A c = σ(FC2(ReLU(FC1(Y avg + Y max )))) (47)
[0095] where FC1 is the first fully connected layer, which maps the input feature vector to a lower-dimensional space. FC2 is the second fully connected layer used to map the output of FC1 back to the original channel dimension. ReLU (Rectified Linear Unit) is an activation function, and its calculation method is shown in Equation (11):
[0096] ReLU = max(0, x) (29)
[0097] Finally, the channel attention map and the spatial attention map are combined to weight the input feature map, and the final output feature map is obtained as shown in Equation (48):
[0098] F out = A c ⊙(A s ⊙X) (49).
[0099] Through these improvements, the MDEA attention mechanism significantly improves the model's ability to capture small target features, especially in the identification of the connection bolt status of tower crane standard sections under complex backgrounds.
[0100] 4) To avoid the influence of overly large differences in the number of samples on the model's tendency and low-quality annotation boxes, the present invention proposes an Adaptive Anchor Precision Loss (AAPL) function. By optimizing the anchor box loss, AAPL improves the model's detection performance for small targets, accelerates the model's training convergence speed, and significantly enhances the positioning accuracy of the anchor boxes.
[0101] The above-mentioned Adaptive Anchor Precision Loss (AAPL) function optimizes the anchor box loss to better focus on defective bolt samples in the case of obvious imbalance of tower crane bolt samples in complex scenarios. The anchor box loss function is implemented based on the center point distance between anchor boxes, and its loss function calculation method is as follows:
[0102] The anchor box loss function is implemented based on the center point distance between anchor boxes, and its loss function calculation method is as follows:
[0103] The accuracy of object detection performance and evaluation is highly correlated with the design of the loss function. The purpose of bounding box regression is to make the bounding box output by the detector approach the true bounding box by comparing it with the true bounding box. IoU (intersection over union) is used for evaluating the loss of prediction boxes in the field of object detection, and its formula is shown in Equation (13):
[0104]
[0105] where B pred is the predicted bounding box region; B gt is the annotated bounding box region;
[0106] The basic IoU assumes that all annotations are good and does not consider the harmful gradients generated during training caused by possible low-quality annotations; instead, it focuses on strengthening the fitting ability of the predicted bounding box region to the annotated bounding box region; at the same time, its ability to distinguish difficult-to-classify samples is weak, and its perception ability in complex environments is insufficient; in addition, the problem of inconsistent numbers of samples for different classes is not considered; therefore, to address these shortcomings, AAPL is designed to reduce the impact of low-quality annotation boxes on model training; the calculation of AAPL is shown in Equation (14):
[0107] L AAPL = ω c ·(1 - L IoU ) γ ·L WIoU (51)
[0108] In the formula, γ is the adjustment hyperparameter for difficult-to-classify samples, ω cis the sample weight for samples of class c, and its expression is shown in Equation (15):
[0109]
[0110] where N is the total number of samples, N c is the number of samples of the class, L IoU is a loss function that measures the overlap degree between the predicted bounding box and the ground truth bounding box, L WIoU is a loss function that introduces additional geometric features based on L IoU The calculation formulas (16) of L IoU and L WIoU are shown in Equation and Equation (17):
[0111] L IoU = 1 - IoU (53)
[0112]
[0113] where (x pred , y pred ) and (x gt , y gt ) represent the center point coordinates of the predicted bounding box and the ground truth bounding box respectively, w c and h c are the width and height of the smallest annotation bounding box. * means separating the smallest bounding box from the gradient calculation to reduce the adverse effects generated by model training; β represents the outlier degree of the predicted bounding box, and the outlier degree β of the predicted bounding box is shown in Equation (18):
[0114]
[0115] The closer the outlier value is to 1, the closer its annotation effect is to the average level. Assigning a smaller gradient gain to samples with a smaller outlier value is beneficial to increasing the generalization ability of the model, while assigning a smaller gradient gain to samples with an overly large outlier value is also beneficial to reducing the large impact of low-quality annotation samples on the model accuracy.
[0116] In this way, AAPL not only considers the overlap degree between the predicted bounding box and the ground truth bounding box, but also dynamically adjusts the loss according to the difficulty and quality of the samples, and at the same time considers the weights of samples of different classes. This loss function is more suitable for dealing with the problems of class imbalance and uneven sample quality, and can significantly improve the performance and robustness of the object detection model. By dynamically adjusting the weights of samples of each class, AAPL helps the model to pay more attention to those difficult-to-detect classes, thereby improving the overall detection accuracy.
[0117] 5) Input the constructed dataset into the proposed MDEA - YOLOv8s model for training to obtain the model weight file of MDEA - YOLOv8s.pt.
[0118] 6) Use the trained MDEA - YOLOv8s to perform status judgment and positioning on the images of the connection bolts of the tower crane standard sections to be detected.
[0119] Example:
[0120] Construct an experimental dataset by collecting on - site with a quad - rotor UAV. At the same time, referring to the common tower crane bolt situations in actual situations, normal, loose, and missing are selected as the main research conditions. Take pictures under different lighting conditions and different tower crane backgrounds. Finally, a dataset of 500 tower crane bolt pictures is built, randomly select 400 as the training set and 100 as the validation set. During the training process, the Epoch is set to 200 rounds, the batch_size is set to 8, and the input image resolution is 640 * 640. Train the models of 4 improvement strategies according to the patent method, conduct ablation experiments on the same dataset to verify the effectiveness of the patent method. To verify the performance of the model, recall (Recall, R), precision (Precision, P), mean average precision (meanAP, mAP), etc. are used as model evaluation indicators. The experiments all use mAP@0.5, that is, a series of IoU thresholds in the range when the IOU threshold is set to 0.5, and the mean is calculated for each class and then mAP is calculated. The effects of the models with different improvement strategies are shown in Table 1.
[0121] Table 1 Ablation Experiments
[0122]
[0123] It can be seen from the experimental results in Table 1 that YOLOv8s represents the original algorithm, and its mAP50 is 90.0%; in Experiment 1, the MDEA attention mechanism is introduced on the basis of YOLOv8. Compared with the original YOLOv8s, the mAP50 is increased by 0.6%; in Experiment 2, the original anchor box loss function is replaced by AAPL, and the mAP is increased by 2.2% without increasing the computational amount; in Experiment 3, both the MDEA attention is introduced and AAPL is replaced, and its mAP is increased by 2.8%. The results of the ablation experiments show that the improved method proposed in the present invention improves the object detection performance of the model to a certain extent, proving the effectiveness of the proposed improved method.
[0124] The present invention also provides an unmanned aerial vehicle tower crane bolt state detection system based on a multi-scale dynamic edge enhancement attention mechanism, which includes a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the method steps described in any of the above can be implemented.
[0125] The present invention also provides a computer-readable storage medium, on which computer program instructions capable of being run by the processor are stored. When the processor runs the computer program instructions, the method steps described in any of the above can be implemented.
[0126] The above are the preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention in terms of the functions and effects produced belong to the protection scope of the present invention.
Claims
1. A method for detecting the state of tower crane bolts of unmanned aerial vehicles based on a multi-scale dynamic edge enhancement attention mechanism, characterized in that, Obtain image samples of the connecting bolts of the tower crane standard section by drone shooting, and construct a sample library under multiple scenarios; based on the YOLOv8 model, propose a multi-scale dynamic edge enhancement attention mechanism MDEA, and propose an adaptive anchor box accuracy loss function AAPL as the loss function during model training; use the trained model to judge the status and locate the status image of the connecting bolts of the tower crane standard section to be detected.
2. The method for detecting the state of the tower crane bolts of the unmanned aerial vehicle based on the multi-scale dynamic edge enhancement attention mechanism according to claim 1, wherein, The method includes the following steps: S1. Use a drone to shoot the status of the connecting bolts of the tower crane standard section at the construction site under different construction scenarios, and obtain tower crane bolt images from different perspectives; S2. Use the labelimg tool to annotate the obtained tower crane bolt image data, label the normal status as normal, the loose status as loose, and the missing status as lack, store the label name and the annotation position to construct a data set and divide it into a training set and a test set; S3. Use the YOLOv8s network as the initial model, and add a multi-scale dynamic edge enhancement attention mechanism MDEA to the backbone part of the YOLOv8 network; S4. Introduce an adaptive anchor box accuracy loss function AAPL as the loss function during model training; S5. Input the constructed data set into the model proposed in steps S3 and S4 for training to obtain a model weight file; S6. Use the trained model to judge the status and locate the status image of the connecting bolts of the tower crane standard section to be detected.
3. The method for detecting the state of the tower crane bolts of the unmanned aerial vehicle based on the multi-scale dynamic edge enhancement attention mechanism according to claim 2, wherein, Different perspectives include front view, upward view, downward view, far distance, and near distance perspectives.
4. The method for detecting the state of tower crane bolts of an unmanned aerial vehicle based on a multi-scale dynamic edge enhancement attention mechanism according to claim 2, wherein, The data set is randomly divided into a training set and a test set according to a ratio of 8:
2.
5. A method for detecting the state of tower crane bolts of an unmanned aerial vehicle based on a multi-scale dynamic edge enhancement attention mechanism according to claim 2, characterized in that, MDEA includes three mechanisms: multi-scale spatial attention, edge enhancement module, and dynamic weight adjustment.
6. A method for detecting the state of bolts of a drone tower crane based on a multi-scale dynamic edge enhancement attention mechanism according to claim 1, 2 or 5, characterized in that, The specific implementation of MDEA is as follows: The multi-scale spatial attention captures spatial features of different scales through multi-scale pooling operations, using max pooling and average pooling of 3×3, 5×5, and 7×7 to enhance the model's ability to extract detailed information of small targets; specifically, for the input feature map X, first perform multi-scale pooling operations as shown in formulas (1) and (2): M ki = MaxPool ki (X) (1) A ki = AvgPool ki (X) (2) where \(k_i\in\{3,5,7\}\) is the size of the pooling kernel, AvgPool(X) is average pooling, MaxPool(X) is global maximum pooling, M ki is the global maximum pooling feature of spatial attention corresponding to the size of the pooling kernel, A ki is the average pooling feature of spatial attention corresponding to the size of the pooling kernel. The multi-scale spatial attention is obtained by concatenating the features after multi-scale pooling along the channels as shown in Equation (3): F scale = Concat(M k1 , A k1 ,..., M kn , A kn ) (3) Subsequently, introduce an edge enhancement module, use the Sobel operator to extract edge features, and splice them with the multi-scale spatial attention. The edge features are extracted using the Sobel operator as shown in formula (4): Edge = Sobel(X) (4) where Edge is the extracted edge feature, and then the edge feature is spliced with the multi-scale spatial attention as shown in formula (5): F edge = Concat(F scale, Edge) (5) Calculate the normalized F through 1×1 convolution and softmax function edge to obtain the dynamic weights of image features as shown in Equation (6): ω s = Softmax(Conv 1×1 (Norm(F edge ))) (6) where Norm is the maximum-minimum normalization calculation; The spatial attention map is generated through a 1×1 convolution and the Sigmoid function, and combined with dynamic weights to further enhance the model's detection ability for small targets; the expression of the spatial attention map is shown in formula (7): A s = σ(Conv 1×1 (F edge )) ⊙ ω s (7) where σ is the Sigmoid function, and ⊙ is the Hadamard product, that is, element-wise multiplication; The channel attention map is obtained through global max pooling and global average pooling, and then the channel attention weights are generated through a fully connected layer to enhance the feature representation of important channels and suppress the feature of unimportant channels. The global max pooling and global average pooling are shown in Equations (8) and (9): Y avg = AvgPool(X) (8) Y max = MaxPool(X) (9) Among them, Y avg is the average pooling feature of channel attention, and Y max is the maximum pooling feature of channel attention; Then the channel attention weights are generated through a fully connected layer as shown in Equation (10): A c = σ(FC2(ReLU(FC1(Y avg +Y max )))) (10) Where, FC1 is the first fully connected layer, which maps the input feature vector to a lower-dimensional space; FC2 is the second fully connected layer used to map the output of FC1 back to the original channel dimension; ReLU is an activation function, and its calculation method is shown in Equation (11): ReLU = max(0, x) (11) Finally, the channel attention map and the spatial attention map are combined to weight the input feature map, and the final output feature map is obtained as shown in Equation (12): F out = A c ⊙ (A s ⊙ X) (12).
7. The method for detecting the state of the tower crane bolts of the unmanned aerial vehicle based on the multi-scale dynamic edge enhancement attention mechanism according to claim 2, wherein, AAPL optimizes the anchor box loss to better focus on defective bolt samples in the case of obvious imbalance of tower crane bolt samples in complex scenarios.
8. A method for detecting the state of tower crane bolts of an unmanned aerial vehicle based on a multi-scale dynamic edge enhancement attention mechanism according to claim 7, characterized in that, The anchor box loss function is implemented based on the center point distance between anchor boxes. The calculation method of the anchor box loss function is as follows: The performance of object detection and the accuracy of evaluation are highly related to the design of the loss function. The purpose of bounding box regression is to make the bounding box output by the detector approach the true bounding box by comparing it with the true bounding box; IoU is used to evaluate the prediction box loss in the field of object detection, and its formula is shown in Equation (13): Among which B pred is the predicted bounding box region; B gt is the annotated bounding box region; The basic IoU assumes that all annotations are good by default, and does not consider the harmful gradients generated during the training process caused by possible low-quality annotations; instead, it focuses on strengthening the fitting ability of the predicted bounding box area to the annotated bounding box area; at the same time, its ability to distinguish difficult-to-classify samples is weak, and its perception ability in complex environments is insufficient; in addition, the problem of inconsistent sample numbers for different classes is not considered; therefore, to solve these shortcomings, AAPL is designed to reduce the impact of low-quality annotation boxes on model training; the calculation of AAPL is shown in Equation (14): L AAPL = ω c ·(1 - L IoU ) γ ·L WIoU (14) where γ is the adjustment hyperparameter for difficult-to-classify samples, and ω c is the sample weight of samples in class c, and its expression is shown in Equation (15): Where N is the total number of samples, N c is the number of class samples, L IoU is the loss function for measuring the overlap between the predicted bounding box and the ground truth bounding box, L WIoU is the loss function that introduces additional geometric features based on L IoU The calculation formulas (16) of L IoU and L WIoU are shown in Equation and Equation (17): L IoU = 1 - IoU(16) where (x pred , y pred ), (x gt , y gt ) represent the center point coordinates of the predicted bounding box and the ground truth bounding box respectively, w c and h c are the width and height of the minimum annotation box. * means separating the minimum bounding box from the gradient calculation to reduce the adverse effects generated by model training; β represents the outlier degree of the predicted box. The outlier degree β of the predicted bounding box is shown in Equation (18): The closer the outlier is to 1, the closer its annotation effect is to the average level. Assigning a smaller gradient gain to samples with smaller outliers is beneficial to increasing the generalization ability of the model, and assigning a smaller gradient gain to samples with too large outliers is also beneficial to reducing the large impact of low-quality annotation samples on the accuracy of the model.
9. An unmanned aerial vehicle tower crane bolt status detection system based on a multi-scale dynamic edge enhancement attention mechanism, characterized in that, Including a memory, a processor, and computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, the method steps described in any one of Claims 1-8 can be implemented.
10. A computer-readable storage medium, on which computer program instructions executable by the processor are stored. When the processor runs the computer program instructions, the method steps described in any one of Claims 1-8 can be implemented.
Citation Information
Cited By
Unmanned aerial vehicle bolt detection model construction method based on image feature guide parameters
CN121884199A