Unmanned aerial vehicle visual angle bolt detection system based on image feature guidance
By introducing the coordinate attention mechanism and weighted bounding box regression loss function in the YOLOv8s model, the problem of low efficiency and accuracy of bolt state detection in the prior art is solved, and efficient and accurate bolt state detection in complex scenarios is achieved.
Patent Information
- Application Number
- CN202510428171.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the efficiency and accuracy of bolt status detection are low, especially in complex scenarios, the detection effect of bolt status of large equipment is poor.
Using a drone viewing bolt detection system based on image feature guidance, the model's ability to capture bolt feature information and the accuracy of prediction boxes are enhanced by introducing coordinate attention mechanism and weighted bounding box regression loss function in the YOLOv8s model.
It significantly improves the efficiency and accuracy of bolt state detection in complex scenarios, optimizes the convergence speed of the model, reduces prediction errors, and improves the overall detection performance.
Smart Images

Figure CN119942387A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bolt detection, and in particular relates to an unmanned aerial vehicle (UAV) perspective bolt detection system based on image feature guidance. Background Art
[0002] With the advancement of technology, machine inspection technology has gradually become the focus of research. It significantly improves the efficiency and accuracy of detection through automation. With the popularization of drone technology in the field of inspection, bolt detection has a new solution. The portability, maneuverability and flight ability of drones in complex environments make them particularly critical in inspection work under extreme climatic conditions. Compared with traditional manual inspections and machine inspections, drone inspections show obvious advantages in personnel safety, maintenance costs and detection efficiency. Therefore, research on drone target detection algorithms has very important practical application value.
[0003] The two major categories of traditional target detection algorithms are two-stage algorithms and single-stage algorithms. In order to improve accuracy, the two-stage detection algorithm performs classification and position regression after generating candidate regions. However, this processing method is relatively slow and not suitable for real-time applications. The single-stage detection algorithm integrates the target detection task into a single neural network model, without the need for a region proposal generation stage, and directly outputs the target category and location information, so it has faster reasoning speed and real-time performance. Among them, YOLOv8 has become the focus of current research with its high accuracy, fast reasoning speed and small number of parameters, and also provides a new solution for drone target detection.
[0004] By adding the coordinate attention mechanism, the model's ability to extract feature information can be enhanced, making better use of feature information, reducing the loss of feature information during the sampling process, and effectively improving the model's detection capability while maintaining its lightweight. In addition, the weighted bounding box regression loss function is optimized to guide the model to pay more attention to the key feature information of the bolts, which has improved the prediction box accuracy and model convergence speed. Summary of the invention
[0005] The purpose of the present invention is to overcome the problems of low efficiency and accuracy in bolt status detection in the prior art. Therefore, a bolt detection system based on image feature guidance from a drone perspective is provided. The system can enhance the model's ability to capture bolt features and improve the detection efficiency and accuracy of the bolt status of large equipment such as lifting machinery in complex scenarios.
[0006] To achieve the above purpose, the technical solution of the present invention is: a bolt detection system based on image feature guidance from the perspective of an unmanned aerial vehicle, comprising: The dataset construction module obtains bolt image samples through drone photography and constructs a sample dataset; The model building module is based on the YOLOv8s model, introduces the coordinate attention mechanism and weighted bounding box regression loss function, and builds an improved YOLOv8s model; Model training module, which trains the improved YOLOv8s model based on the sample data set to obtain the trained bolt status detection model; The bolt state detection module, based on the bolt state detection model, performs identification and detection on the bolt image samples to be identified and outputs the bolt state detection results.
[0007] Furthermore, the data set construction module includes an image acquisition module for acquiring bolt image samples, an image preprocessing module for preprocessing the acquired bolt image samples, and an image annotation module for annotating the preprocessed images.
[0008] Furthermore, the image acquisition module utilizes a high-altitude UAV to photograph the standard sections of the large equipment, and uses the minimum safe distance supported by the optical obstacle avoidance system of the UAV as the preferred shooting distance. By controlling the UAV to a collection point, the camera on the UAV collects images facing the standard sections of the large equipment, and collects images in each direction of the connection between the two standard sections. In each direction, the standard sections are photographed from three angles: looking down, looking straight, and looking up, to obtain bolt image samples.
[0009] Furthermore, the image preprocessing module performs image processing on bolt image samples by an image generation method including random rotation, background replacement, and exposure change, and performs cross-expansion by a sample expansion method to construct a bolt image sample library.
[0010] Furthermore, the image annotation module uses Labelimg data annotation software to select the bolt image range of the bolt image samples in the bolt image sample library one by one and assign corresponding working condition labels. After the annotation is completed, the label files corresponding to the bolt image samples are saved. The label files include the bolt working condition category number, the normalized marking center coordinates and the width and height information of the marking area. Then, based on the bolt image samples in the bolt image sample library and the corresponding label files, a sample data set is constructed and randomly divided into a training set and a test set in a ratio of 4:1.
[0011] Furthermore, the model building module combines the coordinate attention mechanism at the 1st, 3rd, and 4th c2f modules of the neck of the YOLOv8s model, and introduces a weighted bounding box regression loss function.
[0012] Furthermore, the first half of the coordinate attention mechanism targets the annotated bolt image samples. First, it performs one-dimensional average pooling operations on the bolt image samples in the horizontal and vertical directions respectively, extracts the global information of the bolt image samples, and generates two one-dimensional direction-aware feature maps. The specific formula is as follows: In the formula, Represents the feature vector obtained after the nth channel undergoes a one-dimensional average pooling operation in the horizontal direction; Represents the feature vector obtained after the nth channel undergoes a one-dimensional average pooling operation in the vertical direction; W Indicates the width of the bolt image; H Indicates the height of the bolt image; i Represents the pixel position index in the horizontal direction; j Represents the pixel position index in the vertical direction; represents the pixel value of the pixel at position (i, h) of the nth channel; represents the pixel value of the pixel at position (w, j) of the nth channel; The second half of the coordinate attention mechanism focuses on the generation of spatial attention. The second half of the coordinate attention mechanism connects the two perceptual feature maps with global information output by the first half, and compresses the channel dimension to the original C / r through convolution operation, where C is the number of channels of the input feature map and r is a compression ratio; then, batch normalization is used to normalize the feature map after the convolution operation, as follows: In the formula, [,] represents the splicing operation along the spatial dimension; represents the initial perceptual feature map; Represents the convolution operation; Represents the feature map after the convolution operation; Represents the feature map after batch normalization operation; μ and σ 2 Respectively represent the mean and variance of the feature map after the convolution operation; is a constant used for numerical stability to prevent the denominator from being zero; γ and β are learning parameters, which are used to scale and translate the normalized feature maps respectively; The feature map after batch normalization is processed through a nonlinear activation function to generate a feature map of the shape (C / r)×1×(W+H), as follows: In the formula, represents a nonlinear activation function; Will Split into two tensors and , then and Through convolution operation and Sigmoid function respectively, the processing results are as follows: in, and Represents the convolution operation on two tensors. Represents the Sigmoid activation function; i Represents the pixel position index in the horizontal direction, j Represents the pixel position index in the vertical direction; , Represent the attention weights on height and width respectively; The weighted fusion of the spatial information containing bolts is output as follows: in, and Represent the attention weights on the height and width of the nth channel, respectively. represents the pixel value of the pixel at position (i, j) of the nth channel, that is, the input pixel value of the coordinate attention mechanism of the nth channel, is the output result of the nth channel, Weights generated for the coordinate attention mechanism.
[0013] Furthermore, the weighted bounding box regression loss function The loss function is divided into IoU loss, distance loss and edge length loss, weighted bounding box regression loss function The calculation formula is as follows: is the IoU loss, which is used to evaluate the degree of spatial overlap between the real bolt position bounding box and the predicted bolt position bounding box; is the attention-weighted distance loss, which is used to evaluate the Euclidean distance between the center points of the true bolt position bounding box and the predicted bolt position bounding box considering the coordinate attention weight; and is the height and width loss of the attention-weighted prediction box, which is used to evaluate the side lengths of the true bolt position box and the predicted bolt position box considering the coordinate attention weight; , , and The calculation formula is as follows: Among them, A is the border area of the actual bolt position, and B is the border area of the predicted bolt position. In is the weight generated by the coordinate attention mechanism, indicating the degree of attention paid to the center point at the position of the i-th row and the j-th column; and are the x and y coordinates of the true bolt center point; and are the x and y coordinates of the predicted bolt center point; Represents the square of the Euclidean distance between the predicted center point and the true center point; In represents the attention weight of the predicted box height position, is the predicted bolt width; is the width of the real bolt; Represents the absolute difference between the predicted width and the true width; In Represents the attention weight of the predicted box width position; is the predicted bolt height; is the width of the real bolt; Represents the absolute difference between the predicted width and the true width.
[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention proposes a bolt detection system based on image feature guidance from the perspective of a drone. By incorporating coordinate attention weight guidance into the loss function, the model's ability to capture bolt feature information is enhanced, the accuracy of the prediction frame is optimized, and the convergence speed of the model is accelerated. This innovation significantly improves the efficiency and accuracy of the model's detection of bolt status in complex scenarios.
[0015] 2. The present invention adds a coordinate attention mechanism to the neck of the YOLOv8 model, enhances the model's ability to extract bolt feature information based on the input feature image, makes better use of the feature information, reduces the loss of feature information during the sampling process, and effectively improves the model's detection capability while maintaining its lightweight.
[0016] 3. The present invention optimizes the bounding box regression loss function during training and incorporates coordinate attention weight guidance. This improvement provides a more sophisticated optimization mechanism, allowing the model to optimize the width and height of the prediction box more independently, significantly reducing the prediction error and further improving the detection accuracy. Through attention weight guidance, the model can pay more attention to the errors in key positions, thereby more accurately locating and identifying the bolt status in complex scenarios, improving the overall detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a system block diagram of the present invention; Figure 2 This is a schematic diagram of drone flight sampling; Figure 3 is a schematic diagram of image sample collection; Figure 4 Sample expansion diagram; Figure 5 Bolt positioning example diagram based on LabelImg; Figure 6 Bolt detection system UI interface; Figure 7 It is a detection diagram of the system of the present invention; Figure 8 2 is a diagram of the detection results of the system of the present invention. DETAILED DESCRIPTION
[0018] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.
[0019] like Figure 1 As shown, the present invention provides a bolt detection system based on image feature guidance from a drone perspective, comprising: The dataset construction module obtains bolt image samples through drone photography and constructs a sample dataset; The model building module is based on the YOLOv8s model, introduces the coordinate attention mechanism and weighted bounding box regression loss function, and builds an improved YOLOv8s model; Model training module, which trains the improved YOLOv8s model based on the sample data set to obtain the trained bolt status detection model; The bolt state detection module, based on the bolt state detection model, performs identification and detection on the bolt image samples to be identified and outputs the bolt state detection results.
[0020] Each module is implemented as follows: 1. A data set construction module, including an image acquisition module for acquiring bolt image samples, an image preprocessing module for preprocessing the acquired bolt image samples, and an image annotation module for annotating the preprocessed images; wherein, The image acquisition module uses a drone to fly at high altitude to shoot the standard sections of large equipment. The minimum safe distance supported by the drone's optical obstacle avoidance system is used as the preferred shooting distance. The drone is controlled to the collection point so that the camera on the drone can collect images facing the standard sections of the large equipment. Images are collected in various directions at the connection between the two standard sections. In each direction, the standard sections are photographed from three angles: looking down, looking straight, and looking up, to obtain bolt image samples.
[0021] In this example, the image acquisition module uses the drone carrier to Figure 2 , Figure 3 The flight trajectory is used to shoot the physical model of the standard section of the tower crane at a safe distance, and image samples of high-strength bolts in normal standard sections at different angles are collected to construct an experimental data set. At the same time, the common tower crane bolt conditions in actual situations are selected, with normal, loose and missing as the main working conditions. At the same time, the bolts under different lighting conditions and different background areas are photographed.
[0022] The image preprocessing module processes the bolt image samples through image generation methods including random rotation, background replacement, and exposure change, and cross-expansions them through sample expansion methods to build a bolt image sample library.
[0023] In this example, the image preprocessing module processes the image taken by the drone through image generation methods, such as random rotation, background replacement, exposure change, etc. Figure 4 , further expand the image samples, cross-expand the collected original image samples according to the above sample expansion method, and construct a sample library based on original images and expanded images.
[0024] The image annotation module uses Labelimg data annotation software to select the bolt image range of the bolt image samples in the bolt image sample library one by one and assign corresponding working condition labels. After the annotation is completed, the label files corresponding to the bolt image samples are saved. The label files include the bolt working condition category number, the normalized marking center coordinates, and the width and height information of the marking area. Then, based on the bolt image samples in the bolt image sample library and the corresponding label files, a sample data set is constructed and randomly divided into training set and test set in a ratio of 4:1.
[0025] In this example, the image annotation module uses Labelimg data annotation software to annotate the image samples taken by the drone. The image range of the bolts with different working conditions in the standard section image of the sample library is manually selected one by one and the corresponding working condition labels are assigned. The annotation process is as follows: Figure 5 After the labeling is completed, the program will automatically save the label file corresponding to each image. The file contains the bolt condition category number, the normalized mark center coordinates, and the width and height information of the mark area. Then, a sample data set is constructed and randomly divided into a training set and a test set in a ratio of 4:1.
[0026] 2. Model construction module, combining the coordinate attention mechanism at the 1st, 3rd, and 4th c2f modules of the neck of the YOLOv8s model, and introducing the weighted bounding box regression loss function; among them, The first half of the coordinate attention mechanism takes the input labeled bolt image samples and first performs one-dimensional average pooling operations on the bolt image samples in the horizontal and vertical directions respectively, extracts the global information of the bolt image samples, and generates two one-dimensional direction-aware feature maps; specifically, for the nth channel with a width of w and the nth channel with a height of h, the calculation formula is as follows: In the formula, Represents the feature vector obtained after the nth channel undergoes a one-dimensional average pooling operation in the horizontal direction; Represents the feature vector obtained after the nth channel undergoes a one-dimensional average pooling operation in the vertical direction; W Indicates the width of the bolt image; H Indicates the height of the bolt image; i Represents the pixel position index in the horizontal direction; j Represents the pixel position index in the vertical direction; represents the pixel value of the pixel at position (i, h) of the nth channel; represents the pixel value of the pixel at position (w, j) of the nth channel; The second half of the coordinate attention mechanism focuses on the generation of spatial attention. The second half of the coordinate attention mechanism connects the two perceptual feature maps with global information output by the first half, and compresses the channel dimension to the original C / r through convolution operation, where C is the number of channels of the input feature map and r is a compression ratio; then, batch normalization is used to normalize the feature map after the convolution operation, as follows: In the formula, [,] represents the splicing operation along the spatial dimension; represents the initial perceptual feature map; Represents the convolution operation; Represents the feature map after the convolution operation; Represents the feature map after batch normalization operation; μ and σ 2 Respectively represent the mean and variance of the feature map after the convolution operation; is a constant used for numerical stability to prevent the denominator from being zero; γ and β are learning parameters, which are used to scale and translate the normalized feature maps respectively; The feature map after batch normalization is processed through a nonlinear activation function to generate a feature map of the shape (C / r)×1×(W+H), as follows: In the formula, represents a nonlinear activation function; Will Split into two tensors and , then and Through convolution operation and Sigmoid function respectively, the processing results are as follows: in, and Represents the convolution operation on two tensors. Represents the Sigmoid activation function; i Represents the pixel position index in the horizontal direction, j Represents the pixel position index in the vertical direction; , Represent the attention weights on height and width respectively; The weighted fusion of the spatial information containing bolts is output as follows: in, and Represent the attention weights on the height and width of the nth channel, respectively. represents the pixel value of the pixel at position (i, j) of the nth channel, that is, the input pixel value of the coordinate attention mechanism of the nth channel, is the output result of the nth channel, Weights generated for the coordinate attention mechanism.
[0027] Furthermore, the weighted bounding box regression loss function The loss function is divided into IoU loss, distance loss and edge length loss, weighted bounding box regression loss function The calculation formula is as follows: is the IoU loss, which is used to evaluate the degree of spatial overlap between the real bolt position bounding box and the predicted bolt position bounding box; is the attention-weighted distance loss, which is used to evaluate the Euclidean distance between the center points of the true bolt position bounding box and the predicted bolt position bounding box considering the coordinate attention weight; and is the height and width loss of the attention-weighted prediction box, which is used to evaluate the side lengths of the true bolt position box and the predicted bolt position box considering the coordinate attention weight; , , and The calculation formula is as follows: Among them, A is the border area of the actual bolt position, and B is the border area of the predicted bolt position. In is the weight generated by the coordinate attention mechanism, indicating the degree of attention paid to the center point at the position of the i-th row and the j-th column; and are the x and y coordinates of the true bolt center point; and are the x and y coordinates of the predicted bolt center point; Represents the square of the Euclidean distance between the predicted center point and the true center point; In represents the attention weight of the predicted box height position, is the predicted bolt width; is the width of the real bolt; Represents the absolute difference between the predicted width and the true width; In Represents the attention weight of the predicted box width position; is the predicted bolt height; is the width of the real bolt; Represents the absolute difference between the predicted width and the true width.
[0028] 3. Model training module: an improved YOLOv8s model built based on the model building module. Bolt image samples in the training set are used for model training, and bolt image samples in the test set are used for model testing. In training, Epoch is set to 400 rounds, batch_size is set to 64, and the input image resolution is 640*640. After adding the coordinate attention mechanism, the bolt position features in the image are used to enhance the model's recognition of the bolt position features by constructing a loss function. In order to verify the effectiveness of the system of the present invention, the model of 4 improved strategies (v 8-0 -v 8-3 ), and ablation experiments were conducted on the same dataset. Recall (R), precision (P), mean AP (mAP), etc. were used as model evaluation indicators. The experiments used mAP@0.5 and mAP@0.5-0.95, that is, when the IoU threshold was set to 0.5 and a series of IoU thresholds in the range of 0.5 to 0.95, each category was averaged and then mAP was calculated. The effects of the models with different improvement strategies are shown in Table 1.
[0029] As shown in Table 1, experiment v8-0 represents the original model, with mAP0.5 and mAP0.5-0.95 of 92.7% and 64.7% respectively; experiment v8-1 and experiment v8-2 respectively add CA attention mechanism and replace with The comparison experiment of loss function shows that the model performance has been improved to a certain extent from the two indicators of mAP0.5 and mAP0.5-0.95. Each improvement proposed by the system of the present invention has improved the model performance to a certain extent, which can prove the effectiveness and scientificity of the system of the present invention. Experiment v8-3 is to add CA attention mechanism and replace The improved model of the loss function has better effect on model performance than adding a single term.
[0030] 4. Bolt status detection module, based on PyQt5 and other GUI libraries to build the bolt detection system UI interface, improve the operability of the detection system. The system interface content includes: instructions for use, sample set expansion, image detection, video detection and output report, etc. The bolt detection system UI interface is as follows Figure 6 By loading the trained improved YOLOv8 model, the bolt image samples to be identified can be identified and detected, thereby obtaining the detection results and repairing the faulty bolts. Figure 7 , Figure 8 .
[0031] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.
Claims
1. A bolt detection system based on image feature guidance from the perspective of an unmanned aerial vehicle, characterized in that: include: The dataset construction module obtains bolt image samples through drone photography and constructs a sample dataset; The model building module is based on the YOLOv8s model, introduces the coordinate attention mechanism and weighted bounding box regression loss function, and builds an improved YOLOv8s model; Model training module, which trains the improved YOLOv8s model based on the sample data set to obtain the trained bolt status detection model; The bolt state detection module, based on the bolt state detection model, performs identification and detection on the bolt image samples to be identified and outputs the bolt state detection results.
2. The bolt detection system based on image feature guidance from the perspective of an unmanned aerial vehicle according to claim 1 is characterized in that: The data set construction module includes an image acquisition module for acquiring bolt images, an image preprocessing module for preprocessing the acquired bolt images, and an image annotation module for annotating the preprocessed images.
3. The bolt detection system based on image feature guidance from the perspective of an unmanned aerial vehicle according to claim 2 is characterized in that: The image acquisition module uses a drone to fly at high altitude to shoot the standard section of the large equipment, takes the minimum safe distance supported by the drone's optical obstacle avoidance system as the preferred shooting distance, controls the drone to a collection point so that the camera on the drone collects images facing the standard section of the large equipment, and collects images in each direction of the connection between the two standard sections. In each direction, the standard section is photographed at three angles: looking down, looking straight, and looking up, to obtain a bolt image.
4. The bolt detection system based on image feature guidance from the perspective of an unmanned aerial vehicle according to claim 3 is characterized in that: The image preprocessing module processes the bolt image by an image generation method including random rotation, background replacement, and exposure change, and cross-expansions the bolt image by a sample expansion method to construct a bolt image sample library.
5. The bolt detection system based on image feature guidance from the perspective of an unmanned aerial vehicle according to claim 4 is characterized in that: The image annotation module uses Labelimg data annotation software to select the bolt image range of the bolt image samples in the bolt image sample library one by one and assign corresponding working condition labels. After the annotation is completed, the label file corresponding to the bolt image samples is saved. The label file includes the bolt working condition category number, the normalized marking center coordinates and the width and height information of the marking area. Then, based on the bolt image samples in the bolt image sample library and the corresponding label files, a sample data set is constructed and randomly divided into a training set and a test set in a ratio of 4:
1.
6. The bolt detection system based on image feature guidance from the perspective of an unmanned aerial vehicle according to claim 1 is characterized in that: The model building module combines the coordinate attention mechanism at the 1st, 3rd, and 4th c2f modules of the neck of the YOLOv8s model, and introduces a weighted bounding box regression loss function.
7. The bolt detection system based on image feature guidance from the perspective of an unmanned aerial vehicle according to claim 1 or 6, characterized in that: The first half of the coordinate attention mechanism targets the annotated bolt image samples. First, it performs one-dimensional average pooling operations on the bolt image samples in the horizontal and vertical directions respectively, extracts the global information of the bolt image samples, and generates two one-dimensional direction-aware feature maps. The specific formula is as follows: In the formula, Represents the feature vector obtained after the nth channel undergoes a one-dimensional average pooling operation in the horizontal direction; Represents the feature vector obtained after the nth channel undergoes a one-dimensional average pooling operation in the vertical direction; W Indicates the width of the bolt image; H Indicates the height of the bolt image; i Represents the pixel position index in the horizontal direction; j Represents the pixel position index in the vertical direction; represents the pixel value of the pixel at position (i, h) of the nth channel; represents the pixel value of the pixel at position (w, j) of the nth channel; The second half of the coordinate attention mechanism focuses on the generation of spatial attention. The second half of the coordinate attention mechanism connects the two perceptual feature maps with global information output by the first half, and compresses the channel dimension to the original C / r through convolution operation, where C is the number of channels of the input feature map and r is a compression ratio; then, batch normalization is used to normalize the feature map after the convolution operation, as follows: In the formula, [,] represents the splicing operation along the spatial dimension; represents the initial perceptual feature map; Represents the convolution operation; Represents the feature map after the convolution operation; Represents the feature map after batch normalization operation; μ and σ 2 Respectively represent the mean and variance of the feature map after the convolution operation; is a constant used for numerical stability to prevent the denominator from being zero; γ and β are learning parameters, which are used to scale and translate the normalized feature maps respectively; The feature map after batch normalization is processed through a nonlinear activation function to generate a feature map of the shape (C / r)×1×(W+H), as follows: In the formula, represents a nonlinear activation function; Will Split into two tensors and , then and Through convolution operation and Sigmoid function respectively, the processing results are as follows: in, and Represents the convolution operation on two tensors. Represents the Sigmoid activation function; i Represents the pixel position index in the horizontal direction, j Represents the pixel position index in the vertical direction; , Represent the attention weights on height and width respectively; The weighted fusion of the spatial information containing bolts is output as follows: in, and Represent the attention weights on the height and width of the nth channel, respectively. represents the pixel value of the pixel at position (i, j) of the nth channel, that is, the input pixel value of the coordinate attention mechanism of the nth channel, is the output result of the nth channel, Weights generated for the coordinate attention mechanism.
8. The bolt detection system based on image feature guidance from the perspective of an unmanned aerial vehicle according to claim 7 is characterized in that: Weighted Bounding Box Regression Loss Function The loss function is divided into IoU loss, distance loss and edge length loss, weighted bounding box regression loss function The calculation formula is as follows: is the IoU loss, which is used to evaluate the degree of spatial overlap between the real bolt position bounding box and the predicted bolt position bounding box; is the attention-weighted distance loss, which is used to evaluate the Euclidean distance between the center points of the true bolt position bounding box and the predicted bolt position bounding box considering the coordinate attention weight; and is the height and width loss of the attention-weighted prediction box, which is used to evaluate the side lengths of the true bolt position box and the predicted bolt position box considering the coordinate attention weight; , , and The calculation formula is as follows: Among them, A is the border area of the actual bolt position, and B is the border area of the predicted bolt position. In is the weight generated by the coordinate attention mechanism, indicating the degree of attention paid to the center point at the position of the i-th row and the j-th column; and are the x and y coordinates of the true bolt center point; and are the x and y coordinates of the predicted bolt center point; Represents the square of the Euclidean distance between the predicted center point and the true center point; In represents the attention weight of the predicted box height position, is the predicted bolt width; is the width of the real bolt; Represents the absolute difference between the predicted width and the true width; In Represents the attention weight of the predicted box width position; is the predicted bolt height; is the width of the real bolt; Represents the absolute difference between the predicted width and the true width.
Citation Information
Patent Citations
Bolt detection method and system based on improved YOLOv5 model
CN117132566A
Bolt looseness detection method and device based on synthetic data set and deep learning
CN118711035A
Cited By
Unmanned aerial vehicle bolt detection model construction method based on image feature guide parameters
CN121884199A