An infrared image target detection method for unmanned aerial vehicle inspection
Patent Information
- Application Number
- CN202410307803.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-18
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-03-18
AI Technical Summary
[0004]本发明所要解决的是现有深度学习目标检测算法不适用于无人机巡检任务中红外图像目标检测的问题,提供一种面向无人机巡检的红外图像目标检测方法
[0027]1、在Yolov5基准模型中引入CBAM卷积注意力机制,以增强特征图的通道关联性和感兴趣区域特征,同时抑制非感兴趣的背景噪声特征,从而改善模型卷积层之间的特征图质量,提高模型对目标特征的表达能力;
Smart Images

Figure CN118411630B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and specifically to an infrared image target detection method for unmanned aerial vehicle (UAV) inspection. Background Technology
[0002] With the iterative updates of drone technology and infrared thermal imaging technology, drones have acquired advantages such as small size, low cost, and high mobility. The non-contact thermal radiation distribution detection of infrared thermal imaging technology can make up for the inability to image under no-light conditions, providing important information such as the shape and temperature of the target for detection. This improves the feasibility of applying drones to target detection technology in nighttime security patrol missions and further provides technical support for security fields such as emergency rescue.
[0003] While drones and infrared thermal imaging technology offer significant advantages in practice, infrared images obtained from drone inspections generally suffer from low resolution and high background noise, leading to missed or false detections and reduced detection accuracy. However, nighttime security patrols using drones demand stable detection speed and high positioning accuracy. Existing deep learning-based target detection algorithms, such as two-stage algorithms, offer high accuracy but suffer from high computational cost and slow speed; while single-stage algorithms offer fast speed but lower accuracy. Therefore, existing deep learning model-based target detection methods are not entirely suitable for infrared image target detection in drone inspection tasks. Summary of the Invention
[0004] The present invention addresses the problem that existing deep learning target detection algorithms are not applicable to infrared image target detection in UAV inspection tasks, and provides an infrared image target detection method for UAV inspection.
[0005] To solve the above problems, the present invention is achieved through the following technical solution:
[0006] An infrared image target detection method for UAV inspection includes the following steps:
[0007] Step 1: Construct an improved Yolov5 model. This improved Yolov5 model introduces the C3 module in the backbone network of the Yolov5 baseline model into a convolutional attention mechanism module consisting of a channel attention module and a spatial attention module.
[0008] Step 2: Construct an infrared image dataset for drone inspection;
[0009] Step 3: Use the K-Means clustering algorithm to perform a one-time anchor frame size calculation on the infrared image dataset for UAV inspection constructed in Step 2 to determine the optimal anchor frame size;
[0010] Step 4: Input the infrared image dataset for UAV inspection constructed in Step 2 into the improved Yolov5 model constructed in Step 1 for model training. During the model training process, the target is predicted and a prediction box is generated based on the optimal anchor box size determined in Step 3. Then, the localization loss between the prediction box and the labeled box is calculated using the full intersection-union localization loss function. Through repeated forward propagation and backward propagation operations, the infrared image target detection model for UAV inspection is obtained.
[0011] Step 5: Acquire infrared images for drone inspection using a single drone equipped with a thermal imager, and input the acquired infrared images for drone inspection into the infrared image target detection model obtained in Step 4 for target detection. During the target detection process, the target is predicted and a prediction box is generated based on the optimal anchor frame size determined in Step 3, and the final target detection result is obtained.
[0012] In step 1 above, the improved Yolov5 model consists of an input layer, a backbone network, a neck network, and an output layer. The backbone network includes one focusing module, four convolutional modules, four C3 and convolutional attention mechanism modules, and one spatial pyramid pooling module. The neck network includes four convolutional modules, two upsampling modules, four connection modules, and four C3 modules. The input layer is connected to the input of the focusing module. The output of the focusing module is connected to the input of the first convolutional module, the output of the first convolutional module is connected to the input of the first C3 and convolutional attention mechanism module, and the output of the first C3 and convolutional attention mechanism module... The input of the second convolutional module is connected to the input of the second C3 and convolutional attention mechanism module. The output of the second C3 and convolutional attention mechanism module is connected to the input of the third convolutional module. The output of the third convolutional module is connected to the input of the third C3 and convolutional attention mechanism module. The output of the third C3 and convolutional attention mechanism module is connected to the input of the fourth convolutional module. The output of the fourth convolutional module is connected to the input of the spatial pyramid pooling module. The output of the spatial pyramid pooling module is connected to the input of the fourth C3 and convolutional attention mechanism module. The output of the attention mechanism module is connected to the input of the fifth convolutional module. The output of the fifth convolutional module is connected to the input of the first upsampling module. The output of the first upsampling module and the output of the third C3 module and the convolutional attention mechanism module are simultaneously connected to the input of the first connection module. The output of the first connection module is connected to the input of the first C3 module. The output of the first C3 module is connected to the input of the sixth convolutional module. The output of the sixth convolutional module is connected to the input of the second upsampling module. The output of the second upsampling module and the output of the second C3 module and the convolutional attention mechanism module are simultaneously connected to the input of the second connection module. The output of the second connection module is connected to the input of the second C3 module. The output of the second C3 module is connected to the input of the seventh convolutional module. The output of the seventh convolutional module and the output of the sixth convolutional module are simultaneously connected to the input of the third connection module. The output of the third connection module is connected to the input of the third C3 module. The output of the third C3 module is connected to the input of the eighth convolutional module. The output of the eighth convolutional module and the output of the fifth convolutional module are simultaneously connected to the input of the fourth C3 module. The outputs of the second, third, and fourth C3 modules are connected to the output layer.
[0013] The C3 module and the convolutional attention mechanism module consist of a convolutional attention mechanism module and a C3 module. The convolutional attention mechanism module includes a channel attention module and a spatial attention module. The C3 module includes three lightweight convolutional modules, n cascaded bottleneck modules, and one connection module. Here, n represents the number of bottleneck modules and is a positive integer greater than or equal to 1. The input of the channel attention module forms the input of the C3 module and the convolutional attention mechanism module. The output of the channel attention module is connected to the input of the spatial attention module. The output of the spatial attention module is connected to the input of the first lightweight convolutional module. The output of the first lightweight convolutional module is simultaneously connected to the input of the second lightweight convolutional module and the input of the n cascaded bottleneck modules. The output of the second lightweight convolutional module and the output of the n cascaded bottleneck modules are simultaneously connected to the input of the connection module. The output of the connection module is connected to the input of the third lightweight convolutional module. The output of the third lightweight convolutional module forms the output of the C3 module and the convolutional attention mechanism module.
[0014] The number of bottleneck modules, n, in the first and fourth C3 and convolutional attention mechanism modules is 1, while the number of bottleneck modules, n, in the second and third C3 and convolutional attention mechanism modules is 3.
[0015] The specific process of step 2 above is as follows:
[0016] Step 2.1: Acquire an infrared image set for drone inspection using a single drone equipped with a thermal imager;
[0017] Step 2.2: Manually annotate the infrared image set for UAV inspection obtained in Step 2.1 using professional image annotation tools;
[0018] Step 2.3: Perform Mosaic data augmentation on the infrared image set for UAV inspection that has been annotated in Step 2.2;
[0019] Step 2.4: Divide the infrared image set for UAV inspection obtained by Mosaic data augmentation in Step 2.3 into training set, validation set and test set according to a predetermined ratio.
[0020] The specific process of step 3 above is as follows:
[0021] Step 3.1: Extract the width and height of the bounding boxes from the infrared image dataset for UAV inspection as data samples, and determine the number of clusters to be divided.
[0022] Step 3.2: Randomly select a data sample as the initial cluster center for each cluster;
[0023] Step 3.3: For each data sample, calculate its Euclidean distance to the current cluster center of each cluster, and assign each data sample to the cluster to which the current cluster center with the nearest distance belongs;
[0024] Step 3.4: For each cluster, calculate the mean of its internal data samples as the new cluster center;
[0025] Step 3.5: Repeat steps 3.3 and 3.4 until the preset stopping condition is met, that is, the preset threshold for the change in cluster centers or the preset number of iterations is reached. At this time, the cluster center of each cluster is the calculated optimal anchor frame size.
[0026] Compared with the prior art, the present invention has the following characteristics:
[0027] 1. The CBAM convolutional attention mechanism is introduced into the Yolov5 baseline model to enhance the channel correlation and region of interest features of the feature map, while suppressing background noise features that are not of interest, thereby improving the feature map quality between the model's convolutional layers and enhancing the model's ability to express target features.
[0028] 2. The K-Means clustering algorithm was used to determine the optimal anchor box size for the infrared image dataset for UAV inspection, thereby reducing the false negative rate and false positive rate of the model for infrared targets and significantly improving the accuracy of the model in infrared image target detection.
[0029] 3. During the model training process, the complete intersection-union (CIoU) localization loss function is introduced to measure the similarity between the predicted bounding box and the labeled bounding box, which effectively accelerates the convergence speed of localization information in backpropagation and improves the model's prediction accuracy of localization information.
[0030] 4. The target detection method of the present invention has the characteristics of stable detection speed and high positioning accuracy, and can be applied to infrared image target detection in UAV inspection missions. Attached Figure Description
[0031] Figure 1 This is a flowchart of an infrared image target detection method for UAV inspection.
[0032] Figure 2 Improve the model structure diagram for Yolov5;
[0033] Figure 3 Here is a structural diagram of the C3-CBAM module;
[0034] Figure 4 The graph shows the inference verification results of the Yolov5 baseline model.
[0035] Figure 5The graph shows the inference verification results of the improved Yolov5 model. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific examples and the accompanying drawings.
[0037] An infrared image target detection method for UAV inspection, such as Figure 1 As shown, the steps are as follows:
[0038] S1: Construct an improved Yolov5 model.
[0039] See Figure 2 The improved Yolov5 model consists of an input layer, a backbone network, a neck network, and an output layer connected sequentially. The backbone network includes one focus module, four convolutional modules (Conv), four C3 modules with convolutional attention mechanism (C3-CBAM), and one spatial pyramid pooling module (SPP). The neck network includes four convolutional modules (Conv), two upsampling modules, four connection modules (Concat), and four C3 modules (C3).
[0040] The input layer connects to the input of the focusing module. The output of the focusing module connects to the input of the first convolutional module. The output of the first convolutional module connects to the input of the first C3 and convolutional attention mechanism module. The output of the first C3 and convolutional attention mechanism module connects to the input of the second convolutional module. The output of the second convolutional module connects to the input of the second C3 and convolutional attention mechanism module. The output of the second C3 and convolutional attention mechanism module connects to the input of the third convolutional module. The output of the third convolutional module connects to the input of the third C3 and convolutional attention mechanism module. The output of the third C3 and convolutional attention mechanism module connects to the input of the fourth convolutional module. The output of the fourth convolutional module connects to the input of the spatial pyramid pooling module. The output of the spatial pyramid pooling module connects to the input of the fourth C3 and convolutional attention mechanism module. The output of the fourth C3 convolutional attention mechanism module is connected to the input of the fifth convolutional module. The output of the fifth convolutional module is connected to the input of the first upsampling module. The outputs of the first upsampling module and the third C3 convolutional attention mechanism module are simultaneously connected to the input of the first connection module. The output of the first connection module is connected to the input of the first C3 module. The output of the first C3 module is connected to the input of the sixth convolutional module. The output of the sixth convolutional module is connected to the input of the second upsampling module. The outputs of the second upsampling module and the second C3 convolutional attention mechanism module are simultaneously connected to the input of the second connection module. The output of the second connection module is connected to the input of the second C3 module. The output of the second C3 module is connected to the input of the seventh convolutional module. The outputs of the seventh and sixth convolutional modules are simultaneously connected to the input of the third connection module. The output of the third connection module is connected to the input of the third C3 module. The output of the third C3 module is connected to the input of the eighth convolutional module. The outputs of the eighth and fifth convolutional modules are simultaneously connected to the input of the fourth C3 module. The outputs of the second, third, and fourth C3 modules are connected to the output layer. The output of the second C3 module is connected to the first detector head of the output layer, which is the ninth convolutional module. The output of the third C3 module is connected to the second detector head of the output layer, which is the tenth convolutional module. The output of the fourth C3 module is connected to the third detector head of the output layer, which is the eleventh convolutional module.
[0041] The improved Yolov5 model, based on the Yolov5 baseline model, introduces a CBAM attention mechanism module into the C3 module of the backbone network, forming a C3-CBAM module. See also... Figure 3The C3 and Convolutional Attention Mechanism Module (C3-CBAM) consists of a Convolutional Attention Mechanism Module (CBAM) and a C3 module (C3). The Convolutional Attention Mechanism Module (CBAM) includes a channel attention module (Cam) and a spatial attention module (Sam). The C3 module includes three lightweight convolutional modules (CBL), n bottleneck modules, and one concatenation module (Concat). The n bottleneck modules are cascaded sequentially, with one bottleneck module in the first and fourth C3 and Convolutional Attention Mechanism Modules, and three bottleneck modules in the second and third C3 and Convolutional Attention Mechanism Modules.
[0042] The input of the channel attention module forms the input of C3 and the convolutional attention mechanism module. The output of the channel attention module is connected to the input of the spatial attention module. The output of the spatial attention module is connected to the input of the first lightweight convolutional module. The output of the first lightweight convolutional module is simultaneously connected to the input of the second lightweight convolutional module and the input of n cascaded bottleneck modules. The output of the second lightweight convolutional module and the output of the n cascaded bottleneck modules are simultaneously connected to the input of the connection module. The output of the connection module is connected to the input of the third lightweight convolutional module. The output of the third lightweight convolutional module forms the output of C3 and the convolutional attention mechanism module.
[0043] The convolutional attention mechanism module can adaptively weight feature maps to enhance channel correlation and region of interest features, effectively suppress the influence of noisy features, and improve the model's ability to express target features.
[0044] The channel attention module is used to weight the channel features of feature map F to enhance the channel correlation of the feature map. The specific steps are as follows: First, global average pooling and global max pooling are performed on the input feature map F, outputting two feature maps of size C×1×1 (C is the number of channels). Then, these are processed through two fully connected MLP layers. The first fully connected layer has 1 / r the number of channels in the input feature map, where r is the descent rate, and the activation function is ReLU. The second fully connected layer has the same number of channels as the input feature map, and the activation function is a linear function. Next, the results from the two fully connected layers undergo feature transformation and mapping, and are then fed into a Sigmoid function to obtain the weight coefficients M of the channel attention module. C Finally, the feature map F and the weight coefficients M are compared. C Multiplying them together yields a weighted feature map F', which is the output of the channel attention module and the input of the spatial attention module.
[0045] The weight coefficient M of the channel attention module C The calculation method is as follows:
[0046]
[0047] Where σ represents the Sigmoid function; Represents a global average pooling feature map of size C×1×1; This represents a global max-pooling feature map of size C×1×1.
[0048] The spatial attention module is used to weight the spatial location features of feature map F' to enhance the features of the region of interest. The specific steps are as follows: First, global average pooling and global max pooling are performed on the input feature map F' to obtain two feature maps of size 1×H×W (W and H represent the width and height of the feature map, respectively). Then, these two feature maps are concatenated along the channel dimension. Next, a 7×7 convolution operation is performed, and then the result is fed into a sigmoid activation function to obtain the weight coefficients M of the spatial attention module. S Finally, the input feature map F' and the weight coefficients M are... S Multiplying these together yields the output of the spatial attention module, which is the feature map output by the entire convolutional attention mechanism module.
[0049] The weight coefficient M of the spatial attention module S The calculation method is as follows:
[0050]
[0051] Where σ represents the Sigmoid function; ∫ 7×7 This represents a 7×7 convolution operation; This represents a global average pooling feature map of size 1×H×W; This represents a global max-pooling feature map of size 1×H×W.
[0052] S2: Construct an infrared image dataset for drone inspection.
[0053] S2-1 Data Acquisition: A set of infrared images for drone inspections is acquired by a single drone carrying a thermal imager, totaling 3788 images.
[0054] The infrared image acquisition scheme for UAV inspection is as follows: The acquisition scenarios include seven areas from the UAV inspection perspective: grassland, farmland, park, jungle, mountain, lake, and city; the acquisition angles include overhead, upward, frontal, side, and surround views; the target categories are humans and pigs; the target age ranges include young, middle-aged, and elderly; and the body sizes include small, medium, and large. Target postures are primarily standing, walking, and lying down; the acquisition times are morning, noon, dusk, and nighttime on the same day. This acquisition scheme involves variations in scenes, angles, categories, postures, scales, and lighting conditions to meet the requirements of deep learning models for diverse, high-quality, large-scale, and adaptable datasets.
[0055] S2-2 Data Annotation: The infrared image set acquired in S2-1 for UAV inspection was manually annotated using the professional image annotation tool Labelimg.
[0056] Manual annotation includes location information annotation and category label annotation. Location information annotation uses the XYWH format, where XY represents the center coordinates of the annotation box, and WH represents the width and height of the annotation box; category label annotation must be unique. After annotation, data verification and review are performed to ensure that all images in the entire dataset have been correctly labeled, targets of the same category use the same label, and each labeled file correctly corresponds to the corresponding image.
[0057] S2-3 Data Augmentation: Based on S2-2, perform 10,000 Mosaic data augmentation operations on the dataset with completed data annotation.
[0058] Mosaic data augmentation involves randomly selecting four images from the dataset and performing random flipping, scaling, distribution, and color gamut transformation. These images are then stitched together in a 2×2 matrix to form a new image, and the annotation information of the target categories is mapped accordingly.
[0059] S2-4 Dataset Partitioning: The dataset obtained from S2-3 through Mosaic data augmentation is divided into training set, validation set, and test set in a 6:3:1 ratio.
[0060] S3: Use the K-Means clustering algorithm to perform a one-time Anchor Box size calculation on the infrared image dataset for UAV inspection constructed in S2, and determine the optimal Anchor Box size.
[0061] S3-1 extracts the width and height of the bounding boxes as data samples from the infrared image dataset for UAV inspection and determines that the data samples should be divided into 9 clusters.
[0062] Since this embodiment requires determining 3 sets of Anchor Box sizes, and each set of Anchor Box sizes includes 3 sub-Anchor Box sizes (the number of sub-Anchor Boxes determines the accuracy of the confidence score and the computational cost), the sample data is divided into 9 clusters.
[0063] S3-2 randomly selects 9 data samples as the initial cluster centers.
[0064] S3-3 For each data sample, calculate its Euclidean distance to the 9 cluster centers.
[0065] S3-4 assigns each data sample to the cluster to which the nearest cluster center belongs.
[0066] S3-5 For each cluster, calculate the mean of its internal samples as the new cluster center.
[0067] S3-6 repeatedly executes S3-3 to S3-5 until the stopping condition is met.
[0068] Stop condition 1 is that the change in cluster centers is less than the threshold of 10. -4 Stop condition 2 is when the number of iterations is greater than 2000.
[0069] The final cluster centers in S3-7 are the calculated Anchor Box sizes.
[0070] Since the output layer of the improved YOLOv5 model has three detectors, each detector outputs a feature map at a different scale. Since feature maps at different scales need to correspond to different Anchor Box sizes, the Anchor Box sizes obtained in this invention are three sets: the largest Anchor Box size, the medium Anchor Box size, and the smallest Anchor Box size.
[0071] After applying the K-Means clustering algorithm described above, the three optimal Anchor Box sizes obtained in this embodiment are [(4,5),(4,10),(8,6)], [(8,16),(16,9),(17,22)], and [(30,17),(28,37),(61,44)]. That is:
[0072] The 20×20×256 feature map output by the first detector in the output layer corresponds to the largest Anchor Box size [(30,17),(28,37),(61,44)], where the width and height of one sub-Anchor Box are 30 and 17, the width and height of another sub-Anchor Box are 28 and 37, and the width and height of yet another sub-Anchor Box are 61 and 44.
[0073] The 40×40×256 feature map output by the second detector in the output layer corresponds to a medium-sized Anchor Box [(8,16),(16,9),(17,22)], where the width and height of one sub-Anchor Box are 8 and 16, respectively, the width and height of another sub-Anchor Box are 16 and 9, respectively, and the width and height of yet another sub-Anchor Box are 17 and 22, respectively.
[0074] The 80×80×256 feature map output by the third detection head of the output layer corresponds to the smallest Anchor Box size [(4,5),(4,10),(8,6)], where the width and height of one sub-Anchor Box are 4 and 5 respectively, the width and height of another sub-Anchor Box are 4 and 10 respectively, and the width and height of yet another sub-Anchor Box are 8 and 6 respectively.
[0075] S4: The infrared image dataset for UAV inspection constructed in S2 is fed into the improved Yolov5 model constructed in S1 for model training. During the model training process, the target is predicted and the prediction box is generated based on the optimal Anchor Box size determined in S3. Then, the localization loss between the prediction box and the annotation box is calculated using the complete intersection-union (CIoU) localization loss function. Through repeated forward and backward propagation operations, the infrared image target detection model for UAV inspection is obtained.
[0076] S4-1 inputs the training set of S2-4 into the improved Yolov5 model of S1 for forward propagation, and obtains the model outputs as feature maps of 20×20×256, 40×40×256 and 80×80×256 respectively.
[0077] S4-2 uses the optimal Anchor Box size from S3 to predict the target, generates prediction boxes by combining the parameters of the network model, and filters overlapping prediction boxes by non-maximum suppression.
[0078] S4-3 performs backpropagation, first calculating the difference between the predicted bounding box and the ground truth bounding box (i.e., the labeled bounding box in the dataset). The localization loss is calculated using the CIoU localization loss function. Then, the gradient of each parameter with respect to the loss is calculated using the chain rule, and the model parameters are updated using optimization algorithms such as gradient descent, so that the model gradually converges.
[0079] The CIoU localization loss function is used to measure the similarity between predicted and ground truth bounding boxes during backpropagation. The measurement factors include center point distance, diagonal distance of the minimum closure region, and aspect ratio. The introduction of center point distance and diagonal distance of the minimum closure region allows CIoU to more comprehensively consider the position and shape information of the boxes when calculating the overlap area. Simultaneously, the aspect ratio similarity is used to penalize cases with different aspect ratios, enabling CIoU to more accurately measure the distance between predicted and ground truth bounding boxes with significant aspect ratio differences. This can accelerate the convergence speed of localization information during backpropagation. The final CIoU loss is 1 minus the CIoU value; the closer to 0, the more accurate the box prediction.
[0080] The Cross-Union Ratio (CIoU) is calculated as follows:
[0081]
[0082] Where IoU represents the intersection-union ratio, which measures the degree of overlap between predicted and ground truth boxes by calculating the ratio between the intersecting regions of the predicted and ground truth boxes and their union; ρ 2 (b,b gt ) represents the Euclidean distance between the predicted bounding box and the ground truth bounding box; c represents the diagonal distance of the smallest closure region that simultaneously contains both the predicted bounding box and the ground truth bounding box; α represents the weight parameter; v represents the similarity measured by aspect ratio.
[0083] The formula for calculating the aspect ratio similarity v is as follows:
[0084]
[0085] Among them, w gt and h gt represents the width and height of the ground truth bounding box, respectively; w and h represent the width and height of the predicted bounding box, respectively.
[0086] The formula for calculating the weighting coefficient α is as follows:
[0087]
[0088] The final CIoU loss function is defined as follows:
[0089]
[0090] S4-4 iterates through steps S4-1 to S4-3 200 times, and uses the validation set of S2-4 to monitor the generalization performance of the model generated in each iteration. The model version with the highest detection accuracy is retained as the infrared image target detection model for UAV inspection.
[0091] S4-5 uses the test set of S2-4 to test the infrared image target detection model of S4-4 for UAV inspection, and determines whether the model can be effectively applied to infrared image target detection in UAV inspection tasks.
[0092] S5. A single UAV carrying a thermal imager acquires infrared images for UAV inspection. The acquired infrared images for UAV inspection are then fed into the infrared image target detection model obtained in S4 for target detection. During the target detection process, the target is predicted based on the optimal Anchor Box size determined in S3. Predicted bounding boxes are generated by combining the updated parameters of the trained model. Overlapping predicted bounding boxes are filtered out by non-maximum suppression to obtain the final target detection result.
[0093] The following inference verification experiments are conducted on the YOLOv5 baseline model and the improved YOLOv5 model under different infrared thermal image scenarios, such as... Figure 4 The inference verification results are for the Yolov5 baseline model. Figure 5 The inference verification results for the improved Yolov5 model are presented. Experimental results show that, in infrared thermal image scenarios, the infrared image target detection model for UAV inspection has higher positioning accuracy and confidence than the Yolov5 baseline model.
[0094] It should be noted that although the embodiments described above are illustrative, they are not intended to limit the invention. Therefore, the invention is not limited to the specific embodiments described above. Any other embodiments obtained by those skilled in the art under the guidance of this invention without departing from its principles are considered to be within the protection scope of this invention.
Claims
1. A target detection method for infrared images for unmanned aerial vehicle (UAV) inspection, characterized in that, The steps include the following: Step 1: Construct an improved Yolov5 model. This improved Yolov5 model introduces the C3 module in the backbone network of the Yolov5 baseline model into a convolutional attention mechanism module consisting of a channel attention module and a spatial attention module. The improved Yolov5 model consists of an input layer, a backbone network, a neck network, and an output layer. The backbone network includes one focusing module, four convolutional modules, four C3 and convolutional attention mechanism modules, and one spatial pyramid pooling module. The neck network includes four convolutional modules, two upsampling modules, four connection modules, and four C3 modules. The input layer is connected to the input of the focusing module. The output of the focusing module is connected to the input of the first convolutional module, the output of the first convolutional module is connected to the input of the first C3 and convolutional attention mechanism module, and the output of the first C3 and convolutional attention mechanism module is connected to the input of the... The inputs of the two convolutional modules are connected as follows: the output of the second convolutional module is connected to the input of the second C3 and convolutional attention mechanism module; the output of the second C3 and convolutional attention mechanism module is connected to the input of the third convolutional module; the output of the third convolutional module is connected to the input of the third C3 and convolutional attention mechanism module; the output of the third C3 and convolutional attention mechanism module is connected to the input of the fourth convolutional module; the output of the fourth convolutional module is connected to the input of the spatial pyramid pooling module; and the output of the spatial pyramid pooling module is connected to the input of the fourth C3 and convolutional attention mechanism module. The output of the mechanism module is connected to the input of the fifth convolutional module. The output of the fifth convolutional module is connected to the input of the first upsampling module. The output of the first upsampling module and the output of the third C3 module and the convolutional attention mechanism module are simultaneously connected to the input of the first connection module. The output of the first connection module is connected to the input of the first C3 module. The output of the first C3 module is connected to the input of the sixth convolutional module. The output of the sixth convolutional module is connected to the input of the second upsampling module. The output of the second upsampling module and the output of the second C3 module and the convolutional attention mechanism module are simultaneously connected to the input of the second connection module. The output of the second connection module is connected to the input of the second C3 module. The output of the second C3 module is connected to the input of the seventh convolutional module. The output of the seventh convolutional module and the output of the sixth convolutional module are simultaneously connected to the input of the third connection module. The output of the third connection module is connected to the input of the third C3 module. The output of the third C3 module is connected to the input of the eighth convolutional module. The output of the eighth convolutional module and the output of the fifth convolutional module are simultaneously connected to the input of the fourth C3 module. The outputs of the second C3 module, the third C3 module, and the fourth C3 module are connected to the output layer. The aforementioned C3 and convolutional attention mechanism module consists of a convolutional attention mechanism module and a C3 module. The convolutional attention mechanism module includes a channel attention module and a spatial attention module. The C3 module includes three lightweight convolutional modules, n cascaded bottleneck modules, and one connection module. Here, n represents the number of bottleneck modules and is a positive integer greater than or equal to 1. The first and fourth C3 and convolutional attention mechanism modules each have one bottleneck module (n), while the second and third C3 and convolutional attention mechanism modules each have three bottleneck modules (n). The input of the attention module forms the input of C3 and the convolutional attention mechanism module. The output of the channel attention module is connected to the input of the spatial attention module. The output of the spatial attention module is connected to the input of the first lightweight convolution module. The output of the first lightweight convolution module is simultaneously connected to the input of the second lightweight convolution module and the input of n cascaded bottleneck modules. The output of the second lightweight convolution module and the output of the n cascaded bottleneck modules are simultaneously connected to the input of the connection module. The output of the connection module is connected to the input of the third lightweight convolution module. The output of the third lightweight convolution module forms the output of C3 and the convolutional attention mechanism module. Step 2: Construct an infrared image dataset for drone inspection; Step 3: Use the K-Means clustering algorithm to perform a one-time anchor frame size calculation on the infrared image dataset for UAV inspection constructed in Step 2 to determine the optimal anchor frame size; Step 4: Input the infrared image dataset for UAV inspection constructed in Step 2 into the improved Yolov5 model constructed in Step 1 for model training. During the model training process, the target is predicted and a prediction box is generated based on the optimal anchor box size determined in Step 3. Then, the localization loss between the prediction box and the labeled box is calculated using the full intersection-union localization loss function. Through repeated forward propagation and backward propagation operations, the infrared image target detection model for UAV inspection is obtained. Step 5: Acquire infrared images for drone inspection using a single drone equipped with a thermal imager, and input the acquired infrared images for drone inspection into the infrared image target detection model obtained in Step 4 for target detection. During the target detection process, the target is predicted and a prediction box is generated based on the optimal anchor frame size determined in Step 3, and the final target detection result is obtained.
2. The infrared image target detection method for UAV inspection according to claim 1, characterized in that, The specific process of step 2 is as follows: Step 2.1: Acquire an infrared image set for drone inspection using a single drone equipped with a thermal imager; Step 2.2: Manually annotate the infrared image set for UAV inspection obtained in Step 2.1 using professional image annotation tools; Step 2.3: Perform Mosaic data augmentation on the infrared image set for UAV inspection that has been annotated in Step 2.2; Step 2.4: Divide the infrared image set for UAV inspection obtained by Mosaic data augmentation in Step 2.3 into training set, validation set and test set according to a predetermined ratio.
3. The infrared image target detection method for UAV inspection according to claim 1, characterized in that, The specific process of step 3 is as follows: Step 3.1: Extract the width and height of the bounding boxes from the infrared image dataset for UAV inspection as data samples, and determine the number of clusters to be divided. Step 3.2: Randomly select a data sample as the initial cluster center for each cluster; Step 3.3: For each data sample, calculate its Euclidean distance to the current cluster center of each cluster, and assign each data sample to the cluster to which the current cluster center with the nearest distance belongs; Step 3.4: For each cluster, calculate the mean of its internal data samples as the new cluster center; Step 3.5: Repeat steps 3.3 and 3.4 until the preset stopping condition is met, that is, the preset threshold for the change in cluster centers or the preset number of iterations is reached. At this time, the cluster center of each cluster is the calculated optimal anchor frame size.
Citation Information
Patent Citations
Traffic sign detection method based on deep learning
CN116977975A