Lightweight target detection method, device and equipment based on infrared image
By constructing and optimizing infrared image object detection data set and network model, combined with knowledge distillation technology, the problem of existing infrared object detection method model bloated is solved, and the model is lightweight and detection efficiency is improved.
Patent Information
- Application Number
- CN202411936251.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-09
AI Technical Summary
The existing infrared object detection method model is relatively bloated, with large model parameters, which cannot meet the usage needs of some edge computing platforms.
By constructing a drone infrared image object detection dataset and using dual-stream YOLOv5m network and pruned YOLOv5n network for training, combined with knowledge distillation technology, the student model is optimized to adapt to the edge computing platform.
The model is lightweight and speed improvement, the detection efficiency is improved, the computing resources and memory requirements are reduced, and the speed and accuracy of infrared drone image target detection are significantly improved.
Smart Images

Figure CN119964029A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer vision and artificial intelligence, and specifically to a lightweight target detection method, device and equipment based on infrared images. Background Art
[0002] Infrared imaging technology converts the thermal radiation emitted by an object into an electrical signal through an infrared detector, which is then processed and finally displayed in the form of an image. This technology is not affected by lighting conditions, so the infrared target detection system has excellent anti-interference ability and can adapt to harsh climates and nighttime environments. At present, infrared target detection technology has been widely used in many fields such as military reconnaissance, autonomous driving, and security monitoring.
[0003] Traditional infrared target detection methods mainly rely on prior knowledge and manually designed features. Such methods usually extract target features through filters, morphological operations, image segmentation and target extraction, and use classifiers for target discrimination. However, this method often requires manual intervention and expert experience, has high requirements on the shape, size, background and other conditions of the target, and is difficult to adapt to complex and changing scenes.
[0004] With the rise of deep learning, infrared target detection methods based on convolutional neural networks (CNNs) have made significant progress. CNNs can automatically extract features from images and learn the representation of targets by training the network. Some classic CNN architectures, such as Faster R-CNN, YOLO, and SSD, have been successfully applied to infrared target detection tasks. These methods achieve accurate and efficient target detection by dividing images into multiple regions or anchors and using convolutional networks for target classification and position regression. In addition, some improved technologies have also been introduced into infrared target detection. For example, the attention mechanism can help the network better focus on the important areas of the target; feature fusion can improve detection performance by fusing multi-scale or multi-level features; target tracking and target segmentation technologies can achieve real-time target detection in video sequences. However, the models of such methods are often bloated and have a large number of model parameters, which cannot meet the needs of some edge computing platforms.
[0005] Therefore, there is an urgent need for a target detection method for infrared images to solve the problem that the existing infrared target detection method models are often bloated, have a large number of model parameters, and cannot meet the needs of some edge computing platforms. Summary of the invention
[0006] This specification provides a lightweight target detection method, device and equipment based on infrared images, which are used to solve the problem that the existing infrared target detection method model is often bloated, the model parameters are large, and cannot meet the needs of some edge computing platforms.
[0007] In a first aspect, this specification provides a lightweight target detection method based on infrared images, comprising:
[0008] Construction and preprocessing of the UAV infrared image target detection dataset, and dividing the dataset into training set, validation set and test set;
[0009] Build a two-stream YOLOv5m network and a pruned YOLOv5n network;
[0010] Use the drone training set to train the two-stream YOLOv5m network and the pruned YOLOv5n network to obtain the initial teacher model and the initial student model.
[0011] The initial teacher model is used to perform knowledge distillation training on the initial student model to obtain an optimized student model, which is then applied to the drone test set.
[0012] In a second aspect, the present specification provides a lightweight target detection device based on infrared images, including: a data set construction module, a two-stream network construction and network pruning optimization module, a two-stream network and pruning network training module, and a distillation training module; wherein:
[0013] The data set construction module is used for constructing and preprocessing the drone infrared image target detection data set, and dividing the data set into a training set, a validation set, and a test set;
[0014] The dual-stream network construction and network pruning optimization module is used to construct a dual-stream YOLOv5m network and a pruned YOLOv5n network;
[0015] The dual-stream network and pruned network training module is used to respectively train the dual-stream YOLOv5m network and the pruned YOLOv5n network using the drone training set to obtain an initial teacher model and an initial student model;
[0016] The distillation training module is used to perform knowledge distillation training on the initial student model using the initial teacher model to obtain an optimized student model, and apply the optimized student model to the drone test set.
[0017] In a third aspect, this specification also provides a network device, including: a communication interface, a processor, and a memory;
[0018] The processor calls the program instructions in the memory to perform the following actions:
[0019] Construction and preprocessing of the UAV infrared image target detection dataset, and dividing the dataset into training set, validation set and test set;
[0020] Build a two-stream YOLOv5m network and a pruned YOLOv5n network;
[0021] Use the drone training set to train the two-stream YOLOv5m network and the pruned YOLOv5n network to obtain the initial teacher model and the initial student model.
[0022] The initial teacher model is used to perform knowledge distillation training on the initial student model to obtain an optimized student model, which is then applied to the drone test set.
[0023] The beneficial effects of the present invention are as follows:
[0024] This specification discloses a lightweight target detection method based on infrared images. The method divides the data set into a training set, a validation set, and a test set by constructing a drone infrared image target detection data set and image preprocessing; constructing a dual-stream YOLOv5m network and a pruned YOLOv5n network; using the drone training set to train the dual-stream YOLOv5m network and the pruned YOLOv5n network respectively to obtain an initial teacher model and an initial student model; finally, using the initial teacher model to perform knowledge distillation training on the initial student model to obtain an optimized student model, and applying the optimized student model to the drone test set. The method constructs a dual-stream feature extraction network in the backbone network of the YOLOv5m network, so that the model can more accurately identify and locate the target position, and improves the model detection accuracy and recognition ability; performing RTOSS pruning operations on the YOLOv5n network to achieve lightweight models. At the same time, the knowledge distillation technology is used to transfer the knowledge of the YOLOv5m network model to the YOLOv5n network model, and the YOLOv5n model after distillation learning is used to detect the target position, which greatly improves the detection accuracy of the model, realizes the lightweight and speed improvement of the model, improves the detection efficiency, reduces the demand for computing resources and memory, and has high practical value and broad application prospects. This method can effectively reduce the number of network parameters and computational complexity, and significantly improve the speed and accuracy of infrared drone image target detection under limited computing power. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings:
[0026] Figure 1 is a schematic diagram of lightweight target detection based on infrared images provided in an embodiment of this specification;
[0027] Figure 2It is a schematic diagram of a dual-stream YOLOv5m network structure provided in an embodiment of this specification;
[0028] Figure 3 It is a schematic diagram of an RTOSS pruning process provided in an embodiment of this specification;
[0029] Figure 4 It is a schematic diagram of a knowledge distillation training process of an initial teacher model on an initial student model provided in an embodiment of this specification;
[0030] Figure 5 is a schematic diagram of a lightweight target detection device based on infrared images provided in an embodiment of this specification;
[0031] Figure 6 It is a schematic diagram of a network device structure provided in an embodiment of this specification. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solutions and advantages of this specification clearer, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and their corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this document.
[0033] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings. Specific embodiment one:
[0035] This embodiment provides a lightweight target detection method based on infrared images. Figure 1 , including the following steps:
[0036] Step 102: construct and preprocess a UAV infrared image target detection dataset, and divide the dataset into a training set, a validation set, and a test set;
[0037] Specifically, one implementation of step 102 may be:
[0038] S21. Use thermal infrared cameras to capture a large amount of drone infrared image data, and use LabelMe image annotation tools to accurately annotate drone infrared images;
[0039] S22. During the labeling process, the correct category label is assigned to the drone target to ensure that the target's bounding box accurately surrounds the target; after the labeling is completed, a labeling file containing the target category and location information is generated for each image;
[0040] S23. After the labeling task is completed, the UAV infrared image dataset is divided into three parts: training set, validation set and test set;
[0041] Specifically, after the drone category information is marked in detail, this embodiment scientifically divides the data set into a training set, a validation set, and a test set according to the ratio of 80%, 10%, and 10%. Such a division helps to effectively evaluate the performance of the model and optimize it in subsequent model training.
[0042] It should be noted that the main function of the training set is to train the drone detection model; the validation set plays the role of performance evaluation during the training process to timely understand the training effect of the model; and the test set is used for the final comprehensive evaluation of the performance of the model.
[0043] S24. After data preprocessing is completed, in order to improve the accuracy of drone recognition, the target image is divided into the upper left and lower right corner coordinate areas using an image cropping algorithm to remove redundant background information, and then the image resolution is scaled to 416*416*3 using the image scaling algorithm in the OpenCV library to form a target data set, wherein the target data set is used for training one-way network of the YOLOv5m dual-stream feature layer to ensure that the model can focus more on the features of the drone target, thereby improving recognition accuracy.
[0044] Step 104: construct a dual-stream YOLOv5m network and a pruned YOLOv5n network;
[0045] Specifically, one implementation of step 104 may be:
[0046] S41, select a YOLOv5m network with a larger number of parameters, construct a two-way feature extraction layer in its backbone network, and combine it with the Head prediction module in the YOLOv5m backbone network to obtain an optimized two-stream YOLOv5m network;
[0047] S42. Select a lighter model YOLOv5n for RTOSS pruning operation to reduce the number of YOLOv5n model parameters and obtain a pruned YOLOv5n network.
[0048] Furthermore, a specific implementation method of step S41 may be:
[0049] S411, use the Backbone module and Neck module of the backbone network of YOLOv5m as the feature extraction layer to build a two-stream network;
[0050] The first feature extraction layer is used to extract the global features of the drone training set images;
[0051] The global feature information comes from the deep network. As the number of network layers increases, the receptive field becomes larger. Therefore, the global information of the feature map obtained by the deep network will be richer. At this time, the resolution of the feature map is relatively low, and the receptive field of a single pixel is relatively large, which can capture more feature information of medium and large targets.
[0052] The second feature extraction layer is used to extract local features of the image block at the target location;
[0053] The local feature information comes from the shallow network, that is, fine-grained information. At this time, the receptive field is relatively small. Therefore, the feature map obtained by the shallow network has richer local information. The resolution of the feature map at this time is relatively high and can capture more feature information of small targets.
[0054] S412. Combine the output ends of the first feature extraction layer and the second feature extraction layer, and connect them to the convolutional layer in the YOLOv5m prediction network to obtain an optimized dual-stream YOLOv5m network.
[0055] Furthermore, a specific implementation method of step S42 may be:
[0056] S421. Use the depth-first search algorithm to find the parent-child layer coupling in the YOLOv5n model;
[0057] S422, for the parent layer 3×3 kernel and the child layer 1×1 kernel, generate possible pattern masks and apply specific criteria to reduce the number of kernel patterns used;
[0058] S423. Prune the 3×3 and 1×1 kernels in the parent and child layers according to the generated pattern mask, where the 1×1 kernels are first combined into a 3×3 kernel temporary weight matrix and then pruned. After pruning, they are decomposed back into 1×1 kernels and redistributed to their original layers.
[0059] Specific example:
[0060] After image preprocessing, the collected infrared images are sent to the target detection and recognition network framework based on deep learning for target detection. Compared with traditional algorithms, the detection and recognition algorithm based on deep learning has the characteristics of high accuracy and high efficiency and is widely used in battlefield environments.
[0061] The YOLOv5 network is more prominent in deep learning target detection algorithms. It not only maintains high accuracy, but also has a lightweight design, which makes the model faster during detection and better suited to complex environments. Therefore, the YOLOv5 network model is selected as the basic framework for infrared image target detection for research.
[0062] The YOLOv5 backbone network architecture includes four key components, namely the image input module, the Backbone module, the Neck module and the Head prediction module, which work together to achieve end-to-end target detection tasks. YOLOv5 has multiple sub-versions, their overall structure and default resolution are the same, and two parameters are used for transformation: depth_multiple and width_multiple, which determine the network depth and the number of convolution kernels respectively. The scale of the network model can be quickly adjusted by changing this parameter. Among them. The parameters depth_multiple and width_multiple corresponding to YOLOv5m are 0.67 and 0.75 respectively. The parameters depth_multiple and width_multiple corresponding to the YOLOv5n network are 0.33 and 0.25 respectively. In this embodiment, the YOLOv5m network and the YOLOv5n network are used to jointly train the target detection model.
[0063] More specifically, in this embodiment, the drone infrared image training set is fed into the image input module of the network, and an adaptive image scaling mechanism is used to ensure that the image input to the model is adapted to the 416×416 input size requirement of the YOLOv5 model while maintaining content integrity. During the training process, the Mosaic data enhancement technology is used to randomly resize and crop four independent images, and then combine them in a mosaic manner to form a new training sample, expanding the model's understanding and adaptability to diverse scene layouts and target scale changes. Secondly, the adaptive anchor box calculation is used to generate a prediction box, which is compared with the real annotation box (ground truth) to calculate the deviation between them. This difference is then used for reverse optimization to iteratively update the parameters of the network.
[0064] In this embodiment, the Backbone module and the Neck module of the YOLOv5m network are used as feature extraction layers to construct a two-stream network, such as Figure 2 As shown in the figure. The first feature extraction layer is used to extract the global features of the drone infrared image. The second feature extraction layer is used to extract the local features of the infrared target position image. The two features are fused and sent to the convolution layer in the Head prediction module of the YOLOv5m network to obtain the optimized two-stream YOLOv5m network architecture.
[0065] More specifically, the YOLOv5m structure backbone network includes a slice structure Focus, a CBL module, a bottleneck layer C3 and a spatial pyramid pooling SPP. The input image is downsampled multiple times by the backbone network to extract features of several different scales, and then passes through a top-down feature pyramid fusion network structure FPN, and then through a bottom-up path aggregation network structure PAN to fuse feature information of different scales. Finally, the Head prediction module adjusts the feature maps of three different scales of 19×19, 38×38, and 76×76 in the last detection layer to a vector form corresponding to the hard prediction for subsequent training and distillation.
[0066] like Figure 3 As shown, in this embodiment, RTOSS pruning is performed on the YOLOv5n network to compress the model size and reduce the model parameters. An iterative pruning method is used, first using a depth-first search (DFS) algorithm to find parent-child layer couplings in the model, and specific pruning for the size of 3×3 and 1×1 kernels. During the pruning process, a pattern mask is generated and the number of kernel patterns used is reduced by specific criteria.
[0067] More specifically, we first use the pre-trained model as input, and use the DFS algorithm to identify the parent-child layer coupling in the model and build a parent-child graph. This process traverses the model layers and uses DFS search to determine the parent layer of each layer. If a layer has no parent layer, it is considered as a parent layer; if a layer is identified as a child layer of other layers, it is added to the corresponding parent layer group. Doing so can reduce the computational requirements of pruning, because the pruning of the parent layer will be reflected in the child layer.
[0068] Using the 3×3 parent layer kernel weights as input, traverse the 3×3 kernel and apply kernel pattern pruning. This process involves calculating the L2 norm within each layer and finding the most suitable kernel pattern for pruning. The pruned pattern will be applied to the parent layer and child layers, reducing the pruning time of the entire model.
[0069] By performing a 1×1 to 3×3 transformation, connectivity pruning can be removed from kernel pruning, which maintains the accuracy of the model and mitigates the loss caused by connectivity pruning. 1×1 kernel pruning can also speed up inference by grouping similar kernel patterns together. The specific steps include flattening the 1×1 kernel weights, then grouping every 9 weights into a 3×3 temporary weight matrix, applying 3×3 kernel pruning to these matrices, and then converting the output matrix back to a 1×1 kernel and redistributing it to the original 1×1 kernel weights.
[0070] Based on this, by building a dual-stream feature extraction network in the backbone network of the YOLOv5m network, the model can more accurately identify and locate the target position, thereby improving the model's detection accuracy and recognition ability; RTOSS pruning operations are performed on the YOLOv5n network to achieve lightweight models.
[0071] Step 106: Use the drone training set to train the dual-stream YOLOv5m network and the pruned YOLOv5n network to obtain an initial teacher model and an initial student model.
[0072] Specifically, one implementation of step 106 may be:
[0073] S61, send the drone training set to the image input module of the two-stream YOLOv5m network for training to obtain the initial teacher model;
[0074] S62. Send the drone training set to the image input module of the pruned YOLOv5n network for training to obtain the initial student model.
[0075] Specific example:
[0076] The improved two-stream YOLOv5m network and the pruned YOLOv5n network are trained for multiple rounds using the same infrared image training set. In each round of training, the network predicts the target position in each image and calculates the difference between the predicted result and the true label, i.e., the loss function. The network parameters are updated based on the Adam optimization algorithm until the difference between the predicted result and the true label reaches the set threshold, the training is stopped, and the pt model is obtained. The pt model is the initial teacher model and the initial student model.
[0077] Step 108: Use the initial teacher model to perform knowledge distillation training on the initial student model to obtain an optimized student model, and apply the optimized student model to the drone test set.
[0078] Specifically, one implementation of step 108 may be:
[0079] S81, the initial teacher model and the initial student model are respectively subjected to the improved Softmax function to obtain soft labels and soft predictions, and the two values are cross-entropy calculated to obtain the distillation loss;
[0080] S82, the initial student model is hard predicted by the Softmax function, and the student loss is obtained after cross entropy calculation with the true label;
[0081] S83, the total loss is obtained by weighted summation of the distillation loss and the student loss, so that the result of the student model is close to the result of the teacher model, and the optimized student model can be obtained by distillation training;
[0082] S84. Input the drone test set into the optimized student model for target detection.
[0083] Specific example:
[0084] The knowledge distillation method can protect the knowledge learned in the original model by designing the distillation loss. In this embodiment, when training the initial student model, the output of the initial teacher model is used to constrain the initial student model through the distillation loss, thereby protecting the knowledge of the drone infrared image data learned in the initial teacher model.
[0085] In this embodiment, the specific method of using the initial teacher model to perform a distillation operation on the initial student model is:
[0086] First, freeze the initial teacher model parameters. After the input image passes through the initial teacher model and the initial student model respectively, it passes through the last detection layer. The meanings of the output vectors are: the horizontal coordinate of the center point of the target box, the vertical coordinate of the center point of the target box, the width of the target box, the height of the target box, the foreground probability, and the probability of belonging to each category.
[0087] like Figure 4 As shown, the output vector of the initial teacher model and the vector after the sliced output of the initial student model are passed through the improved Softmax function, and the result of the initial teacher model is regarded as a soft label, and the result of the initial student model is regarded as a soft prediction. The soft label and the soft prediction are cross-entropy calculated to obtain the distillation loss, and the hard prediction output by the student model through the Softmax function and the true label are cross-entropy calculated to obtain the student loss. The total loss is obtained by weighted summation of the distillation loss and the student loss, so that the result of the student model is close to the result of the teacher model, that is, the optimized student model is obtained by distillation training, thereby maintaining the student model's ability to recognize infrared images of drones.
[0088] Based on this, by using knowledge distillation technology, the knowledge of the YOLOv5m network model is transferred to the YOLOv5n network model, and the YOLOv5n model after distillation learning is used to detect the target position, which greatly improves the detection accuracy of the model, realizes the lightweight and speed improvement of the model, improves the detection efficiency, and reduces the demand for computing resources and memory.
[0089] In summary, this embodiment constructs a drone infrared image target detection dataset and image preprocessing, divides the dataset into a training set, a validation set, and a test set; constructs a two-stream YOLOv5m network and a pruned YOLOv5n network; uses the drone training set to train the two-stream YOLOv5m network and the pruned YOLOv5n network respectively to obtain an initial teacher model and an initial student model; finally, uses the initial teacher model to perform knowledge distillation training on the initial student model to obtain an optimized student model, and applies the optimized student model to the drone test set. This method can effectively reduce the number of network parameters and computational complexity, and significantly improve the speed and accuracy of infrared drone image target detection under limited computing power. Specific embodiment 2:
[0091] This embodiment provides a lightweight target detection method based on infrared images, and the specific steps are as follows:
[0092] Step 1: Construction and preprocessing of the UAV infrared image target detection dataset, and dividing the dataset into training set, validation set and test set.
[0093] A thermal infrared camera is used to capture a large amount of drone infrared image data, which covers a variety of lighting conditions, backgrounds, shooting angles, and scale changes.
[0094] For the infrared image target detection task, a series of preprocessing steps are required for the dataset. First, the infrared image data is calibrated using the LabelMe image annotation tool, and the drone category information is annotated in detail to ensure the accuracy and completeness of the label. Then, the dataset is scientifically divided into training set, validation set, and test set according to the ratio of 80%, 10%, and 10%. Such a division helps to effectively evaluate the performance of the model and optimize it in subsequent model training.
[0095] After data preprocessing, in order to improve the accuracy of drone recognition, the target image is segmented according to the upper left and lower right coordinate areas using OpenCV technology, effectively removing redundant background information and forming a target data set for training the one-way network of the YOLOv5m dual-stream feature layer to ensure that the model can focus more on the characteristics of the drone target, thereby improving recognition accuracy.
[0096] Step 2: Build a two-stream YOLOv5m network and a pruned YOLOv5n network.
[0097] After image preprocessing, the collected infrared images are sent to the target detection and recognition network framework based on deep learning for target detection. Compared with traditional algorithms, the detection and recognition algorithm based on deep learning has the characteristics of high accuracy and high efficiency and is widely used in battlefield environments. The YOLOv5 network is more prominent in the deep learning target detection algorithm. It not only has the ability to maintain high accuracy, but also has a lightweight design, which makes the model faster in detection and better suitable for complex environments. Therefore, the YOLOv5 network model is selected as the basic framework for infrared image target detection for research.
[0098] The YOLOv5 backbone network architecture includes four key components, namely the image input module, the Backbone module, the Neck module and the Head prediction module, which work together to achieve end-to-end target detection tasks. YOLOv5 has multiple sub-versions, their overall structure and default resolution are the same, and two parameters are used for transformation: depth_multiple and width_multiple, which determine the network depth and the number of convolution kernels respectively. The scale of the network model can be quickly adjusted by changing this parameter. Among them. The parameters depth_multiple and width_multiple corresponding to YOLOv5m are 0.67 and 0.75 respectively. The parameters depth_multiple and width_multiple corresponding to the YOLOv5n network are 0.33 and 0.25 respectively. In this embodiment, the YOLOv5m network and the YOLOv5n network are used to jointly train the target detection model.
[0099] More specifically, in this embodiment, the drone infrared image training set is fed into the image input module of the network, and an adaptive image scaling mechanism is used to ensure that the image input to the model is adapted to the 416×416 input size requirement of the YOLOv5 model while maintaining content integrity. During the training process, the Mosaic data enhancement technology is used to randomly resize and crop four independent images, and then combine them in a mosaic manner to form a new training sample, expanding the model's understanding and adaptability to diverse scene layouts and target scale changes. Secondly, the adaptive anchor box calculation is used to generate a prediction box, which is compared with the real annotation box (ground truth) to calculate the deviation between them. This difference is then used for reverse optimization to iteratively update the parameters of the network.
[0100] In this embodiment, the Backbone module and the Neck module of the YOLOv5m network are used as feature extraction layers to construct a two-stream network, such as Figure 2As shown in the figure. The first feature extraction layer is used to extract the global features of the drone infrared image. The second feature extraction layer is used to extract the local features of the infrared target position image. The two features are fused and sent to the convolution layer in the YOLOv5m prediction module to obtain the optimized two-stream YOLOv5m network architecture.
[0101] More specifically, the YOLOv5m structure backbone network includes a slice structure Focus, a CBL module, a bottleneck layer C3 and a spatial pyramid pooling SPP. The input image is downsampled multiple times by the backbone network to extract features of several different scales, and then passes through a top-down feature pyramid fusion network structure FPN, and then through a bottom-up path aggregation network structure PAN to fuse feature information of different scales. Finally, the Head prediction module adjusts the feature maps of three different scales of 19×19, 38×38, and 76×76 in the last detection layer to a vector form corresponding to the hard prediction for subsequent training and distillation.
[0102] like Figure 3 As shown, in this embodiment, RTOSS pruning is performed on the YOLOv5n network to compress the model size and reduce the model parameters. Current pruning techniques mainly focus on 3×3 convolution kernels, which limits the degree of achievable sparsity and inference acceleration. Most models, such as YOLOv5, RetinaNet, and DETR, are composed of 68.42%, 56.14%, and 63.46% of 1×1 small convolution kernels, respectively. Therefore, in order to increase the sparsity of such models, connection pruning is sometimes used on these 3×3 convolution kernels. However, the "last kernel in each layer" standard used in connection pruning can lead to the loss of important information, thereby affecting the accuracy of the model. Therefore, this embodiment selects RTOSS pruning and adopts an iterative pruning method. First, a depth-first search (DFS) algorithm is used to find the parent-child layer coupling in the model. Tracking DFS, identifying 3×3 and 1×1 kernels in the subgraph, and applying kernel-specific pruning to them, reduces computational cost and time overhead.
[0103] More specifically, the computation graph (G) is computed using the pre-trained model as input and the gradients obtained from back-propagation. An empty list (group_list) is initialized to store parent-child layer groups. The model layers (l) are then traversed and a DFS search is applied on the computation graph G to identify the parent layers of the layer. If a layer does not have any parent layer, then the layer is assigned as its own parent layer (lp), which becomes a group. If a layer is identified as a child layer (lc) of any layer in group_list, the layer now becomes the parent layer (lp) of the child layer (lc) and is added to the group. Each parent layer (lp) can have multiple child layers (lp), but each child layer can only have one parent layer (lp). This process continues until all layers are assigned to a group. Since the layers in each group have coupled channels, they also share their kernel weights and can therefore share the same kernel mode.
[0104] The pattern masks are generated in all possible combinations by the standard combinatorial method, using the following formula:
[0105]
[0106] Where n is the size of the matrix and k is the size of the pattern mask. The following two criteria are then used to reduce the number of kernel patterns used: discarding all patterns without adjacent non-zero weights, maintaining the semi-structured nature of the kernel patterns; and selecting the most frequently used kernel patterns by computing the L2 norm of the kernel using random initialization in the range [-1,1]. The value of k can range from 1 to 8, which can generate 8 different types of pattern groups.
[0107] Connection pruning is removed from kernel pruning by performing a 1×1 to 3×3 transformation. This maintains the model accuracy and mitigates the loss from connection pruning. 1×1 kernel pruning can also speed up inference by grouping similar kernel patterns together.
[0108] Step 3: Use the drone training set to train the two-stream YOLOv5m network and the pruned YOLOv5n network respectively to obtain the initial teacher model and the initial student model.
[0109] The improved two-stream YOLOv5m network and the pruned YOLOv5n network are trained for multiple rounds using the same infrared image training set. In each round of training, the network predicts the target position in each image and calculates the difference between the predicted result and the true label, i.e., the loss function. The network parameters are updated based on the Adam optimization algorithm until the difference between the predicted result and the true label reaches the set threshold, the training is stopped, and the pt model is obtained. The pt model is the initial teacher model and the initial student model.
[0110] Step 4: Use the initial teacher model to perform knowledge distillation training on the initial student model to obtain an optimized student model, and apply the optimized student model to the drone test set.
[0111] The knowledge distillation method can protect the knowledge learned in the original model by designing the distillation loss. In this embodiment, when training the initial student model, the output of the initial teacher model is used to constrain the initial student model through the distillation loss, thereby protecting the knowledge of the drone infrared image data learned in the initial teacher model.
[0112] In this embodiment, the specific method of using the initial teacher model to perform a distillation operation on the initial student model is: first, freeze the initial teacher model parameters, and after the input image passes through the initial teacher model and the initial student model respectively, it passes through the last detection layer, and the meanings of the output vectors are: the horizontal coordinate of the center point of the target box, the vertical coordinate of the center point of the target box, the target box width, the target box height, the foreground probability, and the probability of belonging to each category.
[0113] like Figure 4 As shown, the output vector of the initial teacher model and the vector after the sliced output of the initial student model are passed through the improved Softmax function, and the result of the initial teacher model is regarded as a soft label, and the result of the initial student model is regarded as a soft prediction. The soft label and the soft prediction are cross-entropy calculated to obtain the distillation loss, and the hard prediction output by the student model through the Softmax function and the true label are cross-entropy calculated to obtain the student loss. The total loss is obtained by weighted summation of the distillation loss and the student loss, so that the result of the student model is close to the result of the teacher model, that is, the optimized student model is obtained by distillation training, thereby maintaining the student model's ability to recognize infrared images of drones.
[0114] More specifically, the hard predictions of the student model are obtained through the Softmax function, as follows:
[0115] Among them, q i The probability of each category output, z j Fully connected output for each category.
[0116] Considering that during distillation training, although more attention is paid to the value of the category with the largest probability value in the soft label, for distillation, other small probability values are also the knowledge learned by the teacher network and should also be used. Since the difference between these values is too large, the deformed Softmax function is introduced for smoothing. The specific method is to add the temperature parameter T. The improved Softmax formula is:
[0117] Obviously, when T=1, the function is the original Softmax function, and the calculated result is called a hard prediction. When T>1, as T increases, the difference in the exponential function caused by the input value becomes smaller, and the calculated result is called a soft label. The softer the label, the more obvious the information of the probability of the incorrect category is.
[0118] Figure 4 There are two types of loss functions in : distillation loss and student loss. Assume that the soft labels and soft predictions output by the initial teacher model and the initial student model at temperature T are The distillation loss is obtained by calculating the cross entropy of the two values. The calculation formula is as follows:
[0119] The true label of the data is recorded as c j , then the hard prediction output by the student model is After calculating the cross entropy with the true label, the value of the student loss can be obtained. The formula is as follows:
[0120] The value of the overall loss function is expressed by the distillation loss, the student loss, and the weight coefficients α and β:
[0121] L=αL soft +βL hard (2)
[0122] The drone test set is input into the optimized student model for target detection. Based on the detection results of the student model after knowledge distillation, the performance of the student model after knowledge distillation in terms of precision, recall, and mAP indicators is calculated to evaluate the performance of the student model after knowledge distillation;
[0123] More specifically, precision is the proportion of samples predicted by the model to be positive examples that are truly positive; recall is the proportion of positive examples correctly predicted by the model to all true positive examples; mAP is the average of the average precision under all IoU thresholds.
[0124] In summary, this embodiment constructs a drone infrared image target detection dataset and image preprocessing, divides the dataset into a training set, a validation set, and a test set; constructs a two-stream YOLOv5m network and a pruned YOLOv5n network; uses the drone training set to train the two-stream YOLOv5m network and the pruned YOLOv5n network respectively to obtain an initial teacher model and an initial student model; finally, uses the initial teacher model to perform knowledge distillation training on the initial student model to obtain an optimized student model, and applies the optimized student model to the drone test set. This method can effectively reduce the number of network parameters and computational complexity, and significantly improve the speed and accuracy of infrared drone image target detection under limited computing power. Specific embodiment three:
[0126] This embodiment provides a lightweight target detection device based on infrared images. Figure 5 , including: a data set construction module 501, a two-stream network construction and network pruning optimization module 502, a two-stream network and pruning network training module 503, and a distillation training module 504; wherein:
[0127] The data set construction module 501 is used for constructing and preprocessing a data set for target detection in infrared images of unmanned aerial vehicles, and dividing the data set into a training set, a validation set and a test set;
[0128] The dual-stream network construction and network pruning optimization module 502 is used to construct a dual-stream YOLOv5m network and a pruned YOLOv5n network;
[0129] The dual-stream network and pruned network training module 503 is used to use the drone training set to train the dual-stream YOLOv5m network and the pruned YOLOv5n network respectively to obtain an initial teacher model and an initial student model;
[0130] The distillation training module 504 is used to perform knowledge distillation training on the initial student model using the initial teacher model to obtain an optimized student model, and apply the optimized student model to the drone test set.
[0131] Optionally, the data set construction module 501 is specifically used for:
[0132] Use thermal infrared cameras to capture a large amount of drone infrared image data, and use LabelMe image annotation tools to accurately annotate drone infrared images;
[0133] During the annotation process, the correct category label is assigned to the drone target to ensure that the target's bounding box accurately surrounds the target; after the annotation is completed, a labeling file containing the target category and location information is generated for each image;
[0134] After the labeling task is completed, the drone infrared image dataset is divided into three parts: training set, validation set and test set;
[0135] After data preprocessing, the target image is divided into the upper left corner and lower right corner coordinate areas using the image cropping algorithm to remove redundant background information, and the image resolution is scaled using the image scaling algorithm in the OpenCV library to form a target data set, where the target data set is used to train the one-way network of the YOLOv5m dual-stream feature layer.
[0136] Optionally, the dual-stream network construction and network pruning optimization module 502 includes: a dual-stream network construction unit and a network pruning optimization unit; wherein:
[0137] The dual-stream network construction unit is used to select a YOLOv5m network with a larger parameter amount, construct a two-way feature extraction layer in its backbone network, and combine it with the prediction network to obtain an optimized dual-stream YOLOv5m network;
[0138] The network pruning optimization unit is used to select a lighter small model YOLOv5n to perform RTOSS pruning operation to reduce the model parameter amount of YOLOv5n and obtain a pruned YOLOv5n network.
[0139] Optionally, the dual-stream network construction unit is specifically used to:
[0140] The Backbone module and Neck module of the YOLOv5m backbone network are used as feature extraction layers to build a two-stream network;
[0141] The first feature extraction layer is used to extract the global features of the drone training set images;
[0142] The second feature extraction layer is used to extract local features of the image block at the target location;
[0143] The output ends of the first feature extraction layer and the second feature extraction layer are combined and connected to the convolutional layer in the YOLOv5m prediction network to obtain the optimized two-stream YOLOv5m network.
[0144] Optionally, the network pruning optimization unit is specifically used to:
[0145] Use the depth-first search algorithm to find the parent-child layer coupling in the YOLOv5n model;
[0146] Generate possible pattern masks for the parent 3×3 kernel and the child 1×1 kernel, and apply specific criteria to reduce the number of kernel patterns used;
[0147] According to the generated pattern mask, the 3×3 and 1×1 kernels in the parent and child layers are pruned, where the 1×1 kernels are first combined into a 3×3 kernel temporary weight matrix before being pruned, and then decomposed back into 1×1 kernels after pruning and redistributed to their original layers.
[0148] Optionally, the dual-stream network and pruning network training module 503 is specifically used for:
[0149] The drone training set is fed into the image input module of the two-stream YOLOv5m network for training to obtain the initial teacher model;
[0150] The drone training set is sent to the image input module of the pruned YOLOv5n network for training to obtain the initial student model.
[0151] Optionally, the distillation training module 504 is specifically used to:
[0152] The initial teacher model and the initial student model are respectively subjected to the improved Softmax function to obtain soft labels and soft predictions, and the distillation loss is obtained by cross entropy calculation of the two values;
[0153] The initial student model is subjected to the Softmax function to obtain hard predictions, and the student loss is obtained after cross entropy calculation with the true label;
[0154] The total loss is obtained by weighted summation of the distillation loss and the student loss, so that the result of the student model is close to the result of the teacher model, and the optimized student model can be obtained by distillation training;
[0155] The drone test set is input into the optimized student model for target detection.
[0156] Optionally, during knowledge distillation training, the initial teacher model parameters are frozen, and the input image passes through the initial teacher model and the initial student model respectively, and then passes through the last detection layer. The meanings of the output vectors are: the horizontal coordinate of the center point of the target box, the vertical coordinate of the center point of the target box, the width of the target box, the height of the target box, the foreground probability, and the probability of belonging to each category.
[0157] In summary, this embodiment constructs a drone infrared image target detection dataset and image preprocessing, divides the dataset into a training set, a validation set, and a test set; constructs a two-stream YOLOv5m network and a pruned YOLOv5n network; uses the drone training set to train the two-stream YOLOv5m network and the pruned YOLOv5n network respectively to obtain an initial teacher model and an initial student model; finally, uses the initial teacher model to perform knowledge distillation training on the initial student model to obtain an optimized student model, and applies the optimized student model to the drone test set. This method can effectively reduce the number of network parameters and computational complexity, and significantly improve the speed and accuracy of infrared drone image target detection under limited computing power.
[0158] Specific example 4:
[0159] This manual also provides a network device, see Figure 6 , the network device can implement the details of the method described in the above embodiment and achieve the same effect. Figure 6 As shown, the electronic device 600 includes: a processor 601, a transceiver 602, a memory 603, a user interface 604 and a bus interface, wherein:
[0160] In the embodiment of this specification, the network device 600 further includes: a computer program stored in the memory 603 and executable on the processor 601, and the computer program is executed by the processor 601 to implement the following steps:
[0161] Construction and preprocessing of the UAV infrared image target detection dataset, and dividing the dataset into training set, validation set and test set;
[0162] Build a two-stream YOLOv5m network and a pruned YOLOv5n network;
[0163] Use the drone training set to train the two-stream YOLOv5m network and the pruned YOLOv5n network to obtain the initial teacher model and the initial student model.
[0164] The initial teacher model is used to perform knowledge distillation training on the initial student model to obtain an optimized student model, which is then applied to the drone test set.
[0165] exist Figure 6 In the embodiment, the bus framework may include any number of interconnected buses and bridges, specifically one or more processors represented by processor 601 and various circuits of memory represented by memory 603 are linked together. The bus framework may also link together various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described in this document. The bus interface provides an interface. The transceiver 602 may be a plurality of components, namely, a transmitter and a receiver, providing a unit for communicating with various other devices on a transmission medium. For different user devices, the user interface 604 may also be an interface capable of externally and internally connecting required devices, and the connected devices include but are not limited to a keypad, a display, a speaker, a microphone, a joystick, and the like.
[0166] The processor 601 manages the bus architecture and the usual processing, and the memory 603 can store data used by the processor 601 when performing operations. Optionally, when the computer program is executed by the processor 601, the following steps can also be implemented:
[0167] Optionally, the construction and preprocessing of the drone infrared image target detection dataset, and dividing the dataset into a training set, a validation set, and a test set, include:
[0168] Use thermal infrared cameras to capture a large amount of drone infrared image data, and use LabelMe image annotation tools to accurately annotate drone infrared images;
[0169] During the annotation process, the correct category label is assigned to the drone target to ensure that the target's bounding box accurately surrounds the target; after the annotation is completed, a labeling file containing the target category and location information is generated for each image;
[0170] After the labeling task is completed, the drone infrared image dataset is divided into three parts: training set, validation set and test set;
[0171] After data preprocessing, the target image is divided into the upper left corner and lower right corner coordinate areas using the image cropping algorithm to remove redundant background information, and the image resolution is scaled using the image scaling algorithm in the OpenCV library to form a target data set, where the target data set is used to train the one-way network of the YOLOv5m dual-stream feature layer.
[0172] Optionally, the constructing of a dual-stream YOLOv5m network and a pruned YOLOv5n network includes:
[0173] Select the YOLOv5m network with a larger number of parameters, build two feature extraction layers in its backbone network, and combine them with the Head prediction module in the YOLOv5m backbone network to obtain an optimized two-stream YOLOv5m network;
[0174] Select the lighter model YOLOv5n for RTOSS pruning operation to reduce the number of YOLOv5n model parameters and obtain the pruned YOLOv5n network.
[0175] Optionally, the YOLOv5m network with a larger number of parameters is selected, and two feature extraction layers are constructed in its backbone network, which are combined with the Head prediction module in the YOLOv5m backbone network to obtain an optimized two-stream YOLOv5m network, including:
[0176] The Backbone module and Neck module of the YOLOv5m backbone network are used as feature extraction layers to build a two-stream network;
[0177] The first feature extraction layer is used to extract the global features of the drone training set images;
[0178] The second feature extraction layer is used to extract local features of the image block at the target location;
[0179] The output ends of the first feature extraction layer and the second feature extraction layer are combined and connected to the convolutional layer in the YOLOv5m prediction network to obtain the optimized two-stream YOLOv5m network.
[0180] Optionally, the selecting of a lighter small model YOLOv5n to perform RTOSS pruning operation to reduce the model parameter amount of YOLOv5n, and obtaining a pruned YOLOv5n network includes:
[0181] Use the depth-first search algorithm to find the parent-child layer coupling in the YOLOv5n model;
[0182] Generate possible pattern masks for the parent 3×3 kernel and the child 1×1 kernel, and apply specific criteria to reduce the number of kernel patterns used;
[0183] According to the generated pattern mask, the 3×3 and 1×1 kernels in the parent and child layers are pruned, where the 1×1 kernels are first combined into a 3×3 kernel temporary weight matrix before being pruned, and then decomposed back into 1×1 kernels after pruning and redistributed to their original layers.
[0184] Optionally, the method of using the drone training set to respectively train a dual-stream YOLOv5m network and a pruned YOLOv5n network to obtain an initial teacher model and an initial student model includes:
[0185] The drone training set is fed into the image input module of the two-stream YOLOv5m network for training to obtain the initial teacher model;
[0186] The drone training set is sent to the image input module of the pruned YOLOv5n network for training to obtain the initial student model.
[0187] Optionally, the using the initial teacher model to perform knowledge distillation training on the initial student model to obtain an optimized student model, and applying the optimized student model to the drone test set, includes:
[0188] The initial teacher model and the initial student model are respectively subjected to the improved Softmax function to obtain soft labels and soft predictions, and the distillation loss is obtained by cross entropy calculation of the two values;
[0189] The initial student model is subjected to the Softmax function to obtain hard predictions, and the student loss is obtained after cross entropy calculation with the true label;
[0190] The total loss is obtained by weighted summation of the distillation loss and the student loss, so that the result of the student model is close to the result of the teacher model, and the optimized student model can be obtained by distillation training;
[0191] The drone test set is input into the optimized student model for target detection.
[0192] Optionally, during knowledge distillation training, the initial teacher model parameters are frozen, and the input image passes through the initial teacher model and the initial student model respectively, and then passes through the last detection layer. The meanings of the output vectors are: the horizontal coordinate of the center point of the target box, the vertical coordinate of the center point of the target box, the width of the target box, the height of the target box, the foreground probability, and the probability of belonging to each category.
[0193] In summary, this embodiment can effectively reduce the number of network parameters and computational complexity, and significantly improve the speed and accuracy of infrared UAV image target detection under limited computing power.
[0194] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. For those skilled in the art, this specification may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included in the protection scope of this specification.
Claims
1. A lightweight target detection method based on infrared images, characterized in that: include: Construction and preprocessing of the UAV infrared image target detection dataset, and dividing the dataset into training set, validation set and test set; Build a two-stream YOLOv5m network and a pruned YOLOv5n network; Use the drone training set to train the two-stream YOLOv5m network and the pruned YOLOv5n network to obtain the initial teacher model and the initial student model. The initial teacher model is used to perform knowledge distillation training on the initial student model to obtain an optimized student model, which is then applied to the drone test set.
2. The method according to claim 1, characterized in that The construction and preprocessing of the UAV infrared image target detection dataset, and the division of the dataset into a training set, a validation set, and a test set, include: Use thermal infrared cameras to capture a large amount of drone infrared image data, and use LabelMe image annotation tools to accurately annotate drone infrared images; During the annotation process, the correct category label is assigned to the drone target to ensure that the target's bounding box accurately surrounds the target; after the annotation is completed, a labeling file containing the target category and location information is generated for each image; After the labeling task is completed, the drone infrared image dataset is divided into three parts: training set, validation set and test set; After data preprocessing, the target image is divided into the upper left corner and lower right corner coordinate areas using the image cropping algorithm to remove redundant background information, and the image resolution is scaled using the image scaling algorithm in the OpenCV library to form a target data set, where the target data set is used to train the one-way network of the YOLOv5m dual-stream feature layer.
3. The method according to claim 2, characterized in that The construction of the dual-stream YOLOv5m network and the pruned YOLOv5n network includes: Select the YOLOv5m network with a larger number of parameters, build two feature extraction layers in its backbone network, and combine them with the Head prediction module in the YOLOv5m backbone network to obtain an optimized two-stream YOLOv5m network; Select the lighter model YOLOv5n for RTOSS pruning operation to reduce the number of YOLOv5n model parameters and obtain the pruned YOLOv5n network.
4. The method according to claim 3, characterized in that The YOLOv5m network with a larger number of parameters is selected, and a two-way feature extraction layer is constructed in its backbone network, which is combined with the Head prediction module in the YOLOv5m backbone network to obtain an optimized two-stream YOLOv5m network, including: The Backbone module and Neck module of the YOLOv5m backbone network are used as feature extraction layers to build a two-stream network; The first feature extraction layer is used to extract the global features of the drone training set images; The second feature extraction layer is used to extract local features of the image block at the target location; The output ends of the first feature extraction layer and the second feature extraction layer are combined and connected to the convolutional layer in the YOLOv5m prediction network to obtain the optimized two-stream YOLOv5m network.
5. The method according to claim 3, characterized in that: The method of selecting a lighter model YOLOv5n to perform RTOSS pruning operation to reduce the model parameters of YOLOv5n and obtain a pruned YOLOv5n network includes: Use the depth-first search algorithm to find the parent-child layer coupling in the YOLOv5n model; Generate possible pattern masks for the parent 3×3 kernel and the child 1×1 kernel, and apply specific criteria to reduce the number of kernel patterns used; According to the generated pattern mask, the 3×3 and 1×1 kernels in the parent and child layers are pruned, where the 1×1 kernels are first combined into a 3×3 kernel temporary weight matrix before being pruned, and then decomposed back into 1×1 kernels after pruning and redistributed to their original layers.
6. The method according to claim 5, characterized in that The method uses the drone training set to train the dual-stream YOLOv5m network and the pruned YOLOv5n network to obtain an initial teacher model and an initial student model, including: The drone training set is fed into the image input module of the two-stream YOLOv5m network for training to obtain the initial teacher model; The drone training set is sent to the image input module of the pruned YOLOv5n network for training to obtain the initial student model.
7. The method according to claim 6, characterized in that The method uses the initial teacher model to perform knowledge distillation training on the initial student model to obtain an optimized student model, and applies the optimized student model to the drone test set, including: The initial teacher model and the initial student model are respectively subjected to the improved Softmax function to obtain soft labels and soft predictions, and the distillation loss is obtained by cross entropy calculation of the two values; The initial student model is subjected to the Softmax function to obtain hard predictions, and the student loss is obtained after cross entropy calculation with the true label; The total loss is obtained by weighted summation of the distillation loss and the student loss, so that the result of the student model is close to the result of the teacher model, and the optimized student model can be obtained by distillation training; The drone test set is input into the optimized student model for target detection.
8. The method according to claim 7, characterized in that During knowledge distillation training, the parameters of the initial teacher model are frozen. After the input image passes through the initial teacher model and the initial student model respectively, it passes through the last detection layer. The meanings of the output vectors are: the horizontal coordinate of the center point of the target box, the vertical coordinate of the center point of the target box, the width of the target box, the height of the target box, the foreground probability, and the probability of belonging to each category.
9. A lightweight target detection device based on infrared images, applied to the method according to any one of claims 1 to 8, characterized in that: include: Dataset construction module, two-stream network construction and network pruning optimization module, two-stream network and pruning network training module, distillation training module; among them: The data set construction module is used for constructing and preprocessing the drone infrared image target detection data set, and dividing the data set into a training set, a validation set, and a test set; The dual-stream network construction and network pruning optimization module is used to construct a dual-stream YOLOv5m network and a pruned YOLOv5n network; The dual-stream network and pruned network training module is used to respectively train the dual-stream YOLOv5m network and the pruned YOLOv5n network using the drone training set to obtain an initial teacher model and an initial student model; The distillation training module is used to use the initial teacher model to perform knowledge distillation training on the initial student model to obtain an optimized student model, and apply the optimized student model to the drone test set.
10. A network device, characterized in that: include: Communications interface, processor and memory; The processor calls the program instructions in the memory to perform the following actions: Construction and preprocessing of the UAV infrared image target detection dataset, and dividing the dataset into training set, validation set and test set; Build a two-stream YOLOv5m network and a pruned YOLOv5n network; Use the drone training set to train the two-stream YOLOv5m network and the pruned YOLOv5n network to obtain the initial teacher model and the initial student model. The initial teacher model is used to perform knowledge distillation training on the initial student model to obtain an optimized student model, which is then applied to the drone test set.
Citation Information
Cited By
Infrared image vessel target detection method based on lightweight YOLOv10
CN121074627A