A Real-time Small Target Recognition Method and System for UAV Infrared Detection
Through the improved LCNet network and feature pyramid structure, small target feature extraction is enhanced, combined with loss function optimization, the accuracy and real-time problems of small target recognition in the infrared detection of drones are solved, and efficient infrared small target detection is achieved.
Patent Information
- Application Number
- CN202410413937.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-04-08
AI Technical Summary
The prior art is difficult to achieve high detection rate, low false alarm rate and high real-time performance for small targets in infrared detection of drones, especially in complex backgrounds, small target recognition accuracy is low, and the general neural network model has large parameters and long detection time, making it difficult to meet the real-time requirements.
The improved LCNet network is used as the feature extraction backbone network, combined with the Squeeze-and-Excitation module and the improved feature pyramid structure, and the neck network can be separated by deep separable convolution and feature fusion, enhance the extraction and recognition capabilities of small-target features, and adjust the loss function weight to optimize model performance.
It improves the accuracy of small target recognition of drone infrared detection, reduces false detection rates and missed detection rates, and realizes efficient operation on platforms with limited memory and computing capabilities, meeting real-time requirements.
Smart Images

Figure CN118314477B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of visual detection technologies, and particularly to a method and system for real-time recognition of small targets for unmanned aerial vehicle (UAV) infrared detection. Background Art
[0002] In recent years, UAV detection technologies have become important reconnaissance and detection methods in civilian and military fields due to their good flexibility, concealment, and efficiency. However, due to the interference of punctate high-brightness noise and similar target interference objects on the infrared images collected by UAVs, especially under the scene conditions with complex background radiation and clutter, it poses a major challenge to the task of long-distance recognition of small targets. Moreover, due to the long imaging distance, the targets usually appear as dense small targets, and the proportion of target pixels is small; the radiation intensity of infrared images is weak, the signal-to-noise ratio is low, and the shape and texture information is very limited, which increases the false alarm rate in the detection process. In addition, the currently commonly used YOLO-based detection networks have problems such as large model parameter quantities, long detection times, and high deployment costs on mobile devices. For the task of detecting small targets at a long distance, it is difficult for general image processing technologies to achieve high detection rates, low false alarm rates, and high real-time targets.
[0003] Currently, small target detection algorithms can be divided into neural network methods, spatio-temporal tensor methods, low-rank sparse decomposition, etc. In terms of computational complexity, spatio-temporal tensor methods and low-rank sparse decomposition algorithms involve relatively complex mathematical operations and optimization processes, and the algorithm complexity is high. Especially when dealing with large-scale infrared image data, these problems may be more prominent and it is difficult to meet the real-time requirements. In contrast, deep learning algorithms can automatically learn feature information through network training, can reduce the amount of calculation by adjusting the network structure, and have the advantages of strong feature extraction ability, strong generalization ability, and high real-time performance. Deep learning algorithms are divided into two categories: single-stage algorithms and two-stage algorithms. Two-stage algorithms based on candidate regions, such as R-CNN, can achieve good results in terms of accuracy, but the process of generating candidate regions in R-CNN is more time-consuming compared to SSD and YOLO. In real-time application scenarios, the detection speed of R-CNN may become a bottleneck for its application. The specific process is as Figure 1 shown. Single-stage algorithms include SSD and YOLO, etc. Such algorithms adopt more efficient region generation methods and can achieve faster detection speeds. The specific process is as Figure 2 shown. However, general neural network object detection algorithms are not suitable for the infrared small target recognition application scenario and have problems such as low recognition accuracy. As Figure 3 shown is the flow block diagram of a general object detection algorithm.
[0004] Therefore, it is of great significance and application value to study how to accurately process UAV infrared small target detection technologies in various complex scenarios. Summary of the Invention
[0005] The object of the present invention is to overcome the defects of the prior art, and a small target real-time recognition method and system for UAV infrared detection are proposed.
[0006] To achieve the above object, the present invention proposes a small target real-time recognition method for UAV infrared detection, including:
[0007] Input the infrared image collected by the UAV into the trained recognition model, and output the detection result;
[0008] The recognition model includes: a feature extraction backbone network improved based on LCNet, a feature fusion neck network improved based on LCPAN, and an object detection network; wherein,
[0009] The feature extraction backbone network introduces an SE module for extracting small target features of different scales in the input image;
[0010] The feature fusion neck network is used to fuse small target features of different scales to generate multi-scale features by integrating the feature pyramid structure into lower-level features and improving the downsampling rate in the spatial dimension of the feature map;
[0011] The object detection network is used for object classification and regression and outputs the small target detection result.
[0012] Preferably, the feature extraction backbone network includes 6 Blocks, namely Block1 to Block6. Each Block uses depthwise separable convolution and an optional SE module inside to extract features, and outputs multi-scale feature maps of different levels through the forward propagation function; wherein,
[0013] Block1 includes 1 Conv2D convolutional layer, 1 BatchNorm2D batch normalization layer, and a Hard_Swish activation function;
[0014] Block2 includes 1 layer of depthwise separable convolutional layer with a convolution kernel of 3×3, an input channel number of 16, and an output channel number of 32;
[0015] Block3 includes 2 layers of 3×3 depthwise separable convolutional layers. The first layer has an input channel number of 32, an output channel number of 64, and a stride of 2; the second layer has an input channel number of 64, an output channel number of 64, and a stride of 1;
[0016] Block4 includes 2 layers of 3×3 depthwise separable convolutional layers. The first layer has an input channel number of 64, an output channel number of 128, and a stride of 2; the second layer has an input channel number of 128, an output channel number of 128, and a stride of 1;
[0017] Block 5 includes a 3×3 depthwise separable convolutional layer and five 5×5 depthwise separable convolutional layers. Among them, the 3×3 convolutional layer and the first three 5×5 convolutional layers do not use the SE module, while the SE module is added to the last two 5×5 depthwise separable convolutional layers.
[0018] Block 6 includes two 5×5 depthwise separable convolutions, and the SE module is used in each layer.
[0019] Preferably, the feature fusion neck network changes the feature map levels for the feature pyramid network in LCPAN from three, four, and five layers to two, three, and four layers to enhance the detection ability for small targets; and adjusts the large stride of the feature pyramid network in LCPAN to four, eight, sixteen, and thirty-two small strides to improve the recognition accuracy for small targets.
[0020] Preferably, the object detection network includes: a PicoheadV2 four-detection-head network, an object classification network, and a regression network; where
[0021] The PicoheadV2 four-detection-head network includes four head networks, which respectively receive and process the fusion results of feature maps of different scales and generate a set of detection results including object category and location information;
[0022] The object classification network is used to determine the category of the object in each prediction box;
[0023] The regression network is used to frame the object boundary by adjusting the position and size of the prediction box.
[0024] Preferably, the method further includes the training step of the recognition model; including:
[0025] Establish a data set and perform preprocessing on the data set including data format conversion and data augmentation;
[0026] Input the preprocessed data set into the recognition model, calculate the loss function, and update the weights of the recognition model through the backpropagation algorithm to minimize the loss function and obtain the trained model.
[0027] Preferably, the data format conversion includes: converting the YOLO format of the HIT-UAV data set into the VOC data set format convenient for network processing;
[0028] The data augmentation process includes: randomly flipping the image with a set probability to help the model learn the robustness to directions; randomly resizing the image to any size in [256, 288, 320, 352, 384] for data augmentation; changing the image dimension order through the Permute operation to ensure that the data dimension matches the input dimension expected by the model; and padding the image through the PadGT operation to ensure meeting the size requirements during the model training process.
[0029] Preferably, the loss function includes the target classification network loss function multiplied by the first weight coefficient and the regression network loss function multiplied by the second weight coefficient.
[0030] On the other hand, the present invention proposes a small target real-time recognition system for UAV infrared detection, and the system includes:
[0031] A detection module, configured to input the infrared image collected by the UAV into the trained recognition model and output a detection result;
[0032] The recognition model includes: a feature extraction backbone network improved based on LCNet, a feature fusion neck network improved based on LCPAN, and a target detection network; wherein,
[0033] The feature extraction backbone network introduces an SE module for extracting small target features of different scales in the input image;
[0034] The feature fusion neck network is used to fuse small target features of different scales to generate multi-scale features by integrating the feature pyramid structure into lower-level features and improving the downsampling rate in the spatial dimension of the feature map;
[0035] The target detection network is used for target classification and regression and outputs a small target detection result.
[0036] Compared with the prior art, the advantages of the present invention are:
[0037] 1. Aiming at the problem of detection real-time performance, an improved lightweight LCNet network is proposed as the feature extraction backbone network, which uses depthwise separable convolution to fully extract small target features and reduces the computational amount at the same time, enabling the model to run more effectively on platforms with limited memory and computing power.
[0038] 2. Aiming at the problems of high false detection rate caused by low resolution and few features of small targets, and high missed detection rate caused by large target scale span and coexistence of multiple scales, the Squeeze-and-Excitation module and the improved feature pyramid structure are added to fuse shallow fine-grained features and high-level semantic features, and the receptive field is expanded by fusing different scale features to enhance the accuracy of the network in recognizing small targets.
[0039] 3. According to the characteristics of the image feature information in the HIT-UAV dataset, adjust parameters such as the loss function weight to reduce the false detection rate of small UAV infrared targets. Description of the Drawings
[0040] Figure 1 is a diagram of a two-stage detection network framework;
[0041] Figure 2 is a diagram of a single-stage detection network framework;
[0042] Figure 3 is a flowchart of the general neural network object detection process;
[0043] Figure 4 is a block diagram of the existing unimproved Picodet recognition model;
[0044] Figure 5 is a block diagram of the small target real-time recognition model for UAV infrared detection of the present invention;
[0045] Figure 6(a) is the recognition result of MobilenetV1, Figure 6(b) is the recognition result of the unimproved Picodet model, and Figure 6(c) is the recognition result of the method of the present invention. Detailed Embodiments
[0046] The technical solution of the present invention will be described in detail below with reference to the drawings and embodiments.
[0047] A small target real-time recognition method for UAV infrared detection includes: inputting the infrared image collected by the UAV into a trained recognition model and outputting a detection result;
[0048] The recognition model is an improved Picodet neural network model, including: a feature extraction backbone network improved based on LCnet, a feature fusion neck network improved based on LCPAN, an object detection network, and an improved loss function; among them,
[0049] The backbone network improved based on LCNet is used to extract small target features of different scales in the input image. During model training, the original infrared image collected by the UAV is preprocessed by operations such as cropping and flipping, and then sent into the improved LCnet backbone network for processing.
[0050] The preprocessing operation is used to help the model better learn and generalize. Data preprocessing mainly includes data format conversion and data augmentation processing. Among them, data augmentation includes image standardization, image dimension order adjustment, image padding, random cropping, random flipping of the image, and random distortion or deformation of the image.
[0051] The data format conversion process is used to convert the original YOLO format of the HIT-UAV dataset into the VOC dataset format that is convenient for network processing. It is characterized in that Python scripts are used to batch convert the YOLO format txt annotation files into VOC format XML annotation files. The specific method is as follows: First, create a dictionary dic to map the class numbers in the classes.txt file to the class names in the VOC format while maintaining the order unchanged; then, obtain the list of TXT annotation files in the YOLO format dataset file through os.listdir(txtPath), and for each TXT file, perform the following steps in sequence: ① Create a new XML document and an annotation element, open the corresponding TXT file and read its content; ② Use cv2.imread to read the corresponding image file, obtain the size information (height, width, depth) of the image, create XML elements such as folder, filename, size, etc., and fill in the content. For each line of data in the TXT file, create an object element, and parse out information such as the class and position (bounding box) of the object according to the YOLO annotation format. Create a bndbox element, and calculate the xmin, ymin, xmax, and ymax values of the bounding box. Add the created object element to the annotation; ③ Write the annotation into the XML file and save it to the specified path.
[0052] The data augmentation process described above is used to generate more training samples to increase the generalization ability of the model and prevent overfitting. It is characterized in that the image is randomly flipped with a probability of 50% to help the model learn the robustness to directions; the image is randomly resized to any size in [256, 288, 320, 352, 384] for data augmentation; the dimension order of the image is changed through the Permute operation to ensure that the data dimension matches the input dimension expected by the model; the image is padded through the PadGT operation to ensure that the size requirements during the model training process are met.
[0053] The backbone network improved based on LCnet includes: 6 Blocks, namely Block1 - Block6. Inside each Block, depthwise separable convolution and an optional SE module are used to extract features, and multi-scale feature maps of different levels are output through the forward propagation function; among them,
[0054] The Block1 includes a Conv2D convolutional layer, a BatchNorm2D batch normalization layer, and a Hard_Swish activation function. The convolutional layer slides over the input image data with the convolutional kernel size to extract the local features of the target. The normalization layer is used to standardize the input data to stabilize the training process, accelerate convergence, and improve the generalization ability of the model. The activation function layer introduces non-linearity to learn more complex feature representations.
[0055] The Block2 - Block4 convolutional blocks are configured similarly. Each block is a depthwise separable convolutional layer, and each layer's configuration consists of the convolutional kernel size, the number of input / output channels, the stride, and whether to use the SE module. The convolutional kernel sizes of Block2 - Block4 are all 3×3, and the SE module is not used. The number of input channels of Block2 is 16, the number of output channels is 32, and the stride is 1. Block3 consists of 2 layers. The number of input channels of the first layer is 32, the number of output channels is 64, and the stride is 2; the number of input channels of the second layer is 64, the number of output channels is 64, and the stride is 1. Block4 consists of 2 layers. The number of input channels of the first layer is 64, the number of output channels is 128, and the stride is 2; the number of input channels of the second layer is 128, the number of output channels is 128, and the stride is 1. The convolutional layers with larger strides are used to reduce the spatial size of the feature map to reduce the computational amount and the number of parameters of the network. The above Blocks correspond to different layers of the feature pyramid, and each layer corresponds to different feature sizes and resolutions, which are used to extract information features of different scales of the input data.
[0056] The Block5 includes 1 layer of 3×3 depthwise separable convolutional layer and 5 layers of 5×5 depthwise separable convolutional layers. Among them, the 3×3 convolutional layer and the first 3 layers of 5×5 convolutional layers do not use the SE module, while the last 2 layers of 5×5 convolutional layers add the SE module. Adding the SE module in the later network layers is used to adaptively adjust the importance of different channels to enhance the feature representation ability.
[0057] The Block6 contains 2 layers of 5×5 depthwise separable convolution, and each layer uses the SE module. As the Block number increases (from Block1 to Block6), the size of the feature map gradually decreases, and the number of channels gradually increases. Adding the SE module at deeper levels can make the network better focus on small-scale features.
[0058] As an improvement of the above method, the improved LCnet backbone network enhances the ability to extract small-target features by introducing the SE module in Block5. The network structure before improvement is shown in Table 1. The network structure after improvement is shown in Table 2.
[0059] The neck network improved based on LCPAN is used to fuse different-scale features extracted by the Backbone to generate multi-scale features. The feature maps obtained by feature extraction of the backbone network are processed to obtain feature layers and sent to the neck network. By introducing an improved feature pyramid structure, the features of lower layers are fused with the features of other layers, making the model more sensitive to small targets. The improved neck network fuses the processed N feature layers to obtain N + 1 fused feature layers;
[0060] As an improvement of the above method, the neck network is mainly improved by integrating the features of lower layers into the feature pyramid structure and improving the downsampling rate in the spatial dimension of the feature map. The improved feature pyramid structure is used to capture features from rough to detailed at different levels, so as to effectively process targets of various sizes. In the original model, the feature map levels used for the feature pyramid network are layers 3, 4, and 5. As Figure 4 shown. According to the characteristics of UAV infrared small target data, it is improved to layers 2, 3, and 4. By introducing low-level features, the model can be more sensitive to small targets, thus enhancing the detection ability for small targets. The improvement of the downsampling rate of the feature map is mainly achieved by adjusting the stride of the feature pyramid structure. Adjusting the stride can change the resolution of the feature map and thus affect the detection ability for small targets. According to the characteristics of the HIT-UAV dataset, the large stride of the feature pyramid network in the original model is adjusted to small strides of 4, 8, 16, and 32, so as to retain more detailed information and improve the recognition accuracy for small targets. As Figure 5 shown.
[0061] The object detection network is used for object classification and regression and outputs detection results. It includes the PicoheadV2 detection head network and the classification and regression network. The N + 1 feature layers are sent to the PicoheadV2 four-detection head network for object classification and regression and output detection results;
[0062] The PicoheadV2 four-detection head network is used to receive and process the fusion results of different-scale feature maps, and then send the results to the classification and regression network to achieve accurate detection and positioning of objects. The fusion results corresponding to different-scale feature maps are respectively used as the inputs of 4 head networks (Head1, Head2, Head3, and Head4). Each detection head independently predicts the received feature information and generates a set of detection results, including object category and location information. By processing the features of these different layers, multi-scale information can be fully learned, thus enhancing the detection ability for objects of different sizes.
[0063] The target classification and regression network is used to further process the prediction results of the object category and position information obtained by the head network. Among them, the classification network is responsible for determining the category of the object in each prediction box, while the regression network is responsible for adjusting the position and size of the prediction box to make it more accurately frame the object boundary.
[0064] The loss function is used to calculate the error between the prediction result and the actual label, and update the weights of the model through the backpropagation algorithm to minimize this error.
[0065] As an improvement to the above method, the improvement of the loss function mainly has two aspects. On the one hand, the weight of the classification task loss function VarifocalLoss is adjusted. Specifically, the weight is increased from 1.0 to 1.5. In the calculation of the final total loss, the result of this classification loss function will be multiplied by 1.5 to increase its importance in training, which is used to solve the problem of class imbalance in the object detection task. On the other hand, the weight of the bounding box regression loss function GIoULoss is adjusted. Specifically, the weight is adjusted from 2.5 to 3.0, which is used to optimize the performance of the bounding box regression in the detection task.
[0066] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0067] Embodiment 1
[0068] Embodiment 1 of the present invention proposes a real-time small target recognition method for UAV infrared detection. The specific steps are as follows: First, in the model training stage, the original infrared images collected by the UAV are preprocessed; then the image to be processed is input into the improved LCnet backbone network for small target feature extraction. The improved LCnet backbone network enhances the ability to extract small target features by introducing more SE modules; then the obtained feature map is processed to obtain a feature layer and sent to the neck network. By introducing an improved feature pyramid structure, the features of lower layers are fused with the features of other layers to make the model have a more sensitive representation of small targets. The improved neck network fuses the processed N feature layers to obtain N + 1 fused feature layers. Finally, the N + 1 fused feature layers are sent to the object detection network for object classification and regression. After the model training is completed, the model is converted using the opt tool and exported as a.nb file. The model is deployed to the application platform, and the image to be detected is input. Through the model inference operation, the detection result is output, thereby realizing the detection function of the UAV to real-time recognize infrared small and weak targets.
[0069] As Figure 1 shown, it is a two-stage detection network framework diagram, mainly composed of a backbone network, a neck network, a region proposal network, a region feature encoding module, a detection head network, and a post-processing algorithm module.
[0070] Figure 2 It is a diagram of a single-stage detection network framework, mainly composed of a backbone network, a neck network, a detection head network, and a post-processing algorithm module.
[0071] Figure 3 It is a flowchart of a general neural network object detection process. The specific process is as follows: First, perform preprocessing operations such as data augmentation on the dataset images. Secondly, send them into the backbone network to extract different-scale features of the targets. Then, enter the neck connection layer to fuse different-scale features to generate multi-scale features. Next, generate rectangular boxes for predicting targets according to the set anchor box sizes. After that, complete object detection through regional feature encoding, classification and localization, and loss functions.
[0072] Figure 4 It is a block diagram of an unimproved Picodet recognition model. The ESnet network and the CSP-PAN network are adopted. In the unimproved model, the feature layer levels for the feature pyramid network are 3, 4, and 5 layers.
[0073] Figure 5 It is a block diagram of a small target real-time recognition model for UAV infrared detection according to the present invention. The LCnet network and the LCPAN network are adopted. According to the characteristics of UAV infrared small target data, the feature layer levels of the feature pyramid network are improved to 2, 3, and 4 layers. By introducing low-level features, the model can have a more sensitive representation of small targets, thereby enhancing the detection ability for small targets. The improvement of the feature map downsampling rate is mainly achieved by adjusting the stride of the feature pyramid structure. Adjusting the stride can change the resolution of the feature map and thus affect the detection ability for small targets. According to the characteristics of the HIT-UAV dataset, the large stride in the original feature pyramid network of the model is adjusted to 4, 8, 16, and 32 small strides, so as to retain more detailed information and improve the recognition accuracy for small targets. The neck network adopts LCPAN. Compared with the commonly used CSPPAN network structure, LCPAN effectively improves the network accuracy with a larger receptive field and fewer parameters.
[0074] Table 1 is the original LCNet network structure parameters; Table 2 is the structure parameters of the feature extraction backbone network improved based on LCNet of the present invention. As follows: Table 1
[0075] Operator Kernel Size Stride SE Conv2D 3×3 2 - DepthSepConv 3×3 1 - DepthSepConv 3×3 2 - DepthSepConv 3×3 1 - DepthSepConv 3×3 2 - DepthSepConv 3×3 1 - DepthSepConv 3×3 2 - 5×DepthSepConv 5×5 1 √ DepthSepConv 5×5 2 √ DepthSepConv 5×5 1 - GAP 7×7 1 - Conv2D,NBN 1×1 1 -
[0076] Table 2
[0077]
[0078]
[0079] Figure 6(a) shows the recognition result of MobilenetV1, Figure 6(b) shows the recognition result of the unimproved Picodet model, and Figure 6(c) shows the recognition result of the method of the present invention. This shows the comparison of the recognition method of this application with the recognition results of the other two models.
[0080] Embodiment 2
[0081] Embodiment 2 of the present invention proposes a small target real-time recognition system for UAV infrared detection, which is implemented based on the method of Embodiment 1 and includes:
[0082] A detection module, configured to input the infrared image collected by the UAV into the trained recognition model and output a detection result;
[0083] The recognition model includes: a feature extraction backbone network improved based on LCNet, a feature fusion neck network improved based on LCPAN, and a target detection network; wherein,
[0084] The feature extraction backbone network introduces an SE module for extracting small target features of different scales in the input image;
[0085] The feature fusion neck network is used to fuse small target features of different scales to generate multi-scale features by integrating the feature pyramid structure into lower-level features and improving the downsampling rate in the spatial dimension of the feature map;
[0086] The target detection network is used to perform target classification and regression and output small target detection results.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present invention does not depart from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A real-time small target recognition method for UAV infrared detection, comprising: Inputting the infrared image collected by the UAV into a trained recognition model to output a detection result; The recognition model includes: a feature extraction backbone network improved based on LCNet, a feature fusion neck network improved based on LCPAN, and an object detection network; wherein, The feature extraction backbone network introduces an SE module for extracting small target features of different scales in the input image; The feature fusion neck network is used to fuse small target features of different scales to generate multi-scale features by integrating the feature pyramid structure into lower-level features and improving the downsampling rate in the spatial dimension of the feature map; The object detection network is used for object classification and regression and outputs a small target detection result; The feature extraction backbone network includes 6 Blocks, namely Block1 to Block6. Each Block uses depthwise separable convolution and an optional SE module inside to extract features, and outputs multi-scale feature maps of different levels through a forward propagation function; wherein, Block1 includes 1 Conv2D convolutional layer, 1 BatchNorm2D batch normalization layer, and a Hard_Swish activation function; Block2 includes 1 layer of depthwise separable convolutional layer with a convolution kernel of 3×3, an input channel number of 16, and an output channel number of 32; Block3 includes 2 layers of 3×3 depthwise separable convolutional layers. The first layer has an input channel number of 32, an output channel number of 64, and a stride of 2; the second layer has an input channel number of 64, an output channel number of 64, and a stride of 1; Block4 includes 2 layers of 3×3 depthwise separable convolutional layers. The first layer has an input channel number of 64, an output channel number of 128, and a stride of 2; the second layer has an input channel number of 128, an output channel number of 128, and a stride of 1; Block5 includes 1 layer of 3×3 depthwise separable convolutional layer and 5 layers of 5×5 depthwise separable convolutional layers. The 3×3 convolutional layer and the first 3 layers of 5×5 convolutional layers do not use the SE module, and the last 2 layers of 5×5 depthwise separable convolutional layers add the SE module; Block6 includes 2 layers of 5×5 depthwise separable convolution, and each layer uses the SE module; The feature fusion neck network changes the number of feature layers for the feature pyramid network in LCPAN from 3, 4, and 5 layers to 2, 3, and 4 layers to enhance the detection ability for small targets; adjusts the large stride of the feature pyramid network in LCPAN to 4, 8, 16, and 32 small strides to improve the recognition accuracy for small targets.
2. The real-time small target recognition method for UAV infrared detection according to claim 1, characterized in that, The object detection network includes: a PicoheadV2 four-detection-head network, an object classification network, and a regression network; wherein, The PicoheadV2 four-detection-head network includes 4 head networks, which respectively receive and process the fusion results of different scale feature maps and generate a set of detection results including object category and position information; The object classification network is used to determine the category of the object in each prediction box; The regression network is used to frame the object boundary by adjusting the position and size of the prediction box.
3. The real-time small target recognition method for UAV infrared detection according to claim 1, wherein, The method further includes a training step of the recognition model, including: Establishing a data set and performing preprocessing on the data set, including data format conversion and data augmentation processing; Inputting the preprocessed data set into the recognition model, calculating the loss function, and updating the weights of the recognition model through the backpropagation algorithm to minimize the loss function and obtain a trained recognition model.
4. The real-time small target recognition method for UAV infrared detection according to claim 3, characterized in that The data format conversion includes converting the YOLO format of the HIT-UAV data set into the VOC data set format convenient for network processing; The data augmentation processing includes randomly flipping the image with a set probability to help the model learn the robustness to directions; randomly adjusting the size of the image to any size in [256, 288, 320, 352, 384] for data augmentation; changing the image dimension order through the Permute operation to ensure that the data dimension matches the input dimension expected by the model; padding the image through the PadGT operation to ensure meeting the size requirements during the model training process.
5. The real-time small target recognition method for UAV infrared detection according to claim 3, characterized in that, The loss function includes the target classification network loss function multiplied by the first weight coefficient and the regression network loss function multiplied by the second weight coefficient.
6. A system for a small target real-time recognition method for UAV infrared detection based on claim 1, characterized in that, The system includes: A detection module for inputting the infrared image collected by the UAV into the trained recognition model and outputting a detection result; The recognition model includes a feature extraction backbone network improved based on LCNet, a feature fusion neck network improved based on LCPAN, and an object detection network; where The feature extraction backbone network introduces an SE module for extracting small target features of different scales in the input image; The feature fusion neck network is used to fuse small target features of different scales to generate multi-scale features by integrating the feature pyramid structure into lower-level features and improving the downsampling rate in the spatial dimension of the feature map; The object detection network is used for object classification and regression and outputs small target detection results.
Citation Information
Patent Citations
Brachial plexus ultrasound image recognition method based on target detection
CN115496733A
Infrared small target detection method based on improved Yolov5 network
CN115661611A
Wafer lattice dislocation image detection method and system based on deep learning
CN116168033A