An infrared image small target detection method

Through YOLOIR lightweight detection network and enhanced processing, the shortcomings in accuracy and real-time detection of infrared image micro-objects are solved, and efficient detection of mobile terminals is achieved. It is suitable for smart cars, smart homes and robots.

CN115546500BActive Publication Date: 2025-07-18XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211373188.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-07-18
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

The existing infrared image micro-object detection algorithms have shortcomings in accuracy and real-time performance, especially on mobile devices, which are difficult to achieve efficient detection.

Method used

A lightweight detection network based on YOLOIR is adopted, combined with ShuffleNet backbone network, adaptive feature fusion module, Attention feature fusion module and regression head prediction module, small object detection is carried out for infrared images and enhanced processing is performed.

Benefits of technology

Real-time and accurate infrared micro-object detection on the mobile terminal is realized, and is suitable for natural interactions in fields such as smart cars, smart homes and robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115546500B_ABST
    Figure CN115546500B_ABST
Patent Text Reader

Abstract

A method for lightweight small target detection in infrared images, comprising the following steps: S100: Obtain a small target infrared image by using a lightweight detection network based on YOLOIR. Among them, the lightweight detection network structure based on YOLOIR includes a backbone network, an adaptive feature fusion module, an Attention attention feature fusion module, and a regression head prediction module; S200: Perform enhancement processing on the generated small target infrared image. This method has the characteristics of accurate and clear detection of small target images and supports real-time generation, and can be widely used in natural interaction in fields such as intelligent vehicles, smart homes, and robots.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical fields of computer vision, pattern recognition, and artificial intelligence, and particularly relates to an infrared image small target detection method. Background Art

[0002] With the advent of the intelligent era, people's requirements for the application scenarios of vision systems are becoming more diverse. Since visible light cameras are particularly sensitive to light, there are certain limitations in low-light or dim light environments. Therefore, infrared target detection shows great advantages and value. Infrared images have strong anti-interference capabilities and are sensitive to heat sources, and there is an urgent need in many fields, such as unmanned aerial vehicles, smart homes, robots, medical and national defense, etc. On the other hand, existing algorithms mostly target the detection of people and vehicles at close range. How to obtain micro infrared targets with high precision and high accuracy has become the key to target detection research.

[0003] The general steps of traditional video stream target detection: perform frame-by-frame target detection on the images in the input video stream. First, pass through the feature extraction module to output the feature map of the image, and then pass through the feature fusion module to perform fusion processing on the extracted features to obtain the feature map after the fusion of low-dimensional and high-dimensional information. Finally, perform regression prediction on the feature map. The regression prediction obtains the coordinate parameters of the detection box and the class confidence of the target detection, and finally returns the result to the input image. At present, mainstream target detection algorithms such as YOLO do not have additional designs for the characteristics of infrared and small targets on the one hand, so it is difficult to guarantee the accuracy of detecting small targets directly with infrared data. On the other hand, the number of model parameters and the amount of computation are generally too large to achieve real-time performance on mobile devices. Summary of the Invention

[0004] To solve the above problems, the present disclosure provides a method for lightweight infrared image small target detection based on YOLOIR, including the following steps:

[0005] S100: Obtain a small target infrared image by using a lightweight detection network based on YOLOIR. Among them, the lightweight detection network structure based on YOLOIR includes a backbone network, an adaptive feature fusion module, an Attention attention feature fusion module, and a regression head prediction module;

[0006] S200: Perform enhancement processing on the generated small target infrared image.

[0007] Through the above technical solution, the detection of small targets is realized based on the YOLOIR detection network, which has the characteristics of accurate detection and positioning of small targets, high accuracy, and support for real-time generation. This method is not only applicable to the detection of small targets in infrared images, but also applicable to the detection of dynamic small targets in RGB-IR video streams, and can be widely used in natural interaction in fields such as intelligent vehicles, smart homes, and robots. This method can achieve real-time, accurate, and stable detection of small infrared targets on mobile devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a schematic flowchart of a method for detecting small targets in lightweight infrared images based on YOLOIR provided in an embodiment of the present disclosure;

[0009] Figure 2 is a processing flowchart of the detection of small targets by a lightweight infrared image detection network based on YOLOIR in an embodiment of the present disclosure;

[0010] Figure 3 is a flowchart of the implementation of an adaptive feature fusion in an embodiment of the present disclosure;

[0011] Figure 4 is a flowchart of the implementation of an improved FPN integrating an attention mechanism in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] In order to enable those skilled in the art to understand the technical solutions disclosed in the present disclosure, the following will describe the technical solutions of each embodiment in combination with the embodiments and the relevant attached Figures 1 to 4 , and describe the technical solutions of each embodiment. The described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. The terms "first", "second", etc. used in the present disclosure are used to distinguish different objects, rather than to describe a specific order. In addition, "including" and "having" and any variations thereof are intended to cover and non-exclusively include. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, systems, products, or devices.

[0013] Referring to "embodiment" in this article means that a specific feature, structure, or characteristic described in combination with the embodiment may be included in at least one embodiment of the present disclosure. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art can understand that the embodiments described herein can be combined with other embodiments.

[0014] In one embodiment, as Figure 1As shown, a method for small target detection in lightweight infrared images based on YOLOIR is disclosed, including the following steps:

[0015] S100: Obtain a small target infrared image using a lightweight detection network based on YOLOIR. Among them, the lightweight detection network structure based on YOLOIR includes a backbone network, an adaptive feature fusion module, an Attention attention feature fusion module, and a regression head prediction module;

[0016] S200: Perform enhancement processing on the generated small target infrared image.

[0017] For this embodiment, this method includes two steps: generating the detected small target area image through a detection network based on YOLOIR and performing super-resolution enhancement processing on a specific small target area image. The embodiment can detect small targets in infrared images in real time and perform super-resolution optimization for small targets. For the input infrared image video, small target detection is performed frame by frame, the small target area is extracted, and super-resolution processing is performed on the small target area.

[0018] The infrared small target detection algorithm belongs to a task subclass in general object detection and can follow classical models in object detection. There are mainly two types of general object detection algorithms: single-stage, represented by YOLO and SSD, with a simple model, faster speed, and more suitable for practical applications; two-stage, represented by Faster R-CNN, with a more complex model, higher accuracy but slower speed. Object detection algorithms can also be divided into two types based on whether they require prior anchor boxes: Anchor-base and Anchor-free. Among the Anchor-base series algorithms, the most representative is the YOLO series.

[0019] Considering the lightweight deployment requirements, this method designs and implements an infrared small target detection method based on YOLOIR, which belongs to a single-stage object detection network and is applicable to the lightweight infrared small target detection task of this method.

[0020] The infrared small detection network based on YOLOIR draws on the basis of the YOLOV5 framework in structure and improves and upgrades many of its modules. Specifically, it improves the Feature Pyramid Networks (FPN) to an adaptive feature pyramid, introduces an attention mechanism module in the FPN, and optimizes the loss function, etc.

[0021] The entire process of YOLOIR from input to output is as follows: The input infrared image is 640x480, and the short side is padded with pure black pixels to 640. The image input to YOLOIR is 640×640. After three stages of feature extraction by YOLOIR, feature maps downsampled by 8 times, 16 times, and 32 times are obtained respectively, with sizes of 80x80, 40x40, and 20x20. Each pixel on each layer of these feature maps respectively corresponds to an area of 8x8, 16x16, and 32x32 on the original image. Next, K prediction boxes will be generated respectively at each pixel of these three feature maps according to the pre-set prior box parameters. Each box requires 6 parameters (the horizontal and vertical coordinates of the upper left corner, width, height, and target category (person, vehicle)), so the outputs of these three feature maps after passing through the regression head module will be 80×80×K×(4 + 2), 40×40×K×(4 + 2), 20×20×K×(4 + 2), that is, the positions of 80×80×K prediction boxes and the probabilities belonging to different categories. During training, the loss is calculated using these predictions and the true annotations, and during prediction, the non-maximum suppression algorithm is used to process these prediction boxes to obtain the final prediction results.

[0022] The infrared small target detection network based on YOLOIR conforms to the classic design process of Backbone, Neck, and Head in the target detection algorithm, and its network structure mainly includes three main parts:

[0023] 1) The backbone network for feature extraction, usually called Backbone.

[0024] 2) The module FPN for feature fusion, also called the neck of the network.

[0025] 3) The regression head part, usually called head, which is used to regress information such as the coordinates and class confidence of the target from the features processed by the neck module.

[0026] In another embodiment, the backbone network is the lightweight network ShuffleNet.

[0027] For this embodiment, the original feature extraction network is changed to the lightweight network ShuffleNet as the new backbone network, and the feature extraction network is optimized in strict accordance with its lightweight design concept.

[0028] Specifically, since the C3 Layer uses multi-way separation convolution, it will take up more cache space and reduce the running speed. Therefore, it is necessary to avoid using the C3 Layer multiple times and the C3 Layer with high channels. At the same time, remove the 1024conv and 5×5pooling layers of the ShuffleNetbackbone. Because in the scenario of this article, there are only limited categories. After removing this module, the network speed can be accelerated with limited impact on the accuracy.

[0029] In summary, after replacing the original backbone of YOLO with ShuffleNet, the original 7M parameters can be reduced to about 2M parameters, which greatly optimizes the network's Flops and enables it to achieve real-time on mobile terminals.

[0030] In another embodiment, the detection head in the regression head prediction module is decoupled and a 1×1 convolutional decoupling head is added.

[0031] In this embodiment, compared with the traditional YOLOV5 network model, the branches in the head are decoupled. Specifically, a 1x1 convolution is used to reduce the dimension first, and then two 3x3 convolutions are used in the next two branches, and finally the network parameters are adjusted to increase only a little bit, but the detection box position regression and the target category regression can be decoupled accordingly, which more effectively improves the detection accuracy and the accuracy of small targets.

[0032] In another embodiment, step S100 further includes the following steps:

[0033] S101: Generation of prior anchor boxes and matching of anchor boxes and target boxes;

[0034] S102: Perform end-to-end feature extraction and feature fusion on the input image to finally generate a feature map;

[0035] S103: The obtained feature map is respectively regressed through the target frame coordinate regression branch and the category confidence regression branch to obtain the coordinates of the detected target in the current coordinate system and the maximum confidence of the category to which it belongs.

[0036] In this embodiment, the regression head prediction module includes two parts: a target box coordinate regression branch and a category confidence regression branch, one for regressing coordinates and the other for regressing to obtain probability confidences of different categories. The coordinates of the object in the current image coordinate system are specifically the horizontal and vertical coordinates of the upper left corner and the lower right corner.

[0037] The processing process of small target detection based on YOLOIR lightweight infrared image small target detection network is divided into three steps: Figure 2 shown.

[0038] Step 1: Generation of prior anchor boxes and matching of anchor boxes with ground truth (GT) boxes. The basic principle of all single-stage prior-anchor-box-based object detection algorithms can be summarized as classification and regression after dense sampling of the original image. Therefore, generating anchor boxes is an essential step. Although the geometric meaning of anchor boxes is relative to the original image, their specific generation needs to be combined with the feature map. For YOLOIR here, three feature maps in the network are retained, and the downsampling ratios relative to the original image are 1 / 8, 1 / 16, and 1 / 32 respectively.

[0039] Combined with the characteristics of the infrared image dataset of this method and considering speed, in one instance, the size of the input original infrared image is limited to 640×640. Then the scales of the three feature maps are 80x80, 40x40, and 20x20 respectively. Each pixel point on each layer of the feature map corresponds to an area of 8x8, 16x16, and 32x32 on the original image. For traditional algorithms such as Faster R-CNN, SSD, and YOLO, k different scales and aspect ratios of anchor boxes are generated based on each pixel point on the feature map. Generally, k = 9, representing anchor boxes with 3 different scales and 3 width-to-height ratios. Since this method is for the detection of small targets and the detection boxes are relatively small themselves, in fact, the accuracy of width and height is not important. More attention is paid to the accuracy of localization, that is, the x and y coordinates of the center point of the image, while ignoring the aspect ratio to simplify the design of anchor boxes. Further modify the loss function to increase the proportion of the x and y coordinate losses and reduce the proportion of width and height to further improve the network's attention to the localization accuracy.

[0040] The matching of the anchor box and the ground truth box in step S101 further includes: normalizing the center point of the ground truth box relative to the anchor box using the width and height. To eliminate the influence of the scale of the anchor box itself and treat all anchor boxes equally, it is also necessary to normalize the center point of the ground truth box relative to the anchor box using the width and height. If normalization is not performed, large anchor boxes can tolerate larger deviations, while small anchor boxes are very sensitive to deviations, which is not conducive to the training and learning of the model. Converting the regression of absolute scale to the regression of relative scale can solve this problem.

[0041] After the anchor frame is generated, only the dense sampling of the original image is completed. Further, it is necessary to construct a target for supervised learning for each sample, which specifically indicates the position of the target frame relative to the anchor frame and the category of each anchor frame. In other words, in order to determine whether the anchor frame belongs to a specific target type, it is also necessary to determine its specific position. The position here is represented by the offset of the anchor frame relative to the target frame. The offset here is divided into two parts, the offset of the center point of the target frame relative to the center point of the anchor frame and the conversion of the width and height of the target frame relative to the width and height of the anchor frame. The conversion here specifically indicates the scale ratio of the target frame and the anchor frame after logarithmic transformation.

[0042] Step 2: Perform end-to-end feature extraction and feature fusion on the input image to finally generate a feature map.

[0043] Step 3: These feature maps will be regressed through the target frame coordinate regression branch and the category confidence regression branch to regress the final coordinates and the probabilities of different target classifications. For this method, if the total number of anchor frames is represented by N, then the final output of the classification branch of the network model will be 2N, and the final output of the target frame coordinate regression branch will also be 2N, representing the probability that each anchor frame belongs to the two different categories of people and cars, as well as the offset of the target center point relative to the anchor frame and the logarithmic transformation value of the target width and height relative to the width and height of the anchor frame.

[0044] In another embodiment, the matching of the anchor box and the target box in step S101 further comprises the following step: transforming the width and height of the target box relative to the width and height of the anchor box into a logarithmic space.

[0045] For this embodiment, it is necessary to transform the width and height of the target box relative to the width and height of the anchor box into a logarithmic space. If the transformation is not performed, the output width and height of the model can only be positive values, which increases the requirements for the model and increases the difficulty of optimization. Transforming to a logarithmic space solves this problem.

[0046] In another embodiment, step S102 further includes the following steps:

[0047] S1021: The input image is passed through a backbone network consisting of stacked convolutional layers for feature extraction;

[0048] S1022: extracting the features of two layers in the middle of the backbone network and the features of the last layer and sending them to the adaptive feature fusion module for processing to obtain three adaptive feature maps of different levels;

[0049] S1023: Extract the features of the two middle layers of the backbone network and the features of the last layer and send them to the Attention feature fusion module for processing to obtain feature maps with attention at three different levels;

[0050] S1024: Concatenate and fuse the feature maps obtained in steps S1022 and S1023 to obtain the final feature map.

[0051] For this embodiment, the first step: in the feature extraction process of the entire network from input to output, the input image 3×640×640 first undergoes feature extraction through a backbone network composed of stacked convolutional layers, and the features of each intermediate layer of the network are extracted and sent to the subsequent FPN for processing. Here, the features of the last three layers of the entire backbone network are extracted in total, and the scales of the three feature maps are 256×80×80, 512×40×40, and 1024×20×20 respectively. After FPN feature fusion, 3 layers of features are obtained, and there will be a large number of prior anchor boxes on each layer. To improve the expression ability of the features, the feature maps at this time will also pass through two different modules respectively. One is an adaptive feature fusion module that fuses the three layers of features with different weights to improve the expression ability of the features; the other is an Attention attention feature fusion module that adds an attention mechanism to the features to enhance the receptive ability of the features.

[0052] In another embodiment, the adaptive feature fusion module is an improved adaptive fusion FPN.

[0053] For this embodiment, for the adaptive feature fusion module, in order to make full use of the semantic information of high-level features and the fine-grained information of low-level features, the architecture of FPN is often used for feature fusion. However, the FPN architecture often uses a direct concatenation and addition connection method, which cannot fully and adaptively utilize features of different scales. Therefore, an adaptive structure is added to the traditional FPN architecture. As Figure 3 shown, after adjusting the channels of the features X1, X2, X3 from different feature layers with different strides, they are sent to the adaptive feature fusion AFF module, that is, the feature layers after channel adjustment are multiplied by different weight coefficients a, b, c and added together to obtain the new adaptive weight fusion feature predict. The calculation formula is as follows:

[0054]

[0055] where a, b, c represent different weight coefficients, represents the features after adjustment of different feature layers.

[0056] Since the addition method is adopted, the feature sizes of the outputs of the three feature layers before addition need to be the same, and the number of channels also needs to be the same. It is necessary to upsample or downsample the features of different layers and adjust the number of channels. For the weight parameters a, b, and c, they are obtained by performing a 1×1 convolution on the resized feature map. And after the parameters a, b, and c are concatenated, they pass through the softmax so that their ranges are all within [0, 1] and their sum is 1.

[0057] In another embodiment, the Query in the Attention attention feature fusion module comes from the non-linear transformation of the shallow feature map, and both Key and Value come from the linear transformation of the deep feature map after upsampling.

[0058] For this embodiment, for the Attention attention feature fusion module, it is an Attention-FPN. The feature pyramid can effectively improve the localization ability of the algorithm for targets of different scales. For the small target detection task, because the distance and orientation of the objects being photographed relative to the camera in the actual scene are different, the size of the farthest target is only 16×8, which requires the target detection network to have good detection ability for small targets. The traditional FPN is realized by directly adding the upsampled high-level features and the low-level features. This method designs and realizes an improved FPN that integrates the Attention idea.

[0059] Query, Key, and Value no longer come from the same input. Query comes from the non-linear transformation of the shallow feature map, and both Key and Value come from the linear transformation of the deep feature map after upsampling. The operation using element-wise addition in the original FPN is changed to the fusion using the attention mechanism. From the perspective of the principle of the attention mechanism, this operation can be understood as expressing each pixel in the shallow feature map by weighted summing all the pixels in the deep feature map. The advantage of this is that using the deep attention mechanism to represent the shallow layer can effectively introduce global information for each pixel in the shallow feature map. While convolution focuses more on local information, so the fused feature map retains both global information and local information, which is more conducive to the learning of the model. Finally, after obtaining the new feature map of the shallow feature and the deep feature fused by the attention mechanism, the self-attention mechanism will be used again to further transform this feature map to improve the expression ability of the features.

[0060] Figure 4Shows the complete implementation process of Attention-FPN. The specific operations are as follows: First, upsample the deep feature map, use a 1x1 convolution to align the number of channels with the previous layer, and then, in order to perform the Attention operation on the obtained feature map, first slice the feature map, and perform self-attention operations on all pixels within each slice. The self-attention operation module is specifically as shown in Figure 4 shown on the right. The input is the Query query and the feature vector F. Two separate matrices Key and Value are extracted from F; calculate the attention scores between Key and the Query query, and finally obtain the weighted average according to Value. Then, the obtained weighted average undergoes an inverse transformation to obtain the same shape as the original input feature map, thus realizing one attention calculation process.

[0061] Send the feature maps into the adaptive feature fusion module and the Attention feature fusion module respectively. After feature fusion, use concat to refine the fused features, and finally obtain the fused feature map for the next regression prediction process.

[0062] In another embodiment, the loss function for the regression target box coordinates in step S103 is the Intersection over Union (IoU) loss optimized for small targets.

[0063] For this embodiment, in order to improve the localization accuracy, the loss function for the regression target box coordinates is replaced from the mean absolute error loss to the Intersection over Union (IoU) loss. When using the absolute error to measure the distance between the output and the target, the various geometric quantities obtained by regression are independent of each other and lack the inherent geometric constraints between them. However, if directly optimizing the Intersection over Union between the predicted box and the ground truth box, this geometric relationship can be modeled, which can also be regarded as a direct optimization of the evaluation metric. Since it is optimized for small target detection and the detection box itself is small, in the actual loss calculation, the accuracy of the width and height obtained by regression is not very important. More attention is paid to the accuracy of the x and y coordinate points. Therefore, the IoU Loss is modified to increase the proportion of the x and y coordinate losses and reduce the proportion of the width and height of the regression box to improve the network's attention to the localization accuracy.

[0064] In another embodiment, the enhancement processing in step S200 includes: infrared image denoising, Gamma correction, and super-resolution.

[0065] For this embodiment, the key point features in the enhanced micro-target image will be more prominent, which is beneficial to improving the accuracy of subsequent human key point localization and human action recognition.

[0066] Although the embodiments of the present invention have been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present invention, and all of these fall within the scope of protection of the present invention.

Claims

1. A method for detecting small targets in lightweight infrared images, comprising the following steps: S100: Obtain a small-target infrared image using a lightweight detection network based on YOLOIR. Among them, The YOLOIR-based lightweight detection network structure includes a backbone network, an adaptive feature fusion module, an Attention feature fusion module and a regression head prediction module; S200: performing enhancement processing on the generated small target infrared image; Wherein, step S100 further includes the following steps: S101: Generation of prior anchor boxes and matching of anchor boxes and target boxes; S102: Perform end-to-end feature extraction and feature fusion on the input image to finally generate a feature map; S103: The obtained feature map is regressed through the target frame coordinate regression branch and the category confidence regression branch to obtain the coordinates of the detected target in the current coordinate system and the maximum confidence of the category to which it belongs; Step S102 further includes the following steps: S1021: The input image is passed through a backbone network consisting of stacked convolutional layers for feature extraction; S1022: extract the features of two layers in the middle of the backbone network and the features of the last layer and send them to the adaptive feature fusion module for processing to obtain three adaptive feature maps of different levels; among them, the features X1, X2, and X3 from different feature layers are sent to the adaptive feature fusion AFF module after channel adjustment with different strides; S1023: extract the features of two layers in the middle of the backbone network and the features of the last layer and send them to the Attention feature fusion module for processing to obtain feature maps with attention at three different levels; among them, first upsample the feature map of the deep layer, use 1×1 convolution to align the number of channels with the channels of the previous layer, and then in order to use the obtained feature map for Attention operation, first slice the feature map, and perform self-attention operation on all pixels in each slice; S1024: Concat the feature maps obtained in steps S1022 and S1023 to obtain a final feature map.

2. According to the method of claim 1, the backbone network is a lightweight network ShuffleNet.

3. According to the method of claim 1, the detection head in the regression head prediction module is decoupled and a 1×1 convolution decoupling head is added.

4. According to the method of claim 1, the matching of the anchor frame and the target frame in step S101 further comprises the following step: transforming the width and height of the target frame relative to the width and height of the anchor frame into a logarithmic space.

5. According to the method described in claim 1, the query in the Attention feature fusion module comes from the nonlinear transformation of the shallow feature map, and the key and value both come from the linear transformation of the deep feature map after upsampling.

6. According to the method of claim 1, the loss function for regressing the target frame coordinates in step S103 is an intersection-over-union loss optimized for small targets.

7. The enhancement process in step S200 of the method according to claim 1 includes: Infrared image denoising, gamma correction and super-resolution.