Infrared image minimum target detection method based on semantic information aggregation

By building a semantic information aggregation module in the feature fusion stage of the infrared image extremely small object detection network, the missed detection problem in infrared image extremely small object detection is solved, the semantic information utilization rate is improved, and higher detection accuracy and lower missed detection rate are achieved.

CN120355903APending Publication Date: 2025-07-22ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510500704.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing infrared image minimal object detection method has the problem of high missed detection rate in complex backgrounds, and the existing methods lack attention and aggregation of semantic features of extremely small objects.

Method used

In the feature fusion stage of the existing infrared image minimal object detection network, a semantic information aggregation module is built, including a single feature-weighted semantic information aggregation module and a high and low-order feature-weighted semantic information aggregation module. The weighted fusion weights are generated through convolution and global pooling operations to aggregate the semantic features of the target and improve the utilization rate of semantic information.

Benefits of technology

It effectively reduces the missed detection rate of extremely small targets of infrared images, improves detection accuracy, and maintains a low error detection rate and inter-class confusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355903A_ABST
    Figure CN120355903A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared image minimum target detection method based on semantic information aggregation, and aims to solve the problem that a current infrared image minimum target detection method is high in omission ratio in aerial images with complex backgrounds. According to the method, the problem of high minimum target omission ratio is analyzed and converted into the problem of insufficient semantic information attention from improvement of the semantic information utilization rate, optimization is carried out in combination with deep learning, and infrared image minimum target detection with low omission ratio is completed. The application of the method can effectively improve the performance of the existing infrared image minimum target detection method based on deep learning, so that the omission ratio of the target is effectively reduced in the infrared image minimum target detection application after the method is applied to other existing infrared image minimum target detection methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a method for detecting extremely small targets in infrared images based on semantic information aggregation. Background Art

[0002] The detection of targets in infrared images captured by infrared imaging devices is widely used in fields such as marine rescue, military surveillance, and air defense. However, due to the large imaging field of view, infrared images often have complex backgrounds. Moreover, due to the long shooting distance, the targets in the images often appear extremely small. Therefore, the high-precision detection of extremely small targets in infrared images remains a challenge.

[0003] Due to the complex content of infrared images, traditional methods are difficult to apply. Previous work to address this challenge has mainly been based on deep learning methods. Conventional methods generally directly apply existing advanced detection methods such as Faster R-CNN, SSD, and YOLO to detect extremely small targets in infrared images. However, since conventional detectors are mainly aimed at general-sized targets in natural images, directly using existing popular detectors cannot produce satisfactory performance. Therefore, current research mainly focuses on existing popular detectors and improves methods according to the characteristics of extremely small targets with weak features to improve the accuracy of existing popular detectors in the task of detecting extremely small targets in infrared images. The latest research, such as RDIAN published in "IEEE Transactions on Geoscience and Remote Sensing", designed multiple different modules to enhance the feature expression of extremely small targets from different aspects such as receptive fields to address the problem of weak feature representation of extremely small targets in infrared images on the feature map, and achieved advanced performance in the detection of extremely small targets. Similarly, UIU-Net, DNA-Net, and ALCNet published in "IEEE Transactions on Geoscience and Remote Sensing" and others all improved the intensity of extremely small target features from the aspect of feature expression. These well-optimized methods have achieved satisfactory performance, especially excellent performance in reducing the false detection rate by weakening extremely small targets.

[0004] However, the problem of missed detection of extremely small targets remains the main problem in the task of detecting extremely small targets in infrared images. At the same time, current work still lacks exploration of the problem of missed detection of extremely small targets. Summary of the Invention

[0005] Considering the problem of missed detection in current infrared image tiny target detection methods, this method constructs a semantic information aggregation module using deep learning. Considering that the feature fusion module in existing object detection models lacks attention to semantic information, this method designs a semantic information aggregation module in the feature fusion module of the existing method, and uses this module to focus on the aggregation of semantic features of tiny targets, thereby improving the utilization rate of semantic features of tiny targets. Compared with existing methods, this method can effectively reduce the missed detection rate of tiny targets, thereby further improving the detection accuracy of tiny targets in infrared images.

[0006] The present invention is implemented by the following technical solutions:

[0007] An infrared image tiny target detection method based on semantic information aggregation, which improves the attention to semantic information in feature fusion based on deep learning, and uses a semantic information aggregation module to aggregate the semantic features of the target during the feature fusion process. Specifically, deep learning is used to extract semantic features, so that the infrared image tiny target detection method mainly fuses the semantic features in the features during the process of fusing high-level and low-level features, thereby reducing the missed detection rate of tiny targets.

[0008] In the high-level and low-level feature fusion stage of the network (such as YOLOv5, YOLOv11, and Faster R-CNN, etc.) that can be applied to infrared image tiny target detection, the semantic information aggregation module is constructed. The semantic information aggregation module includes a single feature weighted semantic information aggregation module and a high-level and low-level feature weighted semantic information aggregation module. Using the module proposed by this method for training enables the network to focus on fusing the semantic features in the features during the training process, thereby improving the utilization rate of semantic information and reducing the missed detection rate of infrared image tiny target detection.

[0009] The semantic information aggregation module improves the attention to semantic information in feature fusion based on deep learning, so as to aggregate the semantic features of the target during the feature fusion process. The specific method is as follows:

[0010] The input of the semantic information aggregation module is the features extracted by the backbone networks of N arbitrary networks (such as YOLOv5, YOLOv11, and Faster R-CNN) applied to infrared image tiny target detection where c is the number of channels of the feature map, h and w represent the height and width of the feature map respectively, and n ∈ {1, 2, 3,..., N}.

[0011] The single feature weighted semantic information aggregation module processes each feature F input to the semantic information aggregation module n , and uses convolution and global pooling operations to generate the weighted fusion weight corresponding to each channel according to the different semantic features contained in each channel of the feature And use this weight to fuse the semantic features of each channel The semantic information in to generate a single feature semantic aggregation weight matrix

[0012] ω Ci = GAP(Conv(F n ))

[0013]

[0014] where Conv(·) represents the convolutional layer, GAP(·) represents the global average pooling layer, σ(·) represents the Sigmoid function, and ω Ci represents the weighted fusion weight of channel i; represents the semantic feature of channel i.

[0015] After generating the single feature semantic aggregation weight matrix ω n the single feature weighted semantic information aggregation module uses each feature F input to the semantic information aggregation module n , and uses convolutional and global pooling operations (the parameters in the operation process are different from those in the above process of generating ω Ci ) to generate the global semantic information aggregation weight of each feature

[0016] After generating the single feature semantic aggregation weight matrix ω n and the global semantic information aggregation weight λ n , the single feature weighted semantic information aggregation module uses ω n and λ n to generate the single feature semantic aggregation weight W n . This process can be expressed as

[0017] W n = ω n ⊙ λ n .

[0018] The high and low order feature weighted semantic information aggregation module uses each input feature F n and the single feature semantic aggregation weight W generated by the single feature weighted semantic information aggregation module n , and performs weighted fusion to generate the high and low order semantic information aggregation feature F. This process can be expressed as

[0019]

[0020] Among them, Concat(·, ·,..., ·) represents concatenating features along the channel direction. In this way, by aggregating semantic information, the utilization rate of semantic information of extremely small targets by the detection network is improved, thereby reducing the missed detection rate of extremely small targets.

[0021] Embed the above semantic information aggregation module into the feature fusion stage of the existing extremely small target detection network, replace the original feature concatenation fusion of the existing network, and train the network until convergence. Based on this trained network, the missed detection problem of extremely small targets can be solved.

[0022] The inventive principle of the present invention is as follows:

[0023] Inspired by the fact that the position information of extremely small targets on the deep feature map is manifested as high-light semantic features, the present invention analyzes the phenomenon of feature performance when extremely small targets are missed, and finds that the missed detection of extremely small targets results from insufficient utilization of semantic features on the feature map. Specifically, although the existing work enhances the expression of extremely small target features, it lacks attention and aggregation to the semantic features of extremely small targets in the feature fusion stage. This makes the semantic features that originally contain the position information of extremely small targets underutilized in the feature fusion stage, resulting in relatively low confidence levels of some extremely small targets, thus causing the missed detection of these extremely small targets with relatively low confidence levels under a specific confidence threshold. Therefore, the key to solving this problem lies in paying attention to the aggregation of extremely small target semantic features in the feature fusion stage, thereby improving the confidence level of extremely small targets. For this reason, the present invention designs a new semantic information aggregation module to reduce the missed detection rate of existing infrared image extremely small target detection methods.

[0024] The beneficial effects of the present invention are as follows:

[0025] This method takes into account the missed detection problem of the current infrared image extremely small target detection method. Starting from improving the utilization rate of semantic information, a semantic information aggregation module is constructed to strengthen the aggregation of extremely small target semantic features, thereby reducing the missed detection rate of extremely small targets and completing the detection of extremely small targets in infrared images. This method comprehensively extracts the target semantic information in the features from two aspects: single feature semantic aggregation and global semantic information aggregation. Among them, single feature semantic aggregation extracts the semantic information inside each feature input into the method of the present invention, and global semantic information aggregation synthesizes the semantic information extracted by each group of single feature semantic aggregations, so as to obtain sufficient semantic information from different input features. The application of this method can effectively improve the detection performance of the existing infrared image extremely small target detection method based on deep learning, enabling the existing infrared image extremely small target detection method to have excellent performance in reducing the false detection rate and weakening the inter-class confusion while also having a relatively low missed detection rate of extremely small targets. Description of the Drawings

[0026] Figure 1 is the structural diagram of the semantic information aggregation module of the method of the present invention.

[0027] Figure 2 It is the structural diagram of the single - feature weighted semantic information aggregation module of the method of the present invention.

[0028] Figure 3 It is the effect diagram of reducing the missed detection rate of infrared extremely small targets by improving the utilization rate of semantic information in the embodiment of the present invention. Detailed implementation manners

[0029] The following is further described in conjunction with specific embodiments and the accompanying drawings.

[0030] Embodiment

[0031] The present invention provides an infrared image extremely small target detection method based on semantic information aggregation. In the feature fusion stage of the existing method, a semantic information aggregation module is constructed to strengthen the aggregation of the semantic features of extremely small targets, thereby improving the utilization rate of the semantic features of extremely small targets. Starting from improving the utilization rate of semantic information, this method analyzes and transforms the problem of high missed detection rate of extremely small targets into the problem of insufficient attention to semantic information, and combines deep learning for optimization to complete the detection of extremely small targets in infrared images with low missed detection rate. Compared with the existing method, this method can effectively reduce the missed detection rate of extremely small targets, thereby further improving the detection accuracy of extremely small targets in infrared images. The specific steps are as follows:

[0032] In the high - and low - order feature fusion stage of the existing networks (such as YOLOv5, YOLOv11, and Faster R - CNN, etc.) that can be applied to the detection of extremely small targets in infrared images, a semantic information aggregation module is constructed. This module includes a single - feature weighted semantic information aggregation module and a high - and low - order feature weighted semantic information aggregation module. As shown in the attached Figure 1 figure, the input of this semantic information aggregation module is the features extracted by the backbone networks of N arbitrary networks (such as YOLOv5, YOLOv11, and Faster R - CNN) applied to the detection of extremely small targets in infrared images where c is the number of channels of the feature map, h and w respectively represent the height and width of the feature map, and n ∈ {1, 2, 3,... N}.

[0033] As shown in the attached Figure 2 figure, the single - feature weighted semantic information aggregation module is used to first generate the weighted fusion weight for each channel of each feature F input to the semantic information aggregation module n using convolution and global pooling operations according to the different semantic information contained in each channel of the feature and use this weight to fuse the features of each channel in the semantic information to generate a single - feature semantic aggregation weight matrix This process can be expressed as

[0034] ω ci = GAP(Conv(F n ))

[0035]

[0036] where Conv(·) represents the convolutional layer, GAP(·) represents the global average pooling layer, σ(·) represents the Sigmoid function, and ω Ci represents the weighted fusion weight of channel i; represents the semantic feature of channel i.

[0037] After generating the single-feature semantic aggregation weight matrix ω n the single-feature weighted semantic information aggregation module uses each input feature F n to generate the global semantic information aggregation weight of each feature through convolution and global pooling operations

[0038] After generating the single-feature semantic aggregation weight matrix ω n and the subsequent global semantic information aggregation weight λ n the single-feature weighted semantic information aggregation module uses ω n and λ n to generate the single-feature semantic aggregation weight W n . This process can be expressed as

[0039] W n = ω n ⊙ λ n .

[0040] Then, the high- and low-order feature weighted semantic information aggregation module uses each input feature F n and its single-feature semantic aggregation weight W n generated by the single-feature weighted semantic information aggregation module for weighted fusion to generate the high- and low-order semantic information aggregation feature F. This process can be expressed as

[0041]

[0042] where Concat(·, ·,..., ·) represents concatenating features along the channel dimension. Finally, the semantic information aggregation module is embedded into the feature fusion stage of the existing detection network to replace the feature concatenation layer of the existing network for network training.

[0043] Figure 3Taking YOLOv5m as an example, the detection effect of the basic network YOLOv5m on infrared small targets is shown, as well as the detection effect of the enhanced YOLOv5m on infrared small targets after applying the method of the present invention. It can be seen that YOLOv5m after applying the method of the present invention detects more infrared small targets and reduces the missed detection rate of infrared small targets. At the same time, Figure 3 The superiority of the method of the present invention is numerically shown using the average recall rate. The average recall rate represents the percentage of detected targets among the total number of potential targets. The higher the average recall rate, the lower the missed detection rate of the targets. From Figure 3 It can be seen that after applying this method, the average recall rate of the targets increases by 2%, indicating the effectiveness of this method in reducing the missed detection rate of the targets.

Claims

1. An infrared image extremely small target detection method based on semantic information aggregation, characterized in that, Embed a semantic information aggregation module in the feature fusion stage of the existing extremely small target detection network; the semantic information aggregation module includes a single feature weighted semantic information aggregation module and a high-low order feature weighted semantic information aggregation module; The semantic information aggregation module improves the attention to semantic information in feature fusion based on deep learning, so as to aggregate the semantic features of the target in the feature fusion process. The specific method is as follows: For each feature F input into the semantic information aggregation module n , the single-feature weighted semantic information aggregation module is used to extract the semantic features of each channel in feature F n , and based on the semantic features of each channel, the weighted fusion weight corresponding to each channel is generated. The semantic features of each channel are fused using the weighted fusion weight to generate a single-feature semantic aggregation weight matrix ω n ; The single feature weighted semantic information aggregation module firstly calculates the weighted semantic information based on the feature F n , use convolution and global pooling operations to obtain feature F n The global semantic information aggregation weight λ n , and then use ω n and λ n Get the single feature semantic aggregation weight W n ; The high- and low-order feature weighted semantic information aggregation module is based on the single feature weighted semantic information aggregation weight W n , and aggregates the semantic features in each input feature F n to obtain the high- and low-order semantic information aggregation feature F.

2. The method for detecting extremely small targets in infrared images based on semantic information aggregation according to claim 1, wherein, Train the extremely small target detection network embedded with the semantic information aggregation module, so that the network fuses the semantic features in the features during the training process, thereby improving the utilization rate of semantic information and reducing the missed detection rate of extremely small targets in infrared images.

3. The method for detecting extremely small targets in infrared images based on semantic information aggregation according to claim 1, characterized in that, The input of the semantic information aggregation module is the features extracted by the backbone networks of N arbitrary networks applied to the detection of extremely small targets in infrared images where c is the number of channels of the feature map, h and w respectively represent the height and width of the feature map, and n ∈ {1, 2, 3,... N}.

4. The method for detecting extremely small targets in infrared images based on semantic information aggregation according to claim 3, wherein, Generate the weighted fusion weight corresponding to each channel based on the semantic feature of each channel, and use the weighted fusion weight to fuse the semantic features of each channel to generate a single feature semantic aggregation weight matrix This process is expressed by the formula as follows: ω Ci = GAP(Conv(F n )) Among them, Conv(·) represents the convolutional layer, GAP(·) represents the global average pooling layer, σ(·) represents the Sigmoid function, and ω Ci represents the weighted fusion weight of channel i; represents the semantic feature of channel i.

5. The method for detecting extremely small targets in infrared images based on semantic information aggregation according to claim 1, characterized in that, Using ω b and λ b to obtain the single feature semantic aggregation weight W n , which is expressed by the formula as W n = ω n ⊙ λ n 。 6. The method for detecting extremely small targets in infrared images based on semantic information aggregation according to claim 1, wherein The high- and low-order feature weighted semantic information aggregation module is based on the single feature weighted semantic information aggregation weight W n , aggregates the semantic features in each input feature F n to obtain the high- and low-order semantic information aggregation feature F; this process is expressed by the formula as follows: Among them, Concat(·,·,...,·) means to splice features along the channel direction.