Infrared small target detection method based on semantic prompt
By introducing two-stage feature interaction between CLIP text encoder and CNN-Transformer fusion module, the problem of high false alarm rate in infrared small object detection is solved, and the semantic distinction of different categories of targets is achieved, and the accuracy and robustness of detection are improved.
Patent Information
- Application Number
- CN202510572193.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-15
AI Technical Summary
The existing infrared small object detection method is difficult to effectively distinguish different target semantics and shape characteristics of different categories in multi-objective scenarios, resulting in a high false alarm rate.
Multi-category semantic information is introduced based on CLIP text encoder, combined with CNN and Transformer fusion modules, infrared small object detection is realized through two-stage feature interaction and decoder, enhancing the semantic perception ability of the model.
It reduces the false alarm rate, improves the accuracy and robustness of infrared small target detection, and adapts to the detection needs of multiple targets.
Smart Images

Figure CN120495820A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of infrared small target detection, and in particular relates to an infrared small target detection method based on semantic cues. Background Art
[0002] Infrared imaging technology, which captures infrared radiation emitted by objects and converts it into images, has unique advantages. First, any object above absolute zero will emit infrared radiation, and the greater the temperature difference, the more significant the difference in its appearance in the infrared image. This makes infrared imaging widely used in scenarios related to temperature changes. Second, infrared imaging has passive detection capabilities. Compared with active imaging such as radar imaging, it can capture target information without emitting external signals, and has extremely high concealment, making it particularly suitable for military reconnaissance and nighttime surveillance. Finally, infrared imaging has the characteristics of all-weather operation and is not affected by lighting conditions and weather. It can even penetrate complex environments such as smoke and mist, ensuring reliable detection under adverse conditions. Therefore, infrared imaging is widely used in fire warning, industrial non-destructive testing, medical feature recognition, agricultural production and other fields.
[0003] To detect small infrared targets, researchers have proposed a number of traditional algorithms to address the detection challenges, including methods based on target features, background suppression, and image structural elements. Specifically, target feature-based methods exploit the characteristic differences between the target and its background in a single infrared image frame. For example, local contrast measurement (LCM) exploits the fact that small targets often have higher grayscale values than their surroundings, proposing a visual contrast mechanism. Robust local contrast measurement (RLCM) intentionally enhances infrared targets while suppressing the background. Background suppression methods focus on minimizing image background through filtering, such as spatial and transform domain filtering. For example, TopHat utilizes basic erosion and dilation operations from mathematical morphology to extract images containing residues and faint small targets. The maximum-mean method is a nonlinear filtering technique that processes the image through median filtering and then performs a difference calculation. Methods based on image data structure exploit different structural features in infrared images, such as the sparsity of the target and the low-rank nature of the background. For example, the IPI model reformulates small target detection as the task of separating two different components from the data matrix. Dai et al. improved the IPI model by using a structure tensor and reweighting method to combine local structure weights with sparse enhancement weights, replacing the global constant weight parameters. However, these methods usually rely heavily on prior knowledge, are sensitive to hyperparameters, and generally perform poorly in scenes involving complex backgrounds and images with diverse object sizes.
[0004] Recently, with the development of deep learning methods, numerous infrared target detection methods based on deep learning have been proposed, all of which have achieved impressive performance. Among these methods, ACM introduced an asymmetric context modulation mechanism to trade high-level semantics for low-level details. RepISDNet proposed a simple reparameterized network to balance detection accuracy and inference efficiency. MSAFFNet proposed a spatial pyramid pooling module (ASPPM) to focus on the global context of the target and a dual attention module (DAM) to focus on regions of interest in both shallow and deep feature layers to improve small target detection. SRNet proposed a method for infrared small target detection that learns shape-biased representations and uses a small number of 9×9 convolutions in the encoder to extract reliable shape knowledge from infrared images. DNANet introduced a densely nested attention network to address the problem of target disappearance in deep layers caused by pooling layers. ISNet models single-frame infrared small target (SIRST) detection as a shape detection task and proposes an edge block inspired by Taylor-finite difference (TFD) and two directional attention aggregation (TOAA) blocks to accurately capture the shape characteristics of infrared targets.
[0005] In general, existing deep learning methods for infrared small target detection typically treat this task as a binary segmentation problem, categorizing different targets, such as ships, aircraft, and drones, as "small infrared targets." Binary segmentation struggles to effectively learn the semantic and shape characteristics of diverse targets in tasks involving multiple targets. Consequently, target identification may lack semantic information and instead rely on pixel intensity or local contrast features. This can lead the model to misclassify small, non-target bright spots in the image as targets, increasing the false alarm rate.
[0006] After searching, no prior art documents identical or similar to the present invention were found. Summary of the Invention
[0007] To reduce the false alarm rate of the model, this paper proposes a text-assisted approach for infrared small target detection. This approach incorporates semantic information about the target into infrared small target detection. By semantically distinguishing between different target categories, the model's understanding of target category features is enhanced. Unlike other models that utilize text features and require text annotation for each image, this approach only requires prior knowledge of common infrared small target categories in the dataset, eliminating the need for tedious dataset annotation. Specifically, we use common infrared small targets such as "aircraft," "ship," and "drone" as input text and encode the input text using a pre-trained CLIP text encoder. This approach incorporates multi-category semantic information into the model to accommodate the detection needs of a variety of infrared small targets. Furthermore, to fully exploit key features in the image, we design a fusion extraction module that fuses a convolutional neural network (CNN) and a Transformer to jointly model global image context and local target details. Finally, to fully leverage multimodal features, we design a feature alignment mechanism for image and text. This module consists of two submodules: first, the extracted text features and image features are first passed through a multimodal cross-attention module to achieve interaction between features from different modalities. The interactive features between the modalities are then fused through a dynamic fusion module. This module adaptively adjusts the weights of text and image features based on the complexity of the input scene and the target semantic information contained, achieving efficient fusion of text and image. Finally, the fused features are passed through a decoder to obtain the final output.
[0008] The present invention solves the practical problem by adopting the following technical solutions:
[0009] 1. A method for detecting small infrared targets based on semantic cues, comprising the following steps:
[0010] Step 1: Extract image features through a 5-layer image encoder;
[0011] Step 2: Use the CLIP pre-trained Transformer text encoder to extract the corresponding text embedding features;
[0012] Step 3: Based on the features obtained in steps 1 and 2, the two-stage feature fusion module is used to fuse and interact with each other;
[0013] Step 4: Use the semantically enhanced features from step 3 to detect small infrared targets through the decoder.
[0014] 1. The method for detecting small infrared targets based on semantic cues according to claim 1, wherein the step 1 comprises extracting image features by using a 5-layer image encoder, comprising:
[0015] 1.1. Different convolutional blocks are used to obtain q, k, and v. For q and k, given an input I, we first perform two convolution operations to obtain I1 and I2, respectively. Then, I1 and I2 are reshaped using a reshape operation and then passed through two fully connected layers to obtain q and k, respectively. For v, we calculate it using three dilated convolutional layers.
[0016] The specific calculation formula for q, k, and v is:
[0017] q = Fc1(reshape(Conv1(I)))
[0018] k = Fc2(reshape(Conv2(I)))
[0019] v = DConv(DConv(DConv(I)))
[0020] Among them, () represents ordinary convolution, which means changing the shape of the feature, Fc() represents the fully connected layer, and () represents the dilated convolution.
[0021] 1.2. After obtaining q, k, and v, perform matrix multiplication on q and k, and then perform convolution and softmax operations to obtain the output attention matrix att. Finally, perform matrix multiplication on the attention matrix and the feature matrix v to obtain the output. The final calculation process is shown in the following formula:
[0022] att=softmax(Conv(q×k))
[0023] F a =v×att
[0024] 1.3. The extracted local features are fused with the global features and the final output is calculated through a feed-forward layer. The feed-forward layer consists of a normal convolution layer and a point-by-point convolution layer to ensure that the final output can fully combine local and global features.
[0025] 2. The method for detecting small infrared targets based on semantic cues according to claim 1, wherein step 2 uses a Transformer text encoder pre-trained with CLIP to extract corresponding text embedding features.
[0026] CLIP is a cross-modal model that can understand and correlate image and text information. Key to its cross-modality is its use of cross-modal contrastive learning. Its core concept is to maximize the similarity between semantically related modalities while minimizing the similarity between semantically unrelated modalities through contrastive learning of positive and negative sample pairs. Specifically, in the CLIP model, positive pairs consist of images and text with identical content, while negative pairs consist of unrelated images and text. Through this contrastive learning, the model is able to better learn the deep semantic connections between images and text.
[0027] 3. The method for detecting small infrared targets based on semantic cues according to claim 1 is characterized in that, in step 3, in the task of detecting small infrared targets, traditional methods usually rely on the pixel intensity or local contrast features of the target. Although this method performs well in simple scenarios, it is difficult to fully understand the semantic characteristics of targets of different categories, resulting in a significantly increased false alarm rate in the detection results. In order to allow the image to obtain semantic features, a two-stage feature interaction module is used. By introducing text semantic embedding into deep image features containing motion semantic information, the model can better distinguish between targets and backgrounds in the feature space. The specific steps include:
[0028] 4.1 Based on text and image modality features F t and F i , first generate the corresponding query, key and value, which can be expressed as:
[0029]
[0030] in, It means that the original modal features are only transformed to a certain extent through the convolution or fully connected layer.
[0031] 4.2 Exchange the queries of the two modalities to achieve spatial interaction. The formula is:
[0032]
[0033] Among them, d k is the scaling factor and softmax is the activation function.
[0034] 4.3 Perceive image information and generate dynamic weights for image and text fusion through the weight generation module. The weight generation module consists of two branches: one branch uses global average pooling (GAP) to obtain global context information to emphasize global image information, and the other branch maintains the original feature size to obtain local context information, thereby preventing the target information from being ignored. The specific implementation process of the weight generation module is as follows:
[0035]
[0036] w=σ(w1+w2)
[0037] 4.4 After the weights are generated, the semantic features and fusion features are weighted and fused according to the dynamic weights. The specific implementation process is as follows:
[0038]
[0039] 5. The method for detecting infrared small targets based on semantic cues according to claim 1, wherein step 4 uses the image features after semantic enhancement in step 3 to detect infrared small targets through a decoder.
[0040] Advantages and beneficial effects of the present invention:
[0041] 1. This paper proposes a text-assisted infrared small target detection framework. By introducing the target category information through the CLIP text encoder, the semantic perception ability of the model is enhanced, so that the model no longer tends to judge some small bright spots of non-target classes in the image as targets, thereby reducing the false alarm rate of the model.
[0042] 2. In order to make good use of the local features of the target and the global features of the entire image, this paper proposes a CNN and Transformer fusion module for image feature extraction to improve the performance of infrared small target detection
[0043] 3. The present invention proposes a two-stage multimodal feature interaction module, which integrates the semantic information provided by the text into the extracted image features, thereby enhancing the expressive power of the image features. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 Schematic diagram of the infrared small target detection method based on semantic cues of the present invention;
[0045] Figure 2 It is a schematic diagram of the CNN and Transformer fusion module of the present invention;
[0046] Figure 3 Schematic diagram of the multimodal cross attention module of the present invention;
[0047] Figure 4 It is a schematic diagram of the dynamic semantic fusion module of the present invention. DETAILED DESCRIPTION
[0048] The embodiments of the present invention are further described below in conjunction with the accompanying drawings:
[0049] 1. A method for detecting small infrared targets based on semantic cues, comprising the following steps:
[0050] Step 1: Extract image features through a 5-layer image encoder;
[0051] Step 2: Use the CLIP pre-trained Transformer text encoder to extract the corresponding text embedding features;
[0052] Step 3: Based on the features obtained in steps 1 and 2, the two-stage feature fusion module is used to fuse and interact with each other;
[0053] Step 4: Use the semantically enhanced features from step 3 to detect small infrared targets through the decoder.
[0054] 2. The method for detecting small infrared targets based on semantic cues according to claim 1, wherein the step 1 comprises extracting image features by using a 5-layer image encoder, comprising:
[0055] 1.1 uses different convolutional blocks to obtain q, k, and v. For q and k, given an input I, two convolution operations are first performed to obtain I1 and I2, respectively. Then, I1 and I2 are reshaped using a reshape operation and then passed through two fully connected layers to obtain q and k, respectively. For v, we calculate it using three dilated convolutional layers.
[0056] The specific calculation formula for q, k, and v is:
[0057] q = Fc1(reshape(Conv1(I)))
[0058] k = Fc2(reshape(Conv2(I)))
[0059] v = DConv(DConv(DConv(I)))
[0060] Among them, () represents ordinary convolution, which means changing the shape of the feature, Fc() represents the fully connected layer, and () represents the dilated convolution.
[0061] 1.2 After obtaining q, k, and v, perform matrix multiplication on q and k, and then perform convolution and softmax operations to obtain the output attention matrix att. Finally, perform matrix multiplication on the attention matrix and the feature matrix v to obtain the output. The final calculation process is shown in the following formula:
[0062] att=softmax(Conv(q×k))
[0063] F a =v×att
[0064] 1.3 The extracted local features are fused with the global features and the final output is calculated through a feed-forward layer. The feed-forward layer consists of a normal convolution layer and a point-by-point convolution layer to ensure that the final output can fully combine local and global features.
[0065] 3. The method for infrared small target detection based on semantic cues according to claim 1, wherein step 2 uses a Transformer text encoder pre-trained with CLIP to extract corresponding text embedding features.
[0066] CLIP is a cross-modal model that can understand and correlate image and text information. Key to its cross-modality is its use of cross-modal contrastive learning. Its core concept is to maximize the similarity between semantically related modalities while minimizing the similarity between semantically unrelated modalities through contrastive learning of positive and negative sample pairs. Specifically, in the CLIP model, positive pairs consist of images and text with identical content, while negative pairs consist of unrelated images and text. Through this contrastive learning, the model is able to better learn the deep semantic connections between images and text.
[0067] 4. The method for detecting small infrared targets based on semantic cues according to claim 1, characterized in that, in step 3, in the task of detecting small infrared targets, traditional methods usually rely on the pixel intensity or local contrast features of the target. Although this method performs well in simple scenarios, it is difficult to fully understand the semantic characteristics of targets of different categories, resulting in a significantly increased false alarm rate in the detection results. In order to allow the image to obtain semantic features, a two-stage feature interaction module is used. By introducing text semantic embedding into deep image features containing motion semantic information, the model can better distinguish between targets and backgrounds in the feature space. The specific steps include:
[0068] 4.1 Based on text and image modality features F t and F i , first generate the corresponding query, key and value, which can be expressed as:
[0069]
[0070]
[0071] in, It means that the original modal features are only transformed to a certain extent through the convolution or fully connected layer.
[0072] 4.2 Exchange the queries of the two modalities to achieve spatial interaction. The formula is:
[0073]
[0074] Among them, d k is the scaling factor and softmax is the activation function.
[0075] 4.3 Perceive image information and generate dynamic weights for image and text fusion through the weight generation module. The weight generation module consists of two branches: one branch uses global average pooling (GAP) to obtain global context information to emphasize global image information, and the other branch maintains the original feature size to obtain local context information, thereby preventing the target information from being ignored. The specific implementation process of the weight generation module is as follows:
[0076]
[0077] w=σ(w1+w2)
[0078] 4.4 After the weights are generated, the semantic features and fusion features are weighted and fused according to the dynamic weights. The specific implementation process is as follows:
[0079]
[0080] 5. The method for detecting infrared small targets based on semantic cues according to claim 1, wherein step 4 uses the image features after semantic enhancement in step 3 to detect infrared small targets through a decoder.
[0081] In this embodiment, the overall schematic diagram of the infrared small target detection method based on semantic cues is as follows: Figure 1 As shown, the main It consists of four key components: a text encoder, an image encoder, an image-text interaction module, a feature Levy decoder.
[0082] The CNN and Transformer fusion module constructed by the present invention is as follows Figure 2 As shown, this module redesigns the self-attention The force architecture is mainly composed of convolutional layers, dilated convolutional layers and fully connected layers.
[0083] In addition, the semantic space alignment module such as Figure 3 As shown, its main purpose is to fully integrate characteristics in all dimensions. Cross attention is used to realize the interaction of different modal features.
[0084] Figure 4 The figure shows a schematic diagram of cross-scale feature fusion, which combines semantic features and details according to the characteristics of the image. Features, adaptively adjust the weights of text features and image features according to the complexity of the input scene and the target semantic information contained It emphasizes the characteristics of the target area, suppresses background interference, and realizes adaptive optimization of the feature space.
[0085] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0086] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A method for detecting small infrared targets based on semantic cues, characterized by: The following steps are involved: Step 1: Extract image features through a 5-layer image encoder; Step 2: Use the CLIP pre-trained Transformer text encoder to extract the corresponding text embedding features; Step 3: Based on the features obtained in steps 1 and 2, the two-stage feature fusion module is used to fuse and interact with each other; Step 4: Use the semantically enhanced features from step 3 to detect small infrared targets through the decoder.
2. The method for detecting small infrared targets based on semantic cues according to claim 1, characterized in that: In step 1, image features are extracted through a 5-layer image encoder, including: 1.
1. Different convolutional blocks are used to obtain q, k, and v. For q and k, given an input I, we first perform two convolution operations to obtain I1 and I2, respectively. Then, I1 and I2 are reshaped using a reshape operation and then passed through two fully connected layers to obtain q and k, respectively. For v, we calculate it using three dilated convolutional layers. The specific calculation formula for q, k, and v is: q = Fc1(reshape(Conv1(I))) k = Fc2(reshape(Conv2(I))) v = DConv(DConv(DConv(I))) Among them, () represents ordinary convolution, which means changing the shape of the feature, Fc() represents the fully connected layer, and () represents the dilated convolution. 1.
2. After obtaining q, k, and v, perform matrix multiplication on q and k, and then perform convolution and softmax operations to obtain the output attention matrix att. Finally, perform matrix multiplication on the attention matrix and the feature matrix v to obtain the output. The final calculation process is shown in the following formula: att=softmax(Conv(q×k)) F a =v×att 1.
3. The extracted local features are fused with the global features and the final output is calculated through a feed-forward layer. The feed-forward layer consists of a normal convolution layer and a point-by-point convolution layer to ensure that the final output can fully combine local and global features.
3. The method for detecting small infrared targets based on semantic cues according to claim 1, characterized in that: Step 2 uses the CLIP pre-trained Transformer text encoder to extract corresponding text embedding features. CLIP is a cross-modal model that can understand and correlate image and text information. Key to its cross-modality is its use of cross-modal contrastive learning. Its core concept is to maximize the similarity between semantically related modalities while minimizing the similarity between semantically unrelated modalities through contrastive learning of positive and negative sample pairs. Specifically, in the CLIP model, positive pairs consist of images and text with identical content, while negative pairs consist of unrelated images and text. Through this contrastive learning, the model is able to better learn the deep semantic connections between images and text.
4. The method for detecting small infrared targets based on semantic cues according to claim 1, wherein: In the infrared small target detection task, traditional methods in step 3 usually rely on the pixel intensity or local contrast features of the target. Although this method performs well in simple scenarios, it is difficult to fully understand the semantic characteristics of different categories of targets, resulting in a significantly increased false alarm rate in the detection results. In order to allow the image to obtain semantic features, the two-stage feature interaction module introduces text semantic embedding into deep image features containing motion semantic information, enabling the model to better distinguish between targets and backgrounds in the feature space. The specific steps include: 4.1 Based on text and image modality features F t and F i , first generate the corresponding query, key and value, which can be expressed as: in, It means that the original modal features are only transformed to a certain extent through the convolution or fully connected layer. 4.2 Exchange the queries of the two modalities to achieve spatial interaction. The formula is: in, d k is the scaling factor and softmax is the activation function. 4.3 Perceive image information and generate dynamic weights for image and text fusion through the weight generation module. The weight generation module consists of two branches: one branch uses global average pooling (GAP) to obtain global context information to emphasize global image information, and the other branch maintains the original feature size to obtain local context information, thereby preventing the target information from being ignored. The specific implementation process of the weight generation module is as follows: w=σ(w1+w2) 4.4 After the weights are generated, the semantic features and fusion features are weighted and fused according to the dynamic weights. The specific implementation process is as follows:
5. The method for detecting small infrared targets based on semantic cues according to claim 1, characterized in that: The step 4 uses the image features after semantic enhancement in step 3 to detect small infrared targets through a decoder.
Citation Information
Cited By
Infrared weak and small target detection method based on double-granularity semantic prompt
CN121459031A
Infrared small target detection method based on double alignment and transparency graph optimization
CN122156816A