Infrared small target identification method based on color and shape prompt learning

By introducing color shape prompt learning network ColorPeel and dynamic detection head dynamic head in infrared small object detection, combined with pyramid network FPN and path aggregation network PAN, the accuracy and robustness problems of infrared small object detection in the current technology in complex backgrounds are solved, and a more efficient infrared small object recognition effect is achieved.

CN119942094AActive Publication Date: 2025-05-06NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510423335.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing infrared small object detection methods are difficult to effectively utilize color information under complex backgrounds, resulting in reduced detection accuracy and poor model robustness, which is susceptible to noise and background clutter.

Method used

The infrared small object recognition method based on color shape prompt learning is adopted. The color shape prompt learning network ColorPeel enhances multi-scale features, combines the pyramid network FPN and the path aggregation network PAN to fusion of spatial semantics and significant color shape features, and uses dynamic detection head Dynamic Head and adaptive threshold focus loss ATFL for detection and optimization.

Benefits of technology

It improves the detection accuracy and robustness of infrared small targets in complex backgrounds, and enhances the recognition ability of small targets, especially when there is less data or low signal-to-noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942094A_ABST
    Figure CN119942094A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared small target identification method based on color shape prompt learning. The method comprises the steps of S1, dividing an image data set of an infrared small target into a training set and a test set; s2, inputting the images of the training set into a backbone network of an infrared small target detection network model, and extracting multi-scale features containing the infrared small target images; s3, multi-scale feature enhancement is carried out on the extracted multi-scale features of the infrared small target image; s4, performing different-domain fusion through a dual probability alignment (DPA) fusion framework to obtain multi-scale information features; s5, the multi-scale information features are sent to a dynamic detection head DynamicHead for detection; the detection result is evaluated by the normalized Wasserstein distance NWD and the adaptive threshold focus loss ATFL; and S6, obtaining a final infrared small target detection result. The method has good reliability and robustness, and the detection performance of the infrared small target is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of infrared small target detection, and in particular to an infrared small target recognition method based on color shape cue learning. Background Art

[0002] Infrared images refer to thermal radiation images captured by infrared sensors. Unlike visible light images, infrared images can provide thermal information of targets and are suitable for target detection in low light, complex backgrounds or harsh weather conditions. The common single-frame infrared small target detection method in the early days was a model-based method. It can be summarized as a filter-based method, a human visual system-based method, and a low-rank-based method. Among them, the filter-based method is only applicable to single and uniform scenes. The method based on the human visual system is only fully applicable to situations where the brightness of the object is relatively large and there is a more obvious difference from the surrounding background. Low-rank-based methods, including local low-rank and non-local low-rank, are applicable to almost all types of complex and rapidly changing backgrounds, but in practice, they require GPUs and other acceleration to meet real-time requirements. Model-driven methods are easily affected by clutter and noise, which reduces the robustness of the detection model. In complex backgrounds, it usually fails to find acceptable templates or learn local contrast information of objects, or object modeling is severely affected by model hyperparameters, resulting in poor generalization performance.

[0003] In recent years, the surge in data-driven machine learning methods, especially deep learning methods, has rapidly made it the most widely used method for detecting small infrared objects. Due to the huge imbalance between the target and the background, bounding box regression is extremely sensitive to small infrared targets, and the target information is easily lost at the high-level semantic layer. Existing small target detection methods only focus on the brightness and shape of the target without considering color information. Traditional detection models can only extract target features from the grayscale information of infrared images, and lack effective use of different spectral channels (multi-spectral or pseudo-color infrared). In addition, traditional methods are prone to blurred target boundaries, resulting in reduced detection accuracy. Based on this, this application proposes an infrared small target recognition method based on color shape cue learning. Summary of the invention

[0004] To solve the above problems, the present application discloses an infrared small target recognition method based on color shape cue learning, which can detect infrared small targets more effectively, has good reliability and robustness, and improves the detection performance of infrared small targets.

[0005] The infrared small target recognition method based on color shape cue learning specifically includes the following steps: Step S1, dividing the image data set of small infrared targets into a training set and a test set; Step S2: The images of the training set Input into the backbone network of the infrared small target detection network model to extract multi-scale features of the infrared small target image , , and Respectively represent the height and width of a single-channel image; multi-scale features , The resolutions are ; The images of the training set are single-channel images, Indicated in The factor by which the resolution of the corresponding layer is reduced; Step S3, the multi-scale features are enhanced through the pyramid network FPN, the path aggregation network PAN and the color shape prompt learning network ColorPeel; wherein the multi-scale features enhanced through the pyramid network FPN and the path aggregation network PAN belong to the spatial semantic feature domain, and the multi-scale features enhanced through the color shape prompt learning network ColorPeel belong to the salient color shape feature domain; Step S4, using the dual probability alignment (DPA) fusion framework to perform heterogeneous fusion of the spatial semantic feature domain and the salient color and shape feature domain to obtain multi-scale information features; Step S5, the fused multi-scale information features are sent to the dynamic detection head Dynamic Head for detection; the detection results are evaluated by the normalized Wasserstein distance NWD and the adaptive threshold focus loss ATFL; the normalized Wasserstein distance NWD (Normalized Gaussian Wasserstein Distance) and the adaptive threshold focus loss ATFL calculate the loss and guide the optimization process of the infrared small target detection network model; Step S6: Finally, select the test set image as the input of the trained infrared small target detection network model to obtain the final infrared small target detection result.

[0006] Furthermore, in step S3, the color shape hint learning network ColorPeel is used for color and shape enhancement of multi-scale features, specifically including: a: Given the RGB or color coordinates provided by the user, generate instance images and templates; b: The color shape hint learning network ColorPeel generates a color embedding vector c* and a shape embedding vector s*.

[0007] Furthermore, in step S5, the dynamic detection head Dynamic Head is used to detect and process the fused multi-scale information features, which specifically includes the following steps: Step S51: Given a feature tensor ,in represents all the images in the training set, represents the number of feature pyramid layers, Indicates the size of the feature, = × ; , Respectively represent the height and width of a single-channel image, Indicates the number of channels; Step S52: Perceiving Attention through Scale , spatial perception and attention and task-aware attention Dynamically fuse the features to obtain , the formula is: ; Step S53: Scale-aware attention Dynamic feature fusion is performed based on the multi-scale information features obtained in step S4, and the formula is: ;in It is a 1×1 convolutional layer; is the hard-sigmoid function; spatial perception attention Use deformable convolution to fuse multi-scale information features of different levels in the same spatial position. The formula is: ;in represents the number of sparse sampling locations, represents the number of feature pyramid layers, Indicated in The position in the layer The eigenvalue at is the initial sampling position, is the spatial offset obtained by learning the input features, is in position The scalar is obtained by self-learning and used to adjust the feature weight of the sampling position; task-aware attention Dynamically switch ON and OFF channels to support different tasks. The formula is: ;in is a hyperfunction that controls the threshold, where , Used to perform weighted operations on feature channels. , Used to adjust the bias of feature channels; Through average pooling The dimension is reduced, and then two fully connected layers and a normalization layer are used, and finally normalized by the sigmoid activation function.

[0008] Furthermore, in step S5, the adaptive threshold focus loss ATFL sets a threshold and regards values ​​below the threshold as hard samples, wherein the adaptive threshold focus loss ATFL expression is: ;in represents the current average predicted probability value, Represents the predicted value of the next epoch, is a hyperparameter.

[0009] Furthermore, the formula for the normalized Wasserstein distance NWD in step S5 is: ; where D is a constant related to the data set, , are the Gaussian distributions modeled by bounding box A and bounding box B, respectively. yes and The Wasserstein distance between them.

[0010] Beneficial effects of this application: 1. This application introduces the color shape cue learning network ColorPeel into the infrared small target detection model. The color shape cue learning network ColorPeel helps to enhance the contrast of infrared small targets by separating different colors or spectral information, making them easier to distinguish in complex backgrounds. Traditional infrared small target detection methods may be interfered by noise or background clutter, while the color shape cue learning network ColorPeel can effectively reduce background interference and improve detection accuracy through specific color layer peeling technology. The color shape cue learning network ColorPeel technology can more clearly retain the edge features of the target, reduce the target morphological distortion caused by filtering or background suppression, and improve the integrity of the target. In a variety of complex backgrounds (such as clouds, ocean ripples, ground textures, etc.), the color shape cue learning network ColorPeel improves the adaptability and generalization ability of the detection method through color feature enhancement and hierarchical peeling. Combined with deep learning methods, as a feature enhancement module, it improves the network's recognition ability of small targets, especially when there is less data or low signal-to-noise ratio.

[0011] 2. This application proposes an adaptive threshold focus loss ATFL, and adaptively adjusts the loss value according to the predicted probability value. By increasing the target loss weight and reducing the impact of background loss, the model can be more focused on extracting the features of small infrared targets and improving the detection effect, so as to improve the detection performance of small infrared targets.

[0012] 3. This application proposes the Wasserstein distance between two Gaussian distributions. The traditional IoU (Intersection over Union) is extremely sensitive in small target detection. A small bounding box offset will cause a large IoU change, making the training unstable. NWD remodels the bounding box using the Wasserstein distance to make the regression error smoother and reduce the error amplification problem in small target detection. Since the Wasserstein distance is more robust to changes in target scale, NWD can better adapt to infrared targets of different sizes and improve the generalization ability of detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is an overall flow chart of an embodiment of the present application; Figure 2 This is a flowchart of the implementation of the color feature prompt learning network ColorPeel in the embodiment of the present application; Figure 3 This is a flowchart of the implementation of the dynamic detection head Dynamic Head in an embodiment of the present application; Figure 4 This is a comparison chart of the image detection results of different infrared small target data sets using the mainstream infrared small target detection model and the method of this application in the comparative test of the application embodiment. DETAILED DESCRIPTION

[0014] The present application is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present application and are not used to limit the scope of the present application. It should be noted that the words "front", "rear", "left", "right", "up" and "down" used in the following description refer to directions in the accompanying drawings, and the words "inside" and "outside" refer to directions toward or away from the geometric center of a specific component, respectively.

[0015] like Figure 1 , Figure 2 and Figure 3 As shown, the infrared small target recognition method based on color shape cue learning specifically includes the following steps: Step S1, dividing the image data set of small infrared targets into a training set and a test set; Step S2: The images of the training set Input into the backbone network of the infrared small target detection network model to extract multi-scale features of the infrared small target image , , and Respectively represent the height and width of a single-channel image; multi-scale features , The resolutions are ; The images of the training set are single-channel images, Indicated in The factor by which the resolution of the corresponding layer is reduced; Step S3, multi-scale features are enhanced through the pyramid network FPN, the path aggregation network PAN and the color shape prompt learning network ColorPeel; wherein the multi-scale features enhanced through the pyramid network FPN and the path aggregation network PAN belong to the spatial semantic feature domain, and the multi-scale features enhanced through the color shape prompt learning network ColorPeel belong to the salient color shape feature domain; in step S3, the color shape prompt learning network ColorPeel is used for color and shape enhancement of multi-scale features, specifically including: a: Given the RGB or color coordinates provided by the user, generate instance images and templates; b: The color shape hint learning network ColorPeel generates a color embedding vector c* and a shape embedding vector s*.

[0016] Step S4, using the dual probability alignment (DPA) fusion framework to perform heterogeneous fusion of the spatial semantic feature domain and the salient color and shape feature domain to obtain multi-scale information features; Step S5, the fused multi-scale information features are sent to the dynamic detection head Dynamic Head for detection; the detection results are evaluated by the normalized Wasserstein distance NWD and the adaptive threshold focus loss ATFL; the normalized Wasserstein distance NWD and the adaptive threshold focus loss ATFL calculate the loss and guide the optimization process of the infrared small target detection network model; Step S5 Dynamic Head Dynamic Head is used to detect and process the fused multi-scale information features, specifically including the following steps: Step S51: Given a feature tensor ,in represents all the images in the training set, represents the number of feature pyramid layers, Indicates the size of the feature, = × ; , Respectively represent the height and width of a single-channel image, Indicates the number of channels; Step S52: Perceiving Attention through Scale , spatial perception and attention and task-aware attention Dynamically fuse the features to obtain , the formula is: ; Step S53: Scale-aware attention Dynamic feature fusion is performed based on the multi-scale information features obtained in step S4, and the formula is: ;in It is a 1×1 convolutional layer; is the hard-sigmoid function; spatial perception attention Use deformable convolution to fuse multi-scale information features of different levels in the same spatial position. The formula is: ;in represents the number of sparse sampling locations, represents the number of feature pyramid layers, Indicated in The position in the layer The eigenvalue at is the initial sampling position, is the spatial offset obtained by learning the input features, is in position The scalar is obtained by self-learning and used to adjust the feature weight of the sampling position; task-aware attention Dynamically switch ON and OFF channels to support different tasks. The formula is: ;in is a hyperfunction that controls the threshold, where , Used to perform weighted operations on feature channels. , Used to adjust the bias of feature channels; Through average pooling The dimension is reduced, and then two fully connected layers and a normalization layer are used, and finally normalized by the sigmoid activation function.

[0017] In step S5, the adaptive threshold focus loss ATFL sets a threshold and regards values ​​below this threshold as hard samples, where the adaptive threshold focus loss ATFL expression is: ;in represents the current average predicted probability value, Represents the predicted value of the next epoch, is a hyperparameter.

[0018] In step S5, the bounding box is modeled as a two-dimensional Gaussian distribution, and the Wasserstein distance between the two Gaussian distributions is calculated; two-dimensional Gaussian distribution:

[0019] in , and Represent the random variable vector, mean vector and covariance matrix of Gaussian distribution respectively; when

[0020] Horizontal bounding box , using 𝒩 Modeled as a 2-D Gaussian distribution;

[0021] in,( , ), and Represents the center coordinates, width and height respectively; two 2-D Gaussian distributions )and The 2-D Wasserstein distance between is defined as:

[0022] in is the mean vector and The square of the Euclidean distance between the two distributions, which measures the difference in the locations of the centers of the two distributions. , They are probability distributions and The corresponding covariance matrix. Represents the rank of the matrix.

[0023] Simplified to:

[0024] in is the Frobenius norm, given by the bounding box and bounding box Modeling Gaussian distribution , The distance between them is simplified to , normalize the Wasserstein distance to the range of 0-1, and get the formula of normalized Wasserstein distance NWD as follows: ; where D is a constant related to the data set, , are the Gaussian distributions modeled by bounding box A and bounding box B, respectively. yes and The Wasserstein distance between them.

[0025] Step S6: Finally, select the test set image as the input of the trained infrared small target detection network model to obtain the final infrared small target detection result.

[0026] The experimental simulation content of the infrared small target recognition method based on color shape cue learning in this application is as follows: the experimental platform uses a 64-bit Ubuntu system with a system version of 20.04.4 and a GPU model of GeForce RTX 2080Ti. Python is used as the programming language with a version of Python 3.9, and PyCharm is used as the software development platform. The model is implemented using the deep learning framework Pytorch 2.0, and the input image is a 3-channel RGB image with a size of 512x512. In the training phase, the batch size (batch size) is 6, the number of iterations (epochs) is 500, and the parameters are optimized using the Adam optimizer. The initial learning rate is set to 1e-3. With the acceleration of a GeForce RTX 2080 Ti GPU, the entire network training time is approximately 12 hours.

[0027] Figure 4 The image detection results of two mainstream infrared small target detection models and the method of this application in the comparative test of the embodiment of this application are compared on different infrared small target data sets. Among them, method one and method two use ACM and MDvsFA-cGAN respectively, and method three is GT, which represents the saliency label of infrared small targets. Through comparison, it can be seen that the present application can effectively detect infrared small targets and improve the accuracy of infrared small target detection.

[0028] The technical means disclosed in the present application are not limited to the technical means disclosed in the above-mentioned implementation mode, but also include technical solutions composed of any combination of the above technical features.

Claims

1. An infrared small target recognition method based on color and shape cue learning, characterized in that: The specific steps include: Step S1, dividing the image data set of small infrared targets into a training set and a test set; Step S2: The images of the training set Input into the backbone network of the infrared small target detection network model to extract multi-scale features of the infrared small target image , , and Respectively represent the height and width of a single-channel image; multi-scale features , The resolutions are ; The images of the training set are single-channel images, Indicated in The factor by which the resolution of the corresponding layer is reduced; Step S3, multi-scale features are enhanced through the pyramid network FPN, the path aggregation network PAN and the color shape prompt learning network ColorPeel; Among them, the multi-scale features enhanced by the pyramid network FPN and the path aggregation network PAN belong to the spatial semantic feature domain, and the multi-scale features enhanced by the color shape hint learning network ColorPeel belong to the salient color shape feature domain; Step S4, using the dual probability alignment (DPA) fusion framework to perform heterogeneous fusion of the spatial semantic feature domain and the salient color and shape feature domain to obtain multi-scale information features; Step S5, the fused multi-scale information features are sent to the dynamic detection head Dynamic Head for detection; the detection result is evaluated by the normalized Wasserstein distance NWD and the adaptive threshold focus loss ATFL; Normalized Wasserstein distance NWD and adaptive threshold focal loss ATFL calculate the loss and guide the optimization process of infrared small target detection network model; Step S6: Finally, select the test set image as the input of the trained infrared small target detection network model to obtain the final infrared small target detection result.

2. The infrared small target recognition method based on color and shape cue learning according to claim 1 is characterized in that: In step S3, the color shape hint learning network ColorPeel is used to enhance the color and shape of multi-scale features, specifically including: a: Given the RGB or color coordinates provided by the user, generate instance images and templates; b: The color shape hint learning network ColorPeel generates a color embedding vector c* and a shape embedding vector s*.

3. The infrared small target recognition method based on color shape cue learning according to claim 1 is characterized in that: The dynamic detection head Dynamic Head in step S5 is used to detect and process the fused multi-scale information features, which specifically includes the following steps: Step S51: Given a feature tensor ,in represents all the images in the training set, represents the number of feature pyramid layers, Indicates the size of the feature, = × ; , Respectively represent the height and width of a single-channel image, Indicates the number of channels; Step S52: Perceiving Attention through Scale , spatial perception and attention and task-aware attention Dynamically fuse the features to obtain , the formula is: ; Step S53: Scale-aware attention Dynamic feature fusion is performed based on the multi-scale information features obtained in step S4, and the formula is: ;in It is a 1×1 convolutional layer; is the hard-sigmoid function; spatial perception attention Use deformable convolution to fuse multi-scale information features of different levels in the same spatial position. The formula is: ;in represents the number of sparse sampling locations, represents the number of feature pyramid layers, Indicated in The position in the layer The eigenvalue at is the initial sampling position, is the spatial offset obtained by learning the input features, is in position The scalar is obtained by self-learning and used to adjust the feature weight of the sampling position; task-aware attention Dynamically switch ON and OFF channels to support different tasks. The formula is: ;in is a hyperfunction that controls the threshold, where , Used to perform weighted operations on feature channels. , Used to adjust the bias of feature channels; Through average pooling The dimension is reduced, and then two fully connected layers and a normalization layer are used, and finally normalized by the sigmoid activation function.

4. The infrared small target recognition method based on color shape cue learning according to claim 1 is characterized in that: In step S5, the adaptive threshold focus loss ATFL sets a threshold and regards values ​​below the threshold as hard samples, where the adaptive threshold focus loss ATFL expression is: ;in represents the current average predicted probability value, Represents the predicted value of the next epoch, is a hyperparameter.

5. The infrared small target recognition method based on color and shape cue learning according to claim 1 is characterized in that: The formula for the normalized Wasserstein distance NWD in step S5 is: ; where D is a constant related to the data set, , are the Gaussian distributions modeled by bounding box A and bounding box B, respectively. yes and The Wasserstein distance between them.

Citation Information

Patent Citations

  • Infrared weak and small target identification method based on shape prior segmentation and multi-scale feature aggregation

    CN114842235A

  • Infrared small target identification method based on distraction mining network

    CN117934814A