Neural network infrared small target detection method and system fusing CNN and Mama

By fusing the neural network structure of CNN and Mamba and combining the feature fusion module, the existing infrared small object detection methods in missed and missed detection are solved, and more efficient infrared small object detection is achieved, which has important practical application value.

CN119992064APending Publication Date: 2025-05-13SICHUAN ZHONGKE LANGXING PHOTOELECTRIC TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510153624.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing infrared small object detection methods are prone to missed detection and misdetection when detecting infrared small objects, and deep learning-based methods such as Transformer have high computational complexity, which limits their practical application.

Method used

A method for infrared small-object detection of neural networks that combines CNN and Mamba is proposed. By establishing a neural network structure that combines CNN and Mamba, using the feature extraction capabilities of CNN and Mamba's long-range dependency modeling capabilities, combined with the feature fusion module, the detection of infrared small-objects is achieved.

Benefits of technology

It effectively reduces the misdetection and missed detection of infrared small targets. By fully learning the local detailed information and global context information of infrared small targets, new infrared small target detection methods and ideas are provided, which have important practical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992064A_ABST
    Figure CN119992064A_ABST
Patent Text Reader

Abstract

The invention discloses a neural network infrared small target detection method and system fusing CNN and Mama, and belongs to the field of computer vision. An input image respectively passes through a CNN-based encoder and decoder module and a Mama-based encoder and decoder module to obtain respective output feature maps, then the feature maps are sent to a feature fusion module for feature fusion, an output mask is obtained, and finally a result map is obtained. The method comprises the following steps: extracting local features through a codec by utilizing strong feature extraction and feature interaction capabilities of CNN (Convolutional Neural Network), extracting global features through a codec by utilizing long-range dependence modeling capability of Mama, and then fusing the local features and the global features extracted by the two codec, according to the method, local detail information and global context information of the infrared small target are fully learned by the network, false detection and missing detection of the infrared small target are effectively reduced, a new method and thought are provided for ISTD, and the method has important practical application value and guiding significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a neural network infrared small target detection method and system integrating CNN and Mamba. Background Art

[0002] ISTD (Infrared Small Target Detection) is widely used in various scenarios, such as maritime detection, early warning systems, precision guidance, remote sensing and military tracking systems. Infrared small targets are generally small, dark, lack shape information and are prone to change. When detecting infrared small targets, missed detection and false detection are prone to occur, which brings great challenges to the detection of infrared small targets.

[0003] Current ISTD methods can be divided into traditional methods and deep learning-based methods. In the early days, traditional methods were dominant. However, these methods rely on prior knowledge and artificially designed features. When the image data does not meet the prior knowledge, the detection effect is generally poor. In recent years, deep learning-based methods have received increasing attention in ISTD and achieved remarkable results. These methods automatically learn features in an end-to-end manner, do not rely too much on prior knowledge, and have stronger generalization capabilities in practical applications.

[0004] Specifically, ISTD methods based on deep learning can be divided into two categories: CNN-based methods and Transformer-based methods. CNN-based methods use attention mechanisms, inter-layer interactions between deep features and shallow features, and multiple repeated fusion and enhancement methods to detect infrared small targets. Although CNN has powerful feature extraction capabilities and deep and shallow feature interaction capabilities, it is easy to miss and misdetect small targets in practical applications. This may be related to the fact that CNN focuses on local features but lacks global features. However, global features are very important for ISTD because background pixels and small targets are very similar and global information is needed to distinguish them. Transformer-based methods have powerful long-range dependency modeling capabilities and can effectively extract global features, which to some extent makes up for the shortcomings of CNN-based methods. However, it has quadratic computational complexity and huge computational complexity, which limits its practical application. Summary of the invention

[0005] The object of the present invention is to overcome one or more deficiencies of the prior art and to provide a neural network infrared small target detection method and system integrating CNN and Mamba.

[0006] The objective of the present invention is achieved through the following technical solutions:

[0007] The Transformer-based method has a powerful ability to model long-range dependencies and can effectively extract global features, which makes up for the shortcomings of the CNN-based method to a certain extent. However, it has quadratic computational complexity and huge computational complexity, which limits its practical application. Mamba, which is based on the state-space model, has shown good performance in a variety of long sequence modeling tasks while maintaining linear computational complexity. It is expected to become the next generation basic model after Transformer and provide new solutions for a variety of tasks in the field of computer vision.

[0008] In view of the current status and shortcomings of ISTD, a neural network infrared small target detection method integrating CNN and Mamba is provided, and the method steps include:

[0009] S1. Establish a neural network integrating CNN and Mamba, the neural network structure includes: a CNN-based Encoder and Decoder Module, a Mamba-based Encoder and Decoder Module and a Feature Fusion Module;

[0010] S2. The input image passes through the CNN-based encoder and decoder modules and the Mamba-based encoder and decoder modules respectively to obtain the output feature map;

[0011] S3. Input the output feature map into the feature fusion module for feature fusion to obtain the output mask and the detection result map.

[0012] Furthermore, in step S2, the CNN-based encoder and decoder modules input a 3×H×W image and output four feature maps, namely, C0, C1, C2 and C3;

[0013] Among them, the size of the feature map C0 is , the size of the feature map C1 is , the size of the feature map C2 is , the size of feature map C3 is , where H is the image height and W is the image width.

[0014] Furthermore, in step S2, the input of the Mamba-based encoder and decoder module is a 3×H×W image, and the output is four feature maps M0, M1, M2 and M3;

[0015] Among them, the size of the feature map M0 is , the size of the feature map M1 is , the size of the feature map M2 is , the size of feature map M3 is , where H is the image height and W is the image width.

[0016] Further, in step S3, the input of the feature fusion module is the output feature maps C0, C1, C2 and C3 of the CNN-based encoder and decoder modules and the output feature maps M0, M1, M2 and M3 of the Mamba-based encoder and decoder modules, and the output is a target mask with the same resolution as the original input image and a size of 1×H×W;

[0017] Among them, through CDSM (Concatenation and DySample Module), the channel dimension is first connected, and then the feature maps F0, F1, F2 and F3 are obtained through DySample upsampling; through CCM (Concatenation and Convolution Module), the input feature maps F0, F1, F2 and F3 are first connected in the channel dimension, and then the output mask is obtained through multiple convolution modules to obtain the detection result map.

[0018] A neural network infrared small target detection method system integrating CNN and Mamba is provided to realize infrared small target detection.

[0019] The beneficial effects of the present invention are:

[0020] (1) One encoder / decoder uses CNN’s powerful feature extraction and feature interaction capabilities to extract local features, and another encoder / decoder uses Mamba’s long-range dependency modeling capabilities to extract global features;

[0021] (2) By fusing the local features and global features extracted by the two codecs, the network can fully learn the local detail information and global context information of small infrared targets, effectively reducing the false detection and missed detection of small infrared targets, providing a new method and idea for ISTD, which has important practical application value and guiding significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is the neural network structure diagram that integrates CNN and Mamba;

[0023] Figure 2 This is the structural diagram of the encoder and decoder modules based on CNN;

[0024] Figure 3 This is the CSAM structure diagram;

[0025] Figure 4 This is the structure diagram of the encoder and decoder module based on Mamba;

[0026] Figure 5 It is a diagram of the visual Mamba block structure;

[0027] Figure 6 Feature fusion module structure diagram;

[0028] Figure 7 Output graph for detection results. DETAILED DESCRIPTION

[0029] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0030] A neural network infrared small target detection method integrating CNN and Mamba is provided, and the steps include:

[0031] S1. Establish a neural network structure integrating CNN and Mamba, the neural network structure including: a CNN-based Encoder and Decoder Module, a Mamba-based Encoder and Decoder Module and a Feature Fusion Module;

[0032] S2. The input image passes through the CNN-based encoder and decoder modules and the Mamba-based encoder and decoder modules respectively to obtain the output feature map;

[0033] S3. Input the output feature map into the feature fusion module for feature fusion to obtain the output mask and the detection result map.

[0034] See also Figure 1 By fusing the neural network structure of CNN and Mamba, the input image passes through the encoder and decoder modules based on CNN and the encoder and decoder modules based on Mamba respectively to obtain the feature maps of their respective outputs. These feature maps are then sent to the feature fusion module for feature fusion and the output mask is obtained to obtain the small target detection result image.

[0035] See also Figure 2, based on the CNN encoder and decoder module, the input is a 3×H×W image, and the output is four feature maps C0, C1, C2 and C3. Among them, the size of feature map C0 is , the size of the feature map C1 is , the size of the feature map C2 is , the size of feature map C3 is .

[0036] See also Figure 3 The core components of the encoder and decoder modules based on CNN are the Channel and Spatial Attention Module (CSAM). CSAM realizes the progressive interaction of deep and shallow features through cascading. The cascading methods include down-sampling, up-sampling, up-direction interaction skip connection, down-direction interaction skip connection, and ordinary dense plain skip connection.

[0037] The input of the encoder and decoder modules based on Mamba is a 3×H×W image, and the output is four feature maps M0, M1, M2 and M3. Among them, the size of feature map M0 is , the size of the feature map M1 is , the size of the feature map M2 is , the size of feature map M3 is .

[0038] See also Figure 4 , the encoder and decoder modules based on Mamba are mainly composed of three parts, namely the Visual Patches Generation Module, encoders and decoders. The main function of the Visual Patches Generation Module is to generate visual patches that can be input to the Visual Mamba Block from the input image. Specifically, the input image is evenly divided into n blocks, each of which represents a visual patch; for example, the spatial resolution of each block is (8, 8), that is, a visual patch is represented by an adjacent 8×8 (height×width) pixel area, and the areas between visual patches do not overlap. The sequence composed of these visual patches can be fed to the subsequent Visual Mamba Block to further extract features.

[0039] The encoder has four stages, namely Encoder0, Encoder1, Encoder2 and Encoder3. A single stage of the encoder is mainly composed of a patch merging module and multiple visual Mamba blocks. The patch merging module downsamples the feature map in the spatial dimension and fuses information in the channel dimension to extract higher-level feature representations, reduce the amount of calculation and the number of parameters, and improve the efficiency and performance of the model; the visual Mamba block extracts global information through a sequence of visual sub-blocks.

[0040] See also Figure 5 , LayerNorm represents layer normalization, Linear represents linear layer, DWConv represents depthwise separable convolution, SiLU is the activation function, and SS2D represents two-dimensional selective scanning. The visual Mamba block can effectively extract the spatial features of two-dimensional images through SS2D.

[0041] The decoder has four stages, namely Decoder0, Decoder1, Decoder2 and Decoder3. The single stage of the decoder is mainly composed of a patch expansion module and multiple CNN blocks. The patch expansion module is the inverse process of the patch merging module, which is mainly used to restore the spatial resolution of the feature map and reduce its channel dimension. At the same time, it can fuse multi-scale information and enrich the feature representation; the CNN block will continue to perform deep learning on the feature map obtained by the patch expansion module.

[0042] The feature maps of the corresponding stages of the encoder and decoder are connected through ordinary dense skip connections, and the feature maps of the encoder are directly passed to the decoder to retain more spatial information, thereby improving the accuracy of positioning.

[0043] See also Figure 6 The input of the feature fusion module is the output feature maps C0, C1, C2 and C3 of the CNN-based encoder and decoder modules and the output feature maps M0, M1, M2 and M3 of the Mamba-based encoder and decoder modules, and the output is a target mask with the same resolution as the original input image and a size of 1×H×W.

[0044] Among them, CDSM (Concatenation and DySample Module) represents the channel connection and DySample upsampling module; taking the input feature maps C0 and M0 as an example, they are first connected in the channel dimension, and then the feature map F0 is obtained by DySample upsampling. CCM (Concatenation and Convolution Module) represents the channel connection and convolution module, whose input feature maps F0, F1, F2 and F3 are first connected in the channel dimension, and then the output mask is obtained through multiple convolution modules to obtain the detection result map, see Figure 7 , you can see the small target detected by infrared in the picture.

[0045] The above is only a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not deviate from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.

Claims

1. A neural network infrared small target detection method integrating CNN and Mamba, characterized in that: The method steps include: S1. Establish a neural network integrating CNN and Mamba, wherein the neural network structure includes: an encoder and decoder module based on CNN, an encoder and decoder module based on Mamba, and a feature fusion module; S2. The input image passes through the CNN-based encoder and decoder modules and the Mamba-based encoder and decoder modules respectively to obtain the output feature map; S3. Input the output feature map into the feature fusion module for feature fusion to obtain the output mask and the detection result map.

2. According to claim 1, a neural network infrared small target detection method integrating CNN and Mamba is characterized in that: In step S2, the encoder and decoder modules based on CNN input a 3×H×W image and output four feature maps C0, C1, C2 and C3; Among them, the size of the feature map C0 is , the size of the feature map C1 is , the size of the feature map C2 is , the size of feature map C3 is , where H is the image height and W is the image width.

3. The neural network infrared small target detection method integrating CNN and Mamba according to claim 1 is characterized in that: In step S2, the input of the Mamba-based encoder and decoder module is a 3×H×W image, and the output is four feature maps M0, M1, M2 and M3; Among them, the size of the feature map M0 is , the size of the feature map M1 is , the size of the feature map M2 is , the size of feature map M3 is , where H is the image height and W is the image width.

4. The neural network infrared small target detection method integrating CNN and Mamba according to claim 1 is characterized in that: In step S3, the input of the feature fusion module is the output feature maps C0, C1, C2 and C3 of the CNN-based encoder and decoder modules and the output feature maps M0, M1, M2 and M3 of the Mamba-based encoder and decoder modules, and the output is a target mask with the same resolution as the original input image and a size of 1×H×W; Among them, through the channel connection and DySample upsampling module, the channel dimension is first connected, and then the feature maps F0, F1, F2 and F3 are obtained through DySample upsampling; through the channel connection and convolution module, the input feature maps F0, F1, F2 and F3 are first connected in the channel dimension, and then the output mask is obtained through multiple convolution modules to obtain the detection result map.

5. A neural network infrared small target detection method system integrating CNN and Mamba, characterized in that: Used to execute the neural network infrared small target detection method integrating CNN and Mamba as described in claims 1 to 4, and used to realize infrared small target detection.

Citation Information

Cited By

  • Photoacoustic image enhancement method and device combining Mama and CNN (Convolutional Neural Network)

    CN120707580A

  • A photoacoustic image enhancement method and apparatus combining Mamba and CNN

    CN120707580B