Infrared low, small and slow target detection method combining Mama and edge information

By combining Mamba and edge information, the Mamba-IRS network is constructed based on the U-net network structure, which solves the balance between the calculation burden and detection effect of the existing infrared low-small slow object detection method, and emphasizes the extraction and fusion of edge features, achieving efficient and accurate infrared low-small slow object detection.

CN120107545APending Publication Date: 2025-06-06CHONGQING UNIV OF POSTS & TELECOMM +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510164175.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing infrared low-small and slow target detection methods are difficult to achieve a balance between calculation burden and detection effect, and ignore the edge and shape characteristics of the target, which are of great significance to target recognition.

Method used

Using an infrared low-small slow object detection method combining Mamba and edge information, the Mamba-IRS network is constructed based on the U-net network structure. Through the dual residual path mamba block, decoded block and multi-scale shape adaptive fusion detection head, the advantages of fusion convolution and Mamba are focused on local features and global context information, and the extraction and fusion of edge features are emphasized.

Benefits of technology

It realizes low-calculation and slow target detection with low calculation volume but good results, which can accurately extract and reconstruct the shape characteristics of small targets, improving the accuracy and robustness of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107545A_ABST
    Figure CN120107545A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection, in particular to an infrared low, small and slow target detection method combining Mama and edge information, which comprises the steps that a Mama-IRS network is constructed based on a U-net network structure to realize infrared low, small and slow target detection, and the Mama-IRS network comprises four dual residual path mamba blocks, four decoding blocks and one multi-scale shape adaptive fusion detection head; a lower sampling layer is arranged between every two adjacent dual residual path mamba blocks, an upper sampling layer is arranged between every two adjacent decoding blocks, and the lower sampling layer, the residual attention block, the multi-scale feature adaptive aggregation block, the upper sampling layer and the first decoding block are sequentially connected behind the fourth dual residual path mamba block; the decoding block comprises a residual attention block and a patch expansion 2d layer; according to the method, the Mama is applied to infrared low, small and slow target detection, the calculation amount is low, and the precision is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to an infrared low, small and slow target detection method combining Mamba and edge information. Background Art

[0002] In the detection scenario of infrared low, small and slow targets, algorithms are often deployed in edge intelligent devices. Many algorithms cannot be deployed due to excessive computing burden, which has a great obstacle to the development of the supervision and defense of the entire low, small and slow targets. At the same time, detecting low, small and slow targets in the system is the first step. After detecting low, small and slow targets, it is necessary to restore their contour shape and other detailed features during detection, so that their specific categories can be distinguished in subsequent recognition, so as to make the overall system more robust and minimize false detection. With the continuous development of technology, there are also higher requirements for the shape detection accuracy of infrared low, small and slow targets.

[0003] Traditional infrared low-light, small and slow detection methods are mainly based on image processing and machine learning techniques, including filtering, human visual system (HVS), local contrast and low-rank representation; most traditional infrared low-light, small and slow detection methods have problems such as weak robustness and the need for a large number of hyperparameters. Specifically, the setting and adjustment of hyperparameters take a lot of time, and weak robustness will make the method unable to be applied in multiple scenarios; CNN-based methods have low computational complexity and focus on local features, but ignore global context information; VIT-based methods can more effectively capture global information and long-distance dependencies in images through self-attention mechanisms, but the computational complexity is high and has a heavy computational burden. In addition, these methods often only focus on the existence of small targets, but invariably ignore the edge and shape features of infrared small targets, but these features are of great significance for further target recognition and can effectively distinguish the categories of small targets.

[0004] In summary, although a variety of low-resolution, small, and slow object detection methods have been proposed, the following problems still exist:

[0005] 1) Not paying enough attention to the edge and shape of the target;

[0006] 2) There is no balance between detection effect and computational burden, that is, algorithms with large computational load often have good results, while algorithms with poor results have large computational load;

[0007] 3) CNN-based methods have low computational complexity but focus on local features; while transformers used to process long sequences focus on capturing long-distance dependencies but have high computational complexity. Summary of the invention

[0008] To solve the above problems, the present invention provides an infrared low, small and slow target detection method combining Mamba and edge information. A Mamba-IRS network is constructed based on a U-net network structure to realize infrared low, small and slow target detection. The Mamba-IRS network includes 4 dual residual path mamba blocks, 4 decoding blocks and 1 multi-scale shape adaptive fusion detection head; wherein, a downsampling layer is provided between every two adjacent dual residual path mamba blocks, and an upsampling layer is provided between every two adjacent decoding blocks. The fourth dual residual path mamba block is sequentially connected to the downsampling layer, the residual attention block, the multi-scale feature adaptive aggregation block, the upsampling layer and the first decoding block; the decoding block includes a residual attention block and a patchexpand2d layer.

[0009] The infrared low, small and slow target detection method based on the Mamba-IRS network specifically includes:

[0010] S1. Input the infrared small target image into the Laplacian pyramid to obtain target edge feature maps of 5 different scales;

[0011] S2. Input the infrared small target image into the stem block to obtain the initial feature map, and the initial feature map passes through four dual residual path mamba blocks to obtain encoding feature maps of different scales;

[0012] S3. Fusing the target edge feature map and the encoding feature map of the same scale to obtain a fused feature map;

[0013] S4. Each decoding block takes the up-sampled result of the output of the previous decoding block and the fused feature map of the same scale as its own as input, and outputs a decoding feature map; wherein, for the first decoding block, the output of the previous decoding block adopts the multi-scale feature adaptive aggregation block output;

[0014] S5. Input the output of the multi-scale feature adaptive aggregation block and the decoded feature maps of different scales into the multi-scale shape adaptive fusion detection head to obtain the detection result.

[0015] Furthermore, the processing of the dual residual path mamba block includes:

[0016] S11. The input features are passed through the residual wavelet sampling module to obtain local features, and the local features are divided into LNS input and Mamba input by using a channel division operation;

[0017] S12. The LNS input is unfolded into a visual block of 8×8 size, the visual block is passed through three LNS modules and one sigmoid activation function layer to obtain a first detail feature, and the first detail feature is added to the LNS input to obtain a first local detail feature; the LNS module includes a cascaded Linear layer, a Norm layer and a sigmoid activation function layer;

[0018] S13. Passing the Mamba input through the Mamba module and the sigmoid activation function layer to obtain a second detail feature, and adding the second detail feature to the Mamba input to obtain a second local detail feature;

[0019] S14. Concatenate the local feature, the first local detail feature, and the second local detail feature, and pass them through a convolution layer to obtain an output feature.

[0020] Furthermore, step S11 obtains local features by passing the input features through the residual wavelet sampling module, including:

[0021] The input features are downsampled by wavelet to obtain downsampled features; the downsampled features are passed through three convolutional layers to obtain local features.

[0022] Further, step S13 passes the Mamba input through the Mamba module and the sigmoid activation function layer to obtain the second detail feature, including

[0023] S131.Mamba input passes through the linear layer, one-dimensional convolution layer, sigmoid activation function layer, and SS2D layer in sequence to obtain convolution features;

[0024] S132. Multiply the convolution feature with the Mamba input and pass it through the Norm layer to obtain the Mamba output;

[0025] S133. Pass the Mamba output through the sigmoid activation function layer to obtain the second detail feature.

[0026] Furthermore, the multi-scale feature adaptive aggregation block includes multiple CBRs, each CBR includes a convolution layer, a Norm layer, and a Relu activation function layer; all CBRs are further divided into three types: 1×1 CBR, 3×3 CBR, and 5×5 CBR according to the size of the convolution kernel of the convolution layer; the processing process of the multi-scale feature adaptive aggregation block includes

[0027] The input feature is passed through 1×1 CBR to obtain feature F1, and feature F1 is multiplied by the input feature to obtain feature F11;

[0028] The input feature is passed through the first 3×3 CBR to obtain feature F2, and feature F2 is multiplied by the input feature and passed through the second 3×3 CBR to obtain feature F21;

[0029] The input feature is passed through the first 5×5 CBR to obtain feature F3, feature F3 is multiplied by the input feature and then passed through the second 5×5 CBR to obtain feature F31;

[0030] Add features F11, F21, and F31 to get feature F4, and pass feature F4 through the third 3×3 CBR to get feature F41;

[0031] The input feature is passed through the sigmoid activation function layer and then added to the feature F41 to obtain the feature F42, and the feature F42 is passed through a 1×1 convolution layer to obtain the output feature.

[0032] Furthermore, the residual attention block includes a first CBR, a second CBR, an EFFA and a third CBR which are cascaded in sequence, wherein a residual connection exists between the input feature of the residual attention block and the output of the EFFA; and the processing process of the EFFA includes:

[0033] The output of the second CBR is used as the input of EFFA Getting Input The corresponding target edge feature map

[0034] The target edge feature map The first feature is obtained through the sigmoid activation function layer, and the target edge feature map is With input Multiply to get the second feature;

[0035] The first feature is added to the second feature and then passed through the adaptive average pooling layer to obtain a deep fusion edge information feature map. Deep fusion edge information feature map The third feature is obtained through a one-dimensional convolution layer and a sigmoid activation function layer;

[0036] Combine the third feature with the target edge feature map Multiply the result and the target edge feature map The output features of EFFA are obtained by adding them together.

[0037] Furthermore, the multi-scale shape adaptive fusion detection head includes fusion block 1, fusion block 2, fusion block 3, fusion block 4 and a final fusion block which are cascaded in sequence; wherein, fusion block i, i = 1, 2, 3, 4, the processing process includes

[0038] Get the target edge feature map E of the same scale as fusion block i iAnd the decoded feature map R i , the output of the previous module and the target edge feature map E i And the decoded feature map R i Add them together to get the fusion graph;

[0039] The fusion map is upsampled and input into the convolution layer and BN layer to obtain the feature fusion map;

[0040] The feature fusion map is combined with the target edge feature map E that is one level larger than the fusion block i. i+1 After addition, input the Relu activation function layer to get the output.

[0041] Beneficial effects of the present invention:

[0042] The present invention proposes a hybrid backbone network of Mamba and CNN, and applies Mamba to infrared low, small and slow target detection. The method has low computational complexity and is not inferior to transformers in effect. Specifically, a dual residual path mamba block that integrates convolution and Mamba is designed, and the characteristics of convolution focusing on local features are used to obtain texture and edge features, and global context information is obtained through mamba. At the bottom layer of the encoder, a multi-scale feature adaptive aggregation block is added, and convolution kernels of different sizes are combined to obtain richer feature information. Edge feature focused attention (EFFA) that focuses on edge and shape features is proposed to focus on target edge features, strengthen the shape feature extraction of infrared small targets, and filter out backgrounds similar to the target. A multi-scale shape adaptive fusion detection head (MSAFH) is also proposed. The shape features are deeply fused again in the module, and the feature extraction of small target shapes is once again strengthened. The feature maps output by each layer in the U-Net structure are fused to achieve accurate infrared small target shape reconstruction and accurate target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic diagram of the overall structure of the algorithm of the present invention;

[0044] Figure 2 This is a diagram of the dual residual path mamba block structure of the present invention;

[0045] Figure 3 This is a diagram of the multi-scale feature adaptive aggregation block structure of the present invention;

[0046] Figure 4 This is a diagram of the residual attention block and edge feature focused attention structure of the present invention;

[0047] Figure 5 This is a structural diagram of the multi-scale shape adaptive fusion detection head of the present invention. DETAILED DESCRIPTION

[0048] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0049] The invention provides an infrared low, small and slow target detection method combining Mamba with edge information; wherein the architecture of an encoder and a decoder based on a U-net network combines the advantages of CNN and Mamba, deeply fuses the edge features of the target, and constructs a Mamba-IRS network for realizing infrared low, small and slow target detection, wherein the Mamba-IRS network comprises four dual residual path mamba blocks (DRMBs), four decoding blocks and one multi-scale shape adaptive fusion detection head (MSAFH); wherein a downsampling layer is arranged between every two adjacent DRMBs, an upsampling layer is arranged between every two adjacent decoding blocks, and the fourth DRMB is sequentially connected to the downsampling layer, a residual attention block, a multi-scale feature adaptive aggregation block (MFAAB), an upsampling layer and the first decoding block; the decoding block comprises a residual attention block and a patchexpand2d layer.

[0050] like Figure 1 As shown in the figure, the infrared low, small and slow target detection method based on the Mamba-IRS network specifically includes:

[0051] S1. Input the infrared small target image into the Laplacian pyramid to obtain target edge feature maps of 5 different scales;

[0052] S2. Input the infrared small target image into the stem block to obtain the initial feature map; the initial feature map is downsampled by the encoder to learn the rich multi-scale features of the small target, that is, the encoded feature maps of different scales are obtained through 4 dual residual path mamba blocks;

[0053] S3. The target edge feature map and the encoding feature map of the same scale are fused to obtain a fused feature map; wherein the encoding feature map output by the fourth DRMB is fused with the target edge feature map of the smallest scale after downsampling, and the fusion result is used as the input of the residual attention block connected after the fourth downsampling layer;

[0054] S4. Each decoding block takes the up-sampled result of the output of the previous decoding block and the fused feature map of the same scale as its own as input, and outputs a decoding feature map; among them, for the first decoding block, the output of the previous decoding block adopts the multi-scale feature adaptive aggregation block output; MFAAB is used for multi-scale global context feature fusion, and the expression ability of the model is enhanced by stacking and fusing these features;

[0055] S5. The multi-scale feature adaptive aggregation block output and the decoded feature maps of different scales are input into the multi-scale shape adaptive fusion detection head to further enhance the influence of edge features on the network, make the edge reconstruction of small targets more accurate, and obtain the detection results.

[0056] Preferably, DRMB can effectively utilize the advantages of CNN and Mamba to obtain global context information while also focusing on local detail features. The structure of DRMB is as follows: Figure 2 As shown, the processing process includes:

[0057] The input features are passed through the residual wavelet sampling module (ResWtBlock) to obtain local features, and the channel partitioning operation is used to divide the local features into LNS input and Mamba input.

[0058] Specifically, in ResWtBlock, wavelet down sampling is first used to reduce the spatial resolution of the input features while retaining as much information as possible to obtain downsampled features; compared with traditional downsampling methods, this method can effectively reduce information uncertainty and information loss, thereby improving the accuracy of semantic segmentation. The downsampled features are then passed through three convolutional layers to obtain local features. Specifically, the parameters of these three convolutional layers are set differently. The first convolutional layer is used to double the channels and extract more information. After the first convolutional layer, wavelet down sampling is performed again to extract multi-scale information. The second convolutional layer is used to integrate multi-scale information, and the third convolutional layer is used to reduce the dimensionality of the doubled number of channels.

[0059] In order to reduce the computational cost, the present invention performs channel segmentation on the local features and sends the segmentation results to the LNS module and the Mamba module respectively, so that more diverse features can be captured;

[0060] S12. Unfold the LNS input into a visual block of 8×8 size, which helps the LNS module to obtain more local detail features and reduce the amount of calculation. The visual block passes through three LNS modules and one sigmoid activation function layer to obtain the first detail feature, and the first detail feature is added to the LNS input to obtain the first local detail feature; the LNS module includes a cascaded Linear layer, a Norm layer, and a sigmoid activation function layer. This process can be expressed as

[0061]

[0062] Among them, LNS 3 Indicates that there are three cascaded LNS modules, σ represents the sigmoid activation function, Indicates LNS input, Represents the first local detail feature.

[0063] S13. Pass the Mamba input through the Mamba module and the sigmoid activation function layer to obtain a second detail feature, and add the second detail feature to the Mamba input to obtain a second local detail feature.

[0064] Specifically, the present invention adds residual connection and introduces one-dimensional convolution in the design of Mamba block. It is worth noting that the residual connection here is different from the usual one, and its function is more like strengthening the focus area and weakening the background attention. This provides specific attention to the focus area without increasing the computational complexity. The present invention calls it residual attention connection. Due to its effectiveness, this operation has many applications in subsequent modules. The specific structure of the Mamba module is as follows: Figure 2 As shown in (b), step S13 passes the Mamba input through the Mamba module and the sigmoid activation function layer to obtain the second detail feature, including

[0065] S131.Mamba input passes through the linear layer, one-dimensional convolution layer, sigmoid activation function layer, and SS2D layer in sequence to obtain convolution features;

[0066] The SS2D layer refers to a Mamba module, and its full name is 2D-Selective Operation. The SS2D layer adopts a four-way scanning strategy, that is, scanning from the four corners of the feature map simultaneously to ensure that each element in the feature integrates information from all other positions in different directions, thereby forming a global receptive field without increasing the linear computational complexity.

[0067] S132. Multiply the convolution feature with the Mamba input and pass it through the Norm layer to obtain the Mamba output;

[0068] S133. Pass the Mamba output through the sigmoid activation function layer to obtain the second detail feature.

[0069] The above process can be expressed as

[0070]

[0071] Among them, Linear represents the linear layer, Norm represents LinearNorm, and Conv1D represents one-dimensional convolution. Indicates Mamba input, Represents the second detail feature.

[0072] S14. Concatenate the local feature, the first local detail feature, and the second local detail feature, and pass them through a convolution layer to obtain an output feature.

[0073] The LNS module is better at extracting linear features, while the Mamba module is better at processing complex dependencies. By combining the LNS module and the Mamba module and making full use of their respective advantages, the fusion effect and performance can be significantly improved.

[0074] Preferably, the size of the convolution kernel determines the size of its receptive field. A smaller convolution kernel has a smaller receptive field and can extract local features in the image, such as edges and textures; while a larger convolution kernel has a larger receptive field and can capture global features in the image. Based on this, the present invention is designed as follows Figure 3 The MFAAB structure shown in the figure includes multiple CBRs, each of which includes a convolution layer, a Norm layer and a Relu activation function layer; the present invention uses three sizes of convolution kernels, namely 1, 3 and 5. This is because the target is small, and a smaller convolution kernel is more reasonable, and it also reduces the computational complexity. Here, all CBRs are further divided into three types: 1×1CBR, 3×3CBR and 5×5CBR according to the size of the convolution kernel containing the convolution layer. MFAAB can be divided into four branches: residual attention connection and three convolution paths. The residual attention connection is used to prevent the disappearance of features after convolution. The convolution path with three different sizes of convolution kernels is used to process the input features. After fusion, the receptive field size can be effectively increased, and a wider range of contextual information can also be captured. Residual attention connection is also used in each convolution path to strengthen the expression of features, thereby accelerating the convergence speed of the model. By integrating paths with receptive fields of different sizes, MFAAB improves the model's recognition ability for small targets of different scales for the network, while enhancing robustness, and can better cope with interference from factors such as complex backgrounds and different lighting conditions in images.

[0075] Specifically, the MFAAB process includes:

[0076] The input feature is passed through 1×1 CBR to obtain feature F1, and feature F1 is multiplied by the input feature to obtain feature F11;

[0077] The input feature is passed through the first 3×3 CBR to obtain feature F2, and feature F2 is multiplied by the input feature and passed through the second 3×3 CBR to obtain feature F21;

[0078] The input feature is passed through the first 5×5 CBR to obtain feature F3, feature F3 is multiplied by the input feature and then passed through the second 5×5 CBR to obtain feature F31;

[0079] Feature F11, feature F21, and feature F31 are added together to obtain feature F4, and feature F4 is passed through the third 3×3 CBR to obtain feature F41;

[0080] The input feature is passed through the sigmoid activation function layer and then added to the feature F41 to obtain the feature F42, and the feature F42 is passed through a 1×1 convolution layer to obtain the output feature.

[0081] Preferably, edge information is usually easily ignored in infrared small target detection, and only a few networks notice its value. However, in fact, these methods that combine target edge features often achieve good experimental results. The reason is that edge features, in the process of combining with network features, also help the network locate small targets, and edge features can assist the network to accurately reconstruct small targets. Based on this, the present invention designs a residual attention block (ResAttention block), in which EFFA is introduced to better integrate the edge information of infrared small targets into the network. Figure 4 As shown in (a), the residual attention block includes a first CBR, a second CBR, an EFFA and a third CBR which are cascaded in sequence, wherein there is a residual connection between the input features of the residual attention block and the output of the EFFA.

[0082] Specifically, Figure 4 As shown in (b), the EFFA designed by the present invention is a channel attention with a very simple but effective structure. The edge information of the target is deeply integrated with the feature map from top to bottom to help establish a more accurate target shape. The processing process of the EFFA includes:

[0083] The output of the second CBR is used as the input of EFFA Getting Input The corresponding target edge feature map That is, the target edge feature map with the same processing scale as the current residual attention block; specifically, the target edge information in EFFA is obtained through the Laplacian pyramid, which can capture the details and edge information in the image, and each layer contains the residual information of the original image at different scales. These residual information mainly corresponds to high-frequency details such as edges and textures in the image, which plays an important role in the fusion of EFFA and the subsequent edge information in MSSADH, and can obtain finer edges.

[0084] The target edge feature map The first feature is obtained through the sigmoid activation function layer, and the target edge feature map is With input Multiply to get the second feature;

[0085] The first feature is added to the second feature and then passed through the adaptive average pooling layer to obtain a deep fusion edge information feature map. Deep fusion edge information feature map The third feature is obtained through a one-dimensional convolution layer and a sigmoid activation function layer;

[0086] Combine the third feature with the target edge feature map Multiply the result and the target edge feature map Add together to get the output feature E of EFFA out .

[0087] The above process can be expressed as

[0088]

[0089] Among them, AdAPool represents the adaptive average pooling operation (AdaptiveAvgPool), which adaptively outputs the feature map to a size of 1×1; Conv1D is a one-dimensional convolution operation. EFFA performs two point multiplication and addition operations with edge information, and the edge features are greatly enhanced in the feature map, which helps the network learn more data features and rules, thereby improving the generalization ability.

[0090] Preferably, the detection head is a key component in the target detection algorithm, located at the end of the network, and is directly related to the accuracy and efficiency of the algorithm. Most methods tend to use a lightweight detection head, while the present invention introduces a heavier detection head to further combine edge features and strengthen multi-scale feature fusion. The design concept of the MSAFH proposed in the present invention is similar to that of PANet, and an additional bottom-up path is introduced to integrate the target edge information, the precise positioning information of low-level features, and the rich semantic information of high-level features, thereby enhancing the representation ability of the entire feature hierarchy, such as Figure 5 The multi-scale shape adaptive fusion detection head includes fusion block 1, fusion block 2, fusion block 3, fusion block 4 and the final fusion block, which are cascaded in sequence; wherein, fusion block i, i = 1, 2, 3, 4, the processing process includes

[0091] Get the target edge feature map E of the same scale as fusion block i i And the decoded feature map R i , the output of the previous module and the target edge feature map E i And the decoded feature map R iAdd them together to get the fusion graph;

[0092] The fusion map is upsampled and input into the convolution layer and BN layer to obtain the feature fusion map;

[0093] The feature fusion map is combined with the target edge feature map E that is one level larger than the fusion block i. i+1 After addition, input the Relu activation function layer to get the output.

[0094] The structure of the final fusion block is similar to that of the fusion block, except that the operation of adding the target edge feature map with a larger scale is removed based on the fusion block structure.

[0095] The above process can be expressed as

[0096] F i =Relu(E i+1 +BN(Conv(Up(E i +F i-1 +R i ))))

[0097] H out =Relu(BN(Conv(Up(E 4 +F 3 +R 4 ))))

[0098] Among them, F i represents the output of the i-th fusion block, H out Represents the output of the multi-scale shape adaptive fusion detection head, which outputs the final prediction of the network. In fusion block1, the previous module outputs F 0 R 1 Or ignore. Fusing features during training and removing them during inference can greatly reduce the impact of noise caused by background edges.

[0099] In the present invention, unless otherwise clearly stipulated and limited, the terms such as "installation", "setting", "connection", "fixation" and "rotation" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral one; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium; it can be the internal connection of two elements or the interaction relationship between two elements. Unless otherwise clearly defined, ordinary technicians in this field can understand the specific meanings of the above terms in the present invention according to the specific circumstances.

[0100] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An infrared low, small and slow target detection method combining Mamba and edge information, characterized in that: A Mamba-IRS network is constructed based on the U-net network structure to realize infrared low, small and slow target detection. The Mamba-IRS network includes 4 dual residual path mamba blocks, 4 decoding blocks and 1 multi-scale shape adaptive fusion detection head; wherein, a downsampling layer is provided between every two adjacent dual residual path mamba blocks, and an upsampling layer is provided between every two adjacent decoding blocks. The fourth dual residual path mamba block is sequentially connected to the downsampling layer, the residual attention block, the multi-scale feature adaptive aggregation block, the upsampling layer and the first decoding block; the decoding block includes a residual attention block and a patch expand2d layer; The infrared low, small and slow target detection method based on the Mamba-IRS network specifically includes: S1. Input the infrared small target image into the Laplacian pyramid to obtain target edge feature maps of 5 different scales; S2. Input the infrared small target image into the stem block to obtain the initial feature map, and the initial feature map passes through 4 dual residual path mamba blocks to obtain encoding feature maps of different scales; S3. Fusing the target edge feature map and the encoding feature map of the same scale to obtain a fused feature map; S4. Each decoding block takes the up-sampled result of the output of the previous decoding block and the fused feature map of the same scale as its own as input, and outputs a decoding feature map; wherein, for the first decoding block, the output of the previous decoding block adopts the multi-scale feature adaptive aggregation block output; S5. Input the output of the multi-scale feature adaptive aggregation block and the decoded feature maps of different scales into the multi-scale shape adaptive fusion detection head to obtain the detection result.

2. According to claim 1, a method for detecting infrared low, small and slow targets combining Mamba and edge information is characterized in that: The processing of the dual residual path mamba block includes: S11. The input features are passed through the residual wavelet sampling module to obtain local features, and the local features are divided into LNS input and Mamba input by using a channel division operation; S12. The LNS input is unfolded into a visual block of 8×8 size, the visual block is passed through three LNS modules and one sigmoid activation function layer to obtain a first detail feature, and the first detail feature is added to the LNS input to obtain a first local detail feature; the LNS module includes a cascaded Linear layer, a Norm layer and a sigmoid activation function layer; S13. Passing the Mamba input through the Mamba module and the sigmoid activation function layer to obtain a second detail feature, and adding the second detail feature to the Mamba input to obtain a second local detail feature; S14. Concatenate the local feature, the first local detail feature, and the second local detail feature, and pass them through a convolution layer to obtain an output feature.

3. The infrared low, small and slow target detection method combining Mamba and edge information according to claim 2 is characterized in that: Step S11 obtains local features by passing the input features through the residual wavelet sampling module, including: The input features are downsampled by wavelet to obtain downsampled features; the downsampled features are passed through three convolutional layers to obtain local features.

4. The infrared low, small and slow target detection method combining Mamba and edge information according to claim 2 is characterized in that: Step S13 passes the Mamba input through the Mamba module and the sigmoid activation function layer to obtain the second detail feature, including S131.Mamba input passes through the linear layer, one-dimensional convolution layer, sigmoid activation function layer, and SS2D layer in sequence to obtain convolution features; S132. Multiply the convolution feature with the Mamba input and pass it through the Norm layer to obtain the Mamba output; S133. Pass the Mamba output through the sigmoid activation function layer to obtain the second detail feature.

5. The infrared low, small and slow target detection method combining Mamba and edge information according to claim 1, characterized in that: The multi-scale feature adaptive aggregation block includes multiple CBRs, each of which includes a convolution layer, a Norm layer, and a Relu activation function layer; all CBRs are further divided into three types: 1×1 CBR, 3×3 CBR, and 5×5 CBR according to the size of the convolution kernel of the convolution layer; the processing process of the multi-scale feature adaptive aggregation block includes The input feature is passed through 1×1 CBR to obtain feature F1, and feature F1 is multiplied by the input feature to obtain feature F11; The input feature is passed through the first 3×3 CBR to obtain feature F2, and feature F2 is multiplied by the input feature and passed through the second 3×3 CBR to obtain feature F21; The input feature is passed through the first 5×5 CBR to obtain feature F3, feature F3 is multiplied by the input feature and then passed through the second 5×5 CBR to obtain feature F31; Feature F11, feature F21, and feature F31 are added together to obtain feature F4, and feature F4 is passed through the third 3×3 CBR to obtain feature F41; The input feature is passed through the sigmoid activation function layer and then added to the feature F41 to obtain the feature F42, and the feature F42 is passed through a 1×1 convolution layer to obtain the output feature.

6. The infrared low, small and slow target detection method combining Mamba and edge information according to claim 1, characterized in that: The residual attention block includes a first CBR, a second CBR, an EFFA and a third CBR which are cascaded in sequence, wherein a residual connection exists between the input feature of the residual attention block and the output of the EFFA; the processing process of the EFFA includes: The output of the second CBR is used as the input of EFFA Getting Input The corresponding target edge feature map The target edge feature map The first feature is obtained through the sigmoid activation function layer, and the target edge feature map is With input Multiply to get the second feature; The first feature is added to the second feature and then passed through the adaptive average pooling layer to obtain a deep fusion edge information feature map. Deep fusion edge information feature map The third feature is obtained through a one-dimensional convolution layer and a sigmoid activation function layer; Combine the third feature with the target edge feature map Multiply the result and the target edge feature map The output features of EFFA are obtained by adding them together.

7. The infrared low, small and slow target detection method combining Mamba and edge information according to claim 1, characterized in that: The multi-scale shape adaptive fusion detection head includes fusion block 1, fusion block 2, fusion block 3, fusion block 4 and the final fusion block which are cascaded in sequence; wherein, fusion block i, i = 1, 2, 3, 4, the processing process includes Get the target edge feature map E of the same scale as fusion block i i And the decoded feature map R i , the output of the previous module and the target edge feature map E i And the decoded feature map R i Add them together to get the fusion graph; The fusion map is upsampled and input into the convolution layer and BN layer to obtain the feature fusion map; The feature fusion map is combined with the target edge feature map E that is one level larger than the fusion block i. i+1 After addition, input the Relu activation function layer to get the output.