Infrared Small Target Detection Method, Storage Medium, and Computer Device
The red infrared small target detection method employs a spatial-channel dual-stream self-attention module with multi-scale fusion to address feature expression and environmental interference issues, improving detection accuracy and robustness in complex scenarios.
Patent Information
- Application Number
- CN202510534573.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The prior art has problems such as insufficient detection stability, high false alarm rate, insufficient feature expression ability, and complex background interference affecting detection reliability in infrared small target detection, especially in low signal-to-noise ratio and complex environments. The detection accuracy needs to be improved.
The cascading processing architecture of the space channel dual-flow self-attention module is adopted, combined with layer normalization and multi-layer perceptron, and feature enhancement is performed through the spatial self-attention and channel self-attention mechanism, and the dynamic feature recalibration mechanism of multi-scale information fusion and decoder are used to improve feature representation capabilities.
It significantly improves the accuracy and robustness of infrared small target detection, can effectively detect weak targets in complex backgrounds, and improves the adaptability and generalization performance of the detection algorithm.
Smart Images

Figure CN120047757B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of infrared image detection and computer technology, and particularly relates to an infrared small target detection method, a storage medium, and a computer device. Background Art
[0002] In the technical field of infrared image detection, small target detection, as a key technology, plays an important role in application scenarios such as national defense security, intelligent monitoring, and autonomous driving. Infrared image detection, with its all-weather working ability and sensitivity to temperature differences, demonstrates unique advantages under low illumination and adverse weather conditions, becoming an important supplement to visible light imaging.
[0003] Traditional small target detection methods mainly rely on image processing techniques, including but not limited to: differential methods based on background modeling, enhancement algorithms based on morphological operations, and detection techniques based on edge features. These methods achieve target separation by constructing a background model, enhance target features using mathematical morphological operations, or extract target contours with the help of edge operators. However, when dealing with actual scenarios such as complex background interference and low signal-to-noise ratio of targets, these methods often suffer from problems such as insufficient detection stability and high false alarm rates.
[0004] In recent years, with the rapid development of deep learning technology, object detection algorithms based on convolutional neural networks have gradually become the mainstream. Two-stage and one-stage detectors represented by Faster R-CNN (Fast Region-based Convolutional Neural Network) and YOLO (You Only Look Once) have significantly improved detection performance through end-to-end feature learning and object localization. However, for the special task of infrared small target detection, existing deep learning methods still face many challenges: First, the limited pixel information of small targets leads to insufficient feature expression ability; second, the low contrast between the target and the background increases the difficulty of feature discrimination; in addition, background interference in complex environments also affects the reliability of detection. Summary of the Invention
[0005] In view of this, the present application provides an infrared small target detection method, a storage medium, and a computer device. Through the cascaded processing architecture of the spatial channel dual-stream self-attention module (spatial attention, LN, MLP, channel attention, LN, MLP), hierarchical feature enhancement from local to global is achieved. The spatial branch focuses on pixel-level relationship modeling, and the channel branch strengthens semantic feature interaction. The two cooperate and optimize to significantly improve the feature representation ability, overcoming the problem of insufficient feature expression ability of traditional methods in complex backgrounds. The skip connection part adopts an improved multi-scale information fusion channel multi-head self-attention module, unifies the multi-scale feature representation through block division operations, and establishes cross-scale global dependence relationships using a residual multi-head attention mechanism. Compared with traditional convolution or simple splicing methods, it more effectively fuses local details and global context information, enhancing the adaptability to targets of different sizes.
[0006] According to one aspect of the present application, there is provided an infrared small target detection method, which is applied to an infrared small target detection network. The infrared small target detection network includes a channel multi-head self-attention module for multi-scale information fusion connected to a layer normalization and multi-layer perceptron joint module, an infrared small target detection module, and an encoder and a decoder with a reciprocal structure. The encoder is a cascaded processing architecture composed of four-level spatial channel dual-stream self-attention modules. Each level of the spatial channel dual-stream self-attention module includes a spatial self-attention mechanism and a channel self-attention mechanism. The layer normalization and multi-layer perceptron joint module includes layer normalization and a multi-layer perceptron. The method includes:
[0007] Obtain an infrared image to be detected containing an infrared small target. After using a 3×3 convolutional layer to perform feature embedding on the infrared image to be detected and obtaining an initial feature map, input the initial feature map into the infrared small target detection network to implement infrared small target detection by the infrared small target detection network to perform the following steps:
[0008] Use the spatial self-attention mechanism to project the initial feature map into queries, keys, and values in the spatial dimension, calculate the similarity between the queries and keys in the spatial dimension through the Softmax function to obtain spatial attention weights, and perform weighted aggregation on the values in the spatial dimension according to the spatial attention weights to obtain a spatially enhanced feature map;
[0009] After using layer normalization to perform a primary feature transformation on the spatially enhanced feature map, use a multi-layer perceptron to perform a secondary feature transformation on the spatially enhanced feature map after the primary feature transformation to obtain an intermediate feature map;
[0010] The intermediate feature map after reshaping is projected into queries, keys, and values in the channel dimension using the channel self-attention mechanism. The similarity between the queries and keys in the channel dimension is calculated through the Softmax function to obtain the channel attention weights. The values in the channel dimension are weighted and aggregated according to the channel attention weights to obtain the channel-enhanced feature map;
[0011] After initially transforming the channel-enhanced feature map using layer normalization, a multi-layer perceptron is used to perform a secondary feature transformation on the channel-enhanced feature map after the initial feature transformation, and a multi-scale feature fusion map is output;
[0012] The decoder performs reverse decoding on the multi-scale feature fusion map to obtain a multi-scale feature fusion infrared image, and uses an adaptive weight adjustment mechanism to perform feature recalibration on the multi-scale feature fusion infrared image to generate a binary detection map;
[0013] The infrared small target detection module detects infrared small targets using the binary detection map.
[0014] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the above-mentioned infrared small target detection method is implemented.
[0015] According to still another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the program, the above-mentioned infrared small target detection method is implemented.
[0016] By means of the above technical solutions, an infrared small target detection method, a storage medium, and a computer device provided by the present application can improve the detection accuracy and robustness of infrared small targets by mixing spatial and channel multi-attention during the infrared small target detection process.
[0017] The above description is only an overview of the technical solutions of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are hereinafter specifically exemplified. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0019] Figure 1 A schematic diagram of an infrared small target detection network used in an infrared small target detection method provided by an embodiment of the present application is shown;
[0020] Figure 2 Shows a schematic diagram of a spatial channel dual-stream self-attention module used in an infrared small target detection method provided by an embodiment of the present application;
[0021] Figure 3 Shows a schematic diagram of a channel multi-head self-attention module for multi-scale information fusion used in an infrared small target detection method provided by an embodiment of the present application;
[0022] Figure 4 Shows a schematic diagram of the flow of an infrared small target detection method provided by an embodiment of the present application;
[0023] Figure 5 Shows a schematic diagram of a feature recalibration self-attention module used in an infrared small target detection method provided by an embodiment of the present application. Detailed implementation manners
[0024] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0025] In this embodiment, an infrared small target detection method is provided, which is applied to an infrared small target detection network, such as Figure 1 shown, the infrared small target detection network includes a channel multi-head self-attention module for multi-scale information fusion connected with a layer normalization and multi-layer perceptron joint module (such as Figure 2 shown), an infrared small target detection module, and an encoder and a decoder with a reciprocal structure. The encoder is a cascade processing architecture and is composed of four-level spatial channel dual-stream self-attention modules. Each level of the spatial channel dual-stream self-attention module is as Figure 3 shown, and respectively includes a spatial self-attention mechanism and a channel self-attention mechanism. The layer normalization and multi-layer perceptron joint module includes layer normalization and a multi-layer perceptron. The method includes:
[0026] Obtain a to-be-detected infrared image containing an infrared small target. After using a 3×3 convolutional layer to perform feature embedding on the to-be-detected infrared image to obtain an initial feature map, input the initial feature map into the infrared small target detection network to perform the following steps with reference to Figure 4 shown to achieve infrared small target detection:
[0027] Step 101, use the spatial self-attention mechanism to project the initial feature map into queries, keys, and values in the spatial dimension, calculate the similarity between the queries and keys in the spatial dimension through the Softmax function to obtain a spatial attention weight, and perform weighted aggregation on the values in the spatial dimension according to the spatial attention weight to obtain a spatially enhanced feature map.
[0028] In step 102, after performing an initial feature transformation on the spatially enhanced feature map using layer normalization, a multi-layer perceptron is used to perform a secondary feature transformation on the spatially enhanced feature map after the initial feature transformation, obtaining an intermediate feature map.
[0029] In step 103, the intermediate feature map after reshaping is projected into queries, keys, and values in the channel dimension using a channel self-attention mechanism. The similarity between the queries and keys in the channel dimension is calculated through a Softmax function to obtain channel attention weights. The values in the channel dimension are weighted and aggregated according to the channel attention weights, obtaining a channel-enhanced feature map.
[0030] In step 104, after performing an initial feature transformation on the channel-enhanced feature map using layer normalization, a multi-layer perceptron is used to perform a secondary feature transformation on the channel-enhanced feature map after the initial feature transformation, and a multi-scale feature fusion map is output.
[0031] Currently, infrared target detection algorithms are mainly divided into two categories: traditional algorithms and deep learning methods. Traditional algorithms such as background subtraction, morphological processing, and edge detection rely on handcrafted features and heuristic rules, and have high computational efficiency and real-time processing capabilities. However, they perform poorly in detecting targets in complex backgrounds, under changing lighting conditions, and small targets. Especially when the difference between the target and the background is small or the target size is small, the accuracy drops significantly. Deep learning methods can better handle complex backgrounds and changing environments through automatic feature extraction. However, in infrared small target detection, especially in the case of weak and small targets and complex backgrounds, they still face challenges and the detection accuracy needs to be improved.
[0032] The current technological development shows that although deep learning has brought new solutions to infrared small target detection, it still needs to be further optimized in terms of the effectiveness of feature extraction, multi-scale adaptability, and background suppression ability. Especially when dealing with small targets at long distances and low signal-to-noise ratios, there is still much room for improvement in the detection accuracy and stability of existing methods. This technical bottleneck restricts the practical application effect of infrared image detection systems in complex environments.
[0033] In the above embodiments of the present application, by mixing spatial and channel multi-attention during the infrared small target detection process, the detection accuracy and robustness of infrared small targets can be improved. Specifically, it can be applied to an infrared small target detection network, which consists of three cores: The spatial-channel dual-stream self-attention module adopts a unique dual-stream processing architecture, uses a cascade structure to synergistically enhance the spatial and channel feature representation capabilities, and at the same time introduces a decoder-guided dynamic feature recalibration mechanism to establish an encoder-decoder bidirectional interaction, significantly improving the detection accuracy of small targets in complex backgrounds while ensuring computational efficiency; The channel multi-head self-attention module for multi-scale information fusion realizes the deep integration and optimization of cross-scale channel features by constructing a hierarchical feature interaction mechanism; The feature recalibration self-attention module innovatively introduces an adaptive weight adjustment mechanism to effectively solve the semantic gap problem between the low-level features of the encoder and the high-level features of the decoder. The three cores work together to form a complete spatial-channel feature aggregation system, which can improve the detection accuracy and generalization performance for infrared small target detection in complex environments, providing an efficient and reliable technical solution for the infrared detection field.
[0034] Specifically, first, a 3×3 convolutional layer is used to perform feature embedding on the obtained infrared image to be detected, obtaining an initial feature map. The convolutional layer is the basic unit of a convolutional neural network (CNN) and is used to extract features from the input image. When the 3×3 convolutional kernel performs a convolution operation on the infrared image to be detected, it traverses each position of the image, calculates the dot product of the convolutional kernel and the local area of the image, adds a bias term, and then obtains the output through an activation function. The initial feature map is the result of the convolutional layer extracting features from the input image, retaining the spatial structure information of the image, and the pixel value at each position represents the feature intensity at that position. The initial feature map will be used as the input for subsequent network layers to prepare for further extraction and recognition of infrared small targets.
[0035] Next, the initial feature map is then input into an encoder stacked by four-level spatial-channel dual-stream self-attention modules. The spatial-channel dual-stream self-attention module consists of a spatial self-attention mechanism and a channel self-attention mechanism. Further, the initial feature map is projected into queries, keys, and values in the spatial dimension using the spatial self-attention mechanism. Then, the similarity between the queries and keys in the spatial dimension is calculated through the Softmax function to obtain the spatial attention weights, and the values in the spatial dimension are weighted and aggregated according to the spatial attention weights, thereby obtaining a spatially enhanced feature map. In the process of using the spatial self-attention mechanism, the spatial self-attention mechanism is a technique for enhancing feature representation. It reallocates the weights of features by calculating the similarity between different positions in the feature map (initial feature map), thereby highlighting important features and suppressing unimportant features. In the spatial self-attention mechanism, the initial feature map is projected into three matrices of queries, keys, and values in the spatial dimension. Among them, the query matrix is used to measure the importance of each position in the feature map, the key matrix is used to calculate the similarity with other positions, and the value matrix contains the feature information to be enhanced. The similarity between the queries and keys in the spatial dimension is calculated through the Softmax function to obtain the spatial attention weights. These weights can reflect the correlation between different positions in the feature map. The position with a larger weight indicates a higher correlation with other positions. By reallocating the weights of features according to the importance of different positions in the feature map, important features are enhanced and unimportant features are suppressed. The finally obtained spatially enhanced feature map not only retains the spatial structure information of the initial feature map but also enhances the feature representation ability through the spatial self-attention mechanism, improving the accuracy of subsequent object detection or recognition.
[0036] Next, after the spatially enhanced feature map is initially transformed using layer normalization, a multi-layer perceptron is used to perform a secondary feature transformation on the spatially enhanced feature map after the initial feature transformation to obtain an intermediate feature map.
[0037] In particular, layer normalization is a regularization technique used to improve the training stability and performance of neural networks. It normalizes the input of each layer of the neural network, making the input distribution of each layer relatively stable, thereby alleviating the problem of internal covariate shift, that is, the change in the input distribution between network layers. Layer normalization is used to standardize the output of each layer so that the mean of the output of each layer is 0 and the standard deviation is 1, which helps to speed up the training process and improve the convergence and stability of the model. The specific effects of layer normalization can include:
[0038] 1. Layer normalization can alleviate the problems of gradient vanishing or gradient explosion, making the training process more stable.
[0039] 2. Through normalization, the scale differences between different features are eliminated, which helps the gradient descent algorithm converge faster.
[0040] 3. Layer normalization does not depend on batch statistics, so it is suitable for processing variable-length sequences and dynamic model structures.
[0041] A Multilayer Perceptron (MLP) is a basic artificial neural network model with a multi-layer structure composed of multiple neurons. It is a feedforward neural network, that is, information propagates unidirectionally in the network, from the input layer through one or more hidden layers to the output layer. The MLP consists of multiple layers, and each layer contains multiple neurons for processing input data and performing transformations through non-linear activation functions. The specific functions of the MLP can include:
[0042] 1. The hidden layer of the MLP can automatically extract high-level features of the data, which are particularly important for complex pattern recognition and classification tasks.
[0043] 2. By introducing non-linear activation functions and a multi-layer structure, the MLP can fit complex non-linear relationships.
[0044] 3. The MLP is widely used in various classification, regression, and clustering tasks.
[0045] Specifically, layer normalization is used to perform an initial feature transformation on the spatially enhanced feature map, adjust its distribution to make it more suitable for subsequent processing by the MLP. Subsequently, the layer-normalized features are input into the MLP for a secondary feature transformation. The hidden layer of the MLP MLP transforms the input features through non-linear activation functions (such as ReLU, sigmoid, etc.) to extract higher-level features. After being processed by multiple hidden layers, the MLP finally outputs an intermediate feature map, and these features can be used for subsequent classification, regression, and other tasks. For this reason, the intermediate feature map is the spatially enhanced feature map after being processed by layer normalization and the MLP, containing richer feature information, and can provide more powerful support for subsequent object detection, recognition, and other tasks.
[0046] For this reason, through layer normalization and the MLP for feature transformation, layer normalization provides a good foundation for subsequent feature transformation by stabilizing the training process and improving the convergence speed. The MLP then further transforms the spatially enhanced feature map after the initial feature transformation through its powerful feature extraction and non-linear fitting capabilities to obtain a more representative intermediate feature map. In particular, appropriate network structures and parameter settings can also be selected according to specific tasks and data characteristics to achieve the best performance and effects.
[0047] Next, the intermediate feature map after the reshaping operation is projected into queries, keys, and values in the channel dimension using the channel self-attention mechanism. The similarity between the queries and keys in the channel dimension is calculated through the Softmax function to obtain the channel attention weights. The values in the channel dimension are weighted and aggregated according to the channel attention weights to obtain the channel-enhanced feature map. The reshaping operation is the reshape operation. In data processing and scientific computing, the reshape operation is used to adjust the dimensional structure of an array or matrix without changing its data content. Through the reshape operation, data can be transformed from one shape to another to meet different processing requirements.
[0048] Next, layer normalization is used again to perform an initial feature transformation on the channel-enhanced feature map ( ), and then a multi-layer perceptron is used to perform a secondary feature transformation on the channel-enhanced feature map after the initial feature transformation, and a multi-scale feature fusion map ( ) is output, that is, is processed through an independent LN-MLP module to output features with both spatial structure and channel discriminability .
[0049] Therefore, the entire processing chain uses a cascaded structure of "spatial attention - LN-MLP - channel attention - LN-MLP" to achieve hierarchical feature enhancement from local to global. Among them, the spatial branch focuses on pixel-level relationship modeling, and the channel branch strengthens semantic feature interaction. The two stages are stably connected through LN-MLP, significantly improving the feature representation ability while ensuring computational efficiency.
[0050] Optionally, for "calculating the similarity between the queries and keys in the spatial dimension through the Softmax function to obtain the spatial attention weights, and weighting and aggregating the values in the spatial dimension according to the spatial attention weights to obtain the spatial-enhanced feature map" in step 101, it specifically includes:
[0051] Step 1011, based on the Softmax function, the queries and keys in the spatial dimension, a spatial attention weight calculation formula is constructed, and the similarity between the queries and keys in the spatial dimension is calculated through the spatial attention weight calculation formula to obtain the spatial attention weights.
[0052] Step 1012, based on the spatial-enhanced feature map representation formula and the spatial attention weights, the values in the spatial dimension are weighted and aggregated to obtain the spatial-enhanced feature map, where the spatial attention weight calculation formula is:
[0053] ,
[0054] The spatial-enhanced feature map representation formula is:
[0055] ,
[0056] is the spatial attention weight, is the query in the spatial dimension, is the key in the spatial dimension, is the similarity in the spatial dimension, is the spatially enhanced feature map, is the value in the spatial dimension.
[0057] In the above embodiments of the present application, when the initial feature map is input into the infrared small target detection network, the input initial feature map is first processed by the spatial self-attention mechanism to enhance the features through the self-attention mechanism in the spatial dimension. After being processed by the spatial self-attention mechanism, the initial feature map obtains three key components through linear projection: the query in the spatial dimension ( ), the key ( ), and the value ( ). These three components are the core in the self-attention mechanism and are used to calculate the attention weight and aggregate information.
[0058] Next, the similarity between the query in the spatial dimension ( ) and the key ( ) is calculated through the Softmax function to obtain the spatial attention weight . represents the correlation between different spatial positions and is used to guide the aggregation of information.
[0059] Finally, the value in the spatial dimension ( ) is weighted and aggregated using the spatial attention weight to obtain the spatially enhanced feature map . is the output after being enhanced by the spatial self-attention mechanism and contains more information about the spatial dimension.
[0060] Optionally, in step 102, the "intermediate feature map" is specifically represented as:
[0061] ,
[0062] is the intermediate feature map, is the multi-layer perceptron, is the layer normalization, is the spatially enhanced feature map.
[0063] Optionally, for the description in step 103 "calculating the similarity between the query and the key in the channel dimension through the Softmax function to obtain the channel attention weight, and performing weighted aggregation on the values in the channel dimension according to the channel attention weight to obtain the channel-enhanced feature map", it specifically includes:
[0064] Step 1031, construct a channel attention weight calculation formula based on the Softmax function, the query in the channel dimension, and the key, and calculate the similarity between the query and the key in the channel dimension through the channel attention weight calculation formula to obtain the channel attention weight.
[0065] Step 1032, perform weighted aggregation on the values in the channel dimension based on the channel-enhanced feature map representation formula and the channel attention weight to obtain the channel-enhanced feature map, where the channel attention weight calculation formula is:
[0066] ,
[0067] The channel-enhanced feature map representation formula is:
[0068] ,
[0069] is the channel attention weight, is the query in the channel dimension, is the key in the channel dimension, is the similarity in the channel dimension, is the channel-enhanced feature map, is the value in the channel dimension.
[0070] In the above embodiments of the present application, the following operations can also be performed: For example, the tensor of the input feature map X (the intermediate feature map after reshaping operation) is , where C is the number of channels, H and W are the height and width respectively, and the query and the key are obtained by performing a linear transformation on the input feature map X, that is:
[0071] and , where and are the learnable weight matrices of the query and the key in the channel dimension respectively.
[0072] Then, perform a dot product operation on the query and the key in the channel dimension to obtain the similarity matrix To make the similarity matrix interpretable, for Perform scaling, i.e.:
[0073] ,
[0074] Next, apply the Softmax function to the scaled similarity matrix to obtain the channel attention weight matrix . The Softmax function is used to ensure that the sum of the elements in the weight matrix is 1, thereby representing the relative importance between different channels.
[0075] For this reason, the channel attention weight matrix is obtained. Each element in it represents the similarity between the query and the key on the specific channel in the channel dimension, thus reflecting the importance of this channel in the overall feature representation.
[0076] The value in the channel dimension is also obtained by performing a linear transformation on the input feature map X, i.e.:
[0077] , where is the learnable weight matrix of the value in the channel dimension. Use the channel attention weight matrix to perform weighted aggregation on the value in the channel dimension to obtain the channel-enhanced feature map, that is, . In particular, the above weighted aggregation process adjusts the feature representation according to the importance of each channel, thereby enhancing the response to important features and suppressing the response to unimportant features.
[0078] Step 105, the decoder performs reverse decoding on the multi-scale feature fusion map to obtain a multi-scale feature fusion infrared image, and uses an adaptive weight adjustment mechanism to perform feature recalibration on the multi-scale feature fusion infrared image to generate a binary detection map.
[0079] Next, the decoder performs reverse decoding on the multi-scale feature fusion map to obtain a multi-scale feature fusion infrared image, and uses an adaptive weight adjustment mechanism to perform feature recalibration on the multi-scale feature fusion infrared image to be detected, generating a binary detection map. Specifically, after using the encoder to extract multi-scale features, the extracted features contain context information and detail information of different scales. The multi-scale features are fused to obtain a fused feature map (multi-scale feature fusion map). The fusion method can adopt simple splicing, addition, or more complex attention mechanisms, etc. Then, the decoder performs reverse decoding on the fused multi-scale feature fusion map and gradually restores it to the image space to obtain a multi-scale feature fusion infrared image. Operations such as upsampling and convolution can be used during the decoding process to gradually restore the resolution and details of the image. In particular, the decoder is also designed with an adaptive weight adjustment mechanism that can dynamically adjust the importance of different features according to the image content. Specifically, methods such as attention mechanisms, channel attention, and spatial attention can be used to achieve adaptive weight adjustment. For example, the Softmax function is used to calculate the channel attention weights, and the channel features are weighted and summed according to the weights to achieve feature recalibration. Finally, the image after feature recalibration is binarized to obtain a binary detection map. During the binarization process, the binarization threshold can be adjusted according to specific application scenarios and requirements, or a simple threshold segmentation method can be used, or a more complex image segmentation algorithm can also be used.
[0080] Therefore, through the multi-scale fusion in the early stage of the decoder, the decoder can make full use of feature information of different scales, capture more details and context information, thereby improving the accuracy of target detection. The adaptive weight adjustment mechanism can dynamically adjust the importance of features according to the image content, further highlighting the target features and suppressing background interference. The multi-scale feature fusion map can handle targets of different scales and shapes, improving the robustness of the algorithm. The adaptive weight adjustment mechanism enables the algorithm to adapt to different scenarios and lighting conditions, reducing false detections and missed detections. Through feature recalibration, the algorithm can generate a clearer and more accurate binary detection map, facilitating subsequent target recognition and analysis.
[0081] Optionally, before the step 105 "the decoder performs reverse decoding on the multi-scale feature fusion map to obtain a multi-scale feature fusion infrared image", specifically, it further includes:
[0082] Step 106: The multi-scale information fusion channel multi-head self-attention module converts the multi-scale feature fusion map into a unified feature sequence through block partitioning operations, and through an improved residual multi-head attention mechanism, uses the feature maps output by each level of the spatial-channel two-stream self-attention module as independent queries and shared key-value pairs for interaction to obtain cross-scale global channel dependencies. After integrating the cross-scale global channel dependencies into the unified feature sequence, it is input into the decoder, where the feature maps include the spatially enhanced feature map, the intermediate feature map, the channel-enhanced feature map, and the multi-scale feature fusion map.
[0083] In the above embodiment of the present application, the skip connection part in the infrared small target detection network adopts a multi-scale information fusion channel multi-head self-attention module for processing the outputs of each level of the encoder. This module first converts the multi-scale feature fusion map output by the encoder into a feature sequence with a unified representation (unified feature sequence) through block partitioning operations, while keeping the original channel dimension unchanged. Subsequently, through an improved residual multi-head attention mechanism, each scale feature is used as an independent query and shared key-value pair for interaction to establish cross-scale global channel dependencies. Finally, the processed feature sequence is restored to its original shape and fused with the input features. This module innovatively designs a special attention calculation method to effectively fuse local details and global context information, and realizes deep feature interaction through a multi-layer stacked Transformer (self-attention mechanism) structure. This design significantly improves the model's multi-scale feature representation ability for infrared small targets, especially showing excellent performance when dealing with targets of different sizes and complex background interference, providing an important guarantee for improving the detection accuracy.
[0084] Optionally, the decoder includes a three-level feature recalibration self-attention module (as Figure 5 shown) and a cross-level attention mechanism. The decoder in step 105 performs reverse decoding on the multi-scale feature fusion map to obtain a multi-scale feature fusion infrared image, and uses an adaptive weight adjustment mechanism to perform feature recalibration on the multi-scale feature fusion infrared image to generate a binary detection map, specifically including:
[0085] Step 1051: The decoder performs reverse decoding on the input unified feature sequence incorporating cross-scale global channel dependencies to obtain a multi-scale feature fusion infrared image.
[0086] Step 1052: When the decoder uses each level of the feature recalibration self-attention module to reverse decode the output features, after performing feature transformation on the output features of each level of the feature recalibration self-attention module through layer normalization, the output features after feature transformation are converted into queries, and the multi-scale feature fusion map is mapped into keys and values.
[0087] Step 1053: Calculate the similarity between the query and the key using the cross-level attention mechanism, construct a multi-scale feature correlation matrix based on the calculated similarity, and combine the adaptive weight allocation strategy in the channel dimension. After enhancing by determining the most relevant multi-scale feature fusion map for each output feature, generate a binary detection map, where the binary detection map is generated after being processed by a 1×1 convolutional layer and a sigmoid function.
[0088] In the above embodiments of the present application, the decoder performs reverse decoding on the input unified feature sequence incorporating cross-scale global channel dependencies to obtain a multi-scale feature fusion infrared image. More specifically, the input of the decoder consists of two parts: the output features of each level of the feature recalibration self-attention module in the decoder branch, and the output features of each level of the spatial-channel dual-stream self-attention module in the encoder branch. During the decoder decoding process, a decoder-guided dynamic feature recalibration mechanism is adopted, and its core lies in using high-level semantic information to optimize the low-level feature representation. In the feature transformation stage, first, the high-level semantic features output by the decoder are processed through a LayerNorm layer (layer normalization) and converted into query vectors (Query), while the multi-level spatial features extracted by the encoder are respectively mapped into key-value pairs (Key-Value). This dual-path transformation design not only retains the inherent characteristics of different-level features but also lays a foundation for subsequent feature interaction. In the attention modeling stage, a cross-level attention mechanism is adopted to establish dynamic associations between features. This mechanism calculates the similarity between the query vector and the key vector, constructs a multi-scale feature correlation matrix, and combines the adaptive weight allocation strategy in the channel dimension, enabling each decoder feature to intelligently select the most relevant encoder feature for enhancement. This decoder-guided dynamic feature recalibration mechanism realizes the bidirectional interaction between the encoder and decoder features: the decoder features act as query signals to guide the dynamic adjustment of the encoder features, while the cross-attention mechanism adaptively calculates the correlation between features and can focus on the most discriminative information. This design significantly improves the adaptability to multi-scale targets and exhibits excellent robustness in complex scenarios.
[0089] By introducing a cross-level attention mechanism in the decoding stage, high-level semantic information is used to guide the optimization of low-level features. By dynamically calculating the correlation between the query vector (decoder features) and the key-value pairs (encoder multi-level features), the most discriminative information is adaptively selected for enhancement, significantly improving the localization accuracy of small targets in complex backgrounds, which is superior to the static feature fusion strategy.
[0090] Specifically, the final output of the infrared small target detection network is processed by a 1×1 convolution and a sigmoid function to generate a binary detection map. In the entire data processing flow, the encoder is responsible for gradually extracting and condensing feature information, the skip connections ensure the complete transmission of multi-scale features, and the decoder achieves precise positioning through re-calibration and feature fusion. Residual connections are used for feature transmission between modules, effectively alleviating the problem of gradient disappearance. This carefully designed architecture enables the network to maintain real-time performance while improving the detection performance of small targets in complex backgrounds.
[0091] Step 107, the infrared small target detection module uses the binary detection map to detect infrared small targets.
[0092] In the above embodiments of the present application, the infrared small target detection module uses the binary detection map to detect infrared small targets. Therefore, infrared small target detection is achieved through the infrared small target detection network.
[0093] Optionally, the infrared small target detection module in step 107 uses the binary detection map to detect infrared small targets, specifically including:
[0094] Step 1071, the infrared small target detection module determines the detection threshold of the binary detection map based on the maximum inter-class variance algorithm. After threshold segmentation of the binary detection map based on the detection threshold, the part with pixel values higher than the detection threshold is determined as an infrared small target.
[0095] In the above embodiments of the present application, a suitable threshold can be set according to the characteristics of the binary detection map. The threshold is used to distinguish infrared small targets from background noise. Specifically, the maximum inter-class variance algorithm can be used to determine the detection threshold, and then threshold segmentation is performed on the binary detection map, that is, the part with pixel values higher than the detection threshold is regarded as an infrared small target, and the part lower than the detection threshold is regarded as the background.
[0096] Furthermore, connected component analysis can be performed on the image after threshold segmentation to identify the connected regions in the image. The connected regions correspond to the infrared small targets accordingly, because small targets often appear as continuous pixel blocks in the image. According to the characteristics such as the size and shape of the connected components, the true infrared small targets are further screened and extracted, while possible false detections (such as noise, artifacts, etc.) are removed, and the real targets are retained. The extracted infrared small targets can also be located to determine their specific positions and ranges in the image. For example, bounding boxes, centroid coordinates, etc. can be used to mark the positions of the targets.
[0097] Specifically, further post-processing can also be performed on the detection results according to specific application scenarios and requirements. For example, operations such as target tracking, target classification, and target counting can be carried out.
[0098] Finally, the detection results can be visualized, such as overlaying the target position on the original image or comparing with the ground truth labels or manual detection results to verify the accuracy and reliability of the detection algorithm. Therefore, the binary detection map is used to effectively detect infrared small targets.
[0099] By applying the technical solution of this embodiment, to solve the problem of insufficient detection accuracy in the current infrared target detection method based on deep learning, in the process of infrared small target detection, hybrid spatial-channel multi-attention is mixed, a hierarchical feature processing flow is adopted, and multi-dimensional optimization of features is achieved through a novel attention mechanism. It can be applied to the detection process of complex backgrounds and small targets. Compared with traditional deep learning methods, the above embodiment of the present application has higher efficiency and accuracy in the infrared target detection task, has a good application prospect, and can play an important role in the field of infrared image processing.
[0100] Based on the above as Figures 1 to 5 shown in the method, correspondingly, the embodiment of the present application also provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the infrared small target detection method as Figures 1 to 5 shown above.
[0101] Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present application.
[0102] Based on the above as Figures 1 to 5 shown in the method, to achieve the above object, the embodiment of the present application also provides a computer device, which can specifically be a personal computer, a server, a network device, etc. The computer device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the infrared small target detection method as Figures 1 to 5 shown above.
[0103] Optionally, the computer device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), etc.
[0104] Those skilled in the art will appreciate that the computer device structure provided in this embodiment does not limit the computer device, and may include more or fewer components, or a combination of certain components, or different component arrangements.
[0105] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages and saves the hardware and software resources of the computer device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to realize communication between the components inside the storage medium, and communication with other hardware and software in the physical device.
[0106] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented by means of software plus the necessary general hardware platform, or by hardware through the cascade processing architecture of the spatial channel dual-stream self-attention module (spatial attention, LN, MLP, channel attention, LN, MLP), to achieve hierarchical feature enhancement from local to global. The spatial branch focuses on pixel-level relationship modeling, and the channel branch strengthens the interaction of semantic features. The collaborative optimization of the two significantly improves the feature representation capability and overcomes the problem of insufficient feature expression capability of traditional methods in complex backgrounds. The jump connection part adopts an improved channel multi-head self-attention module with multi-scale information fusion, unifies the multi-scale feature representation through block partitioning operations, and uses the residual multi-head attention mechanism to establish cross-scale global dependencies. Compared with traditional convolution or simple splicing methods, it more effectively integrates local details and global context information, and enhances the adaptability to targets of different sizes.
[0107] Those skilled in the art will appreciate that the accompanying drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the accompanying drawings are not necessarily necessary for implementing the present application. Those skilled in the art will appreciate that the modules in the devices in the implementation scenario can be distributed in the devices of the implementation scenario according to the description of the implementation scenario, or can be changed accordingly and located in one or more devices different from the present implementation scenario. The modules of the above-mentioned implementation scenario can be combined into one module, or can be further split into multiple submodules.
[0108] The above serial numbers of this application are only for description and do not represent the advantages and disadvantages of the implementation scenarios. The above disclosure is only a few specific implementation scenarios of this application, but this application is not limited to them, and any changes that can be made by technicians in this field should fall within the scope of protection of this application.
Claims
1. An infrared small target detection method, characterized in that, Applied to an infrared small target detection network, the infrared small target detection network includes a channel multi-head self-attention module for multi-scale information fusion connected with a layer normalization and multi-layer perceptron joint module, an infrared small target detection module, and an encoder and a decoder with a reciprocal structure. The encoder is a cascaded processing architecture composed of four-level spatial-channel dual-stream self-attention modules. Each level of the spatial-channel dual-stream self-attention module includes a spatial self-attention mechanism and a channel self-attention mechanism. The layer normalization and multi-layer perceptron joint module includes layer normalization and a multi-layer perceptron. The method includes: Obtain a to-be-detected infrared image containing infrared small targets. After using a 3×3 convolutional layer to perform feature embedding on the to-be-detected infrared image to obtain an initial feature map, input the initial feature map into the infrared small target detection network to implement infrared small target detection by the infrared small target detection network to perform the following steps: Use the spatial self-attention mechanism to project the initial feature map into queries, keys, and values in the spatial dimension, calculate the similarity between the queries and keys in the spatial dimension through the Softmax function to obtain spatial attention weights, and perform weighted aggregation on the values in the spatial dimension according to the spatial attention weights to obtain a spatially enhanced feature map; After using layer normalization to perform a primary feature transformation on the spatially enhanced feature map, use a multi-layer perceptron to perform a secondary feature transformation on the spatially enhanced feature map after the primary feature transformation to obtain an intermediate feature map; Use the channel self-attention mechanism to project the reshaped intermediate feature map into queries, keys, and values in the channel dimension, calculate the similarity between the queries and keys in the channel dimension through the Softmax function to obtain channel attention weights, and perform weighted aggregation on the values in the channel dimension according to the channel attention weights to obtain a channel-enhanced feature map; After using layer normalization to perform a primary feature transformation on the channel-enhanced feature map, use a multi-layer perceptron to perform a secondary feature transformation on the channel-enhanced feature map after the primary feature transformation to output a multi-scale feature fusion map; The channel multi-head self-attention module for multi-scale information fusion converts the multi-scale feature fusion map into a unified feature sequence through a block partitioning operation, and through an improved residual multi-head attention mechanism, uses the feature maps output by each level of the spatial-channel dual-stream self-attention module as independent queries and shared key-value pairs for interaction to obtain cross-scale global channel dependencies. After integrating the cross-scale global channel dependencies into the unified feature sequence, input it into the decoder, where the feature maps include the spatially enhanced feature map, the intermediate feature map, the channel-enhanced feature map, and the multi-scale feature fusion map; The decoder performs reverse decoding on the input unified feature sequence integrated with cross-scale global channel dependencies to obtain a multi-scale feature fusion infrared image, and uses an adaptive weight adjustment mechanism to perform feature recalibration on the multi-scale feature fusion infrared image to generate a binary detection map; The infrared small target detection module uses the binary detection map to detect infrared small targets.
2. The method according to claim 1, wherein The decoder includes a three-level feature recalibration self-attention module and a cross-level attention mechanism, and uses an adaptive weight adjustment mechanism to perform feature recalibration on the multi-scale feature fusion infrared image to generate a binary detection map, including: When the decoder uses each level of the feature recalibration self-attention module to decode the output features in reverse, after performing feature transformation on the output features of each level of the feature recalibration self-attention module through layer normalization, the output features after feature transformation are converted into queries, and the multi-scale feature fusion map is mapped into keys and values; The cross-level attention mechanism is used to calculate the similarity between the query and the key, construct a multi-scale feature correlation matrix based on the calculated similarity, and combine the adaptive weight allocation strategy in the channel dimension to determine the most relevant multi-scale feature fusion map for each output feature for enhancement, and then generate a binary detection map.
3. The method according to claim 2, wherein The binary detection map is generated after being processed by a 1×1 convolutional layer and a sigmoid function.
4. The method according to claim 1, wherein The similarity between the query and the key in the spatial dimension is calculated through the Softmax function to obtain the spatial attention weight, and the values in the spatial dimension are weighted and aggregated according to the spatial attention weight to obtain a spatially enhanced feature map, including: Based on the Softmax function, the query and the key in the spatial dimension, a spatial attention weight calculation formula is constructed, and the similarity between the query and the key in the spatial dimension is calculated through the spatial attention weight calculation formula to obtain the spatial attention weight; Based on the spatially enhanced feature map representation formula and the spatial attention weight, the values in the spatial dimension are weighted and aggregated to obtain a spatially enhanced feature map, where the spatial attention weight calculation formula is: , The spatially enhanced feature map representation formula is: , is the spatial attention weight, is the query in the spatial dimension, is the key in the spatial dimension, is the similarity in the spatial dimension, is the spatially enhanced feature map, is the value in the spatial dimension.
5. The method according to claim 1, wherein The similarity between the query and the key in the channel dimension is calculated through the Softmax function to obtain the channel attention weight, and the values in the channel dimension are weighted and aggregated according to the channel attention weight to obtain a channel-enhanced feature map, including: Based on the Softmax function, the query and the key in the channel dimension, a channel attention weight calculation formula is constructed, and the similarity between the query and the key in the channel dimension is calculated through the channel attention weight calculation formula to obtain the channel attention weight; Based on the channel-enhanced feature map representation formula and the channel attention weight, the values in the channel dimension are weighted and aggregated to obtain a channel-enhanced feature map, where the channel attention weight calculation formula is: , The channel-enhanced feature map representation formula is: , is the channel attention weight, is the query on the channel dimension, is the key on the channel dimension, is the similarity on the channel dimension, is the channel enhanced feature map, is the value on the channel dimension.
6. The method according to any one of claims 1 to 5, characterized in that The intermediate feature map is represented as: , is the intermediate feature map, is the multi-layer perceptron, is the layer normalization, is the spatially enhanced feature map.
7. The method according to claim 6, wherein The infrared small target detection module uses the binary detection map to detect infrared small targets, including: The infrared small target detection module determines the detection threshold of the binary detection map based on the maximum inter-class variance algorithm, and after performing threshold segmentation on the binary detection map based on the detection threshold, the part with pixel values higher than the detection threshold is determined as an infrared small target.
8. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the infrared small target detection method according to any one of claims 1 to 7.
9. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the infrared small target detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Infrared and visible light fusion method based on multi-scale feature interaction enhancement
CN119091269A