Infrared small target detection method, storage medium and computer equipment

By using the spatial channel dual-flow self-attention module and the channel multi-head self-attention module that integrates multi-scale information in infrared small object detection, the problem of insufficient feature expression ability in complex backgrounds is solved, and the high accuracy and robustness of infrared small object detection is achieved.

CN120047757AActive Publication Date: 2025-05-27XIAN ORDNANCE IND TECH IND DEV CO LTD

Patent Information

Application Number
CN202510534573.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-27
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing infrared small object detection methods lack the feature expression ability under complex backgrounds, resulting in insufficient detection accuracy and stability.

Method used

The cascading processing architecture of the space channel dual-flow self-attention module is adopted, and the feature representation capability is optimized through the spatial self-attention and channel self-attention mechanism, and the channel multi-head self-attention module is effectively integrated with local details and global context information through the multi-scale information fusion channel multi-head self-attention module.

Benefits of technology

It significantly improves the detection accuracy and robustness of infrared small object detection, and overcomes the problem of insufficient feature expression ability in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047757A_ABST
    Figure CN120047757A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of infrared image detection and computers, and discloses an infrared small target detection method, a storage medium and computer equipment, and the method comprises the steps: achieving the local-to-global hierarchical feature enhancement through a cascading processing architecture of a space channel double-flow self-attention module. The spatial branch focuses on pixel-level relation modeling, the channel branch strengthens semantic feature interaction, the feature representation capability is improved through collaborative optimization of the spatial branch and the channel branch, and the problem that the feature representation capability is insufficient in a complex background in a traditional method is solved. The jump connection part adopts an improved multi-scale information fusion channel multi-head self-attention module, multi-scale feature representation is unified through block division operation, and a cross-scale global dependency relationship is established by using a residual multi-head attention mechanism. Compared with a traditional convolution or simple splicing method, local details and global context information are fused more effectively, and the adaptability to targets of different sizes is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of infrared image detection and computer technology, and in particular, to an infrared small target detection method, a storage medium, and a computer device. Background Art

[0002] In the technical field of infrared image detection, small target detection, as a key technology, plays an important role in application scenarios such as national defense security, intelligent monitoring, and autonomous driving. Infrared image detection, with its all-weather working ability and sensitivity to temperature differences, demonstrates unique advantages under low illumination and adverse weather conditions, becoming an important supplement to visible light imaging.

[0003] Traditional small target detection methods are mainly based on image processing techniques, including but not limited to: differential methods based on background modeling, enhancement algorithms based on morphological operations, and detection techniques based on edge features. These methods achieve target separation by constructing a background model, enhance target features using mathematical morphological operations, or extract the target contour with the help of edge operators. However, when dealing with actual scenarios such as complex background interference and low target signal-to-noise ratio, these methods often suffer from problems such as insufficient detection stability and high false alarm rates.

[0004] In recent years, with the rapid development of deep learning technology, object detection algorithms based on convolutional neural networks have gradually become the mainstream. Two-stage and one-stage detectors represented by Faster R-CNN (Fast Region-based Convolutional Neural Network) and YOLO (You Only Look Once) have significantly improved detection performance through end-to-end feature learning and object localization. However, for the special task of infrared small target detection, existing deep learning methods still face many challenges: First, the limited pixel information of small targets leads to insufficient feature expression ability; second, the low contrast between the target and the background increases the difficulty of feature discrimination; in addition, background interference in complex environments also affects the reliability of detection. Summary of the Invention

[0005] In view of this, the present application provides an infrared small target detection method, a storage medium, and a computer device. Through the cascaded processing architecture of the spatial channel dual-stream self-attention module (spatial attention, LN, MLP, channel attention, LN, MLP), hierarchical feature enhancement from local to global is achieved. The spatial branch focuses on pixel-level relationship modeling, and the channel branch strengthens semantic feature interaction. The two cooperate and optimize to significantly improve the feature representation ability, overcoming the problem of insufficient feature expression ability of traditional methods in complex backgrounds. The skip connection part adopts an improved multi-scale information fusion channel multi-head self-attention module, which unifies multi-scale feature representations through block division operations and establishes cross-scale global dependence relationships using a residual multi-head attention mechanism. Compared with traditional convolution or simple splicing methods, it more effectively fuses local details and global context information and enhances the adaptability to targets of different sizes.

[0006] According to one aspect of the present application, there is provided an infrared small target detection method, which is applied to an infrared small target detection network. The infrared small target detection network includes a channel multi-head self-attention module for multi-scale information fusion connected to a layer normalization and multi-layer perceptron joint module, an infrared small target detection module, and an encoder and a decoder with a reciprocal structure. The encoder is a cascaded processing architecture and is composed of four-level spatial channel dual-stream self-attention modules. Each level of the spatial channel dual-stream self-attention module includes a spatial self-attention mechanism and a channel self-attention mechanism. The layer normalization and multi-layer perceptron joint module includes layer normalization and a multi-layer perceptron. The method includes: Obtain a to-be-detected infrared image containing an infrared small target. After using a 3×3 convolutional layer to perform feature embedding on the to-be-detected infrared image to obtain an initial feature map, input the initial feature map into the infrared small target detection network to implement infrared small target detection by the infrared small target detection network to perform the following steps: Use the spatial self-attention mechanism to project the initial feature map into queries, keys, and values in the spatial dimension, calculate the similarity between the queries and keys in the spatial dimension through the Softmax function to obtain spatial attention weights, and perform weighted aggregation on the values in the spatial dimension according to the spatial attention weights to obtain a spatially enhanced feature map; After using layer normalization to perform a primary feature transformation on the spatially enhanced feature map, use a multi-layer perceptron to perform a secondary feature transformation on the spatially enhanced feature map after the primary feature transformation to obtain an intermediate feature map; Use the channel self-attention mechanism to project the reshaped intermediate feature map into queries, keys, and values in the channel dimension, calculate the similarity between the queries and keys in the channel dimension through the Softmax function to obtain channel attention weights, and perform weighted aggregation on the values in the channel dimension according to the channel attention weights to obtain a channel-enhanced feature map; After performing an initial feature transformation on the channel-enhanced feature map using layer normalization, a multi-layer perceptron is used to perform a secondary feature transformation on the channel-enhanced feature map after the initial feature transformation, and a multi-scale feature fusion map is output; The decoder performs reverse decoding on the multi-scale feature fusion map to obtain a multi-scale feature fusion infrared image, and uses an adaptive weight adjustment mechanism to perform feature recalibration on the multi-scale feature fusion infrared image to generate a binary detection map; The infrared small target detection module uses the binary detection map to detect infrared small targets.

[0007] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above infrared small target detection method is implemented.

[0008] According to still another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the program, the above infrared small target detection method is implemented.

[0009] By means of the above technical solution, an infrared small target detection method, a storage medium, and a computer device provided by the present application can improve the detection accuracy and robustness of infrared small targets by mixing spatial-channel multi-attention during the infrared small target detection process.

[0010] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. Description of the Drawings

[0011] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings: Figure 1 A schematic diagram of an infrared small target detection network used in an infrared small target detection method provided by an embodiment of the present application is shown; Figure 2 A schematic diagram of a spatial-channel dual-stream self-attention module used in an infrared small target detection method provided by an embodiment of the present application is shown; Figure 3 A schematic diagram of a multi-scale information fusion channel multi-head self-attention module used in an infrared small target detection method provided by an embodiment of the present application is shown; Figure 4Shows a schematic flow diagram of an infrared small target detection method provided by an embodiment of the present application; Figure 5 Shows a schematic diagram of a feature recalibration self-attention module used in an infrared small target detection method provided by an embodiment of the present application. Detailed implementation manners

[0012] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.

[0013] In this embodiment, an infrared small target detection method is provided, which is applied to an infrared small target detection network, such as Figure 1 shown, the infrared small target detection network includes a channel multi-head self-attention module for multi-scale information fusion connected to a layer normalization and multi-layer perceptron joint module (such as Figure 2 shown), an infrared small target detection module, and an encoder and a decoder with a reciprocal structure. The encoder is a cascaded processing architecture and is composed of four-level spatial-channel dual-stream self-attention modules. Each level of the spatial-channel dual-stream self-attention module is as Figure 3 shown, and respectively includes a spatial self-attention mechanism and a channel self-attention mechanism. The layer normalization and multi-layer perceptron joint module includes layer normalization and a multi-layer perceptron. The method includes: Obtain a to-be-detected infrared image containing an infrared small target. After using a 3×3 convolutional layer to perform feature embedding on the to-be-detected infrared image to obtain an initial feature map, input the initial feature map into the infrared small target detection network to perform the following steps shown in reference Figure 4 to achieve infrared small target detection: Step 101, use the spatial self-attention mechanism to project the initial feature map into queries, keys, and values in the spatial dimension, calculate the similarity between the queries and keys in the spatial dimension through the Softmax function to obtain a spatial attention weight, and perform weighted aggregation on the values in the spatial dimension according to the spatial attention weight to obtain a spatially enhanced feature map.

[0014] Step 102, after using layer normalization to perform a primary feature transformation on the spatially enhanced feature map, use a multi-layer perceptron to perform a secondary feature transformation on the spatially enhanced feature map after the primary feature transformation to obtain an intermediate feature map.

[0015] Step 103, use the channel self-attention mechanism to project the reshaped intermediate feature map into queries, keys, and values in the channel dimension, calculate the similarity between the queries and keys in the channel dimension through the Softmax function to obtain a channel attention weight, and perform weighted aggregation on the values in the channel dimension according to the channel attention weight to obtain a channel-enhanced feature map.

[0016] In step 104, after performing an initial feature transformation on the channel-enhanced feature map using layer normalization, a multi-layer perceptron is used to perform a secondary feature transformation on the channel-enhanced feature map after the initial feature transformation, and a multi-scale feature fusion map is output.

[0017] Currently, infrared target detection algorithms are mainly divided into two categories: traditional algorithms and deep learning methods. Traditional algorithms such as background subtraction, morphological processing, and edge detection rely on handcrafted features and heuristic rules, and have high computational efficiency and real-time processing capabilities. However, they perform poorly in complex backgrounds, lighting changes, and small target detection. Especially when the difference between the target and the background is small or the target size is small, the accuracy drops significantly. Deep learning methods can better handle complex backgrounds and changing environments through automatic feature extraction. However, in infrared small target detection, especially in weak and small targets and complex backgrounds, they still face challenges and the detection accuracy needs to be improved.

[0018] The current technological development shows that although deep learning has brought new solutions to infrared small target detection, it still needs to be further optimized in terms of the effectiveness of feature extraction, multi-scale adaptability, and background suppression ability. Especially when dealing with small targets at long distances and low signal-to-noise ratios, there is still much room for improvement in the detection accuracy and stability of existing methods. This technical bottleneck restricts the actual application effect of infrared image detection systems in complex environments.

[0019] In the above embodiments of the present application, by mixing spatial-channel multi-attention during the infrared small target detection process, the detection accuracy and robustness of infrared small targets can be improved. Specifically, it can be applied to an infrared small target detection network, which includes three cores: The spatial-channel dual-stream self-attention module adopts a unique dual-stream processing architecture, uses a cascade structure to synergistically enhance the spatial and channel feature representation capabilities, and at the same time introduces a decoder-guided dynamic feature recalibration mechanism to establish an encoder-decoder bidirectional interaction, significantly improving the detection accuracy of weak and small targets in complex backgrounds while ensuring computational efficiency; The channel multi-head self-attention module for multi-scale information fusion realizes the in-depth integration and optimization of cross-scale channel features by constructing a hierarchical feature interaction mechanism; The feature recalibration self-attention module innovatively introduces an adaptive weight adjustment mechanism to effectively solve the semantic gap problem between the low-level features of the encoder and the high-level features of the decoder. The three cores work together to form a complete spatial-channel feature aggregation system, which can improve the detection accuracy and generalization performance for infrared small target detection in complex environments, providing an efficient and reliable technical solution for the infrared detection field.

[0020] Specifically, first, a 3×3 convolutional layer is used to perform feature embedding on the acquired infrared image to be detected, obtaining an initial feature map. The convolutional layer is the basic component of a convolutional neural network (CNN) and is used to extract features from the input image. When the 3×3 convolutional kernel performs a convolution operation on the infrared image to be detected, it traverses each position of the image, calculates the dot product of the convolutional kernel and the local area of the image, adds a bias term, and then obtains the output through an activation function. The initial feature map is the result of the convolutional layer extracting features from the input image, retaining the spatial structure information of the image, and the pixel value at each position represents the feature intensity at that position. The initial feature map will be used as the input for the subsequent network layers to prepare for further extraction and recognition of small infrared targets.

[0021] Next, the initial feature map is then input into an encoder stacked by four-level spatial-channel dual-stream self-attention modules. The spatial-channel dual-stream self-attention module consists of a spatial self-attention mechanism and a channel self-attention mechanism. Further, the initial feature map is projected into queries, keys, and values in the spatial dimension using the spatial self-attention mechanism. Then, the similarity between the queries and keys in the spatial dimension is calculated through the Softmax function to obtain the spatial attention weights, and the values in the spatial dimension are weighted and aggregated according to the spatial attention weights, thereby obtaining a spatially enhanced feature map. In the process of using the spatial self-attention mechanism, the spatial self-attention mechanism is a technique for enhancing feature representation. It reallocates the weights of features by calculating the similarity between different positions in the feature map (initial feature map), thereby highlighting important features and suppressing unimportant features. In the spatial self-attention mechanism, the initial feature map is projected into three matrices of queries, keys, and values in the spatial dimension. Among them, the query matrix is used to measure the importance of each position in the feature map, the key matrix is used to calculate the similarity with other positions, and the value matrix contains the feature information to be enhanced. The similarity between the queries and keys in the spatial dimension is calculated through the Softmax function to obtain the spatial attention weights. These weights can reflect the correlation between different positions in the feature map. The position with a larger weight indicates a higher correlation with other positions. By reallocating the weights of features according to the importance of different positions in the feature map, important features are enhanced and unimportant features are suppressed. The finally obtained spatially enhanced feature map not only retains the spatial structure information of the initial feature map but also enhances the feature representation ability through the spatial self-attention mechanism, improving the accuracy of subsequent object detection or recognition.

[0022] Then, after initially transforming the spatially enhanced feature map using layer normalization, a multi-layer perceptron is used to perform a secondary feature transformation on the spatially enhanced feature map after the initial feature transformation, obtaining an intermediate feature map.

[0023] Specifically, Layer Normalization is a regularization technique used to improve the training stability and performance of neural networks. It normalizes the input of each layer in the neural network, making the input distribution of each layer relatively stable, thereby alleviating the problem of internal covariate shift, that is, the change in the input distribution between network layers. Layer normalization is used to standardize the output of each layer so that the mean of the output of each layer is 0 and the standard deviation is 1, which helps to speed up the training process and improve the convergence and stability of the model. The specific effects of layer normalization can include: 1. Layer normalization can alleviate the problems of gradient vanishing or gradient explosion, making the training process more stable.

[0024] 2. Through normalization, the scale differences between different features are eliminated, which helps the gradient descent algorithm to converge faster.

[0025] 3. Layer normalization does not depend on batch statistics, so it is suitable for processing variable-length sequences and dynamic model structures.

[0026] A Multilayer Perceptron (MLP) is a basic artificial neural network model consisting of multiple layers of neurons. It is a feedforward neural network, that is, information propagates unidirectionally in the network, from the input layer through one or more hidden layers to the output layer. The MLP consists of multiple layers, each layer containing multiple neurons, which are used to process the input data and transform it through a non-linear activation function. The specific effects of the Multilayer Perceptron can include: 1. The hidden layers of the MLP can automatically extract high-level features of the data, which are particularly important for complex pattern recognition and classification tasks.

[0027] 2. By introducing non-linear activation functions and multiple layers, the MLP can fit complex non-linear relationships.

[0028] 3. The MLP is widely used in various classification, regression, and clustering tasks.

[0029] Specifically, layer normalization is used to perform an initial feature transformation on the spatially enhanced feature map, adjusting its distribution to make it more suitable for subsequent multi-layer perceptron processing. Subsequently, the normalized features are input into the multi-layer perceptron for a secondary feature transformation. The hidden layer of the multi-layer perceptron MLP transforms the input features through non-linear activation functions (such as ReLU, sigmoid, etc.) to extract higher-level features. After processing through multiple hidden layers, the MLP finally outputs an intermediate feature map, and these features can be used for subsequent classification, regression, and other tasks. Therefore, the intermediate feature map is the spatially enhanced feature map after layer normalization and MLP processing, containing richer feature information and providing more powerful support for subsequent object detection, recognition, and other tasks.

[0030] Therefore, feature transformation is performed through layer normalization and multi-layer perceptron. Layer normalization provides a good foundation for subsequent feature transformation by stabilizing the training process and improving the convergence speed. The multi-layer perceptron further transforms the spatially enhanced feature map after the initial feature transformation through its powerful feature extraction and non-linear fitting capabilities to obtain a more representative intermediate feature map. In particular, appropriate network structures and parameter settings can also be selected according to specific tasks and data characteristics to achieve the best performance and effects.

[0031] Next, the channel self-attention mechanism is used to project the reshaped intermediate feature map into queries, keys, and values in the channel dimension. The similarity between the queries and keys in the channel dimension is calculated through the Softmax function to obtain channel attention weights, and the values in the channel dimension are weighted and aggregated according to the channel attention weights to obtain a channel-enhanced feature map. The reshaping operation, that is, the reshape operation. In data processing and scientific computing, the reshape operation is used to adjust the dimensional structure of an array or matrix without changing its data content. Through the reshape operation, data can be transformed from one shape to another to meet different processing requirements.

[0032] Then, layer normalization is used again to perform an initial feature transformation on the channel-enhanced feature map ( ), and then the multi-layer perceptron is used to perform a secondary feature transformation on the channel-enhanced feature map after the initial feature transformation, and output a multi-scale feature fusion map ( ), that is, is processed through an independent LN-MLP module to output features with both spatial structure and channel discriminability .

[0033] To this end, the entire processing chain uses a cascaded structure of "spatial attention - LN - MLP - channel attention - LN - MLP" to achieve hierarchical feature enhancement from local to global. Among them, the spatial branch focuses on pixel - level relationship modeling, the channel branch strengthens semantic feature interaction, and the two stages are stably connected through LN - MLP, significantly improving the feature representation ability while ensuring computational efficiency.

[0034] Optionally, for the operation in step 101 of "calculating the similarity between the query and the key in the spatial dimension through the Softmax function to obtain the spatial attention weight, and performing weighted aggregation on the values in the spatial dimension according to the spatial attention weight to obtain the spatially enhanced feature map", it specifically includes: Step 1011: Based on the Softmax function, the query and the key in the spatial dimension, construct a spatial attention weight calculation formula, and calculate the similarity between the query and the key in the spatial dimension through the spatial attention weight calculation formula to obtain the spatial attention weight.

[0035] Step 1012: Based on the spatially enhanced feature map representation formula and the spatial attention weight, perform weighted aggregation on the values in the spatial dimension to obtain the spatially enhanced feature map, where the spatial attention weight calculation formula is: , The spatially enhanced feature map representation formula is: , is the spatial attention weight, is the query in the spatial dimension, is the key in the spatial dimension, is the similarity in the spatial dimension, is the spatially enhanced feature map, is the value in the spatial dimension.

[0036] In the above - mentioned embodiment of the present application, when the initial feature map is input into the infrared small - target detection network, the input initial feature map is first processed by the spatial self - attention mechanism to enhance the features through the self - attention mechanism in the spatial dimension. After being processed by the spatial self - attention mechanism, the initial feature map is linearly projected to obtain three key components: the query in the spatial dimension ( ), the key ( ), and the value ( ), and these three components are the core in the self - attention mechanism, used to calculate the attention weight and aggregate information.

[0037] Then, calculate the similarity between the query ( ) and the key ( ) in the spatial dimension through the Softmax function to obtain the spatial attention weight 。 It represents the correlation between different spatial positions and is used to guide the aggregation of information.

[0038] Finally, using the spatial attention weights weighted aggregation is performed on the values in the spatial dimension ( ) to obtain the spatially enhanced feature map 。 is the output enhanced by the spatial self-attention mechanism and contains more information about the spatial dimension.

[0039] Optionally, in step 102, the "intermediate feature map" is specifically represented as: , is the intermediate feature map, is a multi-layer perceptron, is layer normalization, is the spatially enhanced feature map.

[0040] Optionally, for "calculating the similarity between the query and the key in the channel dimension through the Softmax function to obtain the channel attention weights, and performing weighted aggregation on the values in the channel dimension according to the channel attention weights to obtain the channel enhanced feature map" in step 103, it specifically includes: Step 1031, based on the Softmax function, the query and the key in the channel dimension, construct a channel attention weight calculation formula, and calculate the similarity between the query and the key in the channel dimension through the channel attention weight calculation formula to obtain the channel attention weights.

[0041] Step 1032, based on the channel enhanced feature map representation formula and the channel attention weights, perform weighted aggregation on the values in the channel dimension to obtain the channel enhanced feature map, where the channel attention weight calculation formula is: , The channel enhanced feature map representation formula is: , is the channel attention weight, is the query in the channel dimension, is the key in the channel dimension, is the similarity in the channel dimension, is the channel enhanced feature map, is the value in the channel dimension.

[0042] In the above embodiments of the present application, the following operations can also be performed: For example, the tensor of the input feature map X (the intermediate feature map after reshaping operation) is , where C is the number of channels, H and W are the height and width respectively, and the queries in the channel dimension and keys are obtained by linearly transforming the input feature map X, that is: and , where and are the learnable weight matrices of the queries and keys in the channel dimension respectively.

[0043] Next, perform a dot product operation on the queries and keys in the channel dimension to obtain the similarity matrix . To make the similarity matrix interpretable, can be scaled, that is: , Next, apply the Softmax function to the scaled similarity matrix to obtain the channel attention weight matrix . The Softmax function is used to ensure that the sum of the elements in the weight matrix is 1, thus representing the relative importance between different channels.

[0044] Therefore, the channel attention weight matrix is obtained. Each element in

[0045] represents the similarity between the queries and keys in the channel dimension on a specific channel, thus reflecting the importance of that channel in the overall feature representation. The values in the channel dimension are also obtained by linearly transforming the input feature map X, that is: where is the learnable weight matrix of the values in the channel dimension. Use the channel attention weight matrix to perform weighted aggregation on the values in the channel dimension to obtain the channel-enhanced feature map, that is

[0046] Step 105: The decoder performs reverse decoding on the multi-scale feature fusion map to obtain a multi-scale feature fusion infrared image, and uses an adaptive weight adjustment mechanism to perform feature recalibration on the multi-scale feature fusion infrared image to generate a binary detection map.

[0047] Next, the decoder performs reverse decoding on the multi-scale feature fusion map to obtain a multi-scale feature fusion infrared image, and uses an adaptive weight adjustment mechanism to perform feature recalibration on the multi-scale feature fusion infrared image to be detected, generating a binary detection map. Specifically, after using the encoder to extract multi-scale features, the extracted features contain context information and detail information of different scales. The multi-scale features are fused to obtain a fused feature map (multi-scale feature fusion map). The fusion method can use simple splicing, addition, or more complex attention mechanisms, etc. Then, the decoder performs reverse decoding on the fused multi-scale feature fusion map, gradually restoring it to the image space to obtain a multi-scale feature fusion infrared image. Operations such as upsampling and convolution can be used during the decoding process to gradually restore the resolution and details of the image. In particular, the decoder is also designed with an adaptive weight adjustment mechanism, which can dynamically adjust the importance of different features according to the image content. Specifically, methods such as attention mechanisms, channel attention, and spatial attention can be used to achieve adaptive weight adjustment. For example, the Softmax function is used to calculate the channel attention weights, and the channel features are weighted and summed according to the weights to achieve feature recalibration. Finally, the image after feature recalibration is binarized to obtain a binary detection map. During the binarization process, the binarization threshold can be adjusted according to specific application scenarios and requirements, or a simple threshold segmentation method can be used, or more complex image segmentation algorithms can also be used.

[0048] Therefore, through the multi-scale fusion in the early stage of the decoder, the decoder can make full use of feature information of different scales, capture more details and context information, thereby improving the accuracy of target detection. The adaptive weight adjustment mechanism can dynamically adjust the importance of features according to the image content, further highlighting the target features and suppressing background interference. The multi-scale feature fusion map can handle targets of different scales and shapes, improving the robustness of the algorithm. The adaptive weight adjustment mechanism enables the algorithm to adapt to different scenarios and lighting conditions, reducing false detections and missed detections. Through feature recalibration, the algorithm can generate a clearer and more accurate binary detection map, facilitating subsequent target recognition and analysis.

[0049] Optionally, before the step 105 "the decoder performs reverse decoding on the multi-scale feature fusion map to obtain a multi-scale feature fusion infrared image", it specifically further includes: Step 106: The multi-scale information fusion channel multi-head self-attention module converts the multi-scale feature fusion map into a unified feature sequence through block partitioning operations, and through an improved residual multi-head attention mechanism, uses the feature maps output by each level of the spatial-channel two-stream self-attention module as independent queries and shared key-value pairs for interaction to obtain cross-scale global channel dependencies. After integrating the cross-scale global channel dependencies into the unified feature sequence, it is input into the decoder, where the feature maps include the spatially enhanced feature map, the intermediate feature map, the channel-enhanced feature map, and the multi-scale feature fusion map.

[0050] In the above embodiments of the present application, the skip connection part in the infrared small target detection network adopts a multi-scale information fusion channel multi-head self-attention module for processing the outputs of each level of the encoder. This module first converts the multi-scale feature fusion map output by the encoder into a feature sequence with a unified representation (unified feature sequence) through block partitioning operations, while keeping the original channel dimension unchanged. Subsequently, through an improved residual multi-head attention mechanism, features at each scale are used as independent queries and shared key-value pairs for interaction to establish cross-scale global channel dependencies. Finally, the processed feature sequence is restored to its original shape and fused with the input features. This module innovatively designs a special attention calculation method to effectively fuse local details and global context information, and realizes deep feature interaction through a multi-layer stacked Transformer (self-attention mechanism) structure. This design significantly improves the model's multi-scale feature representation ability for infrared small targets, especially showing excellent performance when dealing with targets of different sizes and complex background interference, providing an important guarantee for improving the detection accuracy.

[0051] Optionally, the decoder includes a three-level feature recalibration self-attention module (as Figure 5 shown) and a cross-level attention mechanism. The decoder in step 105 performs reverse decoding on the multi-scale feature fusion map to obtain a multi-scale feature fusion infrared image, and uses an adaptive weight adjustment mechanism to perform feature recalibration on the multi-scale feature fusion infrared image to generate a binary detection map, specifically including: Step 1051: The decoder performs reverse decoding on the input unified feature sequence incorporating cross-scale global channel dependencies to obtain a multi-scale feature fusion infrared image.

[0052] Step 1052: When the decoder uses each level of the feature recalibration self-attention module to reverse decode the output features, after performing feature transformation on the output features of each level of the feature recalibration self-attention module through layer normalization, the output features after feature transformation are converted into queries, and the multi-scale feature fusion map is mapped into keys and values.

[0053] Step 1053, using a cross-level attention mechanism to calculate the similarity between the query and the key, constructing a multi-scale feature correlation matrix based on the calculated similarity, combining the adaptive weight allocation strategy of the channel dimension, determining the most relevant multi-scale feature fusion map for each output feature, and enhancing it to generate a binary detection map, wherein the binary detection map is generated after being processed by a 1×1 convolution layer and a sigmoid function.

[0054] In the above embodiment of the present application, the decoder reversely decodes the input unified feature sequence that incorporates the cross-scale global channel dependency to obtain a multi-scale feature fusion infrared image. More specifically, the input of the decoder consists of two parts: the output features of the feature recalibration self-attention module at each level in the decoder branch, and the output features of the spatial channel dual-stream self-attention module at each level in the encoder branch. In the decoder decoding process, a decoder-guided dynamic feature recalibration mechanism is adopted, the core of which is to use high-level semantic information to optimize the underlying feature representation. In the feature conversion stage, the high-level semantic features output by the decoder are first processed by the LayerNorm layer (layer normalization) to convert them into query vectors (Query), and the multi-level spatial features extracted by the encoder are respectively mapped to key-value pairs (Key-Value). This dual-path conversion design not only retains the inherent characteristics of features at different levels, but also lays the foundation for subsequent feature interactions. In the attention modeling stage, a cross-level attention mechanism is used to establish dynamic associations between features. This mechanism constructs a multi-scale feature correlation matrix by calculating the similarity between the query vector and the key vector, and combines the adaptive weight allocation strategy of the channel dimension so that each decoder feature can intelligently select the most relevant encoder feature for enhancement. This decoder-guided dynamic feature recalibration mechanism realizes a two-way interaction between encoder and decoder features: the decoder features serve as query signals to guide the dynamic adjustment of encoder features, while the cross-attention mechanism adaptively calculates the correlation between features and can focus on the most discriminative information. This design significantly improves the adaptability to multi-scale targets and exhibits excellent robustness in complex scenarios.

[0055] By introducing a cross-level attention mechanism in the decoding stage, high-level semantic information is used to guide the optimization of underlying features. By dynamically calculating the correlation between the query vector (decoder features) and the key-value pair (encoder multi-level features), the most discriminative information is adaptively selected for enhancement, which significantly improves the positioning accuracy of small targets in complex backgrounds, which is better than the static feature fusion strategy.

[0056] Specifically, the final output of the infrared small target detection network is processed by a 1×1 convolution and a sigmoid function to generate a binary detection map. In the entire data processing flow, the encoder is responsible for gradually extracting and condensing feature information, the skip connection ensures the complete transmission of multi-scale features, and the decoder achieves precise positioning through recalibration and feature fusion. Residual connections are used for feature transmission between modules, effectively alleviating the problem of gradient disappearance. This carefully designed architecture enables the network to maintain real-time performance while improving the detection performance of small and weak targets in complex backgrounds.

[0057] Step 107: The infrared small target detection module uses the binary detection map to detect infrared small targets.

[0058] In the above embodiments of the present application, the infrared small target detection module uses the binary detection map to detect infrared small targets. Therefore, infrared small target detection is achieved through the infrared small target detection network.

[0059] Optionally, the infrared small target detection module in step 107 uses the binary detection map to detect infrared small targets, which specifically includes: Step 1071: The infrared small target detection module determines the detection threshold of the binary detection map based on the maximum inter-class variance algorithm. After threshold segmentation of the binary detection map based on the detection threshold, the part with pixel values higher than the detection threshold is determined as an infrared small target.

[0060] In the above embodiments of the present application, a suitable threshold can be set according to the characteristics of the binary detection map. The threshold is used to distinguish infrared small targets from background noise. Specifically, the maximum inter-class variance algorithm can be used to determine the detection threshold, and then threshold segmentation is performed on the binary detection map, that is, the part with pixel values higher than the detection threshold is regarded as an infrared small target, and the part lower than the detection threshold is regarded as the background.

[0061] Furthermore, connected component analysis can be performed on the image after threshold segmentation to identify the connected regions in the image. The connected regions correspond to the infrared small targets accordingly, because small targets often appear as continuous pixel blocks in the image. According to the characteristics such as the size and shape of the connected components, the true infrared small targets are further screened and extracted, while possible false detections (such as noise, artifacts, etc.) are removed, and the real targets are retained. The extracted infrared small targets can also be located to determine their specific positions and ranges in the image. For example, the position of the target can be marked using methods such as bounding boxes and centroid coordinates.

[0062] Specifically, further post-processing can also be performed on the detection results according to specific application scenarios and requirements. For example, operations such as target tracking, target classification, and target counting can be carried out.

[0063] Finally, the detection results can be visualized, such as overlaying the target positions on the original image or comparing them with the ground truth labels or the results of manual detection to verify the accuracy and reliability of the detection algorithm. For this purpose, binary detection maps are used to effectively detect small infrared targets.

[0064] By applying the technical solution of this embodiment, to address the deficiency in detection accuracy of current deep learning-based infrared target detection methods, spatial-channel multi-attention is mixed during the detection of small infrared targets, a hierarchical feature processing flow is adopted, and multi-dimensional optimization of features is achieved through a novel attention mechanism. It can be applied to the detection of complex backgrounds and small targets. Compared with traditional deep learning methods, the above embodiment of this application has higher efficiency and accuracy in the infrared target detection task, has good application prospects, and can play an important role in the field of infrared image processing.

[0065] Based on the method as Figures 1 to 5 shown above, correspondingly, an embodiment of this application also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the infrared small target detection method as Figures 1 to 5 shown above is implemented.

[0066] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various implementation scenarios of this application.

[0067] Based on the method as Figures 1 to 5 shown above, to achieve the above object, an embodiment of this application also provides a computer device, which can specifically be a personal computer, server, network device, etc. The computer device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the infrared small target detection method as Figures 1 to 5 shown above.

[0068] Optionally, the computer device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen (Display) and an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), etc.

[0069] Those skilled in the art can understand that the structure of a computer device provided in this embodiment does not limit the computer device, and it may include more or fewer components, or combine certain components, or have different component arrangements.

[0070] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages and stores the hardware and software resources of the computer device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components within the storage medium, as well as communication with other hardware and software in the entity device.

[0071] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or can be implemented through hardware by means of a cascaded processing architecture of a spatial channel dual-stream self-attention module (spatial attention, LN, MLP, channel attention, LN, MLP), realizing hierarchical feature enhancement from local to global. The spatial branch focuses on pixel-level relationship modeling, and the channel branch strengthens semantic feature interaction. The two cooperate and optimize to significantly improve the feature representation ability, overcoming the problem of insufficient feature expression ability of traditional methods in complex backgrounds. The skip connection part adopts an improved multi-scale information fusion channel multi-head self-attention module, unifies the multi-scale feature representation through block division operations, and establishes cross-scale global dependence relationships using a residual multi-head attention mechanism. Compared with traditional convolution or simple splicing methods, it more effectively fuses local details and global context information, enhancing the adaptability to targets of different sizes.

[0072] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment scenario, and the modules or processes in the drawings are not necessarily essential for implementing this application. Those skilled in the art can understand that the modules in the device in the embodiment scenario can be distributed in the device in the embodiment scenario according to the description of the embodiment scenario, or can be correspondingly changed and located in one or more devices different from this embodiment scenario. The modules in the above embodiment scenario can be combined into one module, or further split into multiple sub-modules.

[0073] The above serial numbers of this application are only for description and do not represent the advantages or disadvantages of the embodiment scenarios. The above-disclosed are only several specific embodiment scenarios of this application. However, this application is not limited thereto, and any changes made by those skilled in the art should fall within the protection scope of this application.

Claims

1. A method for detecting small infrared targets, characterized in that: The invention is applied to an infrared small target detection network, wherein the infrared small target detection network comprises a channel multi-head self-attention module for multi-scale information fusion connected with a layer normalization and multi-layer perceptron joint module, an infrared small target detection module, and an encoder and a decoder with a reciprocal structure. The encoder is a cascade processing architecture, which is composed of four levels of spatial channel dual-stream self-attention modules. Each level of spatial channel dual-stream self-attention module includes a spatial self-attention mechanism and a channel self-attention mechanism, respectively. The layer normalization and multi-layer perceptron joint module includes layer normalization and a multi-layer perceptron. The method comprises: An infrared image to be detected containing an infrared small target is obtained, and a 3×3 convolutional layer is used to embed features of the infrared image to be detected. After obtaining an initial feature map, the initial feature map is input into the infrared small target detection network, so that the infrared small target detection network performs the following steps to realize infrared small target detection: The spatial self-attention mechanism is used to project the initial feature map into queries, keys, and values ​​in the spatial dimension. The similarity between the query and the key in the spatial dimension is calculated by the Softmax function to obtain the spatial attention weight. The values ​​in the spatial dimension are weighted and aggregated according to the spatial attention weight to obtain the spatial enhanced feature map. After performing a primary feature transformation on the spatial enhancement feature map by using layer normalization, a secondary feature transformation is performed on the spatial enhancement feature map after the primary feature transformation by using a multi-layer perceptron to obtain an intermediate feature map; The channel self-attention mechanism is used to project the intermediate feature map after the reshaping operation into the query, key and value in the channel dimension. The similarity between the query and the key in the channel dimension is calculated by the Softmax function to obtain the channel attention weight. The values ​​in the channel dimension are weighted and aggregated according to the channel attention weight to obtain the channel enhanced feature map. After performing a primary feature transformation on the channel enhancement feature map by using layer normalization, a secondary feature transformation is performed on the channel enhancement feature map after the primary feature transformation by using a multi-layer perceptron, and a multi-scale feature fusion map is output; The decoder reversely decodes the multi-scale feature fusion image to obtain a multi-scale feature fusion infrared image, and uses an adaptive weight adjustment mechanism to perform feature recalibration on the multi-scale feature fusion infrared image to generate a binary detection image; The infrared small target detection module detects the infrared small target using a binary detection image.

2. The method according to claim 1, characterized in that Before the decoder reversely decodes the multi-scale feature fusion image to obtain the multi-scale feature fusion infrared image, the method further includes: The channel multi-head self-attention module of the multi-scale information fusion converts the multi-scale feature fusion map into a unified feature sequence through a block partitioning operation, and uses an improved residual multi-head attention mechanism to interact the feature maps output by each of the spatial channel dual-stream self-attention modules at each level as independent queries with shared key-value pairs to obtain a cross-scale global channel dependency relationship, and after integrating the cross-scale global channel dependency relationship into the unified feature sequence, the feature map is input into the decoder, wherein the feature map includes a spatial enhancement feature map, an intermediate feature map, a channel enhancement feature map, and a multi-scale feature fusion map; Accordingly, the decoder reversely decodes the multi-scale feature fusion image to obtain a multi-scale feature fusion infrared image, including: The decoder reversely decodes the input unified feature sequence integrated with the cross-scale global channel dependency to obtain a multi-scale feature fused infrared image.

3. The method according to claim 1, characterized in that: The decoder includes a three-level feature recalibration self-attention module and a cross-level attention mechanism. The adaptive weight adjustment mechanism is used to perform feature recalibration on the multi-scale feature fusion infrared image to generate a binary detection map, including: When the decoder reversely decodes the output features of the feature recalibration self-attention modules at each level, the output features of the feature recalibration self-attention modules at each level are respectively transformed by layer normalization, the transformed output features are converted into queries, and the multi-scale feature fusion graph is mapped into keys and values; The cross-level attention mechanism is used to calculate the similarity between the query and the key, and a multi-scale feature correlation matrix is ​​constructed based on the calculated similarity. Combined with the adaptive weight allocation strategy in the channel dimension, the most relevant multi-scale feature fusion map is determined for each output feature to enhance it and generate a binary detection map.

4. The method according to claim 3, characterized in that: The binary detection map is generated after being processed by a 1×1 convolutional layer and a sigmoid function.

5. The method according to claim 1, characterized in that The similarity between the query and the key in the spatial dimension is calculated by the Softmax function to obtain the spatial attention weight, and the values ​​in the spatial dimension are weighted and aggregated according to the spatial attention weight to obtain the spatial enhanced feature map, including: Based on the Softmax function, the query and key in the spatial dimension, a spatial attention weight calculation formula is constructed. The similarity between the query and the key in the spatial dimension is calculated by the spatial attention weight calculation formula to obtain the spatial attention weight. Based on the spatial enhancement feature map representation formula and the spatial attention weight, the values ​​on the spatial dimension are weighted and aggregated to obtain the spatial enhancement feature map, where the spatial attention weight calculation formula is: , The spatial enhancement feature map characterization formula is: , is the spatial attention weight, For queries in the spatial dimension, is the key in the spatial dimension, is the similarity in the spatial dimension, is the spatial enhancement feature map, is the value of the spatial dimension.

6. The method according to claim 1, characterized in that The Softmax function is used to calculate the similarity between the query and the key in the channel dimension to obtain the channel attention weight, and the values ​​in the channel dimension are weighted and aggregated according to the channel attention weight to obtain the channel enhanced feature map, including: Based on the Softmax function, the query and key in the channel dimension, a channel attention weight calculation formula is constructed. The similarity between the query and key in the channel dimension is calculated through the channel attention weight calculation formula to obtain the channel attention weight. Based on the channel enhancement feature map representation formula and the channel attention weight, the values ​​on the channel dimension are weighted and aggregated to obtain the channel enhancement feature map, where the channel attention weight calculation formula is: , The channel enhancement feature map characterization formula is: , is the channel attention weight, For queries on the channel dimension, is the key on the channel dimension, is the similarity in the channel dimension, is the channel enhanced feature map, is the value in the channel dimension.

7. The method according to any one of claims 1 to 6, characterized in that The intermediate feature map is expressed as: , is the intermediate feature map, is a multi-layer perceptron, is layer normalization, It is a spatial enhancement feature map.

8. The method according to claim 7, characterized in that The infrared small target detection module detects the infrared small target using a binary detection image, including: The infrared small target detection module determines the detection threshold of the binary detection image based on the maximum inter-class variance algorithm, performs threshold segmentation on the binary detection image based on the detection threshold, and determines the part with pixel value higher than the detection threshold as an infrared small target.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the infrared small target detection method according to any one of claims 1 to 8 is implemented.

10. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the infrared small target detection method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Absolute pose regression method based on cascade attention module

    CN118657831A

  • Infrared and visible light fusion method based on multi-scale feature interaction enhancement

    CN119091269A

  • Infrared small target detection method based on multi-scale self-attention motion background modeling

    CN119478349A

  • Underground coal mine pedestrian detection method based on image fusion and feature enhancement

    WO2024037408A1

  • Joint modeling method and apparatus for enhancing local features of pedestrians

    WO2024060321A1

Cited By

  • Unmanned aerial vehicle RGB-infrared progressive fusion target detection method and device based on state space model

    CN121482655A