An infrared dim small target detection method based on hybrid spatial modulation feature convolutional neural network

CN116486102BActive Publication Date: 2026-09-25BEIHANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310406665.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-17
Publication Date
2026-09-25
Estimated Expiration
2043-04-17

AI Technical Summary

Technical Problem

[0006]1、目的:针对复杂背景下弱小目标检测精度低、虚警率高、实时性差的问题,本发明提出了一种基于混合空间调制特征卷积神经网络的红外弱小目标检测方法,模型充分提取红外弱小目标具有高斯分布特性的多方向特征、灰度突变的局部特征并基于多层次跨尺度的特征融合思路进行网络设计,在提高检测精度,降低模型参数量、虚警率以及运行时间上有明显改善

Benefits of technology

[0028]本发明提出一种基于混合空间调制特征卷积神经网络的红外弱小目标检测方法,从弱小目标具有高斯分布特性这一多方向特征出发,利用全局注意力机制和卷积操作构造多方向固定高斯注意力以抑制背景并增强目标特征;从弱小目标局部亮度较高且与背景存在较大的突变这一局部灰度特性出发,构造混合感受野骨干网络进一步利用弱小目标的局部邻域特性,实现更适于本任务的特征提取;构造交叉窗口注意力机制融合低中高层特征,更好地保留小目标相关特征的同时提取多尺度特征。模型设计从红外弱小目标的特性出发,在可解释性和性能方面有较好表现,应用前景广泛。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116486102B_ABST
    Figure CN116486102B_ABST
Patent Text Reader

Abstract

The application provides an infrared dim small target detection method based on a hybrid space modulation feature convolutional neural network, and the steps are as follows: 1. constructing a multi-direction fixed Gaussian kernel attention, using global attention for background suppression, and then using a fixed weight Gaussian kernel to extract target multi-direction features for target feature enhancement; 2. constructing a backbone network based on a hybrid receptive field convolution block series for three-group feature extraction on the enhanced shallow layer features; 3. constructing a cross sliding attention mechanism, fusing the three groups of features extracted by the backbone network through the cross sliding window attention mechanism, and splicing in the channel dimension; using the multi-direction Gaussian kernel attention and the convolution layer for pixel-by-pixel prediction to obtain a pixel-level probability prediction map of the whole image; 4. sequentially connecting the modules to build a convolutional neural network, and constructing a loss function to train the network; using the prediction result and the pixel-level label to calculate the loss, so as to realize the training of the network parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an infrared weak target detection method based on a hybrid spatial modulation feature convolutional neural network, belonging to the fields of digital image processing and computer vision, mainly involving deep learning and target detection technology, and has broad application prospects in various image-based application systems. Background Technology

[0002] Currently, infrared low-target detection is widely used in precision weapon guidance, forest fire monitoring and early warning, and UAV target detection and identification. The accuracy, stability, and real-time performance of its detection algorithms are crucial indicators for evaluating the performance of infrared low-target detection systems. In recent years, with the rapid development of the UAV industry, the high speed and small size of UAVs have posed challenges to visible light detection, and their high threat level necessitates real-time detection capabilities in the corresponding algorithms. In infrared images, the batteries, cameras, and other equipment carried by UAVs are relatively bright, making it possible to detect UAV targets using infrared images. However, in practical applications, infrared low-target detection is subject to numerous interferences: for example, clouds, fog, and clumps in the sky typically have high brightness in infrared images, easily obscuring the presence of low-target targets or causing them to be misidentified by detectors. In mountainous forest environments, low-target targets can often be well concealed due to reflections and interference from natural light. In ocean environments, changes in sea waves can interfere with infrared radiation, and there is also clutter interference from sea surface reflections. These complex factors significantly increase the difficulty of infrared low-target detection. Therefore, developing a real-time infrared low-target detection algorithm for complex backgrounds is a highly challenging and meaningful task.

[0003] Early infrared weak target detection algorithms can be roughly divided into three categories: methods based on background consistency estimation, methods based on saliency detection of the human visual system, and methods based on image patches. The background consistency estimation method assumes that the background is a continuous and smooth region. The appearance of weak targets will cause the continuous smoothness of the background region to be locally destroyed. Therefore, filters can be designed or morphological methods can be used for detection. Deng et al. proposed an adaptive infrared small target detection algorithm based on top-hat transformation (see reference: Deng Lizhen et al., Infrared small target detection based on adaptive M-estimator ring top-hat transformation, Pattern Recognition, 2021, 112: 107729. (Deng L, Zhang J, Xu G, et al. Infrared Small Target Detection via Adaptive M-estimator Ring Top-Hat Transformation[J]. Pattern Recognition, 2021, 112: 107729.)). Yao et al. designed an infrared small target detection algorithm based on manually designed filters (see reference: Yao Qin et al., Infrared small target detection based on multiple kernel filters and random walk, IEEE Transactions on Geography and Remote Sensing, 2019, 57(9): 7104-7118. (Qin Y, Bruzzone) L, Gao C, et al. Infrared Small Target Detection Based on Facet Kernel and RandomWalker[J].IEEE Transactions on Geoscience and Remote Sensing,2019,57(9):7104-7118.)) This type of method is limited by fixed handmade feature design, has low accuracy, and is extremely limited in application. Saliency detection methods based on the human visual system are mainly based on the contrast between the target and the background for algorithm design. Chen et al. proposed a detection method based on local contrast calculation (Reference: Chen Chunping et al.).A local contrast method for small infrared target detection, IEEE Transactions on Geoscience and Remote Sensing, 2013, 52(1):574-581. (Chen CLP, Li H, Wei Y, et al. A Local Contrast Method for Small Infrared Target Detection[J].IEEE Transactions on Geoscience and Remote Sensing, 2013, 52(1):574-581.) Deng et al. proposed a method for small infrared target detection based on local characteristics by designing a saliency measure (see reference: Deng et al., Small Infrared Target Detection Based on Weighted Local Difference Measure[J].IEEE Transactions on Geoscience and Remote Sensing, 2016, 54(7):4204-4214. (Deng H, Sun X, Liu M, et al. Small Infrared Target Detection Based on Weighted Local Difference Measure[J].IEEE Transactions on Geoscience and Remote Sensing, 2016, 54(7):4204-4214.) Sensing, 2016, 54(7): 4204-4214.)). In addition, Han et al. divided small targets and their neighborhoods into core layers, retention layers and background layers based on saliency, and used this to divide windows to construct local contrast. (See reference: Han Jinhui et al., A local contrast method for infrared small-target detection using a three-layer window, IEEE Geoscience and Remote Sensing Letters, 2019, 17(10): 1822-1826. Han J, Moradi S, Faramarzi I, et al. A LocalContrast Method for Infrared Small-target Detection Utilizing a Tri-layerWindow[J].IEEE Geoscience and Remote Sensing Letters, 2019, 17(10): 1822-1826.)) Although the algorithm based on the human visual system has a fast detection speed, it has poor robustness and is easily affected by local bright backgrounds, noise and other interferences. The image patch-based method takes advantage of the small proportion and sparse distribution of weak targets, divides the entire infrared image into multiple image patches, and uses an optimization algorithm to separate the targets from the background.Gao et al. first used this idea to propose a detection method based on a patch model (see reference: Gao Chenqiang et al. Infrared Patch-image Model for Small Target Detection in a SingleImage[J].IEEE Transactions on Image Processing,2013,22(12):4996-5009.), but the proposed model is relatively complex, resulting in excessive computation and poor practicality. Zhang et al. introduced a partial sum of tensor nuclear norm (PSTNN) combined with weighted l1 norm in the above IPI model to suppress background and preserve target (see reference: Zhang L, Peng Z. Infrared Small Target Detection Based on Partial Sum of the Tensor NuclearNorm[J].Remote Sensing,2019,11(4):382.)) but still could not solve the problem of computational burden caused by large image patches.

[0004] In recent years, deep learning technology has been widely used in computer vision, object detection and recognition, which has also promoted the integration of infrared weak target detection and deep learning technology. Liu et al. proposed a multi-layer convolutional network based on correlation filters, which treats the detection problem as a binary classification problem, cascades multiple weak classifiers and obtains relatively accurate results (Liu Qiang et al., Deep Convolutional Neural Networks for Thermal InfraredObject Tracking[J].Knowledge-Based Systems,2017,134:189-198.).Furthermore, attention mechanisms are considered an effective means to enhance the network's attention to regions of interest. Various attention extraction methods exist, such as the self-attention mechanism proposed by Vaswani et al. (see: Vaswani et al., Self-Attention Model, Conference and Workshop on Neural Information Processing Systems, 2017, 30. (Vaswani A, Shazeer N, Parmar N, et al. Attention Is All You Need[J]. Advances in Neural Information Processing Systems, 2017, 30.)) and the convolutional spatial attention mechanism proposed by Li et al. (see: Li Shangyu et al., Cbam: Convolutional Block Attention Module, European Conference on Computer Vision, 2018: 3-19. (Woo S, Park J, Lee JY, et al. Cbam: Convolutional Block Attention Module[C] / / Proceedings of the European Conference on Computer Vision)). Vision (ECCV). 2018:3-19.) and the global attention mechanism proposed by Cao et al. (see: Cao et al., Gcnet: Non-local networks meet squeeze-excitation networks and beyond[C] / / Proceedings of the IEEE / CVF international conference on computervision workshops.2019:0-0.) has been widely applied in various neural network models. Combining top-down task-driven attention and bottom-up saliency attention, Dai et al. proposed a non-modulated symmetric context mechanism, which fuses local target information and semantic information to achieve the detection of weak targets.(See reference: Dai Y, Wu Y, Zhou F, et al. Asymmetric Contextual Modulation for Infrared Small Target Detection[C] / / Proceedings of the IEEE / CVF Winter Conference on Applications of ComputerVision.2021:950-959.)). Li et al. proposed a densely linked target detection network to alleviate the difficulty of losing target information in the deep layers of the network for weak target features. (See reference: Li B, Xiao C, Wang L, et al. Dense Nested Attention Network for Infrared Small Target Detection[J].IEEE Transactions on Image Processing,2022.)).

[0005] While deep learning methods offer advantages in accuracy, current methods rarely consider the characteristics of small targets and suffer from poor real-time performance due to complex network architectures, limiting their effectiveness in small target detection tasks. To achieve fast and effective small target detection, this invention designs a deep learning network model based on the scale and grayscale distribution characteristics of small targets, proposing an infrared small target detection method based on a hybrid spatial modulation feature convolutional neural network. Summary of the Invention

[0006] 1. Objective: To address the problems of low detection accuracy, high false alarm rate, and poor real-time performance of weak targets in complex backgrounds, this invention proposes an infrared weak target detection method based on a hybrid spatial modulation feature convolutional neural network. The model fully extracts multi-directional features with Gaussian distribution characteristics and local features with gray-scale abrupt changes of infrared weak targets, and designs the network based on a multi-level cross-scale feature fusion approach. This method significantly improves detection accuracy, reduces model parameter count, false alarm rate, and running time.

[0007] 2. Technical Solution: To achieve the above objectives, the overall approach of this invention is based on the statistical result that small targets have relatively high local brightness and a significant difference from the background. From three perspectives—multi-directional, hybrid receptive field, and multi-scale feature fusion—it designs a multi-directional Gaussian kernel attention network, a lightweight hybrid receptive field backbone network, and a cross-sliding window attention mechanism to extract multi-scale information and fuse low, medium, and high-level features. This constructs a lightweight neural network for infrared small target detection, while ensuring fast detection speed and high target feature extraction capability. The technical approach of this invention is mainly reflected in the following three aspects:

[0008] 1) Based on the statistical result that weak targets can usually be approximated as a two-dimensional Gaussian distribution of noise, we designed a multi-directional fixed Gaussian kernel attention to extract spatial attention, which can perform background suppression while fully integrating features from various directions, thereby enhancing the extraction effect of targets.

[0009] 2) Based on the local grayscale characteristics of small targets with high local brightness and large abrupt changes with the background, a hybrid receptive field convolutional block is designed. By connecting several convolutional units of different sizes, expansion coefficients and group numbers in series and parallel, the local features of the target and the difference features between the target and its neighboring background are fully extracted, thereby further enhancing the target.

[0010] 3) Based on the characteristic that details such as edge, shape, and texture of weak targets are concentrated in low-level features and semantic information including spatial location and background suppression is concentrated in high-level features, a cross-sliding window attention mechanism is designed. The low-level, mid-level and high-level features of the backbone network are divided into windows of different sizes. Combined with sliding window attention, the corresponding details and semantic information are fully integrated, extracting multi-scale features while ensuring low computational complexity, thereby achieving better segmentation and detection results.

[0011] This invention relates to a method for detecting weak infrared targets based on a hybrid spatial modulation feature convolutional neural network. The specific steps of this method are as follows:

[0012] Step 1: Extract shallow features and construct a multi-directional fixed Gaussian kernel attention. Use global attention for background suppression, and then use a fixed-weight Gaussian kernel to extract multi-directional features of the target for target feature enhancement.

[0013] Step 2: Construct a backbone network based on concatenated convolutional blocks with hybrid receptive fields to extract three sets of features from the enhanced shallow features;

[0014] Step 3: Construct a cross-sliding attention mechanism to fuse the three sets of features extracted from the backbone network through the cross-sliding window attention mechanism and stitch them together in the channel dimension; and then use multi-directional Gaussian kernel attention and convolutional layers with a kernel size of 3×3, expansion coefficient of 1, grouping of 1, and stride of 1 to perform pixel-by-pixel prediction to obtain the pixel-level probability prediction map of the entire image.

[0015] Step 4: Concatenate the above modules sequentially to build a convolutional neural network, and construct a loss function to train the network. Use the prediction results and pixel-level labels to calculate the loss, thereby training the network parameters.

[0016] Output: The infrared image is processed using the trained neural network; after sufficient iterative training of the hybrid spatial modulation feature-based convolutional neural network using the training data, the trained network is used to detect target pixels.

[0017] Specifically, step one is as follows:

[0018] 1.1: Extracting shallow features and using multi-directional fixed Gaussian kernel attention for target feature enhancement. The network mainly uses convolutional units as basic components. Each convolutional unit consists of one convolutional layer, a batch normalization layer, and a Selu activation function operation. Parameters such as kernel size, expansion coefficients, number of groups, stride, and activation function type in the convolutional layer are adjusted as needed. First, the input image passes through a convolutional unit with a kernel size of 7×7, expansion coefficient of 1, number of groups of 1, and stride of 1, generating a shallow feature F with 16 channels. s Typically, an infrared image of a weak target can be considered to consist of three parts: the target, the background, and noise. I = B + T + N, where I represents the original image matrix, B represents the background matrix, T represents the target matrix, and N represents noise and other error matrices. To accurately separate the background and target, this invention proposes that weak targets can be modeled as unusual bright spots in the image with high contrast to the background, whose grayscale distribution exhibits characteristics similar to a two-dimensional Gaussian function, such as... Figure 1 As shown in region c, a multi-directional fixed Gaussian kernel is designed to effectively locate the target. Furthermore, the background may contain bright clouds or fog, which can easily interfere with the detection of weak targets, such as... Figure 1 As shown in regions s1 to s3, their grayscale distribution characteristics resemble those of weak targets, thus requiring the introduction of a background suppression mechanism to reduce interference from the background. To address these issues, this invention proposes a multi-directional fixed Gaussian kernel attention mechanism to focus on the extracted shallow features F. s Background suppression and target enhancement are performed to obtain the enhanced shallow features F. e The specific structure is as follows Figure 2As shown. Considering the widespread distribution of clouds and fog in the background and the sparse distribution of targets, a global attention mechanism GCBlock is first introduced for background suppression. Then, a multi-directional fixed Gaussian kernel is constructed. The probability that a pixel is a target is measured by the magnitude of the gray-level difference between a pixel and its neighboring pixels in multiple directions. Spatial attention is extracted to enhance target features. This invention's multi-directional fixed Gaussian kernel attention first applies the input feature map F... s Background suppression was performed using GCBlock, a global attention mechanism with a channel dimension compression ratio of 0.25, to obtain the feature map F. c-attn Next, a pointwise convolutional layer with a kernel size of 1×1, a spread factor of 1, a group number of 1, and a stride of 1 is used to reduce the channel dimension of the feature image after background suppression to 8, resulting in feature map F. a Then the feature map F a The system is divided into 8 groups along the channel dimension, and 8 fixed convolutional kernels are used in parallel to compute the directional feature maps F. d The kernel size is fixed at 5×5. The remaining convolution kernels d i It is obtained by rotating d1 counterclockwise by i×45°. Then, two sets of cascaded pointwise convolutional units with a kernel size of 1×1, an expansion coefficient of 1, a group number of 1, and a stride of 1 are used to fully fuse the orientation feature map F. d The directional information from different channels is used to obtain the fused directional feature map F′. d Finally, F d With F′ d Pointwise multiplication is performed, and pointwise convolutional units with a kernel size of 1×1, a spread factor of 1, a group number of 1, a stride of 1, and an activation function of sigmoid are used to reduce the channel dimension to 1, resulting in the multidirectional attention feature map F. d-attn Multi-directional attention feature map F d-attn With F c-attn The enhanced input feature map F is obtained by performing point-by-point multiplication. e .

[0019] Step two is as follows:

[0020] 2.1: Constructing a backbone network to extract features from the enhanced low-level features; the backbone network consists of three sets of hybrid receptive field convolutional blocks and convolutional units responsible for downsampling, alternating between them. Each hybrid receptive field convolutional block consists of a certain number of hybrid receptive field convolutional units and a global attention mechanism GCBlock responsible for background suppression, connected in series. Each hybrid receptive field convolutional unit consists of convolutional units with increasing kernel size, expansion coefficient of 1, number of groups equal to the number of input channels, and stride of 1; grouped expanded convolutional units with kernel size of 3×3, expansion coefficient of 2, number of groups of 4, and stride of 1; pointwise convolutional layers with kernel size of 1×1, expansion coefficient of 1, number of groups of 1, and stride of 1; and residual connections, in order to extract features at different scales. Its specific structure is as follows: Figure 3 As shown. The hybrid receptive field convolutional unit designed in this invention first divides the input features into four groups along the channel dimension. For each group, features are extracted using convolutional units with kernel sizes of 1×1, 3×3, 5×5, and 7×7, a scaling factor of 1, a grouping number equal to the number of input channels, and a stride of 1. These extracted features are then concatenated along the channel dimension. The processed features are then sequentially processed using a grouped extended convolutional unit with a kernel size of 3×3, a scaling factor of 2, a grouping of 4, and a stride of 1, and a pointwise convolutional layer with a kernel size of 1×1, a scaling factor of 1, a grouping number of 1, and a stride of 1. Finally, a residual connection is made with the input features to obtain the output features. The hybrid receptive field convolutional blocks are connected by a set of convolutional units with a kernel size of 3×3, a scaling factor of 2, a grouping number equal to the number of input channels, and a stride of 2, responsible for downsampling. The enhanced shallow feature F e The first set of hybrid receptive field convolutional blocks G1, consisting of a hybrid receptive field convolutional unit and a global attention mechanism GCBlock, yields the corresponding output feature map F1. This is then downsampled by a set of convolutional units with a kernel size of 3×3, a scaling factor of 1, grouping based on the number of input channels, and a stride of 2, doubling the channel dimension to 32. Next, the second set of hybrid receptive field convolutional blocks G2, consisting of a hybrid receptive field convolutional unit and a global attention mechanism GCBlock, yields the corresponding output feature map F2. This is then downsampled by another set of convolutional units with a kernel size of 3×3, a scaling factor of 1, grouping based on the number of input channels, and a stride of 2, doubling the channel dimension to 64. Finally, the third set of hybrid receptive field convolutional blocks G3, consisting of three hybrid receptive field convolutional units and a global attention mechanism GCBlock, yields the corresponding output feature map F3. The specific implementation of the feature extraction process is as follows... Figure 4 As shown.

[0021] Step three is as follows:

[0022] 3.1: Cross-Sliding Window Attention Mechanism. In infrared small target detection tasks, low-level features reflecting details such as edges, shape, and texture are related to target edge segmentation, such as... Figure 5b As shown; mid-to-high-level features containing more semantic information are related to target location determination and background suppression, such as... Figure 5c , 5d As shown. This invention designs a cross-sliding window attention mechanism for weak target features, and combines window partitioning to fuse low-level features with mid- and high-level features at multiple scales, resulting in a multi-scale output feature map F. m The main part of the cross-sliding window attention mechanism, the cross-window attention module, is implemented as follows: Figure 6 As shown.

[0023] 3.2: Detection is performed using a lightweight output layer consisting of multi-directional fixed Gaussian kernel attention and convolutional layers with a kernel size of 3×3, a spread factor of 1, a grouping of 1, and a stride of 1. The multi-scale output feature map F... m After further attention enhancement using a multi-directional fixed Gaussian kernel, the channel dimension is reduced to 1 using a convolutional layer with a kernel size of 3×3, a spread factor of 1, a grouping of 1, and a stride of 1. Then, after processing with a sigmoid activation function, a probability prediction map at the pixel level of the entire image is output.

[0024] Step four is as follows:

[0025] 4.1: Concatenate the modules proposed in steps one through three to build a convolutional neural network, such as... Figure 4 As shown. The loss function consists of the Intersection over Union (IOU) loss, L = L IOU The intersection-union ratio (IU) refers to the overlap rate between the predicted and ground truth regions; it's the ratio of their intersection to their union. When training a network for object detection, the ideal scenario is that the predicted and ground truth regions completely overlap, meaning the IU equals 1. Therefore, in practice, the IU value is always between 0 and 1, and a higher value indicates more accurate detection. Thus, the IU loss is defined. Where area(predict) is the target region predicted by the method of this invention, area(trut) is the area of ​​the real target region, ∩ is the intersection operation of sets, and ∪ is the union operation of sets. After giving the above definition of the loss function, the infrared image is input into the convolutional neural network to obtain the probability prediction map and perform a pixel-by-pixel multiplication with the labeled real result map to obtain the overlap result of the predicted target region and the real target region, i.e., area(predict)∩area(trut); based on this, the number of pixels in the real target region, the predicted target region, and the overlapping area of ​​the two are summed to calculate the intersection-union ratio loss.

[0026] 4.2: This invention uses the Adamw optimizer for optimization. The initial learning rate of the network is 0.0002, and the weight decay coefficient is 10. -3 During training, the learning rate is adaptively updated, and the network parameters are adjusted by gradient backpropagation combined with a moving exponential average to reduce the corresponding loss function.

[0027] 3. Advantages and effects:

[0028] This invention proposes an infrared weak target detection method based on a hybrid spatial modulation feature convolutional neural network. Starting from the multi-directional characteristic of weak targets exhibiting Gaussian distribution, a multi-directional fixed Gaussian attention mechanism is constructed using global attention and convolutional operations to suppress background and enhance target features. Based on the local grayscale characteristic of weak targets with high local brightness and significant abrupt changes from the background, a hybrid receptive field backbone network is constructed to further utilize the local neighborhood characteristics of weak targets, achieving feature extraction more suitable for this task. A cross-window attention mechanism is constructed to fuse low, medium, and high-level features, better preserving relevant features of small targets while extracting multi-scale features. The model design, based on the characteristics of infrared weak targets, demonstrates good interpretability and performance, and has broad application prospects. Attached Figure Description

[0029] Figure 1 This diagram illustrates the local characteristics of the target in this invention and the background regions that are easily interfered with by the detector. Region c represents the target region, and regions s1 to s3 are background regions that are easily interfered with by the detector.

[0030] Figure 2 This is the basic structure of a multi-directional fixed Gaussian kernel attention module.

[0031] Figure 3 This is a schematic diagram of the basic structure of a hybrid receptive field convolutional block, specifically a hybrid convolutional unit.

[0032] Figure 4 This is a flowchart illustrating the principle of the infrared weak target detection method based on a hybrid spatial modulation feature convolutional neural network proposed in this invention.

[0033] Figures 5a-5d The diagram illustrates the low-level and high-level features extracted by this invention. Figure 5a To input the raw infrared image, Figure 5b , 5c Figures 5d and 5d are schematic diagrams of the low-level, mid-level, and high-level features extracted by the three sets of hybrid receptive field convolutional blocks of this invention.

[0034] Figure 6 This is the basic structure of the cross-window attention module.

[0035] Figures 7a-7h The detection results of this invention in a real-world scenario are demonstrated; wherein, Figure 7a , 7b 7e and 7f are the original infrared images, with small targets marked by white squares. Figure 7c , 7d 7g and 7h are the detection results of the method of this invention. Detailed Implementation

[0036] To better understand the technical solution of the present invention, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0037] This invention relates to a method for detecting weak infrared targets based on a hybrid spatial modulation feature convolutional neural network. The specific steps of this method are as follows:

[0038] Step 1: Construct a multi-directional fixed Gaussian kernel attention mechanism to suppress the background while extracting target features from multiple directions for target feature enhancement;

[0039] Step 2: Construct a backbone network based on concatenated convolutional blocks with hybrid receptive fields to extract three sets of features from the enhanced shallow features;

[0040] Step 3: Construct a cross-sliding attention mechanism to fuse the three sets of features extracted from the backbone network through the cross-sliding window attention mechanism and stitch them together in the channel dimension; and then use multi-directional Gaussian kernel attention and convolutional layers with a kernel size of 3×3, expansion coefficient of 1, grouping of 1, and stride of 1 to perform pixel-by-pixel prediction to obtain the pixel-level probability prediction map of the entire image.

[0041] Step 4: Connect the above modules in sequence to build a convolutional neural network; and construct a loss function to train the network;

[0042] Output: The infrared image is processed using the trained neural network; after sufficient iterative training of the hybrid spatial modulation feature-based convolutional neural network using the training data, the trained network is used to detect target pixels.

[0043] Specifically, step one is as follows:

[0044] 1.1: Extracting shallow features and using multi-directional fixed Gaussian kernel attention for target feature enhancement. The network mainly uses convolutional units as basic components. Each convolutional unit consists of one convolutional layer, a batch normalization layer, and a Selu activation function operation. Parameters such as kernel size, expansion coefficients, number of groups, stride, and activation function type in the convolutional layers are adjusted as needed. First, the input image passes through a convolutional unit with a kernel size of 7×7, expansion coefficient of 1, number of groups of 1, and stride of 1, generating a shallow feature F with 16 channels. sTypically, an infrared image of a weak target can be considered to consist of three parts: the target, the background, and noise. I = B + T + N, where I represents the original image matrix, B represents the background matrix, T represents the target matrix, and N represents noise and other error matrices. To accurately separate the background and target, this invention proposes that weak targets can usually be modeled as unusual bright spots in the image with high contrast to the background, whose grayscale distribution exhibits characteristics similar to a two-dimensional Gaussian function, such as... Figure 1 As shown in region c, a multi-directional fixed Gaussian kernel is designed to effectively locate the target. Furthermore, the background may contain bright clouds or fog, which can easily interfere with the detection of weak targets, such as... Figure 1 As shown in regions s1 to s3, their grayscale distribution characteristics resemble those of weak targets, thus requiring the introduction of a background suppression mechanism to reduce interference from the background. To address these issues, this invention proposes a multi-directional fixed Gaussian kernel attention mechanism to focus on the extracted shallow features F. s Background suppression and target enhancement are performed to obtain the enhanced shallow features F. e The specific structure is as follows Figure 2 As shown. Considering the widespread distribution of clouds and fog in the background and the sparse distribution of targets, a global attention mechanism GCBlock is first introduced for background suppression. Then, a multi-directional fixed Gaussian kernel is constructed. The probability that a pixel is a target is measured by the magnitude of the gray-level difference between a pixel and its neighboring pixels in multiple directions. Spatial attention is extracted to enhance target features. This invention's multi-directional fixed Gaussian kernel attention first applies the input feature map F... s Background suppression was performed using GCBlock, a global attention mechanism with a channel dimension compression ratio of 0.25, to obtain the feature map F. c-attn Next, a pointwise convolutional layer with a kernel size of 1×1, a spread factor of 1, a group number of 1, and a stride of 1 is used to reduce the channel dimension of the feature image after background suppression to 8, resulting in feature map F. a Then the feature map F a The system is divided into 8 groups along the channel dimension, and 8 fixed convolutional kernels are used in parallel to compute the directional feature maps F. d The kernel size is fixed at 5×5. The remaining convolution kernels d i It is obtained by rotating d1 counterclockwise by i×45°. Then, two sets of cascaded pointwise convolutional units with a kernel size of 1×1, an expansion coefficient of 1, a group number of 1, and a stride of 1 are used to fully fuse the orientation feature map F. d The directional information from different channels is used to obtain the fused directional feature map F′. d Finally, F d With F′ dPointwise multiplication is performed, and pointwise convolutional units with a kernel size of 1×1, a spread factor of 1, a group number of 1, a stride of 1, and an activation function of sigmoid are used to reduce the channel dimension to 1, resulting in the multidirectional attention feature map F. d-attn Multi-directional attention feature map F d-attn With F c-attn The enhanced input feature map F is obtained by performing point-by-point multiplication. e .

[0045] Step two is as follows:

[0046] 2.1: Constructing a backbone network to extract features from the enhanced low-level features; the backbone network consists of three sets of hybrid receptive field convolutional blocks and convolutional units responsible for downsampling, alternating between them. Each hybrid receptive field convolutional block consists of a certain number of hybrid receptive field convolutional units and a global attention mechanism GCBlock responsible for background suppression, connected in series. Each hybrid receptive field convolutional unit consists of convolutional units with increasing kernel size, expansion coefficient of 1, number of groups equal to the number of input channels, and stride of 1; grouped expanded convolutional units with kernel size of 3×3, expansion coefficient of 2, number of groups of 4, and stride of 1; pointwise convolutional layers with kernel size of 1×1, expansion coefficient of 1, number of groups of 1, and stride of 1; and residual connections, in order to extract features at different scales. Its specific structure is as follows: Figure 3 As shown. The hybrid receptive field convolutional unit designed in this invention first divides the input features into four groups along the channel dimension. For each group, features are extracted using convolutional units with kernel sizes of 1×1, 3×3, 5×5, and 7×7, a scaling factor of 1, a grouping number equal to the number of input channels, and a stride of 1. These extracted features are then concatenated along the channel dimension. The processed features are then sequentially processed using a grouped extended convolutional unit with a kernel size of 3×3, a scaling factor of 2, a grouping of 4, and a stride of 1, and a pointwise convolutional layer with a kernel size of 1×1, a scaling factor of 1, a grouping number of 1, and a stride of 1. Finally, a residual connection is made with the input features to obtain the output features. The hybrid receptive field convolutional blocks are connected by a set of convolutional units with a kernel size of 3×3, a scaling factor of 2, a grouping number equal to the number of input channels, and a stride of 2, responsible for downsampling. The enhanced shallow feature F eThe first set of hybrid receptive field convolutional blocks G1, consisting of a hybrid receptive field convolutional unit and a global attention mechanism GCBlock, yields the corresponding output feature map F1. This is then downsampled by a set of convolutional units with a kernel size of 3×3, a scaling factor of 2, grouping based on the number of input channels, and a stride of 2, doubling the channel dimension to 32. Next, the second set of hybrid receptive field convolutional blocks G2, consisting of a hybrid receptive field convolutional unit and a global attention mechanism GCBlock, yields the corresponding output feature map F2. This is then downsampled by another set of convolutional units with a kernel size of 3×3, a scaling factor of 2, grouping based on the number of input channels, and a stride of 2, doubling the channel dimension to 64. Finally, the third set of hybrid receptive field convolutional blocks G3, consisting of three hybrid receptive field convolutional units and a global attention mechanism GCBlock, yields the corresponding output feature map F3. The specific implementation of the feature extraction process is as follows... Figure 4 As shown.

[0047] Step three is as follows:

[0048] 3.1: Cross-Sliding Window Attention Mechanism. In infrared small target detection tasks, low-level features reflecting details such as edges, shape, and texture are related to target edge segmentation, while high-level features containing more semantic information are related to target location determination and background suppression. For example, for a target like... Figure 5a The infrared image input shown is output as feature map F1, which mainly reflects the backbone network's extraction of low-level features. It retains relatively clear and accurate edge and texture features in the target, mountain background, and sea background in the image, such as... Figure 5b As shown; the output feature map F2 reflects the feature extraction by the intermediate layers of the backbone network. It retains the general outline of the target and background in the image, while further extracting and enhancing the target location information, such as... Figure 5c As shown; the output feature map F3 reflects the backbone network's extraction of high-level features. Its different channels contain parts with different semantics from the original infrared image, enabling the differentiation between target and background regions and between different background regions. However, its detailed information is blurred, such as... Figure 5d As shown. Therefore, it is necessary to effectively fuse low-level features with mid- and high-level features, and fully combine semantic and detailed information to achieve target segmentation and extraction. Considering that the local neighborhood of a weak target contains relatively rich multi-scale features, this invention designs a cross-sliding window attention mechanism for weak target features. This mechanism achieves feature fusion of different layers by dividing the target into windows of different sizes and calculating the cross-window attention. Its main component, the cross-window attention module, follows the formula Attn(X, Y) = softmax(norm(X)norm(Y)). TThe cross-window attention function CWA(X,Y) is calculated using the formulas: Linear(Y) = X + Mlp(Attn(X,Y)). Here, norm is the normalization function, softmax is the softmax activation function, B is the relative position offset, Linear is the linear projection function, Mlp is the multilayer perceptron function, and X and Y are the input feature matrices, respectively. T For the transpose of Y, such as Figure 6 As shown, the cross-sliding window attention mechanism first uses a pointwise convolutional unit with a kernel size of 1×1, an expansion coefficient of 1, a group number of 1, and a stride of 1 to compress the channel dimension of the backbone network output feature maps F3 and F3 to 16, obtaining the corresponding feature maps F′2 and F′3. Then, bilinear interpolation is used to restore F′2 and F′3 to their original input size, obtaining the corresponding feature maps F″2 and F″3. Subsequently, the cross-window attention module is used to calculate the cross-window attention of F′2 and F′3 with respect to F1, respectively; the inputs F′2 and F′3 are then divided into 8×8 non-overlapping windows F′2 and F′3. 2-window A non-overlapping 4×4 window F′ 3-window Divide the input F1 into 16×16 non-overlapping windows F′ 1window ;Calculate F′ using the cross-window attention module respectively 1window With F′ 2window and F′ 3window Attention-enhanced feature map CWA(F′) 1window F′ 2window ) and CWA(F′ 1window F′ 3window Then, the attention-enhanced feature map CWA(F′) is applied. 1window F′ 2window ), CWA(F′ 1window F′ 3window ) Move 8 pixels to the lower right and divide it into 16×16 non-overlapping windows F′ 1-2window F′ 1-3window ;Shift F′2 and F′3 downwards and to the right by 4 and 2 pixels respectively, and divide them into 8×8 non-overlapping windows F′ 2shifted-window A non-overlapping 4×4 window F′ 3shifted-window ;Calculate F′ using the cross-window attention module respectively 1-2window With F′ 2s hi fted-window Cross-window attention CWA(F′) 1-2window F′ 2shifted-window ) and F′ 1-3window With F′ 3shifted-window Cross-window attention CWA(F′) 1-3window F′ 1shifted-window), shift 8 pixels to the upper left to obtain the corresponding cross-sliding window attention feature map. Finally, the attention feature map of the cross-sliding window is used. The fused feature map is obtained by performing residual connections with F″2 and F″3. F1, Multi-scale output feature map F is obtained by concatenating along the channel dimension. m .

[0049] 3.2: Detection is performed using a lightweight output layer consisting of multi-directional fixed Gaussian kernel attention and convolutional layers with a kernel size of 3×3, a spread factor of 1, a grouping of 1, and a stride of 1. The multi-scale output feature map F... m After further attention enhancement using a multi-directional fixed Gaussian kernel, the channel dimension is reduced to 1 using a convolutional layer with a kernel size of 3×3, a spread factor of 1, a grouping of 1, and a stride of 1. Then, after processing with a sigmoid activation function, a probability prediction map at the pixel level of the entire image is output.

[0050] Step four is as follows:

[0051] 4.1: Concatenate the modules proposed in steps one through three to build a convolutional neural network, such as... Figure 4 As shown. The loss function consists of the Intersection over Union (IOU) loss, L = L IOU The intersection-union ratio (IU) refers to the overlap rate between the predicted and ground truth regions; it's the ratio of their intersection to their union. When training a network for object detection, the ideal scenario is that the predicted and ground truth regions completely overlap, meaning the IU equals 1. Therefore, in practice, the IU value is always between 0 and 1, and a higher value indicates more accurate detection. Thus, the IU loss is defined. Where area(predict) is the target region predicted by the method of this invention, area(trut) is the area of ​​the real target region, ∩ is the intersection operation of sets, and ∪ is the union operation of sets. After giving the above definition of the loss function, the infrared image is input into the convolutional neural network to obtain the probability prediction map and perform a pixel-by-pixel multiplication with the labeled real result map to obtain the overlap result of the predicted target region and the real target region, i.e., area(predict)∩area(trut); based on this, the number of pixels in the real target region, the predicted target region, and the overlapping area of ​​the two are summed to calculate the intersection-union ratio loss.

[0052] 4.2: This invention uses the Adamw optimizer for optimization. The initial learning rate of the network is 0.0002, and the weight decay coefficient is 10. -3During training, the learning rate is adaptively updated, and network parameters are adjusted to reduce the corresponding loss function through gradient backpropagation combined with a moving exponential average. In this process, gradient descent is used for backpropagation, and the chain rule is used to update the parameters by taking the partial derivative of the loss function with respect to a specific network parameter. Where θ i The network parameters before backpropagation, θ′ i Here, η represents the network parameters updated after backpropagation, η is the learning rate, and L is the loss function.

[0053] Figures 7a-7h This is an application of the invention in a real infrared scene; the location of weak targets is marked with a white frame. Figure 7c , 7d 7g and 7h represent the corresponding detection results. The images used in the experiment came from different infrared scenes, most of which contained very faint and small targets, making it difficult to extract effective texture information. Furthermore, the backgrounds contained complex interference such as clouds, vegetation, and noise. However, the experimental results not only effectively eliminated noise interference and accurately detected the position and shape of the targets, but also demonstrated advantages in computation time, achieving rapid and accurate target detection. This fully illustrates the high efficiency of the invention, which can be widely applied to various infrared weak target detection systems, possessing broad market prospects and application value.

Claims

1. A method for detecting weak infrared targets based on a hybrid spatial modulation feature convolutional neural network, characterized in that, Includes the following steps: Step 1: Construct a multi-directional fixed Gaussian kernel attention mechanism to suppress the background while extracting target features from multiple directions for target feature enhancement; Step 2: Construct a backbone network based on concatenated convolutional blocks with hybrid receptive fields to extract three sets of features from the enhanced shallow features; Step 3: Construct a cross-sliding attention mechanism to fuse the three sets of features extracted from the backbone network through the cross-sliding window attention mechanism and stitch them together in the channel dimension; and then use multi-directional Gaussian kernel attention and convolutional layers with a kernel size of 3×3, expansion coefficient of 1, grouping of 1, and stride of 1 to perform pixel-by-pixel prediction to obtain the pixel-level probability prediction map of the entire image. Step 4: Connect the modules sequentially to build a convolutional neural network; and construct a loss function to train the network; Output: The infrared image is processed using the trained neural network; after sufficient iterative training of the convolutional neural network based on hybrid spatial modulation features using the training data, the trained network is used to detect target pixels. Step three is as follows: A cross-sliding window attention mechanism is designed for weak target features. This mechanism achieves feature fusion across different layers by dividing the screen into windows of different sizes and calculating the cross-window attention. The cross-window attention module follows the formula... and Calculate cross-window attention In the formula For normalization function, The softmax activation function is used. This is a relative position offset. It is a linear projection function. For multilayer perceptron functions, , These are the input feature matrices, for Transpose of; The cross-sliding window attention mechanism first utilizes a convolution kernel size of Pointwise convolutional units with an expansion factor of 1, a group size of 1, and a stride of 1 output feature maps from the backbone network. , The channel dimension is compressed to 16 to obtain the corresponding feature map. , Recovery using bilinear interpolation , The corresponding feature map is obtained by adjusting the original input size. , Subsequently, the attention module of the cross window was used to calculate... , for Cross-window attention; input , Divide into 8x8 non-overlapping windows. 4x4 non-overlapping windows , will input Divided into Non-overlapping windows ;Calculate using the cross-window attention module respectively and and Attention-enhanced feature maps and Then enhance the attention feature map , Shift 8 pixels to the lower right and divide into Non-overlapping windows , ;Will , Shift 4 and 2 pixels to the lower right respectively and divide them into two groups. Non-overlapping windows and Non-overlapping windows ;Calculate using the cross-window attention module respectively and Cross-window attention and and Cross-window attention Shifting the image 8 pixels to the upper left yields the corresponding cross-sliding window attention feature map. , Finally, the attention feature map of the cross sliding window is... , and , By performing residual connections, a fused feature map is obtained. , ,Will , , Multi-scale output feature maps are obtained by concatenating them along the channel dimension. .

2. The infrared weak target detection method based on a hybrid spatial modulation feature convolutional neural network according to claim 1, characterized in that: Step one is as follows: Shallow features are extracted and target feature enhancement is performed using multi-directional fixed Gaussian kernel attention. The network mainly uses convolutional units as basic components. Each convolutional unit consists of a convolutional layer, a batch normalization layer, and a Selu activation function operation. The parameters of the convolutional kernel size, expansion coefficient, number of groups, stride, and activation function type in the convolutional layer are adjusted as needed.

3. The infrared weak target detection method based on a hybrid spatial modulation feature convolutional neural network according to claim 1 or 2, characterized in that: First, the input image is processed by a convolution kernel of size [size missing]. A convolutional unit with a spread factor of 1, a group size of 1, and a stride of 1 generates shallow features with 16 channels. ; for the extracted shallow features Background suppression and target enhancement are performed to obtain enhanced shallow features. First, a global attention mechanism GCBlock is introduced to suppress the background. Then, a multi-directional fixed Gaussian kernel is constructed. The probability of a point being a target is measured by the gray-level difference between a pixel and its neighboring pixels in multiple directions. Spatial attention is extracted to enhance the target features.

4. The infrared weak target detection method based on a hybrid spatial modulation feature convolutional neural network according to claim 3, characterized in that: Multi-directional fixed Gaussian kernel attention, firstly for the input feature map Background suppression was performed using GCBlock, a global attention mechanism with a channel dimension compression ratio of 0.25, to obtain the feature map. Secondly, the kernel size is used. A pointwise convolutional layer with a spread factor of 1, a group number of 1, and a stride of 1 reduces the channel dimension of the feature image after background suppression to 8, resulting in a feature map. Then the feature map The system is divided into 8 groups along the channel dimension, and directional feature maps are calculated in parallel using 8 fixed convolutional kernels. Fixed kernel size convolution kernel The remaining convolution kernels Depend on Rotate counterclockwise We obtain; then use a convolution kernel with a size of Two concatenated pointwise convolutional units with an expansion factor of 1, a group number of 1, and a stride of 1 fully fuse the directional feature maps. The directional information from different channels is used to obtain a fused directional feature map. Finally, and Perform pointwise multiplication and use a convolution kernel size of A pointwise convolutional unit with an expansion coefficient of 1, a group size of 1, a stride of 1, and an activation function of sigmoid reduces the channel dimension to 1, resulting in a multi-directional attention feature map. Multi-directional attention feature map and The enhanced input feature map is obtained by performing point-by-point multiplication. .

5. The infrared weak target detection method based on a hybrid spatial modulation feature convolutional neural network according to claim 1, characterized in that: Step two is as follows: A backbone network is constructed to extract features from the enhanced low-level features. The backbone network consists of three sets of hybrid receptive field convolutional blocks and convolutional units responsible for downsampling, alternating between them. Each hybrid receptive field convolutional block comprises a certain number of hybrid receptive field convolutional units and a global attention mechanism (GCBlock) responsible for background suppression, cascaded together. Each hybrid receptive field convolutional unit consists of convolutional units with increasing kernel size, a expansion coefficient of 1, a grouping number equal to the number of input channels, and a stride of 1; and grouped expanded convolutional units with a kernel size of 3×3, a expansion coefficient of 2, a grouping number of 4, and a stride of 1. It consists of pointwise convolutional layers with an expansion factor of 1, a group number of 1, a stride of 1, and residual connections, in order to extract features at different scales.

6. The infrared weak target detection method based on a hybrid spatial modulation feature convolutional neural network according to claim 5, characterized in that: The hybrid receptive field convolutional unit first divides the input features into four groups along the channel dimension. For each group, features are extracted using convolutional units with kernel sizes of 1×1, 3×3, 5×5, and 7×7, a spread factor of 1, a group number equal to the corresponding number of input channels, and a stride of 1. These extracted features are then concatenated along the channel dimension. The processed features are then sequentially expanded using a grouped convolutional unit with a kernel size of 3×3, a spread factor of 2, a group number of 4, and a stride of 1, and a convolutional unit with a kernel size of... The input features are processed by a pointwise convolutional layer with a scaling factor of 1, a group size of 1, and a stride of 1, followed by a residual connection with the input features to obtain the output features. The convolutional blocks with mixed receptive fields are connected by a set of 3×3 convolutional units with a scaling factor of 2, grouped according to the number of input channels, and a stride of 2, responsible for downsampling. The enhanced shallow features are then processed. The first set of hybrid receptive field convolutional blocks consists of a hybrid receptive field convolutional unit and a global attention mechanism GCBlock. Obtain the corresponding output feature map Then, it is downsampled through a set of convolutional units with a kernel size of 3×3, a spread factor of 2, a grouping of input channels, and a stride of 2, doubling the channel dimension to 32. Subsequently, it passes through a second set of hybrid receptive field convolutional blocks consisting of a hybrid receptive field convolutional unit and a global attention mechanism GCBlock. Obtain the corresponding output feature map Then, it is downsampled through a set of convolutional units with a kernel size of 3×3, a spread factor of 2, a grouping of input channels, and a stride of 2, doubling the channel dimension to 64. Finally, it passes through a third set of hybrid receptive field convolutional blocks consisting of three hybrid receptive field convolutional units and a global attention mechanism GCBlock. Obtain the corresponding output feature map .

7. The infrared weak target detection method based on a hybrid spatial modulation feature convolutional neural network according to claim 1, characterized in that: Detection is performed using a lightweight output layer consisting of multi-directional fixed Gaussian kernel attention and convolutional layers with a kernel size of 3×3, a spread factor of 1, a grouping of 1, and a stride of 1; multi-scale output feature maps are then processed. After further attention enhancement using a multi-directional fixed Gaussian kernel, the channel dimension is reduced to 1 using a convolutional layer with a kernel size of 3×3, a spread factor of 1, a grouping of 1, and a stride of 1. Then, after processing with a sigmoid activation function, a probability prediction map at the pixel level of the entire image is output.

8. The infrared weak target detection method based on a hybrid spatial modulation feature convolutional neural network according to claim 1, characterized in that: Step four is as follows: The loss function consists of the Intersection over Union (IOU) loss. The intersection-union ratio (IU) refers to the overlap rate between the predicted and ground truth regions; it is the ratio of their intersection to their union. When training a network for object detection, the ideal situation is for the predicted and ground truth regions to completely overlap, i.e., the IU equals 1. Therefore, the IU value is always between 0 and 1, and a larger value indicates a more accurate detection performance. Thus, the IU loss is defined. ,in For the predicted target area, This represents the actual area of ​​the target region. For set intersection operation, For set union operations; after defining the loss function above, the infrared image is input into the convolutional neural network to obtain the probability prediction map. This prediction map is then multiplied pixel-by-pixel with the labeled ground truth map to obtain the overlap result between the predicted target region and the ground truth target region. ; The number of pixels in the real target region, the predicted target region, and the overlapping region of the two are summed separately, and then the crossover ratio loss is calculated.

9. The infrared weak target detection method based on a hybrid spatial modulation feature convolutional neural network according to claim 8, characterized in that: The Adamw optimizer was used for optimization, with an initial learning rate of 0.0002 and a weight decay factor of [value missing]. During training, the learning rate is adaptively updated, and network parameters are adjusted to reduce the corresponding loss function through gradient backpropagation combined with a moving exponential average. In this process, gradient descent is used for backpropagation, and the chain rule is used to update the parameters by taking the partial derivative of the loss function with respect to a specific network parameter. ,in These are the network parameters before backpropagation. For the network parameters updated via backpropagation, For learning rate, This is the loss function.

Citation Information

Patent Citations

  • Infrared weak and small target detection method for constructing convolutional neural network by using multi-directional features

    CN114821018A