Infrared dim small target detection method combining single-scale network feature extraction and fusion
Patent Information
- Application Number
- CN202411148284.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-08-21
AI Technical Summary
相比之下,在单尺度架构没有使用下采样,能够保留细节信息,但是感受野十分有限
[0041]本发明的有益效果为:本发明提供的一种结合单尺度网络特征提取与融合的红外弱小目标检测方法能够有效解决小目标在下采样过程中容易淹没在背景中以及容易丢失细节信息的问题。不同与传统但阶段的模型,该模型有效地结合了编码器-解码器架构和单尺度无上采样的子网络,以整合不同框架的优势,完成精确的红外弱小目标检测。方法构建了一个类似U-Net的网络,并设计了一个基于大核卷积和扩展卷积的多尺度特征提取模块来优化特征提取并引入了不进行下采样的子网有效地保留细节信息,并完成特征融合。方法在SIRST和IRSTD-1k两个公共数据集上都能够达到良好的小目标检测性能,相较于其他前沿红外弱小目标检测方法具有明显的性能提升。
Smart Images

Figure CN119152228B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, specifically relating to an infrared weak target detection method that combines single-scale network feature extraction and fusion. Background Technology
[0002] Infrared technology is widely used in various fields, including infrared guidance, missile defense systems, urban security monitoring and early warning, ship monitoring, and intelligent maritime defense, due to its strong adaptability to changes in lighting and environment, as well as its unique cloud-penetrating capabilities. Given its enormous potential in various fields, infrared technology has become a focus of attention for both academia and industry.
[0003] However, since infrared images are usually obtained through long-range imaging systems, the objects of interest typically occupy only a small portion of the image. They are often easily obscured by background noise, and the lack of clear color, texture, and structural features makes the detection of small infrared targets extremely challenging, with false positives and false negatives being particularly pronounced.
[0004] Currently, deep learning-based infrared weak target detection models have demonstrated superior performance in infrared weak target detection. Based on their output, they can be divided into two categories: detection-based methods and segmentation-based methods. Segmentation-based IRSTD methods typically produce binary masks as output. In these binary masks, 1 represents pixels belonging to the target, and 0 represents background pixels. The result of detection-based infrared weak target detection models is a bounding box that captures the location and size of the detected target in the image.
[0005] However, regardless of the deep learning-based model, the encoder-decoder architecture is extremely common. But because the targets are very small, after undergoing downsampling operations in mainstream encoder-decoder networks, small targets not only lose their spatial information but may also become blurred in the background, leading to missed detections. Current networks focused on infrared weak target detection utilize various methods to mitigate the problems caused by downsampling and reduce target feature loss, but downsampling still exists in existing networks, failing to fundamentally solve the problem. In contrast, single-scale architectures do not use downsampling and can preserve detailed information, but their receptive field is very limited. Overcoming the inherent shortcomings of these architectures remains a problem that needs to be solved. Summary of the Invention
[0006] To address the shortcomings of current deep learning-based infrared weak target detection methods, this invention provides an infrared weak target detection method that combines single-scale network feature extraction and fusion. This method effectively combines an encoder-decoder architecture and a single-scale architecture to leverage the advantages of both, extracting more valuable information from infrared images and enabling more accurate detection of small targets.
[0007] The technical solution of this invention is:
[0008] An infrared weak target detection method combining single-scale network feature extraction and fusion includes the following steps:
[0009] S1. Collect infrared images for training;
[0010] S2. Construct an object detection model, including an encoder, decoder, feature enhancement module, and single-scale network. Before inputting training data into the object detection model, it needs to be preprocessed. The preprocessing method is to divide the original training image into four non-overlapping image blocks to obtain the first input image I1, divide the original training image into two non-overlapping image blocks to obtain the second input image I2, and use the original training image as the third input image I3.
[0011] The model iterates through the encoder, feature enhancement module, and decoder. In the first iteration, the encoder's input is the first input image I1, and in the second iteration, the input is the decoder's output from the first iteration and the second input image I2. In these two iterations, the encoder downsamples the input image to capture high-level features. The encoder consists of three downsampling blocks, each consisting of a two-channel attention module and a downsampling operation. The encoder's output is then enhanced by the feature enhancement module and input to the decoder for decoding.
[0012] In the first iteration, the feature image patch input from the encoder to the feature enhancement module is defined as X. i , i∈{1,2,3,4}, X i The following processing is performed in the feature enhancement module:
[0013] X i O is obtained after the original convolution. i,1 :
[0014] O i,1 =f 3×3 (X i )
[0015] Where f 3×3 (·) represents the original 3x3 convolution;
[0016] X i O is obtained after dilation convolution. i,2 :
[0017] O i,2 =g 3×3 (X i dilation=3)
[0018] Where g 3×3(·, dilation=3) represents a dilated convolution with a dilation rate of 3;
[0019] X i The O function is obtained by sequentially applying a 1×1 convolution, a large kernel attention module, GELU, and another 1×1 convolution. i,3 :
[0020]
[0021] Here, LKA represents a large kernel attention module, and GELU indicates that the activation function used is a Gaussian error linear unit. In LKA, the idea of decomposing the large kernel convolution is used to decompose the 21x21 convolution into a combination of 5x5 depthwise convolution, 7x7 depthwise dilated convolution, and 1x1 convolution to capture long-range dependencies. Let ψ represent the i-th 1×1 convolutional layer. 5×5 (·) represents a 5×5 depthwise convolution, φ 7×7 (·) represents a 7×7 depthwise dilated convolution with a dilation rate of 3;
[0022] Then O i,1 O i,2 O i,3 and X i Connect to get F i :
[0023] F i =concatenate(O i,1 O i,2 O i,3 ,X i )
[0024] Finally F i After a 1×1 convolution, it is then combined with X. i Fusion yields P i :
[0025] P i =X i +f 1×1 (F i )
[0026] Since the input images I1 and I2 are different, in the second iteration, the image block M input by the encoder to the feature enhancement module is... i For each image patch, i∈{1,2}, the operation performed and the operation performed on X in the first iteration. i The operation is the same. The output P of the feature enhancement module iThe input is fed into the decoder, which consists of three upsampling blocks. Each upsampling block includes two channel attention modules and an upsampling operation. The output of the first decoder is added element-wise to the second input image I2 and used as the input of the second encoder. The output of the second decoder is then fed into a single-scale non-downsampling subnet together with the third input image I3 for feature extraction and fusion.
[0027] The single-scale network is a subnetwork that does not perform downsampling at all. This subnetwork consists of four feature extraction and fusion modules. The input X of the first feature extraction and fusion module is... L0 The third input image is I3, and the input is X. H0 The output of the second decoder is X. L1 and X H1 :
[0028] X LH0 =Conv1(X L0 )
[0029] X HL0 =Conv2(X H0 )
[0030] X HH0 =Conv3(X H0 )
[0031]
[0032] X L1 =BN(ReLU(X) LL0 +X HL0 ))
[0033] X H1 =BN(ReLU(X) LH0 +X HH0 ))
[0034] Where BN represents batch normalization, and ReLU refers to the correction linear unit that introduces nonlinearity into the model. i∈{1,2,3} represents a 1x1 convolution, while Conv i (·), i∈{1,2,3} represents a 3x3 convolution;
[0035] The input X of the second feature extraction and fusion module L1 and X H1 This is the output of the first feature extraction and fusion module, and the output X of the second feature extraction and fusion module. L2 X H2 and X L0 X H0 The input X to the third feature extraction and fusion module is obtained by adding elements one by one.L3 X H3 The output X of the third feature extraction and fusion module L4 X H4 It is the input to the fourth feature extraction and fusion module; the final output H of the spatial and semantic feature extraction and fusion module is:
[0036] H = concatenate(X) L5 +X L3 ,X H5 +X H3 )
[0037] Where X L5 X H5 This is the output of the fourth feature extraction and fusion module;
[0038] The output features are processed using the sigmoid function. Pixels larger than the threshold are considered target pixels, and pixels smaller than the threshold are considered background pixels.
[0039] S3. Use the training data to train the constructed target detection model to obtain the trained target detection model;
[0040] S4. Input the image to be detected into the trained target detection model to perform infrared weak target detection.
[0041] The beneficial effects of this invention are as follows: The infrared weak target detection method combining single-scale network feature extraction and fusion provided by this invention can effectively solve the problems of small targets being easily submerged in the background and losing detailed information during downsampling. Unlike traditional single-stage models, this model effectively combines an encoder-decoder architecture and a single-scale sub-network without upsampling to integrate the advantages of different frameworks and achieve accurate infrared weak target detection. The method constructs a network similar to U-Net and designs a multi-scale feature extraction module based on large kernel convolution and extended convolution to optimize feature extraction and introduces a sub-network that does not perform downsampling to effectively preserve detailed information and complete feature fusion. The method achieves good small target detection performance on both the SIRST and IRSTD-1k public datasets, showing a significant performance improvement compared to other cutting-edge infrared weak target detection methods. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of the infrared weak target detection model structure combining single-scale network feature extraction and fusion in an embodiment of the present invention;
[0043] Figure 2 This is a schematic diagram of the feature enhancement module based on large kernel convolution in an embodiment of the present invention;
[0044] Figure 3 This is a schematic diagram of the spatial and semantic feature extraction and fusion module in a single-scale subnet in an embodiment of the present invention;
[0045] Figure 4 This is a schematic diagram of the infrared weak target detection results in an embodiment of the present invention. Detailed Implementation
[0046] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.
[0047] refer to Figure 1 In one embodiment of the present invention, the present invention provides an infrared weak target detection method combining single-scale network feature extraction and fusion, comprising the following steps:
[0048] Step 1: Collect relevant public datasets (e.g., select and collect images from the SIRST and IRSTD-1K public datasets, and perform data augmentation on the images, such as flipping and cropping, to enrich the training set and improve the generalization ability of the method), and train the model. The images in the dataset are normalized by calculating the mean and variance of each dataset, and then subtracting the mean from the pixel value of each image by channel and dividing by the variance. This can accelerate the convergence speed of the model.
[0049] Step 2: This invention employs a multi-image patch hierarchical structure to further process the image preprocessed in Step 1. This method divides the input infrared image into multiple non-overlapping image patches of different sizes, and inputs them separately into the model. This allows the model to sequentially focus on regions of different sizes, from smaller image patches to the entire original image. This design enables the model to gradually extract features from local details and ultimately combine global information for comprehensive analysis.
[0050] Step 3: Encode and decode the input infrared image patches using a U-Net-like encoder-decoder network to extract image information over a wider receptive field. This includes the following steps:
[0051] Step 3.1: Before the image patches are fed into the backbone network, the features of small infrared targets are initially extracted and enhanced using original convolution and channel attention modules, while simultaneously capturing the basic structural information of the image. This preprocessing preserves the image's original resolution but increases the number of channels, providing richer information for subsequent feature fusion.
[0052] Step 3.2: A U-Net-like network is used as the backbone. The encoder progressively downsamples the input image to expand the receptive field and capture global information; while the decoder progressively upsamples the features and combines low-level and high-level features through skip connections. Unlike the traditional U-Net, this network uses channel attention modules in both the encoder and decoder to replace some convolutional layers, further improving the efficiency and accuracy of feature extraction. The encoder and decoder each consist of three downsampling blocks and three upsampling blocks, each containing two channel attention modules and one sampling operation (downsampling or upsampling).
[0053] Step 3.3: A feature enhancement module based on large kernel convolution is specifically integrated, and its specific structure is as follows: Figure 2 As shown, this module captures and fuses contextual features at different scales for each input image patch through multi-scale convolution operations. Specifically, for each image patch, the module first uses standard convolution to capture short-range dependencies, then dilated convolution to capture mid-range dependencies, and finally a strategy of decomposing large kernel convolutions to capture long-range dependencies. These features at different scales are concatenated, dimensionality-reduced through 1x1 convolution, and fused with the original feature map to generate updated image patch features. Through this multi-scale feature enhancement and fusion strategy, the module can more effectively extract and integrate key information in the image, thereby improving the detection performance of infrared weak targets.
[0054] Step 4: By means of Figure 3 The single-scale subnet shown uses spatial and semantic feature extraction and fusion modules to further extract and fuse image features, thereby avoiding the loss of detail information in the traditional downsampling process, and finally generating output features, such as... Figure 4 As shown, the specific processing of this single-scale subnet is as follows:
[0055] Step 4.1: In this stage, the input image first undergoes a preprocessing step. This includes preliminary feature extraction of the image using original convolution and channel attention modules.
[0056] Step 4.2: The third-stage subnet consists of four feature extraction and fusion modules. These modules are responsible for further extracting and fusing feature information from different sources while maintaining high resolution. The features from Step 3 are rich in semantic information due to the expansion of the receptive field, while the original image is considered to be rich in spatial information.
[0057] Step 4.3: Process the feature map output from the third stage using the sigmoid function. The sigmoid function maps feature values to the range of 0 to 1, and then segments pixels into target pixels and background pixels by setting a threshold. Pixels with values greater than the threshold are considered targets, while pixels with values less than the threshold are considered background.
[0058] Step 5: Train the network models built in S2 to S4 end-to-end using the dataset provided in S1, and optimize the model by setting the loss function and introducing a supervision mechanism until its performance is stable. Then, use the test set to evaluate the trained model. The detection results of this invention are compared with the detection results of the other six methods as follows:
[0059] Table 1 compares the detection results of the infrared weak target detection method used with six other methods.
[0060]
[0061] The results show that the infrared weak target detection method proposed in this invention, which combines single-scale network feature extraction and fusion, has higher accuracy and lower false detection rate.
Claims
1. A method for detecting weak infrared targets by combining single-scale network feature extraction and fusion, characterized in that, Includes the following steps: S1. Collect infrared images for training; S2. Construct an object detection model, including an encoder, decoder, feature enhancement module, and single-scale network; Before the training data is input into the object detection model, it needs to be preprocessed. The preprocessing method is to divide the original training image into four non-overlapping image blocks to obtain the first input image I1, divide the original training image into two non-overlapping image blocks to obtain the second input image I2, and use the original training image as the third input image I3. The model iterates through the encoder, feature enhancement module, and decoder. In the first iteration, the encoder's input is the first input image I1, and in the second iteration, the input is the decoder's output from the first iteration and the second input image I2. In these two iterations, the encoder downsamples the input image to capture high-level features. The encoder consists of three downsampling blocks, each consisting of a two-channel attention module and a downsampling operation. The encoder's output is then enhanced by the feature enhancement module and input to the decoder for decoding. In the first iteration, the feature image patch input from the encoder to the feature enhancement module is defined as X. i , i∈{1,2,3,4},X i The following processing is performed in the feature enhancement module: X i O is obtained after the original convolution. i,1 : O i,1 =f 3×3 (X i ) Where f 3×3 (·) represents the original 3x3 convolution; X i O is obtained after dilation convolution. i,2 : O i,2 =g 3×3 (X i ,dilation=3) Where g 3×3 (·, dilation=3) represents a dilated convolution with a dilation rate of 3; X i The O function is obtained by sequentially applying a 1×1 convolution, a large kernel attention module, GELU, and another 1×1 convolution. i,3 : Here, LKA represents a large kernel attention module, and GELU indicates that the activation function used is a Gaussian error linear unit. In LKA, the idea of decomposing the large kernel convolution is used to decompose the 21x21 convolution into a combination of 5x5 depthwise convolution, 7x7 depthwise dilated convolution, and 1x1 convolution to capture long-range dependencies. Let ψ represent the i-th 1×1 convolutional layer. 5×5 (·) represents a 5×5 depthwise convolution, φ 7×7 (·) represents a 7×7 depthwise dilated convolution with a dilation rate of 3; Then O i,1 O i,2 O i,3 and X i The connection yields F i : F i =concatenate(O i,1 ,O i,2 ,O i,3 ,X i ) Finally F i After a 1×1 convolution, it is then combined with X. i Fusion yields P i : P i =X i +f 1×1 (F i ) Since the input images I1 and I2 are different, in the second iteration, the encoder inputs image block M to the feature enhancement module. i For each image patch, i∈{1,2}, the operation performed and the operation performed on X in the first iteration. i The operation is the same; the output P of the feature enhancement module is the same. i The input is fed into the decoder, which consists of three upsampling blocks. Each upsampling block includes two channel attention modules and an upsampling operation. The output of the first decoder is added element-wise to the second input image I2 and used as the input of the second encoder. The output of the second decoder is then fed into a single-scale non-downsampling subnet together with the third input image I3 for feature extraction and fusion. The single-scale network is a subnetwork that does not perform downsampling at all. This subnetwork consists of four feature extraction and fusion modules. The input X of the first feature extraction and fusion module is... L0 The third input image is I3, and the input is X. H0 The output of the second decoder is X. L1 and X H1 : X LH0 =Conv1(X L0 ) X HL0 =Conv2(X H0 ) X HH0 =Conv3(X H0 ) X L1 =BN(ReLU(X LL0 +X HL0 )) X H1 =BN(ReLU(X LH0 +X HH0 )) Where BN represents batch normalization, and ReLU refers to the correction linear unit that introduces nonlinearity into the model. i∈{1,2,3} represents a 1x1 convolution, while Conv i (·), i∈{1,2,3} represents a 3x3 convolution; The input X of the second feature extraction and fusion module L1 and X H1 This is the output of the first feature extraction and fusion module, and the output X of the second feature extraction and fusion module. L2 X H2 and X L0 X H0 The input X to the third feature extraction and fusion module is obtained by adding elements one by one. L3 X H3 The output X of the third feature extraction and fusion module L4 X H4 It is the input to the fourth feature extraction and fusion module; the final output H of the spatial and semantic feature extraction and fusion module is: H=concatenate(X L5 +X L3 ,X H5 +X H3 ) Where X L5 X H5 This is the output of the fourth feature extraction and fusion module; The output features are processed using the sigmoid function. Pixels larger than the threshold are considered target pixels, and pixels smaller than the threshold are considered background pixels. S3. Use the training data to train the constructed target detection model to obtain the trained target detection model; S4. Input the image to be detected into the trained target detection model to perform infrared weak target detection.