An infrared small target detection method using band context and hole enhancement deep unfolding network
Patent Information
- Application Number
- CN202610380640.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-26
- Publication Date
- 2026-09-08
AI Technical Summary
[0006]然而,现有深度展开方法在内部模块设计上仍存在明显局限
[0023]To address the aforementioned technical problems, this invention proposes an infrared small target detection method based on strip context and dilated context sparse (DCSM) deep unfolding networks. This method comprises a Strip Context Background Module (SCBM), a Dilated Context Sparse Module (DCSM), and an Adaptive Residual Fusion Module (ARFM). SCBM uses direction-aware striped pooling to capture a wide range of anisotropic background clutter along both horizontal and vertical directions, achieving more thorough background suppression. DCSM introduces dilated convolution instead of standard convolution, expanding the receptive field from 13×13 to 25×25 without adding any additional parameters, fully utilizing the local contrast features of small targets to achieve accurate target recovery. ARFM replaces traditional element-wise addition with channel-level stitching, completely preserving the independent discriminative features of both background and target components, preventing information loss and cumulative degradation across unfolding stages.
Smart Images

Figure CN122714741A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method for infrared small target detection using strip context and hole-enhanced depth unrolling networks. Background Technology
[0002] Infrared Small Target Detection (IRSTD) is a core technology in applications such as remote sensing monitoring, sea surface surveillance, early warning, and environmental monitoring [1-3]. Due to the unique advantages of thermal imaging, such as being passive, light-independent, and capable of long-distance detection, infrared detection systems can operate reliably under complex all-weather conditions [4-5]. However, the core challenge of IRSTD lies in the inherent visual characteristics of targets: distant targets often occupy only a few pixels, have extremely low signal-to-noise ratios, and lack recognizable texture or shape features [6]. At the same time, they are deeply embedded in complex and heterogeneous background clutter (such as ground buildings, vegetation, and urban scenes), which makes IRSTD a difficult and persistent research problem [7-9].
[0003] Before the era of deep learning, IRSTD was mainly dominated by traditional model-driven methods, which can be roughly divided into three paradigms. Filter-based methods [10-11] apply morphological or spatial frequency transformations to suppress background noise, but are sensitive to non-uniform clutter. Local contrast-based methods [12-13] enhance target saliency by measuring the relative intensity difference between candidate regions and their spatial neighborhoods. Although computationally simple, they have limited adaptability to target scale changes. Low-rank sparse decomposition (LRSD)-based methods [4,8,9] reformulate the detection problem as a matrix factorization problem, taking advantage of the low rank of the background and the sparsity of point targets to achieve stronger robustness under low signal-to-clutter conditions. Although these paradigms are physically interpretable and mathematically tractable, they all have a fundamental bottleneck: they rely on fixed hand-crafted priors and statically tuned hyperparameters, which makes them inherently inflexible. As real-world infrared scenes become structurally complex and heterogeneous, the growing mismatch between rigid model assumptions and dynamic background statistics leads to inevitable performance degradation and poor cross-scene generalization ability
[14] .
[0004] To combine the interpretability of mathematical models with the powerful expressiveness of deep learning, Deep Unfolding Networks (DUNs) were introduced into the IRSTD field. DUNs unfold iterative optimization algorithms into trainable neural network stages, ensuring a one-to-one correspondence between the network structure and the optimization process.
[0005] In the prior art, a deep unfolding network (RPCANet) based on RPCA (Robust Principal Component Analysis) [4] unfolds the RPCA optimization process into multiple learnable stages, and for the first time realizes explicit iterative separation of background and target in infrared images. On this basis, another dynamic robust principal component analysis network (DRPCA-Net) [5] further introduces dynamic residual groups and adaptive iterative parameter estimation mechanism, which improves decomposition accuracy and network expressive power. The above deep unfolding methods successfully embed low-rank sparse priors into the network topology, and achieve competitive detection performance while maintaining high interpretability and compact model size.
[0006] However, existing deep unrolling methods still have significant limitations in their internal module design. In the background estimation stage, existing methods use isotropic standard local convolutions, which cannot effectively model the large-scale anisotropic directional structures commonly found in infrared backgrounds, such as cloud edges and horizons, leading to incomplete background suppression and residual clutter. In the target recovery stage, the receptive field provided by stacked small kernel convolutions is too limited (for example, a 6-layer 3×3 convolution has a theoretical receptive field of only 13×13), making it difficult to fully utilize the local contrast information between small targets and surrounding areas. This makes the network unable to effectively distinguish between real targets and high-frequency noise, easily triggering false alarms. In the data reconstruction stage, existing methods merge the separated background component B and target component T using element-wise addition. This irreversible "many-to-one" operation results in the irrecoverability of the independent discrimination information of the two components, and this information loss accumulates and amplifies with each unrolling stage, limiting the upper limit of overall detection accuracy.
[0007] To address the aforementioned issues, there is an urgent need for an infrared small target detection method that can simultaneously achieve directional background modeling, wide receptive field target perception, and lossless feature fusion.
[0008] References:
[0009] [1]Cheng G, Han J. A survey on object detection in optical remotesensing images[J]. ISPRS journal of photogrammetry and remote sensing, 2016,117: 11-28.
[0010] [2]Zou Z, Chen K, Shi Z, et al. Object detection in 20 years: Asurvey[J]. Proceedings of the IEEE, 2023, 111(3): 257-276.
[0011] [3]Wu F, Zhang T, Li L, et al. RPCANet: Deep unfolding RPCA basedinfrared small target detection[C] / / Proceedings of the IEEE / CVF WinterConference on Applications of Computer Vision. 2024: 4809-4818.
[0012] [4]Gao C, Meng D, Yang Y, et al. Infrared patch-image model for smalltarget detection in a single image[J]. IEEE transactions on image processing,2013, 22(12): 4996-5009.
[0013] [5]Xia C, Li X, Zhao L, et al. Infrared small target detection basedon multiscale local contrast measure using local energy factor[J]. IEEEGeoscience and Remote Sensing Letters, 2019, 17(1): 157-161.
[0014] [6]Liu Q, Liu R, Zheng B, et al. Infrared small target detection withscale and location sensitivity[C] / / Proceedings of the IEEE / CVF Conference onComputer Vision and Pattern Recognition. 2024: 17490-17499.
[0015] [7]Zhu C, Zhu L, Zhou F. Infrared Small Target Detection Based onAdaptive Balance Structure Tensor Indicator[J]. IEEE Geoscience and RemoteSensing Letters, 2025, 22: 1-5.
[0016] [8]Zhang L, Peng L, Zhang T, et al. Infrared small target detectionvia non-convex rank approximation minimization joint l 2, 1 norm[J]. RemoteSensing, 2018, 10(11): 1821.
[0017] [9]Dai Y, Wu Y. Reweighted infrared patch-tensor model with bothnonlocal and local priors for single-frame small target detection[J]. IEEEjournal of selected topics in applied earth observations and remote sensing,2017, 10(8): 3752-3767.
[0018]
[10] Bai X, Zhou F. Analysis of new top-hat transformation and theapplication for infrared dim small target detection[J]. Pattern Recognition,2010, 43(6): 2145-2156.
[0019]
[11] Deshpande S D, Er M H, Venkateswarlu R, et al. Max-mean and max-median filters for detection of small targets[C] / / Signal and Data Processing of Small Targets 1999. SPIE, 1999, 3809: 74-83.
[0020]
[12] Chen C L P, Li H, Wei Y, et al. A local contrast method for small infrared target detection[J]. IEEE transactions on geoscience and remote sensing, 2013, 52(1): 574-581.
[0021]
[13] Wei Y, You X, Li H. Multiscale patch-based contrast measure for small infrared target detection[J]. Pattern recognition, 2016, 58: 216-226.
[0022]
[14] Zhao M, Li W, Li L, et al. Single-frame infrared small-target detection: A survey[J]. IEEE geoscience and remote sensing magazine, 2022, 10(2): 87-119. Invention Summary
[0023] To address the aforementioned technical problems, this invention proposes an infrared small target detection method based on strip context and dilated context sparse (DCSM) deep unfolding networks. This method comprises a Strip Context Background Module (SCBM), a Dilated Context Sparse Module (DCSM), and an Adaptive Residual Fusion Module (ARFM). SCBM uses direction-aware striped pooling to capture a wide range of anisotropic background clutter along both horizontal and vertical directions, achieving more thorough background suppression. DCSM introduces dilated convolution instead of standard convolution, expanding the receptive field from 13×13 to 25×25 without adding any additional parameters, fully utilizing the local contrast features of small targets to achieve accurate target recovery. ARFM replaces traditional element-wise addition with channel-level stitching, completely preserving the independent discriminative features of both background and target components, preventing information loss and cumulative degradation across unfolding stages.
[0024] This invention proposes an infrared small target detection method using striped context and hole-enhanced deep unfolded networks, including preparing training and testing sets for the IRSTD-1K dataset, and further including the following steps:
[0025] Step 1: Use the images in the training set to train the SCD-Net network and generate a training model;
[0026] Step 2: Save the trained model to a local folder, test the effect of the trained model using the images in the test set, and if satisfied, save the trained model as a satisfactory trained model;
[0027] Step 3: Save the satisfactory training model to a local folder and use the satisfactory training model to test infrared images.
[0028] Preferably, the SCD-Net network includes a striped context background module, a holed context sparsity module, and an adaptive residual fusion module.
[0029] In any of the above schemes, preferably, the strip context background module uses a strip pooling block (SPB) as its core building block. The SPB processes the input features through three parallel paths: the local path uses 3×3 convolution to extract fine-grained local features; the horizontal stripe path uses adaptive average pooling to compress the feature map height to a single row (1×W) to encode the global horizontal context, and then restores it to the original resolution through bilinear interpolation after 1×1 convolution dimensionality reduction; the vertical stripe path symmetrically compresses the width to a single column (H×1) to encode the global vertical context, and also restores it through dimensionality reduction and upsampling; the horizontal and vertical context features are concatenated along the channels and then concatenated with the local features again, fused through 1×1 convolution, and output with residual connections. The complete SCBM consists of a head convolution, three cascaded SPBs, and a tail convolution. The input is the difference between the observed image and the target estimate from the previous stage, and residual background estimation is achieved through skip connections.
[0030] In any of the above schemes, preferably, the DCSM includes a dilated context network, which consists of a standard 3×3 convolution at the head, L layers of 3×3 dilated convolutions with a dilation rate of d=2, and a 3×3 convolution at the tail, where L=6 and the theoretical receptive field is 25×25. The DCSM also includes two lightweight parameter generators, each consisting of global average pooling, a 1×1 bottleneck convolution, and a sigmoid output; the first parameter generator dynamically generates a hybrid weight γ based on the target estimation of the previous stage, which is used to adaptively fuse the historical target estimation with the current residual into an intermediate state; the second parameter generator dynamically generates an update step size ε based on the intermediate state, updates and outputs the sparse target estimation of the current stage using proximal gradient descent.
[0031] In any of the above schemes, the preferred embodiment is that the ARFM concatenates the background estimate B and the target estimate T along the channel dimension to form a dual-channel input, which is then fed into the mapping network M for reconstruction. The mapping network M is composed of a 1×1 projective convolution (including batch normalization and LeakyReLU), a dynamic residual group (DRG), and a 1×1 output convolution in sequence. The DRG is composed of N cascaded residual channel-spatial attention blocks (RCSABs). Each RCSAB extracts features through two layers of convolution and then undergoes adaptive feature calibration through channel attention and dynamic spatial attention, and contains residual connections.
[0032] SCD-Net refers to SCD-Net: Strip-Context and Dilation-Enhanced DeepUnfolding
[0033] Network for Infrared Small Target Detection, namely, a striped context and hole-enhanced deep unfolding network. Attached Figure Description
[0034] Figure 1 This is a flowchart of a preferred embodiment of an infrared small target detection method using striped context and hole-enhanced depth unfolding networks according to the present invention.
[0035] Figure 2 This is a schematic diagram of the overall network structure of an embodiment of the SCD-Net network used in the infrared small target detection method according to the present invention, which employs strip context and hole enhancement depth unfolding network.
[0036] Figure 3 This is a schematic diagram of the strip context background module structure of the SCD-Net network in the infrared small target detection method using strip context and hole enhancement depth unfolding network according to the present invention.
[0037] Figure 4 This is a schematic diagram of the strip context background submodule structure of the SCD-Net network in the infrared small target detection method using strip context and hole enhancement depth unfolding network according to the present invention.
[0038] Figure 5 This is a schematic diagram of the hole context sparse module structure of the SCD-Net network in the infrared small target detection method using strip context and hole enhancement depth unfolding network according to the present invention.
[0039] Figure 6 This is a schematic diagram of the adaptive fusion module structure of the SCD-Net network in the infrared small target detection method using strip context and hole enhancement depth unfolding network according to the present invention. Detailed Implementation
[0040] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0041] Example 1
[0042] like Figure 1 As shown, perform step 100 to prepare the training and test sets for the IRSTD-1K dataset.
[0043] Perform step 110 to train the SCD-Net network using the images in the training set and generate a training model.
[0044] The SCD-Net network includes a striped context background module, a holed context sparse module, and an adaptive fusion module.
[0045] The input to the SCD-Net network is a single-channel infrared grayscale image. The all-zero matrix is used as the initial estimate of the target components. The process iterates in stages. In each stage, SCBM estimates the low-rank background component B by sensing the anisotropic directional structure in the background through strip pooling based on the difference between the current observation and the target estimate; DCSM expands the receptive field using dilated convolution and adaptively recovers the sparse target component T using proximal gradient descent; ARFM losslessly feeds B and T into the mapping network through channel concatenation to reconstruct the updated observations for use in the next stage. The target component T output in the final stage is the infrared small target detection mask.
[0046] Training employed the Adam optimizer with an initial learning rate of 1e-5, using a poly learning rate decay strategy, a batch size of 8, and 400 training epochs. The loss function consisted of a weighted combination of the Soft-IoU segmentation loss and the mean squared error reconstruction loss, with a weight coefficient of 0.1.
[0047] Perform step 120 to save the training model to a local folder, test the effect of the training model using images in the test set, and if satisfied, save the training model as a satisfactory training model.
[0048] Perform step 130 to save the satisfactory training model to a local folder, and use the satisfactory training model to test infrared images.
[0049] Example 2
[0050] This invention is a strip-context and dilation-enhanced deep unfolding network for infrared small target detection, named "Strip-Context and Dilation-Enhanced Deep Unfolding Network" (SCD-Net), a trainable end-to-end detection network.
[0051] SCD-Net consists of a Striped Context Background Module (SCBM), a Hole Context Sparse Module (DCSM), and an Adaptive Residual Fusion Module (ARFM).
[0052] Training SCD-Net requires a paired "infrared image-target mask" labeled dataset. By inputting the network hyperparameters from the trained model into the network model, it achieves infrared small target detection, resulting in accurate target detection. It effectively suppresses background clutter and recovers high-quality target detection results.
[0053] Steps for using SCD-Net (implemented using the Python programming language):
[0054] 1. Prepare the training and test sets for the IRSTD-1K dataset, and set the image format to ".png";
[0055] 2. Start training the network using the training file "train.py" and adjust parameters such as batch size and learning rate as needed;
[0056] 3. Save the trained model to a local folder, use the test set to evaluate the performance of the trained model, and if you are not satisfied, you can use the "checkpoint" technique to continue training a satisfactory model;
[0057] 4. Save the satisfactory training model to a local folder and use the satisfactory training model to test infrared images.
[0058] Example 3
[0059] Existing infrared small target detection methods still face three core challenges in complex scenes: incomplete suppression of anisotropic background clutter, limited receptive field for target recovery, and information loss during cross-stage feature fusion. Current deep unfolding methods primarily rely on isotropic standard convolutions for background estimation. However, infrared backgrounds commonly contain large-scale directional extensions such as cloud edges and horizons. Isotropic convolutions severely limit the network's ability to model directional clutter, meaning the network cannot effectively distinguish between directional background clutter and point-like small targets. Simultaneously, the target recovery module uses small-kernel convolutions with limited receptive fields, making it difficult to fully capture the local contrast information between the small target and its surrounding neighborhood. The data reconstruction module merges background and target components element-wise, an irreversible operation that results in the loss of independent discriminative information between the two components. To address these challenges, this invention proposes a striped context-hole-enhanced deep unfolding network for infrared small target detection, called SCD-Net.
[0060] like Figure 2 As shown, the SCD-Net network structure includes a Striped Context Background Module (SCBM), a Hollow Context Sparse Module (DCSM), and an Adaptive Residual Fusion Module (ARFM). The overall process of SCD-Net is as follows: starting with the input single-channel infrared image D and the all-zero target estimate, it undergoes K=6 cascaded decomposition stages for iterative processing. Each stage sequentially executes SCBM to estimate the low-rank background component, DCSM to recover the sparse target component, and ARFM to reconstruct and update the observations. These three modules jointly simulate a one-step alternating minimization iteration of RPCA optimization. The target estimate in the final stage is the target detection mask output. The entire process achieves end-to-end high-precision infrared small target detection through the guidance of low-rank sparse priors, striped orientation sensing, and hole receptive field expansion.
[0061] The functions of each module are as follows:
[0062] 1. Strip Context Background Module (SCBM) (e.g.) Figure 3 , 4 (As shown)
[0063] To address the challenge of effectively modeling anisotropic directional structures such as cloud edges and horizons prevalent in infrared backgrounds using standard isotropic convolution, this invention proposes a striped context background module (SCBM), such as... Figure 3 As shown, SCBM uses strip pooling blocks (SPBs) as its core building blocks, cascading three SPBs between the head and tail convolutions. The input is the currently observed image. With target estimation The difference R is used to output residual background estimation via a skip connection. .
[0064] The detailed structure of SPB is as follows: Figure 4 As shown, this module processes the input feature map X through three parallel paths: the local path uses 3×3 convolution to extract fine-grained local features F_local; the horizontal strip path uses adaptive average pooling to compress the feature map height to a single row (1×W) to encode global horizontal structural information, and after 1×1 convolution dimensionality reduction, bilinear interpolation restores it to the original resolution, resulting in F_hor; the vertical strip path symmetrically compresses the width to a single column (H×1) to encode global vertical structural information, and is also restored through dimensionality reduction and upsampling, resulting in F_ver. The horizontal and vertical features are concatenated along the channels to form a directional context, and then concatenated again with the local features, fused by 1×1 convolution, and output with residual connections. Compared to global average pooling (GAP), strip pooling preserves the spatial distribution along one axis, effectively distinguishing directional background clutter from small point targets.
[0065] 2. Diffuse Context Sparse Modules (DCSM) (e.g.) Figure 5 (As shown)
[0066] To address the problem of limited receptive field and difficulty in fully utilizing local contrast features of small targets during the target recovery stage, this invention proposes a Hollow Context Sparse Module (DCSM), such as... Figure 4 As shown. DCSM replaces the intermediate convolutional layers in the target recovery subnetwork with dilation coefficients. The dilated convolution, without adding any extra parameters, increases the theoretical receptive field of 6 stacked convolutions from... Expand to It fully covers the neighborhood range required for local contrast detection. DCSM also includes two lightweight parameter generators. and Each is achieved through global average pooling, It consists of bottleneck convolution and Sigmoid output. Based on the previous stage target estimate Dynamically generate mixed weights : The historical target estimate is adaptively fused with the current residual into an intermediate state:
[0067] according to Dynamically generate update step size: Output the sparse target estimate for the current stage using proximal gradient descent. ,in This represents element-wise multiplication. For head convolution, Layer expansion rate A dilated context network consisting of dilated convolutions and tail convolutions has a theoretical receptive field of [missing information]. .
[0068] 3. Adaptive Residual Fusion Module (ARFM) (e.g.) Figure 6 (As shown)
[0069] To address the problem of irreversible information loss caused by the element-wise additive merging of background and target components in existing unfolded networks, this invention proposes an Adaptive Residual Fusion Module (ARFM), such as... Figure 6 As shown. ARFM will estimate the background at the current stage. With target estimation A dual-channel input is formed by concatenating along the channel dimension and fed into the mapping network M to preserve the independent discriminative features of the two components. The mapping network M consists of a 1×1 projective convolution (including batch normalization and LeakyReLU), a dynamic residual group (DRG), and a 1×1 output convolution in sequence. The DRG is composed of N=5 cascaded residual channel-spatial attention blocks (RCSABs). Each RCSAB performs adaptive feature calibration through channel attention and dynamic spatial attention, and contains residual connections to effectively prevent information accumulation and degradation across the unfolding stage.
[0070] To better understand this invention, specific embodiments have been described in detail above, but these are not intended to limit the invention. Any simple modifications made to the above embodiments based on the technical essence of this invention still fall within the scope of this invention. Each embodiment in this specification focuses on its differences from other embodiments; similar or identical parts between embodiments can be referred to mutually. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
Claims
1. An infrared small target detection method using striped context and hole-enhanced deep unfolded networks, comprising preparing training and testing sets for the IRSTD-1K dataset, characterized in that, It also includes the following steps: Step 1: Use the images in the training set to train the SCD-Net network and generate a training model; Step 2: Save the trained model to a local folder, test the effect of the trained model using the images in the test set, and if satisfied, save the trained model as a satisfactory trained model; Step 3: Save the satisfactory training model to a local folder and use the satisfactory training model to test infrared images.
2. The infrared small target detection method using striped context and hole-enhanced depth unrolling network as described in claim 1, characterized in that, The SCD-Net network includes multiple learnable iterative stages based on a deep unfolding framework. Each stage sequentially includes a striped context background module, a holed context sparse module, and an adaptive residual fusion module.
3. The infrared small target detection method using striped context and hole-enhanced depth unrolling network as described in claim 2, characterized in that, The strip context background module, by introducing a direction-aware strip pooling mechanism, aggregates global context information in the horizontal and vertical directions, solving the problem that traditional isotropic convolution is difficult to model anisotropic directional clutter in infrared scenes.
4. The infrared small target detection method using striped context and hole-enhanced depth unrolling network as described in claim 3, characterized in that, The hollow context sparse module expands the receptive field through hollow convolution to capture a wide range of local contrast cues, achieving accurate recovery and background separation of extremely dark and weak infrared small targets without adding additional parameters.
5. The infrared small target detection method using striped context and hole-enhanced depth unrolling network as described in claim 4, characterized in that, The adaptive residual fusion module, through channel-level feature stitching, adaptively retains the discriminative features of the separated low-rank background and sparse targets, effectively preventing irreversible loss and degradation of information caused by traditional pixel-by-pixel addition fusion.
6. The infrared small target detection method using striped context and hole-enhanced depth unrolling network as described in claim 5, characterized in that, The SCD-Net first uses the raw infrared image as the initial state. Then, the network expands the low-rank sparse decomposition process into K optimization stages. In each stage, the striped context background module is alternately executed for background extraction, and the hollow context sparse module is executed for target recovery. After hierarchical processing, the features are collaboratively processed by the adaptive residual fusion module to handle the interaction information between the background and the target, dynamically balancing the reconstruction fidelity. Finally, deep feature fusion is completed through progressive residual optimization, outputting the final result. Mask image of the detected small infrared target.
7. The infrared small target detection method using striped context and hole-enhanced depth unrolling network as described in claim 6, characterized in that, The strip context background module has a core strip pooling block (SPB) consisting of three parallel paths: the local path extracts fine-grained local features through 3×3 convolution, while the horizontal stripe path compresses the features to 1×W and the vertical stripe path compresses the features to H×1; then the module upsamples the horizontal and vertical stripe features and concatenates them with the local features through channels, and finally outputs background features sensitive to directional clutter through 1×1 convolution and residual connection.
8. The infrared small target detection method using striped context and hole-enhanced depth unrolling network as described in claim 7, characterized in that, The cross-modal attention, sparse hole context module establishes local contrast correlations, generates sparse features by stacking multiple dilated convolutional layers with different dilation rates, dynamically generates adaptive thresholds through a parameter generator, and obtains the final target estimation features through nonlinear threshold shrinkage.
9. The infrared small target detection method using striped context and hole-enhanced depth unrolling network as described in claim 8, characterized in that, The adaptive residual fusion module first concatenates the background features and target features along the channel dimension, then adaptively calibrates the multimodal information through Residual Channel Spatial Attention Block (RCSAB), and finally reconstructs the data flow of the current stage through a convolutional layer with skip connections, and passes it to the next iteration stage.