Infrared sequence weak and small target detection method and device based on frequency domain information
By designing a frequency domain feature extraction module and an unsupervised learning strategy for infrared sequence small target detection, the problems of reduced positioning accuracy and high computing resources of existing models in complex environments are solved, and efficient and real-time infrared small target detection is achieved.
Patent Information
- Application Number
- CN202510962747.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-14
AI Technical Summary
The existing infrared small target detection model has reduced positioning accuracy in complex environments, is difficult to adapt to unknown backgrounds and fast-moving targets, consumes high computing resources, and lacks real-time performance, which limits its application in air surveillance and missile warning missions.
A small target detection method for infrared sequences based on frequency domain information is designed. It adopts a five-layer frequency domain feature extraction module and a multi-head self-attention mechanism, combined with three-dimensional convolutional decoding and refinement modules. The network is trained through an unsupervised learning strategy, integrating spatiotemporal and frequency domain features to improve detection accuracy and adaptability.
It achieves high-quality infrared small target detection under unsupervised conditions, enhances the adaptability and detection accuracy of the model in complex scenes, reduces dependence on computing resources, and is suitable for real-time applications.
Smart Images

Figure CN120472150B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a method and device for detecting small infrared targets based on frequency domain information. Background Art
[0002] With the development of deep learning technology, neural networks can now distinguish between infrared background and small infrared targets. As a data-driven approach, infrared small target detection models based on deep learning can be roughly divided into two categories: detection models and segmentation models.
[0003] Detection-based models aim to locate small infrared targets and generate bounding boxes containing their location information. Xu et al. proposed the RMT-YOLOv9s model for infrared small target detection; Zhang et al. proposed MPFFNet, which integrates traditional methods; Dahri et al. designed OSTD-YOLOv8 for small target detection; Duan et al. proposed a three-domain joint strategy based on frequency-aware memory enhancement; Chen et al. designed a sliced spatiotemporal network (SSTNet) that outputs infrared target locations; Bai et al. proposed a cross-connected bidirectional pyramid network (CBP-Net) for target localization; and Han et al. proposed a context-aware network (KCPNet) that integrates prior knowledge for detecting and localizing small ships at sea. However, when there are multiple targets in a scene or when targets are densely distributed, these models struggle to accurately separate them due to overlapping targets or blurred boundaries.
[0004] Generally speaking, detection-based models offer faster inference speeds and are more suitable for real-time applications. However, their localization accuracy often degrades in complex environments. In contrast, segmentation-based models provide pixel-level detection results and are particularly effective for improving the accuracy of small infrared target detection. Typical methods include U-Net in U-Net (UIU-Net), Dense Nested Attention Network (DNANet), Orthogonal Input Perceptual Fusion Framework (OIPF), Deep Unfolding RPCA Network (RPCANet), Multi-Scale Attention Fusion U-Net (AMFUNet), and Multi-branch Mutual Learning Network (MMLNet). Although segmentation-based methods offer more refined detection, they typically come with higher computational overhead and slower inference speeds. Most models focus on single-frame detection, relying on large amounts of annotated data and powerful computing resources. Due to the extremely small size of infrared targets, many models employ deeper network structures to enhance detection capabilities, but this can result in weak targets being suppressed or lost during feature extraction.
[0005] In general, most existing networks are typically trained on specific datasets and lack the ability to adaptively learn significant differences between targets and backgrounds. Consequently, their performance degrades significantly when encountering unknown backgrounds, rapidly moving targets, or changing sensor conditions. Due to the small size and low contrast of infrared targets, many deep networks are also prone to target loss or false detection, especially when the representation of detail is weakened by the deep structure. Furthermore, these models generally suffer from insufficient real-time performance, limiting their potential for application in time-sensitive tasks such as air surveillance and missile warning. Summary of the Invention
[0006] In view of the deficiencies in the prior art, the present invention aims to provide a method and device for detecting small targets in an infrared sequence based on frequency domain information.
[0007] This method fully integrates spatiotemporal and frequency domain features, and its innovations are mainly reflected in the following aspects: First, a frequency domain feature extraction module is designed, which adopts a five-layer structure and introduces a dual-frequency domain transformation structure and a multi-head self-attention mechanism in each layer to effectively mine the frequency-sensitive features and local-global attention information in the infrared image sequence; secondly, a multi-level attention feature aggregation mechanism is constructed, and the attention features of different levels are stacked according to the channel dimension to form a multi-scale feature set, providing rich hierarchical information support for subsequent refined reconstruction; thirdly, a background reconstruction structure based on three-dimensional convolutional decoding and refinement modules is proposed, which upsamples and fuses shallow features layer by layer, and further enhances the detail reconstruction effect through two refinement modules, and finally obtains a high-quality background estimation consistent with the original spatiotemporal tensor; in addition, a loss function is designed, and an unsupervised learning strategy is adopted for network training. Network training can be completed without labeled data, thereby enhancing the model's adaptability in different scenarios; the excellent performance of the infrared sequence weak target detection method based on frequency domain information is comprehensively verified in the implementation examples.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] On the one hand, the present invention discloses a method for detecting small targets in an infrared sequence based on frequency domain information, which comprises the following steps:
[0010] Step 1): Preprocess each frame of the thermal infrared image sequence to be detected to obtain initial features; at the same time, stack all infrared images in sequence according to the time dimension to obtain a spatiotemporal tensor;
[0011] Step 2): Design a frequency domain feature extraction module, which includes five layers of feature extraction modules. Each module contains two frequency domain transformation structures and a multi-head self-attention mechanism structure. The initial features obtained in step 1) are sequentially subjected to frequency domain transformation calculations and multi-head self-attention mechanism calculations to obtain attention features.
[0012] Step 3): the operations of steps 1) to 2) are performed on N image frames to obtain N attention features; the N attention features obtained by the five-layer feature extraction module are stacked along the channel dimension to obtain the attention features aggregated by each layer feature extraction module;
[0013] Step 4): a background reconstruction module is designed to perform upsampling on the attention features aggregated by the fifth layer feature extraction module layer by layer, and sequentially perform channel splicing with the attention features aggregated by the first four layer feature extraction modules, and then gradually reconstruct the background by using a three-dimensional convolution decoding manner; two refining modules are further used to improve the details to obtain a reconstructed background tensor consistent in scale with the original spatio-temporal tensor;
[0014] Step 5): a loss function is designed to train the infrared sequence weak target detection network based on frequency domain information composed of the frequency domain feature extraction module and the background reconstruction module in an unsupervised learning manner;
[0015] Step 6): the infrared image sequence to be detected is input into the infrared sequence weak target detection network trained in step 5), the residual between the corresponding spatio-temporal tensor and the reconstructed background tensor is calculated, a filter is designed to filter the information of the residual, and a final target tensor is obtained, and the infrared small target detection result sequence is reconstructed to realize infrared small target detection.
[0016] The application also discloses an infrared small target detection device based on frequency domain information for implementing the method, which comprises:
[0017] A preprocessing and spatio-temporal tensor construction module is used to pre-process each infrared image in the infrared image sequence to obtain initial features, and construct the infrared image sequence into a spatio-temporal tensor;
[0018] A frequency domain feature extraction module is used to extract features of each infrared image, and the frequency domain feature extraction module comprises five layer feature extraction modules, each of which comprises two frequency domain transformation structures and a multi-head self-attention mechanism calculation structure;
[0019] A feature aggregation module is used to aggregate N attention features obtained by the five layer feature extraction module for each image to obtain an attention feature set;
[0020] A background reconstruction module is used to perform layer-by-layer upsampling calculation on the attention feature set, and two refining modules are designed to further improve the details to obtain a reconstructed background tensor consistent in scale with the original spatio-temporal tensor;
[0021] A loss function design module is used to design a background reconstruction loss function, including a mean square error loss and a multi-scale structural similarity loss, to train the network in an unsupervised learning manner;
[0022] The information filtering module designs a filter and performs information filtering on the residual between the original spatiotemporal tensor and the reconstructed background tensor to obtain the final target tensor;
[0023] The target detection result output module is used to output the infrared small target detection result map of each frame of infrared image in the infrared image sequence.
[0024] Compared with the prior art, the present invention has the following beneficial effects:
[0025] 1) This paper proposes a method for detecting small targets in infrared sequences that integrates frequency domain information. The method constructs an unsupervised detection network consisting of a five-layer frequency domain feature extraction module and a background reconstruction module. This network fully integrates the frequency characteristics and spatiotemporal structure information in infrared image sequences, helping to enhance small target response and improve background reconstruction accuracy. Furthermore, by constructing a joint loss function to achieve unsupervised training, it avoids reliance on large-scale labeled data, improving the method's adaptability and generalizability in complex scenarios.
[0026] 2) To address the problem of frequency domain feature modeling, the present invention designs a feature extraction module based on dual-frequency domain transformation and multi-head self-attention mechanism, which can fully extract frequency-sensitive global semantic information; to address the problem of effective fusion of features at different levels, the present invention proposes a multi-scale attention feature stacking mechanism to integrate deep and shallow frequency domain information and improve feature expression capabilities; in addition, the present invention constructs a refined background decoder and information filtering module to screen the reconstructed residual signal to suppress pseudo-responses and highlight the real target, thereby achieving accurate small target detection effect under the unsupervised learning framework; experimental results show that the network structure and information filter designed by the present invention can achieve satisfactory infrared small target detection performance under sufficient data volume and unsupervised learning conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 The frequency domain feature extraction module in the present invention includes five layers of feature extraction modules;
[0028] Figure 2 Each frequency domain feature extraction module in the present invention includes two frequency domain transformation structures (a) and (b), and a multi-head self-attention mechanism structure (c);
[0029] Figure 3 Schematic diagram of the background reconstruction module structure (a) and the schematic diagram of each layer of background reconstruction structure (b);
[0030] Figure 4 Schematic diagram of the structure of the infrared small target detection device based on frequency domain information in the present invention;
[0031] Figure 5This is an example frame image of the thermal infrared image sequence used for experimental testing;
[0032] Figure 6 is the detection result diagram of the example frame image;
[0033] Figure 7 The original image of an example frame of a thermal infrared image and the thermal infrared small target detection result diagram detected by the method of the present invention and the comparative method. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention is described in detail below with reference to specific embodiments. Specific embodiments are described below to simplify the present invention. However, it should be understood that the present invention is not limited to the illustrated embodiments, and that various modifications of the present invention are possible without departing from the underlying principles, and these equivalent forms also fall within the scope defined by the appended claims.
[0035] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0036] The infrared sequence small target detection method based on frequency domain information of the present invention mainly includes:
[0037] Step 1: For each frame of the thermal infrared image sequence to be detected Perform preprocessing to obtain initial features ; At the same time, all infrared images are stacked in sequence according to the time dimension to obtain the space-time tensor ;
[0038] Specifically, for the thermal infrared image sequence to be detected, including N image frames, each frame is recorded as , , , Respectively represent the height, width and number of channels of the infrared image frame; Perform preprocessing to obtain initial features :
[0039] (1)
[0040] in, , p represents the position encoding parameter, Drop represents dropout regularization, and PE represents two-dimensional patch embedding;
[0041] At the same time, all infrared images are stacked in sequence according to the time dimension to obtain the space-time tensor , , , , Represent the space-time tensor The number of forward slices, height, width, and number of channels;
[0042] Specifically in the embodiment, and Both are 256, Take it as 16, is 1, the constructed space-time tensor The size is .
[0043] Step 2: Design a frequency domain feature extraction module, such as Figure 1 As shown, the frequency domain feature extraction module of this embodiment includes five layers of feature extraction modules, each of which is as follows: Figure 2 As shown, it contains two frequency domain transformation structures and a multi-head self-attention mechanism structure. The initial features obtained in step 1) are sequentially subjected to frequency domain transformation calculations and multi-head self-attention mechanism calculations to obtain attention features.
[0044] Specifically, for the i-th layer feature extraction module, the calculation method is as follows:
[0045] (2)
[0046] in, represents the input features of the i-th layer feature extraction module, represents the features obtained by the feature extraction module at layer i, represents the i-th layer feature extraction module, Represents the computational operation of the multi-head self-attention mechanism in the feature extraction module of the i-th layer, Frequency domain transformation operation in the i-th layer feature extraction module.
[0047] Input features , first perform a custom normalization operation DyT on the input:
[0048] (3)
[0049] in, Represents input features The output features after custom normalization operation, represents the tanh activation function, , , represents the network learning parameters;
[0050] Features Perform a two-dimensional real fast Fourier transform , and then multiply by a learnable complex domain weight , through the two-dimensional real inverse fast Fourier transform Transform to the spatial domain to obtain the frequency domain enhancement result :
[0051] (4)
[0052] in, Represents the element-wise multiplication symbol;
[0053] Carry out residual enhancement processing, first Normalization, MLP and dropout regularization are performed, and then The residual connection itself gives:
[0054] (5)
[0055] in, represents the output features obtained by residual enhancement processing, represents the element addition symbol, Represents multi-layer fully connected layer computing operations;
[0056] After two consecutive frequency domain transformation structures are processed, the frequency domain features are obtained .
[0057] The obtained frequency domain features After normalization, multi-head attention calculation is carried out:
[0058] (6)
[0059] in, Represents the attention feature, MSAttn represents the multi-head self-attention mechanism operation, and is calculated as:
[0060] (7)
[0061] in, is the Softmax activation function, Indicates query, K indicates key, V indicates value, represents the position code, represents the scaling factor, express The transpose of , K, V are transformed linearly 、 、 get, , , A tensor representing the weights learned by the network.
[0062] Step 3: Perform steps 1) to 2) on N image frames to obtain N attention features , ; And stack the N attention features obtained by the five-layer feature extraction module along the channel dimension to obtain the attention features after aggregation of the feature extraction modules of each layer , the five-layer feature extraction module finally obtains a multi-level attention feature set represented as ;
[0063] Specifically, repeat steps 1) to 2) for N frames in the thermal infrared image sequence to be detected to obtain N attention features ; For the i-th background reconstruction structure, the N attention features corresponding to the N image frames are stacked along the channel dimension to obtain the aggregated attention features ; Finally, a multi-level attention feature set can be obtained .
[0064] Step 4: Design a background reconstruction module, upsample the attention features aggregated by the fifth-layer feature extraction module layer by layer, and sequentially perform channel splicing with the attention features aggregated by the first four layers of feature extraction modules, and gradually reconstruct the background using a three-dimensional convolutional decoding method; further enhance the details through two refinement modules to obtain the original spatiotemporal tensor. The scale-consistent reconstructed background tensor ;
[0065] Specifically, during the background reconstruction process, the feature map is upsampled to have the same resolution as the original infrared sequence; the attention feature Input to the first background reconstruction layer , get the initial background estimate :
[0066] (8)
[0067] For each subsequent scale , the reconstruction result Attention features after scale alignment Perform channel stitching and input the next background reconstruction layer , gradually complete the background reconstruction, background reconstruction layer Structure such as Figure 3 As shown; among them, the attention feature The calculation formula for scale alignment is:
[0068] (9)
[0069] wherein, , denotes a three-dimensional point convolution, denotes an up-sampling operation, denotes a batch normalization operation, denotes a leakyReLU activation function, denotes the attention feature obtained after scale alignment;
[0070] background reconstruction layer The background reconstruction process is:
[0071] (10)
[0072] wherein, denotes a channel-wise concatenation operation, denotes the input of the i-th background reconstruction layer , which is also the reconstructed quantity output by the i-th background reconstruction layer The background reconstruction layer is specifically
[0073] (11)
[0074] wherein, denotes a three-dimensional transpose convolution of the i-th background reconstruction layer ;
[0075] The reconstructed quantity calculated by the last background reconstruction layer is sequentially input into two background refinement modules for detail recovery, and a reconstructed background consistent in scale with the original spatio-temporal tensor is obtained through normalization and an activation function, and the calculation formula is as follows:
[0076] (12)
[0077] wherein, denotes a Sigmoid function, and denote background refinement modules, and .
[0078] Step 5: design a loss function including a mean square error loss and a multi-scale structural similarity loss to train the infrared sequence weak and small target detection network based on frequency domain information composed of the frequency domain feature extraction module and the background reconstruction module in an unsupervised learning manner;
[0079] Specifically, in order to improve the quality of background reconstruction, the reconstruction results are constrained from two aspects: pixel accuracy and structure preservation, and the loss function is designed. as follows:
[0080] First, define the mean squared error loss :
[0081] (13)
[0082] in, represents the square of the Frobenius norm;
[0083] Second, define the multi-scale structural similarity loss :For the Frame infrared image and the corresponding reconstruction background , through Gaussian blur processing and step-by-step downsampling operations, we get K pairs of images with decreasing resolution and ; At each scale, calculate the local mean of the image and , local standard deviation and , local covariance , calculate the structural similarity at this scale Similarity with contrast , the calculation formula is:
[0084] (14)
[0085] (15)
[0086] in, and is a positive constant. Specifically in the embodiment, , ;
[0087] Take the weighted combination of SSIM and MCS at all scales to obtain the structural similarity loss as follows:
[0088] (16)
[0089] in, is a weight coefficient. Specifically in the embodiment, Taken as 0.0448, 0.2856, 0.3001, 0.2363, 0.1333;
[0090] Combining formula (13) to formula (16), the following loss function is designed: :
[0091] (17)
[0092] wherein, , denote weight coefficients, which are specific to embodiments, , ;
[0093] Thus, the background reconstruction network is trained in an unsupervised learning manner by using the loss function .
[0094] In network training, the SGD algorithm is optimized with a momentum of 0.9 and a weight decay coefficient of 0.00005, the initial learning rate is , the cosine learning rate is used to dynamically adjust the learning rate, the expression is , the coefficient , x is the current training iteration round, is the total number of training iteration rounds, which is set to 1000, and the batch size is set to 32 in the training and testing process.
[0095] Step 6: input the to-be-detected thermal infrared image sequence into the infrared sequence weak target detection network trained in step 5), calculate the corresponding spatio-temporal tensor , and obtain the residual between the residual and the reconstructed background tensor , and obtain the residual as the target tensor , design a filter to filter , obtain the final target tensor , and reconstruct into an infrared small target detection result sequence T, and realize infrared small target detection.
[0096] Specifically, input the to-be-detected thermal infrared image sequence into the network trained in step 5), calculate the spatio-temporal tensor , and obtain the residual between the residual and the reconstructed background tensor , and obtain the residual as the target tensor , design a filter to filter , specifically, introduce a sparse base selector and a tubular fiber filtering operation, first, use the sparse base selector to retain the first high-response pixels of each frontal section of , and set the pixel values of the remaining positions to zero to obtain a temporary target tensor ;
[0097] Specifically, in embodiments, ;
[0098] second, design a tubular fiber filter to filter All pixel positions in , calculate the time-averaged response value:
[0099] (18)
[0100] Constructing a mask ,when hour, represents the threshold value, If it is 0, otherwise it is 1, the mask Perform element-wise multiplication on the temporary target tensor Each front slice of , thereby eliminating the interference points or dead spots that may come from the high brightness of the sensor, and thus obtaining the final target tensor ;Will Reconstructed into infrared small target detection result sequence T, infrared small target detection is achieved. Specifically in the embodiment, .
[0101] Corresponding to the aforementioned embodiment of an infrared sequence small target detection method based on frequency domain information, the present invention also provides an embodiment of an infrared small target detection device based on frequency domain information.
[0102] Figure 4 1 is a block diagram of a device for detecting small infrared targets based on frequency domain information according to an exemplary embodiment. The device includes:
[0103] The preprocessing and spatiotemporal tensor construction module preprocesses each frame of the infrared image sequence to obtain initial features and constructs the infrared image sequence into a spatiotemporal tensor;
[0104] A frequency domain feature extraction module is used to extract features of each frame of infrared image. The frequency domain feature extraction module includes five layers of feature extraction modules, each of which contains two frequency domain transformation structures and a multi-head self-attention mechanism calculation structure;
[0105] The feature aggregation module aggregates the N attention features obtained by the five-layer feature extraction module for each frame image to obtain an attention feature set;
[0106] The background reconstruction module performs layer-by-layer upsampling on the attention feature set and designs two refinement modules to further enhance the details, obtaining a reconstructed background tensor with the same scale as the original spatiotemporal tensor;
[0107] Loss function design module, which designs background reconstruction loss functions, including mean square error loss and multi-scale structural similarity loss, to train the network in an unsupervised learning manner;
[0108] The information filtering module designs a filter and performs information filtering on the residual between the original spatiotemporal tensor and the reconstructed background tensor to obtain the final target tensor;
[0109] The target detection result output module is used to output the infrared small target detection result map of each frame of infrared image in the infrared image sequence.
[0110] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0111] As for the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is only schematic. The various modules in the device are a kind of logical function division. There may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another unit. Another point is that the connection between the modules shown or discussed can be a communication connection through some interfaces, which can be electrical or other forms. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present application scheme. Ordinary technicians in this field can understand and implement it without paying creative labor. The following uses the disclosed real thermal infrared image sequence as an example to illustrate the specific implementation method to reflect the technical effect of the present invention. The specific steps in the embodiment will not be repeated.
[0112] Example
[0113] The accompanying drawings illustrating the embodiments of the present invention serve to more clearly illustrate the objectives, technical solutions, and advantages of the present invention. It should be noted that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. Any equivalent substitutions, modifications, and the like made within the methodologies and principles provided by the present invention are intended to be included within the scope of protection of the present invention.
[0114] In this embodiment, the effectiveness of the thermal infrared small target detection method will be verified using a public infrared image sequence. The infrared image sequence contains 399 frames of infrared images, each with a size of 256×256. The sequence is set against a complex jungle background with linear, high-brightness, strong interference structures. The slow-flying aircraft is affected by clutter and noise in the image. A typical image frame of the sequence is as follows: Figure 5 As shown, Figure 6A result image obtained by the infrared sequence weak small target detection method based on frequency domain information of the present application. In order to more accurately and objectively evaluate the superiority of the proposed infrared sequence weak small target detection method based on frequency domain information, the latest traditional and deep learning-based thermal infrared small target detection methods are selected, and the target detection result images of the method in the infrared image sequence instance frame are qualitatively compared and analyzed. From the quantitative point of view, the performance of the detection method is evaluated by the 3D-ROC evaluation index system, in which AUC (D, τ) , AUC TD is used to evaluate the target detection ability of the method; AUC (F, τ) , AUC BS , AUC SNPR is used to evaluate the background suppression ability of the method; AUC (D,F) , AUC TDBS , AUC ODP is used to evaluate the comprehensive ability of the detector.
[0115] In order to more objectively verify the effectiveness of the method of the present application, 13 latest thermal infrared small target detection methods are selected for comparison with the method of the present application, and the comparison methods include MDWCM, DGRL, STT-TRNR, TWTVR, FGLR-MCP, ALCNet, ISTDUNet, DNANet, AGPCNet, AMFUNet, RDIANet, RPCANet and MSHNet. Table 1 shows the index of the 3D-ROC evaluation system of the infrared weak small target detection results of the thermal infrared image sequence using the comparison method and the method of the present application. The bold and underlined values represent the best and second best performance, respectively.
[0116] MDWCM is derived from R. Lu, X. Yang, W. Li, J. Fan, D. Li, and X. Jing, “Robust infrared small target detection via multidirectional derivative-based weighted contrast measure,” IEEE Geosci. Remote Sens. Lett., vol. 19, pp. 1–5, 2022.
[0117] DGRL源自F. Zhou, Y. Wu, Y. Dai, and K. Ni, “Robust infrared smalltarget detection via jointly sparse constraint of l1 / 2-metric and dual-graphregularization,” Remote Sens., vol. 12, no. 12, 2020. [Online]. Available:https: / / www.mdpi.com / 2072-4292 / 12 / 12 / 1963.
[0118] STT-TRNR源自H. Yi, C. Yang, R. Qie, J. Liao, F. Wu, T. Pu, and Z.Peng, “Spatialtemporal tensor ring norm regularization for infrared smalltarget detection,” IEEE Geosci. Remote Sens. Lett., vol. 20, pp. 1–5, 2023.
[0119] TWTVR源自E. Zhao, L. Dong, C. Li, and Y. Ji, “Infrared maritimetarget detection based on temporal weight and total variation regularizationunder strong wave interferences,” IEEE Trans. Geosci. Remote Sens., vol. 62,pp. 1–19, 2024.
[0120] FGLR-MCP源自T. Liu, Y. Liu, J. Yang, B. Li, Y. Wang, and W. An,“Graph laplacian regularization for fast infrared small target detection,”Pattern Recognit., vol. 158, p. 111077, 2025. [Online]. Available: https: / / www.sciencedirect.com / science / article / pii / S0031320324008288.
[0121] ALCNet源自Y. Dai, Y. Wu, F. Zhou, and K. Barnard, “Attentional localcontrast networks for infrared small target detection,” IEEE Trans. Geosci.Remote Sens., vol. 59, no. 11, pp. 9813–9824, 2021.
[0122] ISTDUNet源自Q. Hou, L. Zhang, F. Tan, Y. Xi, H. Zheng, and N. Li,“Istdu-net: Infrared small-target detection u-net,” IEEE Geosci. Remote Sens.Lett., vol. 19, pp. 1–5, 2022.
[0123] DNANet源自B. Li, C. Xiao, L. Wang, Y. Wang, Z. Lin, M. Li, W. An, andY. Guo, “Dense nested attention network for infrared small target detection,”IEEE Trans. Image Process., vol. 32, pp. 1745–1758, 2023.
[0124] AGPCNet源自T. Zhang, L. Li, S. Cao, T. Pu, and Z. Peng, “Attention-guided pyramid context networks for detecting infrared small target undercomplex background,” IEEE Trans. Aerosp. Electron. Syst., vol. 59, no. 4, pp.4250–4261, 2023.
[0125] AMFUNet源自W. Y. Chung, I. H. Lee, and C. G. Park, “Lightweightinfrared small target detection network using full-scale skip connection u-net,” IEEE Geosci. Remote Sens. Lett., vol. 20, pp. 1–5, 2023.
[0126] RDIANet源自H. Sun, J. Bai, F. Yang, and X. Bai, “Receptive-field anddirection induced attention network for infrared dim small target detectionwith a large-scale dataset irdst,” IEEE Trans. Geosci. Remote Sens., vol. 61,pp. 1–13, 2023.
[0127] RPCANet源自F. Wu, T. Zhang, L. Li, Y. Huang, and Z. Peng, “Rpcanet:Deep unfolding rpca based infrared small target detection,” in Proc. IEEEWinter Conf. Appl. Comput. Vis. (WACV), 2024, pp. 4809–4818.
[0128] MSHNet is derived from Q. Liu, R. Liu, B. Zheng, H. Wang, and Y. Fu, “Infraredsmall target detection with scale and location sensitivity,” in Proc. IEEEConf. Comput. Vis. Pattern Recognit. (CVPR), 2024, pp. 17 490–17 499.
[0129] Figure 7 The original image of the infrared image example frame and the thermal infrared small target detection result image detected by the method of the present invention, MDWCM, DGRL, STT-TRNR, TWTVR, FGLR-MCP, ALCNet, ISTDUNet, DNANet, AGPCNet, AMFUNet, RDIANet, RPCANet and MSHNet. It can be seen intuitively from the infrared small target detection result image that the present invention can effectively complete the infrared small target detection and obtain a pure target detection image. Figure 7 It can be seen that MDWCM, ALCNet, ISTDUNet, and AGPCNet have target loss problems, STT-TRNR, TWTVR, and FGLR-MCP have a large amount of background components, and the detection results of other methods have clutter interference, such as Figure 7 As shown in the area marked with a yellow dotted circle in the middle; in contrast, the infrared sequence weak target detection method based on frequency domain information proposed in the present invention can enhance the saliency of small targets and suppress non-target components. In the field of target detection, target enhancement capability and background suppression capability are a pair of contradictory indicators. Improving target enhancement capability will lead to a weakening of background suppression capability, and vice versa. Therefore, it is necessary to comprehensively evaluate the performance of infrared weak target detection methods. According to the indicator results based on the 3D-ROC evaluation system shown in Table 1, the method proposed in the present invention achieves the best performance in target detection capability, background suppression capability and comprehensive capability, followed by DGRL; other methods basically do not meet the needs of practical applications. Combining the above qualitative and quantitative analysis, the infrared sequence weak target detection method based on frequency domain information proposed in the present invention has superior target detection capability, background suppression capability and comprehensive effectiveness.
[0130] Table 1 - Quantitative indicators of detection results of thermal infrared image example sequences using MDWCM, DGRL, STT-TRNR, TWTVR, FGLR-MCP, ALCNet, ISTDUNet, DNANet, AGPCNet, AMFUNet, RDIANet, RPCANet, MSHNet and the method of the present invention
[0131]
[0132] The accompanying drawings illustrating the embodiments of the present invention serve to more clearly illustrate the objectives, technical solutions, and advantages of the present invention. It should be noted that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. Any equivalent substitutions, modifications, and the like made within the methodologies and principles provided by the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for detecting small targets in infrared sequences based on frequency domain information, characterized in that: The steps include: Step 1): Preprocess each infrared image in the thermal infrared image sequence to be detected to obtain initial features; at the same time, stack all infrared images in sequence according to the time dimension to obtain a spatiotemporal tensor; Step 2): Design a frequency domain feature extraction module, which includes five layers of feature extraction modules. Each module contains two frequency domain transformation structures and a multi-head self-attention mechanism structure. The initial features obtained in step 1) are sequentially subjected to frequency domain transformation calculations and multi-head self-attention mechanism calculations to obtain attention features. Step 3): Perform steps 1) to 2) on N image frames to obtain N attention features; stack the N attention features obtained by the five-layer feature extraction module along the channel dimension to obtain the attention features after aggregation of the feature extraction modules of each layer; Step 4): Design a background reconstruction module to upsample the attention features aggregated by the fifth-layer feature extraction module layer by layer, and then sequentially perform channel splicing with the attention features aggregated by the first four layers of feature extraction modules, and then gradually reconstruct the background using a three-dimensional convolutional decoding method; further enhance the details through two refinement modules to obtain a reconstructed background tensor with the same scale as the original spatiotemporal tensor; Step 5): Design a loss function and train the infrared sequence small target detection network based on frequency domain information, which consists of a frequency domain feature extraction module and a background reconstruction module, in an unsupervised learning manner; Step 6): Input the thermal infrared image sequence to be detected into the infrared sequence weak target detection network trained in step 5), calculate the residual between its corresponding spatiotemporal tensor and the reconstructed background tensor, design a filter to filter the residual information, obtain the final target tensor, and reconstruct it into an infrared small target detection result sequence to achieve infrared small target detection.
2. The infrared sequence small target detection method based on frequency domain information according to claim 1 is characterized in that: The step 1) includes: For a thermal infrared image sequence consisting of N image frames, each frame , , , Respectively represent the height, width, and number of channels of the infrared image frame; Perform preprocessing to obtain initial features : ; in, , p represents the position encoding parameter, Drop represents dropout regularization, and PE represents two-dimensional patchembedding; At the same time, all infrared images are stacked in sequence according to the time dimension to obtain the space-time tensor , , , , Represent the space-time tensor The number of forward slices, height, width, and number of channels of .
3. The infrared sequence small target detection method based on frequency domain information according to claim 1 is characterized in that: In step 2), for the i-th layer feature extraction module, the calculation method is as follows: ; in, represents the input features of the i-th layer feature extraction module, represents the features obtained by the feature extraction module at layer i, represents the i-th layer feature extraction module, Represents the computational operation of the multi-head self-attention mechanism in the feature extraction module of the i-th layer, Frequency domain transformation operation in the i-th layer feature extraction module.
4. The infrared sequence small target detection method based on frequency domain information according to claim 3 is characterized in that: The calculation process of the frequency domain transformation structure processing in step 2) is specifically as follows: For input features , first perform a custom normalization operation DyT on the input: ; in, Represents input features The output features after custom normalization operation, represents the tanh activation function, , , represents the network learning parameters; Pair Features Perform a two-dimensional real fast Fourier transform , and then multiply by a learnable complex domain weight , through the two-dimensional real inverse fast Fourier transform Transform to the spatial domain to obtain the frequency domain enhancement result : ; in, Represents the element-wise multiplication symbol; Carry out residual enhancement processing, first Normalization, MLP and dropout regularization are performed, and then The residual connection itself gives: ; in, represents the output features obtained by residual enhancement processing, represents the element addition symbol, Represents multi-layer fully connected layer computing operations; After two consecutive frequency domain transformation structures are processed, the frequency domain features are obtained .
5. The infrared sequence small target detection method based on frequency domain information according to claim 4 is characterized in that: The calculation process of the multi-head self-attention mechanism structure in step 2) is specifically as follows: The obtained frequency domain features After normalization, multi-head attention calculation is carried out: ; in, Represents the attention feature, and MSAttn represents the multi-head self-attention mechanism operation.
6. The infrared sequence small target detection method based on frequency domain information according to claim 1 is characterized in that: The step 3) includes: Repeat steps 1) to 2) for each of the N image frames in the thermal infrared image sequence to be detected, and obtain N attention features. , ; For the i-th feature extraction module, the N attention features corresponding to the N image frames are stacked along the channel dimension to obtain the aggregated attention features ; Finally, a multi-level attention feature set is obtained .
7. The infrared sequence small target detection method based on frequency domain information according to claim 6 is characterized in that: The step 4) includes: In the background reconstruction process, the attention features aggregated by the fifth-layer feature extraction module are upsampled to have the same resolution as the original infrared sequence; the attention features are Input to the first background reconstruction layer , get the initial background estimate : ; For each subsequent scale , the reconstruction result Attention features after scale alignment Perform channel stitching and input the next background reconstruction layer , gradually complete the background reconstruction; among them, the attention feature The calculation formula for scale alignment is: ; in, , represents three-dimensional point convolution, represents the upsampling operation, represents the batch normalization operation, represents the leakyReLU activation function, Represents the attention features obtained after scale alignment; Background reconstruction layer The background reconstruction process is: ; in, ; Indicates that the features are spliced according to the channel dimension. Represents the i-th background reconstruction layer The input is also the Layer Background Reconstruction Layer Output reconstruction amount; background reconstruction layer Specifically ; in, Represents the i-th layer 3D transposed convolution of ; Reconstruct the last background layer The calculated reconstruction amount Input two background refinement modules in sequence to restore details, and obtain the original spatiotemporal tensor through normalization and activation function. Scale-consistent reconstruction background , the calculation formula is as follows: ; in, represents the Sigmoid function, and Represents the background refinement module, and has: 。 8. The infrared sequence small target detection method based on frequency domain information according to claim 1 is characterized in that: In step 5), the loss function is set as follows: First, define the mean squared error loss : ; in, represents the square of the Frobenius norm; Second, define the multi-scale structural similarity loss :For the Frame infrared image and the corresponding reconstruction background , through Gaussian blur processing and step-by-step downsampling operations, we get K pairs of images with decreasing resolution and ; At each scale, calculate the local mean of the image and , local standard deviation and , local covariance , calculate the structural similarity at this scale Similarity with contrast , the calculation formula is: ; ; in, and Is a positive constant; take the weighted combination of SSIM and MCS at all scales to obtain the structural similarity loss as follows: ; in, is the weight coefficient; Designing the total loss function : ; in 、 represents the weight coefficient; Therefore, using the total loss function The background reconstruction network is trained in an unsupervised learning manner.
9. The infrared sequence small target detection method based on frequency domain information according to claim 7, characterized in that: The step 6) is specifically as follows: Input the thermal infrared image sequence to be detected into the network trained in step 5) and calculate the spatiotemporal tensor and reconstructed background tensor The residual between them is obtained as the target tensor , design filter pair Information filtering is performed by introducing sparse cardinality selector and tubular fiber filtering operation. First, the sparse cardinality selector is used to retain Each front section of the High response pixels, and the pixel values of other positions are set to zero to obtain a temporary target tensor ; Secondly, design a tubular fiber filter to All pixel positions in , calculate the time-averaged response value: ; Constructing a mask ,when hour, represents the threshold value, If it is 0, otherwise it is 1, the mask Perform element-wise multiplication on the temporary target tensor Each front slice of , thereby eliminating interference points or dead points, and obtaining the final target tensor ; Will Reconstruct it into infrared small target detection result sequence T to realize infrared small target detection.
10. An infrared small target detection device based on frequency domain information for implementing the method of claim 1, characterized in that: include: The preprocessing and spatiotemporal tensor construction module preprocesses each frame of the infrared image sequence to obtain initial features and constructs the infrared image sequence into a spatiotemporal tensor; A frequency domain feature extraction module is used to extract features of each frame of infrared image. The frequency domain feature extraction module includes five layers of feature extraction modules, each of which contains two frequency domain transformation structures and a multi-head self-attention mechanism calculation structure; The feature aggregation module aggregates the N attention features obtained by the five-layer feature extraction module for each frame image to obtain an attention feature set; The background reconstruction module performs layer-by-layer upsampling on the attention feature set and designs two refinement modules to further enhance the details, obtaining a reconstructed background tensor with the same scale as the original spatiotemporal tensor; Loss function design module, which designs background reconstruction loss functions, including mean square error loss and multi-scale structural similarity loss, to train the network in an unsupervised learning manner; The information filtering module designs a filter and performs information filtering on the residual between the original spatiotemporal tensor and the reconstructed background tensor to obtain the final target tensor; The target detection result output module is used to output the infrared small target detection result map of each frame of infrared image in the infrared image sequence.
Citation Information
Patent Citations
Infrared small target detection method and device based on spatio-temporal information completion model
CN118196658A
Infrared small target detection method based on frequency domain decomposition
CN120163973A