An Infrared Low-Speed Slow-Moving Small Target Enhancement Method Based on Frequency Domain Enhancement and Spatiotemporal Prior
By constructing a super-resolution reconstruction model that incorporates infrared radiation characteristic coding, sparse optical flow alignment, and multi-branch high-frequency enhancement, the problem of inaccurate inter-frame and intra-frame feature matching in infrared low-speed small target images is solved. This improves the image's texture detail restoration and feature alignment capabilities, and enhances the visibility and contrast of small targets.
Patent Information
- Application Number
- CN202511107530.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-08-08
AI Technical Summary
In existing technologies, the inter-frame and intra-frame feature matching and alignment accuracy of infrared low-speed small target images is not high, which makes it easy to miss detections and false detections during the detection process, affecting the reliability of subsequent tasks.
An infrared low-speed small target enhancement method based on frequency domain enhancement and spatiotemporal prior is adopted. The super-resolution reconstruction model is constructed by four parts: infrared radiation characteristic encoding, coarse-grained inter-frame sparse optical flow alignment, coarse-grained intra-frame multi-branch high-frequency enhancement, and fine-grained inter-frame-intra-frame prior attention fusion, which improves the accuracy of feature matching and alignment.
It improves the texture detail recovery quality and feature alignment capability of infrared low-speed small target images, enhances the visibility and contrast of small targets, and adapts to infrared low-speed small target reconstruction tasks under different complex backgrounds.
Smart Images

Figure CN120612559B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image super-resolution reconstruction and is an infrared low-speed small target enhancement method based on frequency domain enhancement and spatiotemporal prior. Background Technology
[0002] Image super-resolution reconstruction is an important research direction in computer vision, aiming to recover high-resolution images from low-resolution images using reference information. Image super-resolution reconstruction techniques are commonly used in remote sensing image processing, medical image analysis, video enhancement, security monitoring, satellite navigation, and industrial inspection. Improving image resolution has a significant impact on tasks such as target recognition, feature extraction, diagnostic accuracy, and automated analysis. However, infrared images of small, slow-moving targets inherently suffer from low resolution, blurred details, and missing target texture information, leading to frequent missed and false detections during detection, severely affecting the reliability of subsequent tasks. Therefore, in recent years, researchers have begun to focus on utilizing the temporal information in consecutive frame infrared image sequences to improve image quality through super-resolution reconstruction techniques, thereby enhancing the visibility and contrast of small targets. Thus, fully exploiting the local spatial and temporal redundancy information in image sequences to reconstruct high-resolution, high-contrast infrared images has become a current research hotspot.
[0003] The Chinese patent publication number is "CN112102163A", entitled "A Super-Resolution Reconstruction Method for Continuous Multi-Frame Images Based on a Multi-Scale Motion Compensation Framework and Recursive Learning". This method constructs a deep neural network that includes feature extraction, non-local attention, multi-scale alignment, recursive upsampling, and reconstruction modules by setting the first frame in a series of consecutive multi-frame images as the reference frame and recursively processing adjacent frames. In each recursion, image features are extracted and aligned, and high-resolution feature reconstruction is achieved through learning upsampling filters and local convolution operations, ultimately outputting a clear image.
[0004] However, this method does not address the problem of poor granularity of inter-frame and intra-frame features in the detection field, and is prone to feature alignment misalignment. Therefore, designing a super-resolution reconstruction method for infrared low-speed, small target images that can reduce target feature misalignment, enhance intra-frame feature utilization efficiency, and improve inter-frame feature alignment capability is a problem that this invention urgently needs to solve. Summary of the Invention
[0005] (a) Technical problems to be solved
[0006] To address the shortcomings of existing technologies, this invention provides an infrared low-speed small target enhancement method based on frequency domain enhancement and spatiotemporal prior, which solves the problem of low matching and alignment accuracy between consecutive frame images. Furthermore, it alleviates the phenomenon of insufficient utilization of global correlation between consecutive frame images and local intra-frame structure, thus solving the problems mentioned in the background technology.
[0007] (II) Technical Solution
[0008] To achieve the above objectives, the present invention specifically adopts the following technical solution:
[0009] An infrared low-speed, small target enhancement method based on frequency domain enhancement and spatiotemporal prior includes the following steps:
[0010] S1. Prepare datasets: Prepare three continuous frame infrared low-speed small target detection datasets, and make a custom dataset. Dataset 1 and Dataset 4 are used for network training and model fine-tuning, while Dataset 2 and Dataset 3 are used for model testing.
[0011] S2, Construct an infrared low-speed small target super-resolution reconstruction model: The reconstruction model consists of four parts: infrared radiation characteristic encoding, coarse-grained inter-frame sparse optical flow alignment, coarse-grained intra-frame multi-branch high-frequency enhancement, and fine-grained inter-frame-intra-frame prior attention fusion.
[0012] S3, Train the network model: Train the infrared low-speed small target image super-resolution reconstruction model by inputting the dataset 1 and dataset 4 prepared in step S1 into the reconstruction model constructed in step S2 and training it.
[0013] S4. Select a suitable loss function and determine the optimal evaluation metric for this method: Select a suitable loss function that minimizes the difference between the output reconstructed image and the input reference image, and minimizes the loss. Set a training loss threshold and iteratively optimize the model until the number of training iterations reaches the set threshold or the value of the loss function reaches the set threshold range. The model parameters are then considered to have been pre-trained and saved. Select test images from datasets 2 and 3 and input them into the fixed model to obtain the target prediction results. Use the optimal evaluation metric for super-resolution reconstruction effect to measure the accuracy and performance of the model.
[0014] S5, fine-tuning the model: The model was trained and fine-tuned using Infrared Low Slow Small Target Detection Dataset 2 and Dataset 3 to optimize model parameters, further improve the performance of the reconstruction network, and obtain reconstruction results with clearer edges and more accurate textures.
[0015] S6, save the model. After the fine-tuning training in S5 is completed, solidify the fine-tuned network parameters and determine the final super-resolution reconstruction model.
[0016] Furthermore, in S1, dataset one is the SAITD dataset; dataset two is the Hui dataset; dataset three is the Anti-UAV dataset; and dataset four is a self-made dataset.
[0017] Furthermore, in S2, the infrared low-speed small target image super-resolution reconstruction model includes four parts: infrared radiation characteristic encoding, coarse-grained inter-frame sparse optical flow alignment, coarse-grained intra-frame multi-branch high-frequency enhancement, and fine-grained inter-frame-intra-frame prior attention fusion. The infrared radiation characteristic encoding part provides the model with a leap from "image pixel-driven" to "physical attribute-driven" capabilities, providing a more interpretable modeling foundation for subsequent intelligent reconstruction and target detection. The coarse-grained inter-frame sparse optical flow alignment part can provide more stable prior information for optical flow estimation in the shallow feature stage, improving the matching accuracy of low-speed small target motion. The coarse-grained intra-frame multi-branch high-frequency enhancement part enhances the local texture expression capability through multi-branch structure, and improves the texture detail recovery quality of low-speed small targets by combining wavelet inverse reconstruction and feature fusion. The fine-grained inter-frame-intra-frame prior attention fusion part can enhance the local correlation of small target motion features in the deep feature stage, optimizing the feature alignment and information aggregation capabilities of low-speed small targets in complex motion scenes.
[0018] Furthermore, in S2, the infrared radiation characteristic encoding consists of a linear layer, an activation function, a normalization layer, and Dropout;
[0019] Furthermore, in S2, the coarse-grained inter-frame sparse optical flow alignment consists of thermal gradient enhanced feature matching, multi-scale optical flow fusion estimation, and constrained deformable alignment convolution.
[0020] The thermal gradient enhancement feature matching consists of thermal gradient operator design, Harris corner point extraction, and thermal sensing feature descriptor; by analyzing the thermal radiation characteristics of infrared images, an adaptive thermal gradient operator is constructed to enhance target edge information.
[0021] The multi-scale optical flow fusion estimation consists of a multi-scale pyramid, an optical flow calculation module, and an optical flow confidence assessment. By constructing a pyramid structure to estimate optical flow layer by layer from top to bottom and implementing a confidence assessment mechanism, it achieves coarse-to-fine and explicit-to-implicit multi-source information fusion.
[0022] The constrained deformable aligned convolution consists of multiple constraint mechanisms and an adaptive masking mechanism. By integrating multiple constraint mechanisms such as spatial range restriction, motion consistency, neighborhood smoothness and temporal consistency, the stability and physical rationality of inter-frame feature alignment are improved.
[0023] Furthermore, in S2, the coarse-grained intra-frame multi-branch high-frequency enhancement part consists of wavelet decomposition, SE attention mechanism, 3 sets of mutually independent learnable parameters, 2 sets of standard convolutional blocks, and 2 sets of activatable convolutional blocks; each set of standard convolutional blocks has the same structure, consisting of 1 1×1 convolutional layer and 1 BN layer; each set of activatable convolutional blocks has the same structure, consisting of 1 3×3 convolutional layer, 1 BN layer, and L-shaped activation layer;
[0024] Furthermore, the fine-grained inter-frame and intra-frame prior attention fusion part consists of 4 alignment modules and 3 central difference residual dense block groups; each alignment module has the same structure, consisting of 2 central difference convolutions and element-wise dot multiplication operations; each central difference residual dense block group has the same structure, consisting of 4 central difference convolution residual blocks and dense connections; each central difference convolution residual block has the same structure, consisting of central difference convolution, BN layer, L-shaped activation layer and residual connections;
[0025] Furthermore, in step S4, the loss function is the mean squared error loss function;
[0026] The mean square error loss function is used to measure the pixel-level difference between the super-resolution image and the real image, thereby achieving the infrared image super-resolution task of high-fidelity reconstruction.
[0027] Furthermore, in step S4, the performance of the algorithm's prediction results is evaluated using evaluation metrics during the training of the network model.
[0028] (III) Beneficial Effects
[0029] Compared with existing technologies, this invention provides an infrared low-speed, slow-moving, small target enhancement method based on frequency domain enhancement and spatiotemporal prior, which has the following beneficial effects.
[0030] (1) The infrared radiation characteristic encoding method proposed in this invention reduces the dimension of the original thermal radiation information and enhances its discriminability. The output encoding vector is embedded as a physical prior into the subsequent alignment and enhancement module, thereby improving the thermal perception capability of feature modeling and the distinguishability of low, slow and small target regions.
[0031] (2) The coarse-grained inter-frame sparse optical flow alignment module proposed in this invention combines optical flow estimation and feature matching. It performs sparse feature matching in the initial stage of the pyramid, providing more stable prior information for optical flow estimation and improving the matching accuracy of small target motion. Furthermore, it further realizes the first pixel-level feature correction and alignment through deformable convolution, thereby improving the utilization efficiency of inter-frame temporal information in complex motion scenarios.
[0032] (3) The intra-frame multi-branch high-frequency enhancement module proposed in this invention uses wavelet transform to separate high and low frequency information of the image, and adaptively controls the high frequency response through attention mechanism and learnable parameters. Furthermore, it enhances the local texture expression capability through multi-branch structure, and improves the texture detail restoration quality of small targets by combining wavelet inverse reconstruction and feature fusion.
[0033] (4) The fine-grained inter-frame-intra-frame prior attention fusion module proposed in this invention combines self-attention mechanism and central difference convolution. While extracting inter-frame temporal information, it enhances the local correlation of motion features of small targets. Furthermore, with the help of prior-guided feature fusion strategy, it optimizes the feature alignment and information aggregation capabilities of small targets in complex motion scenarios.
[0034] (5) The super-resolution reconstruction model proposed in this invention has shown good results in the SAITD dataset, Hui dataset and Anti-UAV dataset. Both quantitative evaluation indicators have been greatly improved, indicating that the image reconstruction method proposed in this paper has a very strong generalization ability when facing different complex backgrounds and different types of low, slow and small targets, and can adapt to most infrared low, slow and small target reconstruction tasks. Attached Figure Description
[0035] Figure 1 A flowchart of an infrared low-speed small target enhancement method based on frequency domain enhancement and spatiotemporal prior;
[0036] Figure 2 This is a schematic diagram of the infrared low-speed small target super-resolution reconstruction model constructed in this invention.
[0037] Figure 3 This is a structural diagram of the coarse-grained inter-frame sparse optical flow alignment module described in this invention;
[0038] Figure 4 The diagram shows the structure of the multi-scale optical flow fusion estimation module of the present invention, wherein (a) is the structure of the multi-scale optical flow fusion estimation module of the present invention, (b) is the structure of the optical flow calculation module of the present invention, and (c) is the structure of the optical flow basic module of the present invention.
[0039] Figure 5 This is a structural diagram of the coarse-grained intra-frame multi-branch high-frequency enhancement module of the present invention;
[0040] Figure 6 This is a structural diagram of the fine-grained inter-frame and intra-frame prior attention fusion module of the present invention;
[0041] Figure 7 This is a structural diagram of the alignment module of the present invention;
[0042] Figure 8 This is a diagram of the central difference convolution structure of the present invention;
[0043] Figure 9 This is a structural diagram of the central difference convolution residual block of the present invention;
[0044] Figure 10 This is a structural diagram of the central difference residual dense group module of the present invention;
[0045] Figure 11This is a qualitative comparison of the infrared low-speed small target super-resolution reconstruction method of the present invention with existing methods.
[0046] Figure 12 This diagram illustrates a comparison of evaluation metrics between the infrared low-speed small target super-resolution reconstruction method of this invention and existing methods. Detailed Implementation
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] Example 1:
[0049] like Figure 1 As shown in the figure, Embodiment 1 of the invention provides a flowchart of an infrared low-speed small target enhancement method based on frequency domain enhancement and spatiotemporal prior. The method specifically includes the following steps:
[0050] S1, Prepare the datasets: Prepare SAITD dataset 1 and self-made dataset 4 for training the entire detection network; prepare Hui dataset 2 for model fine-tuning; prepare Anti-UAV dataset 3 for end-to-end model testing.
[0051] S2, Constructing an Infrared Low-Slow-Small Target Super-Resolution Reconstruction Model: Constructing a super-resolution reconstruction network model consisting of infrared radiation characteristic encoding, coarse-grained inter-frame sparse optical flow alignment, coarse-grained intra-frame multi-branch high-frequency enhancement, and fine-grained inter-frame-intra-frame prior attention fusion.
[0052] The infrared super-resolution reconstruction model for low-speed, small targets comprises four parts: infrared radiation characteristic encoding, coarse-grained inter-frame sparse optical flow alignment, coarse-grained intra-frame multi-branch high-frequency enhancement, and fine-grained inter-frame-intra-frame prior attention fusion. The infrared radiation characteristic encoding part provides a leap from "image pixel-driven" to "physical attribute-driven" capabilities, offering a more interpretable modeling foundation for subsequent intelligent reconstruction and target detection. The coarse-grained inter-frame sparse optical flow alignment part provides more stable prior information for optical flow estimation at the shallow feature stage, improving the matching accuracy of low-speed, small target motion. The coarse-grained intra-frame multi-branch high-frequency enhancement part enhances local texture representation through a multi-branch structure and, combined with wavelet inverse reconstruction and feature fusion, improves the quality of texture detail restoration for low-speed, small targets. The fine-grained inter-frame-intra-frame prior attention fusion part enhances the local correlation of small target motion features at the deep feature stage, optimizing feature alignment and information aggregation capabilities for low-speed, small targets in complex motion scenes.
[0053] Infrared radiation characteristic encoding consists of a linear layer, an activation function, a normalization layer, and Dropout.
[0054] Coarse-grained inter-frame sparse optical flow alignment consists of thermal gradient-enhanced feature matching, multi-scale optical flow fusion estimation, and constrained deformable alignment convolution. Thermal gradient-enhanced feature matching analyzes the thermal radiation characteristics of infrared images to construct an adaptive thermal gradient operator to enhance target edge information. Multi-scale optical flow fusion estimation constructs a pyramid structure to estimate optical flow layer by layer from top to bottom and implements a confidence evaluation mechanism to achieve coarse-to-fine and explicit-to-implicit multi-source information fusion. Constrained deformable alignment convolution integrates multiple constraint mechanisms, including spatial range limitations, motion consistency, neighborhood smoothness, and temporal consistency, to improve the stability and physical rationality of inter-frame feature alignment. The coarse-grained intra-frame multi-branch high-frequency enhancement part consists of wavelet decomposition, SE attention mechanism, and... The system consists of 3 sets of independent learnable parameters, 2 sets of standard convolutional blocks, and 2 sets of activatable convolutional blocks. Each set of convolutional blocks has the same structure, consisting of one 1×1 convolutional layer and one batch normalization (BN) layer. Each set of activatable convolutional blocks has the same structure, consisting of one 3×3 convolutional layer, one BN layer, and an L-shaped activation layer. The fine-grained inter-frame and intra-frame prior attention fusion part consists of 4 alignment modules and 3 sets of central difference residual dense blocks. Each set of alignment modules has the same structure, consisting of two central difference convolutions and element-wise dot product operations. Each set of central difference residual dense blocks has the same structure, consisting of four central difference convolutional residual blocks and dense connections. Each set of central difference convolutional residual blocks has the same structure, consisting of central difference convolutions, BN layers, L-shaped activation layers, and residual connections.
[0055] S3, Train the network model: Train the infrared low-speed small target super-resolution reconstruction model by inputting the dataset 1 and dataset 4 prepared in step S1 into the reconstruction model constructed in step S2 and training it.
[0056] S4. Select a suitable loss function and determine the optimal evaluation metric for this method: Select a suitable loss function that minimizes the difference between the output reconstructed image and the input reference image, and minimizes the loss. Set a training loss threshold and iteratively optimize the model until the number of training iterations reaches the set threshold or the value of the loss function reaches the set threshold range. The model parameters are then considered to have been pre-trained and saved. Select test images from datasets 2 and 3 and input them into the solidified model to obtain the target prediction results. Use the optimal evaluation metric for super-resolution reconstruction to measure the accuracy and performance of the model. During training, the mean squared error loss function is selected to measure the pixel-level difference between the super-resolution image and the real image, thereby achieving high-fidelity reconstruction of infrared images for super-resolution tasks. Suitable evaluation metrics are Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM).
[0057] S5, fine-tuning the model: The model was trained and fine-tuned using Infrared Low-Slow Small Target Detection Dataset 2 and Dataset 3 to optimize model parameters, further improve the performance of the reconstruction network, and obtain reconstruction results with clearer edges and more accurate textures.
[0058] S6, save the model. After the fine-tuning training in S5 is completed, solidify the fine-tuned network parameters and determine the final super-resolution reconstruction model. If performing a continuous frame infrared low-speed small target super-resolution reconstruction task, you can directly input 7 consecutive frames of infrared low-speed small target images into the trained end-to-end network model to obtain the final reconstruction result.
[0059] Example 2:
[0060] like Figure 1 As shown, an infrared low-speed small target enhancement method based on frequency domain enhancement and spatiotemporal prior is presented. The method specifically includes the following steps:
[0061] S1. Prepare the datasets. Prepare Dataset 1 and Dataset 4 for training the super-resolution reconstruction network. The entire super-resolution reconstruction network will be trained. Dataset 1 is the SAITD dataset, which includes small targets such as drones and bright spots against complex backgrounds such as sky, water, vegetation, and buildings, totaling 350 data segments and 150,185 images. Dataset 4 is a self-made dataset, which includes small targets such as quadcopter drones, fixed-wing drones, and low-altitude balloons, totaling 10,000 infrared images and corresponding infrared radiation text data for each frame. Prepare Hui Dataset 2 for model fine-tuning. The dataset includes small targets such as multi-rotor drones and fixed-wing drones against complex backgrounds such as sky, ground, and ground-to-air junctions, totaling 22 data segments, 30 tracks, 16,177 images, and 16,944 targets. Prepare the Anti-UAV dataset, which includes complex backgrounds such as urban rail transit and urban buildings, with targets such as kites, balloons, and birds, totaling more than 724,000 images.
[0062] S2, Construct a super-resolution reconstruction model for infrared low-speed small target images, such as Figure 2 The diagram shown is a schematic of the infrared low-speed small target super-resolution reconstruction model constructed in this invention, which includes four parts: infrared radiation characteristic encoding, coarse-grained inter-frame sparse optical flow alignment, coarse-grained intra-frame multi-branch high-frequency enhancement, and fine-grained inter-frame-intra-frame prior attention fusion.
[0063] The infrared radiation characteristic encoding consists of 3 linear layers, 2 activation functions, 2 normalization layers, and 1 Dropout.
[0064] Coarse-grained inter-frame sparse optical flow alignment consists of thermal gradient enhanced feature matching, multi-scale optical flow fusion estimation, and constrained deformable alignment convolution.
[0065] Thermal gradient-enhanced feature matching consists of thermal gradient operator design, Harris corner extraction, and thermal sensing feature descriptor.
[0066] ① Thermal gradient operator design
[0067] Based on the spatial consistency and local contrast characteristics of temperature distribution in infrared images, this paper designs a thermal gradient operator. This operator incorporates information such as temperature sensitivity factors and thermal distribution-guided weights during the calculation process, effectively enhancing the response of target edges and regions with abrupt local temperature differences in infrared images. The mathematical expression of this operator is as follows:
[0068]
[0069] in, and They represent (Horizontal) and The thermal gradient components in the longitudinal direction are mathematically expressed as follows:
[0070]
[0071]
[0072] Among them, thermal gradient kernel and The mathematical expression is as follows:
[0073]
[0074]
[0075] in, This represents the thermal sensitivity coefficient, which adaptively adjusts based on the local and global temperature variance of the image to enhance thermal gradient contrast. The mathematical expression of this coefficient is as follows:
[0076]
[0077] in, This represents a weighting function based on temperature distribution, designed according to that distribution. The mathematical expression of this function is shown below:
[0078]
[0079] Among them, adaptive weights The design is based on local temperature distribution, and its mathematical expression is as follows:
[0080]
[0081] in, and These represent the mean and variance within the local window, respectively. This represents the temperature contrast enhancement factor.
[0082] ② Harris Corner Extraction
[0083] Secondly, an improved Harris corner detector is used to extract feature points from the enhanced image based on thermal gradients. Based on the directional derivative information generated by the thermal gradient image, the corresponding structure tensor is defined as follows:
[0084]
[0085] in, and These represent the images in and Thermal gradient response in the direction. Structure tensor. The covariance information that reflects the changes in gray level (in this case, thermal gradient intensity) within a local window essentially describes the directional features and gradient distribution patterns of a local region of the image.
[0086] Meanwhile, the two eigenvalues corresponding to the structure tensor in the corner region are both large, indicating that there are significant changes in the local region in both directions. The mathematical expression of the above corner response function is as follows:
[0087]
[0088] in, The determinant of the local gradient covariance matrix reflects the overall strength of gradient changes within the region. Represents the total amount of gradient change; The sensitivity adjustment coefficient, set empirically, is used to control the response difference between corner points and edge points.
[0089] ③ Thermal sensing feature descriptor
[0090] Then, to improve the accuracy of feature matching, this chapter designs a feature descriptor that integrates thermal radiation information. Based on the traditional SIFT descriptor, temperature gradient direction and amplitude information are introduced:
[0091]
[0092] Among them, temperature gradient descriptor Record the direction of the dominant temperature gradient:
[0093]
[0094] Temperature histogram descriptor Statistical analysis of local temperature distribution:
[0095]
[0096] ④ Robust feature matching
[0097] Finally, bidirectional nearest neighbor matching and RANSAC geometric consistency verification are employed, combined with initial optical flow estimation values for confidence weighting, to improve matching reliability. Matching confidence is evaluated using multiple factors, including descriptor distance ratio, geometric consistency, and thermal gradient direction consistency. The mathematical expression for bidirectional nearest neighbor matching combined with the RANSAC algorithm is shown below:
[0098]
[0099] The mathematical expression for the combined assessment of match confidence using descriptor distance ratio and geometric consistency is shown below:
[0100]
[0101] in, and Indicates the distance between the nearest neighbor and the second nearest neighbor. This represents the estimated geometric transformation.
[0102] The multi-scale optical flow fusion estimation consists of a multi-scale pyramid, an optical flow calculation module, and an optical flow confidence assessment.
[0103] 1) Constructing multi-scale pyramids
[0104] The three-tiered pyramid structure corresponds to image scales with downsampling factors of 1, 2, and 4, respectively, ensuring a coarse-to-fine motion estimation process. The mathematical expression of the above process is as follows:
[0105]
[0106] The downsampling process employs global max pooling with an adaptive kernel size.
[0107] 2) Optical Flow Calculation Module
[0108] The optical flow module consists of four basic optical flow modules, one 3×3 standard convolutional layer, and bilinear interpolation. The basic optical flow modules include one 3×3 standard convolutional layer, the Mish activation function, group normalization, and Dropout regularization. The Mish activation function provides better gradient flow.
[0109] 3) Optical flow confidence assessment
[0110] At each scale, three confidence indices, including forward and backward consistency, feature matching consistency, and optical flow smoothness, are designed and weighted to improve the confidence of the final optical flow estimate.
[0111] The mathematical expression for the forward-backward consistency confidence is as follows:
[0112]
[0113] The mathematical expression for the confidence level of optical flow smoothness is as follows:
[0114]
[0115] The mathematical expression for feature matching confidence is as follows:
[0116]
[0117] The mathematical expression for the overall confidence level is as follows:
[0118]
[0119] Constrained deformable aligned convolution consists of multiple constraint mechanisms and adaptive masking mechanisms.
[0120] I. Multiple Constraint Mechanisms
[0121] This paper introduces physical rationality and spatiotemporal consistency constraints from four dimensions: spatial range, motion consistency, neighborhood smoothness, and temporal continuity. The aim is to regulate the behavior of offset prediction, suppress non-physical deformation, and improve the positioning accuracy and temporal stability of low, slow, and small targets.
[0122] Spatial Range Constraint. The goal of this constraint is to limit the maximum spatial span of the offset, preventing the offset sampling points from deviating excessively from the effective target area. Since small, slow-moving targets in infrared images typically occupy only a small spatial range, allowing large-scale free deformation can easily cause sampling points to fall into invalid background areas, thus introducing noise or misalignment. Therefore, this chapter uses the tanh nonlinear mapping function to normalize the original offset and introduces a maximum offset range. This achieves dynamic suppression of the offset amplitude. The mathematical expression of the above constraints is shown below:
[0123]
[0124] in, Adaptive settings based on target size:
[0125]
[0126] Motion consistency constraint. This constraint ensures that the predicted offset is consistent with the direction and amplitude of the optical flow estimate. Since the optical flow estimate usually contains some global dynamic information and can serve as a physical reference for offset learning, introducing optical flow as a "weak supervision signal" to constrain the offset helps guide it towards a more reasonable alignment direction. The mathematical expression of the above constraint is as follows:
[0127]
[0128] in, This represents the initial optical flow estimation result. This represents a weighted coefficient that incorporates multiple confidence factors to enhance the impact of high-confidence regions.
[0129] Neighborhood smoothing constraint. This constraint aims to maintain a continuous and smooth change in offset between spatially adjacent pixels, avoiding drastic jumps or discontinuous sampling, thereby improving the spatial consistency and structural integrity of feature alignment. This paper adopts an edge-preserving smoothing regularization term, which ensures strong smoothness in non-edge regions while allowing for appropriate relaxation of the constraint at image edges. The mathematical expression of the above constraint is shown below:
[0130]
[0131] in, This represents the gradient of the offset in space. This represents the edge strength of the image. By using an edge-guided exponential decay function, necessary structural changes can be preserved at image edges, while forced smoothing is applied to smooth areas, improving the spatial stability and visual consistency of the alignment results.
[0132] Temporal consistency constraint. This constraint ensures good temporal continuity in offset predictions across frames; that is, the offset of the current frame should be temporally consistent with the alignment result of the previous frame. This constraint is particularly suitable for aligning consecutive frames in video sequences, helping to improve the overall temporal stability of the sequence. The mathematical expression of this constraint is shown below:
[0133]
[0134] in, It is a feature transformation function based on optical flow, used to project the offset field of the previous frame into the coordinate system of the current frame.
[0135] II. Adaptive Masking Mechanism
[0136] This paper proposes an adaptive masking mechanism that integrates multiple factors to dynamically adjust the weight allocation of each sampling point in deformable convolution. The specific multi-factor adaptive masking design is as follows:
[0137]
[0138] in, This represents the initial mask learned by the basic convolutional network. The graph represents the overall confidence level of the optical flow estimation. The following two terms are introduced physical prior adjustment factors, defined as follows:
[0139] Feature similarity mask This mask is used to measure whether there are consistent feature responses among the sampling points of the supporting frames in the current frame, reflecting the consistency of thermal radiation structure between the two frames. This chapter uses a Euclidean distance metric based on thermal sensing features:
[0140]
[0141] in, Indicates the features of the reference frame. This indicates support for frame features. This indicates that frame features will be supported via offset. The result after transformation This represents the hyperparameter used to adjust similarity sensitivity.
[0142] Time Consistency Mask This mask is used to characterize the degree of offset change at the same location in adjacent frames, and is weighted by calculating the difference between the current frame offset and the predicted offset of the previous frame:
[0143]
[0144] in, This represents the predicted value projected onto the current frame using the offset from the previous frame through optical flow transformation. This mechanism ensures the continuity of the offset in the time dimension, avoiding unstable sampling caused by inter-frame jitter or mismatches.
[0145] Through the combined regulation of the above factors, the system can adaptively suppress sampling behavior in untrusted regions while enhancing attention to target regions or weakly salient regions, ultimately achieving more reliable and stable inter-frame alignment, especially significantly improving the modeling of low, slow, small targets and blurred boundary regions.
[0146] III. Constrained Deformable Convolution Implementation
[0147] After completing the design of the offset constraint and mask adjustment mechanism, this paper further implements Constrained Deformable Aligned Convolution (CDCN). Based on the standard deformable convolution, this convolution introduces adaptive masks and constraint offsets to achieve spatial control and weight adjustment of the sampling region. The mathematical expression of CDCN is shown below:
[0148]
[0149] in, Indicates the number of sampling points. This represents the feature value sampled by offset on the supporting frames. This represents the convolution weights corresponding to the sampling positions. This represents the multi-factor weighted mask at that position. This represents the offset value after final constraint optimization.
[0150] IV. Optimization Objective Function for Constraint Offset
[0151] To obtain the optimal offset prediction, this chapter combines multiple constraint terms into a single optimization objective function, which is then trained using backpropagation. The mathematical expression of this objective function is as follows:
[0152]
[0153] in, This represents the alignment loss, which measures the consistency of features after alignment. This indicates a loss of motion consistency; Indicates neighborhood smoothness regularity; Indicates time consistency constraints; , and These represent the weight coefficients of the three regularization terms.
[0154] The coarse-grained intra-frame multi-branch high-frequency enhancement part consists of wavelet decomposition, SE attention mechanism, 3 sets of independent learnable parameters, 2 sets of standard convolutional blocks and 2 sets of activatable convolutional blocks; each set of standard convolutional blocks has the same structure, consisting of 1 1×1 convolution and 1 BN layer; each set of activatable convolutional blocks has the same structure, consisting of 1 3×3 convolution, 1 BN layer and L-shaped activation layer.
[0155] Wavelet transform is used for multi-directional expansion to more comprehensively capture the high-frequency characteristics of low, slow, and small targets. The input features are downsampled using Haar wavelet transform, dividing the input features into... The non-overlapping blocks are used to compute wavelet subbands. The input features are divided into odd rows and odd columns based on the non-overlapping blocks. ), odd rows and even columns ( Even rows and odd columns ( ) and even rows and even columns ( Four types of regional characteristics. Further, the information from these four types of regions is used to obtain four different wavelet sub-bands in different directions through linear calculations, mainly including... (Infrared image structural information) (High-frequency information in the horizontal direction, such as edges) (High-frequency information in the vertical direction, such as texture) (High-frequency information in the diagonal direction, such as complex details). The linear calculation rules are as follows:
[0156]
[0157]
[0158]
[0159]
[0160] The present invention applies SE attention mechanism to the high-frequency components in three directions and multiplies them with corresponding independent learnable parameters to adaptively adjust the weights of high-frequency information in different directions.
[0161] This invention selects high-frequency components using adaptive weights. , and with low-frequency components The high and low frequency components are input into multiple frequency branches and summed along the channel dimension. This sum is then subtracted from the low-frequency component by pixels to obtain cleaner background noise information. Next, the high and low frequency components are fed into a multi-branch convolutional structure. The independent branch design avoids interference between different frequency components, improving feature extraction efficiency. Specifically, the low-frequency component multi-branch convolutional structure focuses on modeling global structural information, ensuring the coherence of the reconstructed background; the high-frequency component multi-branch structure enhances edge, detail, and diagonal information, providing fine-grained feature information for subsequent feature recovery. Finally, four enhanced sub-bands are calculated through linear inverse transform, and the reconstructed feature image is obtained using Haar inverse transform.
[0162] The fine-grained inter-frame and intra-frame prior attention fusion part consists of 4 alignment modules, 3 central difference residual dense block groups, and 3 convolutional layers. Each alignment module has the same structure, consisting of 2 central difference convolutions and element-wise multiplication. Each central difference residual dense block group has the same structure, consisting of 4 central difference convolution residual blocks and dense connections. Each central difference convolution residual block has the same structure, consisting of central difference convolution, BN layer, L-shaped activation layer, and residual connections.
[0163] The fine-grained inter-frame and intra-frame prior attention fusion part first calculates the attention weights between adjacent frames and the reference frame through the alignment module, achieving more accurate inter-frame alignment and information fusion. Subsequently, it uses central difference convolution and dense connections to obtain the local region associations of the intra-frame image.
[0164] The alignment module generates a query by performing a center difference convolution on the reference frame. The supporting frames generate the key through the same central difference convolution. ) and Value ( Then, in the attention calculation phase, Query( ) and Key( The inter-frame similarity is calculated through element-wise dot product operations, thus obtaining the inter-frame matching score matrix:
[0165]
[0166] in, and These represent the feature indices of the reference frame and the supporting frame, respectively. The calculated similarity matrix... go through Normalization yields the attention weights:
[0167]
[0168] This normalization process ensures that the information fusion weights between different frames are reasonably allocated, so that highly correlated regions occupy a larger proportion in the alignment process, while mismatched information in the background region is effectively suppressed.
[0169] In the feature alignment and fusion stage, the calculated attention weights are used. For Value ( Perform a weighted summation to generate aligned, fine-grained feature maps:
[0170]
[0171] The alignment module ensures that the information from the supporting frames is accurately aligned to the coordinate space of the reference frames, while avoiding the cumulative errors that might arise from directly relying on optical flow offset. Finally, the aligned feature maps are fed into subsequent reconstruction steps to enhance the sharpness and discernibility of small targets. The alignment module optimizes inter-frame relationships through an adaptive attention mechanism, enabling the super-resolution model to fully utilize temporal information and improve the reconstruction quality and stability of small, slow-moving targets.
[0172] The central difference residual dense block consists of four central difference residual blocks, which fuse features from different levels through a dense connection mechanism. Within the central difference residual dense block, the output of each block is passed to subsequent layers and concatenated with the input of the current layer to ensure sufficient interaction between lower and higher level features. The mathematical expression of this operation is as follows:
[0173]
[0174] in, Indicates the first The output of the layer, Indicates the first Central difference residual block transformation of the layer, This represents the stack of inputs from all preceding layers. Furthermore, the central difference residual dense block employs global residual connections to ensure information flow between the input and output of the entire block, thereby avoiding the gradient vanishing problem. Finally, the features generated by the central difference residual dense block are fed into the subsequent super-resolution reconstruction module to enhance the sharpness and detail recovery of small targets.
[0175] The central difference residual block consists of central difference convolution, BN layer, L-shaped activation layer and residual connection.
[0176] S3, train the detection network model by inputting the dataset 1 and dataset 4 prepared in step S1 into the detection network model constructed in step S2 for training.
[0177] S4. Select a suitable loss function and determine the optimal evaluation metric for this method: Select a suitable loss function that minimizes the difference between the output prediction result and the input target image and minimizes the loss. Set a training loss threshold and iteratively optimize the model until the number of training iterations reaches the set threshold or the value of the loss function reaches the set threshold range. The model parameters can then be considered to have been pre-trained and saved. Select the test image from dataset 3 and input it into the network model to obtain the target prediction result. Use the optimal evaluation metric for the target prediction result to measure the accuracy and performance of the model.
[0178] In S4, the loss function calculated using the network output and reference image is the mean squared error loss function. It is used to measure the pixel-level difference between a super-resolution image and a real image, thereby enabling high-fidelity reconstruction of infrared images for super-resolution tasks. This can be expressed using the following formula:
[0179]
[0180] in, The first output of the model represents the... pixel value, The first reference image pixel value, This indicates the total number of pixels in the image.
[0181] Appropriate evaluation metrics include peak signal-to-noise ratio (PSNR) and structural similarity (SSIM).
[0182] In S5, the appropriate evaluation metrics are peak signal-to-noise ratio (PSNR) and structural similarity (SSIM).
[0183] Peak Signal-to-Noise Ratio (PSNR) measures the difference between the reconstructed image and the original image by calculating the difference between corresponding pixels in two images. PSNR can be expressed by the following formula:
[0184]
[0185] Where MAX is the highest pixel value in the image (usually expressed as dynamic range), MSE is the mean square error of the two compared images, and PSNR is in decibels (dB).
[0186] Structural similarity is determined by comparing three relatively independent components: intensity, contrast, and structure. The closer the SSIM result is to 1, the higher the similarity between the reconstructed image and the corresponding ground truth image. SSIM can be expressed by the following formula:
[0187]
[0188] in, and This represents the average gray level of all pixels in the two images. and The standard deviation of grayscale values. Represents the image covariance. and It represents any constant.
[0189] During network training, the learning rate was set to 0.001, the batch size to 8, and a total of 500 iterations were performed. The Adam optimizer was used to continuously update the network parameters, with its exponential decay rate and eps values set to (0.9, 0.999) and 1e-08, respectively. The entire training process lasted approximately 20 hours. To ensure that the four loss values in the total loss function were as close to the same order of magnitude as possible, the four loss weights were set to [values to be inserted here]. =0.1, =0.001, =1.
[0190] S5, Fine-tuning the model: The model was trained and fine-tuned again using Hui dataset II and Anti-UAV dataset III, with the learning rate set to 0.005 and a total of 500 iterations. Other parameters remained unchanged to further improve the performance of the image super-resolution reconstruction network.
[0191] S6. After the fine-tuning training in S5 is completed, solidify the fine-tuned network parameters and determine the final super-resolution reconstruction model. If performing a continuous frame infrared low-speed small target super-resolution reconstruction task, seven consecutive frames of infrared low-speed small target images can be directly input into the trained end-to-end network model to obtain the final reconstruction result.
[0192] The implementation of convolution, central difference convolution, activation functions, addition operations, etc., are algorithms known to those skilled in the art, and the specific processes and methods can be found in relevant textbooks or technical documents.
[0193] This invention constructs an infrared low-speed, slow-moving, small target enhancement method based on frequency domain enhancement and spatiotemporal prior, which can improve the super-resolution reconstruction effect of infrared low-speed, slow-moving, small targets. It eliminates intermediate steps and avoids the need for manually designed spatiotemporal feature extraction methods, simplifying and improving the efficiency of image super-resolution reconstruction. A qualitative comparison of the prediction results of existing technologies and the method proposed in this invention is as follows: Figure 11 As shown, under the same conditions, the correlation index between the image super-resolution reconstruction results and the reference image obtained by calculation and existing methods was further verified to demonstrate the feasibility and superiority of the proposed method.
[0194] A comparative diagram of evaluation indicators between existing technologies and the method proposed in this invention is shown below. Figure 12 As shown in the figure, the method proposed in this invention has a higher peak signal-to-noise ratio and structural similarity than existing methods. These indicators further demonstrate that the method proposed in this invention achieves superior infrared low-speed small target super-resolution reconstruction performance and obtains the expected results.
[0195] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for enhancing infrared low-speed, small targets based on frequency domain enhancement and spatiotemporal prior, characterized in that: The method specifically includes the following steps: S1, Prepare the dataset: Prepare three continuous frame infrared low-slow-moving small target detection datasets, and make a custom dataset. Dataset 1 and Dataset 4 are used for network training and model fine-tuning, and Dataset 2 and Dataset 3 are used for model testing. S2, Constructing a super-resolution reconstruction model for infrared low-speed small target images: The reconstruction model includes thermal gradient enhancement feature matching, which consists of thermal gradient operator design, Harris corner point extraction, and thermal sensing feature descriptor. ① Thermal gradient operator design Based on the spatial consistency and local contrast characteristics of temperature distribution in infrared images, a thermal gradient operator was designed. This operator incorporates information such as temperature sensitivity factors and thermal distribution-guided weights during the calculation process, effectively enhancing the response of target edges and regions with abrupt local temperature differences in infrared images. The mathematical expression of this operator is as follows: in, and Let x (horizontal) and y (vertical) represent the thermal gradient components, respectively. The mathematical expressions for the thermal gradient components in different directions are as follows: Among them, thermal gradient kernel and The mathematical expression is as follows: Where α represents the thermal sensitivity coefficient, which is adaptively adjusted based on the local and global temperature variances of the image to enhance thermal gradient contrast; the mathematical expression of this coefficient is shown below: Among them, W thermal (x, y) represents a weighting function based on the temperature distribution, which is designed based on the temperature distribution; the mathematical expression of this function is shown below: The adaptive weight A_{thermal}(x,y) is designed based on the local temperature distribution, and its mathematical expression is as follows: Where, μ local and σ local These represent the mean and variance within the local window, respectively, and β = 0.3 is the temperature contrast enhancement coefficient. ② Harris Corner Extraction Secondly, an improved Harris corner detector is used to extract feature points from the enhanced image based on thermal gradients. Based on the directional derivative information formed by the thermal gradient image, the corresponding structure tensor is defined as follows: Among them, G x and G y These represent the thermal gradient responses of the image in the x and y directions, respectively; the structure tensor M thermal The covariance information that reflects the gray-level (here, thermal gradient intensity) changes within the local window essentially describes the directional features and gradient distribution patterns of the local region of the image. Meanwhile, the two eigenvalues corresponding to the structure tensor in the corner region are both large, indicating that there are significant changes in the local region in both directions; the mathematical expression of the above corner response function is as follows: R(x,y)=det(M thermal )-k·trace 2 (M thermal ) in, The determinant of the local gradient covariance matrix reflects the overall strength of gradient changes within the region; This represents the total amount of gradient change; k is an empirically set sensitivity adjustment coefficient used to control the response difference between corner points and edge points; ③ Thermal sensing feature descriptor Then, to improve the accuracy of feature matching, this invention designs a feature descriptor that integrates thermal radiation information; based on the traditional SIFT descriptor, it introduces temperature gradient direction and amplitude information: D thermal =[D SIFT ,D temp_grad ,D temp_hist ] Among them, the temperature gradient descriptor D temp_grad Record the direction of the dominant temperature gradient: D temp_grad [i]=sum {(x,y)in R_i} G thermal(x,y) ×cos(θ thermal(x,y) -θ i ) Temperature histogram descriptor D temp_hist Statistical analysis of local temperature distribution: D temp_hist [j]=(1 / |R|)×sum {(x,y)in R} I[I(x,y)in Bin j ]; S3, Train the network model: Train the infrared low-speed small target image super-resolution reconstruction model by inputting the dataset 1 and dataset 4 prepared in step S1 into the reconstruction model constructed in step S2 and training it. S4. Select a suitable loss function and determine the optimal evaluation metric for this method: Select a suitable loss function that minimizes the difference between the output reconstructed image and the input reference image, and minimizes the loss. Set a training loss threshold and iteratively optimize the model until the number of training iterations reaches the set threshold or the value of the loss function reaches the set threshold range. The model parameters are then considered to have been pre-trained and saved. Select test images from datasets 2 and 3 and input them into the fixed model to obtain the target prediction results. Use the optimal evaluation metric for super-resolution reconstruction effect to measure the accuracy and performance of the model. S5, fine-tuning the model: The model was trained and fine-tuned using Infrared Low Slow Small Target Detection Dataset 2 and Dataset 3 to optimize model parameters, improve the performance of the reconstruction network, and obtain reconstruction results with clearer edges and more accurate textures. S6, save the model. After the fine-tuning training in S5 is completed, solidify the fine-tuned network parameters and determine the final super-resolution reconstruction model.
2. The infrared low-speed, slow-moving, small target enhancement method based on frequency domain enhancement and spatiotemporal prior as described in claim 1, characterized in that: In S1, dataset one is the SAITD dataset; dataset two is the Hui dataset; dataset three is the Anti-UAV dataset; and dataset four is a self-made dataset.
3. The infrared low-speed, slow-moving small target enhancement method based on frequency domain enhancement and spatiotemporal prior as described in claim 1, characterized in that: S2 also includes infrared radiation characteristic encoding, which consists of a linear layer, an activation function, a normalization layer, and Dropout.
4. The infrared low-speed, slow-moving, small target enhancement method based on frequency domain enhancement and spatiotemporal prior as described in claim 1, characterized in that: S2 also includes coarse-grained intra-frame multi-branch high-frequency enhancement, which consists of wavelet decomposition, SE attention mechanism, 3 sets of mutually independent learnable parameters, 2 sets of standard convolutional blocks and 2 sets of activatable convolutional blocks; each set of standard convolutional blocks has the same structure, consisting of 1 1×1 convolution and 1 BN layer; each set of activatable convolutional blocks has the same structure, consisting of 1 3×3 convolution, 1 BN layer and L-shaped activation layer.
5. The infrared low-speed, slow-moving small target enhancement method based on frequency domain enhancement and spatiotemporal prior as described in claim 1, characterized in that: S2 further includes fine-grained inter-frame to intra-frame prior attention fusion, which consists of 4 alignment modules and 3 central difference residual dense block groups. Each alignment module has the same structure, consisting of 2 central difference convolutions and element-wise multiplication operations. Each central difference residual dense block group has the same structure, consisting of 4 central difference convolution residual blocks and dense connections. Each central difference convolution residual block group has the same structure, consisting of central difference convolution, BN layer, L-shaped activation layer and residual connections.
6. The infrared low-speed, slow-moving, small target enhancement method based on frequency domain enhancement and spatiotemporal prior as described in claim 1, characterized in that: In step S4, the loss function is the mean squared error loss function; The mean squared error loss function is used to measure the pixel-level difference between the super-resolution image and the real image, thereby achieving the infrared image super-resolution task of high-fidelity reconstruction.
7. The infrared low-speed, slow-moving, small target enhancement method based on frequency domain enhancement and spatiotemporal prior as described in claim 1, characterized in that: In step S4, the process of training the network model also includes evaluating the algorithm reconstruction performance through evaluation metrics.
Citation Information
Patent Citations
Continuous multi-frame image super-resolution reconstruction method based on multi-scale motion compensation framework and recursive learning
CN112102163A
Passive infrared sensor-based indoor behavior track movement semantic parsing method
CN107657215A
Sparse optical flow estimation
CN116235209A