An infrared image small target detection method and system based on parallel optimization and redundancy suppression network
Patent Information
- Application Number
- CN202610808847.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-25
AI Technical Summary
[0007]本发明的目的是提供一种基于并行优化与冗余抑制网络的红外影像小目标检测方法及系统,为了解决红外图像中小目标检测困难的问题
[0104]本发明是一种基于并行优化与冗余抑制网络的红外影像小目标检测方法。本发明通过多向池化注意力并行优化增强弱目标与背景信息的区分能力;本发明通过通过聚类压缩像素令牌得到语义中心,降低全局建模计算量,同时保留局部细节;本发明通过冗余抑制特征融合模块改善多尺度融合中的上采样失真和特征混叠;本发明通过高斯回归一致性损失函数提升红外小目标回归精度与训练稳定性;本发明整体方案可端到端训练,适合复杂场景下红外小目标实时检测。实验表明,与现有的红外小目标检测方法相比,本发明可以有效应对不同场景的背景杂波干扰,提高小目标细节表达,显著提高了红外小目标的检测精度。
Smart Images

Figure CN122821080A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of infrared image detection technology, specifically to a method and system for detecting small targets in infrared images based on parallel optimization and redundancy suppression networks. Background Technology
[0002] Infrared small target detection is widely used in scenarios such as early warning and surveillance, maritime search and rescue, urban security, and low-altitude target identification. Infrared small targets are typically characterized by their small size, sparse texture, low signal-to-noise ratio, and strong background interference, making them easily obscured by the background in images, significantly increasing the difficulty of detection compared to general target detection tasks.
[0003] Existing infrared small target detection methods mainly include traditional image processing methods and deep learning-based methods. Traditional methods typically rely on prior assumptions such as local contrast enhancement, background suppression, low-rank sparse decomposition, filtering, and saliency analysis. While these methods can achieve certain results in specific scenarios, they are sensitive to background complexity and imaging conditions, have limited generalization ability, and are prone to false positives and false negatives in complex interference environments. With the development of deep learning, detection methods based on convolutional neural networks have gradually become mainstream. These methods improve the automation and robustness of infrared small target detection through end-to-end feature representation learning. However, the small size and weak information of infrared small targets make it difficult for conventional convolutional networks to fully preserve their fine-grained features. At the same time, the network is prone to losing shallow positional information during downsampling, causing the target features to gradually weaken in high-level semantics.
[0004] On the other hand, multi-scale feature fusion is a common method to improve the detection performance of small targets. By fusing high-level semantic information with low-level detail information, the expressive power of targets at different scales can be enhanced. However, in infrared small target scenarios, multi-scale fusion is often accompanied by problems such as upsampling distortion, feature alignment deviation, and cross-layer information aliasing. In particular, during feature transfer, background redundancy information is easily mixed with target features, leading to a decrease in target saliency and further affecting the accuracy of the detection results.
[0005] Therefore, how to effectively enhance the weak feature representation capability of infrared small targets, suppress background interference, and improve the quality of multi-scale feature fusion under complex background and low signal-to-noise ratio conditions has become a technical problem that urgently needs to be solved in the current field of infrared small target detection. Summary of the Invention
[0006] The technical problem to be solved by this invention is:
[0007] The purpose of this invention is to provide a method and system for detecting small targets in infrared images based on parallel optimization and redundancy suppression networks, in order to solve the problem of difficulty in detecting small targets in infrared images.
[0008] The technical solution adopted by the present invention to solve the above problems is as follows:
[0009] The method includes the following steps:
[0010] (1) Preprocessing the infrared image and inputting it into the established detection network: Use Cubic interpolation to uniformly adjust the infrared image and corresponding label to a size of 640×640, label the minimum bounding box and category number of the target, construct a parallel optimization and redundancy suppression network, and then input the image into the network.
[0011] (2) Constructing a parallel optimization module: A dual-branch attention mechanism is adopted, which models the global background distribution and local target details through different pooling operations. Average pooling captures smoother and more blurred background information, while max pooling is used to preserve sharp details in the image, thereby enhancing the target representation.
[0012] (3) Constructing a clustering semantic attention module: global semantic centers are adaptively constructed through iterative clustering to drive attention computation. Two linear layers are mapped to global keys and values, and pixel tokens are used as queries to compute dot product attention. This is then integrated with a local attention mechanism to preserve local details, thereby achieving efficient global modeling and feature enhancement.
[0013] (4) Construct a redundancy suppression feature fusion module: Convolution downsample the input features, use frequency domain closed solution and learnable regularization to achieve top-down feature fusion, then fuse multi-scale features according to bottom-up path, preprocess the original features through lightweight convolution and reconstruct them, crop and align the fused features, use linear subtraction to suppress redundancy, and output three scale feature maps for bounding box decoupling.
[0014] (5) Model training and optimization using Gaussian regression consistency loss function: First, calculate the Euclidean distance between the center of the predicted box and the center of the real box, model it as a two-dimensional Gaussian distribution, calculate the distribution distance by combining the covariance matrix, and obtain the Gaussian regression consistency loss by fusing the position and distribution distance. At the same time, use the distribution focal loss and CIoU to optimize the regression, and use the binary cross-entropy for the classification loss; the final loss is a multi-weighted combination.
[0015] (6) Input the image to be detected into the trained model to obtain the detection result: Select the optimal network with the best test set index, input the infrared image to be detected into the network, and output the final infrared small target detection result.
[0016] Furthermore, in step (1) above, the infrared image is acquired and input into the established detection network:
[0017] Specifically, we used Cubic interpolation to unify the infrared image and labels to 640×640. During label creation, the minimum bounding box and category number of the target were annotated. Then, a parallel optimization and redundancy suppression network was constructed, and the infrared image was input into it. The network backbone consists of convolutional layers, a parallel optimization module, and a multi-cluster semantic attention module. The convolutional layers are used for downsampling, the parallel optimization module for feature extraction, and the multi-cluster semantic attention module for feature enhancement of the network backbone's P3 layer. After extracting feature maps at different scales, a redundancy suppression feature fusion module was used to achieve multi-scale feature fusion. Finally, the fused feature map was input into the YOLO detection head to obtain the detection result.
[0018] Furthermore, the construction of the parallel optimization module in step (2) above can be carried out according to the following steps:
[0019] Two parallel attention processing branches were added to the bottleneck structure. Each attention branch processes the feature map of the image. The process is divided into two directions: height and width. Each branch generates two sets of attention maps in different directions.
[0020]
[0021] Here, AvgPool represents average pooling, MaxPool represents maximum pooling, and the subscript indicates the pooling direction. and These represent the pooling results of the first branch along the h-axis and w-axis dimensions, respectively. and These represent the pooling results of the second branch along the h-axis and w-axis dimensions, respectively.
[0022] Subsequently, the pooling results in the width direction are subjected to dimension permutation, i.e. The permutation results are concatenated along the dimensions to obtain the intermediate feature Y.
[0023] Concat represents the concatenation operation.
[0024] Subsequently, 1×1 convolutional kernels are used to reduce the dimensionality of the intermediate feature Y and fuse cross-channel features, followed by batch normalization layers and the ReLU activation function. The processed intermediate features are then split, and dimensional permutations are performed to match the pooling shape of the original width dimension. Attention weights in different directions are obtained based on the features from the two branches.
[0025] (7)
[0026] (8)
[0027] (9)
[0028] (10)
[0029] in, This represents the sigmoid function. This represents a 1×1 convolution operation. It is The high-dimensional features obtained after splitting, It is The dimension permutation result of the width dimension feature obtained after splitting. It is The high-dimensional features obtained after splitting, It is Dimension permutation result of the width dimension feature obtained after splitting
[0030] Subsequently, the attention weights in each branch are applied together to the original feature map, and residual links are added.
[0031] (11)
[0032] Convolutional layers with different receptive fields (e.g., 1×1, 3×3, and 5×5) are used to fully fuse the information from these features.
[0033] Furthermore, the construction of the clustering semantic attention module in step (3) above can be performed according to the following steps:
[0034] First, the feature map X input to the module is expanded into a token sequence and then normalized by the layer. The token sequence is retained as a residual term for residual stacking.
[0035] (12)
[0036] in, and These represent the mean and variance of the token sequence along the channel dimension, respectively. To prevent tiny values with a denominator of 0, and These are learnable scaling and offset parameters.
[0037] A large number of pixel tokens are then compressed through clustering operations to form k semantic tokens. Pixel tokens With cluster center The similarity between them can be expressed as:
[0038] (13)
[0039] in, The L2 norm is used. Then, the mean of all tokens within each cluster is calculated to generate new cluster centers. The iterative update formula is:
[0040] (14)
[0041] in, The center of the k-th cluster in the (t+1)-th round, For the set of pixel tokens assigned to the k-th cluster in the t-th round, This is an L2 normalization operation.
[0042] To accommodate the small sample size of the infrared small target dataset, an exponential moving average (EMA) strategy is used to update cluster centers, preventing oscillations caused by sample fluctuations. The EMA update formula is:
[0043] (15)
[0044] in, For the updated cluster centers, As a historical cluster center, The new cluster centers obtained in the current iteration, This is the EMA attenuation coefficient.
[0045] The cluster centers obtained after three iterations of the module are used as global semantic tokens to construct global attention. Specifically, the cluster centers are mapped to global keys k and global values v through two linear layers. The pixel token is then used as the query (q), and a dot product attention score is calculated with k, as follows:
[0046] (16)
[0047] Here, c represents the number of channels, and h represents the number of attention heads. The attention scores are then normalized using the Softmax function and weighted and summed with v to obtain the global attention enhancement features for each pixel token.
[0048] Subsequently, the pixel tokens are divided into N / G groups. If the groups are not divisible, a flip-fill strategy is used to fill them. Self-attention calculations are performed within each group to capture the spatial relationships between pixels within the group. Simultaneously, global semantic tokens obtained from clustering are used for attention calculations within each group to fuse local details and global semantics. The fused result is then added element-wise to the self-attention result to generate the final local attention-enhanced feature.
[0049] Finally, the global attention enhancement features and local attention enhancement features are added together and then converted into feature tensors through dimensionality rearrangement. Subsequently, 1×1 convolutional kernels are applied to adjust the channels, and a convolutional feedback network (ConvFFN) is added to enhance feature representation.
[0050] Furthermore, the construction of the redundancy suppression feature fusion module in step (4) above can be performed according to the following steps:
[0051] For input features The convolution and downsampling process can be represented as:
[0052] (17)
[0053] in, It outputs the feature map, and performs a downsampling operation. It is applied per channel.
[0054] In this case, the original input feature map X is recovered by minimizing the error between the observed output Y and the reconstructed output. Therefore, the optimization problem can be expressed as:
[0055] (18)
[0056] in, and These are the input and output feature maps of the c-th channel, respectively. (Frobenius norm) It is still an error metric calculated separately for each channel.
[0057] In the frequency domain, the recovered solution is obtained by performing a Fourier transform on each channel. Let ∀f{F} denote the Fourier transform. and These are the observation output and the convolution kernel's representations in the frequency domain, respectively. Therefore, the closed-form solution in the frequency domain can be expressed as:
[0058] (19)
[0059] in, It is the frequency domain information of the c-th channel. It is the Fourier transform of the c-th channel after upsampling. The complex conjugate of the Fourier transform of convolution kernel K. This represents the element-wise squared magnitude of the Fourier transform of the convolution kernel. This represents element-wise multiplication on a d×d block. This indicates that the average downsampling is performed on the d×d block. This indicates a standard d-fold upsampling operation, which increases the spatial dimension by inserting zeros.
[0060] Through the aforementioned closed-form solution in the frequency domain and learnable regularization, a mathematically invertible learnable frequency-constrained (LFC) upsampler was obtained, and this operator was used for top-down feature fusion. Then, multi-scale features were further fused following a bottom-up path. This operation fuses high-level features with low-level features, further optimizing the transmission of contextual information. Subsequently, background and redundant features were suppressed.
[0061] Let the current fusion feature be The original features at the same scale are First, perform light preprocessing on the original features:
[0062] (20)
[0063] Here, BN represents batch normalization operation.
[0064] The features are then reconstructed using the LFC upsampler U:
[0065] (twenty one)
[0066] The upsampling results are then cropped to restore them to the same spatial size as the fused features.
[0067] (twenty two)
[0068] Where C2f represents the feature extraction module in the YOLOv8 model, Crop represents nearest neighbor interpolation, and [H, W] represents... The size of the redundancy is ultimately achieved using linear subtraction.
[0069] (twenty three)
[0070] in, This represents the suppressed output. 'a' represents the suppression factor. The entire suppression process is hierarchically optimized and added to the feature fusion process. The entire redundant suppression feature fusion module ultimately outputs three feature maps at different scales for bounding box decoupling.
[0071] Furthermore, the model training and optimization in step (5) above, using the Gaussian regression consistency loss function, can be performed as follows:
[0072] During training, the Gaussian regression consistency loss function is used to optimize the regression branch. First, the Euclidean distance between the predicted bounding box and the center of the ground truth bounding box is calculated:
[0073] (twenty four)
[0074] Among them, (x p , y p), (x g , y g The coordinates of the predicted bounding box and the ground truth bounding box are respectively. The predicted and ground truth bounding boxes are then modeled as two-dimensional Gaussian distributions, and the distribution distance is calculated using the covariance matrix.
[0075] (25)
[0076] (26)
[0077] (27)
[0078] in, Represents the covariance matrix. Tr is the trace of the covariance matrix obtained by solving. det represents the determinant of the covariance matrix calculated using Cholesky decomposition.
[0079] The Gaussian regression consistency loss function is obtained by fusing location distance and distribution distance:
[0080] (28)
[0081] in, This represents the balance coefficient. To maintain numerical stability, the fused loss value is normalized and smoothed:
[0082] (29)
[0083] (30)
[0084] The regression loss is calculated using a combination of distributed focal loss and CIoU. The formula for calculating distributed focal loss is as follows:
[0085] (31)
[0086] in, and These are the predicted values output by the network, and the nearest neighbor predicted values. , , The values are the actual value of the label, the label integral value, and the integral value of the neighboring labels.
[0087] The formula for calculating CIoU is as follows:
[0088] (32)
[0089] (33)
[0090] (34)
[0091] in, and This represents the center point of the prediction box and the truth box. Represents Euclidean distance. and Represents the width of the prediction box and the ground truth box. and This represents the height of the prediction box and the truth box.
[0092] The classification loss is calculated using binary cross-entropy. The formula is as follows:
[0093] (35)
[0094] in, It's a binary tag, meaning either 0 or 1. It outputs the probability of belonging to the label. This represents the number of groups of objects the model predicts. The final loss function calculation formula is as follows:
[0095] (36)
[0096] in, , , , These are the weighting coefficients for each type of loss.
[0097] The constructed network can be trained using a pre-made dataset. The dataset is divided into training, validation, and test sets in a 7:1:2 ratio. The SGD optimizer is selected as the network optimizer. The initial learning rate is 0.01, which decays exponentially in each iteration. The model is trained for 200 generations, and gradient updates are performed during iterations using the aforementioned loss to guide the direction of network parameter updates.
[0098] Furthermore, the detection using the trained model in step (6) above can be performed as follows:
[0099] The optimal network with the highest test set index obtained during training is selected. The infrared image to be detected is then input into the optimal network to obtain the detection result.
[0100] The infrared image to be detected is input into the trained network to obtain the detection result. This process can be represented as:
[0101] (37)
[0102] in, This represents the optimal network for small target detection in infrared images. and These represent the input infrared image and the output infrared small target detection result, respectively.
[0103] The present invention has the following beneficial technical effects:
[0104] This invention presents a method for small target detection in infrared images based on parallel optimization and redundancy suppression networks. The invention enhances the ability to distinguish weak targets from background information through multi-directional pooling attention parallel optimization; it reduces the computational cost of global modeling by obtaining semantic centers through clustering and compressing pixel tokens, while preserving local details; it improves upsampling distortion and feature aliasing in multi-scale fusion through a redundancy suppression feature fusion module; and it improves the regression accuracy and training stability of infrared small targets through a Gaussian regression consistency loss function. The overall scheme of this invention can be trained end-to-end and is suitable for real-time detection of infrared small targets in complex scenes. Experiments show that, compared with existing infrared small target detection methods, this invention can effectively cope with background clutter interference in different scenes, improve the representation of small target details, and significantly improve the detection accuracy of infrared small targets. Attached Figure Description
[0105] Figure 1 This is a flowchart of the method of the present invention.
[0106] Figure 2 This is a schematic diagram of the overall structure of the infrared small target detection network proposed in this invention (a schematic diagram of the overall structure of the infrared image small target detection method based on parallel optimization and redundancy suppression network proposed in this invention).
[0107] Figure 3 This is a schematic diagram of the parallel optimization module structure;
[0108] Figure 4 This is a schematic diagram of the clustering semantic attention module structure;
[0109] Figure 5 This is a schematic diagram of the redundancy suppression feature fusion structure;
[0110] Figure 6 This is a comparison chart of the detection results of the method of this invention with other currently advanced methods; Detailed Implementation
[0111] The present invention will now be described in detail with reference to the accompanying drawings and examples.
[0112] See the flowchart of the method of this invention. Figure 1 The specific implementation steps are as follows:
[0113] (1) Preprocessing the infrared image and inputting it into the established detection network: Use Cubic interpolation to uniformly adjust the infrared image and corresponding label to a size of 640×640, label the minimum bounding box and category number of the target, construct a parallel optimization and redundancy suppression network, and then input the image into the network.
[0114] (2) Constructing a parallel optimization module: A dual-branch attention mechanism is adopted, which models the global background distribution and local target details through different pooling operations. Average pooling captures smoother and more blurred background information, while max pooling is used to preserve sharp details in the image, thereby enhancing the target representation.
[0115] (3) Constructing a clustering semantic attention module: global semantic centers are adaptively constructed through iterative clustering to drive attention computation. Two linear layers are mapped to global keys and values, and pixel tokens are used as queries to compute dot product attention. This is then integrated with a local attention mechanism to preserve local details, thereby achieving efficient global modeling and feature enhancement.
[0116] (4) Construct a redundancy suppression feature fusion module: Convolution downsample the input features, use frequency domain closed solution and learnable regularization to achieve top-down feature fusion, then fuse multi-scale features according to bottom-up path, preprocess the original features through lightweight convolution and reconstruct them, crop and align the fused features, use linear subtraction to suppress redundancy, and output three scale feature maps for bounding box decoupling.
[0117] (5) Model training and optimization using Gaussian regression consistency loss function: First, calculate the Euclidean distance between the center of the predicted box and the center of the real box, model it as a two-dimensional Gaussian distribution, calculate the distribution distance by combining the covariance matrix, and obtain the Gaussian regression consistency loss by fusing the position and distribution distance. At the same time, use the distribution focal loss and CIoU to optimize the regression, and use the binary cross-entropy for the classification loss; the final loss is a multi-weighted combination.
[0118] (6) Input the image to be detected into the trained model to obtain the detection result: Select the optimal network with the best test set index, input the infrared image to be detected into the network, and output the final infrared small target detection result.
[0119] The above step (1) shall be performed as follows:
[0120] Specifically, we used Cubic interpolation to unify the infrared image and labels to 640×640. During label creation, the minimum bounding box and category number of the target were annotated. Then, a parallel optimization and redundancy suppression network was constructed, and the infrared image was input into it. The network backbone consists of convolutional layers, a parallel optimization module, and a multi-cluster semantic attention module. The convolutional layers are used for downsampling, the parallel optimization module for feature extraction, and the multi-cluster semantic attention module for feature enhancement of the network backbone's P3 layer. After extracting feature maps at different scales, a redundancy suppression feature fusion module was used to achieve multi-scale feature fusion. Finally, the fused feature map was input into the YOLO detection head to obtain the detection result.
[0121] Step (2) above shall be performed as follows:
[0122] We added two parallel attention processing branches to the bottleneck structure. Each attention branch processes the feature map of the image. The process is divided into two directions: height and width. Each branch generates two sets of attention maps in different directions.
[0123] Here, AvgPool represents average pooling, MaxPool represents maximum pooling, and the subscript indicates the pooling direction. and These represent the pooling results of the first branch along the h-axis and w-axis dimensions, respectively. and These represent the pooling results of the second branch along the h-axis and w-axis dimensions, respectively.
[0124] Subsequently, the pooling results in the width direction are subjected to dimension permutation, i.e. The permutation results are concatenated along the dimensions to obtain the intermediate feature Y.
[0125]
[0126] Concat represents the concatenation operation.
[0127] Subsequently, 1×1 convolutional kernels are used to reduce the dimensionality of the intermediate feature Y and fuse cross-channel features, followed by batch normalization layers and the ReLU activation function. The processed intermediate features are then split, and dimensional permutations are performed to match the pooling shape of the original width dimension. Attention weights in different directions are obtained based on the features from the two branches.
[0128] (7)
[0129] (8)
[0130] (9)
[0131] (10)
[0132] in, This represents the sigmoid function. This represents a 1×1 convolution operation. It is The high-dimensional features obtained after splitting, It is The dimension permutation result of the width dimension feature obtained after splitting. It is The high-dimensional features obtained after splitting, It is Dimension permutation result of the width dimension feature obtained after splitting
[0133] Subsequently, the attention weights in each branch are applied together to the original feature map, and residual links are added.
[0134] (11)
[0135] Convolutional layers with different receptive fields (e.g., 1×1, 3×3, and 5×5) are used to fully fuse the information from these features.
[0136] Step (3) above shall be performed as follows:
[0137] First, the feature map X input to the module is expanded into a token sequence and then normalized by the layer. The token sequence is retained as a residual term for residual stacking.
[0138] (12)
[0139] in, and These represent the mean and variance of the token sequence along the channel dimension, respectively. To prevent tiny values with a denominator of 0, and These are learnable scaling and offset parameters.
[0140] A large number of pixel tokens are then compressed through clustering operations to form k semantic tokens. Pixel tokens With cluster center The similarity between them can be expressed as:
[0141] (13)
[0142] in, The L2 norm is used. Then, the mean of all tokens within each cluster is calculated to generate new cluster centers. The iterative update formula is:
[0143] (14)
[0144] in, The center of the k-th cluster in the (t+1)-th round, For the set of pixel tokens assigned to the k-th cluster in the t-th round, This is an L2 normalization operation.
[0145] To accommodate the small sample size of the infrared small target dataset, an exponential moving average (EMA) strategy is used to update cluster centers, preventing oscillations caused by sample fluctuations. The EMA update formula is:
[0146] (15)
[0147] in, For the updated cluster centers, As a historical cluster center, The new cluster centers obtained in the current iteration, This is the EMA attenuation coefficient.
[0148] The cluster centers obtained after three iterations of the module are used as global semantic tokens to construct global attention. Specifically, the cluster centers are mapped to global keys k and global values v through two linear layers. The pixel token is then used as the query (q), and a dot product attention score is calculated with k, as follows:
[0149] (16)
[0150] Here, c represents the number of channels, and h represents the number of attention heads. The attention scores are then normalized using the Softmax function and weighted and summed with v to obtain the global attention enhancement features for each pixel token.
[0151] Subsequently, the pixel tokens are divided into N / G groups. If the groups are not divisible, a flip-fill strategy is used to fill them. Self-attention calculations are performed within each group to capture the spatial relationships between pixels within the group. Simultaneously, global semantic tokens obtained from clustering are used for attention calculations within each group to fuse local details and global semantics. The fused result is then added element-wise to the self-attention result to generate the final local attention-enhanced feature.
[0152] Finally, the global attention enhancement features and local attention enhancement features are added together and then converted into feature tensors through dimensionality rearrangement. Subsequently, 1×1 convolutional kernels are applied to adjust the channels, and a convolutional feedback network (ConvFFN) is added to enhance feature representation.
[0153] The above step (4) shall be performed as follows:
[0154] For input features The convolution and downsampling process can be represented as:
[0155] (17)
[0156] in, It outputs the feature map, and performs a downsampling operation. It is applied per channel.
[0157] In this case, the original input feature map X is recovered by minimizing the error between the observed output Y and the reconstructed output. Therefore, the optimization problem can be expressed as:
[0158] (18)
[0159] in, and These are the input and output feature maps of the c-th channel, respectively. (Frobenius norm) It is still an error metric calculated separately for each channel.
[0160] In the frequency domain, the recovered solution is obtained by performing a Fourier transform on each channel. Let ∀f{F} denote the Fourier transform. and These are the observation output and the convolution kernel's representations in the frequency domain, respectively. Therefore, the closed-form solution in the frequency domain can be expressed as:
[0161] (19)
[0162] in, It is the frequency domain information of the c-th channel. It is the Fourier transform of the c-th channel after upsampling. The complex conjugate of the Fourier transform of convolution kernel K. This represents the element-wise squared magnitude of the Fourier transform of the convolution kernel. This represents element-wise multiplication on a d×d block. This indicates that the average downsampling is performed on the d×d block. This indicates a standard d-fold upsampling operation, which increases the spatial dimension by inserting zeros.
[0163] Through the aforementioned closed-form solution in the frequency domain and learnable regularization, a mathematically invertible learnable frequency-constrained (LFC) upsampler was obtained, and this operator was used for top-down feature fusion. Then, multi-scale features were further fused following a bottom-up path. This operation fuses high-level features with low-level features, further optimizing the transmission of contextual information. Subsequently, background and redundant features were suppressed.
[0164] Let the current fusion feature be The original features at the same scale are First, perform light preprocessing on the original features:
[0165] (20)
[0166] Here, BN represents batch normalization operation.
[0167] The features are then reconstructed using the LFC upsampler U:
[0168] (twenty one)
[0169] The upsampling results are then cropped to restore them to the same spatial size as the fused features.
[0170] (twenty two)
[0171] Where C2f represents the feature extraction module in the YOLOv8 model, Crop represents nearest neighbor interpolation, and [H, W] represents... The size of the redundancy is ultimately achieved using linear subtraction.
[0172] (twenty three)
[0173] in, This represents the suppressed output. 'a' represents the suppression factor. The entire suppression process is hierarchically optimized and added to the feature fusion process. The entire redundant suppression feature fusion module ultimately outputs three feature maps at different scales for bounding box decoupling.
[0174] Step (5) above shall be performed as follows:
[0175] During training, the Gaussian regression consistency loss function is used to optimize the regression branch. First, the Euclidean distance between the predicted bounding box and the center of the ground truth bounding box is calculated:
[0176] (twenty four)
[0177] Among them, (x p , y p ), (x g , y g The coordinates of the predicted bounding box and the ground truth bounding box are respectively. The predicted and ground truth bounding boxes are then modeled as two-dimensional Gaussian distributions, and the distribution distance is calculated using the covariance matrix.
[0178] (25)
[0179] (26)
[0180] (27)
[0181] in, Represents the covariance matrix. Tr is the trace of the covariance matrix obtained by solving. det represents the determinant of the covariance matrix calculated using Cholesky decomposition.
[0182] The Gaussian regression consistency loss function is obtained by fusing location distance and distribution distance:
[0183] (28)
[0184] in, This represents the balance coefficient. To maintain numerical stability, the fused loss value is normalized and smoothed:
[0185] (29)
[0186] (30)
[0187] The regression loss is calculated using a combination of distributed focal loss and CIoU. The formula for calculating distributed focal loss is as follows:
[0188] (31)
[0189] in, and These are the predicted values output by the network, and the nearest neighbor predicted values. , , The values are the actual value of the label, the label integral value, and the integral value of the neighboring labels.
[0190] The formula for calculating CIoU is as follows:
[0191] (32)
[0192] (33)
[0193] (34)
[0194] in, and This represents the center point of the prediction box and the truth box. Represents Euclidean distance. and Represents the width of the prediction box and the ground truth box. and This represents the height of the prediction box and the truth box.
[0195] The classification loss is calculated using binary cross-entropy. The formula is as follows:
[0196] (35)
[0197] in, It's a binary tag, meaning either 0 or 1. It outputs the probability of belonging to the label. This represents the number of groups of objects the model predicts. The final loss function calculation formula is as follows:
[0198] (36)
[0199] in, , , , These are the weighting coefficients for each type of loss.
[0200] The constructed network can be trained using a pre-made dataset. The dataset is divided into training, validation, and test sets in a 7:1:2 ratio. The SGD optimizer is selected as the network optimizer. The initial learning rate is 0.01, which decays exponentially in each iteration. The model is trained for 200 generations, and gradient updates are performed during iterations using the aforementioned loss to guide the direction of network parameter updates.
[0201] The above step (6) shall be performed as follows:
[0202] The optimal network with the highest test set index obtained during training is selected. The infrared image to be detected is then input into the optimal network to obtain the detection result.
[0203] The infrared image to be detected is input into the trained network to obtain the detection result. This process can be represented as:
[0204] (37)
[0205] in, This represents the optimal network for small target detection in infrared images. and These represent the input infrared image and the output infrared small target detection result, respectively.
[0206] To quantitatively evaluate the performance of the proposed method, three quantitative evaluation metrics—precision, recall, and average precision (mAP)—were used to assess the infrared small target detection effect. Precision is the percentage of actually correct detections out of all detected positive examples. Recall refers to the percentage of correctly detected positive examples. Average precision (mAP) is the area enclosed by a graph plotted with precision on the y-axis and recall on the x-axis, typically calculated by integrating the area under the precision-recall (PR) curve. 50 This represents the average AP value for each target class when the IoU threshold is 50%. 50:95 Represents mAP 50 to mAP 95 The average value of all mAPs obtained after each 5% increase in the threshold.
[0207] Figure 6 This diagram presents a comparison of the detection results of the method proposed in this invention with other state-of-the-art methods. From left to right, it shows the input image containing the label, and the detection results of YOLOv10, YOLOv12, IR-DETR, IRSTD-YOLO, ABRNet, and the method proposed in this invention. Red circles represent missed detections, and yellow circles represent false positives. As shown, the method proposed in this invention achieves the best detection results in different scenarios, adapts to interference from various backgrounds, and significantly improves the robustness of detection.
[0208] Table 1. Quantitative comparison results of different methods in infrared small target detection on the test dataset.
[0209]
[0210] Table 1 lists the quantitative comparison results of the method of the present invention with other currently advanced detection methods, including accuracy, recall, and mean precision (mAP). 50 and average accuracy mAP 50:95 As the quantitative comparison results show, the method proposed in this invention achieves optimal accuracy and can significantly improve the detection accuracy of small infrared targets.
Claims
1. A method for detecting small targets in infrared images based on parallel optimization and redundancy suppression networks, characterized in that, The method utilizes parallel optimization of the background and target to prevent small target features from being masked by complex background clutter, and employs redundancy suppression to address the feature mixing problem caused by multi-scale fusion, thereby improving the accuracy of infrared small target detection. The method includes the following steps: (1) Preprocessing the infrared image and inputting it into the established detection network: Use Cubic interpolation to uniformly adjust the infrared image and corresponding label to a size of 640×640, label the minimum bounding box and category number of the target, construct a parallel optimization and redundancy suppression network, and then input the image into the network. (2) Constructing a parallel optimization module: A dual-branch attention mechanism is adopted, which models the global background distribution and local target details through different pooling operations. Average pooling captures smoother and more blurred background information, while max pooling is used to preserve sharp details in the image, thereby enhancing the target representation. (3) Constructing a clustering semantic attention module: global semantic centers are adaptively constructed through iterative clustering to drive attention computation. Two linear layers are mapped to global keys and values, and pixel tokens are used as queries to compute dot product attention. This is then integrated with a local attention mechanism to preserve local details, thereby achieving efficient global modeling and feature enhancement. (4) Construct a redundancy suppression feature fusion module: Convolution downsample the input features, use frequency domain closed solution and learnable regularization to achieve top-down feature fusion, then fuse multi-scale features according to bottom-up path, preprocess the original features through lightweight convolution and reconstruct them, crop and align the fused features, use linear subtraction to suppress redundancy, and output three scale feature maps for bounding box decoupling. (5) Model training and optimization using Gaussian regression consistency loss function: First, calculate the Euclidean distance between the center of the predicted box and the center of the real box, model it as a two-dimensional Gaussian distribution, calculate the distribution distance by combining the covariance matrix, and obtain the Gaussian regression consistency loss by fusing the position and distribution distance. At the same time, use the distribution focal loss and CIoU to optimize the regression, and use the binary cross-entropy for the classification loss; the final loss is a multi-weighted combination. (6) Input the image to be detected into the trained model to obtain the detection result: Select the optimal network with the best test set index, input the infrared image to be detected into the network, and output the final infrared small target detection result.
2. The method as described in claim 1, characterized in that, The acquisition of infrared images and input into the established detection network in step (1) specifically involves: using Cubic interpolation to unify the infrared images and labels to 640×640; labeling the minimum bounding box and category number of the target; then constructing a parallel optimization and redundancy suppression network and inputting the infrared images into it; the network backbone consists of convolutional layers, a parallel optimization module, and a multi-cluster semantic attention module; the convolutional layers are used for downsampling, the parallel optimization module is used for feature extraction, and the multi-cluster semantic attention module is used to enhance the features of the P3 layer of the network backbone; after extracting feature maps of different scales, the redundancy suppression feature fusion module is used to achieve multi-scale feature fusion; finally, the fused feature maps are input into the YOLO detection head to obtain the detection results.
3. The method as described in claim 1 or 2, characterized in that, Step (2) is performed as follows: Two parallel attention processing branches were added to the bottleneck structure, each of which processed the feature map of the image. The process is divided into two directions: height and width. Each branch generates two sets of attention maps in different directions. Where AvgPool represents average pooling, MaxPool represents maximum pooling, and the subscript represents the pooling direction; and These represent the pooling results along the h-axis and w-axis dimensions of the first branch, respectively; and These represent the pooling results along the h-axis and w-axis dimensions of the second branch, respectively; Subsequently, the pooling results in the width direction are subjected to dimension permutation, i.e. The intermediate feature Y is obtained by concatenating the permutation results through dimensional concatenation. Here, Concat represents the concatenation operation; Subsequently, 1×1 convolutional kernels are used to reduce the dimensionality of the intermediate feature Y and fuse cross-channel features, followed by batch normalization layers and ReLU activation function. The processed intermediate features are then split, and dimensional permutation is performed to match the pooling shape of the original width dimension. Attention weights in different directions are obtained based on the features of the two branches. (7) (8) (9) (10) in, This represents the sigmoid function. This represents a 1×1 convolution operation. It is The high-dimensional features obtained after splitting, It is The dimension permutation result of the width dimension feature obtained after splitting. It is The high-dimensional features obtained after splitting, It is The dimension permutation result of the width dimension feature obtained after splitting; Subsequently, the attention weights in each branch are applied together to the original feature map, and residual links are added; (11) Convolutional layers with different receptive fields (e.g., 1×1, 3×3, and 5×5) are used to fully fuse the information from these features.
4. The method as described in claim 3, characterized in that, Step (3) is performed as follows: First, the feature map X input to the module is expanded into a token sequence and then normalized by the layer. The token sequence is retained as a residual term for residual stacking. (12) in, and These represent the mean and variance of the token sequence along the channel dimension, respectively. To prevent tiny values with a denominator of 0, and These are learnable scaling and offset parameters; Subsequently, a large number of pixel tokens are compressed through clustering operations to form k semantic tokens; pixel tokens With cluster center The similarity between them can be expressed as: (13) in, Given the L2 norm, we calculate the mean of all tokens within each cluster to generate new cluster centers. The iterative update formula is: (14) in, The center of the k-th cluster in the (t+1)-th round, For the set of pixel tokens assigned to the k-th cluster in the t-th round, This is an L2 normalization operation; To accommodate the small sample size of the infrared small target dataset, an exponential moving average (EMA) strategy is used to update the cluster centers, preventing them from oscillating due to sample fluctuations. The EMA update formula is as follows: (15) in, For the updated cluster centers, As a historical cluster center, The new cluster centers obtained in the current iteration, The EMA attenuation coefficient; The cluster centers obtained after three iterations of the module are used as global semantic tokens to construct global attention. Two linear layers are used to map the cluster centers to global keys k and global values v, respectively. The pixel token is used as the query (q), and a dot product attention score is calculated with k, as shown in the formula: (16) Where c represents the number of channels and h represents the number of attention heads, the attention score is then normalized using the Softmax function and weighted and summed with v to obtain the global attention enhancement feature for each pixel token; Subsequently, the pixel tokens are divided into N / G groups. If they are not divisible, a flip-fill strategy is used to fill them. Self-attention calculations are performed in each group to capture the spatial relationships between pixels within the group. At the same time, attention calculations are performed in each group using the global semantic tokens obtained from clustering to achieve the fusion of local details and global semantics. The fused result is added element-wise to the self-attention result to generate the final local attention-enhanced feature. Finally, the global attention enhancement features and local attention enhancement features are added together and then converted into feature tensors through dimensionality rearrangement. Subsequently, 1×1 convolutional kernels are applied to adjust the channels, and a convolutional feedback network (ConvFFN) is added to enhance the feature representation.
5. The method as described in claim 4, characterized in that, Step (4) is performed as follows: For input features The convolution and downsampling process can be represented as: (17) in, It outputs the feature map, and performs a downsampling operation. It is applied per channel; In this case, the original input feature map X is recovered by minimizing the error between the observed output Y and the reconstructed output. The optimization problem is expressed as: (18) in, and These are the input and output feature maps of the c-th channel, respectively, and the Frobenius norm. It is still an error metric calculated separately for each channel; The recovered solution is obtained by performing a Fourier transform on each channel in the frequency domain; let ∀f{F} denote the Fourier transform. and Let these be the observed output and the convolution kernel, respectively, represented in the frequency domain; then, the closed-form solution in the frequency domain can be expressed as: (19) in, It is the frequency domain information of the c-th channel. It is the Fourier transform of the c-th channel after upsampling. The complex conjugate of the Fourier transform of convolution kernel K. This represents the element-wise squared magnitude of the Fourier transform of the convolution kernel. This represents element-wise multiplication on a d×d block. This indicates that the average downsampling is performed on the d×d block. This indicates a standard d-fold upsampling operation, which increases the spatial dimension by inserting zeros; Through the above frequency domain closed-form solution and learnable regularization, a mathematically invertible learnable frequency domain constrained (LFC) upsampler is obtained, and this operator is used for top-down feature fusion. Then, multi-scale features are further fused according to the bottom-up path. This operation uses low-level features to fuse high-level features, which further optimizes the transmission of context information. Subsequently, background and redundant features are suppressed. Let the current fusion feature be The original features at the same scale are First, perform light preprocessing on the original features: (20) Where BN represents batch normalization operation; The features are then reconstructed using the LFC upsampler U: (21) The upsampling results are then cropped to restore them to the same spatial size as the fused features. (22) Where C2f represents the feature extraction module in the YOLOv8 model, Crop represents nearest neighbor interpolation, and [H, W] represents... The size; ultimately, linear subtraction is used to suppress redundancy: (23) in, 'a' represents the suppressed output, and 'a' represents the suppression factor. The entire suppression is added to the feature fusion process and is optimized hierarchically. The entire redundant suppression feature fusion module finally outputs three feature maps of different scales for bounding box decoupling.
6. The method as described in claim 1 or 5, characterized in that, Step (5) is performed as follows: During training, the Gaussian regression consistency loss function is used to optimize the regression branch. First, the Euclidean distance between the predicted box and the center of the ground truth box is calculated: (24) Among them, (x p , y p ), (x g , y g The coordinates of the predicted bounding box and the ground truth bounding box are given, respectively. The predicted and ground truth bounding boxes are then modeled as two-dimensional Gaussian distributions, and the distribution distance is calculated using the covariance matrix. (25) (26) (27) in, Tr represents the covariance matrix, det represents the trace of the covariance matrix obtained by solving the covariance matrix, and det represents the determinant of the covariance matrix calculated using Cholesky decomposition. The Gaussian regression consistency loss function is obtained by fusing location distance and distribution distance: (28) in, The balancing coefficient is used to normalize and smooth the fused loss value to maintain numerical stability. (29) (30) The regression loss is calculated using a combination of distributed focal loss and CIoU. The formula for calculating distributed focal loss is as follows: (31) in, and The network outputs predicted values, and the nearest neighbor predicted values. , , The actual value of the label, the label integral value, and the integral value of the neighboring labels; The formula for calculating CIoU is as follows: (32) (33) (34) in, and This represents the center point of the prediction box and the truth box; Represents European distance. and Represents the width of the predicted bounding box and the ground truth bounding box. and This represents the height of the prediction box and the truth box; The classification loss uses binary cross-entropy, and the calculation formula is as follows: (35) in, It is a binary tag, that is, 0 or 1. It outputs the probability of belonging to the label. This represents the number of groups of objects predicted by the model. Finally, the formula for calculating the loss function is as follows: (36) in, , , , These are the weighting coefficients for each type of loss; The constructed network can be trained using a pre-made dataset, which is divided into training, validation, and test sets in a 7:1:2 ratio. The network optimizer is SGD, with an initial learning rate of 0.01 that decays exponentially in each iteration. The model is trained for 200 generations, and the aforementioned loss guides the direction of network parameter updates during the iteration process, resulting in gradient updates.
7. The method as described in claim 1, characterized in that, Step (6) is performed as follows: The optimal network with the highest test set index obtained during training is selected, and the infrared image to be detected is input into the optimal network to obtain the detection result. The infrared image to be detected is input into the trained network to obtain the detection result. This process can be represented as: (37) in, This represents the optimal network for small target detection in infrared images. and These represent the input infrared image and the output infrared small target detection result, respectively.
8. A small target detection system for infrared images based on parallel optimization and redundancy suppression networks, characterized in that: The system has a program module corresponding to the steps of any one of the claims 1-7 above, and executes the steps in the infrared image small target detection method based on parallel optimization and redundancy suppression network described above when running.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program configured to, when invoked by a processor, implement the steps of the infrared image small target detection method based on a parallel optimization and redundancy suppression network as described in any one of claims 1-7.