Infrared small object target detection method based on adaptive frequency domain separation enhancement network
Through adaptive frequency domain separation enhancement network, combined with frequency domain feature separation and gradient positioning enhancement, the problem of feature information loss in infrared small object target detection is solved, achieving higher detection accuracy and accuracy.
Patent Information
- Application Number
- CN202510700399.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-02
AI Technical Summary
The existing infrared object detection method ignores the possibility of missing targets in network depth design, and the scarce feature information of infrared targets during the downsampling of convolutional neural networks is compressed or lost, resulting in low detection accuracy.
Adaptive frequency domain separation enhancement network is adopted, including resnet residual network, frequency domain adaptive separation module, encoder, gradient positioning enhancement module and decoder, and through frequency domain feature separation and gradient positioning enhancement, feature representation and object detection accuracy is improved.
It effectively improves the feature representation and detection accuracy of small infrared targets, solves the problem of compressing scarce feature information, and improves the accuracy and precision of the model positioning of small infrared targets.
Smart Images

Figure CN120580412A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an infrared small object target detection method, and in particular to an infrared small object target detection method based on an adaptive frequency domain separation enhancement network. Background Art
[0002] Infrared images maintain clarity in both low-light and high-light conditions, offering significant advantages in complex environments. Compared to traditional visible light images, infrared images can effectively penetrate smoke, haze, and other meteorological obstructions, providing more stable and reliable visual information. This characteristic makes infrared images particularly suitable for use at night or in adverse weather conditions. Therefore, infrared small target detection (IRSTD) technology is crucial in a variety of applications. It aims to extract and identify low-contrast, small targets using infrared images. The unique advantage of infrared small target detection lies in its ability to effectively capture low-contrast and small targets in complex environments. By improving image processing algorithms, the system's target capture capabilities in low-contrast environments can be enhanced, reducing false alarms and missed detections, ensuring efficient and accurate target detection in a variety of practical application scenarios. This technology is widely used in fields such as nighttime surveillance, aerospace imaging, and security patrols, enabling rapid identification of potential threats and providing real-time feedback.
[0003] In recent years, research on IRSTD has shifted from traditional methods to deep learning approaches. Traditional methods include filtering, local contrast, and low-rank decomposition. These methods reduce background interference and extract small target features through morphological operations, grayscale difference analysis, and low-rank background separation, respectively. Leveraging the superior feature extraction capabilities of convolutional neural networks, deep learning methods have demonstrated superior performance when handling diverse targets and complex scenes. Deep learning methods can be broadly categorized as detector-based, image segmentation-based, and generative adversarial network (GAN)-based. Detector-based methods typically leverage the deep feature extraction capabilities of convolutional neural networks, combined with region proposals and object regression strategies, to directly predict the target's location and category in an end-to-end manner. Image segmentation-based methods, on the other hand, transform the infrared small target detection task into a pixel-level region segmentation problem. These methods typically utilize a feature reconstruction strategy that gradually recovers spatial information from the image, ensuring accurate pixel-level localization of small targets and effectively improving the detail representation of target detection. Generative adversarial network-based methods employ an adversarial learning mechanism between a generator and a discriminator to improve model performance by enhancing the features of the target region or synthesizing high-quality training data.
[0004] Most current IRSTD methods suffer from two weaknesses: First, they focus solely on network depth design, ignoring the possibility of missing targets during actual detection. Second, small infrared targets typically occupy only a small number of pixels in an image. During the downsampling process of convolutional neural networks, scarce feature information is further compressed or even lost. Summary of the Invention
[0005] The purpose of the present invention is to solve the technical problems that the current infrared small object target detection method only focuses on the design of network depth, ignores the possibility of losing the target in actual detection, and because small infrared targets usually only occupy a very small number of pixels in the image, resulting in the scarce feature information being further compressed or even lost during the downsampling process of the convolutional neural network, and provides an infrared small object target detection method based on an adaptive frequency domain separation enhancement network.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for infrared small object target detection based on an adaptive frequency domain separation enhancement network is characterized in that it includes the following steps:
[0008] Step 1: extracting original image data containing infrared small objects from the database;
[0009] Step 2: Construct an adaptive frequency domain separation enhancement network, which includes a resnet residual network, a frequency domain adaptive separation module, an encoder, a gradient positioning enhancement module, and a decoder;
[0010] The resnet residual network is used to extract the spatial domain features of the image through convolution operations;
[0011] The frequency domain adaptive separation module is used to perform fast Fourier transform on the image, convert the image from the spatial domain to the frequency domain, generate frequency domain features, and adaptively separate the frequency domain features to obtain multiple frequency domain features;
[0012] The encoder is used to fuse the spatial domain features and multi-band frequency domain features of the image to obtain coding features;
[0013] The gradient positioning enhancement module calculates the gradient direction and gradient magnitude of the input image through the Sobel operator, and then calculates the saliency map, and integrates the saliency map into the decoder;
[0014] The decoder is used to decode the encoded features to obtain a prediction result;
[0015] Step 3: The adaptive frequency domain separation enhancement network is trained using the original image data to obtain an optimized adaptive frequency domain separation enhancement network;
[0016] Step 4: Input the image to be detected into the optimized adaptive frequency domain separation enhancement network to obtain the infrared small object target detection result.
[0017] In step 1, the original image data containing the infrared small object includes the original image containing the infrared small object and the true label of the infrared small object in the original image.
[0018] Step 3 is as follows:
[0019] Step 3.1, initialize the parameters of the adaptive frequency domain separation enhancement network, set the iterative convergence conditions, and randomly select multiple original images containing infrared small objects from the original data image I input And the original image I input The true labels of small mid-infrared objects are divided into multiple batches of raw data images;
[0020] Step 3.2: The original images I containing infrared small objects in the same batch input Input to the resnet residual network, and extract the shallow spatial domain features F of the image through convolution operation shallow , mid-level spatial domain features F middle , deep spatial domain features F deep ;
[0021] At the same time, the original image containing the infrared small object I input Input to the frequency domain adaptive separation module, the original image I input Convert from spatial domain to frequency domain to generate frequency domain features, and perform adaptive separation on the frequency domain features to obtain low frequency domain features F' L , mid-frequency domain feature F' B , high-frequency domain features F' H ;
[0022] Step 3.3: Encode the shallow spatial domain features F shallow , mid-level spatial domain features F middle , deep spatial domain features F deep and low-frequency domain features F' L , mid-frequency domain feature F' B , high-frequency domain features F' H Perform splicing and fusion to obtain coding features;
[0023] Step 3.4: Convert the original image containing the infrared small object I input Input the gradient positioning enhancement module to calculate the gradient direction and gradient amplitude of the input image, and then obtain the saliency map S grad , and the saliency map S grad Integrate into decoder;
[0024] Step 3.5: Incorporate the encoded feature input into the saliency map S grad The decoder decodes the image to obtain the predicted label of the infrared small object in the original image;
[0025] Step 3.6, combine the predicted labels of the infrared small objects in the original image with the true labels of the infrared small objects in the original image, calculate the loss value, use the back propagation algorithm to calculate the gradient of the adaptive frequency domain separation enhancement network parameters, and use the Adam optimizer to update the adaptive frequency domain separation enhancement network parameters, return to step 3.2, and use another batch of original images containing infrared small objects I input The input is fed into the resnet residual network until the iterative convergence condition is reached, and the adaptive frequency domain separation enhancement network parameters corresponding to the minimum loss value are selected to form an optimized adaptive frequency domain separation enhancement network.
[0026] In step 3.2, the shallow spatial domain feature F of the image is extracted by convolution operation shallow , mid-level spatial domain features F middle , deep spatial domain features F deep The formulas are:
[0027] F shallow =Conv1(I input )
[0028] F middle =Conv2(MaxPool(F shallow ))
[0029] F deep =ReLU(Conv3(F middle ));
[0030] Among them, Conv1, Conv2, and Conv3 represent convolution operations, MaxPool represents the maximum pooling operation, and ReLU represents the activation function.
[0031] In step 3.2, the original image I containing the infrared small object is input Input to the frequency domain adaptive separation module, the original image I input Converting from the spatial domain to the frequency domain, the formula for generating frequency domain features is:
[0032]
[0033] Among them, F(u,v) represents the frequency domain characteristics of the image, f(x,y) represents the spatial domain characteristics, that is, the pixel value of the original image, (u,v) represents the frequency domain obtained after the image is subjected to a two-dimensional Fourier transform, u and v represent the horizontal and vertical component positions of each frequency in the image, respectively, M and N are the number of rows and columns of the image, respectively, x represents the horizontal pixel position, and y represents the vertical pixel position.
[0034] In step 3.2, the adaptive separation of frequency domain features is specifically as follows:
[0035] Step a: Calculate the amplitude spectrum |F(u,v)| based on the frequency domain characteristics. The calculation formula is:
[0036]
[0037] Step b: Determine the cutoff frequency D of the low-pass filter and high-pass filter based on the amplitude spectrum |F(u,v)| L and D H , respectively use high-pass filter, low-pass filter and band-pass filter to separate the frequency domain features F(u,v) to obtain the high-frequency domain features F' H , low-frequency domain features F' L And the mid-frequency domain feature F' B , the separation formulas are:
[0038]
[0039] F' B =F' L -F' H ,
[0040] Where D L and D H Represent the cutoff frequencies of the low-pass filter and high-pass filter respectively; D(u,v) represents the frequency point (u, v) to the original image I input The distance to the center of mass.
[0041] Step 3.4 is as follows:
[0042] Step 3.4.1, the original image containing the infrared small object I input Input the gradient positioning enhancement module, and convolve the image with the horizontal Sobel convolution kernel to obtain the horizontal gradient value G of each pixel in the image. x By convolving the image with the Sobel convolution kernel in the vertical direction, the gradient value G of each pixel in the image in the vertical direction is obtained. y , calculate the input original image I input The gradient direction θ and gradient amplitude G are calculated as follows:
[0043]
[0044] Step 3.4.2: Input the original image I according to the representation input The gradient direction θ and gradient magnitude G are used to calculate the variance C of the gradient direction in the local area of each pixel. θ and the variance of the gradient magnitude C G , the calculation formula is:
[0045]
[0046] Where n represents the number of pixels in the local area, θ i represents the gradient direction of the i-th pixel in the local area, Represents the mean value of the gradient direction of all pixels in the local area; G i represents the gradient magnitude of the i-th pixel in the local area, Represents the mean value of the gradient magnitude of all pixels in the local area, Among them G x,i and G y,i are the horizontal and vertical gradients of the i-th pixel;
[0047] Step 3.4.3. Calculate the saliency map S grad , the formula is:
[0048] S grad =exp(-α·C θ -β·C G ),
[0049] Where α and β are weight coefficients used to adjust the consistency of gradient direction and amplitude;
[0050] The saliency map S grad Integrate into the decoder.
[0051] Step 3.5 is as follows:
[0052] The encoding feature input is integrated into the saliency map S grad The decoder is decoded, and the decoding formula is:
[0053]
[0054] Where I represents the predicted label of the infrared small object in the original image, Represents the splicing and fusion operation, represents the upsampling Up operation, ⊙ represents the supplementary operation, and F represents the spatial domain feature, which is the shallow spatial domain feature F shallow , mid-level spatial domain features F middle , deep spatial domain features F deep; f represents the frequency domain feature, that is, the low frequency domain feature F' L , mid-frequency domain feature F' B , high-frequency domain features F' H .
[0055] In step 3.6, the loss value is calculated specifically as follows: the binary cross entropy loss value, intersection-over-union loss value, and Dice loss value are calculated respectively based on the predicted label of the infrared small object in the original image and the true label of the infrared small object in the original image, and the binary cross entropy loss value, intersection-over-union loss value, and Dice loss value are weightedly summed to obtain the loss value.
[0056] The iterative convergence condition is:
[0057] After running a round, calculate the IoU of the current round and compare it with the historical best IoU. If the IoU of the current round is greater than the historical best IoU, update the IoU of the current round to the historical best IoU; otherwise, keep the historical best IoU.
[0058] If the intersection-over-union ratio for 10 consecutive rounds is not greater than the historical best intersection-over-union ratio, the iteration is terminated.
[0059] Beneficial effects of the present invention:
[0060] (1) The present invention provides an infrared small object target detection method based on an adaptive frequency domain separation enhancement network. An adaptive frequency domain separation enhancement network is designed. On the one hand, by applying adaptive frequency domain separation technology, the frequency domain features are accurately decomposed into high-frequency, medium-frequency, and low-frequency components. Combined with the shallow, medium, and deep spatial domain features, the overall feature representation of the infrared small target is effectively improved. On the other hand, the gradient characteristics of the infrared small object in the image are used to define the direction consistency function and the amplitude consistency function. This function effectively identifies the edge features of the target and significantly improves the accuracy of the model in infrared small target positioning.
[0061] (2) The present invention provides an infrared small object target detection method based on an adaptive frequency domain separation enhancement network. Based on the encoder and decoder, a frequency domain adaptive separation network is proposed to solve the problem of further compression or even loss of scarce feature information during infrared small object target detection. By extracting and separating frequency domain features, the saliency of low-contrast and small targets is enhanced, thereby better retaining key feature information. In addition, through gradient positioning, the module can effectively identify the target center, improving detection accuracy and the target's fine positioning capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1This is a flow chart of training the adaptive frequency domain separation enhancement network using original image data in an embodiment of an infrared small object target detection method based on the adaptive frequency domain separation enhancement network provided by the present invention;
[0063] Figure 2 This is a diagram showing the finalization results of multiple data sets in an embodiment of an infrared small object target detection method based on an adaptive frequency domain separation enhancement network provided by the present invention. DETAILED DESCRIPTION
[0064] The following will clearly and completely describe the technical solution of this embodiment in conjunction with the accompanying drawings and embodiments. Obviously, the embodiments described are only part of the embodiments of this embodiment, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of this embodiment without creative work are within the scope of protection of this embodiment.
[0065] This embodiment provides an infrared small object target detection method based on an adaptive frequency domain separation enhancement network, comprising the following steps:
[0066] Step 1: extracting original image data containing infrared small objects from the database;
[0067] The original image data containing the small infrared object includes the original image containing the small infrared object and the true label of the small infrared object in the original image.
[0068] Step 2: Construct an adaptive frequency domain separation enhancement network, which includes a resnet residual network, a frequency domain adaptive separation module, an encoder, a gradient positioning enhancement module, and a decoder.
[0069] Resnet residual network is used to extract the spatial domain features of the image through convolution operation;
[0070] The frequency domain adaptive separation module is used to perform fast Fourier transform on the image, convert the image from the spatial domain to the frequency domain, generate frequency domain features, and adaptively separate the frequency domain features to obtain multi-segment frequency domain features;
[0071] The encoder is used to fuse the spatial domain features and multi-band frequency domain features of the image to obtain the coding features;
[0072] The gradient localization enhancement module calculates the gradient direction and gradient magnitude of the input image through the Sobel operator, and then calculates the saliency map, and integrates the saliency map into the decoder;
[0073] The decoder is used to decode the encoded features to obtain the prediction results;
[0074] Step 3: The adaptive frequency domain separation enhancement network is trained using the original image data to obtain an optimized adaptive frequency domain separation enhancement network; Figure 1 As shown, specifically:
[0075] Step 3.1, initialize the parameters of the adaptive frequency domain separation enhancement network, set the iterative convergence conditions, and randomly select multiple original images containing infrared small objects from the original data image I input And the original image I input The true labels of small mid-infrared objects are divided into multiple batches of raw data images;
[0076] Step 3.2: The original images I containing infrared small objects in the same batch input Input to the resnet residual network, and extract the shallow spatial domain features F of the image through convolution operation shallow , mid-level spatial domain features F middle , deep spatial domain features F deep ; The formulas are:
[0077] F shallow =Conv1(I input )
[0078] F middle =Conv2(MaxPool(F shallow ))
[0079] F deep =ReLU(Conv3(F middle ));
[0080] Among them, Conv1, Conv2, and Conv3 represent convolution operations, MaxPool represents the maximum pooling operation, and ReLU represents the activation function.
[0081] At the same time, the original image containing the infrared small object I input Input to the frequency domain adaptive separation module, the original image I input Converting from the spatial domain to the frequency domain generates frequency domain features. The formula for generating frequency domain features is:
[0082]
[0083] Among them, F(u,v) represents the frequency domain characteristics of the image, f(x,y) represents the spatial domain characteristics, that is, the pixel value of the original image, (u,v) represents the frequency domain obtained after the image is subjected to a two-dimensional Fourier transform, u and v represent the horizontal and vertical component positions of each frequency in the image, respectively, M and N are the number of rows and columns of the image, respectively, x represents the horizontal pixel position, and y represents the vertical pixel position.
[0084] Adaptively separate the frequency domain features, specifically:
[0085] Step a: Calculate the amplitude spectrum |F(u,v)| based on the frequency domain characteristics. The calculation formula is:
[0086]
[0087] Step b: Determine the cutoff frequency D of the low-pass filter and high-pass filter based on the amplitude spectrum |F(u,v)| L and D H , respectively use high-pass filter, low-pass filter and band-pass filter to separate the frequency domain features F(u,v) to obtain the high-frequency domain features F' H , low-frequency domain features F' L And the mid-frequency domain feature F' B , the separation formulas are:
[0088]
[0089] F' B =F' L -F' H ,
[0090] Where D L and D H Represent the cutoff frequencies of the low-pass filter and high-pass filter respectively; D(u,v) represents the frequency point (u, v) to the original image I input The distance to the center of mass.
[0091] In actual use, clustering algorithms can also be used to classify the amplitude spectrum. Statistical characteristics such as mean, standard deviation, and peak value can be combined to calculate the initial centroid of the clustering algorithm, which helps to quickly converge the clustering algorithm. High-pass filters, low-pass filters, and band-pass filters can be designed to extract low-frequency domain features, mid-frequency domain features, and high-frequency domain features.
[0092] In this embodiment, low-frequency features are used to capture background information, mid-frequency features contain the object outline and its transition area, and high-frequency features emphasize edge details. This separation can more effectively preserve the details in the image.
[0093] Step 3.3: Encode the shallow spatial domain features F shallow , mid-level spatial domain features F middle , deep spatial domain features F deep and low-frequency domain features F' L , mid-frequency domain feature F' B , high-frequency domain features F' H Perform splicing and fusion to obtain coding features;
[0094] Step 3.4: Convert the original image containing the infrared small object I input Input the gradient positioning enhancement module to calculate the gradient direction and gradient amplitude of the input image, and then obtain the saliency map S grad , and the saliency map S grad Integrate into the decoder; specifically:
[0095] Step 3.4.1, the original image containing the infrared small object I input Input the gradient positioning enhancement module, and convolve the image with the horizontal Sobel convolution kernel to obtain the horizontal gradient value G of each pixel in the image. x By convolving the image with the Sobel convolution kernel in the vertical direction, the gradient value G of each pixel in the image in the vertical direction is obtained. y , calculate the input original image I input The gradient direction θ and gradient amplitude G are calculated as follows:
[0096]
[0097] Step 3.4.2: Input the original image I according to the representation input The gradient direction θ and gradient magnitude G are used to calculate the variance C of the gradient direction in the local area of each pixel. θ and the variance of the gradient magnitude C G , to measure the consistency of the gradient direction. A smaller variance indicates that the gradient direction consistency of the region is higher. The calculation formula is:
[0098]
[0099] Where n represents the number of pixels in the local area, θ i represents the gradient direction of the i-th pixel in the local area, Represents the mean value of the gradient direction of all pixels in the local area; G i represents the gradient magnitude of the i-th pixel in the local area, Represents the mean value of the gradient magnitude of all pixels in the local area, Among them G x,i and G y,i are the horizontal and vertical gradients of the i-th pixel;
[0100] Step 3.4.3: Based on the consistency function, construct the computational significance map S using the exponential function. grad , the formula is:
[0101] S grad =exp(-α·C θ -β·C G ),
[0102] Where α and β are weight coefficients used to adjust the consistency of gradient direction and amplitude;
[0103] The saliency map S grad Integrate into the decoder.
[0104] Step 3.5: Incorporate the encoded feature input into the saliency map S grad The decoder is decoded, and the decoding formula is:
[0105]
[0106] Where I represents the predicted label of the infrared small object in the original image, Represents the splicing and fusion operation, represents the upsampling Up operation, ⊙ represents the supplementary operation, and F represents the spatial domain feature, which is the shallow spatial domain feature F shallow , mid-level spatial domain features F middle , deep spatial domain features F deep ; f represents the frequency domain feature, that is, the low frequency domain feature F' L , mid-frequency domain feature F' B , high-frequency domain features F' H .
[0107] Step 3.6, combine the predicted labels of the infrared small objects in the original image with the true labels of the infrared small objects in the original image, calculate the loss value, use the back propagation algorithm to calculate the gradient of the adaptive frequency domain separation enhancement network parameters, and use the Adam optimizer to update the adaptive frequency domain separation enhancement network parameters, return to step 3.2, and use another batch of original images containing infrared small objects I input The input is fed into the resnet residual network until the iterative convergence condition is reached, and the adaptive frequency domain separation enhancement network parameters corresponding to the minimum loss value are selected to form an optimized adaptive frequency domain separation enhancement network.
[0108] The specific calculation of the loss value is as follows: the binary cross entropy loss value, intersection-over-union loss value, and Dice loss value are calculated respectively based on the predicted label of the infrared small object in the original image and the true label of the infrared small object in the original image, and the binary cross entropy loss value, intersection-over-union loss value, and Dice loss value are weighted summed to obtain the loss value.
[0109] The iterative convergence condition is:
[0110] After running a round, calculate the IoU of the current round and compare it with the historical best IoU. If the IoU of the current round is greater than the historical best IoU, update the IoU of the current round to the historical best IoU; otherwise, keep the historical best IoU.
[0111] If the intersection-over-union ratio for 10 consecutive rounds is not greater than the historical best intersection-over-union ratio, the iteration is terminated.
[0112] Step 4: Input the image to be detected into the optimized adaptive frequency domain separation enhancement network to obtain the infrared small object target detection result.
[0113] In this embodiment, a method for infrared small object target detection based on an adaptive frequency domain separation enhancement network utilizes a frequency domain adaptive decomposition module and a gradient positioning enhancement module to solve the problems of further compression or even loss of scarce feature information and insufficient exploration of frequency domain information during the infrared small object target detection process, while also having certain advantages in speed.
[0114] This embodiment uses an infrared small object target detection method based on an adaptive frequency domain separation enhancement network to conduct experimental verification:
[0115] (1) Experimental conditions:
[0116] In this example, we used PyTorch, an open source deep learning framework developed by Facebook's artificial intelligence research team, to build the experimental environment. We selected Python as the development language, used the compiler and related packages in Anaconda3, selected Pycharm as the integrated development environment, and used a server equipped with a 1080Ti graphics card for model calculations.
[0117] ①Comparison method:
[0118] In order to demonstrate the superiority of the method of this embodiment, the algorithms TopHat, PSTNN, ALCNet, ACM, ISNet, AGPCNet, EGPNet, and DNANet are selected as comparison algorithms, as follows:
[0119] TopHat: This network introduces a modified top-hat transform and uses the calculated judgment value to effectively suppress clutter and significantly enhance the visibility of small targets with low contrast.
[0120] PSTNN: By introducing low-rank constraints, improved local prior maps and efficient tensor decomposition methods, it effectively suppresses background and enhances target detection capabilities while significantly improving computational efficiency.
[0121] ALCNet: Utilizes labeled data and domain knowledge. Designs a feature map cyclic shift scheme and a parameter-free nonlinear feature optimization layer, combined with bottom-up attention modulation to retain and enhance small target features.
[0122] ACM: A novel model-driven deep network is proposed, which significantly improves the performance of infrared small target detection by combining discriminative networks with traditional model-driven methods as well as labeled data and domain knowledge.
[0123] ISNet: By designing the Taylor finite difference edge module and the bidirectional attention aggregation module, the contrast between the target and the background is effectively enhanced and the shape features of small infrared targets are extracted.
[0124] AGPCNet: By exploring contextual information, an attention-guided context module is designed, and pixel associations at specific scales are perceived through local semantic associations and global contextual attention, thereby improving feature representation capabilities.
[0125] EGPNet: It extracts features through a multi-scale feature progressive fusion encoder to enhance semantic information and contextual association. It also introduces an edge-guided image optimization module to improve the integrity of the target shape and uses a local target amplifier to improve the visibility and performance of the target and suppress background clutter.
[0126] DNANet: By designing densely nested interaction modules, it realizes the progressive interaction of high-level and low-level features, maintains the deep mid-infrared small target information, and introduces cascade channels and spatial attention modules to adaptively enhance multi-layer features.
[0127] ②Dataset:
[0128] IRSTD-1k dataset: The IRSTD-1k dataset contains 1,001 real-world images covering a wide range of object shapes, sizes, and complex backgrounds. Object types include drones, ships, and vehicles. Each image has a resolution of 512×512. The dataset is divided into a training set (800 images) and a test set (201 images).
[0129] SIRST Aug Dataset: The SIRST Aug dataset involves augmenting the original image by first cropping the 512×512 image into a 256×256 pixel target region, ensuring that the target is located in the corner and center, and that the cropping is done based on the original aspect ratio. The training set of this dataset contains 8,525 images and the test set contains 545 images.
[0130] MDFA Dataset: The MDFA dataset includes both real infrared images and synthetic images. The training set and test set contain 9,978 and 100 samples, respectively. To enhance the robustness of the model, noise, flipping, and rotation augmentation are applied to the images in the training set.
[0131] ㈡Experimental results
[0132] The evaluation indicators of the eight algorithms on the IRSTD-1k remote sensing dataset are shown in Table 1:
[0133] Table 1
[0134]
[0135] Table 1 shows that this embodiment achieves relatively superior performance on the IRSTD-1k dataset compared to other algorithms. First, regarding the IoU performance metric, the method of this embodiment achieves an IoU of 0.6854, the best among all compared solutions. This demonstrates that this embodiment accurately locates targets. nIoU, which accounts for the influence of varying target sizes and is normalized to facilitate comparison of detection performance for targets of varying sizes, achieves a nIoU of 0.68, higher than most other methods. This demonstrates that this embodiment maintains relatively good performance in detecting small infrared targets of varying sizes. EGPNet ranks first in this performance metric, likely because nIoU is a normalized metric that accounts for target size and is affected by factors such as inconsistent target size and prediction accuracy of the detection bounding box. DNANet performs best in terms of AUC, but this embodiment remains close. Models with high AUCs generally maintain good detection capabilities across different thresholds, adapting to varying environmental variations and noise. ALCNet and ACM perform best in Fa, but their corresponding Pd are quite low. This suggests that ALCNet and ACM may over-mark target areas, resulting in an increase in false positives. Due to the excessive number of false positives, the actual target cannot be accurately detected, resulting in a decrease in the detection probability. In contrast, this embodiment has a relatively balanced Fa and Pd, which can effectively improve the target detection rate while ensuring a low false positive rate, thereby achieving more accurate target detection in complex backgrounds. The evaluation indicators of the eight algorithms on the SIRST-Aug remote sensing dataset are shown in Table 2:
[0136] Table 2
[0137]
[0138]
[0139] Table 2 shows that compared to Table 1, this embodiment performs significantly better on the SIRST-Aug dataset. Interference over Union (IoU) measures the overlap between the predicted and ground-truth bounding boxes. A higher IoU indicates a more accurate object localization. For IRSTD, an improved IoU directly reflects the accuracy and detail of the detected bounding boxes. This embodiment achieves an IoU of 0.5038, which compares favorably to other methods and is significantly higher than methods such as TopHat's 0.2438 and PSTNN's 0.2965. Furthermore, for nIoU, this embodiment maintains a top-tier performance, demonstrating its balanced ability to handle objects of varying sizes. It also significantly outperforms all other models except the EGPNet algorithm. The same holds true for Area under Conditional Uncertainty (AUC). AUC measures the overall performance of a model at different decision thresholds. A higher AUC indicates a model's ability to maintain high accuracy in distinguishing between positive and negative samples and exhibits strong robustness. For both Fa and Pd, this embodiment ranks second only to EGPNet. The results show that it performs well in terms of false alarm control and target detection capabilities, and can successfully identify most targets while ensuring a low false alarm rate. The evaluation indicators of the eight algorithms on the MDFA remote sensing dataset are shown in Table 3:
[0140] Table 3
[0141]
[0142]
[0143] As can be seen from Table 3, the IoU of this embodiment is 0.7629, ranking the highest, higher than the second-ranked EGPNet's 0.7614, and its performance is very outstanding. However, compared with other methods, such as TopHat's 0.1399 and PSTNN's 0.2076, it shows a significant improvement in positioning accuracy, indicating that it has almost reached the best level in target positioning. nIoU is the normalized IoU, which helps to eliminate the impact of target size on model performance. The nIoU of this embodiment is 0.7328, second only to EGPNet's 0.7259. This shows that the model can maintain high accuracy in the detection of targets of different sizes and performs well, especially when dealing with smaller targets. On the MDFA data, except for TopHat's 0.6065 and PSTNN's 0.6078, the other algorithms have relatively high performance AUCs, and are generally higher than 0.9. The difference between this embodiment and other algorithms is not large. Fa is second only to EGPNet, performing very well, but significantly lower than DNANet's 265.6 and ALCNet's 181.5. This indicates that the model is better able to handle background interference and reduce the chances of misidentifying background areas as targets. While PD is lower than EGPNet and ISNet, it is significantly higher than the other algorithms, demonstrating that this embodiment remains highly capable of detecting targets, especially when dealing with small objects.
[0144] Qualitative analysis:
[0145] In view of the fact that infrared small targets are extremely small and difficult to detect, different colors are used in the image to mark the detection results. Figure 2 As shown. Figure 2 It can be seen that traditional methods often produce false positives and missed detections when dealing with very small targets. Figure 2 In the ALCNet column, missed detection occurred on (2), in the ACM column, missed detection occurred on (2), and in the ISNet column, missed detection occurred on (2). In the ALCNet column, false detection occurred on (3), (4), and (5), in the ACM column, false detection occurred on (4), and in the DNANet column, false detection occurred on (3). The above-mentioned competing methods showed significant false negatives and false positives. Although the EGPNet column and the AFSENet column (this embodiment) did not show missed detection or false detection on all networks. However, when the prediction results of each network were compared with the GT (true result), the prediction results of the method in this embodiment were closer to the GT in shape, which was a relatively accurate detection result, but there were still slight differences. In contrast, the method proposed in this embodiment showed better performance in processing complex backgrounds and various target shapes. Accurate small target detection was achieved, and false positives and missed positives were effectively reduced, thereby bringing higher detection accuracy and clearer visual results.
[0146] The above quantitative and qualitative results show that the infrared small object target detection method based on the adaptive frequency domain separation enhancement network proposed in this embodiment designs a frequency domain adaptive decomposition module and a gradient positioning enhancement module, effectively integrating frequency domain and spatial domain information, providing a more accurate solution for small target detection. Moreover, after verification on multiple data sets, the average accuracy is better than other algorithms, which is sufficient to demonstrate the performance advantages and robustness of the algorithm in this embodiment.
[0147] The above description is merely a specific embodiment of the present invention, and a comparison of the effects of the specific embodiment with the relevant comparative examples. However, the scope of protection of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present invention shall be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be based on the scope of protection of the claims.
Claims
1. A method for infrared small object detection based on adaptive frequency domain separation enhancement network, characterized in that: The following steps are involved: Step 1: extracting original image data containing infrared small objects from the database; Step 2: Construct an adaptive frequency domain separation enhancement network, which includes a resnet residual network, a frequency domain adaptive separation module, an encoder, a gradient positioning enhancement module, and a decoder; The resnet residual network is used to extract the spatial domain features of the image through convolution operations; The frequency domain adaptive separation module is used to perform fast Fourier transform on the image, convert the image from the spatial domain to the frequency domain, generate frequency domain features, and adaptively separate the frequency domain features to obtain multiple frequency domain features; The encoder is used to fuse the spatial domain features and multi-band frequency domain features of the image to obtain coding features; The gradient positioning enhancement module calculates the gradient direction and gradient magnitude of the input image through the Sobel operator, and then calculates the saliency map, and integrates the saliency map into the decoder; The decoder is used to decode the encoded features to obtain a prediction result; Step 3: The adaptive frequency domain separation enhancement network is trained using the original image data to obtain an optimized adaptive frequency domain separation enhancement network; Step 4: Input the image to be detected into the optimized adaptive frequency domain separation enhancement network to obtain the infrared small object target detection result.
2. The infrared small object target detection method based on the adaptive frequency domain separation enhancement network according to claim 1 is characterized in that: In step 1, the original image data containing the infrared small object includes the original image containing the infrared small object and the true label of the infrared small object in the original image.
3. The infrared small object target detection method based on the adaptive frequency domain separation enhancement network according to claim 4 is characterized in that: Step 3 is as follows: Step 3.1, initialize the parameters of the adaptive frequency domain separation enhancement network, set the iterative convergence conditions, and randomly select multiple original images containing infrared small objects from the original data image I input And the original image I input The true labels of small mid-infrared objects are divided into multiple batches of raw data images; Step 3.2: The original images I containing infrared small objects in the same batch input Input to the resnet residual network, and extract the shallow spatial domain features F of the image through convolution operation shallow , mid-level spatial domain features F middle , deep spatial domain features F deep ; At the same time, the original image containing the infrared small object I input Input to the frequency domain adaptive separation module, the original image I input Convert from spatial domain to frequency domain to generate frequency domain features, and perform adaptive separation on the frequency domain features to obtain low frequency domain features F′ L , mid-frequency domain feature F′ B , high-frequency domain features F′ H ; Step 3.3: Encode the shallow spatial domain features F shallow , mid-level spatial domain features F middle , deep spatial domain features F deep and low-frequency domain features F′ L , mid-frequency domain feature F′ B , high-frequency domain features F′ H Perform splicing and fusion to obtain coding features; Step 3.4: Convert the original image containing the infrared small object I input Input the gradient positioning enhancement module to calculate the gradient direction and gradient amplitude of the input image, and then obtain the saliency map S grad , and the saliency map S grad Integrate into decoder; Step 3.5: Incorporate the encoded feature input into the saliency map S grad The decoder decodes the image to obtain the predicted label of the infrared small object in the original image; Step 3.6, combine the predicted labels of the infrared small objects in the original image with the true labels of the infrared small objects in the original image, calculate the loss value, use the back propagation algorithm to calculate the gradient of the adaptive frequency domain separation enhancement network parameters, and use the Adam optimizer to update the adaptive frequency domain separation enhancement network parameters, return to step 3.2, and use another batch of original images containing infrared small objects I input The input is fed into the resnet residual network until the iterative convergence condition is reached, and the adaptive frequency domain separation enhancement network parameters corresponding to the minimum loss value are selected to form an optimized adaptive frequency domain separation enhancement network.
4. The infrared small object detection method based on the adaptive frequency domain separation enhancement network according to claim 3 is characterized by: In step 3.2, the shallow spatial domain feature F of the image is extracted by convolution operation shallow , mid-level spatial domain features F middle , deep spatial domain features F deep The formulas are: F shallow =Conv1(I input ) F middle =Conv2(MaxPool(F shallow )) F deep =ReLU(Conv3(F middle )) Among them, Conv1, Conv2, and Conv3 represent convolution operations, MaxPool represents the maximum pooling operation, and ReLU represents the activation function.
5. The infrared small object detection method based on the adaptive frequency domain separation enhancement network according to claim 3 is characterized by: In step 3.2, the original image I containing the infrared small object is input Input to the frequency domain adaptive separation module, the original image I input Converting from the spatial domain to the frequency domain, the formula for generating frequency domain features is: Among them, F(u,v) represents the frequency domain characteristics of the image, f(x,y) represents the spatial domain characteristics, that is, the pixel value of the original image, (u,v) represents the frequency domain obtained after the image is subjected to a two-dimensional Fourier transform, u and v represent the horizontal and vertical component positions of each frequency in the image, respectively, M and N are the number of rows and columns of the image, respectively, x represents the horizontal pixel position, and y represents the vertical pixel position.
6. The infrared small object detection method based on the adaptive frequency domain separation enhancement network according to claim 5, characterized in that: In step 3.2, the adaptive separation of frequency domain features is specifically as follows: Step a: Calculate the amplitude spectrum |F(u,v)| based on the frequency domain characteristics. The calculation formula is: Step b: Determine the cutoff frequency D of the low-pass filter and high-pass filter based on the amplitude spectrum |F(u,v)| L and D H , respectively use high-pass filter, low-pass filter and band-pass filter to separate the frequency domain features F(u,v) to obtain the high-frequency domain features F′ H , low-frequency domain features F′ L And the mid-frequency domain feature F′ B , the separation formulas are: F′ B =F′ L -F′ H ; Where D L and D H Represent the cutoff frequencies of the low-pass filter and high-pass filter respectively; D(u,v) represents the frequency point (u, v) to the original image I input The distance to the center of mass.
7. The infrared small object target detection method based on adaptive frequency domain separation enhancement network according to claim 3 is characterized in that: Step 3.4 is as follows: Step 3.4.1, the original image containing the infrared small object I input Input the gradient positioning enhancement module, and convolve the image with the horizontal Sobel convolution kernel to obtain the horizontal gradient value G of each pixel in the image. x By convolving the image with the Sobel convolution kernel in the vertical direction, the gradient value G of each pixel in the image in the vertical direction is obtained. y , calculate the input original image I input The gradient direction θ and gradient amplitude G are calculated as follows: Step 3.4.2: According to the input original image I input The gradient direction θ and gradient magnitude G are used to calculate the variance C of the gradient direction in the local area of each pixel. θ and the variance of the gradient magnitude C G , the calculation formula is: Where n represents the number of pixels in the local area, θ i represents the gradient direction of the i-th pixel in the local area, Represents the mean value of the gradient direction of all pixels in the local area; G i represents the gradient magnitude of the i-th pixel in the local area, Represents the mean value of the gradient magnitude of all pixels in the local area, Among them G x,i and G y,i are the horizontal and vertical gradients of the i-th pixel; Step 3.4.
3. Calculate the saliency map S grad , the formula is: S grad =exp(-α·C θ -β·C G ), Where α and β are weight coefficients used to adjust the gradient direction consistency and amplitude consistency respectively; The saliency map S grad Integrate into the decoder.
8. The infrared small object target detection method based on adaptive frequency domain separation enhancement network according to claim 3 is characterized in that: Step 3.5 is as follows: The encoding feature input is integrated into the saliency map S grad The decoder is decoded, and the decoding formula is: Where I represents the predicted label of the infrared small object in the original image, Represents the splicing and fusion operation, represents the upsampling Up operation, ⊙ represents the supplementary operation, and F represents the spatial domain feature, which is the shallow spatial domain feature F shallow , mid-level spatial domain features F middle , deep spatial domain features F deep ; f represents the frequency domain feature, that is, the low-frequency domain feature F′ L , mid-frequency domain feature F′ B , high-frequency domain features F′ H .
9. The infrared small object target detection method based on adaptive frequency domain separation enhancement network according to claim 3 is characterized in that: In step 3.6, the loss value is calculated specifically as follows: the binary cross entropy loss value, intersection-over-union loss value, and Dice loss value are calculated respectively based on the predicted label of the infrared small object in the original image and the true label of the infrared small object in the original image, and the binary cross entropy loss value, intersection-over-union loss value, and Dice loss value are weightedly summed to obtain the loss value.
10. The infrared small object target detection method based on adaptive frequency domain separation enhancement network according to claim 3, characterized in that: In steps 3.1 and 3.6, the iterative convergence condition is: After running a round, calculate the IoU of the current round and compare it with the historical best IoU. If the IoU of the current round is greater than the historical best IoU, update the IoU of the current round to the historical best IoU; otherwise, keep the historical best IoU. If the intersection-over-union ratio for 10 consecutive rounds is not greater than the historical best intersection-over-union ratio, the iteration is terminated.
Citation Information
Cited By
Infrared small target detection method supporting set driving and significance priori guidance
CN122289827A
A method for infrared small target detection that supports set-driven and saliency-prior-guided approaches.
CN122289827B