Infrared small target detection method and system
By adopting spatial and frequency domain dual branch detection methods in infrared small object detection, and using multi-scale void contrast convolution and dynamic high-pass filter modules, the problems of variable size, dull appearance and complex background in infrared small object detection are solved, and more efficient object detection performance is achieved.
Patent Information
- Application Number
- CN202510062080.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-15
AI Technical Summary
Infrared small target detection faces problems such as variable size, dull appearance, low signal-to-noise ratio and complex background, resulting in false alarms and missed detection.
The dual-branch detection method based on spatial domain and frequency domain is adopted, and the target perception is enhanced by using a multi-scale hollow contrast convolution module, and the low-frequency background information is suppressed through the dynamic high-pass filter module, and the spatial domain and frequency domain prediction map are fused to improve detection performance.
It effectively enhances the performance of infrared small object detection, especially when dealing with small, variable-size targets and low signal-to-noise ratio scenarios, reducing false alarms and missed detection, significantly improving contrast and detection accuracy.
Smart Images

Figure CN120047665A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of target detection, and particularly to an infrared small target detection method and system. Background Art
[0002] Infrared Small Target Detection (IRSTD) aims to identify small targets in infrared images and separate them from complex backgrounds. It has a wide range of applications, including military reconnaissance and traffic monitoring. Compared with the detection of general objects in natural scene images, the detection of small targets in infrared images faces two main challenges: their small and variable sizes, and their dim appearance and low signal-to-noise ratio. Due to the different distances between the camera and the target, infrared small targets only occupy a few to dozens of pixels, making it difficult to accurately separate them. In addition, there are a large number of slowly changing low-frequency background regions in infrared images, which reduce the visibility and contrast of the targets and make them easily occluded. All these obstacles may lead to false alarms or missed detections of targets, making IRSTD still a challenging task.
[0003] Current infrared small target detection methods are mainly divided into two categories: traditional methods and deep learning-based methods. Traditional methods, such as filter-based methods, low-rank representation-based methods, and local contrast-based methods, aim to enhance local contrast or highlight target regions through manually designed priors or structural assumptions. Infrared small target images in real scenes are much more complex, with large variations in target sizes and cluttered backgrounds, making it difficult for traditional methods to handle such variations with handcrafted features and fixed hyperparameters. In contrast, deep learning-based methods use an end-to-end learning paradigm and use gradient descent algorithms to learn features from a large number of complex scenes, achieving better performance than traditional methods. These methods rely on model design to achieve more effective feature extraction. However, the inherent ambiguity of small targets often misleads the model, making it vulnerable to interference from complex backgrounds, and targets of different scales may lead to false alarms.
[0004] In recent years, with the development of deep learning, especially Convolutional Neural Networks (CNNs), end-to-end networks such as convolutional neural networks are used to directly learn features from data without complex manual feature design. Although these deep learning methods have achieved satisfactory performance in most simple scenarios, they still lack comprehensive discrimination and separation capabilities for small, variable-sized, dimly lit, and low signal-to-noise ratio targets. Especially when the background and the target have similar frequency components, the model may misidentify the background as the target, resulting in false alarms and missed detections of the target. Therefore, accurately identifying infrared small targets of different sizes under a low-frequency background and low signal-to-noise ratio is still a difficult task that has not been deeply explored. Summary of the Invention
[0005] To solve the above problems, the present disclosure proposes an infrared small target detection method and system, which is based on two branches in the spatial domain and the frequency domain, and uses the multi-scale target perception ability in the spatial domain and the low-frequency information suppression ability in the frequency domain to enhance the performance of infrared small target detection.
[0006] According to some embodiments, the present disclosure adopts the following technical solutions:
[0007] An infrared small target detection method, comprising:
[0008] Using a spatial domain detection model, performing multi-scale small target detection on the infrared image to be detected, and generating a spatial domain prediction map;
[0009] Based on the spatial domain prediction map, through a frequency domain enhancement model, performing multi-level high-frequency information enhancement on the frequency domain map corresponding to the infrared image, and generating a frequency domain prediction map;
[0010] Fusing the spatial domain prediction map and the frequency domain prediction map to obtain a final small target detection map;
[0011] Wherein, the spatial domain detection model enhances the perception of infrared small targets through a multi-scale atrous contrast convolution module, using multiple parallel atrous contrast convolutions with different kernel sizes; the frequency domain enhancement model dynamically removes low-frequency information by calculating the low-frequency signal energy layer by layer through a dynamic high-pass filter module to retain high-frequency image details.
[0012] According to some embodiments, the present disclosure adopts the following technical solutions:
[0013] An infrared small target detection system, comprising:
[0014] A spatial domain module, configured to: use a spatial domain detection model to perform multi-scale small target detection on the infrared image to be detected, and generate a spatial domain prediction map;
[0015] A frequency domain module, configured to: based on a spatial domain prediction map, through a frequency domain enhancement model, perform multi-level high-frequency information enhancement on the frequency domain map corresponding to the infrared image to generate a frequency domain prediction map;
[0016] A fusion module, configured to: fuse the spatial domain prediction map and the frequency domain prediction map to obtain a final small target detection map;
[0017] Wherein, the spatial domain detection model enhances the perception of infrared small targets through a multi-scale dilated contrast convolution module, using multiple parallel dilated contrast convolutions with different kernel sizes; the frequency domain enhancement model dynamically removes low-frequency information by calculating the low-frequency signal energy layer by layer through a dynamic high-pass filter module to retain high-frequency image details.
[0018] According to some embodiments, the present disclosure adopts the following technical solutions:
[0019] A computer program product, including a computer program, which when executed by a processor implements the infrared small target detection method described above.
[0020] According to some embodiments, the present disclosure adopts the following technical solutions:
[0021] A non-transitory computer-readable storage medium, which is used to store computer instructions, and when the computer instructions are executed by a processor, the infrared small target detection method described above is implemented.
[0022] According to some embodiments, the present disclosure adopts the following technical solutions:
[0023] An electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements the infrared small target detection method described above.
[0024] Compared with the prior art, the beneficial effects of the present disclosure are:
[0025] 1. The present invention provides an infrared small target detection method and system, based on two branches of the spatial domain and the frequency domain, using the multi-scale target perception ability of the spatial domain and the low-frequency information suppression ability of the frequency domain to enhance the performance of infrared small target detection.
[0026] 2. The present invention proposes a spatial domain multi-scale dilated contrast convolution module, with multiple parallel dilated contrast convolutions of three specially designed different kernel sizes, which improves the contrast between the target and the cluttered background and enhances the perception ability of small and variable-sized targets.
[0027] 3. The present invention proposes a frequency-domain dynamic high-pass filter module, which calculates the energy of low-frequency signals in layers and dynamically removes specific low-frequency information to retain high-frequency image details, effectively filtering out the slow-changing low-frequency background interference and highlighting small targets. Description of the Drawings
[0028] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of the disclosure. The illustrative embodiments and descriptions thereof of the disclosure are used to explain the disclosure and do not constitute an improper limitation of the disclosure.
[0029] Figure 1 is the model framework diagram proposed in Embodiment 1;
[0030] Figure 2 is the internal structure diagram of the multi-scale dilated contrast convolution module in Embodiment 1;
[0031] Figure 3 is the internal structure diagram of the dynamic high-pass filter module in Embodiment 1;
[0032] Figure 4 is the visualization result diagram of the method in Embodiment 1 compared with five existing networks. Detailed Embodiments
[0033] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0034] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.
[0035] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "comprising" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0036] Embodiment 1
[0037] An infrared small target detection method is provided in an embodiment of the present disclosure, including:
[0038] Step 1: Using a spatial domain detection model, perform multi-scale small target detection on the infrared image to be detected to generate a spatial domain prediction map;
[0039] Step 2: Based on the spatial domain prediction map, through the frequency domain enhancement model, perform multi-level high-frequency information enhancement on the frequency domain map corresponding to the infrared image to generate a frequency domain prediction map;
[0040] Step 3: Fuse the spatial domain prediction map and the frequency domain prediction map to obtain the final small target detection map;
[0041] Among them, the spatial domain detection model uses a multi-scale atrous contrast convolution module, and utilizes multiple parallel atrous contrast convolutions with different kernel sizes to enhance the perception of infrared small targets; the frequency domain enhancement model uses a dynamic high-pass filter module to hierarchically calculate the low-frequency signal energy and dynamically remove the low-frequency information to retain the high-frequency image details.
[0042] As an embodiment, an infrared small target detection method of the present disclosure is based on two branches of the spatial domain and the frequency domain, and utilizes the multi-scale target perception ability in the spatial domain and the low-frequency information suppression ability in the frequency domain to enhance the performance of infrared small target detection, especially enhancing the perception of small and variable-sized targets. The specific implementation process is described in detail below.
[0043] This embodiment designs an infrared small target detection network based on the hybrid domain, as Figure 1 shown, including a spatial domain detection model, a frequency domain enhancement model, and a fusion model. Among them, the spatial domain detection model and the frequency domain enhancement model correspond to the spatial domain branch and the frequency domain branch respectively.
[0044] The spatial domain detection model adopts a classic encoder-decoder framework. The encoder uses a specially designed multi-scale atrous contrast convolution (MAC) module as its core component, while the decoder uses ordinary convolution blocks to gradually upsample and restore the target information. The encoder output features are fused with the decoder through skip connections to enhance the capture of context information. Four decoder levels generate multi-level prediction maps, which help to suppress the complex background interference in the following frequency domain.
[0045] Each stage of the frequency domain enhancement model integrates a specially designed dynamic high-pass filter (DHPF) module to gradually remove the low-frequency information of the image.
[0046] The fusion model merges the final spatial prediction map and the frequency domain prediction map to generate the final infrared small target prediction map.
[0047] Provide a specific example of detection using the infrared small target detection network. The specific steps are as follows:
[0048] S1: Use the multi-scale atrous contrast convolution (MAC) module to perform feature extraction on the input image in the encoder stage of the spatial domain detection model; and generate the final spatial prediction map and the decoder stage prediction map.
[0049] The spatial domain detection model, as shown in Figure 1 the spatial domain part of it, includes several layers of encoders and several layers of decoders; the encoder includes a multi-scale dilated contrast convolution module, a pooling layer, a residual block, and an upsampling block, which are used for multi-scale feature extraction and provide skip connection inputs for the decoder; the decoder, based on the extracted multi-scale features, obtains prediction maps of different scales through convolution blocks.
[0050] Specifically, the five-layer encoder performs multi-scale feature extraction through the MAC module and provides skip connection inputs for the decoder stage; the four-layer decoder obtains four prediction maps P 1 , P 2 , P 3 , P 4 , and obtains the final spatial domain prediction map P through the following formula 0 :
[0051] P 0 = Conv 1×1 (Concat(P 1 , BU-2(P 2 ), BU-4(P 4 ), BU-8(P 8 ))),
[0052] where Conv 1×1 (·) is a 1×1 convolution block, Concat(·) refers to the channel fusion operation, and BU-X(·) refers to upsampling the feature map by X times.
[0053] The MAC module, as shown in Figure 2 , includes four branches: a direct connection branch and three parallel dilated contrast branches. The specific processing steps are as follows:
[0054] S11: Given the input feature map F in , use 1×1 convolution to expand the channel dimension and generate the feature map X:
[0055] X = Conv 1×1 (F in )
[0056] where Conv 1×1 (·) represents a 1×1 convolution block, aiming to achieve channel dimension expansion.
[0057] S12: The feature map X is evenly divided into four groups along the channel dimension, denoted as X = [X 0 , X 1 , X 2 , X 3 .
[0058] X 0 As a direct connection and immediate output, marked as Z 0 , that is, the direct connection branch.
[0059] X 1 ,X 2 ,X 3 Input in parallel to three different dilated contrast convolution branches, namely ACC-1, ACC-2, and ACC-3, to generate contrast features with different receptive fields, denoted as Z 1 ,Z 2 ,Z 3 , specifically calculated as follows:
[0060]
[0061] Among them, ACC 1 (·), ACC 2 (·), ACC 3 (·) represent the calculations of three parallel multi-scale dilated contrast convolution kernels. The specific calculation method is:
[0062]
[0063] Among them, (u,v) represents the pixel coordinates in the feature map X i . Y represents the set of central pixels marked in yellow, B represents the set of surrounding pixels marked in blue within a contrast convolution kernel ( Figure 2 in ACC-1, ACC-2, ACC-3), and n and m represent the total number of central or surrounding pixels in a contrast kernel.
[0064] S13: All features Z 0 、Z 1 、Z 2 、Z 3 are concatenated along the channel dimension to obtain the contrast feature Z = [Z 0 ,Z 1 ,Z 2 ,Z 3 , and feature mixing is performed on Z through a 1×1 convolution block.
[0065] S14: The feature map X on the direct connection branch and the contrast feature Z after feature mixing are passed through a residual block by the following formula to obtain a contrast-enhanced feature map:
[0066] F out =Res_Block(Conv 1×1 (Z)+X),
[0067] Among them, Conv 1×1(·) refers to a 1×1 convolution block, and Res_Block(·) consists of Conv 3×3 (·), channel attention CA(·), and spatial attention SA(·).
[0068] S2: Using the dynamic high-pass filter (DHPF) module, in the frequency domain enhancement model, filter and reduce the low-frequency energy of the spatial prediction map generated by the spatial domain detection model and the prediction map in the decoder stage to obtain the frequency domain feature map.
[0069] The frequency domain enhancement model, as shown in the frequency domain part of Figure 1 is composed of several cascaded dynamic high-pass filter (DHPF) modules. The target enhanced feature map output by the last dynamic high-pass filter module is used as the frequency domain prediction map.
[0070] The structures of the four dynamic high-pass filter (DHPF) modules are the same. Taking the first DHPF module as an example, as shown in (a) of Figure 3 , the specific processing steps are as follows:
[0071] S21: Generate a frequency domain map through the fast Fourier transform.
[0072] First, perform 8-fold upsampling on the prediction map P 4 output by the decoder, and obtain the prediction map P' through the Sigmoid function 4 . Perform initial filtering on the prediction map P' 4 to effectively reduce background interference and generate a target enhanced image. Then, the target enhanced image is converted into a frequency domain map, that is, a frequency feature map, through the fast Fourier transform (FFT), which is expressed by the formula:
[0073]
[0074] where IRI ∈ R H×W×1 and P' 4 ∈ R H×W×1 represent the original infrared image and the prediction maps generated by the corresponding decoder stages respectively. ⊙ represents element-wise multiplication, refers to the frequency feature map, and FFT(·) refers to the fast Fourier transform algorithm.
[0075] S22: Gradually filter a preset proportion of low-frequency information from the frequency domain map.
[0076] Figure 3 The frequency feature map shown in (a) of It shows that when approaching the center, the frequency decreases and the amplitude increases, while when moving away from the center, the frequency increases and the amplitude decreases, indicating that the infrared image contains rich low-frequency background information; since both the small target area and the edge area belong to high-frequency information, in order to effectively eliminate the low-frequency background and high-frequency edge background information, this embodiment proposes a method for gradually filtering out a certain proportion of low-frequency information, and the specific steps are as follows:
[0077] (1) Calculate the energy of the frequency feature map :
[0078]
[0079] where (u, v) represents the pixel coordinates in the frequency feature map, is the amplitude value at pixel (u, v), and EC represents the total frequency domain energy of the image.
[0080] (2) According to the preset energy removal ratio, calculate the dynamic filtering mask radius of the frequency domain map, which is expressed by the formula:
[0081]
[0082] where E removed represents the low-frequency energy to be suppressed in the DHPF stage, λ 1 refers to the preset energy removal ratio, and the dynamic filtering mask radius d takes the maximum radius value that satisfies the above equation.
[0083] (3) According to the radius, obtain the dynamic filtering mask as follows:
[0084]
[0085] where Mask(u, v) ∈ R H×W×1 represents the dynamic filtering mask.
[0086] Since the initial infrared image contains rich low-frequency information, and as the decoder transitions to shallower layers, its prediction map shows a reduction in the amount of low-frequency background. Therefore, gradually reduce the energy filtering ratio and preset it as λ = [λ 1 , λ 2 , λ 3 , λ 4 = [0.8, 0.4, 0.2, 0.1].
[0087] (4) Use the dynamic filtering mask to process the frequency domain map to obtain the target enhanced feature map:
[0088]
[0089] where iFFT(·) represents the inverse fast Fourier transform, Represents the output of the first DHPF module. The frequency-domain visualization feature map is as shown in Figure 3 (d) in, it can be seen that as λ in different DHPF modules i decreases, the low-frequency information decreases. Figure 3 The spatial-domain prediction maps shown in columns 2-5 of (b) and (c) in show that as λ i decreases, the large-area low-frequency background information and the edge background regions of high-frequency information are gradually suppressed.
[0090] Finally, after passing through four DHPFs, the frequency-domain enhancement model obtains the prediction map in the frequency-domain space
[0091] S3: Merge the spatial prediction map and the frequency-domain prediction map to generate the final infrared small target prediction map.
[0092] Specifically, the final prediction map P is obtained by fusing the prediction map P in the spatial domain 0 and the prediction map in the frequency domain and is expressed by the formula:
[0093]
[0094] where sig(·) is the Sigmoid activation function, and P represents the infrared small target prediction map, which can effectively suppress the background region.
[0095] Train the infrared small target detection network composed of the spatial domain detection model, the frequency-domain enhancement model, and the fusion model, and use the position-sensitive loss function to supervise the training process. The position-sensitive loss function is composed of the scale-sensitive loss and the position-sensitive loss.
[0096] Specifically, during the training process, in order to ensure the effective learning and optimization of the model, a deep learning supervision mechanism is adopted to supervise the prediction maps (including the intermediate prediction maps of the spatial domain detection model and the final prediction maps output by the infrared small target detection network) in multiple aspects, and the infrared small target detection performance is improved by calculating the loss functions of all intermediate prediction maps and final prediction maps generated by each decoder.
[0097] The loss function uses the position-sensitive loss function (Scale and Location Sensitive Loss, SLSLoss) to comprehensively improve the detection accuracy of the model. The formula of the position-sensitive loss function is:
[0098] L SLS = L S + L L
[0099] where L Sand L L respectively refer to the scale-sensitive loss and the position-sensitive loss.
[0100] The scale-sensitive loss L S is defined as:
[0101]
[0102] where P and G represent the predicted map and the ground truth map, |·| represents the count of the pixel set, and Var(·) calculates the variance of the given scalar.
[0103] The position-sensitive loss L L is defined as:
[0104]
[0105] where c p =(x p , y p ) and c gt =(x gt , y gt ) are the center points of the predicted target pixel set in the predicted map P and the ground truth target pixel set in the ground truth map G, respectively. They are calculated by averaging the coordinates of all pixels in each pixel set. Then, we convert the coordinates of these center points to the polar coordinate system. (d p , θ p ) and (d gt , θ gt ) represent the distance and angle of c p and c gt , respectively.
[0106] Based on the position-sensitive loss function L SLS , this embodiment constructs a multi-scale ground truth map to supervise the predicted maps generated at different stages, that is, the four intermediate predicted maps P 1 , P 2 , P 3 , P 4 of the spatial domain detection model and the final predicted map P output by the infrared small target detection network. The specific loss design is as follows:
[0107]
[0108] where G is the ground truth map, is the operation of spatially downsampling the first parameter with the second parameter as the factor, is the constructed multi-scale ground truth map, and P i (i∈[1,2,3,4]) refers to the predicted maps obtained from the four decoder stages, and P refers to the final predicted map.
[0109] In this embodiment, a comparative experiment on the effects of this method and existing method models (PBT, SCTransNet, MSHNet, RPCANet, and MTUNet) was conducted on the same dataset (three international standard datasets IRSTD-1K, NUAA-SIRST, and NUDT-SIRST). The detection accuracy was evaluated using the metrics IoU, Pd, and Fa. The experimental results are shown in Table 1 and Figure 4 as follows:
[0110] Table 1 Comparison Table of Detection Accuracy
[0111]
[0112] In Table 1, among the 9 metrics of the three datasets, the method of this embodiment achieved seven best performances, one second performance, and one third performance. Especially in the NUAA-SIRST dataset, the target-level Pd accuracy of 100% was achieved; Figure 4 is the visualization result of the method of this embodiment compared with the five existing networks. The correctly detected targets, missed targets, and misdetected targets are framed by red, blue, and yellow boxes respectively, and the enlarged view of the target is shown in the corner of the image; it can be seen that there are obvious misdetections and missed detections in the existing networks (the 4th - 8th columns), resulting in a higher Fa. However, the method of this embodiment can accurately identify and segment these small targets; therefore, this embodiment can effectively enhance the performance of infrared small target detection.
[0113] Embodiment 2
[0114] In an embodiment of the present disclosure, an infrared small target detection system is provided, including:
[0115] A spatial domain module, configured to: use a spatial domain detection model to perform multi-scale small target detection on the infrared image to be detected, and generate a spatial domain prediction map;
[0116] A frequency domain module, configured to: based on the spatial domain prediction map, perform multi-level high-frequency information enhancement on the frequency domain map corresponding to the infrared image through a frequency domain enhancement model, and generate a frequency domain prediction map;
[0117] A fusion module, configured to: fuse the spatial domain prediction map and the frequency domain prediction map to obtain a final small target detection map;
[0118] Wherein, the spatial domain detection model enhances the perception of infrared small targets through a multi-scale dilated contrast convolution module, using multiple parallel dilated contrast convolutions with different kernel sizes; the frequency domain enhancement model dynamically removes low-frequency information by calculating the low-frequency signal energy layer by layer through a dynamic high-pass filter module to retain high-frequency image details.
[0119] Embodiment 3
[0120] In one embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the infrared small target detection method described above is implemented.
[0121] Embodiment 4
[0122] In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, and the non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the infrared small target detection method described above is implemented.
[0123] Embodiment 5
[0124] In one embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements the infrared small target detection method described above.
[0125] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0126] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0127] Although the specific embodiments of the present disclosure are described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.
Claims
1. A method for detecting small infrared targets, characterized in that: include: Using the spatial domain detection model, multi-scale small target detection is performed on the infrared image to be detected, and a spatial domain prediction map is generated; Based on the spatial domain prediction map, a frequency domain enhancement model is used to perform multi-level high-frequency information enhancement on the frequency domain map corresponding to the infrared image to generate a frequency domain prediction map; The spatial domain prediction map and the frequency domain prediction map are fused to obtain the final small target detection map; Among them, the spatial domain detection model uses a multi-scale hole contrast convolution module and multiple parallel hole contrast convolutions with different kernel sizes to enhance the perception of small infrared targets; the frequency domain enhancement model uses a dynamic high-pass filter module to hierarchically calculate the low-frequency signal energy and dynamically remove the low-frequency information to retain the high-frequency image details.
2. A method for detecting small infrared targets as claimed in claim 1, characterized in that: The spatial domain detection model includes several layers of encoders and several layers of decoders; The encoder includes a multi-scale hole contrast convolution module, a pooling layer, a residual block, and an upsampling block, which are used to perform multi-scale feature extraction and provide a skip connection input for the decoder; The decoder obtains prediction images of different scales through convolution blocks based on the extracted multi-scale features.
3. The infrared small target detection method according to claim 1, characterized in that: The multi-scale hole contrast convolution module includes a direct connection branch and a plurality of hole contrast branches. The hole contrast branch uses a multi-scale hole contrast convolution kernel to extract the contrast features of the feature map. The specific formula is: Among them, (u,v) represents the feature map X i Y represents the central pixel set, B represents the surrounding pixel set, n and m represent the total number of central pixels and surrounding pixels respectively.
4. The infrared small target detection method according to claim 1, characterized in that: The frequency domain enhancement model is composed of a plurality of cascaded dynamic high-pass filter modules, and the target enhancement feature map output by the last dynamic high-pass filter module is used as the frequency domain prediction map; The specific operation of the dynamic high-pass filter module is: Based on the spatial domain prediction image and the infrared image, a frequency domain image is generated through fast Fourier transform; Calculate the dynamic filter mask radius of the frequency domain image according to the preset energy removal ratio; According to the radius, a dynamic filtering mask is obtained; The frequency domain image is processed using a dynamic filter mask to obtain the target enhanced feature map.
5. The infrared small target detection method according to claim 1, characterized in that: The fusion of the spatial domain prediction map and the frequency domain prediction map is performed by first activating the frequency domain prediction map through a fusion model, then multiplying the frequency domain prediction map pixel by pixel with the spatial domain prediction map, and finally adding the frequency domain prediction map to the spatial domain prediction map.
6. A method for detecting small infrared targets as claimed in claim 5, characterized in that: It also includes training the spatial domain detection model, the frequency domain enhancement model and the fusion model as a whole, and supervising the training process using a position sensitive loss function, wherein the position sensitive loss function is composed of a scale sensitive loss and a position sensitive loss.
7. An infrared small target detection system, characterized in that: include: The spatial domain module is configured to: use the spatial domain detection model to perform multi-scale small target detection on the infrared image to be detected and generate a spatial domain prediction map; The frequency domain module is configured to: based on the spatial domain prediction map, perform multi-level high-frequency information enhancement on the frequency domain map corresponding to the infrared image through a frequency domain enhancement model to generate a frequency domain prediction map; The fusion module is configured to: fuse the spatial domain prediction map and the frequency domain prediction map to obtain a final small target detection map; Among them, the spatial domain detection model uses a multi-scale hole contrast convolution module and multiple parallel hole contrast convolutions with different kernel sizes to enhance the perception of small infrared targets; the frequency domain enhancement model uses a dynamic high-pass filter module to hierarchically calculate the low-frequency signal energy and dynamically remove the low-frequency information to retain the high-frequency image details.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, an infrared small target detection method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, an infrared small target detection method as described in any one of claims 1-6 is implemented.
10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes an infrared small target detection method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Encoder and decoder large intestine polyp detection method based on cavity convolution
CN112200773A
Infrared small target detection method based on double-flow enhanced network
CN115565034A
Aerial infrared dim small moving target detection method
CN119169262A
Infrared weak and small target detection method based on wavelet guide state space model
CN119251618A
Method, apparatus, and system for wireless motion monitoring based on classified sliding time windows
EP4295760A1
Cited By
Infrared small target detection supervision enhancement method
CN120580417A