Infrared small target detection method based on frequency domain decomposition

Through the detection method based on frequency domain decomposition, the infrared image is adaptively decomposed into high-frequency and low-frequency components, enhance the target characteristics and suppress noise, solving the problem of low detection accuracy of small infrared targets in complex environments, and achieving higher detection robustness and accuracy.

CN120163973AActive Publication Date: 2025-06-17NINGBO UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
CN202510645481.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-17
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

In complex environments, accurate detection of small infrared targets still faces many challenges, including small size, low contrast and low signal-to-noise ratio, resulting in bottlenecks in target enhancement, background suppression and robust feature extraction.

Method used

The detection method based on frequency domain decomposition is adopted, through three stages: adaptive frequency domain decomposition, frequency domain information perception and frequency domain information fusion, the input infrared image is decomposed into high-frequency and low-frequency components, which automatically enhances the effective frequency domain components and suppresses invalid noise information, thereby improving the robustness and accuracy of target detection.

Benefits of technology

It significantly improves the detection performance of small infrared targets, especially in complex backgrounds and low contrast environments, target features can be extracted more accurately and accurately detect and real-time tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163973A_ABST
    Figure CN120163973A_ABST
Patent Text Reader

Abstract

The invention provides an infrared small target detection method based on frequency domain decomposition, and aims at solving the problem that a traditional spatial domain method is insufficient in detection performance under the conditions of complex background, low contrast and noise interference. Aiming at the problem that an infrared small target generally has relatively small spatial distribution and high-frequency characteristics, an input image is decomposed into high-frequency and low-frequency components through a self-adaptive frequency domain decomposition technology, frequency domain information perception and a spatial information aggregation module are combined, and frequency domain characteristics are weighted and extracted by adopting a sparse attention mechanism; the contributions of different frequency components are dynamically adjusted. Through efficient feature fusion, the expression of target information is optimized, and especially in a low-contrast and complex background environment, the detection precision and robustness are improved. An experimental verification result on an NUAA data set shows that compared with an existing mainstream method, the method has remarkable advantages on multiple evaluation indexes, and the effectiveness and superiority of the method in infrared small target detection are verified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of infrared target detection, and in particular relates to research on a small infrared target detection method based on frequency domain decomposition. Background Art

[0002] Compared with traditional RGB imaging technology, infrared imaging uses its detection capability beyond the visible spectrum to effectively capture the infrared radiation signal of the target at night and in bad weather conditions, realizing target detection and identification in all weather and all environments. This feature makes it play an indispensable role in natural disaster warning, military reconnaissance and precision navigation. However, in complex environments, the accurate detection of small infrared targets still faces many challenges. These small targets are usually small in size, low in contrast and low in signal-to-noise ratio. Due to the long imaging distance and the limitation of device resolution, these targets account for a very low proportion in the image, usually not more than 0.2%, and are easily masked by background noise. In addition, affected by the uneven distribution of environmental thermal radiation, the target outline is blurred and lacks texture information, resulting in bottlenecks in target enhancement, background suppression and robust feature extraction for detection methods based on traditional methods. Therefore, how to effectively extract target features, enhance contrast, and realize accurate detection and real-time tracking of small infrared targets under high noise and high dynamic backgrounds is still a key technical problem that needs to be solved urgently.

[0003] However, due to the lack of color information and texture features of small infrared targets, existing infrared small target detection methods mainly focus on the existence of targets and attempt to extract features by building deeper neural networks. However, as the network deepens, the detailed features of small targets are easily lost, while in segmentation tasks, edge details and contour information are crucial. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a detection method based on frequency domain decomposition, which is applied to the field of infrared small target detection. Most of the existing infrared small target detection methods rely on traditional spatial domain feature extraction methods. Although certain effects can be achieved in some simple scenarios, in complex backgrounds, especially in the case of low contrast and strong noise interference, problems such as difficulty in effectively distinguishing the target from the background and low detection accuracy are often faced. In addition, many methods lack flexibility in processing high-frequency and low-frequency information and fail to make full use of the discrimination ability of different frequency domain components for the target, resulting in the neglect of the detailed information of small targets or being submerged by background noise. This method introduces a frequency domain analysis mechanism. First, it adaptively decomposes the input infrared image in the frequency domain, decomposing the image into high-frequency and low-frequency components, giving full play to the expression ability of the high-frequency components for the boundary details of the target and the suppression effect of the low-frequency components on the background information. Therefore, the present invention proposes an infrared small target detection method based on frequency domain decomposition. Through frequency domain weighting and information perception, the method can automatically enhance effective frequency domain components while suppressing invalid noise information, thereby improving the robustness and accuracy of target detection. In addition, by using the spatial information aggregation module, high-frequency and low-frequency features are efficiently fused, further optimizing the expression of target information. Especially in complex backgrounds and low-contrast environments, the detection performance of infrared small targets can be significantly improved. Its core lies in the following four parts:

[0005] S1. Construct a network model based on frequency domain decomposition. This model is divided into three stages: adaptive frequency domain decomposition, frequency domain information perception, and frequency domain information fusion, which are used to dynamically decompose the input infrared image into high-frequency and low-frequency components, extract effective information from the decomposed high-frequency and low-frequency components, and fuse the features of different frequency components, and finally obtain a clean target mask;

[0006] S2. Construct an adaptive frequency domain selection module and a frequency domain decomposition module. Through the learnable discrete cosine transform, different frequency components of the input features are adaptively enhanced, and the enhanced features are dynamically decomposed into high-frequency and low-frequency components, so as to facilitate subsequent steps to more easily capture the key frequency domain features of the target object;

[0007] S3. Design a spatial sparse attention network. According to the different spatial sparsity of the image space dimension, this module establishes a long-distance spatial dependence relationship by efficiently calculating the similarity of the non-local space in the input features, thereby significantly improving the network's perception ability of effective targets;

[0008] S4. Design a spatial information aggregation module. Through bidirectional attention information extraction of rows and columns, effective integration of high-frequency and low-frequency features is achieved.

[0009] Furthermore, the network model based on frequency domain decomposition includes three stages: adaptive frequency domain decomposition, frequency domain information perception, and frequency domain information fusion. Among them, in the adaptive frequency domain decomposition stage, a learnable frequency domain filter is used to adaptively enhance specific frequency components, and the enhanced features are dynamically decomposed into high-frequency and low-frequency components to highlight the target information and suppress background interference; in the frequency domain information perception stage, a spatial sparse attention network is adopted to effectively extract the effective information for the characteristics of high-frequency and low-frequency components respectively, so that the model can adapt to the target expression differences at different frequencies; in the frequency domain information fusion stage, a spatial information aggregation module is designed for the efficient fusion of high-frequency and low-frequency features to enhance the integrity of the target features. On this basis, the output of the spatial information aggregation module passes through a residual block (Res Block) to obtain the finally predicted target mask, denoted as Output; at the same time, the high-frequency features extracted in the frequency domain information perception stage pass through a residual block to obtain the boundary of the predicted target mask, denoted as Edge.

[0010] The loss function in the training process of this network model consists of two parts: BceLoss and EdgeLoss. Among them, BceLoss is used to constrain the final small target detection result, and binary cross-entropy loss is used to measure the gap between the predicted target mask and the true mask; EdgeLoss is used to constrain the accurate extraction of high-frequency information, and the true boundary information of the true target mask is extracted by the Canny operator to effectively constrain the extracted high-frequency components, and binary cross-entropy loss is also used.

[0011] Furthermore, in the adaptive frequency domain decomposition stage, a residual block is first used to perform preliminary feature extraction on the input infrared image to obtain rich low-level semantic information. Subsequently, an adaptive frequency domain selection module is used to perform frequency domain screening on the features to enhance the expression of the key information frequency components, while suppressing the invalid or interfering frequency components, thereby improving the distinguishability between the target and the background. Finally, through the frequency domain decomposition module, the enhanced features are dynamically divided into high-frequency and low-frequency components for more accurate extraction of the key frequency domain features of the target object in the subsequent stage.

[0012] The adaptive frequency domain selection module uses the two-dimensional discrete cosine transform (2D DCT) as a selective filtering mechanism to adaptively modulate the importance of different frequency components, thereby enhancing the spatial information recognition ability. Specifically, for the input feature , its 2D DCT representation is: . Among them , represent the spatial dimensions, , , represents a set of basis functions of the 2D DCT.

[0013] For a specific frequency component , the input The feature representation after frequency-domain transformation is denoted as . Then, the features after frequency-domain transformation in each group are concatenated (Concat) along the channel dimension, and global average pooling (AvgPooling) and global max pooling (MaxPooling) operations are respectively applied to the concatenated results for channel dimension modulation. After adding the output results of the two, channel information integration and reshaping are performed through two 1×1 convolutional layers.

[0014] Finally, the integrated features are normalized through the Sigmoid function and multiplied element-wise with the input features to achieve weighted adjustment of different frequency components, thereby adaptively strengthening the contribution of key frequency components and suppressing invalid or redundant frequency-domain information.

[0015] The frequency-domain decomposition module is based on the features weighted in the frequency domain, denoted as . First, downsampling is performed through a convolution with a stride of 4 and a kernel size of 4×4 to reduce the spatial resolution and capture low-frequency information, and then upsampling is performed to simply filter out high-frequency information and obtain low-frequency features, denoted as .

[0016] Subsequently, the input features are processed through a convolutional layer with a stride of 1 and a kernel size of 1×1, and the low-frequency features are subtracted to dynamically extract the high-frequency features of the effective features, denoted as .

[0017] Furthermore, the frequency-domain information perception stage includes two branches composed of spatial sparse attention networks for low-frequency and high-frequency features respectively. Since the spatial sparsity of low-frequency and high-frequency features is different, different sparsity level parameters and are set for the two branches.

[0018] Given the input feature , the spatial sparse attention network first performs window partitioning in the spatial dimension, and the size of each window is . Therefore can be divided into non-overlapping image patches , where . Then, for each feature patch , considering the spatial sparsity, using the ordinary self-attention calculation method will inevitably lead to waste of computing resources. Therefore, a reshape operation is first performed on it and it is divided into pixel blocks, sort them according to the numerical values of each pixel block, and select the top pixel blocks containing valid information for subsequent self-attention calculation, where , represents the sparsity, thus effectively reducing the model calculation complexity.

[0019] Next, map the selected valid information blocks to query: , key: , and value: respectively, and perform self-attention calculation: , where represents the feature after attention weighting. Add it to the position encoding information, and at the same time, obtain the output of the self-attention network through a linear projection operation. Finally, rearrange the output according to the original spatial position and obtain the final output result through a reshape operation, thus realizing efficient feature perception based on sparse spatial attention.

[0020] Furthermore, in the frequency domain information fusion stage, through the designed spatial information aggregation module, high-frequency and low-frequency features are efficiently fused, and the integrity of the target features is enhanced.

[0021] The spatial information aggregation module first records the high-frequency and low-frequency features as and respectively, concatenate them along the channel dimension and perform preliminary fusion feature extraction through a 3×3 convolution. Then, extract the guiding information of the fusion feature along the row and column directions respectively through two parallel convolutions with kernel sizes of 1×5 and 5×1 to generate feature maps and representing different direction information, and add the two to achieve feature fusion and generate the corresponding attention weights through the Sigmoid function. Finally, perform an element-wise multiplication operation with the low-frequency feature , and further strengthen the effective fusion of high-frequency and low-frequency information through a residual connection, and finally obtain the fused feature representation, denoted as . BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a schematic diagram of the infrared small target detection model based on frequency domain decomposition of the invention;

[0023] Figure 2 is a schematic diagram of the adaptive frequency domain selection module and the frequency domain decomposition module of the invention;

[0024] Figure 3 is a schematic diagram of the spatial sparse attention network module of the invention;

[0025] Figure 4 Schematic diagram of the spatial information aggregation module of the invention;

[0026] Figure 5 Schematic diagram of the comparison results between the invention embodiment and the prior art on the NUAA dataset. Detailed implementation manners

[0027] Various exemplary embodiments, features and methods of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.

[0028] In addition, for better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0029] In this embodiment, an infrared small target detection model is provided. As Figure 1 shown, this model includes three stages: adaptive frequency domain decomposition, frequency domain information perception, and frequency domain information fusion. Among them, as Figure 2 shown, in the adaptive frequency domain decomposition stage, a learnable frequency domain filter is used to adaptively enhance specific frequency components, and the enhanced features are dynamically decomposed into high-frequency and low-frequency components to highlight target information and suppress background interference; as Figure 3 shown, in the frequency domain information perception stage, a spatial sparse attention network is adopted to effectively extract effective information for the characteristics of high-frequency and low-frequency components respectively, so that the model can adapt to the target expression differences at different frequencies; as Figure 4 shown, in the frequency domain information fusion stage, a spatial information aggregation module is designed for efficient fusion of high-frequency and low-frequency features to enhance the integrity of target features. On this basis, the output of the spatial information aggregation module passes through a residual block (Res Block) to obtain the finally predicted target mask, denoted as Output; at the same time, the high-frequency features extracted in the frequency domain information perception stage pass through a residual block to obtain the boundary of the predicted target mask, denoted as Edge. This residual block consists of three groups of convolutional layers and activation functions. The first two convolutional layers use 3×3 convolutional kernels, with a stride of 1, padding of 1, and include bias terms. After these two convolutional layers, a residual connection is set to add the input features to the output of the residual block. The last convolutional layer uses 1×1 convolutional kernels, with a stride of 1, padding of 0, and includes bias terms. To enhance the non-linear fitting ability of the network, a ReLU activation function is applied after each convolutional layer.

[0030] The loss function of the training process of this network model consists of two parts, BceLoss and EdgeLoss. Among them, BceLoss is used to constrain the final detection results of small targets, and binary cross-entropy loss is adopted to measure the gap between the predicted target mask and the true mask; EdgeLoss is used to constrain the accurate extraction of high-frequency information. The true boundary information of the true target mask extracted by the Canny operator is used to effectively constrain the extracted high-frequency components, and binary cross-entropy loss is also adopted.

[0031] In the adaptive frequency-domain decomposition stage, first, a residual block is used to perform preliminary feature extraction on the input infrared image to obtain rich low-level semantic information. Subsequently, the adaptive frequency-domain selection module is used to perform frequency-domain screening on the features, enhance the expression of the key information frequency components, and suppress the invalid or interfering frequency components simultaneously, thereby improving the distinguishability between the target and the background. Finally, through the frequency-domain decomposition module, the enhanced features are dynamically divided into high-frequency and low-frequency components for more accurate extraction of the key frequency-domain features of the target object in the subsequent stage.

[0032] The adaptive frequency-domain selection module uses the two-dimensional discrete cosine transform (2D DCT) as a selective filtering mechanism to adaptively modulate the importance of different frequency components, thereby enhancing the spatial information recognition ability. Specifically, for the input feature , its 2D DCT representation is: . Among them , represent the spatial dimensions, , , represent a set of basis functions of the 2D DCT. For a specific frequency component , the feature representation of the input after frequency-domain transformation is . Then, the features after frequency-domain transformation for each group are concatenated (Concat) along the channel dimension, and global average pooling (AvgPooling) and global maximum pooling (MaxPooling) operations are respectively applied to the concatenated results for channel dimension modulation. After adding the output results of the two, two 1×1 convolutional layers are used for channel information integration and reshaping.

[0033] Finally, the integrated features are normalized through the Sigmoid function and multiplied element-wise with the input features to achieve weighted adjustment of different frequency components, thereby adaptively strengthening the contribution of the key frequency components and suppressing the invalid or redundant frequency-domain information.

[0034] The frequency-domain decomposition module is based on the features weighted in the frequency domain, denoted as , first, perform downsampling operations through convolution with a stride of 4 and a convolution kernel size of 4×4 to reduce the spatial resolution and capture low-frequency information, and then perform upsampling operations to simply filter out high-frequency information and obtain low-frequency features, denoted as . Subsequently, process the input features through a convolution layer with a stride of 1 and a convolution kernel size of 1×1, and subtract the low-frequency features , dynamically extract the high-frequency features of the effective features, denoted as .

[0035] The frequency-domain information perception stage contains two branches composed of spatial sparse attention networks, which are used for low-frequency and high-frequency features respectively. Since the spatial sparsity of low-frequency and high-frequency features is different, different sparsity level parameters and are set for the two branches. Here, and are set to 0.3 and 0.1 respectively.

[0036] Given the input feature , the spatial sparse attention network first performs window partitioning in the spatial dimension, and the size of each window is , so can be divided into non-overlapping image patches , where . Then, for each feature patch , considering the spatial sparsity, using the ordinary self-attention calculation method will inevitably lead to a waste of computing resources. Therefore, first perform a reshape operation on it, divide it into pixel blocks, sort them according to the numerical size of each pixel block, and select the first pixel blocks containing valid information for subsequent self-attention calculation, where , represents the sparsity, thus effectively reducing the model calculation complexity.

[0037] Next, map the selected valid information blocks to query: , key: , and value: respectively, and perform self-attention calculation: , where It represents the features after attention weighting, which are added to the position encoding information and, at the same time, passed through a linear projection operation to obtain the output of the self-attention network. Finally, the output is rearranged according to the original spatial positions and the final output result is obtained through a reshape operation, thus realizing efficient feature perception based on sparse spatial attention.

[0038] In the frequency-domain information fusion stage, through the designed spatial information aggregation module, efficient fusion of high- and low-frequency features is achieved, and the integrity of the target features is enhanced.

[0039] The spatial information aggregation module first denotes the high-frequency and low-frequency features as and , respectively. They are concatenated along the channel dimension and preliminary fusion feature extraction is performed through a 3×3 convolution. Then, through two parallel convolutions with kernel sizes of 1×5 and 5×1, the guiding information of the fusion features is extracted along the row and column directions respectively to generate feature maps and representing different direction information, and the two are added to achieve feature fusion and the corresponding attention weights are generated through the Sigmoid function. Finally, through element-wise multiplication with the low-frequency feature , and through residual connection, the effective fusion of high- and low-frequency information is further strengthened, and finally the fused feature representation is obtained, denoted as .

[0040] We finally selected methods such as IPI, NRAM, PSTNN, MDvsFA, ACM, and ALCNet to verify the advancement and effectiveness of the present invention on the NUAA dataset. The visualization results in typical scenarios are as shown in Figure 5 . The present invention uses common evaluation metrics such as IoU (Intersection over Union), nIoU (Normalized Intersection over Union), and Pd, and conducts quantitative comparison and analysis with several other infrared small target detection methods. Among them, IoU represents the intersection over union, with a value range from 0 to 1, and the closer the value is to 1, the better the detection effect; nIoU is the normalized form of IoU, also with a value range from 0 to 1, and the closer the value is to 1, the better the detection effect; Pd represents the detection rate, which refers to the ratio of correctly detected targets to all targets, with a value range from 0 to 100%, and the closer the value is to 100%, the better the detection effect. The experimental results are shown in Table 1, and the visualization results are as shown in Figure 5 .

[0041] Method Iou nIoU Pd IPI 2.62 4.16 84.40 NRAM 45.68 55.49 85.32 PSTNN 51.95 62.66 82.57 MDvsFA 60.28 58.16 76.15 ACM 71.96 71.05 96.25 ALCNet 73.43 71.44 97.23 The present invention 75.31 73.24 97.83

[0042] Through the comparison of experimental data and visualization results, the results show that the method of the present invention is significantly superior to the comparative algorithm in terms of detection effect, verifying its excellent performance.

Claims

1. A method for detecting small infrared targets based on frequency domain decomposition, characterized in that: The following steps are involved: S1. Build a network model based on frequency domain decomposition. The model is divided into three stages: adaptive frequency domain decomposition, frequency domain information perception, and frequency domain information fusion. It is used to dynamically decompose the input infrared image into high-frequency and low-frequency components, extract effective information from the decomposed high- and low-frequency components, and fuse the features of different frequency components to finally obtain a clean target mask. S2. Construct an adaptive frequency domain selection module and a frequency domain decomposition module, adaptively enhance different frequency components of input features through learnable discrete cosine transform, and dynamically decompose the enhanced features into high-frequency and low-frequency components, so that the subsequent steps can more easily capture the key frequency domain features of the target object; S3. Design a spatial sparse attention network. This module establishes spatial long-distance dependencies by efficiently calculating the similarity of non-local spaces in input features according to the different sparsity of image space dimensions, thereby significantly improving the network's perception of valid targets. S4. Design a spatial information aggregation module to achieve effective integration of high- and low-frequency features through bidirectional attention information extraction between rows and columns.

2. The infrared small target detection method based on frequency domain decomposition according to claim 1 is characterized in that: The step S1 comprises the following steps: S1.1 The network model based on frequency domain decomposition includes three stages: adaptive frequency domain decomposition, frequency domain information perception, and frequency domain information fusion. Among them, specific frequency components are enhanced by learnable filters and divided into high frequency and low frequency to highlight the target and suppress the background. In the frequency domain information perception stage, a spatial sparse attention network is used to extract effective information of high and low frequencies respectively to adapt to the expression differences of targets at different frequencies. In the frequency domain information fusion stage, a spatial information aggregation module is designed to efficiently fuse high and low frequency features to enhance the integrity of target features. The final output is passed through the residual block to obtain the target mask (Output); the high frequency features are passed through the residual block to obtain the boundary (Edge). S1.2 The loss function of this model consists of two parts: BceLoss and EdgeLoss. BceLoss is used to constrain the final small target detection results and measure the gap between the predicted target mask and the true mask; EdgeLoss is used to constrain the accurate extraction of high-frequency information. Both use binary cross entropy loss.

3. The infrared small target detection method based on frequency domain decomposition according to claim 1 is characterized in that: The step S2 comprises the following steps: In the S2.1 adaptive frequency domain decomposition stage, the low-level semantic information of the input image is extracted through the residual block; then the adaptive frequency domain selection module is used to enhance the key frequency components and suppress the interference information; finally, the frequency domain decomposition module dynamically divides the features into high-frequency and low-frequency components; S2.2 Adaptive frequency domain selection module uses two-dimensional discrete cosine transform (DCT) as a selective filtering mechanism to improve spatial recognition ability; for input features , its two-dimensional discrete cosine transform is expressed as: ;in , Indicates the size of the space, , , A set of basis functions representing the two-dimensional discrete cosine transform; The S2.3 frequency domain decomposition module captures low-frequency information through convolution operations based on weighted features and dynamically extracts high-frequency features, providing a basis for subsequent fine extraction.

4. The infrared small target detection method based on frequency domain decomposition according to claim 1 is characterized in that: The step S3 comprises the following steps: The frequency domain information perception stage S3.1 consists of two spatial sparse attention network branches, which process low-frequency and high-frequency features respectively. Considering the difference in spatial sparsity between the two, different sparsity parameters are set respectively. and ; S3.2 Given input features , the spatial sparse attention network first divides the window in the spatial dimension, and the size of each window is ,therefore Divided into non-overlapping image patches, where ; Then, for each feature block Considering the sparsity of the space, we first reshape it and divide it into pixel blocks, sort them by value, and select the top The pixel blocks containing valid information are used for subsequent self-attention calculations, where , Indicates sparsity; The selected pixel blocks are mapped to queries respectively: , key: , and value: The three items are calculated by self-attention: ,in It represents the features after attention weighting, which are added to the position code and linearly mapped as the attention result; finally, the output is rearranged and reshaped according to the original spatial position to achieve efficient feature perception under sparse spatial attention.

5. The infrared small target detection method based on frequency domain decomposition according to claim 1 is characterized in that: The step S4 comprises the following steps: S4.1 frequency domain information fusion stage uses spatial information aggregation module to fuse high and low frequency features to enhance the integrity of target representation; S4.2 Spatial information aggregation module first records the high-frequency and low-frequency features as and , splicing along the channel dimension and performing preliminary fusion feature extraction through a 3×3 convolution; then, two parallel convolutions with kernel sizes of 1×5 and 5×1 are used to extract guidance information from the fusion features along the row and column directions respectively and generate feature maps representing information in different directions and , and add the two to achieve feature fusion and generate the corresponding attention weights through the Sigmoid function; finally, by combining with the low-frequency features Perform element-by-element point multiplication operation, and further enhance the effective fusion of high-frequency and low-frequency information through residual connection, and finally obtain the fused feature representation, which is recorded as .

Citation Information

Patent Citations

  • Infrared weak and small target detection method based on wavelet guidance

    CN118587507A

  • Infrared small target detection network based on feature enhancement

    CN118736214A

  • Infrared small target detection method based on diffusion model

    CN119313913A

  • Air infrared weak and small target detection method based on space-frequency domain feature fusion

    CN119478375A

Cited By

  • Infrared sequence weak and small target detection method and device based on frequency domain information

    CN120472150A

  • Infrared sequence weak and small target detection method and device based on frequency domain information

    CN120472150B

  • Signal detection and processing method suitable for distributed optical fiber acoustic sensing

    CN120974261A

  • A signal detection and processing method suitable for distributed fiber optic acoustic sensing

    CN120974261B