An infrared small target detection method based on frequency domain decomposition

Through the combination of frequency domain decomposition and sparse attention network, the problem of difficulty in distinguishing between target and background in infrared small object detection is solved, and high-precision detection and robustness improvement in complex environments are achieved.

CN120163973BActive Publication Date: 2025-07-18NINGBO UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510645481.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-07-18
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The existing infrared small object detection methods are difficult to effectively distinguish between targets and backgrounds under complex background and low contrast conditions, and the high-frequency and low-frequency information processing lacks flexibility, resulting in low detection accuracy and poor robustness.

Method used

Using a detection method based on frequency domain decomposition, through adaptive frequency domain decomposition and sparse attention network, infrared images are dynamically decomposed into high-frequency and low-frequency components, target information is enhanced and background noise is suppressed, and feature fusion is combined with spatial information aggregation module to optimize target detection performance.

Benefits of technology

In complex backgrounds and low contrast environments, the detection accuracy and robustness of infrared small targets are significantly improved, and the accuracy of target detection and real-time tracking capabilities are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163973B_ABST
    Figure CN120163973B_ABST
Patent Text Reader

Abstract

The present invention proposes an infrared small target detection method based on frequency domain decomposition, aiming to solve the insufficient detection performance of traditional spatial domain methods under complex backgrounds, low contrast, and noise interference conditions. In view of the problem that infrared small targets usually have small spatial distributions and high-frequency characteristics, we decompose the input image into high-frequency and low-frequency components through an adaptive frequency domain decomposition technique, combine the frequency domain information perception and spatial information aggregation modules, and use a sparse attention mechanism to weight and extract the frequency domain features to dynamically adjust the contributions of different frequency components. Through efficient feature fusion, the expression of target information is optimized, especially in low-contrast and complex background environments, improving the detection accuracy and robustness. The experimental verification results on the NUAA dataset show that compared with existing mainstream methods, the present invention shows significant advantages in multiple evaluation indicators, verifying its effectiveness and superiority in infrared small target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of infrared target detection, and in particular relates to research on a small infrared target detection method based on frequency domain decomposition. Background Art

[0002] Compared with traditional RGB imaging technology, infrared imaging uses its detection capability beyond the visible spectrum to effectively capture the infrared radiation signal of the target at night and in bad weather conditions, realizing target detection and identification in all weather and all environments. This feature makes it play an indispensable role in natural disaster warning, military reconnaissance and precision navigation. However, in complex environments, the accurate detection of small infrared targets still faces many challenges. These small targets are usually small in size, low in contrast and low in signal-to-noise ratio. Due to the long imaging distance and the limitation of device resolution, these targets account for a very low proportion in the image, usually not more than 0.2%, and are easily masked by background noise. In addition, affected by the uneven distribution of environmental thermal radiation, the target outline is blurred and lacks texture information, resulting in bottlenecks in target enhancement, background suppression and robust feature extraction for detection methods based on traditional methods. Therefore, how to effectively extract target features, enhance contrast, and realize accurate detection and real-time tracking of small infrared targets under high noise and high dynamic backgrounds is still a key technical problem that needs to be solved urgently.

[0003] However, due to the lack of color information and texture features of small infrared targets, existing infrared small target detection methods mainly focus on the existence of targets and attempt to extract features by building deeper neural networks. However, as the network deepens, the detailed features of small targets are easily lost, while in segmentation tasks, edge details and contour information are crucial. Summary of the invention

[0004] In view of this, the object of the present invention is to provide a detection method based on frequency domain decomposition, which is applied to the field of infrared small target detection. Most of the existing infrared small target detection methods rely on traditional spatial domain feature extraction methods. Although certain effects can be achieved in some simple scenarios, in complex backgrounds, especially in the case of low contrast and strong noise interference, problems such as difficulty in effectively distinguishing the target from the background and low detection accuracy are often faced. In addition, many methods lack flexibility in processing high-frequency and low-frequency information, and fail to make full use of the discrimination ability of different frequency domain components for the target, resulting in the neglect of the detailed information of small targets or being submerged by background noise. This method introduces a frequency domain analysis mechanism. First, it adaptively decomposes the input infrared image into high-frequency and low-frequency components, giving full play to the expression ability of high-frequency components for the boundary details of the target and the suppression effect of low-frequency components on background information. Therefore, the present invention proposes an infrared small target detection method based on frequency domain decomposition. Through frequency domain weighting and information perception, the method can automatically enhance effective frequency domain components while suppressing invalid noise information, thereby improving the robustness and accuracy of target detection. In addition, by using a spatial information aggregation module, high-frequency and low-frequency features are efficiently fused, further optimizing the expression of target information. Especially in complex backgrounds and low-contrast environments, the detection performance of infrared small targets can be significantly improved. Its core lies in the following four parts:

[0005] S1. Construct a network model based on frequency domain decomposition. This model is divided into three stages: adaptive frequency domain decomposition, frequency domain information perception, and frequency domain information fusion, which are used to dynamically decompose the input infrared image into high-frequency and low-frequency components, extract effective information from the decomposed high-frequency and low-frequency components, and fuse the features of different frequency components, and finally obtain a clean target mask;

[0006] S2. Construct an adaptive frequency domain selection module and a frequency domain decomposition module. Through a learnable discrete cosine transform, different frequency components of the input features are adaptively enhanced, and the enhanced features are dynamically decomposed into high-frequency and low-frequency components, so as to facilitate subsequent steps to more easily capture the key frequency domain features of the target object;

[0007] S3. Design a spatial sparse attention network. According to the different spatial sparsity of the image, this module establishes a long-distance spatial dependence relationship by efficiently calculating the similarity of the non-local space in the input features, thereby significantly improving the network's perception ability of effective targets;

[0008] S4. Design a spatial information aggregation module. Through bidirectional attention information extraction of rows and columns, effective integration of high-frequency and low-frequency features is achieved.

[0009] Furthermore, the network model based on frequency domain decomposition includes three stages: adaptive frequency domain decomposition, frequency domain information perception, and frequency domain information fusion. Among them, in the adaptive frequency domain decomposition stage, a learnable frequency domain filter is used to adaptively enhance specific frequency components, and the enhanced features are dynamically decomposed into high-frequency and low-frequency components to highlight target information and suppress background interference; in the frequency domain information perception stage, a spatial sparse attention network is adopted to effectively extract effective information for the characteristics of high- and low-frequency components respectively, enabling the model to adapt to the target expression differences at different frequencies; in the frequency domain information fusion stage, a spatial information aggregation module is designed for the efficient fusion of high- and low-frequency features to enhance the integrity of target features. On this basis, the output of the spatial information aggregation module passes through a residual block (Res Block) to obtain the finally predicted target mask, denoted as Output; at the same time, the high-frequency features extracted in the frequency domain information perception stage pass through a residual block to obtain the boundary of the predicted target mask, denoted as Edge.

[0010] The loss function in the training process of this network model consists of two parts: BceLoss and EdgeLoss. Among them, BceLoss is used for the constraint of the final small target detection result, and binary cross-entropy loss is used to measure the gap between the predicted target mask and the true mask; EdgeLoss is used to constrain the accurate extraction of high-frequency information, and the true boundary information of the true target mask is extracted through the Canny operator to effectively constrain the extracted high-frequency components, and binary cross-entropy loss is also used.

[0011] Furthermore, in the adaptive frequency domain decomposition stage, a residual block is first used to perform preliminary feature extraction on the input infrared image to obtain rich low-level semantic information. Subsequently, an adaptive frequency domain selection module is used to perform frequency domain screening on the features, enhance the expression of key information frequency components, and at the same time suppress invalid or interfering frequency components, thereby improving the distinguishability between the target and the background. Finally, through the frequency domain decomposition module, the enhanced features are dynamically divided into high-frequency and low-frequency components for more accurate extraction of the key frequency domain features of the target object in the subsequent stage.

[0012] The adaptive frequency domain selection module uses the two-dimensional discrete cosine transform (2D DCT) as a selective filtering mechanism to adaptively modulate the importance of different frequency components, thereby enhancing the spatial information recognition ability. Specifically, for the input feature , its 2D DCT representation is: . Among them , represent the spatial dimensions, , , represent a set of basis functions of the 2D DCT.

[0013] For a specific frequency component , the input The feature representation after frequency-domain transformation is denoted as . Then, the features after frequency-domain transformation in each group are concatenated (Concat) along the channel dimension, and global average pooling (AvgPooling) and global max pooling (MaxPooling) operations are respectively applied to the concatenated result for channel dimension modulation. After adding the output results of the two, two 1×1 convolutional layers are used to integrate and reshape the channel information.

[0014] Finally, the integrated features are normalized through the Sigmoid function and element-wise multiplied with the input features to achieve weighted adjustment of different frequency components, thereby adaptively strengthening the contribution of key frequency components and suppressing invalid or redundant frequency-domain information.

[0015] The frequency-domain decomposition module is based on the features weighted in the frequency domain, denoted as . First, a downsampling operation is performed through a convolution with a stride of 4 and a kernel size of 4×4 to reduce the spatial resolution and capture low-frequency information, and then an upsampling operation is used to simply filter out high-frequency information and obtain low-frequency features, denoted as .

[0016] Subsequently, the input features are processed through a convolutional layer with a stride of 1 and a kernel size of 1×1, and the low-frequency features are subtracted to dynamically extract the high-frequency features of the effective features, denoted as .

[0017] Furthermore, the frequency-domain information perception stage contains two branches composed of spatial sparse attention networks for low-frequency and high-frequency features respectively. Since the spatial sparsity of low-frequency and high-frequency features is different, different sparsity level parameters and are set for the two branches.

[0018] Given the input feature , the spatial sparse attention network first performs window partitioning in the spatial dimension, and the size of each window is , so can be divided into non-overlapping image patches , where . Then, for each feature patch , considering the spatial sparsity, using the ordinary self-attention calculation method will inevitably lead to waste of computing resources. Therefore, a reshape operation is first performed on it and divided into pixel blocks, sort them according to the numerical values of each pixel block, and select the top pixel blocks containing valid information for subsequent self-attention calculation, where , represents the sparsity, thus effectively reducing the computational complexity of the model.

[0019] Next, map the selected valid information blocks to query: , key: , and value: respectively, and perform self-attention calculation: , where represents the features after attention weighting. Add it to the position encoding information, and at the same time, obtain the output of the self-attention network through a linear projection operation. Finally, rearrange the output according to the original spatial position and obtain the final output result through a reshape operation, thus realizing efficient feature perception based on sparse spatial attention.

[0020] Furthermore, in the frequency domain information fusion stage, through the designed spatial information aggregation module, efficient fusion of high-frequency and low-frequency features is achieved, and the integrity of target features is enhanced.

[0021] The spatial information aggregation module first records the high-frequency and low-frequency features as and respectively, concatenate them along the channel dimension and perform preliminary fusion feature extraction through a 3×3 convolution. Then, extract guiding information for the fusion features along the row and column directions respectively through two parallel convolutions with kernel sizes of 1×5 and 5×1 to generate feature maps and representing different direction information, and add the two to achieve feature fusion and generate corresponding attention weights through the Sigmoid function. Finally, perform an element-wise multiplication operation with the low-frequency feature , and further strengthen the effective fusion of high-frequency and low-frequency information through a residual connection, and finally obtain the fused feature representation, denoted as . Brief Description of the Drawings

[0022] Figure 1 is a schematic diagram of the infrared small target detection model based on frequency domain decomposition of the invention;

[0023] Figure 2 is a schematic diagram of the adaptive frequency domain selection module and the frequency domain decomposition module of the invention;

[0024] Figure 3 is a schematic diagram of the spatial sparse attention network module of the invention;

[0025] Figure 4 Schematic diagram of the spatial information aggregation module of the invention;

[0026] Figure 5 Schematic diagram of the comparison results between the invention embodiment and the prior art on the NUAA dataset. Detailed implementation manners

[0027] Various exemplary embodiments, features and methods of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.

[0028] In addition, for better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0029] In this embodiment, an infrared small target detection model is provided. As Figure 1 shown, this model includes three stages: adaptive frequency domain decomposition, frequency domain information perception, and frequency domain information fusion. Among them, as Figure 2 shown, in the adaptive frequency domain decomposition stage, a learnable frequency domain filter is used to adaptively enhance specific frequency components, and the enhanced features are dynamically decomposed into high-frequency and low-frequency components to highlight target information and suppress background interference; as Figure 3 shown, in the frequency domain information perception stage, a spatial sparse attention network is adopted to effectively extract effective information for the characteristics of high-frequency and low-frequency components respectively, so that the model can adapt to the target expression differences at different frequencies; as Figure 4 shown, in the frequency domain information fusion stage, a spatial information aggregation module is designed for efficient fusion of high-frequency and low-frequency features to enhance the integrity of target features. On this basis, the output of the spatial information aggregation module passes through a residual block (Res Block) to obtain the finally predicted target mask, denoted as Output; at the same time, the high-frequency features extracted in the frequency domain information perception stage pass through a residual block to obtain the boundary of the predicted target mask, denoted as Edge. This residual block consists of three groups of convolutional layers and activation functions. The first two convolutional layers use 3×3 convolutional kernels, with a stride of 1, padding of 1, and include bias terms. After these two convolutional layers, a residual connection is set to add the input features to the output of the residual block. The last convolutional layer uses 1×1 convolutional kernels, with a stride of 1, padding of 0, and includes bias terms. To enhance the non-linear fitting ability of the network, a ReLU activation function is applied after each convolutional layer.

[0030] The loss function in the training process of this network model consists of two parts: BceLoss and EdgeLoss. Among them, BceLoss is used to constrain the final detection results of small targets, and binary cross-entropy loss is adopted to measure the gap between the predicted target mask and the true mask; EdgeLoss is used to constrain the accurate extraction of high-frequency information. The true boundary information of the true target mask extracted by the Canny operator is used to effectively constrain the extracted high-frequency components, and binary cross-entropy loss is also adopted.

[0031] In the adaptive frequency-domain decomposition stage, a residual block is first used to perform preliminary feature extraction on the input infrared image to obtain rich low-level semantic information. Subsequently, the adaptive frequency-domain selection module is used to perform frequency-domain screening on the features, enhance the expression of the key information frequency components, and suppress the invalid or interfering frequency components simultaneously, so as to improve the distinguishability between the target and the background. Finally, through the frequency-domain decomposition module, the enhanced features are dynamically divided into high-frequency and low-frequency components for more accurate extraction of the key frequency-domain features of the target object in the subsequent stage.

[0032] The adaptive frequency-domain selection module uses the two-dimensional discrete cosine transform (2D DCT) as a selective filtering mechanism to adaptively modulate the importance of different frequency components, thereby enhancing the spatial information recognition ability. Specifically, for the input feature , its 2D DCT representation is: . Among them , represent the spatial dimensions, , , represent a set of basis functions of the 2D DCT. For a specific frequency component , the feature representation of the input after frequency-domain transformation is . Then, the features after frequency-domain transformation in each group are concatenated (Concat) along the channel dimension, and global average pooling (AvgPooling) and global maximum pooling (MaxPooling) operations are respectively applied to the concatenated results for channel dimension modulation. After adding the output results of the two, two 1×1 convolutional layers are used to integrate and reshape the channel information.

[0033] Finally, the integrated features are normalized through the Sigmoid function and multiplied element-wise with the input features to achieve weighted adjustment of different frequency components, thereby adaptively strengthening the contribution of the key frequency components and suppressing the invalid or redundant frequency-domain information.

[0034] The frequency-domain decomposition module is based on the features weighted in the frequency domain, denoted as , first, a downsampling operation is performed through a convolution with a stride of 4 and a convolution kernel size of 4×4 to reduce the spatial resolution and capture low-frequency information, and then an upsampling operation is used to simply filter out high-frequency information and obtain low-frequency features, denoted as . Subsequently, the input features are processed through a convolution layer with a stride of 1 and a convolution kernel size of 1×1, and the low-frequency features are subtracted , and the high-frequency features of the effective features are dynamically extracted, denoted as .

[0035] The frequency-domain information perception stage contains two branches composed of spatial sparse attention networks, which are used for low-frequency and high-frequency features respectively. Since the spatial sparsity of low-frequency and high-frequency features is different, different sparsity level parameters and are set for the two branches. Here, and are set to 0.3 and 0.1 respectively.

[0036] Given the input features , the spatial sparse attention network first performs window partitioning in the spatial dimension, and the size of each window is , so can be divided into non-overlapping image patches , where . Then, for each feature patch , considering the spatial sparsity, using the ordinary self-attention calculation method will inevitably lead to a waste of computing resources. Therefore, a reshape operation is first performed on it, dividing it into pixel patches, sorting them according to the numerical size of each pixel patch, and selecting the first pixel patches containing valid information for subsequent self-attention calculation, where , represents the sparsity, thus effectively reducing the model's computational complexity.

[0037] Next, the selected valid information blocks are respectively mapped to query: , key: , and value: three items and perform self-attention calculation: , where It represents the features after attention weighting. Add it to the position encoding information, and at the same time obtain the output of the self-attention network through a linear projection operation. Finally, rearrange the output according to the original spatial positions and obtain the final output result through a reshape operation, thus realizing efficient feature perception based on sparse spatial attention.

[0038] In the frequency domain information fusion stage, through the designed spatial information aggregation module, efficient fusion of high-frequency and low-frequency features is achieved, and the integrity of target features is enhanced.

[0039] The spatial information aggregation module first denotes the high-frequency and low-frequency features as and , concatenate them along the channel dimension and perform preliminary fusion feature extraction through a 3×3 convolution. Then, through two parallel convolutions with kernel sizes of 1×5 and 5×1 respectively, extract guiding information from the fusion features along the row and column directions and generate feature maps and representing different direction information, and perform an addition operation on the two to achieve feature fusion and generate corresponding attention weights through the Sigmoid function. Finally, perform an element-wise multiplication operation with the low-frequency feature , and further strengthen the effective fusion of high-frequency and low-frequency information through a residual connection. Finally, obtain the fused feature representation, denoted as .

[0040] We finally selected methods such as IPI, NRAM, PSTNN, MDvsFA, ACM, and ALCNet to verify the advancement and effectiveness of the present invention on the NUAA dataset. The visualization results in typical scenarios are as Figure 5 shown. The present invention uses common evaluation metrics such as IoU (Intersection over Union), nIoU (Normalized Intersection over Union), and Pd, and conducts quantitative comparison and analysis with several other infrared small target detection methods. Among them, IoU represents the intersection over union, with a value range of 0 to 1. The closer the value is to 1, the better the detection effect; nIoU is the normalized form of IoU, also with a value range of 0 to 1. The closer the value is to 1, the better the detection effect; Pd represents the detection rate, which refers to the ratio of correctly detected targets to all targets, with a value range of 0 to 100%. The closer the value is to 100%, the better the detection effect. The experimental results are shown in Table 1, and the visualization results are as Figure 5 shown.

[0041] Method Iou nIoU Pd IPI 2.62 4.16 84.40 NRAM 45.68 55.49 85.32 PSTNN 51.95 62.66 82.57 MDvsFA 60.28 58.16 76.15 ACM 71.96 71.05 96.25 ALCNet 73.43 71.44 97.23 The present invention 75.31 73.24 97.83

[0042] Through the comparison of experimental data and visualization results, the results show that the method of the present invention is significantly superior to the comparative algorithm in terms of detection effect, verifying its excellent performance.

Claims

1. An infrared small target detection method based on frequency domain decomposition, characterized in that It includes the following steps: S1. Construct a network model based on frequency-domain decomposition, which is divided into three stages: adaptive frequency-domain decomposition, frequency-domain information perception, and frequency-domain information fusion, respectively used to dynamically decompose the input infrared image into high-frequency and low-frequency components, effectively extract the high- and low-frequency components decomposed, and fuse the features of different frequency components, and finally obtain a clean target mask; S2. Construct an adaptive frequency-domain selection module and a frequency-domain decomposition module, adaptively enhance different frequency components of the input features through a learnable discrete cosine transform, and dynamically decompose the enhanced features into high-frequency and low-frequency components, so as to facilitate subsequent steps to more easily capture the key frequency-domain features of the target object; S3. Design a spatial sparse attention network, which, according to the different spatial sparsity of the image, establishes a long-range spatial dependence relationship by efficiently calculating the similarity of the non-local space of the input features, thereby significantly improving the network's perception ability of effective targets; S4. Design a spatial information aggregation module to effectively integrate high- and low-frequency features through bidirectional attention information extraction of rows and columns.

2. The infrared small target detection method based on frequency domain decomposition according to claim 1, wherein, The step S1 includes the following steps: S1.1 The network model based on frequency-domain decomposition includes three stages: adaptive frequency-domain decomposition, frequency-domain information perception, and frequency-domain information fusion; among them, specific frequency components are enhanced through a learnable filter and divided into high-frequency and low-frequency to highlight the target and suppress the background; in the frequency-domain information perception stage, a spatial sparse attention network is used to extract the effective information of high- and low-frequency respectively to adapt to the target expression differences at different frequencies; in the frequency-domain information fusion stage, a spatial information aggregation module is designed to efficiently fuse high- and low-frequency features and enhance the integrity of target features; finally, the output passes through a residual block to obtain the target mask (Output); the high-frequency features pass through a residual block to obtain the boundary (Edge); S1.2 The loss function of this model consists of two parts: BceLoss and EdgeLoss, where BceLoss is used to constrain the final small target detection result and measure the gap between the predicted target mask and the true mask; EdgeLoss is used to constrain the accurate extraction of high-frequency information, and both use binary cross-entropy loss.

3. A method for detecting infrared small targets based on frequency domain decomposition according to claim 1, characterized in that, The step S2 includes the following steps: S2.1 In the adaptive frequency-domain decomposition stage, the low-level semantic information of the input image is extracted through a residual block; then, the adaptive frequency-domain selection module is used to enhance the key frequency components and suppress the interference information; finally, the features are dynamically divided into high- and low-frequency components by the frequency-domain decomposition module; The S2.2 adaptive frequency domain selection module uses two-dimensional discrete cosine transform (DCT) as the selective filtering mechanism to enhance the spatial recognition ability; for the input features , its two-dimensional discrete cosine transform is expressed as: ; where , represent the spatial dimensions, , , represent a set of basis functions of the two-dimensional discrete cosine transform; S2.3 The frequency-domain decomposition module captures the low-frequency information through convolution operations based on the weighted features and dynamically extracts the high-frequency features, providing a basis for subsequent fine extraction.

4. A method for detecting infrared small targets based on frequency domain decomposition according to claim 1, characterized in that The step S3 includes the following steps: The frequency-domain information perception stage in S3.1 consists of two spatially sparse attention network branches, which process low-frequency and high-frequency features respectively; considering the differences in their spatial sparsity, different sparsity parameters are set respectively and ; S3.2 Given input features , the spatial sparse attention network first performs window partitioning in the spatial dimension, and the size of each window is , so is divided into non-overlapping image patches, where ; then, for each feature patch , considering the spatial sparsity, first perform a reshape operation on it, divide it into pixel patches, sort them according to the numerical size, and select the first pixel patches containing valid information for subsequent self-attention calculation, where , represents the sparsity; The selected pixel blocks are respectively mapped to query: , key: , and value: These three items are used to calculate self-attention: , where represents the feature after attention weighting. It is added to the positional encoding and linearly mapped to obtain the attention result. Finally, it is rearranged according to the original spatial position and reshaped to obtain the output, realizing efficient feature perception under sparse spatial attention.

5. A method for detecting infrared small targets based on frequency domain decomposition according to claim 1, characterized in that The step S4 includes the following steps: S4.1 In the frequency-domain information fusion stage, a spatial information aggregation module is used to fuse high- and low-frequency features and enhance the integrity of target representation; The S4.2 spatial information aggregation module first denotes the high-frequency and low-frequency features as and , concatenates them along the channel dimension and performs preliminary fused feature extraction through a 3×3 convolution; then, extracts guiding information for the fused features along the row and column directions respectively through two parallel convolutions with kernel sizes of 1×5 and 5×1 to generate feature maps and representing different direction information, and adds the two to achieve feature fusion and generates corresponding attention weights through the Sigmoid function; finally, performs an element-wise multiplication operation with the low-frequency feature , and further strengthens the effective fusion of high-frequency and low-frequency information through residual connection, and finally obtains the fused feature representation, denoted as .

Citation Information

Patent Citations

  • Infrared weak and small target detection method based on wavelet guidance

    CN118587507A

  • Infrared small target detection network based on feature enhancement

    CN118736214A