Infrared small target detection method based on U-Net cross-scale cross fusion dense connection network
By introducing a cross-scale cross-fusion dense connection structure and feature fusion module in the U-Net network, the problem of missing infrared small target features in the deep network is solved, and high-precision and efficient detection of infrared small target detection is achieved.
Patent Information
- Application Number
- CN202510211402.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
In the existing deep learning infrared small object detection technology, infrared small object features are easily lost in deep networks due to the increase in downsampling times, resulting in low detection accuracy and serious background interference.
A densely connected network across scale cross-fusion based on U-Net is adopted. By adding intermediate nodes and additional decoding nodes between specific levels of U-Net, and densely jumping connections between multiple nodes at each layer, combining the cross-scale cross-fusion channel attention module and the multi-input feature fusion module, we ensure that the shallow and deep feature information of small infrared targets can be effectively retained and utilized in the deep network.
Effectively balance the shallow and deep information of infrared small target images, suppress background clutter interference, improve detection performance and information utilization, and significantly improve the accuracy and efficiency of infrared small target detection.
Smart Images

Figure CN120147798A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of infrared detection, and particularly relates to an infrared small target detection method based on a U-Net cross-scale cross-fusion dense connection network. Background Art
[0002] Existing single-frame infrared small target detection methods have problems such as low detection accuracy and being greatly affected by the background. With the development of deep learning, compared with traditional target detection technologies, convolutional neural network (CNN) has shown obvious advantages in image feature representation, discriminative ability, and semantic parsing of scenes. CNN can not only automatically integrate region selection and feature classification, extract deep semantic features, but also support the end-to-end integration of feature extraction and model training, so that CNN has strong generalization ability and anti-interference ability, and its generalization ability and anti-interference ability are crucial for infrared small target detection.
[0003] However, the infrared small target detection technology based on CNN still has problems such as difficult extraction of infrared small target features and insufficient interaction between shallow and deep features. Moreover, with the increase of downsampling operations in the network, small target features often disappear in the deep network. Summary of the Invention
[0004] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide an infrared small target detection method based on a U-Net cross-scale cross-fusion dense connection network to solve the problem that infrared small target features in the deep network will be lost with the increase of the number of downsampling times in existing deep learning infrared small target detection.
[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] An infrared small target detection method based on a U-Net cross-scale cross-fusion dense connection network includes the following steps:
[0007] Step 1, input an initial infrared image into an improved U-Net. The improved U-Net adds a number of cascaded intermediate nodes between the encoding nodes and decoding nodes of the 0th layer to the 2nd layer of the original U-Net, adds cascaded additional decoding nodes after the decoding node of the 3rd layer and after the encoding node of the 4th layer respectively, and performs dense skip connections between multiple nodes in each layer;
[0008] Step 2, based on the cross-scale cross-fusion channel attention module, fuse shallow features and corresponding deep features in any one or several of the 1st layer to the 4th layer of the improved U-Net;
[0009] Step 3: Based on the multi-input feature fusion module, fuse the output features of the last node of each layer of the improved U-Net to obtain the finally output detection result.
[0010] Compared with the prior art, the present invention constructs a densely connected network for SIRST detection. This network combines shallow and deep feature fusion modules, enabling a good balance between the shallow information and deep information of infrared small target images, effectively cross-fusing the shallow information into the deep information and suppressing the interference of background clutter. In addition, the densely connected structure also provides an optimized path for feature transmission, and the multi-input feature fusion module (MIFM) ensures that each unit in the deep layer of the network can receive useful information from the previous layer and the current layer. This design greatly improves the information utilization rate and thus enhances the detection performance. Therefore, compared with the existing infrared small target detection methods based on convolutional neural networks, the detection ability of the present invention is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is the network structure diagram of U-Net.
[0012] Figure 2 is the structure diagram of LCAM module.
[0013] Figure 3 is the schematic diagram of DGCAM structure.
[0014] Figure 4 is the schematic diagram of CsHCAM module structure.
[0015] Figure 5 is the schematic diagram of MIFM network structure.
[0016] Figure 6 is the schematic diagram of the densely connected network structure based on CsHCAM.
[0017] Figure 7 is the network structure diagram of the CsH_DCNet of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The embodiments of the present invention will be described in detail below with reference to the drawings and embodiments.
[0019] There are usually two major problems in the infrared small target detection network: (1) The problem that the feature extraction ability of the network is insufficient, resulting in the inability to accurately locate the position and contour of small targets; (2) In order to improve the feature extraction ability of the network, the number of network layers is continuously deepened, but small targets often gradually disappear as the network depth increases.
[0020] Based on this, the present invention proposes a DenseConnectivity Networks based on Cross-scale Hybrid (CsH_DCNet), which ensures that the features of infrared small targets in the deep network will not be lost. Furthermore, an infrared small target detection method based on the U-Net cross-scale hybrid dense connection network is provided. By using the dense connection network U-Net, the ability of the convolutional neural network to extract multi-scale features is improved through continuous repeated upsampling and downsampling. Moreover, through the dense skip connections between multiple nodes, the network can more fully mine and utilize the shallow and deep feature information in the infrared small target image, significantly improving the information utilization rate of the network.
[0021] To verify the effectiveness of the method of the present invention, experiments are carried out using the NUAA-SIRST dataset proposed in 2022 and the NUDT-SIRST dataset proposed in 2021. The evaluation metrics are the mean intersection over union (mIoU), recall rate (R) a and false alarm rate (F). a. .
[0022] The model of this method is built in the PyTorch 1.12.0 environment, combined with the CUDA technology of the cu113 version. The experiment uses a CPU configured as i7-7700K (main frequency 4.20GHz) and a GPU of RTX3080Ti. The system memory specification is 24G, and the operating environment is Ubuntu 22.04. During the model training process, the number of iterations is set to 1600 times, the initial learning rate is set to 0.05, 8 samples are processed in each batch, and the Adagrad algorithm is used for optimization.
[0023] (1) The intersection over union (IoU) is an index to measure the coincidence degree of two bounding boxes by calculating the ratio of the overlapping area between the predicted bounding box Y pred and the ground truth bounding box Y true to the union area of them. It is a commonly used accuracy evaluation method in the field of object detection. The IoU calculation formula is shown as follows
[0024]
[0025] The position of the ground truth bounding box of the target is represented by the square Y true , and the target bounding box predicted by the model is represented by the square Y pred Y true and Y predThe Intersection over Union represents the proportion of the correctly predicted region in the total region, that is, the similarity between two bounding boxes. The larger the IoU value, the higher the similarity between the two bounding boxes and the more accurate the result. By definition, its value usually ranges from 0 to 1. To determine whether the output of an infrared small target recognition is accurate, a specific threshold is usually set, such as 0.5. If the IoU value exceeds this threshold, the recognition result is considered acceptable. In the ideal case, the predicted bounding box perfectly aligns with the true bounding box of the target, and the IoU value reaches 1.
[0026] In the exploration of infrared small target detection, the change in target size is a key consideration. The number of pixels between targets can vary greatly, with a difference range from 20 to 100 times. Such scale differences mean that the traditional Intersection over Union (IoU) metric may be more inclined to reflect the performance of the model in dealing with larger infrared targets rather than its balanced performance on targets of all sizes. Most model-based methods focus more on target recognition rather than the accuracy of segmentation, resulting in lower IoU values on larger targets.
[0027] To fully evaluate the performance of the model on infrared small target datasets and considering the special attributes of small targets, the present invention adopts a new evaluation metric: mean Intersection over Union (mIoU). This metric aims to provide a more balanced performance evaluation, especially when considering the diversity of target sizes. By adjusting the IoU, mIoU can more fairly evaluate the detection and segmentation capabilities of the model for infrared targets of different scales.
[0028] mIoU achieves a better balance between the model and data-driven methods by first calculating the IoU on each image and then averaging these values over the entire dataset to avoid the influence of large targets on small targets in the IoU scoring.
[0029] (2) Recall rate R a is an evaluation metric at the target level. It measures the number of correctly predicted targets Y true and the ratio to the total number of targets Y all . Its mathematical expression is as follows.
[0030]
[0031] If the centroid deviation of the target is greater than the predefined deviation threshold, these pixels can be regarded as mispredicted pixels. In the present invention, the predefined deviation threshold is set to 3.
[0032] (3) False alarm rate (F a ) is another target-level evaluation metric. It is used to measure the mispredicted pixels Pfalse With all image pixels P all The ratio is as follows. Its mathematical expression is as follows.
[0033]
[0034] If the centroid deviation of the target is greater than the predefined deviation threshold, these pixels can be regarded as mispredicted pixels. In the present invention, the predefined deviation threshold is set to 3.
[0035] The process of infrared dim and small target detection using the method of the present invention is as follows:
[0036] Step 1: Split the infrared dim and small target data set, and divide it into a training set, a test set and a validation set according to the ratio of 8:1:1. Obviously, the data set division ratio can be adjusted according to requirements.
[0037] Step 2: Train the split data set with the improved U-Net network, and obtain the model and evaluation index data.
[0038] Step 3: Extract the trained model to perform the final effect detection on the validation set.
[0039] Step 2 is the key step of the present invention, and its training process reflects the core idea of the infrared small target detection method of the present invention. It is specifically described as follows:
[0040] Step 21: Use U-Net as the basic network. A typical U-Net is as Figure 1 shown, which is a five-layer structure. From shallow to deep, they are the 0th, 1st, 2nd, 3rd, and 4th layers. The 0th to 3rd layers have encoding nodes and decoding nodes, and the 4th layer only has encoding nodes. The U-Net network improves the ability of the convolutional neural network to extract multi-scale features by continuously repeating upsampling and downsampling between layers, as Figure 1 shown.
[0041] In this step, the existing typical U-Net is improved, and the initial infrared image is input into the improved U-Net. The improvement in this step mainly includes two aspects. First, a number of cascaded intermediate nodes are added between the encoding nodes and decoding nodes of the 0th to 2nd layers of the original U-Net, and cascaded additional decoding nodes are added after the decoding node of the 3rd layer and after the encoding node of the 4th layer respectively. Second, dense skip connections are made between multiple nodes in each layer. After improvement, the U-Net is still a five-layer structure, and the number of channels in each layer is still twice that of the previous layer, but it has more nodes. Referring to Figure 6 and Figure 7 shown, its nodes and output features include:
[0042] The 0th layer encoding node, three intermediate nodes of the 0th layer, and the 0th layer decoding node, and the output features are successively L0,0 , L 0 ,1 , L 0,2 , L 0,3 , L 0,4 ;
[0043] The first - layer encoding node, the three intermediate nodes of the first layer, and the first - layer decoding node, with the output features being L 1,0 , L 1 ,1 , L 1,2 , L 1,3 , L 1,4 ;
[0044] The second - layer encoding node, the two intermediate nodes of the second layer, and the second - layer decoding node, with the output features being L 2,0 , L 2 ,1 , L 2,2 , L 2,3 ;
[0045] The third - layer encoding node, the third - layer decoding node, and the first additional decoding node, with the output features being L 3,0 , L 3 ,1 , L 3,2 ;
[0046] The fourth - layer encoding node, the second additional decoding node, and the third additional decoding node, with the output features being L 4,0 , L 4,1 , L 4,2 .
[0047] Step 22, in this step, the improved U - Net in the previous step is further improved. Specifically, based on the cross - scale cross - fusion channel attention module (CsHCAM), in any one or several of the first to fourth layers of the improved U - Net in the previous step, the shallow - layer features are fused with the corresponding deep - layer features.
[0048] The cross - scale cross - fusion channel attention module (CsHCAM) of the present invention mainly includes two major parts: the shallow - layer channel attention module (LCAM) and the deep - layer global channel attention module (DGCAM). The output features of the shallow - layer channel attention module and the deep - layer global channel attention module are cross - fused, which is the output of the cross - scale cross - fusion channel attention module (CsHCAM).
[0049] Specifically, the present invention uses the shallow - layer channel attention module (Low Channel Attention Modulation, LCAM) to fuse the shallow - layer features in the infrared small - target image into the deep - layer features. The calculation process of the shallow - layer channel attention module is as Figure 2As shown, for the convenience of describing the present invention, the output feature of the j-th node in the i-th layer of the improved U-Net is defined as the shallow feature L i,j , and the output feature of the j-th node in the (i + k)-th layer is defined as the deep feature L i+k,j . The calculation formula of the shallow channel attention module for L i,j is as follows:
[0050]
[0051]
[0052] In the formula, Lc(·) represents the calculation operation of the shallow channel attention module, R is the ReLU non-linear activation function, B is the batch normalization (BN) layer, is the dilated convolution with a dilation rate of d. In this embodiment, d = 2, and f 1 (·) is the micro-processing unit, and σ is the Sigmoid activation function, is the pointwise convolution.
[0053] Then, the output of the shallow channel attention module (LCAM) can be expressed as follows:
[0054]
[0055] is the element-wise multiplication operation after scale alignment.
[0056] In this module, let r be the channel compression ratio. By default, r = 4 in the embodiment. Then, for the first f 1 (·), the input of PWC 1×1 to L i,j has C channels, and the output channels are C / r. For the second f 1 (·), the input channels are C / r, and the output channels are C. The kernel sizes of the two pointwise convolutions are both 1×1.
[0057] At the same time, the present invention uses the Deep Global Channel Attention Module (DGCAM). By using global average pooling from a global perspective to compress the spatial information of the deep features, the operation efficiency is improved. And by calculating the corresponding weights for each channel of the feature layer, the important channels can be assigned higher weights. During this period, the module uses a one-dimensional convolution module with different kernel sizes (typically, for example, 1, 3, 5) to obtain more interaction information between channels. The calculation process of the shallow channel attention module is as Figure 3 shown. First, for L i+k,j Perform global average pooling (GAP) to obtain a feature representation of C×1×1. This process is not carried out by setting a specific pooling window size, but rather taking the average value of the entire feature map as the output. By globally averaging the feature values of each channel on a feature map, the pooled feature values corresponding to each channel are generated. The finally obtained pooled feature vector can be regarded as the global information representation of the entire feature map. Then, perform Squeeze compression on this feature representation to obtain a feature map of C×1, and subsequently perform a Transpose operation to obtain a feature map of 1×C for subsequent convolution operations. The expression is as follows:
[0058] f 2 (L i+k,j )=T r {S q [GAP(L i+k,j )]}
[0059] In the formula, f 2 (·) is the microprocessing unit, S q is the dimensionality reduction operation, T r is the transpose operation, and GAP is global average pooling.
[0060] Subsequently, perform one-dimensional convolution operations on the obtained feature map f 2 (L i+k,j ) with different convolution kernels and add them to obtain the fused feature map The expression is as follows:
[0061]
[0062] In the formula, and are one-dimensional convolution operations with convolution kernel sizes of a, b, and c respectively. In this embodiment, specifically, it is input into one-dimensional convolution modules with convolution kernels of 1, 3, and 5 respectively for convolution operations, is the tensor addition operation.
[0063] Finally, perform a transpose operation through T r on the penultimate and antepenultimate dimensions to obtain a feature map of scale C×1, and then perform an Unsqueeze operation to expand this feature map on the last dimension to obtain an output of C×1×1. Finally, multiply it with the original feature map to obtain the weighted deep feature The expression is as follows:
[0064]
[0065] In the formula, US q is the dimensionality increase operation;
[0066] Finally, the output of the Deep Global Channel Attention Module (DGCAM) is expressed as follows:
[0067]
[0068] wherein, is the output of the DGAM module.
[0069] Thus, referring to Figure 4 shown, the CsHCAM module aligns the shallow features and the deep features in scale through upsampling and channel number adjustment, and uses the cross-fusion method to fuse the shallow features and the deep features. Its construction method is as follows: for the output feature maps obtained after LCAM and DGCAM, a cross-fusion method is used to fuse the shallow features and the deep features. Specifically, the result of element-wise multiplication of L c (L i,j ) and the original L i+k,j is added to the result of element-wise multiplication of and the original L i,j to obtain That is:
[0070] In a specific embodiment of the present invention, the shallow feature L i,j is defined as the feature L 0,0 output by the coding node of the 0th layer, and the deep feature L i+k,j is defined as the feature L 3,0 output by the coding node of the 3rd layer and the feature L 4,0 output by the coding node of the 4th layer. That is, based on the Cross-scale Cross-fusion Channel Attention Module (CsHCAM) arranged between the coding node and the decoding node of the 3rd layer, and between the coding node of the 4th layer and the first additional decoding node of this layer, the corresponding deep features L 3,0 and L 4,0 are respectively fused with L 0,0 .
[0071] Step 23, based on the Multi-Input Feature Fusion Module (MIFM), fuse the output features of the last node of each layer of the improved U-Net of the present invention, so as to obtain the final output detection result.
[0072] The multi-input feature fusion module (MIFM) of the present invention aims to effectively fuse multiple input features from different feature layers. In the previous step, the shallow global channel attention module was used to extract the interaction relationship between channels from the local perspective through pointwise convolution of shallow features, and then the deep global channel attention module was used to extract the interaction relationship between the global channels of deep features by performing weighted operations on each channel of the deep features using a set of one-dimensional convolution modules from the global perspective. In this step, the multi-input feature fusion module is used to fuse the extracted shallow channel information and deep channel information in a cross-fusion manner.
[0073] Further, the dense skip connections between multiple nodes in each layer in step 1 of the present invention are also implemented through this multi-input feature fusion module. Among them, the dense skip connections of the last two layers are to use two CsHCAM structures to fuse the shallow feature L 0,0 with the deep feature L 3,0 、L 4,0 to obtain the deep features L 3,1 、L 4,1 integrating deep information and shallow information. In addition, skip connections are made in the manner shown in the attached Figure 6 to continue adding the L 3,2 、L 4,2 nodes. The dense skip connections of each of the remaining layers are such that the input received by each node is the output of all previous nodes in the current layer and the downsampling of the output of the same-level nodes in the adjacent upper level and the upsampling of the output of the nearest previous node in the next layer. The specific connection method is as shown in Figure 6 .
[0074] Specifically, referring to Figure 5 shown, the multi-input feature fusion module of the present invention performs the following steps:
[0075] Step (1), stack the multiple inputs by channels to obtain a relatively "thick" feature L 1 ', which is expressed as follows:
[0076]
[0077] Both gh, nm, and xy represent the input features received by the current node from the current level or the inputs after upsampling or downsampling to the same scale from the adjacent levels, represents stacking by channels, and L 1 ' is the output after stacking by channels, which retains all the input feature information.
[0078] Step (2), apply L 1It is divided into two paths. One path first passes through two depthwise separable convolution blocks to change the number of channels to that of the feature map of the current layer, then performs normalization, and then obtains the weighted feature one through the convolutional attention module. The other path first passes through a two-dimensional convolution with a kernel size of 1, then performs normalization, adds it to the weighted feature one, and then obtains the mixed feature L through the activation function. 2 ', which is expressed as follows:
[0079]
[0080] In the formula, R is the ReLU non-linear activation function, B is the batch normalization layer (Batch Normalization, BN), DSConv1 and DSConv2 are both depthwise separable convolutions, and Φ is the convolutional attention module (Convolutional Block Attention Module, CBAM), and its output can be expressed as follows:
[0081]
[0082] In the formula, L is the input of the convolutional attention module, and M c () represents the output of the channel attention module CA, and its expression is as follows:
[0083] M c (L) = σ(MLP(P avg (L)) + MLP(P max (L)))
[0084] In the formula, MLP() is the shared fully connected layer, and P avg () is the average pooling layer, and P max () is the max pooling layer;
[0085] M s () represents the output of the channel attention module SA, and the expression is as follows:
[0086]
[0087] In the formula, σ is the Sigmoid activation function, is the two-dimensional convolution with a kernel size of 7.
[0088] Step (3), divide L 2 ' into two paths. One path first passes through two depthwise separable convolution blocks to change the number of channels to that of the feature map of the current layer, then performs normalization, and then obtains the weighted feature two through the convolutional attention module. The other path first passes through a two-dimensional convolution with a kernel size of 1, then performs normalization, adds it to the weighted feature two, and then obtains the final output feature L through the activation function.new , which is shown as follows:
[0089]
[0090] Both DSConv3 and DSConv4 are depthwise separable convolutions.
[0091] After the above process, the improved U-net of the present invention outputs L 0,4 , L 1,4 , L 2,3 , L 3,2 , L 4,2 five features. This design of multi-scale feature information fusion improves the model's ability to capture subtle information. The output feature L 0,4 of the last node in the 0th layer remains unchanged. The output features L 1,4 , L 2,3 , L 3 ,2 , L 4,2 of the last nodes in the remaining layers are adjusted to the size of the original image through upsampling to obtain Final_L 0,4 , Final_L 1,4 , Final_L 2,3 , Final_L 3,2 , Final_L 4,2 . Then, they are fused based on the multi-input feature fusion module. By stacking these five feature layers in the channel dimension, the representational power of the features is further enhanced. This process ensures that the model can effectively integrate information at different levels. Finally, the number of image channels is adjusted through a 1×1 convolution to generate the final output detection result. The complete architecture and process of the present invention can be referred to Figure 7 as shown.
[0092] To illustrate the effectiveness of this method, CsH_DCNet is compared and analyzed with traditional algorithms and deep learning algorithms. Among them, the traditional algorithms include algorithms based on local features. The present invention selects the WSLCM algorithm, RLCM algorithm, and TLLCM algorithm; it also includes algorithms based on low-rank sparse matrices. The present invention selects the IPI algorithm and PSTNN algorithm. For the deep learning-based algorithms, the present invention selects the DNANet algorithm, ALCNet algorithm, and AGPCNet algorithm.
[0093] To prove the effectiveness of the algorithm, objective index comparison and analysis are carried out in the NUAA and NUDT datasets for the comparative experiment.
[0094] Table 1 shows the performance of objective evaluation indexes of different SIRST algorithms on the NUDT-SIRST dataset and the NUAA-SIRST dataset.
[0095] As can be seen from Table 1, on both the relatively large-scale dataset NUDT-SIRST and NUAA-SIRST, the CsH_DCNet proposed in this chapter has achieved the best performance. In contrast, traditional methods are usually designed for specific scenarios (e.g., specific target sizes and complex backgrounds), and the manually selected parameters (e.g., the structure size in Tophat and the patch size in IPI) limit the generalization performance of these methods. The CsH_DCNet proposed in this chapter is insensitive to scene changes.
[0096] In the comparative analysis between traditional algorithms and CsH_DCNet, a significant phenomenon can be observed: compared with the recall rate R a , the improvement of deep learning algorithms in the mean intersection over union (mIoU) is more obvious. This difference mainly stems from the emphasis of traditional methods on target localization. They tend to focus on overall target localization rather than the precise matching of target shapes. This finding further validates the previous discussion on the selection of evaluation metrics: using pixel-level evaluation metrics (such as intersection over union) may lead to unfair comparisons and thus inaccurate conclusions.
[0097] Table 1 Evaluation Metrics of Different Algorithms on Dual Datasets
[0098]
[0099] On the NUDT dataset, compared with the second-place DNANet, CsH_DCNet increased its mIoU by 1.77% and R a by 1.12%, and F a decreased by 5.366×10 -6 .
[0100] On the NUAA dataset, compared with the second-place AGPCNet, CsH_DCNet increased its mIoU by 2.11% and R a by 1.05%, and F a also reached the second place. Compared with the first-place DNANet, the false alarm rate was only 4.822×10 -6 higher, but its overall performance was better than that of DNANet.
[0101] Compared with other methods based on convolutional neural networks, CsH_DCNet has achieved obvious improvements. This is because this chapter redesigned a new densely connected network tailored for SIRST detection. This densely connected network combines the CsHCAM shallow and deep feature fusion module, enabling it to well balance the shallow information and deep information of infrared small target images, effectively cross-fusing the shallow information into the deep information and suppressing the interference of background clutter. In addition, the densely connected structure of CsH_DCNet also provides an optimized path for feature transmission, and the MIFM module ensures that each unit in the deep layer of the network can receive useful information from the previous layer and the current layer. This design greatly improves the information utilization rate and thus enhances the detection performance.
[0102] mIoU is an important metric to measure the overall performance of a network on a specific dataset, and CsH_DCNet has achieved the best results on two different datasets, which fully demonstrates its excellent ability to extract infrared small targets in various scenarios. In addition, through further comparative analysis, it can be found that CsH_DCNet has achieved the highest R a value, indicating that under the same total number of targets, the number of targets correctly extracted by CsH_DCNet leads all the comparison algorithms. At the same time, its lower F a means that CsH_DCNet can significantly reduce false alarms while maintaining the detection of a large number of infrared small targets, showing superior background suppression ability.
[0103] This excellent performance is attributed to the advanced design of CsH_DCNet. It not only shows advantages in the number of target extractions, but more importantly, the pixel matching degree between the targets it extracts and the actual targets is extremely high, indicating a significant improvement in detail accuracy. This is particularly crucial in the field of infrared small target detection because correct target recognition and precise shape matching directly affect the efficiency and accuracy of subsequent processing and applications.
[0104] In summary, the present invention uses a densely connected network with cross-scale cross-fusion based on U-net (CsH_DCNet) for precise localization and detail extraction of infrared small targets. First, by introducing a cross-scale cross-fusion channel attention module, this network uses pointwise convolution to extract small-scale detail information and uses one-dimensional convolutions with a set of different receptive fields to mine large-scale semantic information, and then cross-fuses this information, so as to retain the detail features of small targets in the deep network. Second, the present invention adopts a feature interaction strategy of dense nodes, and through upsampling and downsampling operations between different levels and dense skip connections between the same levels, it reduces the semantic gap between the encoder and the decoder and fully mines and utilizes the information in the infrared image. In addition, the present invention also constructs a multi-input feature fusion module, which combines a dual residual network and an attention module and can effectively identify and fuse important information in multiple input features. Experimental results show that the overall performance of CsH_DCNet on the NUDT-SIRST and NUAA-SIRST datasets is better than that of a variety of mainstream algorithms, demonstrating good generalization ability. The structure of the present invention is reasonably designed and is applicable to the field of small target detection in infrared images.
Claims
1. A method for infrared small target detection based on U-Net cross-scale cross-fusion densely connected network, characterized in that: The steps include: Step 1, inputting the initial infrared image into an improved U-Net, wherein the improved U-Net adds several cascaded intermediate nodes between the encoding nodes and the decoding nodes of the 0th to 2nd layers of the original U-Net, adds cascaded additional decoding nodes after the decoding nodes of the 3rd layer and after the encoding nodes of the 4th layer, and performs dense skip connections between multiple nodes in each layer; Step 2, based on the cross-scale cross-fusion channel attention module, in any one or several layers from the first layer to the fourth layer of the improved U-Net, the shallow features are fused with the corresponding deep features; Step 3: Based on the multi-input feature fusion module, the output features of the last node of each layer of the improved U-Net are fused to obtain the final output detection result.
2. According to claim 1, the infrared small target detection method based on U-Net cross-scale cross-fusion dense connection network is characterized in that: In the improved U-Net, the number of channels in each layer is twice that of the previous layer, and the nodes include: The output features of the encoding node at layer 0, the three intermediate nodes at layer 0, and the decoding node at layer 0 are L 0,0 , L 0,1 , L 0 ,2 , L 0,3 , L 0,4 ; The output features of the first layer encoding node, the first layer three intermediate nodes and the first layer decoding node are L 1,0 , L 1,1 , L 1 ,2 , L 1,3 , L 1,4 ; The output features of the second layer encoding node, the second layer two intermediate nodes and the second layer decoding node are L 2,0 , L 2,1 , L 2 ,2 , L 2,3 ; The output features of the third layer encoding node, the third layer decoding node and the first additional decoding node are L 3,0 , L 3,1 , L 3 ,2 ; The output features of the 4th layer encoding node, the second additional decoding node, and the third additional decoding node are L 4,0 , L 4,1 , L 4,2 .
3. According to claim 1, the infrared small target detection method based on U-Net cross-scale cross-fusion dense connection network is characterized in that: The cross-scale cross-fusion channel attention module includes a shallow channel attention module and a deep global channel attention module; the output features of the shallow channel attention module and the deep global channel attention module are cross-fused to obtain the output of the cross-scale cross-fusion channel attention module.
4. According to claim 3, the infrared small target detection method based on U-Net cross-scale cross-fusion dense connection network is characterized in that: Define the output feature of the jth node of the i-th layer of the improved U-Net as the shallow feature L i,j , the output feature of the jth node in the i+kth layer is the deep feature L i+k,j ; The output of the shallow channel attention module is as follows: In the formula, is the element-by-element point multiplication operation after scale alignment, Lc(·) represents the calculation operation of the shallow channel attention module, and L i,j The calculation formula is as follows: In the formula, is the ReLU nonlinear activation function, B is the batch processing unit layer, is a dilated convolution with a dilation rate of d, f1(·) is a microprocessing unit, σ is a Sigmoid activation function, It is point-wise convolution; The output of the deep global channel attention module is as follows: is the weighted deep feature, the expression is as follows: f2(L i+k,j )=T r {S q [GAP(L i+k,j )]} Where, US q For the dimension-raising operation, S q For dimensionality reduction operation, T r is the transposition operation, GAP is the global average pooling, f2(·) is the micro-processing unit, To convert f2(L i+k,j ) are the sum of the one-dimensional convolution operations of different convolution kernels, and the expression is as follows: In the formula, and are one-dimensional convolution operations with kernel sizes of a, b, and c respectively, and ⊕ is a tensor addition operation; The output of the cross-scale cross-fusion channel attention module is:
5. According to claim 4, the infrared small target detection method based on U-Net cross-scale cross-fusion dense connection network is characterized in that: The shallow channel attention module, assuming r is the channel compression ratio, then the PWC in the first f1(·) 1×1 Input L i,j The number of channels is C, the number of output channels is C / r, and the second time f1(·) The number of input channels is C / r, the number of output channels is C, and the convolution kernel size of the two point-by-point convolutions is 1×1; The deep global channel attention module, a=1, b=3, c=5, By T r The transpose operation transposes across the first-last and second-last dimensions.
6. The infrared small target detection method based on U-Net cross-scale cross-fusion dense connection network according to any one of claims 1 to 5 is characterized in that: The shallow feature is the feature L output by the encoding node at the 0th layer. 0,0 The cross-scale cross-fusion channel attention module (CsHCAM) is arranged between the encoding node and the decoding node of the third layer, and between the encoding node of the fourth layer and the first additional decoding node of the layer, and the corresponding deep feature is the output feature L of the encoding node of the third layer. 3,0 , and the output feature L of the 4th layer encoding node 4,0 , respectively with L 0,0 Fusion.
7. According to claim 6, the infrared small target detection method based on U-Net cross-scale cross-fusion dense connection network is characterized in that: The improved U-Net uses a multi-input feature fusion module to perform dense skip connections between multiple nodes in each layer, wherein the dense skip connections of the last two layers use two cross-scale cross-fusion channel attention modules to fuse shallow features with deep features, and the dense skip connections of each of the remaining layers are inputs received by each node, which are down-sampled outputs of all previous nodes in the current layer and the outputs of the same-level nodes in the adjacent previous layer, and up-sampled outputs of the previous nodes of the nearest neighbors in the next layer.
8. The infrared small target detection method based on U-Net cross-scale cross-fusion dense connection network according to claim 1 or 7 is characterized in that: The multi-input feature fusion module performs the following steps: Step (1), stack multiple inputs by channel to obtain feature L1′; Step (2), L1′ is divided into two paths, one path first passes through two depth-separable convolution blocks to change the number of channels to the number of channels of the feature map of the current layer, then normalizes, and then passes through the convolution attention module to obtain the weighted feature one, and the other path first passes through a two-dimensional convolution with a convolution kernel size of 1, then normalizes, and adds it to the weighted feature one, and then passes through the activation function to obtain the mixed feature L2′; Step (3), L2′ is divided into two paths, one path first passes through two depth-separable convolution blocks, the number of channels is changed to the number of channels of the feature map of the current layer, then normalized, and then passed through the convolution attention module to obtain the weighted feature 2, the other path first passes through a two-dimensional convolution with a convolution kernel size of 1, then normalized, and added to the weighted feature 2, and then through the activation function to obtain the final output feature L new .
9. The infrared small target detection method based on U-Net cross-scale cross-fusion dense connection network according to claim 8 is characterized in that: The output of the convolutional attention module is expressed as follows: Where L is the input of the convolutional attention module, M c () represents the output of the channel attention module CA, which is expressed as follows: M c (L)=σ(MLP(P avg (L))+MLP(P max (L))) Where MLP() is a shared fully connected layer, P avg () is the average pooling layer, P max () is the maximum pooling layer; M s () represents the output of the channel attention module SA, which is expressed as follows: Where σ is the Sigmoid activation function, It is a two-dimensional convolution with a convolution kernel size of 7.
10. The infrared small target detection method based on U-Net cross-scale cross-fusion dense connection network according to claim 4 is characterized in that: In step 3, the output features of the last node of the 0th layer of the improved U-Net are kept unchanged, and the output features of the last nodes of the remaining layers are adjusted to the original image size by upsampling, and then fused, and the fusion result is adjusted to the number of image channels through a 1×1 convolution, so as to generate the final output detection result.
Citation Information
Cited By
Infrared weak and small target detection network under complex background based on multi-level feature combination
CN120852794A
Spatial faint target detection method based on multi-scale wavelet features
CN121458966A