Infrared Target Anti-Jamming Detection Algorithm Based on Feature-Enhanced Semantic Segmentation Network

Through position attention feature fusion and mixed hollow space pyramid pooling module, infrared target features are enhanced, and the detection performance problem of infrared target detection under weak features and interference occlusion is solved, and efficient anti-interference detection of infrared targets is achieved.

CN118674918BActive Publication Date: 2025-07-08HENAN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410820709.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2025-07-08
Estimated Expiration
2044-06-24

AI Technical Summary

Technical Problem

The existing infrared target detection algorithms are affected when infrared target characteristics are weak and interference occlusion, and anti-interference detection cannot be effectively realized.

Method used

The infrared target anti-interference detection algorithm based on feature-enhanced semantic segmentation network is adopted to enhance the target features through the position attention feature fusion network and the mixed hollow space pyramid pooling module, including the design of the position attention module and the mixed hollow space pyramid pooling module, respectively, by obtaining the correlation weights between pixels at different locations and expanding the convolution kernel receptive field to enhance feature expression.

Benefits of technology

When the infrared target radiation characteristics are weak and there is interference occlusion, the performance of infrared target detection is significantly improved, and the detection accuracy and recall rate are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118674918B_ABST
    Figure CN118674918B_ABST
Patent Text Reader

Abstract

An infrared target anti-interference detection algorithm based on a feature-enhanced semantic segmentation network enhances target features by combining position attention feature fusion and hybrid atrous spatial pyramid pooling. The position attention feature fusion network obtains the correlation weights between pixels at different positions to enhance the current-level feature map, and fuses it with the high-level feature map rich in semantic features to enhance the expression ability of infrared target features. The hybrid atrous spatial pyramid pooling expands the receptive field of the convolutional kernel on the feature map without losing feature information through the tandem structure of dilated convolutions with small dilation rates, obtains the context information of the feature map, and further enhances the expression ability of infrared target features. The present invention combines the position attention feature fusion network, the hybrid atrous spatial pyramid pooling module and the network for maintaining the resolution of the feature map, and can effectively realize the anti-interference detection of infrared targets in the case of weak infrared target radiation features and interference occlusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning infrared target detection methods, and in particular to an infrared target anti-interference detection algorithm based on a feature-enhanced semantic segmentation network. Background Art

[0002] The semantic segmentation-based target detection task is to assign a class label to each pixel in the image through a convolutional neural network to obtain a segmentation mask to distinguish the target from the background, and the target detection can be achieved by the position and class of the pixels. It mainly includes technical routes such as fully convolutional-based, encoder-decoder-based, feature fusion-based, and optimized convolutional structure-based. HRNetv2 is a semantic segmentation network that maintains the feature map resolution based on the encoder-decoder structure, including four processing stages: high-resolution stage, low-resolution stage, horizontal stage, and merging stage. The high-resolution stage contains multiple residual modules to extract target feature information and outputs a feature map with the same resolution as the input feature map. The low-resolution stage downsamples the feature map output by the high-resolution stage to reduce the memory occupancy of the model. The horizontal stage introduces multiple parallel branches on feature maps with different resolutions for feature map interaction. The merging stage restores feature maps with different resolutions to the same size as the high-resolution feature map through convolution and upsampling, and then fuses all feature maps to generate the final segmentation result. However, when the infrared target features are weak and when the target is interfered and occluded, the learnable features will be further reduced, resulting in insufficient learning of features by the convolutional neural network, thereby affecting the detection performance. Summary of the Invention

[0003] The object of the present invention is to provide an infrared target anti-interference detection algorithm based on a feature-enhanced semantic segmentation network, which can effectively achieve anti-interference detection of infrared targets in the case of weak infrared target radiation features and interference occlusion.

[0004] The technical solution adopted by the present invention to solve the above technical problems is: an infrared target anti-interference detection algorithm based on a feature-enhanced semantic segmentation network, including the following steps:

[0005] Step 1, obtain a high-resolution feature map rich in detail information and semantic information through a position attention feature fusion network. The specific method is as follows:

[0006] First, obtain a feature map C1 with the same resolution as the input image, a feature map C2 downsampled by 4 times, a feature map C3 downsampled by 8 times, and a feature map C4 downsampled by 16 times through the HRNetv2 network;

[0007] Then, the channel number of the downsampled 8-fold feature map C3 is adjusted through a 1×1 convolution, and then feature enhancement is performed through the position attention module. Then, the feature map with enhanced expression is fused with the downsampled 16-fold feature map C4 after upsampling to obtain the feature map P3;

[0008] Then, the channel number of the downsampled 4-fold feature map C2 is adjusted through a 1×1 convolution, and then feature enhancement is performed through the position attention module. Then, the feature map with enhanced expression is fused with the feature map P3 after upsampling to obtain the feature map P2;

[0009] Then, the channel number of the feature map C1 with the same resolution as the input image is adjusted through a 1×1 convolution, and then feature enhancement is performed through the position attention module. Then, the feature map with enhanced expression is fused with the feature map P2 after upsampling to obtain the feature map P1, which is a high-resolution feature map rich in detailed information and semantic information;

[0010] The above calculation process for fusing the feature maps is as follows:

[0011] ...... Equation (1);

[0012] In Equation (1), Upsample represents upsampling, f LAM represents the feature enhancement operation, Conv 1×1 represents the 1×1 convolution, F i is the feature map with enhanced expression at the i-th layer;

[0013] The above process of performing feature enhancement operation through the position attention module is as follows:

[0014] First, the feature maps B, C, D ∈ R are obtained from the original feature map A that needs to perform feature enhancement through the use of a 1×1 convolution. C×H×W Then, B, C, D are reshaped into B', C', D' ∈ R C×N , where N = H × W;

[0015] Then, the transpose of B' is dot-multiplied with C' to obtain the correlation score, and then the softmax function is used to calculate the position attention weight S ij ∈ R N×N , and the calculation formula is:

[0016] ...... Equation (2);

[0017] In Equation (2), B' T is the transpose of B', B' T i is the i-th position element of B' T , C' jis the element at the j-th position of C', S ij is the weight of the element at the i-th position with respect to the element at the j-th position;

[0018] Then, multiply D' by the position attention weight S ij to obtain the feature map E ∈ R C×N , where N = H × W, and the formula for the j-th position element E j is as follows:

[0019] ...... Equation (3);

[0020] In Equation (3), D' i is the i-th position element of D';

[0021] Then, after reshaping the feature map E that enhances pixel dependencies and adding it to the original feature map A, F = E + A, we obtain the enhanced feature map F ∈ R C×H×W ;

[0022] Step 2: Extract the context information of the feature map P1 through a hybrid atrous spatial pyramid pooling module. The hybrid atrous spatial pyramid pooling module includes multiple parallel atrous convolution branches, and each atrous convolution branch has an atrous convolution tandem structure formed by sequentially connecting multiple layers of atrous convolution. The dilation rate r i of the i-th layer of atrous convolution in any one of the atrous convolution branches is determined according to the following method:

[0023] For an atrous convolution branch formed by sequentially connecting n layers of atrous convolution, first, a set of preset dilation rates {r1, r2, r3...... r n} is given, where r1 = 1, that is, the dilation rate of the first layer of atrous convolution is 1;

[0024] Then, starting from the n-th layer of this atrous convolution branch, calculate the maximum distance M i between the dilation rates of two adjacent layers in the atrous convolution branch successively. The calculation formula is:

[0025] ...... Equation (4);

[0026] Let M n = r n for the first calculation, substitute M n and r n into Equation (4) to calculate M n-1 , and then substitute M n-1 and r n-1 into Equation (4) successively until M2 is obtained;

[0027] Compare M2 with the convolution kernel size k of the second-layer dilated convolution in the dilated convolution branch of the cavity. When M2 ≤ k, directly adopt the preset dilation rates {r1, r2, r3...... r n}, when M2 > k, re-give the preset dilation rate and calculate again;

[0028] Step 3: Concatenate the output of the first-layer 1×1 convolution of the hybrid dilated spatial pyramid pooling module, the outputs of multiple dilated convolution branches, and the pooling output to obtain the output feature map;

[0029] Step 4: Pass the feature map in Step 3 to the Softmax classifier for pixel-level semantic segmentation to achieve anti-interference detection of infrared targets.

[0030] According to the above technical solution, the beneficial effects of the present invention are:

[0031] The anti-interference detection algorithm for infrared targets of the present invention can enhance target features by combining position attention feature fusion and hybrid dilated spatial pyramid pooling in the case of weak infrared target radiation features and interference occlusion, and then effectively achieve anti-interference detection of infrared targets. Among them, the position attention feature fusion network obtains the correlation weights between pixels at different positions to enhance the feature map of the current layer, and fuses it with the high-level feature map rich in semantic features to enhance the expression ability of infrared target features. The hybrid dilated spatial pyramid pooling expands the receptive field of the convolution kernel on the feature map without losing feature information through the series structure of small dilation rate dilated convolutions, obtains the context information of the feature map, and further enhances the expression ability of infrared target features. Description of the Drawings

[0032] Figure 1 It is a schematic structural diagram of an anti-interference detection algorithm for infrared targets based on a feature-enhanced semantic segmentation network;

[0033] Figure 2 It is a schematic structural diagram of a position attention feature fusion network (LAFFN);

[0034] Figure 3 It is a schematic structural diagram of a position attention module (LAM);

[0035] Figure 4 It is a schematic diagram of the calculation of hybrid dilated spatial pyramid pooling with the series structure of dilation rates r = 1 and r = 3;

[0036] Figure 5 It is a schematic structural diagram of a dilated convolution with dilation rates {6, 12, 18} in the existing atrous spatial pyramid pooling (ASPP);

[0037] Figure 6Schematic diagram of the cascaded structure of dilated convolutions with dilation rates of {[1, 2, 3], [1, 2, 3, 7], [1, 3, 5, 9]} in the Hybrid Atrous Spatial Pyramid Pooling (H-ASPP) of the embodiment;

[0038] Figure 7 Partial images of the infrared aircraft dataset under interference;

[0039] Figure 8 Comparison chart of the input and output of the position attention;

[0040] Figure 9 Comparison chart of the segmentation results of different algorithms. Detailed implementation manners

[0041] The infrared target anti-interference detection algorithm based on the feature-enhanced semantic segmentation network includes the following steps:

[0042] Step 1: Obtain a high-resolution feature map rich in detailed information and semantic information through the Location Attention Feature Fusion Networks (LAFFN), specifically as Figure 2 shown.

[0043] First, obtain the feature maps C1 with the same resolution as the input image, the feature map C2 downsampled by 4 times, the feature map C3 downsampled by 8 times, and the feature map C4 downsampled by 16 times through the HRNetv2 network.

[0044] Then, adjust the number of channels of the feature map C3 downsampled by 8 times through a 1×1 convolution, then perform feature enhancement through the Location Attention Module (LAM), and then fuse the enhanced feature map with the feature map C4 downsampled by 16 times after upsampling to obtain the feature map P3.

[0045] Then, adjust the number of channels of the feature map C2 downsampled by 4 times through a 1×1 convolution, then perform feature enhancement through the Location Attention Module (LAM), and then fuse the enhanced feature map with the feature map P3 after upsampling to obtain the feature map P2.

[0046] Then, adjust the number of channels of the feature map C1 with the same resolution as the input image through a 1×1 convolution, then perform feature enhancement through the Location Attention Module (LAM), and then fuse the enhanced feature map with the feature map P2 after upsampling to obtain the feature map P1, which is the high-resolution feature map rich in detailed information and semantic information.

[0047] The above calculation process for fusing the feature maps is:

[0048] ...... Equation (1);

[0049] In formula (1), Upsample represents upsampling, and f LAM represents the feature enhancement operation, and Conv 1×1 represents a 1×1 convolution, and F i is the feature map of the i-th layer of enhanced expression.

[0050] The process of performing feature enhancement operations through the Location Attention Module (LAM) is as Figure 3 shown.

[0051] First, feature maps B, C, D ∈ R are obtained from the original feature map A that needs feature enhancement by using a 1×1 convolution. C×H×W , and then B, C, D are reshaped into B', C', D' ∈ R through shape reshaping reshape. C×N , where N = H×W.

[0052] Then, the transpose of B' is dot-multiplied with C' to obtain the correlation score, and the softmax function is used to calculate the location attention weight S ij ∈ R N×N , and the calculation formula is:

[0053] ...... Formula (2);

[0054] In formula (2), B' T is the transpose of B', B' T i is the i-th position element of B' T , C' j is the j-th position element of C', and S ij is the weight of the i-th position element for the j-th position element.

[0055] Then, D' is multiplied by the location attention weight S ij to obtain the feature map E ∈ R that enhances the pixel dependency relationship. C×N , where N = H×W, and the calculation formula for the j-th position element E j is:

[0056] ...... Formula (3);

[0057] In formula (3), D' i is the i-th position element of D'.

[0058] Then, after reshaping the feature map E that enhances pixel dependencies through reshape, it is added to the original feature map A, and the enhanced feature map F ∈ R is obtained. C×H×W .

[0059] Step 2: Extract the context information of the feature map P1 through the Hybrid Atrous Spatial Pyramid Pooling (H-ASPP) module. The Hybrid Atrous Spatial Pyramid Pooling (H-ASPP) module is designed based on the Atrous Spatial Pyramid Pooling (ASPP) module. The Atrous Spatial Pyramid Pooling (ASPP) module uses atrous convolutions with different dilation rates to capture the context information of different receptive fields of the feature map, improving the detection accuracy of the algorithm, but prone to losing feature information. To overcome this shortcoming, the Hybrid Atrous Spatial Pyramid Pooling (H-ASPP) module includes multiple parallel atrous convolution branches, and each atrous convolution branch has an atrous convolution tandem structure formed by sequentially connecting multiple layers of atrous convolutions.

[0060] Taking the tandem structure with atrous convolution dilation rates r1 = 1 and r2 = 3 as an example, the detailed calculation process of feature extraction is as Figure 4 shown. Figure 4 (a) is a 9×9 grayscale image. Figure 4 (b) is the first-layer feature map obtained by using a normal convolution with a convolution kernel k = 3 and a dilation rate r1 = 1 on Figure 4 (a). Figure 4 (c) is the second-layer feature map obtained by using an atrous convolution with a convolution kernel k = 3 and a dilation rate r2 = 3 on Figure 4 (b).

[0061] The calculation formulas for the size of the convolution kernel and the receptive field after dilation by atrous convolution are shown as follows:

[0062] ;

[0063] ;

[0064] In the formula, k is the size of the convolution kernel before dilation, k' is the size of the convolution kernel after dilation, l n-1 is the size of the receptive field of the (n - 1)-th layer, k n is the size of the convolution kernel of the n-th layer, and s is the stride of the convolution kernel.

[0065] Figure 4 (a) The receptive field of each element of the input image is 1, so l1 is 1, the convolution kernel k1 = 3, the convolution kernel stride s1 = 1, and it is calculated that l2 = 3, that is Figure 4Each element in (b) is Figure 4 The receptive field in (a) is 3×3, and Figure 4 Any element in (b) is Figure 4 Calculated from the nine elements in (a) with the same color as this element.

[0066] For the second-layer convolutional kernel k2 = 3 and dilation rate r2 = 3, the size of the dilated convolutional kernel k' = 7 is calculated, and then l3 = 9 is calculated, that is, Figure 4 Each element in (c) is Figure 4 The receptive field in (a) is 9×9, and in Figure 4 The receptive field in (b) is 7×7, and only Figure 4 Nine elements with the number 1 symbol in (b) are involved in the calculation.

[0067] Analysis Figure 4 It can be found from (a) that when using ordinary convolution with convolutional kernel k1 = 3 in the first layer and dilated convolution with convolutional kernel k2 = 3 and dilation rate r2 = 3 in the second layer, the feature information of the input image is not lost exactly.

[0068] When multiple dilated convolutions are cascaded, the calculation process is the same as above. However, the third-layer feature map and subsequent feature maps are obtained by using dilated convolution on the previous-layer feature map, and there is loss of the feature information of the previous-layer feature map. The first-layer feature map is obtained by using ordinary convolution on the input image without loss of feature information. Therefore, if the second-layer feature map is obtained by using dilated convolution on all elements of the first-layer feature map, it can be inferred that the dilated convolution in the cascaded structure does not lose the feature information of the input image during the feature extraction process.

[0069] To ensure that the dilated convolution in the cascaded structure does not lose the feature information of the input image during the feature extraction process, the dilation rate r of the i-th layer of dilated convolution in any dilated convolution branch in the hybrid dilated spatial pyramid pooling module is i Determined according to the following method.

[0070] For a dilated convolution branch formed by cascading n layers of dilated convolutions in sequence, first a set of preset dilation rates {r1, r2, r3...... r n} is given, where r1 = 1, that is, the dilation rate of the first layer of dilated convolution is 1.

[0071] Then starting from the n-th layer of this dilated convolution branch, the maximum distance M between the dilation rates of adjacent two layers in the dilated convolution branch is calculated successively i , and the calculation formula is:

[0072] ...... Equation (4).

[0073] Let M be calculated for the first timen = r n , substitute M n and r n into Equation (4) to calculate M n-1 , and then substitute M n-1 and r n-1 into Equation (4 successively until M2 is obtained.

[0074] Compare M2 with the kernel size k of the second dilated convolution in the dilated convolution branch of this hole. When M2 ≤ k, directly adopt the preset dilation rates {r1, r2, r3...... r n}, when M2 > k, re - specify the preset dilation rate and calculate again.

[0075] Under the above principle, this embodiment can convert the dilated convolution with dilation rates {6, 12, 18} in the Atrous Spatial Pyramid Pooling module (ASPP) into a cascaded - structure dilated convolution with dilation rates {[1, 2, 3], [1, 2, 3, 7], [1, 3, 5, 9]} in the Hybrid Atrous Spatial Pyramid Pooling module (H - ASPP). Figure 5 is the dilated convolution with dilation rates {6, 12, 18} of ASPP. Figure 6 is the cascaded - structure dilated convolution with dilation rates {[1, 2, 3], [1, 2, 3, 7], [1, 3, 5, 9]} of H - ASPP.

[0076] Step 3: Concatenate the output of the 1×1 convolution in the first layer of the Hybrid Atrous Spatial Pyramid Pooling module, the outputs of multiple dilated convolution branches, and the pooling output to obtain the output feature map.

[0077] Step 4: Pass the feature map in Step 3 to the Softmax classifier for pixel - level semantic segmentation to achieve anti - interference detection of infrared targets.

[0078] The complete structural schematic diagram of the infrared target anti - interference detection algorithm based on the feature - enhanced semantic segmentation network is as Figure 1 shown.

[0079] Experimental verification

[0080] Hardware platform: The CPU model is Intel(R) Core(TM) i7-8700K CPU @ 3.70 GHz, and the running memory is 32 GB; the GPU model is NVIDIA 3090 Ti, and the video memory size is 24 GB. Software platform: The operating system is windows10, the Pytorch 1.13 deep learning framework, the programming language is Python 3.8, and CUDA 11.3.1 and CUDNN 8.2.1 are used to accelerate the GPU. During the algorithm training stage, the image resolution is uniformly set to 640×640, the batch-size is set to 8, the network uses the SGD optimizer, the learning rate is 0.004, and the weight decay is set to 0.0001. A total of 100 rounds of training are performed, and the learning rate gradually decreases linearly in each round.

[0081] 1. Dataset

[0082] The experimental data is from some data in the publicly available infrared dataset LSOTB-TIR, a total of 3984 images, but there is no interference. Based on this data, interference is artificially added to create infrared aircraft data under interference, as Figure 7 shown. The first row represents the original infrared aircraft data in LSOTB-TIR, and the second row represents the infrared aircraft data after artificially adding interference. The added interference follows the following rules: (1) The radiation of the interference is modeled according to a Gaussian distribution, that is, the gray value is the largest at the center and gradually weakens outward; (2) After the interference is released, it first increases from small to large, then decreases from large to small, and gradually moves away from the aircraft; (3) The shape is modeled by the equation of a circle. Using the labelme software, annotation is achieved by generating a JSON file containing target pixel category information and location information. For the infrared aircraft target under interference occlusion, a total of four categories are divided during annotation: handpiece, empennage, airscrew, interference, and the dataset is divided into a training set, a validation set, and a test set according to a ratio of 8:1:1.

[0083] 2. Data Analysis

[0084] 2.1 Evaluation Metrics

[0085] The present invention uses the mean intersection over union (MIoU), mean pixel accuracy (MPA), and mean recall (MR) as evaluation metrics.

[0086] (1) MIoU

[0087] The mean Intersection over Union (mIoU) is calculated by first computing the IoU for each class and then averaging the IoU values across all classes. The formula is as follows:

[0088] ;

[0089] (2) MPA

[0090] The mean pixel accuracy (MPA) is calculated by first computing the pixel accuracy for each class and then averaging the pixel accuracy values across all classes. The formula is as follows:

[0091] ;

[0092] (3) MR

[0093] The mean recall (MR) is calculated by first computing the recall for each class and then averaging the recall values across all classes. The formula is as follows:

[0094] ;

[0095] In the above formulas, i and j represent the classes in the dataset, n represents the number of classes in the dataset, n + 1 represents the total number of classes including the label and the background, p ii represents the number of pixels that are actually of class i and are predicted to be of class i, p ij represents the number of pixels that are actually of class i and are predicted to be of class j, p ji represents the number of pixels that are actually of class j and are predicted to be of class i.

[0096] 2.2 Ablation Analysis

[0097] To verify the effectiveness of the Location Attention Feature Fusion Network and the Hybrid Atrous Spatial Pyramid Pooling in the semantic segmentation algorithm, the present invention conducts ablation experiments with the HRNetv2 semantic segmentation algorithm as the basic framework. The experimental results are shown in Table 1.

[0098] Table 1 Ablation Analysis

[0099]

[0100] (1) To verify the effectiveness of the LAFFN module, experiments are conducted by adding only the LAFFN to the HRNetv2 basic algorithm. Compared with the HRNetv2 basic algorithm, the mIoU increases by 0.22 percentage points, the MPA increases by 0.7 percentage points, and the MR increases by 0.21 percentage points, indicating that the LAFFN module can improve the segmentation performance of the algorithm.

[0101] (2) To verify the effectiveness of the H-ASPP module, experiments were conducted by adding H-ASPP only to the basic HRNetv2 algorithm. Compared with the basic HRNetv2 algorithm, the MIoU increased by 0.25 percentage points, the MPA increased by 0.97 percentage points, and the MR increased by 0.30 percentage points, indicating that the H-ASPP module can improve the segmentation performance of the algorithm.

[0102] (3) To verify the effectiveness of the combined action of the two measures, both of the above two measures were added to the basic HRNetv2 algorithm for experiments. Compared with the basic HRNetv2 algorithm, under the combined action of the two measures, the MIoU increased by 0.56 percentage points, the MPA increased by 1.12 percentage points, and the MR increased by 0.67 percentage points, and it was higher than when each measure was used alone, indicating that the combined action of the two measures designed in the present invention can better improve the detection performance of the algorithm.

[0103] (4) To verify the enhancement effect on the target feature information after introducing the position attention module, visual analysis was performed on the input and output feature maps of the position attention, as Figure 8 shown. The first row represents the original image, the second row represents the feature map input to the position attention, and the third row represents the output feature map of the position attention. It can be found from the figure that after introducing the position attention module, the nose and tail target features of the data in the first column are enhanced, the nose target feature of the data in the second column is enhanced, the nose, tail, and propeller target features of the data in the third column are enhanced, and the tail and propeller target features of the data in the fourth column are enhanced. Therefore, it can be concluded that after introducing the position attention, the representation of the target area is strengthened and the feature information of the target is highlighted.

[0104] (5) To verify that H-ASPP has better detection performance compared with the ASPP module, these two modules were separately added to the HRNetv2 algorithm for comparative experiments, and the experimental results are shown in Table 2. It can be found from it that when adding H-ASPP to the basic HRNetv2 algorithm compared with adding ASPP to the basic HRNetv2 algorithm, the MIoU increased by 0.2 percentage points, the MPA increased by 0.82 percentage points, and the MR increased by 0.21 percentage points. It shows that H-ASPP designed in the present invention can better improve the detection performance of the algorithm compared with ASPP.

[0105] Table 2 Verification Results of ASPP and H-ASPP Modules

[0106]

[0107] 2.3 Comparative Experiments of Different Semantic Segmentation Algorithms

[0108] The algorithm of the present invention was compared with semantic segmentation algorithms such as FCN, U-Net, PSPNet, DeepLabv3+, Segformer, and HRNetv2. The results are shown in Table 3. It can be found that in the infrared aircraft target segmentation task, the MIoU of the invented algorithm reaches 90.61%, the MPA reaches 94.78%, and the MR reaches 95.29%, all of which are higher than those of the other semantic segmentation algorithms.

[0109] Table 3 Segmentation Results of Different Algorithms

[0110]

[0111] The visualization results of different semantic segmentation algorithms are as Figure 9 shown. The first and second rows are the detection results of the infrared aircraft propeller target. It can be seen from these two rows of data that all algorithms can detect the infrared aircraft propeller target, but the segmentation result of the invented algorithm for the aircraft propeller is closer to the ground truth label. The third and fourth rows are the detection results of the infrared aircraft tail target. It can be seen from the third row of data that the invented algorithm can detect the infrared aircraft tail target, while the other algorithms cannot. It can be seen from the fourth row of data that all algorithms can detect the infrared aircraft tail target, but the segmentation result of the invented algorithm for the aircraft tail is closer to the ground truth label. The fifth and sixth rows are the detection results of the infrared aircraft head target. It can be seen from the fifth row of data that the invented algorithm can detect the head target of the infrared aircraft below, while the other algorithms cannot detect it; it can be seen from the sixth row of data that the invented algorithm can effectively segment the nose of the infrared aircraft from the propeller, while the segmentation results of the nose and propeller by the other algorithms have a large difference from the ground truth label.

[0112] In summary, the present invention first constructs a position attention feature fusion network to obtain the correlation weights between pixels at different positions to enhance the feature map of the current layer, and fuses it with the high-level feature map rich in semantic features to enhance the expression ability of infrared target features; then constructs a constraint criterion for the dilation rate, uses the cascaded structure of dilated convolutions with small dilation rates to solve the problem of losing feature information when dilated spatial pyramid pooling expands the receptive field, and designs a hybrid dilated spatial pyramid pooling module on this basis to expand the receptive field of the convolutional kernel on the feature map to obtain the context information of the feature map, further enhancing the expression ability of infrared target features; combines the position attention feature fusion network, the hybrid dilated spatial pyramid pooling module with the network for maintaining the resolution of the feature map to construct an anti-interference detection algorithm for infrared targets based on a feature-enhanced semantic segmentation network. The present invention has been fully experimentally verified on an infrared dataset, and the results verify the effectiveness of the present invention.

Claims

1. An infrared target anti-interference detection algorithm based on a feature-enhanced semantic segmentation network, characterized in that It includes the following steps: Step 1: Obtain a high-resolution feature map rich in detailed information and semantic information through a position attention feature fusion network. The specific method is as follows: First, obtain feature maps C1 with the same resolution as the input image, downsampled feature map C2 by 4 times, downsampled feature map C3 by 8 times, and downsampled feature map C4 by 16 times through the HRNetv2 network; Then, adjust the number of channels of the downsampled feature map C3 through a 1×1 convolution, enhance the features through a position attention module, and then fuse the enhanced feature map with the upsampled downsampled feature map C4 to obtain feature map P3; Then, adjust the number of channels of the downsampled feature map C2 through a 1×1 convolution, enhance the features through a position attention module, and then fuse the enhanced feature map with the upsampled feature map P3 to obtain feature map P2; Then, adjust the number of channels of the feature map C1 with the same resolution as the input image through a 1×1 convolution, enhance the features through a position attention module, and then fuse the enhanced feature map with the upsampled feature map P2 to obtain feature map P1, which is the high-resolution feature map rich in detailed information and semantic information; The above calculation process for fusing feature maps is as follows: ...... formula (1); In formula (1), Upsample represents upsampling, and f LAM represents the feature enhancement operation, and Conv 1×1 represents a 1×1 convolution, and F i is the feature map of the enhanced expression of the i-th layer; The above process for performing feature enhancement operations through the position attention module is as follows: First, obtain feature maps B, C, D ∈ R from the original feature map A that needs feature enhancement by using 1×1 convolution C ×H×W , and then reshape B, C, D into B', C', D' ∈ R through shape reshaping C×N , where N = H × W; Then, take the matrix dot product of the transpose of B' and C' to obtain the correlation scores, and then use the softmax function to calculate the position attention weights S ij ∈R N×N , and the calculation formula is: ......Formula (2); In formula (2), B' T is the transpose of B', and B' T i is the element at the i-th position of B', C' T is the element at the j-th position of C', and S j is the weight of the element at the i-th position with respect to the element at the j-th position; ij ​ Then multiply D' by the position attention weight S ij to obtain a feature map E ∈ R C×N , where N = H × W, and the formula for the j-th position element E j is as follows: ......Formula (3); In formula (3), D' i is the i-th position element of D'; Then, after reshaping the feature map E that enhances pixel dependencies, it is added to the original feature map A, and thus the enhanced feature map F ∈ R is obtained. C×H×W ; Step 2: Extract the context information of the feature map P1 through a hybrid dilated spatial pyramid pooling module. The hybrid dilated spatial pyramid pooling module includes a plurality of parallel dilated convolution branches, and each dilated convolution branch has a dilated convolution tandem structure formed by sequentially connecting multiple layers of dilated convolutions in series. The dilation rate r of the i-th layer of dilated convolution in any one of the dilated convolution branches is i determined according to the following method: For a dilated convolution branch formed by sequentially connecting n layers of dilated convolutions in series, first, a set of preset dilation rates {r1, r2, r3...... r n} is given, where r1 = 1, that is, the dilation rate of the first layer of dilated convolution is 1; Then, starting from the nth layer of the dilated convolution branch, successively calculate the maximum distance M between the dilation rates of two adjacent layers in the dilated convolution branch i , and the calculation formula is as follows: ......Formula (4); For the first calculation, let M n = r n , substitute M n and r n into Equation (4) to calculate M n-1 , and then substitute M n-1 and r n-1 successively into Equation (4) until M2 is obtained; Compare M2 with the convolution kernel size k of the second dilated convolution in the dilated convolution branch of the hole. When M2 ≤ k, directly adopt the preset dilation rates {r1, r2, r3...... r n}, when M2 > k, re-give the preset dilation rate and calculate again; Step 3: Concatenate the output of the first 1×1 convolution of the hybrid atrous spatial pyramid pooling module, the outputs of multiple atrous convolution branches, and the pooling output to obtain an output feature map; Step 4: Pass the feature map in Step 3 to a Softmax classifier for pixel-level semantic segmentation to achieve anti-interference detection of infrared targets.

Citation Information

Patent Citations

  • Target detection method and device based on pyramid integration and attention enhancement

    CN116740376A

  • Image region localization method, image region localization apparatus, and medical image processing device

    US20210225027A1