Infrared Small Target Detection Method Based on Collaborative Fusion Comparison of Multiple Mechanism Attention

Through the multi-mechanical attention synergistic fusion structure and standardized loss function, the problems of low accuracy and poor robustness in infrared small object detection are solved, and high-precision small object detection is achieved in complex backgrounds.

CN115546610BActive Publication Date: 2025-07-25NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211310514.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2025-07-25
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

The existing infrared small object detection methods have low detection accuracy and poor robustness in complex backgrounds, making it difficult to effectively separate small objects from backgrounds. The traditional methods rely on prior knowledge and their performance decreases when the environment changes.

Method used

The infrared small object detection method of multi-mechanical attention collaborative fusion is adopted. Through multi-layer feature extraction, multi-scale local contrast and multi-mechanical attention fusion structure, combined with weak target channel attention and pixel attention mechanism, the underlying detail features and deep semantic features are enhanced, and sub-pixel convolution upsampling and standardized loss functions are used to improve detection accuracy.

Benefits of technology

It significantly improves the detection accuracy and robustness of infrared small targets, and can accurately separate small targets in complex backgrounds, reduce false alarms, enhance feature utilization, and achieve more accurate target shape and position prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115546610B_ABST
    Figure CN115546610B_ABST
Patent Text Reader

Abstract

This application relates to an infrared small target detection method based on multi-mechanism attention collaborative fusion and contrast. The method includes: proposing an infrared small target detection network, which includes a multi-layer feature extraction structure that extracts multiple feature maps from the input infrared image from the bottom layer to the deep layer, a multi-scale local contrast structure that converts each feature map into a feature map with local contrast, and a multi-mechanism attention collaborative fusion structure that fuses the bottom-layer detail features and the deep-layer semantic features through a weak target channel attention mechanism and a pixel attention mechanism. When training this network, the output image is probabilistically processed and then the loss function is calculated. Using this method can calculate the loss more accurately, solve the imbalance problem between the target and the background, and significantly improve the detection effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of infrared image detection technology, and particularly to an infrared small target detection method based on multi-mechanism attention collaborative fusion and contrast. Background Art

[0002] Infrared small target detection is a key technology in infrared search and tracking systems, which is widely used in many fields such as precision guidance, early warning, and maritime surveillance. Compared with other imaging methods, infrared imaging has the advantages of long detection range, strong anti-interference ability, and clear images. However, due to the long detection range, the detected targets have problems such as few occupied pixels, lack of shape, texture, and color information. At the same time, the targets are usually submerged in complex backgrounds and are greatly affected by changes in the surrounding environment. Therefore, how to separate infrared small targets from complex background clutter remains a challenging topic.

[0003] Infrared small target detection is mainly divided into two categories: model-driven methods and data-driven methods. Model-driven methods have made relatively in-depth development in the past few decades, and several types of systematic methods have been proposed, mainly including methods based on background suppression, methods based on the human visual system, methods based on optimization, etc. Methods based on background suppression assume that infrared images are continuously changing, and the presence of small targets breaks the original continuity. Filtering and morphological methods are used for background suppression and target separation. Methods based on the human visual system assume that there is a large local contrast between the target and the background, and the target can be detected through the local difference between the target and the background. Methods based on optimization adopt the idea of matrices, assuming that the target is a sparse matrix and the background is a low-rank matrix, and transform target detection into an optimization problem of separating low-rank matrices and sparse matrices. However, these traditional model-driven methods focus on the physical characteristics of the target, and various methods adopt different assumptions, relying heavily on prior knowledge and artificially set functions, and are difficult to handle when the characteristics of the real scene change. When the real scene does not meet the detection conditions, the detection performance will be greatly reduced. Although new algorithms emerge in an endless stream, there are still problems such as low detection accuracy, large influence by the environment, and poor robustness. Different from model-driven methods, data-driven methods have only been developed to a certain extent in recent years due to the lack of public datasets for infrared small targets, and have achieved better results than model-driven methods. However, most methods still make fine-tuning based on general object detection algorithms, ignoring the inherent physical characteristics of infrared small targets. Moreover, the pixels occupied by infrared small targets are usually much smaller than those of general targets. Directly applying these methods without combining the physical characteristics of infrared small targets for detection is likely to result in the loss of deep small targets. Summary of the Invention

[0004] Based on this, it is necessary to provide an infrared small target detection method based on multi-mechanism attention collaborative fusion and contrast for the above-mentioned technical problems, which can improve the accuracy of small target detection.

[0005] An infrared small target detection method based on multi-mechanism attention collaborative fusion and contrast, the method comprising:

[0006] Obtain an infrared image training set, where the infrared image training set includes multiple infrared training images containing small targets;

[0007] Input each of the infrared training images into an infrared small target detection network, the infrared small target detection network including a multi-layer feature extraction structure, a multi-scale local contrast structure, and a multi-mechanism attention collaborative fusion structure. Among them, the multi-layer feature extraction structure extracts features from the input image from the bottom layer to the deep layer to obtain multiple feature maps, the multi-scale local contrast structure converts each of the feature maps into a feature map with local contrast, and the multi-mechanism attention collaborative fusion structure enhances and fuses the bottom-layer detail features and deep-layer semantic features through a weak target channel attention mechanism and a pixel attention mechanism to obtain an output image;

[0008] Perform probability processing on the output image to obtain a probability output image, and construct a loss function according to the probability output image;

[0009] Train the infrared small target detection network according to the loss function to obtain a trained infrared small target detection network;

[0010] Obtain an infrared image to be detected, input the infrared image into the trained infrared small target detection network, detect small targets in the infrared image, and predict the shape and location of the small targets.

[0011] In one embodiment, the multi-layer feature extraction structure adopts an improved ResNet-20 network;

[0012] The multi-layer feature extraction structure includes three feature extraction units, which respectively output a bottom-layer feature map, a middle-layer feature map, and a high-layer feature map with different sizes.

[0013] In one embodiment, the multi-scale local contrast structure includes three local contrast units with different sizes, which respectively process the bottom-layer feature map, the middle-layer feature map, and the high-layer feature map to correspondingly obtain a bottom-layer local contrast feature map, a middle-layer local contrast feature map, and a high-layer local contrast feature map.

[0014] In one embodiment, the multi-mechanism attention collaborative fusion structure includes a first multi-mechanism fusion unit and a second multi-mechanism fusion unit;

[0015] The high-level local contrast feature map and the middle-level local contrast feature map are input into the first multi-mechanism fusion unit for fusion, and an intermediate fusion feature map is output;

[0016] The intermediate fusion feature map and the low-level local contrast feature map are input into the second multi-mechanism fusion unit for fusion, and a fusion feature map is output.

[0017] In one embodiment, in the first multi-mechanism fusion unit and the second multi-mechanism fusion unit, the high-level local contrast feature map and the intermediate fusion feature map are processed by the weak target channel attention mechanism, and the middle-level local contrast feature map and the low-level local contrast feature map are processed by the pixel attention mechanism.

[0018] In one embodiment, in the first multi-mechanism fusion unit and the second multi-mechanism fusion unit, the following formula is used to process the input features:

[0019] F(x,y) = y + (x·σ(C1D(τ(σ(x)))))·σ(β(PWConv2(δ(β(PWConv1(γ(y)))))))

[0020] In the above formula, x represents the high-level local contrast feature map or the intermediate fusion feature map, y represents the middle-level local contrast feature map or the low-level local contrast feature map, σ represents the Sigmoid function, τ represents global max pooling, C1D represents 1D convolution, β represents batch normalization, PWConv represents Point-Wise convolution, δ represents the ReLU function, and γ represents global average pooling.

[0021] In one embodiment, the infrared small target detection network further includes a sub-pixel convolutional upsampling structure, and the sub-pixel convolutional upsampling structure includes a first upsampling unit and a second upsampling unit;

[0022] The high-level local contrast feature map is upsampled by the first upsampling unit and then input into the first multi-mechanism fusion unit;

[0023] The intermediate fusion feature map is upsampled by the second upsampling unit and then input into the second multi-mechanism fusion unit.

[0024] In one embodiment, the infrared small target detection network is trained according to the loss function, and the trained infrared small target detection network obtained includes:

[0025] The gradient of the loss function is calculated, and the parameters of the infrared small target detection network are corrected according to the direction of the calculation result until convergence, and the trained infrared small target detection network is obtained.

[0026] An infrared small target detection device based on multi - mechanism attention collaborative fusion and contrast, the device comprising:

[0027] A training set acquisition module, configured to acquire an infrared image training set, where the infrared image training set includes multiple infrared training images containing small targets;

[0028] A feature fusion module, configured to input each of the infrared training images into an infrared small target detection network, the infrared small target detection network including a multi - layer feature extraction structure, a multi - scale local contrast structure, and a multi - mechanism attention collaborative fusion structure. Among them, the multi - layer feature extraction structure extracts features from the input image from the bottom layer to the deep layer to obtain multiple feature maps, the multi - scale local contrast structure converts each of the feature maps into a feature map with local contrast, and the multi - mechanism attention collaborative fusion structure enhances and fuses the bottom - layer detail features and the deep - layer semantic features through a weak target channel attention mechanism and a pixel attention mechanism to obtain an output image;

[0029] A loss function construction module, configured to perform probability processing on the output image to obtain a probability - processed output image, and construct a loss function according to the probability - processed output image;

[0030] An infrared small target detection network training module, configured to train the infrared small target detection network according to the loss function to obtain a trained infrared small target detection network;

[0031] A small target detection module, configured to acquire an infrared image to be detected, input the infrared image into the trained infrared small target detection network, detect small targets in the infrared image, and predict the shape and location of the small targets.

[0032] A computer device, comprising a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0033] Acquire an infrared image training set, where the infrared image training set includes multiple infrared training images containing small targets;

[0034] Input each of the infrared training images into an infrared small target detection network, which includes a multi-layer feature extraction structure, a multi-scale local contrast structure, and a multi-mechanism attention collaborative fusion structure. Among them, the multi-layer feature extraction structure extracts features from the input image from the bottom layer to the deep layer to obtain multiple feature maps. The multi-scale local contrast structure converts each of the feature maps into a feature map with local contrast. The multi-mechanism attention collaborative fusion structure enhances and fuses the bottom-layer detail features and deep-layer semantic features through a weak target channel attention mechanism and a pixel attention mechanism to obtain an output image;

[0035] Perform probability processing on the output image to obtain a probability output image, and construct a loss function based on the probability output image;

[0036] Train the infrared small target detection network according to the loss function to obtain a trained infrared small target detection network;

[0037] Obtain an infrared image to be detected, input the infrared image into the trained infrared small target detection network, detect small targets in the infrared image, and predict the shape and location of the small targets.

[0038] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the following steps are implemented:

[0039] Obtain an infrared image training set, which includes multiple infrared training images containing small targets;

[0040] Input each of the infrared training images into an infrared small target detection network, which includes a multi-layer feature extraction structure, a multi-scale local contrast structure, and a multi-mechanism attention collaborative fusion structure. Among them, the multi-layer feature extraction structure extracts features from the input image from the bottom layer to the deep layer to obtain multiple feature maps. The multi-scale local contrast structure converts each of the feature maps into a feature map with local contrast. The multi-mechanism attention collaborative fusion structure enhances and fuses the bottom-layer detail features and deep-layer semantic features through a weak target channel attention mechanism and a pixel attention mechanism to obtain an output image;

[0041] Perform probability processing on the output image to obtain a probability output image, and construct a loss function based on the probability output image;

[0042] Train the infrared small target detection network according to the loss function to obtain a trained infrared small target detection network;

[0043] Obtain the infrared image to be detected, input the infrared image into the trained infrared small target detection network, detect the small targets in the infrared image, and predict the shape and location of the small targets.

[0044] The above infrared small target detection method based on multi-mechanism attention collaborative fusion and contrast trains by inputting each infrared training image into the infrared small target detection network. The infrared small target detection network includes a multi-layer feature extraction structure, a multi-scale local contrast structure, and a multi-mechanism attention collaborative fusion structure. Among them, the multi-layer feature extraction structure extracts features from the input image from the bottom layer to the deep layer to obtain multiple feature maps. The multi-scale local contrast structure converts each feature map into a feature map with local contrast. The multi-mechanism attention collaborative fusion structure enhances and fuses the bottom-layer detail features and deep-layer semantic features through the weak target channel attention mechanism and the pixel attention mechanism to obtain an output image, performs probability processing on the output image to obtain a probability output image, constructs a loss function based on the probability output image, and trains the infrared small target detection network according to the loss function to obtain a trained infrared small target detection network. In actual application, only need to input the infrared image to be detected into the trained infrared small target detection network, then the shape and location of the small target can be predicted. In this method, a multi-mechanism attention collaborative fusion structure is proposed to dynamically modulate the bottom-layer detail features by enhancing the deep-layer features through the weak target channel attention mechanism and the pixel attention mechanism, enhancing the infrared small targets while improving the feature utilization rate. And the standardized loss function performs maximum and minimum normalization on the network output to reflect the relative difference, calculates the loss more accurately, and significantly improves the detection effect. Description of the Drawings

[0045] Figure 1 It is a schematic flowchart of the infrared small target detection method based on multi-mechanism attention collaborative fusion and contrast in an embodiment;

[0046] Figure 2 It is a schematic structural diagram of the infrared small target detection network in an embodiment;

[0047] Figure 3 It is a schematic structural diagram of the multi-mechanism attention collaborative fusion structure in an embodiment;

[0048] Figure 4 It is a schematic diagram of the structure of the weak target channel attention mechanism in an embodiment;

[0049] Figure 5 It is a schematic diagram of the structure of the pixel attention mechanism in an embodiment;

[0050] Figure 6 It is a schematic diagram of the sub-pixel convolutional upsampling process in an embodiment;

[0051] Figure 7 Schematic diagram showing a partial display of the SIRST dataset in an experiment;

[0052] Figure 8 Schematic diagram showing a partial display of the IDTAT dataset in an experiment;

[0053] Figure 9 Schematic diagram of the qualitative results of different detection methods for scenario 1 in an experiment;

[0054] Figure 10 Schematic diagram of the qualitative results of different detection methods for scenario 2 in an experiment;

[0055] Figure 11 Schematic diagram of the qualitative results of different detection methods for scenario 3 in an experiment;

[0056] Figure 12 Schematic diagram of the 3D visualization results of different detection methods for scenario 1 in an experiment;

[0057] Figure 13 Schematic diagram of the 3D visualization results of different detection methods for scenario 2 in an experiment;

[0058] Figure 14 Schematic diagram of the 3D visualization results of different detection methods for scenario 3 in an experiment;

[0059] Figure 15 Block diagram of an infrared small target detection device based on multi - mechanism attention collaborative fusion and comparison in an embodiment;

[0060] Figure 16 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0061] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0062] As Figure 1 shown, a method for infrared small target detection based on multi - mechanism attention collaborative fusion and comparison is provided, including the following steps:

[0063] Step S100, obtain an infrared image training set, where the infrared image training set includes multiple infrared training images containing small targets;

[0064] Step S110: Input each infrared training image into the infrared small target detection network. The infrared small target detection network includes a multi-layer feature extraction structure, a multi-scale local contrast structure, and a multi-mechanism attention collaborative fusion structure. Among them, the multi-layer feature extraction structure extracts features from the input image from the bottom layer to the deep layer to obtain multiple feature maps. The multi-scale local contrast structure converts each feature map into a feature map with local contrast. The multi-mechanism attention collaborative fusion structure enhances and fuses the bottom-layer detail features and deep-layer semantic features through a weak target channel attention mechanism and a pixel attention mechanism to obtain an output image;

[0065] Step S120: Perform probability processing on the output image to obtain a probability output image, and construct a loss function based on the probability output image;

[0066] Step S130: Train the infrared small target detection network according to the loss function to obtain a trained infrared small target detection network;

[0067] Step S140: Obtain the infrared image to be detected, input the infrared image into the trained infrared small target detection network, detect the small targets in the infrared image, and predict the shape and location of the small targets.

[0068] In this method, steps S100 to S130 are the process of training the infrared small target detection network, while step S140 is the process of using the trained infrared small target detection network for actual small target detection.

[0069] In step S100, the infrared training images used to train the network all include small-sized targets, and there are also label images corresponding to various infrared training images. The label images are used to train the infrared small target detection network subsequently.

[0070] In step S110, the composition of the infrared small target detection network (MAFCNet) is further described, including a multi-layer feature extraction structure, a multi-scale local contrast structure, and a multi-mechanism attention collaborative fusion structure, as Figure 2 shown.

[0071] Specifically, the multi-layer feature extraction structure uses an improved ResNet-20 network. The multi-layer feature extraction structure includes three feature extraction units, Stage1, Stage2, and Stage3, which perform two downsamplings on the input image (infrared training image) in sequence, and respectively output a bottom-layer feature map, a middle-layer feature map, and a high-layer feature map with different sizes.

[0072] Specifically, the multi-scale local contrast structure (MLC) includes three local contrast units with different sizes, which respectively process the underlying feature map, the middle-level feature map, and the high-level feature map to obtain the underlying local contrast feature map X1, the middle-level local contrast feature map X2, and the high-level local contrast feature map X3.

[0073] Specifically, the multi-mechanism attention collaborative fusion structure includes a first multi-mechanism fusion unit (MAFM) and a second multi-mechanism fusion unit (MAFM). The high-level local contrast feature map and the middle-level local contrast feature map are input into the first multi-mechanism fusion unit for fusion, and an intermediate fusion feature map is output. Then, the intermediate fusion feature map and the underlying local contrast feature map are input into the second multi-mechanism fusion unit for fusion, and a fusion feature map is output.

[0074] In the first multi-mechanism fusion unit and the second multi-mechanism fusion unit, the high-level local contrast feature map and the intermediate fusion feature map are processed by the weak target channel attention mechanism, and the middle-level local contrast feature map and the underlying local contrast feature map are processed by the pixel attention mechanism.

[0075] Furthermore, in this method, a multi-mechanism attention collaborative fusion structure is used to fuse the underlying detail information and the deep semantic information. The multi-mechanism fusion unit includes two main parts: the weak target channel attention mechanism and the pixel attention mechanism, and element-wise addition is used for feature fusion. The overall structure F(x, y) = R H×W×C , such as Figure 3 shown, and its calculation formula is as follows:

[0076] F(x, y) = y + (x · σ(C1D(τ(σ(x))))) · σ(β(PWConv2(δ(β(PWConv1(γ(y))))))) (1)

[0078] In formula (1), x represents the high-level local contrast feature map or the intermediate fusion feature map, that is, the feature map output after high-level convolution, y represents the middle-level local contrast feature map or the underlying local contrast feature map, σ represents the Sigmoid function, τ represents global max pooling, C1D represents 1D convolution, β represents batch normalization, PWConv represents Point-Wise convolution, δ represents the ReLU function, and γ represents global average pooling. Formula (1) can also be described in a modular idea as follows:

[0079]

[0080] The motivation of this module is to fuse the underlying detail features and deep semantic features of images. Due to the lack of intrinsic features of infrared small targets, they will be submerged in the deeper layers of the network. Therefore, the underlying feature map Y is selected as the fusion benchmark, and then the weak target channel attention mechanism is used to process the deep features Figure X Global max pooling is used for semantic enhancement. The obtained result P(X) is regarded as the probability of the small target appearing within the receptive field, and used as the weight for the underlying feature L(Y) to guide the network to dynamically select detail features from the bottom layer. After adding local contrast to the original underlying feature map and deep feature map respectively, MAFM enhances the underlying detail features through pixel attention and enhances the deep semantic features through weak target channel attention. The enhanced semantic features guide the network to dynamically select detail features. Finally, a complete fused feature map Z′ is obtained, where ψ is sub-pixel convolution:

[0081]

[0082] The weak target channel attention mechanism is used to process the deep semantic feature map. As the network depth increases, the neural network can better understand the meaning of the scene and extract better semantic features, which can help the network distinguish background clutter from targets. However, as the network deepens, the possibility of losing target information increases. To solve this problem, a channel attention mechanism (WCA) is proposed in this method to enhance the deep semantic information and guide the network to dynamically select the underlying detail information. Its structure is as Figure 4 shown

[0083] Denote the output of the previous convolutional block as x ∈ R H×W×C , where W, H, and C are the width, height, and number of channels of the feature map. Then the weight p(x) of the WCA channel can be obtained by the following formula:

[0084] p(x) = σ(C1D(τ(σ(x))))) (4)

[0085] In formula (4), τ is global channel max pooling, τ(x) = max(x i,j,k ), i ∈ H, j ∈ W, k ∈ C, σ is the sigmoid function, and C1D is 1D convolution

[0086] Since the existing complex attention mechanisms are not applicable to infrared small target detection, and lightweight attention mechanisms such as Efficient Channel Attention (ECA) cannot effectively form attention for small targets lacking features, WCA combines the idea of lightweight attention mechanisms, normalizes the feature map using σ(x), and utilizes global maximum pooling to highlight deep semantic information, aiming to form effective attention for infrared small targets. At the beginning of the module, since the discrimination basis based on local contrast is the relative difference rather than the absolute value between pixels, σ(x) is used to normalize the feature map in this method. At the same time, due to the few intrinsic features and weak semantic information of infrared small targets, it is not appropriate to use global average pooling. Under the idea of regarding small targets as local sparse matrices, small targets should form singular values in the local image, so global maximum pooling is adopted to enhance infrared small targets.

[0087] In addition, based on the existing methods, it is noted that it is more important to avoid reducing the number of channels to preserve information than to consider channel compression without linear relationships. Therefore, in this method, in the weak target channel attention mechanism, the number of channels is not compressed.

[0088] The pixel attention mechanism is used to process the underlying feature map. Considering that in the infrared small target detection task, the target occupies a small proportion in the whole image and lacks shape and texture features, pointwise convolution (PWConv) is selected as the local up and down aggregator, which interacts the spatial information of the local channels, making the network pay more attention to the information features with local high contrast, so as to highlight infrared small targets. As Figure 5 shown, the pixel attention mechanism L aggregates the local feature context of pixels as:

[0089] L(Y) = σ(β(PWConv2(δ(β(PWConv1(γ(y))))))) (5)

[0090] In formula (5), σ represents the Sigmoid function, β represents batch normalization, PWConv represents point-wise convolution, δ represents the ReLU function, and γ represents global average pooling.

[0091] Among them, the kernel sizes of PWConv1 and PWConv2 are (C / r)×C×1×1 and C×(C / r)×1×1 respectively, forming a bottleneck structure. The pixel information of the local channels is aggregated through pointwise convolution, strengthening infrared small targets at the pixel level and effectively solving the problems of small target pixel proportion and lack of shape and texture features.

[0092] Since the attention weight map L(Y) has the same shape as the deep reinforcement feature map P(X), the method of element-wise multiplication can be used to fuse it with P(X), guiding the network to dynamically select the underlying detailed information with deep semantic information, and further enhancing the underlying small targets:

[0093]

[0094] In the infrared small target detection network, finally, the complete fused feature map is output by the second multi-mechanism fusion unit, and finally, the output image, that is, the predicted image, is obtained by restoration (Pridict) according to the fused feature map.

[0095] Since the feature maps X1, X2, and X3 generated by the multi-scale local contrast structure (MLC) have different sizes, upsampling is needed to adjust the size of the feature maps so that the sizes of the feature maps input to MAFA are the same.

[0096] Therefore, in this embodiment, the infrared small target detection network further includes a sub-pixel convolutional upsampling structure (Sub-Pixel Conv). Common upsampling methods include bilinear interpolation method, nearest neighbor method, mean interpolation method, etc. Since the infrared small target detection task has high requirements for target detailed information, the interpolation method has poor ability to retain image detailed information. Therefore, in this embodiment, the idea of super-resolution is introduced, and the sub-pixel convolution technology is used to implement the sampling process. This method takes a low-resolution feature map as input and obtains a high-resolution feature map through convolution and multi-channel pixel recombination. Specifically, it combines a single pixel point on the low-resolution multi-channel feature map into a high-resolution feature map unit, and each pixel point on the low-resolution feature map is equivalent to a sub-pixel on the high-resolution feature map. It can convert the low-resolution feature map N * (C * r * r * ) * H * W into a high-resolution feature map N * C * (H * r) * (W * r). Among them, N, C, W, H, and r respectively represent the upsampling number, channels, width, height, and multiple of the image. The sub-pixel convolutional upsampling method is as Figure 6 shown.

[0097] Specifically, the sub-pixel convolutional upsampling structure includes a first upsampling unit and a second upsampling unit. The high-level local contrast feature map is upsampled by the first upsampling unit and then input into the first multi-mechanism fusion unit, and the intermediate fused feature map is upsampled by the second upsampling unit and then input into the second multi-mechanism fusion unit.

[0098] In the infrared small target detection task based on local contrast, the recognition of small targets and the recognition of the background essentially rely on a relative value relationship. However, the traditional loss function uses the loss caused by the absolute value difference, which is not conducive to the contrast-based method. The traditional loss for the image output value p i,j is defined as follows:

[0099] p i,j =(σ(MAF(f,θ))) (7)

[0100] p i,j has the same distribution as the sigmoid activation function, and its loss response to the background is relatively large. Therefore, in this embodiment, a new measurement method, namely the normalized loss function, is proposed. First, the p i,j output by the last layer of the infrared small target detection network is probabilized. The specific calculation method is as follows:

[0101]

[0102] In formula (8), the maximum and minimum normalization of probabilities is performed on MAF(f,θ). The maximum and minimum normalization reflects the relative difference here, is not affected by the absolute value, and can make the background in the image tend to 0 and the infrared small target tend to 1, reflecting the relative difference between the target and the background, so as to calculate the loss more accurately.

[0103] After probabilizing p i,j , combined with solving the imbalance problem between the infrared small target and the background, the relative Soft-IoU loss function is adopted in this embodiment, and its definition is as follows:

[0104]

[0105] In formula (9), y i.j =R H×W is the label map, whose value is 0 or 1, and p i,j =R H×W is the probabilized output image. By converting predict into probability, the formed relative Soft-IoU loss function can accurately adapt to the detection method based on local contrast, has a certain universality, and can accurately reflect the position and shape of the segmented small target.

[0106] Based on the Soft-IoU loss function, the predicted value is normalized by the maximum and minimum probabilities, reflecting the relative difference between the target and the background, better applicable to the neural network integrating local contrast, solving the imbalance problem between the infrared small target and the background, and achieving excellent performance.

[0107] Finally, the gradient of the loss function is calculated, and the parameters of the infrared small target detection network are corrected according to the direction of the calculation result until convergence, and the trained infrared small target detection network is obtained.

[0108] In this paper, experiments are also conducted to verify the effectiveness of the infrared small target detection network (MAFCNet). First, the experimental settings are described, including the dataset, comparison networks, evaluation metrics, and implementation details. Then, MAFCNet is visually and numerically compared with other data-driven methods and model-driven methods to further prove that MAFCNet can accurately detect infrared small targets and verify its effectiveness.

[0109] Dataset: In this experiment, the open-source infrared small target dataset SIRST and the infrared image weak small aircraft target detection and tracking dataset open-sourced by National University of Defense Technology are used. SIRST contains original images and pixel-level label images, with a total of 427 different images and 480 scene instances from hundreds of real videos. In this paper, it is divided into three sets: the train set, the validation set, and the test set, and their proportions are 50%, 20%, and 30% respectively. The SIRST dataset is as Figure 7 shown. It can be seen from (c), (g), (h), (i), and (j) that the overall image is relatively dark, the small targets are in a complex background, the background is relatively blurred, and there are clutter interferences such as clouds and Rayleigh noise. In the infrared image weak small aircraft target detection and tracking dataset under the ground / air background, folders 6-12, 15, 17, and 21 with relatively complex backgrounds are selected, with a total of 5993 pictures. The IDTAT dataset is as Figure 8 shown. The scenes of this dataset are diverse and complex. The IDTAT dataset is divided into the train and test sets, with a ratio of 8:2. To make the dataset applicable to the evaluation metrics of this network, the dataset is improved, and the selected pictures are pixel-level labeled. Experiments show that the improved infrared image weak small aircraft target detection and tracking dataset under the ground / air background can be effectively applied to this network.

[0110] Comparison networks: To prove the effectiveness of MAFCNet, the method proposed in this paper is compared with other data-driven methods and traditional model-driven methods. Among the data-driven methods, the Feature Pyramid Network (FPN), the Attention Local Contrast Network (ALCNet), and the Attention-Guided Pyramid Context Network (APGCNet) are selected in this paper. These methods have the same loss function, optimizer, and other hyperparameters as the proposed method. For the traditional model-driven methods, two classic methods are selected in this paper, namely the MPCM algorithm based on the human visual system and the IPI algorithm based on optimization. Their detailed hyperparameter settings are listed in Table 1.

[0111] Table 1: Hyperparameter table of model-driven method

[0112]

[0113] Evaluation indicators: This algorithm is a segmentation-based infrared small target detection algorithm, and the detection result is a binary image after segmentation. Traditional infrared small target detection evaluation indicators such as signal-to-noise ratio and background suppression factor are not applicable here, so the segmentation task indicators IoU and nIoU are used to objectively evaluate the performance of the proposed network. IoU is an important indicator for evaluating the shape detection ability of the algorithm in pixel segmentation tasks. nIoU is a proposed infrared small target detection evaluation metric that can better balance the metrics between model-driven methods and data-driven methods, and is defined as:

[0114]

[0115] In formula (10), TP, T and P represent the true positive target, true target and positive target respectively.

[0116] Implementation details: This method is implemented based on PyTorch. The optimizer uses the ADAM optimizer, in which the weight decay coefficient is set to 0.0001. The initial learning rate is 0.001, and the decay strategy of poly is used. The batch size is set to 8, and the loss function of each model is the standardized Soft-IoU. The SIRST dataset is trained for 300 epochs, and the infrared image weak aircraft target detection and tracking dataset under ground / air background is trained for 40 epochs. In terms of hardware, Nvidia GeForceRTX2080Super GPU is used for training.

[0117] To demonstrate the superiority of this method, MAFCNet is compared with other state-of-the-art data-driven methods and traditional model-driven methods both quantitatively and qualitatively. The results are shown in Tables 2 and Figure 9 , Figure 10 and Figure 11 shown.

[0118] Table 2: Comparative test

[0119]

[0120] Quantitative results: Table 2 shows the IoU and nIoU values of the six methods. Obviously, the MAFCNet network proposed in this paper has achieved the best results compared with other networks, and the improvement effect is obvious. From the quantitative results, the performance of the data-driven algorithm is better than that of the model-driven algorithm, while the performance of the model-driven algorithm is not ideal.

[0121] The reasons are twofold:

[0122] 1. Model-driven algorithms have strong assumptions about the environment. However, the environments of the SIRST dataset and the IDTAT dataset are complex and variable, and the assumptions of model-driven algorithms are not fully met, resulting in poor performance.

[0123] 2. Traditional classical infrared small target detection methods pay more attention to the detection of target positions while ignoring the importance of complete target segmentation. The targets detected by these methods are often incomplete, so the performance is very poor.

[0124] Moreover, model-driven algorithms will bring a more serious problem - the generation of false alarms. Therefore, direct application in the early warning field will cause many problems. In contrast, data-driven methods, due to being unrestricted by prior knowledge, the results of the network depend on the target and background features it has learned, and can achieve a balance between accurate segmentation and false alarm reduction. In addition, data-driven methods can perform feature fusion. Existing networks basically all have feature fusion modules, only the feature fusion methods are different. It can be seen that the multi-mechanism attention collaborative fusion module designed in this paper for the characteristics of few inherent features and small target ratios of infrared small targets is effective. Experimental data show that the network proposed in this paper is superior to other advanced deep learning methods in suppressing the background, accurately detecting objects, and segmenting objects.

[0125] Qualitative results: As Figures 9 - 11 shown, the detection results of extremely weak targets in the dataset are presented, and the detection results of six methods are compared. The target area is magnified in the upper right corner for better display. To facilitate the observation of false alarms in the detection results, Figures 12 - 14 a three-dimensional display of the image is given.

[0126] It can be seen from the results that the MAFCNet proposed in this method achieves accurate target location output and shape segmentation. Model-driven methods are sensitive to noise and detect more false alarm areas. Other state-of-the-art data-driven methods will have missed detections and false alarm areas when the target is very weak. In addition, from the perspective of shape segmentation, MAFCNet produces more accurate shape segmentation and achieves better performance than other advanced data-driven methods.

[0127] In the above infrared small target detection method based on multi-mechanism attention collaborative fusion and contrast, a multi-mechanism attention collaborative fusion module is proposed. This module consists of a weak target channel attention mechanism and a pixel attention mechanism. It enhances small targets on the underlying detailed feature map and fuses deep semantic information with the underlying detailed information to ensure that the network detects enhanced infrared small targets on the underlying feature map under the guidance of deep semantic features. It solves the problem of small target annihilation on the deep feature map in infrared small target detection. At the same time, aiming at the problem that the loss function in the general infrared small target detection network is no longer applicable to the end-to-end detection method combining model-driven and data-driven, a normalized loss function is proposed to perform probability maximum and minimum normalization on the network output, enabling the loss function to focus on the relative values between pixel points, improving the convergence efficiency and detection accuracy of the network, which is also the key factor to improve the efficiency of the end-to-end detection method combining model-driven and data-driven. Also, aiming at the problem that infrared small targets occupy few pixels in the image and it is difficult to extract high-resolution effective features for target detection, the sub-pixel convolutional upsampling technology is adopted to enhance the feature representation of small targets, converting the low-resolution feature representation into a high-resolution recognizable feature representation, and improving the detection accuracy of the network for infrared small targets.

[0128] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0129] In one embodiment, as Figure 15 shown, an infrared small target detection device based on multi-mechanism attention collaborative fusion and contrast is provided, including: a training set acquisition module 200, a feature fusion module 210, a loss function construction module 220, an infrared small target detection network training module 230, and a small target detection module 240, where:

[0130] The training set acquisition module 200 is used to acquire an infrared image training set, and the infrared image training set includes multiple infrared training images containing small targets;

[0131] A feature fusion module 210 is configured to input each of the infrared training images into an infrared small target detection network, where the infrared small target detection network includes a multi-layer feature extraction structure, a multi-scale local contrast structure, and a multi-mechanism attention collaborative fusion structure. Among them, the multi-layer feature extraction structure extracts features from the input image from the bottom layer to the deep layer to obtain multiple feature maps, the multi-scale local contrast structure converts each of the feature maps into a feature map with local contrast, and the multi-mechanism attention collaborative fusion structure enhances and fuses the bottom-layer detail features and deep-layer semantic features through a weak target channel attention mechanism and a pixel attention mechanism to obtain an output image;

[0132] A loss function construction module 220 is configured to perform probability processing on the output image to obtain a probability output image, and construct a loss function according to the probability output image;

[0133] An infrared small target detection network training module 230 is configured to train the infrared small target detection network according to the loss function to obtain a trained infrared small target detection network;

[0134] A small target detection module 240 is configured to obtain an infrared image to be detected, input the infrared image into the trained infrared small target detection network, detect small targets in the infrared image, and predict the shapes and positions of the small targets.

[0135] For the specific limitations of the infrared small target detection device based on multi-mechanism attention collaborative fusion contrast, reference can be made to the limitations of the infrared small target detection method based on multi-mechanism attention collaborative fusion contrast in the above text, which will not be elaborated here. Each module in the above infrared small target detection device based on multi-mechanism attention collaborative fusion contrast can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0136] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as shown in Figure 16As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements an infrared small target detection method based on multi-mechanism attention collaborative fusion and contrast. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0137] Those skilled in the art can understand that Figure 16 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0138] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0139] Obtain an infrared image training set, where the infrared image training set includes multiple infrared training images containing small targets;

[0140] Input each of the infrared training images into an infrared small target detection network. The infrared small target detection network includes a multi-layer feature extraction structure, a multi-scale local contrast structure, and a multi-mechanism attention collaborative fusion structure. Among them, the multi-layer feature extraction structure extracts features from the input image from the bottom layer to the deep layer to obtain multiple feature maps. The multi-scale local contrast structure converts each of the feature maps into a feature map with local contrast. The multi-mechanism attention collaborative fusion structure enhances and fuses the bottom-layer detail features and deep-layer semantic features through a weak target channel attention mechanism and a pixel attention mechanism to obtain an output image;

[0141] Perform probability processing on the output image to obtain a probability output image, and construct a loss function according to the probability output image;

[0142] Train the infrared small target detection network according to the loss function to obtain a trained infrared small target detection network;

[0143] Obtain an infrared image to be detected, input the infrared image into a trained infrared small target detection network, detect small targets in the infrared image, and predict the shape and location of the small targets.

[0144] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0145] Obtain an infrared image training set, where the infrared image training set includes multiple infrared training images containing small targets;

[0146] Input each of the infrared training images into an infrared small target detection network. The infrared small target detection network includes a multi-layer feature extraction structure, a multi-scale local contrast structure, and a multi-mechanism attention collaborative fusion structure. Among them, the multi-layer feature extraction structure extracts features from the input image from the bottom layer to the deep layer to obtain multiple feature maps. The multi-scale local contrast structure converts each of the feature maps into a feature map with local contrast. The multi-mechanism attention collaborative fusion structure enhances and fuses the bottom-layer detail features and deep-layer semantic features through a weak target channel attention mechanism and a pixel attention mechanism to obtain an output image;

[0147] Perform probability processing on the output image to obtain a probability output image, and construct a loss function according to the probability output image;

[0148] Train the infrared small target detection network according to the loss function to obtain a trained infrared small target detection network;

[0149] Obtain an infrared image to be detected, input the infrared image into a trained infrared small target detection network, detect small targets in the infrared image, and predict the shape and location of the small targets.

[0150] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0151] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0152] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. An infrared small target detection method based on the collaborative fusion contrast of multi-mechanism attention, characterized in that The method includes: Obtaining an infrared image training set, where the infrared image training set includes multiple infrared training images containing small targets; Inputting each of the infrared training images into an infrared small target detection network, the infrared small target detection network includes a multi-layer feature extraction structure, a multi-scale local contrast structure, and a multi-mechanism attention collaborative fusion structure. Among them, the multi-layer feature extraction structure extracts features from the input image from the bottom layer to the deep layer to obtain multiple feature maps, the multi-scale local contrast structure converts each of the feature maps into a feature map with local contrast, and the multi-mechanism attention collaborative fusion structure includes a first multi-mechanism fusion unit and a second multi-mechanism fusion unit. In the first multi-mechanism fusion unit and the second multi-mechanism fusion unit, the weak target channel attention mechanism processes the high-level local contrast feature map and the intermediate fusion feature map, and the pixel attention mechanism processes the middle-level local contrast feature map and the bottom-level local contrast feature map to obtain an output image, where the intermediate fusion feature map is obtained by fusing the high-level local contrast feature map and the middle-level local contrast feature map; Performing probability processing on the output image to obtain a probability output image, and constructing a loss function according to the probability output image; Training the infrared small target detection network according to the loss function to obtain a trained infrared small target detection network; Obtaining an infrared image to be detected, inputting the infrared image into the trained infrared small target detection network, detecting small targets in the infrared image, and predicting the shape and location of the small targets.

2. The infrared small target detection method according to claim 1, characterized in that The multi-layer feature extraction structure uses an improved ResNet-20 network; The multi-layer feature extraction structure includes three feature extraction units, which respectively output a bottom-level feature map, a middle-level feature map, and a high-level feature map with different sizes.

3. The infrared small target detection method according to claim 2, characterized in that, The multi-scale local contrast structure includes three local contrast units with different sizes, which respectively process the bottom-level feature map, the middle-level feature map, and the high-level feature map to obtain a bottom-level local contrast feature map, a middle-level local contrast feature map, and a high-level local contrast feature map.

4. The infrared small target detection method according to claim 3, characterized in that, The multi-mechanism attention collaborative fusion structure includes a first multi-mechanism fusion unit and a second multi-mechanism fusion unit; The high-level local contrast feature map and the middle-level local contrast feature map are input into the first multi-mechanism fusion unit for fusion, and the intermediate fusion feature map is output; The intermediate fusion feature map and the bottom-level local contrast feature map are input into the second multi-mechanism fusion unit for fusion, and the fusion feature map is output.

5. The infrared small target detection method according to claim 4, wherein In the first multi-mechanism fusion unit and the second multi-mechanism fusion unit, the following formula is used to process the input features: ; In the above formula, represents the high-level local contrast feature map or the intermediate fusion feature map, represents the middle-level local contrast feature map or the low-level local contrast feature map, represents the Sigmoid function, represents global max pooling, represents 1D convolution, represents batch normalization, represents Point-Wise convolution, represents the ReLU function, represents global average pooling.

6. The infrared small target detection method according to claim 5, wherein The infrared small target detection network further includes a sub-pixel convolutional upsampling structure, and the sub-pixel convolutional upsampling structure includes a first upsampling unit and a second upsampling unit; The high-level local contrast feature map is upsampled by the first upsampling unit and then input into the first multi-mechanism fusion unit; The intermediate fusion feature map is upsampled by the second upsampling unit and then input into the second multi-mechanism fusion unit.

7. The infrared small target detection method according to claim 6, wherein Training the infrared small target detection network according to the loss function, the trained infrared small target detection network obtained includes: Calculating the gradient of the loss function, correcting the parameters of the infrared small target detection network according to the direction of the calculation result until convergence, and obtaining the trained infrared small target detection network.