Infrared small target detection method based on multi-branch interactive guidance
By adopting multi-branch interactive guidance method in infrared small object detection, the depth characteristics of infrared images are extracted using lossless encoder, edge detection and positioning branches, and transformed through multi-branch guidance fusion module, the problem of insufficient accuracy and robustness of infrared small object detection in complex environments in the prior art is solved, and more efficient and reliable infrared small object detection is achieved.
Patent Information
- Application Number
- CN202510053387.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
AI Technical Summary
The existing infrared small object detection method has unstable detection effect under complex background and noise interference, especially in low contrast and blurred images, and it is difficult to extract effective target features. In the detection of small objects, the deep learning methods have problems such as low contrast between the target and background, large scale differences, and serious noise interference, resulting in insufficient detection accuracy and robustness.
The infrared small object detection method based on multi-branch interactive guidance is adopted. By obtaining the infrared image to be detected, and processing images based on the pre-constructed lossless encoder branch, edge detection branch and positioning branch, the depth features of multi-scale and multi-information are extracted respectively. The depth features are transformed using the multi-branch guide fusion module to obtain the small object detection results.
It significantly improves the accuracy and robustness of infrared small object detection, can accurately identify small objects under complex background and noise interference, improves detection reliability, and maintains a reasonable balance in model complexity and computing overhead.
Smart Images

Figure CN120014326A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer vision and image processing, and in particular to a method and device for detecting small infrared targets based on multi-branch interactive guidance. Background Art
[0002] In recent years, the application of infrared small target detection in military surveillance, security monitoring, disaster relief and other fields has gradually increased, especially at night and in complex environments. It is of irreplaceable importance. Since infrared images can provide clear image information under low light or dark conditions, the detection of infrared small targets can effectively supplement the deficiencies of traditional visible light target detection and has broad application prospects. In modern battlefields and complex monitoring scenes, the real-time detection of infrared small targets is not only an important means to ensure safety, but also a key factor in improving combat effectiveness. Therefore, how to accurately and real-time detect infrared small targets has become the current research focus in the field of target detection.
[0003] However, existing infrared small target detection methods face many challenges. Traditional infrared small target detection methods mostly rely on manual features and traditional image processing technology. Although these methods can identify targets to a certain extent, the detection effect is unstable under complex backgrounds and noise interference, especially in low-contrast and blurred images, it is difficult to extract effective target features. Although deep learning-based detection methods have made great progress, most methods often have problems such as low contrast between targets and backgrounds, large scale differences, and severe noise interference when facing small targets, resulting in insufficient detection accuracy and robustness. Existing methods generally lack the ability to fully utilize multi-scale and multi-information, especially in complex scenes, it is difficult to take into account the detection of small targets of different scales, especially in the critical moment of accurate target recognition. Summary of the invention
[0004] The main purpose of the present application is to provide a method and device for infrared small target detection based on multi-branch interactive guidance, aiming at improving the accuracy of infrared target detection in the case of lack of clear morphological and texture features in the image.
[0005] To achieve the above-mentioned purpose, the present application provides an infrared small target detection method based on multi-branch interactive guidance, including: obtaining an infrared image to be detected; processing the infrared image to be detected based on a pre-constructed lossless encoder branch to obtain a first infrared feature map, wherein the first encoder of the lossless encoder branch is constructed based on a downsampling module after Haar wavelet transform processing; processing the infrared image to be detected based on a pre-constructed edge detection branch to obtain a second infrared feature map, wherein the edge detection branch includes a second encoder and a decoder that are sequentially communicated and connected, and the encoder and the decoder are both constructed based on global average pooling and a multi-layer perceptron; processing the infrared image to be detected based on a pre-constructed positioning branch to obtain a third infrared feature map, wherein the positioning branch is constructed based on a multi-core pooling module and multiple first convolution modules; using a pre-constructed multi-branch guided fusion module to respectively extract the respective depth features of the first infrared feature map, the second infrared feature map and the third infrared feature map, and transforming each of the depth features based on the interactive guidance module to obtain a small target detection result in the infrared image.
[0006] Optionally, the lossless encoder branch includes: a lossless encoder and a multi-level feature fusion module that are communicatively connected in sequence.
[0007] Optionally, the multi-level feature fusion module includes: at least one cascaded residual module and multiple cascaded first upsampling modules; wherein, the input end of the at least one cascaded residual module is communicatively connected to the output end of the first encoder, the output end is communicatively connected to the input end of the cascaded multiple first upsampling modules, and the output end of the first encoder is communicatively connected to the input end of the multiple first upsampling modules, and the number of the first upsampling modules is the same as the number of downsampling in the downsampling module.
[0008] Optionally, the downsampling module includes: a feature pyramid module, which is used to extract a multi-scale feature map of the infrared image to be detected, and in the downsampling stage of the feature pyramid module, use Haar wavelet transform to process each of the multi-scale feature maps to obtain four sub-feature maps, and combine the four sub-feature maps to obtain a multi-scale feature map with the same number of channels as the original multi-scale feature map. Figure 4 times the multi-scale feature maps.
[0009] Optionally, the infrared image to be detected is processed based on the pre-built positioning branch to obtain a third infrared feature map, including: using multiple pooling modules with different kernel sizes to perform maximum pooling operations on the input infrared image respectively, and correspondingly obtaining target detection images of different sizes; using the difference between the infrared image and each target detection image to obtain each target detection image after background noise suppression; performing aggregation processing on each target detection image to obtain an aggregated target detection image; and sequentially using each of the first convolution modules with residual connection to extract positioning features from the aggregated target detection image to obtain a third infrared feature map.
[0010] Optionally, a fusion module is provided between the lossless encoder branch and the edge detection branch, and two input ends of the fusion module are respectively connected to the first encoder and the second encoder, the first encoder outputs a second sub-feature map, and the pre-built edge detection branch is used to process the infrared image to obtain a second infrared feature map, including: using global average pooling and a multi-layer perceptron to protect the edge details of the small target in the infrared image to obtain a first sub-feature map; using the fusion module to perform feature fusion on the second sub-feature map and the first sub-feature map to obtain a third sub-feature map, and performing residual connection on the first sub-feature map and the third sub-feature map to obtain a fourth sub-feature map; using the decoder to perform a decoding operation on the fourth sub-feature map to obtain a second infrared feature map.
[0011] Optionally, the decoder includes a plurality of second upsampling modules, convolution modules and activation modules that are communicatively connected in sequence, and each of the upsampling modules is communicatively connected to the encoder, and the use of the decoder to perform a convolution operation on the fourth sub-feature map to obtain a second infrared feature map includes: using a plurality of the second upsampling modules to upsample the input fourth sub-feature map to obtain a fifth sub-feature map of the same size as the infrared image; using the convolution module to extract edge features from the fifth sub-feature map to obtain a sixth sub-feature map; and using the activation module to perform edge feature enhancement on the sixth sub-feature map to obtain a second infrared feature map.
[0012] Optionally, the pre-built lossless encoder branch processes the infrared image to be detected to obtain a first infrared feature map, including: using the lossless encoder to losslessly downsample the infrared image to obtain a lossless coding image; residually connecting the lossless coding image and the third sub-feature map to obtain a seventh sub-feature map; using multiple first upsampling modules to sequentially perform size restoration processing on the seventh sub-feature map to obtain an infrared feature map.
[0013] Optionally, the multi-branch guided fusion module includes a convolution guided submodule and a Sobel guided submodule, and the pre-built multi-branch guided fusion module is used to respectively extract the respective depth features of the first infrared feature map, the second infrared feature map and the third infrared feature map, and each of the depth features is transformed based on the interactive guided module to obtain the small target detection result in the infrared image, including: splicing the first infrared feature map, the second infrared feature map and the third infrared feature map to obtain a fourth infrared feature map with depth features; using the convolution guided submodule to process the fourth infrared feature map to obtain a first guided feature map; calculating the difference between the first guided feature map and the first infrared feature map to obtain a fifth infrared feature map; processing the fourth infrared feature map based on the Sobel guided submodule to obtain a second guided feature map; calculating the second guided feature map and the fifth infrared feature map to obtain a sixth infrared feature map; using the residual module to process the sixth feature map to obtain the small target detection result in the infrared image.
[0014] In addition, to achieve the above-mentioned purpose, the present application also provides an infrared small target detection device based on multi-branch interactive guidance, including: an acquisition module, used to acquire an infrared image to be detected; a lossless detection module, used to process the infrared image to be detected based on a pre-constructed lossless encoder branch to obtain a first infrared feature map, wherein the first encoder of the lossless encoder branch is constructed based on a downsampling module after Haar wavelet transform processing; an edge detection module, used to process the infrared image to be detected based on a pre-constructed edge detection branch to obtain a second infrared feature map, wherein the edge detection branch includes a second encoder and a decoder that are sequentially communicated and connected, and the encoder and the decoder are both constructed based on global average pooling and a multi-layer perceptron; a positioning module, used to process the infrared image to be detected based on a pre-constructed positioning branch to obtain a third infrared feature map, wherein the positioning branch is constructed based on a multi-core pooling module and multiple first convolution modules; a fusion module, used to use a pre-constructed multi-branch guided fusion module to respectively extract the respective depth features of the first infrared feature map, the second infrared feature map and the third infrared feature map, and transform each of the depth features based on the interactive guidance module to obtain a small target detection result in the infrared image.
[0015] The embodiment of the present application proposes an infrared small target detection method and device based on multi-branch interactive guidance, which obtains an infrared image to be detected; processes the infrared image to be detected based on a pre-constructed lossless encoder branch to obtain a first infrared feature map, wherein the lossless encoder is constructed based on a downsampling module and a multi-level feature fusion module; processes the infrared image to be detected based on a pre-constructed edge detection branch to obtain a second infrared feature map, wherein the edge detection branch includes a second encoder and a decoder that are sequentially communicated and connected, and the second encoder and the decoder are both constructed based on global average pooling and a multi-layer perceptron; processes the infrared image to be detected based on a pre-constructed positioning branch to obtain a third infrared feature map, wherein the positioning branch is constructed with a multi-core pooling module and multiple convolution modules; uses a pre-constructed multi-branch guided fusion module to respectively extract the respective depth features of the first infrared feature map, the second infrared feature map and the third infrared feature map, and transforms each depth feature based on the interactive guidance module to obtain a small target detection result in the infrared image. The method consists of three branches: edge, positioning and detection, and each branch designs a special module from a unique perspective. In the detection branch, a multi-dimensional lossless encoder is introduced to mitigate feature loss in small targets. In the localization branch, a target localization strategy is proposed to explicitly identify candidate targets from images through a learnable multi-kernel model. In the edge branch, a simple pooling architecture is adopted to enhance the model's ability to preserve target shapes. In order to effectively utilize the knowledge of different branches, a mutually guided fusion module is introduced to adjust the information within and between branches, so that the branches can effectively share information, thereby improving the accuracy and robustness of infrared target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A flowchart diagram of an embodiment of an infrared small target detection method based on multi-branch interactive guidance is provided for this application;
[0017] Figure 2 A diagram of the architecture of an infrared small target detection model based on multi-branch interactive guidance provided in an embodiment of an infrared small target detection method based on multi-branch interactive guidance of the present application;
[0018] Figure 3 A detection result diagram of a lossless encoder branch provided in an embodiment of an infrared small target detection method based on multi-branch interactive guidance of the present application;
[0019] Figure 4 An infrared small target image detection result diagram provided for an embodiment of an infrared small target detection method based on multi-branch interactive guidance of the present application;
[0020] Figure 5 A structural block diagram is provided for an embodiment of an infrared small target detection device based on multi-branch interactive guidance of the present application.
[0021] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0022] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0023] Figure 1 This is a flowchart of an infrared small target detection method based on multi-branch interactive guidance provided in an embodiment of the present application, refer to Figure 1 The method may be executed by a processor of a server or a terminal, and the method may include the following execution process:
[0024] S10, obtaining an infrared image to be detected;
[0025] S20, processing the infrared image to be detected based on a pre-constructed lossless encoder branch to obtain a first infrared feature map, wherein the first encoder of the lossless encoder branch is constructed based on a downsampling module after Haar wavelet transform processing;
[0026] S30, processing the infrared image to be detected based on the pre-built edge detection branch to obtain a second infrared feature map, wherein the edge detection branch includes a second encoder and a decoder that are sequentially connected in communication, and the encoder and the decoder are both constructed based on global average pooling and a multi-layer perceptron;
[0027] S40, processing the infrared image to be detected based on the pre-constructed positioning branch to obtain a third infrared feature map, wherein the positioning branch is constructed based on a multi-core pooling module and a plurality of first convolution modules;
[0028] S50, using a pre-built multi-branch guided fusion module to respectively extract the respective depth features of the first infrared feature map, the second infrared feature map and the third infrared feature map, and transform each depth feature based on the interactive guided module to obtain the small target detection result in the infrared image.
[0029] The first encoder is a lossless encoder, and the second encoder is a common encoder.
[0030] The infrared small target detection method based on multi-branch interactive guidance provided in this embodiment effectively solves the limitations of the existing infrared small target detection technology in complex environments, and significantly improves the accuracy and robustness of small target detection by introducing a multi-branch interactive guidance mechanism. The method makes full use of information interaction at different scales, enhances the model's perception of small targets, and can maintain a high detection accuracy, especially in the case of low contrast and strong noise interference. The present application optimizes the model structure, especially the design of the interactive guidance mechanism, so that the network can accurately distinguish targets when facing various complex backgrounds, thereby improving the accuracy of target detection. Compared with traditional methods, the multi-branch interactive guidance network of the present application can not only effectively solve the scale problem in small target detection, but also overcome the influence of background noise on the detection results. Experimental results show that the optimized network has improved indicators such as detection accuracy, especially when dealing with small targets under complex backgrounds and weak contrast, the performance is more stable, and the reliability of detection is significantly improved. In addition, the method maintains a reasonable balance between model complexity and computational overhead, and can achieve efficient real-time detection. This technology provides a more accurate and reliable solution for infrared small target detection tasks, and is particularly suitable for application scenarios with high precision requirements such as military surveillance and security monitoring.
[0031] refer to Figure 2 In an embodiment of the present application, step S40 may include the following execution process:
[0032] S401, using a plurality of pooling modules with different kernel sizes to perform maximum pooling operations on the input infrared image respectively, and correspondingly obtaining target detection images of different sizes;
[0033] S402, using the difference between the infrared image and each target detection image, correspondingly obtaining each target detection image after background noise suppression;
[0034] S403, performing aggregation processing on each target detection image to obtain an aggregated target detection image;
[0035] S404, sequentially using the first convolution modules of the residual connection to extract positioning features from the aggregated target detection image to obtain a third infrared feature map.
[0036] refer to Figure 2 , the processor can retain the edge details of small targets and improve the accuracy of detection by executing steps S401 to S404. Specifically, for the construction of the positioning branch, the processor first introduces the maximum pool with different kernel sizes on the positioning branch to weaken unnecessary details. At the same time, it also reduces the brightness of the potential target. Then, using the I input of the original image, subtracting the processed image, the potential target existing in the original image is obtained, that is:
[0037] I c (i)=I input -M(I input ,i)
[0038] Where M(,) represents the maximum pooling operation under different kernel sizes. The maximum pooling operation is a learnable nonlinear filter. The nonlinear filter can analyze objects of different sizes based on the input image, thereby highlighting areas that are brighter than the contours in the original image and effectively suppressing background noise. On this basis, the filtered image I c Aggregation is performed to form an aggregated target detection image. Under the condition of minimal background interference, the aggregated target detection image is subjected to location feature extraction through several convolution blocks connected by residuals to obtain a third infrared feature map. It can be understood that the use of the pre-built localization branch can retain the edge details of small targets, thereby improving the accuracy of detection.
[0039] Continue to refer Figure 2 In an embodiment of the present application, a fusion module is provided between the lossless encoder branch and the edge detection branch, and two input ends of the fusion module are respectively connected to the first encoder and the second encoder. Step S30 may specifically include the following execution process:
[0040] S301, using global average pooling and a multi-layer perceptron to protect edge details of small targets in the infrared image, to obtain a first sub-feature map;
[0041] S302, using a fusion module to perform feature fusion on the second sub-feature map and the first sub-feature map to obtain a third sub-feature map, and performing residual connection on the first sub-feature map and the third sub-feature map to obtain a fourth sub-feature map;
[0042] S303: Use a decoder to decode the fourth sub-feature map to obtain a second infrared feature map.
[0043] It should be noted that global average pooling and multi-layer perceptron refer to the feature map output by the convolutional layer in the convolutional neural network (CNN) architecture, which usually contains rich spatial information and is large in size. Global average pooling is a dimensionality reduction operation that averages all pixel values of each channel of the feature map, thereby compressing the two-dimensional (or multi-dimensional) feature map into a one-dimensional vector. The edge detection branch constructed by the present application using global average pooling, multi-layer perceptron and decoder can perform edge detection on the pre-constructed original image. This operation is mainly used to retain the edge details of small targets, thereby improving the accuracy of detection. On the basis of the main branch of non-destructive detection, the positioning branch and the edge branch are further introduced, and each branch is responsible for processing intermediate features of different dimensions and different features.
[0044] The decoder includes a plurality of second upsampling modules, convolution modules and activation modules that are sequentially connected in communication, and each upsampling module is connected in communication with the encoder, and the decoder is used to perform a convolution operation on the fourth sub-feature map to obtain a second infrared feature map, including:
[0045] Using a plurality of second up-sampling modules to perform up-sampling processing on the input fourth sub-feature map to obtain a fifth sub-feature map having the same size as the infrared image;
[0046] The convolution module is used to extract edge features from the fifth sub-feature map to obtain the sixth sub-feature map;
[0047] The activation module is used to perform edge feature enhancement processing on the sixth sub-feature map to obtain a second infrared feature map.
[0048] In an embodiment of the present application, step S20 may include the following execution process:
[0049] S201, using a lossless encoder to perform lossless downsampling processing on the infrared image to obtain a lossless coded image;
[0050] S202, residually connect the lossless coding map and the third sub-feature map to obtain a seventh sub-feature map;
[0051] S203, using a plurality of first up-sampling modules to sequentially perform size restoration processing on the seventh sub-feature map to obtain an infrared feature map.
[0052] The lossless encoder is described in detail below. The lossless encoder branch may include: a lossless encoder and a multi-level feature fusion module that are sequentially communicatively connected. The multi-level feature fusion module may include at least one cascaded residual module and multiple cascaded first upsampling modules. If the residual module is not unique, the multiple residual modules are cascaded; wherein the input end of the at least one cascaded residual module is communicatively connected to the output end of the first encoder, and the output end is communicatively connected to the input end of the cascaded multiple first upsampling modules, and the output end of the first encoder is communicatively connected to the input end of the multiple first upsampling modules, and the number of first upsampling modules is the same as the number of downsampling in the downsampling module.
[0053] Specifically, the downsampling module can be a feature pyramid module, which is used to extract the multi-scale feature map of the infrared image to be detected, and in the downsampling stage of the feature pyramid module, the Haar wavelet transform is used to process each multi-scale feature map to obtain four sub-feature maps, and the four sub-feature maps are combined to obtain a channel number equal to the original multi-scale feature map. Figure 4 Multiple-scale feature maps.
[0054] For example, suppose an intermediate feature X∈R of the feature pyramid module C×W×H, then first use Haar wavelet transform to transform the original features into intermediate features X∈R C×W×H , where C, W, and H represent the number of channels, width, and height of the feature map, respectively, resulting in four components. The four components have low-frequency components and high-frequency components, including horizontal, vertical, and diagonal directions, namely:
[0055] X L ,X H ,X V ,X D =WT(X),
[0056] X′=[X L ;X H ;X V ;X D ],
[0057] Wherein, WT() is the Haar wavelet transform function, through this transformation, the size of each component is W / 2×H / 2. Then the four components are combined to form a new feature map, and the number of channels of the new feature map is increased four times. This process can effectively encode spatial information into channel dimensions. This operation realizes lossless transmission during the downsampling process of the feature pyramid and avoids the loss of image texture and details. Similarly, for other downsampling involved in the feature pyramid, this operation is performed to ensure that detail features are not easily lost. The present application introduces a lossless downsampling strategy and a multi-level feature fusion strategy in the detection branch to construct a multi-dimensional lossless detection encoder, in which wavelet transforms are used to extract and fuse multiple dimensions of features to ensure the retention of detail features during the downsampling process. Figure 2 The visualization results of the detection branch are given. Figure 3 It is a test result diagram of the nondestructive testing branch provided by the present invention.
[0058] In an embodiment of the present application, step S50 may include the following execution process: the multi-branch guided fusion module includes a convolution guided submodule and a Sobel guided submodule, and the pre-built multi-branch guided fusion module is used to respectively extract the depth features of the first infrared feature map, the second infrared feature map, and the third infrared feature map, and the depth features are transformed based on the interactive guided module to obtain the small target detection result in the infrared image, including:
[0059] S501, performing splicing processing on the first infrared feature map, the second infrared feature map and the third infrared feature map to obtain a fourth infrared feature map having a depth feature;
[0060] S502, using a convolution guidance submodule to process the fourth infrared feature map to obtain a first guidance feature map;
[0061] S503, calculating the difference between the first guide characteristic map and the first infrared characteristic map to obtain a fifth infrared characteristic map;
[0062] S504: Process the fourth infrared feature map based on the Sobel guidance submodule to obtain a second guidance feature map
[0063] S505, calculating the second guide characteristic map and the fifth infrared characteristic map to obtain a sixth infrared characteristic map;
[0064] S506: Process the sixth feature map using a residual module to obtain a small target detection result in the infrared image.
[0065] Among them, since the algorithm has three branches: detection, localization, and detection, different branches have different functions and cannot be directly combined. Therefore, this paper proposes a multi-branch guided fusion module. This module aims to optimize the integration of features and ensure that the specific information from each branch is fully utilized. The multi-branch guided fusion module is specifically composed of three different inputs: detection branch D, localization branch L, and edge branch e. Figure 2 As shown in the figure, the module first uses the deep features of these three parallel branches. Then, each feature is transformed using the interactive guidance block, which mainly includes two methods: convolution guidance and Sobel guidance, and finally the detection result is output after passing through the residual module. The multi-branch guidance fusion module introduced in this application enables each branch to effectively share information, thereby improving the accuracy and robustness of infrared target detection.
[0066] In addition, in the embodiment of the present application, the training of the infrared small target detection model based on multi-branch interactive guidance is also included. The optimizer used in the training of the present application adopts Adam, and the stochastic gradient descent method is used to train the network parameters by minimizing the loss function. Among them, the loss function is defined as:
[0067] Loss=||I pred -I gt ||1
[0068] Among them, I pred The predicted results, I gt represents the true value and ||·||1 is the L1 norm.
[0069] In the specific training, the public infrared small target data set is first loaded, and the data set is mainly divided into a training set and a test set. For the training set, the training set data is first enhanced by flipping and rotating. The true value is directly loaded, and the test set is directly used as input for prediction.
[0070] According to the results of the test set, the network parameters are adjusted. For example, the learning rate, momentum and other parameters are modified to continuously optimize the network and obtain the parameter configuration with the best effect on the test set as the final model parameters.
[0071] The network finally outputs the detection results as follows Figure 4 As shown in the figure, it can be seen that the method provided by the present application can well detect the specific position of the infrared target.
[0072] The designed multi-branch interactive guidance algorithm is trained through the public infrared small target data set. And by optimizing the loss function, the network can accurately identify infrared small targets under complex background and noise interference, and improve the positioning accuracy of the target.
[0073] Evaluation and optimization. Use an independent test set to evaluate the detection effect of the trained network model. Use multiple indicators to quantitatively evaluate the detection performance, optimize the shortcomings of the algorithm model, and improve its detection performance in practical applications.
[0074] The embodiment of the present application also provides a simulation process of an infrared small target detection model based on multi-branch interactive guidance. Specifically, the present invention is a process in which the central processor is The simulation was performed on an i7-6800K 3.40GHz CPU, NVIDIA GeForce GTX3090GPU, and Ubuntu operating system using Python software and PyTorch deep learning framework. We selected the IRSTD-1K dataset for algorithm verification, which comes from https: / / github.com / RuiZhang97 / ISNet. The IRSTD-1k dataset was collected under real conditions at a long distance and consists of 1001 images, each with a resolution of 512×512 pixels. These images are divided into a training set of 800 images and a test set of 201 images.
[0075] In order to prove the effectiveness of the infrared small target improvement method of the present invention, IoU, nIoU, AUC, Fa and Pd are used as evaluation criteria and tested on the public dataset IRSTD-1k. We selected AGPCNet, ISNet network and EGPNet network as comparison algorithms. By comparing the values of IoU, nIoU, AUC, Fa and Pd, we can intuitively evaluate the performance of various algorithms on target detection tasks. The comparison results are shown in Table 1:
[0076] Table 1 Comparison results of different detection algorithms on infrared small target dataset
[0077] method IoU iGO AUC Fa Pd AGPCNet 0.5602 0.5344 0.9356 17.1 0.9150 ISNet 0.6538 0.6372 0.8879 18.0 0.9226 EGPNet 0.6662 0.6800 0.8962 24.2 0.9495 Proposed method 0.6721 0.6804 0.9009 14.0 0.9428
[0078] Generally speaking, the higher the values of IoU, nIoU, AUC, Fa and Pd, the better the algorithm performs in the detection task. As can be seen from Table 1, on the public dataset, the method proposed in this invention can outperform other algorithms in several indicators. Figure 4 It is a partial detection result diagram of the method of the present invention.
[0079] refer to Figure 5 On the basis of the above-mentioned embodiments, the present application further provides an infrared small target detection device based on multi-branch interactive guidance, the infrared small target detection device 100 comprises an acquisition module 1001, a non-destructive detection module 1002, an edge detection module 1003, a positioning module 1004 and a fusion module 1005, wherein the acquisition module 1001 is used to acquire an infrared image to be detected; the non-destructive detection module 1002 is used to process the infrared image to be detected based on a pre-built lossless encoder branch to obtain a first infrared feature map, wherein the first encoder of the lossless encoder branch is constructed based on a downsampling module after Haar wavelet transform processing; the edge detection module 1003 is used to process the infrared image to be detected based on a pre-built edge detection branch to obtain a first infrared feature map The infrared image to be detected is processed to obtain a second infrared feature map, wherein the edge detection branch includes a second encoder and a decoder which are sequentially communicated and connected, and the encoder and the decoder are constructed based on global average pooling and a multi-layer perceptron; the positioning module 1004 is used to process the infrared image to be detected based on a pre-constructed positioning branch to obtain a third infrared feature map, wherein the positioning branch is constructed based on a multi-core pooling module and multiple first convolution modules; the fusion module 1005 is used to use the pre-constructed multi-branch guided fusion module to respectively extract the respective depth features of the first infrared feature map, the second infrared feature map and the third infrared feature map, and transform each depth feature based on the interactive guided module to obtain the small target detection result in the infrared image.
[0080] To achieve the above objectives, the present application also provides a computer-readable storage medium, which includes instructions, which, when executed on a computer, enables the computer to execute the infrared small target detection method based on multi-branch interactive guidance provided in the above embodiment.
[0081] To achieve the above-mentioned objectives, the present application also provides an electronic device, which includes: at least one processor, a memory and an input and output unit; wherein the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the infrared small target detection method based on multi-branch interactive guidance provided in any of the aforementioned embodiments.
[0082] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for infrared small target detection based on multi-branch interactive guidance, characterized in that: include: Acquire an infrared image to be detected; Processing the infrared image to be detected based on a pre-constructed lossless encoder branch to obtain a first infrared feature map, wherein a first encoder of the lossless encoder branch is constructed based on a downsampling module after Haar wavelet transform processing; Processing the infrared image to be detected based on a pre-built edge detection branch to obtain a second infrared feature map, wherein the edge detection branch includes a second encoder and a decoder that are sequentially connected in communication, and the encoder and the decoder are both constructed based on global average pooling and a multi-layer perceptron; Processing the infrared image to be detected based on a pre-constructed positioning branch to obtain a third infrared feature map, wherein the positioning branch is constructed based on a multi-core pooling module and a plurality of first convolution modules; The pre-built multi-branch guided fusion module is used to extract the respective depth features of the first infrared feature map, the second infrared feature map and the third infrared feature map, and each of the depth features is transformed based on the interactive guided module to obtain the small target detection result in the infrared image.
2. The infrared small target detection method based on multi-branch interactive guidance as claimed in claim 1 is characterized in that: The lossless encoder branch comprises: The multi-level feature fusion module has an input end communicatively connected to the output end of the first encoder.
3. The infrared small target detection method based on multi-branch interactive guidance as claimed in claim 2 is characterized in that: The multi-level feature fusion module includes: at least one cascaded residual module and a plurality of cascaded first upsampling modules; Among them, the input end of at least one cascaded residual module is communicatively connected to the output end of the first encoder, and the output end is communicatively connected to the input end of multiple cascaded first upsampling modules, and the output ends of the first encoder are all communicatively connected to the input ends of multiple first upsampling modules, and the number of the first upsampling modules is the same as the number of downsampling in the downsampling module.
4. The infrared small target detection method based on multi-branch interactive guidance as claimed in claim 1 is characterized in that: The downsampling module comprises: The feature pyramid module is used to extract the multi-scale feature map of the infrared image to be detected, and in the downsampling stage of the feature pyramid module, use Haar wavelet transform to process each of the multi-scale feature maps to obtain four sub-feature maps, and combine the four sub-feature maps to obtain each of the multi-scale feature maps with a channel number four times that of the original multi-scale feature map.
5. The infrared small target detection method based on multi-branch interactive guidance as claimed in claim 1, characterized in that: The method of processing the infrared image to be detected based on the pre-built positioning branch to obtain a third infrared feature map includes: Using a plurality of pooling modules with different kernel sizes to perform maximum pooling operations on the input infrared image, respectively, to obtain target detection images of different sizes; Using the difference between the infrared image and each of the target detection images, correspondingly obtaining each of the target detection images after background noise suppression; Aggregating the target detection images to obtain an aggregated target detection image; The first convolution modules connected by residual connections are used in sequence to extract positioning features from the aggregated target detection image to obtain a third infrared feature map.
6. The infrared small target detection method based on multi-branch interactive guidance as claimed in claim 3 is characterized in that: A fusion module is provided between the lossless encoder branch and the edge detection branch, two input ends of the fusion module are respectively connected to the first encoder and the second encoder, the first encoder outputs a second sub-feature map, and the pre-built edge detection branch processes the infrared image to obtain a second infrared feature map, including: Protecting edge details of the small target in the infrared image by using global average pooling and a multi-layer perceptron to obtain a first sub-feature map; Using the fusion module to perform feature fusion on the second sub-feature map and the first sub-feature map to obtain a third sub-feature map, and performing residual connection on the first sub-feature map and the third sub-feature map to obtain a fourth sub-feature map; The decoder is used to perform a decoding operation on the fourth sub-feature map to obtain a second infrared feature map.
7. The infrared small target detection method based on multi-branch interactive guidance as claimed in claim 6 is characterized in that: The decoder includes a plurality of second upsampling modules, a second convolution module and an activation module which are sequentially communicatively connected, and each of the upsampling modules is communicatively connected to the encoder, and the decoder is used to perform a convolution operation on the fourth sub-feature map to obtain a second infrared feature map, including: Using a plurality of the second up-sampling modules to perform up-sampling processing on the input fourth sub-feature map to obtain a fifth sub-feature map having the same size as the infrared image; Using the second convolution module to extract edge features from the fifth sub-feature graph to obtain a sixth sub-feature graph; The activation module is used to perform edge feature enhancement processing on the sixth sub-feature map to obtain a second infrared feature map.
8. The infrared small target detection method based on multi-branch interactive guidance as claimed in claim 6, characterized in that: The pre-built lossless encoder branch processes the infrared image to be detected to obtain a first infrared feature map, including: Using the lossless encoder to perform lossless downsampling processing on the infrared image to obtain a lossless coded image; Residual connecting the lossless coding image and the third sub-feature image to obtain a seventh sub-feature image; The plurality of first up-sampling modules are used to sequentially perform size restoration processing on the seventh sub-feature map to obtain an infrared feature map.
9. The infrared small target detection method based on multi-branch interactive guidance as claimed in claim 1, characterized in that: The multi-branch guided fusion module includes a convolution guided submodule and a Sobel guided submodule. The pre-built multi-branch guided fusion module is used to extract the depth features of the first infrared feature map, the second infrared feature map and the third infrared feature map, respectively, and each of the depth features is transformed based on the interactive guided module to obtain the small target detection result in the infrared image, including: The first infrared characteristic map, the second infrared characteristic map and the third infrared characteristic map are spliced to obtain a fourth infrared characteristic map having a depth feature; Processing the fourth infrared feature map using the convolution guidance submodule to obtain a first guidance feature map; Calculating a difference between the first guide characteristic map and the first infrared characteristic map to obtain a fifth infrared characteristic map; Processing the fourth infrared feature map based on the Sobel guidance submodule to obtain a second guidance feature map Calculating the second guide characteristic map and the fifth infrared characteristic map to obtain a sixth infrared characteristic map; The sixth feature map is processed by using a residual module to obtain the small target detection result in the infrared image.
10. An infrared small target detection device based on multi-branch interactive guidance, characterized in that: include: An acquisition module, used for acquiring an infrared image to be detected; A lossless detection module, used for processing the infrared image to be detected based on a pre-constructed lossless encoder branch to obtain a first infrared feature map, wherein the first encoder of the lossless encoder branch is constructed based on a downsampling module after Haar wavelet transform processing; An edge detection module, used for processing the infrared image to be detected based on a pre-built edge detection branch to obtain a second infrared feature map, wherein the edge detection branch comprises a second encoder and a decoder which are sequentially communicatively connected, and the encoder and the decoder are both constructed based on global average pooling and a multi-layer perceptron; A positioning module, used for processing the infrared image to be detected based on a pre-built positioning branch to obtain a third infrared feature map, wherein the positioning branch is built based on a multi-core pooling module and a plurality of first convolution modules; A fusion module is used to use a pre-built multi-branch guided fusion module to respectively extract the respective depth features of the first infrared feature map, the second infrared feature map and the third infrared feature map, and transform each of the depth features based on an interactive guided module to obtain a small target detection result in the infrared image.