Infrared target detection method and device

By introducing a multi-layer encoder-decoder structure and feature fusion module into the infrared object detection model, the problem of insufficient infrared object detection accuracy is solved, and efficient identification of infrared small objects is achieved.

CN120259681APending Publication Date: 2025-07-04HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510314438.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing infrared small object detection method based on deep learning has insufficient detection accuracy, especially due to the sparse characteristics of small objects and the separation of deep semantic information from shallow detail information.

Method used

The infrared object detection model is adopted, including a multi-layer encoder-decoder structure, and the feature maps output by the encoder of adjacent layers and the decoder are characterized by the feature fusion module, and the output fusion module is weighted and fused to each decoder's feature map to enhance the detection performance of small objects.

Benefits of technology

The detection accuracy and performance of small infrared targets are improved, and the accuracy of target recognition is enhanced by retaining shallow detail information and deep semantic information in infrared images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259681A_ABST
    Figure CN120259681A_ABST
Patent Text Reader

Abstract

The invention relates to an infrared target detection method and device, and belongs to the technical field of infrared detection.The method comprises the steps that a first encoder conducts feature extraction on a first feature map of an infrared image to obtain a second feature map; the second encoder performs feature extraction on the second feature map to obtain a third feature map; the first decoder decodes the third feature map to obtain a fourth feature map; a first feature fusion module performs feature fusion on the second feature map, the third feature map and the fourth feature map to obtain a first fusion feature map; the second decoder decodes the first fusion feature map to obtain a fifth feature map; and the output fusion module performs weighted fusion on the feature maps output by the decoders, and determines an infrared target detection result according to the feature maps obtained by weighted fusion. According to the method, the three feature maps are fused, detail features can be reserved, and the target recognition accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of infrared detection technology, and in particular, to an infrared target detection method and device. Background Art

[0002] Infrared small target detection technology is widely used in scenarios such as infrared search, tracking, and guidance. Among them, the detection method based on deep learning has great advantages in detection speed and accuracy. However, when detecting infrared small targets, due to the very small existence area and scarce features of small targets, in deep learning methods, as the network deepens, the detailed features of the detection target are severely lost, and the deep semantic information and shallow detailed information are split, resulting in difficulty in improving the detection accuracy of small targets. Summary of the Invention

[0003] In view of this, it is necessary to provide an infrared target detection method and device to solve the problem of low detection accuracy of the existing infrared small target detection method based on deep learning.

[0004] To solve the above problems, in a first aspect, the present invention provides an infrared target detection method, including: Inputting an infrared image into a trained infrared target detection model; wherein, the infrared target detection model includes a first encoder, a second encoder, a first decoder, a second decoder, a first feature fusion module, and an output fusion module; Extracting features of a first feature map of the infrared image through the first encoder to obtain a second feature map; Extracting features of the second feature map through the second encoder to obtain a third feature map; Decoding the third feature map through the first decoder to obtain a fourth feature map; Performing feature fusion on the second feature map, the third feature map, and the fourth feature map through the first feature fusion module to obtain a first fusion feature map; Decoding the first fusion feature map through the second decoder to obtain a fifth feature map; Performing weighted fusion on the feature maps output by each decoder in the infrared target detection model through the output fusion module, and determining an infrared target detection result according to the feature map obtained by weighted fusion.

[0005] Optionally, the infrared target detection model further includes a third encoder, a third decoder, and a second feature fusion module; the decoding the third feature map through the first decoder to obtain a fourth feature map includes: Extracting features of the third feature map through the third encoder to obtain a sixth feature map; The sixth feature map is decoded by the third decoder to obtain a seventh feature map; The third feature map, the sixth feature map, and the seventh feature map are subjected to feature fusion by the second feature fusion module to obtain a second fused feature map; The second fused feature map is decoded by the first decoder to obtain a fourth feature map.

[0006] Optionally, the infrared target detection model further includes a fourth encoder and a fourth decoder; the step of performing weighted fusion on the feature maps output by each decoder in the infrared target detection model through the output fusion module and determining the infrared target detection result based on the feature map obtained by the weighted fusion includes: The infrared image is subjected to feature extraction by the fourth encoder to obtain the first feature map; The third fused feature map is decoded by the fourth decoder to obtain an eighth feature map; wherein, the third fused feature map is obtained by performing feature fusion on the first feature map and the fifth feature map; The fourth feature map, the fifth feature map, the seventh feature map, and the eighth feature map are subjected to weighted fusion by the output fusion module, and the infrared target detection result is determined based on the feature map obtained by the weighted fusion.

[0007] Optionally, the step of obtaining the first fused feature map by performing feature fusion on the second feature map, the third feature map, and the fourth feature map through the first feature fusion module includes: After the second feature map and the fourth feature map are respectively subjected to three-dimensional coordinate attention enhancement by the first feature fusion module, they are subjected to feature fusion with the third feature map to obtain a fused feature map.

[0008] Optionally, the step of performing three-dimensional coordinate attention enhancement on the second feature map by the first feature fusion module includes: The second feature map is subjected to max pooling processing and average pooling processing on the channels by the first feature fusion module to obtain first global key data and first basic data; the first weight and the second weight are determined according to the first global key data and the first basic data; the second feature map is subjected to two-dimensional coordinate attention enhancement on the height and width to obtain a ninth feature map; based on the first weight and the second weight, the ninth feature map and the second feature map are subjected to feature fusion to obtain a second feature map with completed three-dimensional coordinate attention enhancement.

[0009] Optionally, the step of determining the first weight and the second weight according to the first global key data and the first basic data includes: The first weight and the second weight are obtained through the following formula: ,

[0010] wherein, and are the first weight and the second weight respectively, and are the first global key data and the first basic data respectively; Feature fusion is performed on the ninth feature map and the second feature map based on the first weight and the second weight to obtain a second feature map with enhanced three-dimensional coordinate attention, including: The second feature map with enhanced three-dimensional coordinate attention is obtained through the following formula:

[0011] wherein, represents the second feature map with enhanced three-dimensional coordinate attention, and represent the ninth feature map and the second feature map respectively.

[0012] Optionally, the three-dimensional coordinate attention enhancement of the fourth feature map by the first feature fusion module includes: Performing max pooling processing and average pooling processing on the fourth feature map in the channel by the first feature fusion module to obtain second global key data and second basic data; determining a third weight and a fourth weight according to the second global key data and the second basic data; performing two-dimensional coordinate attention enhancement on the fourth feature map in height and width to obtain a tenth feature map; performing feature fusion on the fourth feature map and the tenth feature map based on the third weight and the fourth weight to obtain a fourth feature map with enhanced three-dimensional coordinate attention.

[0013] Optionally, the feature fusion with the third feature map to obtain the first fusion feature map includes: Performing three-dimensional coordinate adjustment on the third feature map by the first feature fusion module to make the three-dimensional coordinates of the third feature map consistent with the three-dimensional coordinates of the second feature map and the fourth feature map; performing feature fusion on the third feature map with completed three-dimensional coordinate adjustment, the second feature map with enhanced three-dimensional coordinate attention, and the fourth feature map with enhanced three-dimensional coordinate attention to obtain the first fusion feature map.

[0014] Optionally, the weights of the fourth feature map, the fifth feature map, the seventh feature map, and the eighth feature map are determined according to the size information of the fourth feature map, the fifth feature map, the seventh feature map, and the eighth feature map.

[0015] In a second aspect, the present invention further provides an infrared target detection device, comprising: An image input functional component for inputting an infrared image into a trained infrared target detection model; wherein, the infrared target detection model includes a first encoder, a second encoder, a first decoder, a second decoder, a first feature fusion module, and an output fusion module; A first encoding functional component for extracting features from a first feature map of the infrared image through the first encoder to obtain a second feature map; A second encoding functional component for extracting features from the second feature map through the second encoder to obtain a third feature map; A first decoding functional component for decoding the third feature map through the first decoder to obtain a fourth feature map; A feature fusion functional component for performing feature fusion on the second feature map, the third feature map, and the fourth feature map through the first feature fusion module to obtain a first fusion feature map; A second decoding functional component for decoding the first fusion feature map through the second decoder to obtain a fifth feature map; An output fusion functional component for performing weighted fusion on the feature maps output by each decoder in the infrared target detection model through the output fusion module, and determining an infrared target detection result based on the feature map obtained by the weighted fusion.

[0016] The beneficial effects of the present invention are: The present invention inputs an infrared image into a trained infrared target detection model. The infrared target detection model includes a first encoder, a second encoder, a first decoder, a second decoder, a first feature fusion module, and an output fusion module. The first encoder extracts features from the first feature map of the infrared image to obtain a second feature map. The second encoder extracts features from the second feature map to obtain a third feature map. The first decoder decodes the third feature map to obtain a fourth feature map. The first feature fusion module performs feature fusion on the second feature map, the third feature map, and the fourth feature map to obtain a first fusion feature map. By performing feature fusion on the feature maps output by adjacent layer encoders and the first decoder through the first feature fusion module, the present invention can retain the shallow detail information and deep semantic information in the infrared image in the first fusion feature, so that the recognition accuracy of infrared small targets can be improved based on the first fusion feature. In addition, the present invention performs weighted fusion on the feature maps output by each decoder in the infrared target detection model through the output fusion module, and determines the infrared target detection result according to the feature map obtained by weighted fusion, which can further improve the detection performance of the model and make the detection of infrared small targets more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic flowchart of an embodiment of the infrared target detection method provided by the present invention; Figure 2 is an architecture diagram of an infrared target detection model provided by the present invention; Figure 3 is a schematic working flowchart of an output fusion module provided by the present invention; Figure 4 is a schematic diagram of an encoder structure provided by the present invention; Figure 5 is a schematic working diagram of a decoder provided by the present invention; Figure 6 is a schematic working flowchart of a feature fusion module provided by the present invention; Figure 7 is a schematic structural diagram of an embodiment of the infrared target detection device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0019] In the description of the embodiments of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more. In the embodiments of the present invention, the "first", "second", etc. involved are used to distinguish similar objects, rather than to describe a specific order or sequence, nor to indicate or imply their relative importance or implicitly specify the quantity of the indicated technical features. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and do not limit the number of objects. For example, the first object can be one or multiple.

[0020] Reference herein to "embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0021] Refer to Figure 1 , which shows a schematic flowchart of an embodiment of the infrared target detection method provided by the present invention. The method includes: S101, input the infrared image into the trained infrared target detection model; wherein, the infrared target detection model includes a first encoder, a second encoder, a first decoder, a second decoder, a first feature fusion module, and an output fusion module; S102, extract features from the first feature map of the infrared image through the first encoder to obtain a second feature map.

[0022] S103, extract features from the second feature map through the second encoder to obtain a third feature map.

[0023] S104, decode the third feature map through the first decoder to obtain a fourth feature map.

[0024] S105, perform feature fusion on the second feature map, the third feature map, and the fourth feature map through the first feature fusion module to obtain a first fusion feature map.

[0025] S106, decode the first fusion feature map through the second decoder to obtain a fifth feature map.

[0026] S107, perform weighted fusion on the feature maps output by each decoder in the infrared target detection model through the output fusion module, and determine the infrared target detection result according to the feature map obtained by the weighted fusion.

[0027] The infrared target detection model can be a model for detecting small targets in infrared images. A small target can refer to a target whose pixel area or pixel diameter or pixel size in the image is less than a preset threshold.

[0028] The infrared target detection model can include a first encoder, a second encoder, a first decoder, a second decoder, a first feature fusion module, and an output fusion module. The first encoder, the second encoder, and the first decoder are connected in sequence. The first encoder is used to extract features from the first feature map of the infrared image to obtain a second feature map. It should be noted that there can be at least one other encoder before the first encoder. The first encoder can be used to extract features from the first feature map obtained by other encoders extracting features from the infrared image first, and then obtain the second feature map. The second encoder is used to extract features from the second feature map to obtain a third feature map. The first decoder is used to decode the third feature map to obtain a fourth feature map.

[0029] The first feature fusion module can be connected to the first encoder, the second encoder, the first decoder, and the second decoder respectively. The first feature fusion module is used to perform feature fusion on the second feature map output by the first encoder, the third feature map output by the second encoder, and the fourth feature map output by the first decoder to obtain a first fusion feature map, and input the first fusion feature map into the second decoder.

[0030] The second decoder can be connected to the first feature fusion module and is used to decode the first fusion feature map to obtain a fifth feature map.

[0031] It should be noted that the first encoder, the second encoder, the first decoder, the second decoder, and the first feature fusion module can be regarded as a unit, and one or more such units can be included in the infrared target detection model.

[0032] The output fusion module can be connected to each decoder in the infrared target detection model respectively, and is used to perform weighted fusion on the feature maps output by each decoder, and determine the infrared target detection result according to the feature map obtained by the weighted fusion. The infrared target detection model can contain multiple above-mentioned units, that is, it contains multiple first decoders and second decoders, and in addition, it can also contain other decoders except the first decoder and the second decoder. The output fusion module is used to perform weighted fusion on the feature maps output by each decoder in the infrared target detection model, and then determine the infrared target detection result according to the feature map obtained by the weighted fusion.

[0033] The feature map obtained by the weighted fusion can be a pixel-probability map. When the probability value corresponding to a pixel is greater than 0.5, it can be determined that the pixel belongs to the infrared target.

[0034] In the present invention, the first feature fusion module performs feature fusion on the feature maps output by adjacent-layer encoders and the first decoder to obtain a first fusion feature map, which can retain the shallow detail information and deep semantic information in the infrared image. Furthermore, the fifth feature map obtained by decoding the first fusion feature map through the second decoder will also retain the shallow detail information and deep semantic information. Finally, only weighted fusion is performed on the feature maps output by each encoder in the infrared target detection model, and the infrared target detection result is determined based on the weighted fusion feature map, which can further improve the detection performance and make the detection of infrared small targets more accurate.

[0035] In one embodiment, the infrared target detection model further includes a third encoder, a third decoder, and a second feature fusion module; decoding the third feature map through the first decoder to obtain a fourth feature map, including: extracting features from the third feature map through the third encoder to obtain a sixth feature map; decoding the sixth feature map through the third decoder to obtain a seventh feature map; performing feature fusion on the third feature map, the sixth feature map, and the seventh feature map through the second feature fusion module to obtain a second fusion feature map; and decoding the second fusion feature map through the first decoder to obtain a fourth feature map.

[0036] In the previous embodiment, the first encoder, the second encoder, the first feature fusion module, the first decoder, and the second decoder constitute a unit. In this embodiment, the second encoder, the third encoder, the second feature fusion module, the third decoder, and the first decoder also constitute a unit. In this embodiment, a method of introducing multiple such units into the infrared target detection model is shown. Referring to this embodiment, a feature fusion module can also be inserted between subsequent encoders and decoders according to requirements to introduce this unit, and the introduction method will not be elaborated here.

[0037] Refer to Figure 2 , which shows an architecture diagram of an infrared target detection model provided by the present invention. The infrared target detection model is a model established based on the U-Net architecture. The infrared target detection model includes a multi-layer encoder-decoder structure, with skip connections between the encoder and the decoder. Multiple encoders are used to perform feature extraction on the input infrared image step by step, and multiple decoders are used to perform size recovery on the extracted features step by step. To balance the detection performance, resource consumption, and detection frame rate, this embodiment finally selects a four-layer encoder-decoder structure. And a feature fusion module (TDCAFF) is added to the encoder-decoder structures of the second layer (the first encoder - the second decoder) and the third layer (the Two encoder - the first decoder).

[0038] In one embodiment, the infrared target detection model further includes a fourth encoder, a fourth decoder, and an output fusion module; S107 may include: extracting features from the infrared image through the fourth encoder to obtain a first feature map; decoding the third fusion feature map through the fourth decoder to obtain an eighth feature map; wherein, the third fusion feature map is obtained by performing feature fusion on the first feature map and the fifth feature map; performing weighted fusion on the fourth feature map, the fifth feature map, the seventh feature map, and the eighth feature map through the output fusion module, and determining the infrared target detection result according to the feature map obtained by the weighted fusion.

[0039] In this embodiment, the fourth encoder is connected to the fourth decoder. Since the fourth encoder is located in the shallow network structure and the detailed information is significant, there is no need to insert a feature fusion module between the fourth encoder and the fourth decoder. Refer to Figure 2 , in the first layer of the encoding and decoding structure (the fourth encoder - the fourth decoder), the detailed information of the shallow network is significant, and the fourth encoder and the fourth decoder generally adopt a residual connection. In the second and third layers of the encoding and decoding structure, a feature fusion module is inserted. Since the feature fusion module needs to use the encoder and decoder of the subsequent layer for feature feedback, the feature fusion module cannot be inserted into the last layer of the encoding and decoding structure.

[0040] In this embodiment, the final output result of the infrared target recognition model does not completely depend on the output of the last decoder (the fourth decoder), but fully utilizes the feature maps output by each decoder, performs weighted fusion on the feature maps output by each decoder, and obtains the infrared target detection result. Among them, the weights of the feature maps output by each decoder can be determined according to the size information of the feature maps output by each decoder, and the size information may include height and width.

[0041] Refer to Figure 3 , which shows a schematic diagram of the working process of an output fusion module provided by the present invention. The spatial sizes of the feature maps output by each decoder are different. Therefore, it is first necessary to directly upsample the feature maps output by each decoder to the final output target size . The feature map output by the fourth decoder with the largest size contains more detailed information. The normalized probability map of the upsampled feature map F4 of it is used to guide the fusion of other decoded feature maps. The probability map of F4 is obtained through the sigmoid function, and then it is convolved point by point with the feature maps of other layers.

[0042] Refer to Figure 4, which shows a schematic structural diagram of an encoder provided by the present invention. The encoder uses VisionTransformer as the basic structure. Vision Transformer has demonstrated excellent performance in the field of image processing, but it has always been limited by its quadratic complexity characteristics. As the processing sequence increases, the computational consumption increases sharply. To balance its processing performance and resource consumption, the LSR module in PVT v2 is used in this embodiment to reduce spatial attention and thus reduce computational resource consumption.

[0043] Referring to Figure 5 , which shows a schematic diagram of the operation of a decoder provided by the present invention. The decoder first performs 3x3 convolution to enable the decoder to capture rich local features with a relatively small amount of computation. This step effectively extracts the spatial details of the input feature map, laying the foundation for subsequent operations. Then, bicubic interpolation is used to restore the resolution of the feature map, gradually restoring the details of the image. Bicubic interpolation can maintain the smooth continuity of features, avoid generating jagged edges, and ensure the improvement of image quality. Finally, through 1x1 convolution, the number of channels of the feature map is adjusted to a preset value. This step adjusts the feature expression ability of the model at different stages while maintaining the spatial resolution. After each convolution, BN batch normalization and ReLU activation function are used to further enhance the non-linear expression ability of the features.

[0044] In one embodiment, S105 may include: respectively performing three-dimensional coordinate attention enhancement on the second feature map and the fourth feature map through the first feature fusion module, and then performing feature fusion with the third feature map to obtain a first fusion feature map.

[0045] Since the small target features are scarce and the edges are blurred, directly fusing the feature models between adjacent encoders results in unsatisfactory recognition effects. Moreover, the feature similarity between adjacent encoders is high, the feature redundancy is large, and the fusion efficiency is low. Therefore, in this embodiment, the first feature fusion module is used to perform three-dimensional coordinate attention enhancement on the second feature map and the fourth feature map respectively, enhancing the detailed features of small targets in the second feature map and the fourth feature map. Then, after fusing with the third feature map, the features of small targets in the shallow-layer detailed information can be highlighted, making the recognition accuracy higher when performing small target recognition based on the first fusion feature.

[0046] The three-dimensional coordinate attention enhancement can be based on the two-dimensional coordinate (width W and height H) attention and use the information between channels to enhance the image features.

[0047] In one embodiment, the steps of performing three-dimensional coordinate attention enhancement on the second feature map may include: performing max pooling and average pooling on the second feature map in the channel dimension through a first feature fusion module to obtain first global key data and first basic data; determining a first weight and a second weight based on the first global key data and the first basic data; performing two-dimensional coordinate attention enhancement on the second feature map in the height and width dimensions to obtain a ninth feature map; and performing feature fusion on the ninth feature map and the second feature map based on the first weight and the second weight to obtain the second feature map with completed three-dimensional coordinate attention enhancement.

[0048] Specifically, the pooling process can be represented by the following formula:

[0049]

[0050] In the formula, and are the first global key data and the first basic data of the second feature map respectively, represents performing max pooling on the second feature map , represents performing average pooling on the second feature map .

[0051] The first weight and the second weight are obtained through the following formula: ,

[0052] In the formula, and are the first weight and the second weight respectively, and are the first global key data and the first basic data respectively.

[0053] The second feature map with completed three-dimensional coordinate attention enhancement is obtained through the following formula:

[0054] In the formula, represents the second feature map with completed three-dimensional coordinate attention enhancement, and represent the ninth feature map and the second feature map respectively.

[0055] In this embodiment, based on the first weight and the second weight, the before two-dimensional coordinate attention enhancement and the after two-dimensional coordinate attention enhancement are fused, which can enhance the key information representing the image detail contours and strong edges in the second feature map and retain the average information in the second feature map.

[0056] In one embodiment, the steps of performing three-dimensional coordinate attention enhancement on the fourth feature map may include: performing max pooling and average pooling on the fourth feature map in the channel dimension through a feature fusion module to obtain second global key data and second basic data; determining a third weight and a fourth weight based on the second global key data and the second basic data; performing two-dimensional coordinate attention enhancement on the fourth feature map in the height and width dimensions to obtain a tenth feature map; and performing feature fusion on the tenth feature map and the fourth feature map based on the third weight and the fourth weight to obtain the fourth feature map with completed three-dimensional coordinate attention enhancement.

[0057] The steps of performing three-dimensional coordinate attention enhancement on the fourth feature map are similar to those of performing three-dimensional coordinate attention enhancement on the second feature map, and the specific implementation process of this embodiment can be referred to the specific implementation process of the previous embodiment.

[0058] In one embodiment, the steps of performing feature fusion with the third feature map to obtain the first fusion feature map may include: performing three-dimensional coordinate adjustment on the third feature map through a first feature fusion module to make the three-dimensional coordinates of the third feature map consistent with the three-dimensional coordinates of the second feature map and the fourth feature map; and performing feature fusion on the third feature map with completed three-dimensional coordinate adjustment, the second feature map with completed three-dimensional coordinate attention enhancement, and the fourth feature map with completed three-dimensional coordinate attention enhancement to obtain the first fusion feature map.

[0059] The second feature map and the fourth feature Figure 3 have the same three-dimensional size, while the three-dimensional size of the third feature map is inconsistent with these two feature maps. Therefore, in this embodiment, it is necessary to perform three-dimensional size adjustment on the third feature map and then perform feature fusion with the second feature map and the fourth feature map with completed three-dimensional coordinate attention enhancement to obtain the first fusion feature map.

[0060] Refer to Figure 6 , which shows a schematic diagram of the working process of a feature fusion module provided by the present invention. , and are respectively the second feature map output by the first encoder, the third feature map output by the second encoder, and the fourth feature map output by the first decoder. The second feature map and the fourth feature Figure 3 have the same three-dimensional size, both being , the number of channels of the third feature map is and twice that of and half of the two-dimensional size of and will undergo three-dimensional coordinate attention enhancement through TDCA, and for The channels will be adjusted through depthwise separable convolution (DSConv) and upsampled to a three-dimensional size of . To reduce resource consumption, depthwise separable convolution is used instead of ordinary convolution. Here, the DSConv is a complete convolution block that includes a normalization BN layer and an activation ReLU function.

[0061] In one embodiment, the step of obtaining the second fusion feature map by performing feature fusion on the third feature map, the sixth feature map, and the seventh feature map through the second feature fusion module may include: after respectively performing three-dimensional coordinate attention enhancement on the third feature map and the seventh feature map through the second feature fusion module, performing feature fusion with the sixth feature map to obtain the second fusion feature map. Among them, the step of performing three-dimensional coordinate attention enhancement on the third feature map and the seventh feature map, and the step of performing feature fusion with the sixth feature map to obtain the second fusion feature map may refer to the above embodiment and will not be elaborated here.

[0062] In one embodiment, in addition to inserting the first feature fusion module between the first encoder and the second decoder and inserting the second feature fusion module between the second encoder and the first decoder, feature fusion modules can also be inserted between subsequent encoders and decoders according to requirements. The insertion method and the working principle of the inserted feature fusion module can both refer to the above embodiment and will not be elaborated here.

[0063] The present invention proposes a feature fusion module, aiming to retain the low-level feature information of small targets through information fusion and target enhancement. In addition, a weighted output fusion strategy is used to fuse the outputs of different decoding stages, further improving the detection performance. Comparative experiments prove that the overall effect of the algorithm proposed by the present invention is better than that of existing algorithms, and ablation experiments prove that the feature fusion module of the present invention has better effects than existing feature fusion modules, especially in pixel-level detection.

[0064] In one embodiment, the infrared target recognition model can be trained based on sample infrared images. InfiRay flip ph35+, InfiRay AFFO AH25 long-wave infrared cameras, and a self-made LEO cooled mid-wave infrared camera can be used to collect sample infrared images containing different types of small targets in a variety of different scenarios. The scenarios include the edges of seas, rivers, and lakes, forests, urban buildings, the sky, etc. The small targets include airplanes, drones, ships, cars, artificial heat sources, flying birds, and other small targets. Then, the sample infrared image set is finely labeled, and appropriate sample infrared images are selected to form a new sample infrared image set. The resolution of the new sample infrared image set is adjusted to 512×512. For the part with insufficient resolution, the border is directly extended, and the pixel values at the extended positions are the mean of the original pixel values. For the part with excessive resolution, the redundant background is cropped. The image set is uniformly named FD_0001~FD1000, and the image set is divided into three groups by a random allocation program, with a ratio of 5:3:2, namely the training set, the validation set, and the test set. The training set is used to train the infrared target detection model, the test set is used to test the training effect of the model in real time during the training process, and the validation set is used to verify the performance of the model after training is completed.

[0065] The configuration used for training the infrared target recognition model can be a server with 2 Intel(R) Xeon 6326 16-core @2.9GHz CPUs, 4 NVIDIA GeForce RTX 3090 GPUs, and 4 32G 3200MHz DDR4 memory modules. The system is Ubuntu 22.04.4, and Python 3.9.19 and Pytorch 1.10.0 are used.

[0066] Referring to Figure 7 , a schematic structural diagram of an embodiment of the infrared target detection device provided by the present invention is shown. The device 70 includes: An image input functional component 701 for inputting an infrared image into the trained infrared target detection model. Among them, the infrared target detection model includes a first encoder, a second encoder, a first decoder, a second decoder, a first feature fusion module, and an output fusion module; A first encoding functional component 702 for extracting features from the first feature map of the infrared image through the first encoder to obtain a second feature map; A second encoding functional component 703 for extracting features from the second feature map through the second encoder to obtain a third feature map; A first decoding functional component 704 for decoding the third feature map through the first decoder to obtain a fourth feature map; The feature fusion functional component 705 is configured to perform feature fusion on the second feature map, the third feature map, and the fourth feature map through the first feature fusion module to obtain a first fused feature map; The second decoding functional component 706 is configured to decode the fused feature map through the second decoder to obtain a fifth feature map; The output fusion functional component 707 is configured to perform weighted fusion on the feature maps output by each decoder in the infrared target detection model through the output fusion module, and determine the infrared target detection result according to the feature map obtained by the weighted fusion.

[0067] It should be noted that: the implementation principle or implementation process of the above modules can refer to the embodiments of the foregoing infrared target detection method, which will not be elaborated here one by one.

[0068] Those skilled in the art can understand that all or part of the processes for implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.

[0069] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. An infrared target detection method, characterized in that, Including: Inputting the infrared image into a trained infrared target detection model; wherein, the infrared target detection model includes a first encoder, a second encoder, a first decoder, a second decoder, a first feature fusion module, and an output fusion module; Extracting features from the first feature map of the infrared image through the first encoder to obtain a second feature map; Extracting features from the second feature map through the second encoder to obtain a third feature map; Decoding the third feature map through the first decoder to obtain a fourth feature map; Performing feature fusion on the second feature map, the third feature map, and the fourth feature map through the first feature fusion module to obtain a first fusion feature map; Decoding the first fusion feature map through the second decoder to obtain a fifth feature map; Performing weighted fusion on the feature maps output by each decoder in the infrared target detection model through the output fusion module, and determining the infrared target detection result according to the feature map obtained by the weighted fusion.

2. The infrared target detection method according to claim 1, wherein The infrared target detection model further includes a third encoder, a third decoder, and a second feature fusion module; the step of decoding the third feature map through the first decoder to obtain a fourth feature map includes: Extracting features from the third feature map through the third encoder to obtain a sixth feature map; Decoding the sixth feature map through the third decoder to obtain a seventh feature map; Performing feature fusion on the third feature map, the sixth feature map, and the seventh feature map through the second feature fusion module to obtain a second fusion feature map; Decoding the second fusion feature map through the first decoder to obtain a fourth feature map.

3. The infrared target detection method according to claim 2, wherein, The infrared target detection model further includes a fourth encoder and a fourth decoder; the step of performing weighted fusion on the feature maps output by each decoder in the infrared target detection model through the output fusion module, and determining the infrared target detection result according to the feature map obtained by the weighted fusion includes: Extracting features from the infrared image through the fourth encoder to obtain a first feature map; Decoding the third fusion feature map through the fourth decoder to obtain an eighth feature map; wherein, the third fusion feature map is obtained by performing feature fusion on the first feature map and the fifth feature map; Performing weighted fusion on the fourth feature map, the fifth feature map, the seventh feature map, and the eighth feature map through the output fusion module, and determining the infrared target detection result according to the feature map obtained by the weighted fusion.

4. The infrared target detection method according to claim 1, wherein The step of performing feature fusion on the second feature map, the third feature map, and the fourth feature map through the first feature fusion module to obtain a first fusion feature map includes: Performing three-dimensional coordinate attention enhancement on the second feature map and the fourth feature map respectively through the first feature fusion module, and then performing feature fusion with the third feature map to obtain a first fusion feature map.

5. The infrared target detection method according to claim 4, wherein The step of performing three-dimensional coordinate attention enhancement on the second feature map through the first feature fusion module includes: The first feature fusion module performs max pooling processing and average pooling processing on the second feature map in the channel dimension to obtain first global key data and first basic data; determines a first weight and a second weight according to the first global key data and the first basic data; performs two-dimensional coordinate attention enhancement on the second feature map in the height and width dimensions to obtain a ninth feature map; and performs feature fusion on the ninth feature map and the second feature map based on the first weight and the second weight to obtain a second feature map with three-dimensional coordinate attention enhancement completed.

6. The infrared target detection method according to claim 5, characterized in that, Determining the first weight and the second weight according to the first global key data and the first basic data includes: Obtaining the first weight and the second weight through the following formula: , Among them, and are the first weight and the second weight respectively, and are the first global key data and the first basic data respectively; Performing feature fusion on the ninth feature map and the second feature map based on the first weight and the second weight to obtain a second feature map with three-dimensional coordinate attention enhancement completed includes: Obtaining the second feature map with three-dimensional coordinate attention enhancement completed through the following formula: Among them, represents the second feature map that completes three-dimensional coordinate attention enhancement, , respectively represent the ninth feature map and the second feature map.

7. The infrared target detection method according to claim 4, wherein, Performing three-dimensional coordinate attention enhancement on the fourth feature map through the first feature fusion module includes: The first feature fusion module performs max pooling processing and average pooling processing on the fourth feature map in the channel dimension to obtain second global key data and second basic data; determines a third weight and a fourth weight according to the second global key data and the second basic data; performs two-dimensional coordinate attention enhancement on the fourth feature map in the height and width dimensions to obtain a tenth feature map; and performs feature fusion on the fourth feature map and the tenth feature map based on the third weight and the fourth weight to obtain a fourth feature map with three-dimensional coordinate attention enhancement completed.

8. The infrared target detection method according to claim 4, characterized in that Performing feature fusion with the third feature map to obtain a first fused feature map includes: The first feature fusion module performs three-dimensional coordinate adjustment on the third feature map to make the three-dimensional coordinates of the third feature map consistent with the three-dimensional coordinates of the second feature map and the fourth feature map; and performs feature fusion on the third feature map with three-dimensional coordinate adjustment completed, the second feature map with three-dimensional coordinate attention enhancement completed, and the fourth feature map with three-dimensional coordinate attention enhancement completed to obtain a first fused feature map.

9. The infrared target detection method according to claim 3, wherein The weights of the fourth feature map, the fifth feature map, the seventh feature map, and the eighth feature map are determined according to the size information of the fourth feature map, the fifth feature map, the seventh feature map, and the eighth feature map.

10. An infrared target detection device, characterized in that, Including: An image input functional component for inputting an infrared image into a trained infrared target detection model; wherein, the infrared target detection model includes a first encoder, a second encoder, a first decoder, a second decoder, a first feature fusion module, and an output fusion module; A first encoding functional component for extracting features of a first feature map of the infrared image through the first encoder to obtain a second feature map; A second encoding functional component for extracting features of the second feature map through the second encoder to obtain a third feature map; The first decoding functional component is used to decode the third feature map through the first decoder to obtain a fourth feature map; The feature fusion functional component is used to perform feature fusion on the second feature map, the third feature map, and the fourth feature map through the first feature fusion module to obtain a first fused feature map; The second decoding functional component is used to decode the first fused feature map through the second decoder to obtain a fifth feature map; The output fusion functional component is used to perform weighted fusion on the feature maps output by each decoder in the infrared target detection model through the output fusion module, and determine the infrared target detection result according to the feature map obtained by the weighted fusion.