Multi-frame Infrared Small Target Detection Method, Device, Computer Equipment and Storage Medium

By using the U-shaped network and the timing Transformer module in multi-frame infrared small object detection, the problem of insufficient detection accuracy and positioning accuracy in the existing technology in complex backgrounds is solved, and more efficient small object detection and positioning is achieved.

CN120070871BActive Publication Date: 2025-06-27NAT UNIV OF DEFENSE TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510524425.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-06-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing multi-frame infrared small object detection method is difficult to effectively eliminate background interference under complex backgrounds and extract accurate target features, resulting in insufficient detection accuracy and positioning accuracy of small object.

Method used

The encoder and decoder of the U-shaped network are adopted, and the timing channel cross-transformer submodule and the timing spatial cross-transformer submodule in the timing Transformer module are combined to cross-fusion of spatial features and timing features in memory. The dual-output ConvLSTM module is used to process timing features and update memory, breaking through the limitations of the fixed time window and extracting richer space-time features.

Benefits of technology

The performance of multi-frame infrared small object detection is significantly improved, the accuracy of small object detection and the accuracy of small object positioning is improved, and the target positioning can be positioned and the target shape can be estimated more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070871B_ABST
    Figure CN120070871B_ABST
Patent Text Reader

Abstract

The present application relates to a multi-frame infrared small target detection method, device, computer device, and storage medium. The method includes: obtaining a video frame sequence sample; constructing a multi-frame infrared small target detection model, which includes an encoder, a temporal Transformer module, a decoder, and an output module, and the encoder and the decoder are connected through convolutional layers to form a U-shaped network; the temporal Transformer module includes a temporal channel cross Transformer sub-module and a temporal spatial cross Transformer sub-module, and the output module processes the features output by the decoder to obtain a target segmentation map, training the multi-frame infrared small target detection model according to the video frame sequence sample to obtain a trained multi-frame infrared small target detection model, and using the trained multi-frame infrared small target detection model for target detection. Using this method can significantly improve the performance of multi-frame infrared small target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of target detection, and particularly to a multi-frame infrared small target detection method, device, computer device, and storage medium. Background Art

[0002] Infrared small target detection aims to accurately locate small-sized targets in infrared images and videos. Benefiting from the excellent performance of infrared imaging under all-weather and low-light conditions, this technology is widely used in tasks such as intrusion warning, target guidance, and maritime rescue. According to the number of infrared image frames used, it can be classified into single-frame infrared small target detection and multi-frame infrared small target detection. Theoretically, compared with single-frame infrared small target detection that can only utilize the spatial information in one image, multi-frame infrared small target detection can utilize the spatio-temporal information between frames to better detect weak and small targets in complex backgrounds.

[0003] Existing multi-frame infrared small target detection methods usually have two characteristics: one is to adopt instance-based detection frameworks such as YOLO and Faster-RCNN, and the other is to extract spatio-temporal features in a time window of a fixed size. The former has a large development gap compared with the segmentation-based detection frameworks that currently perform excellently in single-frame infrared small target detection, and the latter limits the algorithm to extract more spatio-temporal features from long sequences to achieve better performance. Due to the above limitations, in complex scenes such as the sky, vegetation, and buildings, when the detection algorithm faces small-sized and weak-feature targets, it is difficult to effectively exclude background interference and extract accurate target features, which not only makes accurate detection and shape estimation more complex, but also greatly increases the difficulty of accurately locating the position of small targets, and it is difficult to meet the requirements of high-precision detection of small targets in practical applications. Summary of the Invention

[0004] Based on this, it is necessary to provide a multi-frame infrared small target detection method, device, computer device, and storage medium for the above technical problems.

[0005] A multi-frame infrared small target detection method, the method comprising:

[0006] Obtaining a video frame sequence sample; the video frame sequence sample is obtained by preprocessing an infrared video;

[0007] Constructing a multi-frame infrared small target detection model; the multi-frame infrared small target detection model includes an encoder, a temporal Transformer module, a decoder, and an output module; the encoder and the decoder are connected by a convolutional layer to form a U-shaped network; the temporal Transformer module includes a temporal channel cross Transformer sub-module and a temporal spatial cross Transformer sub-module;

[0008] The current video frame is downsampled step by step through the encoder to obtain spatial features at multiple different scales, and the spatial features output by the last-stage downsampling are convolved and then output to the decoder;

[0009] Each spatial feature is mapped to the same size through the temporal-channel cross Transformer sub-module and then concatenated along the channel dimension to obtain a first concatenated feature. The first concatenated feature is processed by a dual-output ConvLSTM to obtain an updated channel-level memory. Cross-attention processing is performed on the temporal features included in the updated channel-level memory and the convolved first concatenated feature to obtain a first fusion feature. The first fusion feature is split along the channel dimension to obtain multiple channel-level fusion features;

[0010] Each channel-level fusion feature is concatenated along the batch dimension through the temporal-spatial cross Transformer sub-module to obtain a second concatenated feature. The second concatenated feature is processed by a dual-output ConvLSTM to obtain an updated spatial-level memory. Cross-attention processing is performed on the temporal features included in the updated spatial-level memory and the convolved second concatenated feature to obtain a second fusion feature. The second fusion feature is split along the batch dimension to obtain multiple spatial-level fusion features. Each spatial-level fusion feature is reconstructed into features at different scales to obtain corresponding spatio-temporal fusion features;

[0011] The decoder performs upsampling step by step on the received convolved spatial features and each spatio-temporal fusion feature, and outputs decoded features;

[0012] The output module processes the decoded features to obtain a target segmentation map;

[0013] The multi-frame infrared small target detection model is trained according to the video frame sequence samples to obtain a trained multi-frame infrared small target detection model, and the trained multi-frame infrared small target detection model is used for target detection.

[0014] A multi-frame infrared small target detection device, the device includes:

[0015] A sample acquisition module for acquiring a video frame sequence sample; the video frame sequence sample is obtained by preprocessing an infrared video;

[0016] A model construction module for constructing a multi-frame infrared small target detection model; the multi-frame infrared small target detection model includes an encoder, a temporal Transformer module, a decoder, and an output module; the encoder and the decoder are connected by a convolutional layer to form a U-shaped network; the temporal Transformer module includes a temporal-channel cross Transformer sub-module and a temporal-spatial cross Transformer sub-module;

[0017] An encoding module, configured to perform progressive downsampling on the current video frame through the encoder to obtain spatial features of multiple different scales, and output the spatial features output by the last-stage downsampling after convolution to the decoder;

[0018] A channel feature fusion module, configured to map each spatial feature to the same size through the temporal channel cross Transformer sub-module and then splice them according to the channel dimension to obtain a first spliced feature, perform dual-output ConvLSTM processing on the first spliced feature to obtain an updated channel-level memory, perform cross-attention processing on the temporal features included in the updated channel-level memory and the convolved first spliced feature to obtain a first fusion feature, and split the first fusion feature according to the channel dimension to obtain multiple channel-level fusion features;

[0019] A spatial feature fusion module, configured to splice each channel-level fusion feature according to the batch dimension through the temporal spatial cross Transformer sub-module to obtain a second spliced feature, perform dual-output ConvLSTM processing on the second spliced feature to obtain an updated spatial-level memory, perform cross-attention processing on the temporal features included in the updated spatial-level memory and the convolved second spliced feature to obtain a second fusion feature, split the second fusion feature according to the batch dimension to obtain multiple spatial-level fusion features, and reconstruct each spatial-level fusion feature into features of different scales to obtain corresponding spatio-temporal fusion features;

[0020] A decoding module, configured to perform progressive upsampling on the received convolved spatial features and each spatio-temporal fusion feature through the decoder to output decoded features;

[0021] A result output module, configured to process the decoded features through the output module to obtain a target segmentation map;

[0022] A target detection module, configured to train the multi-frame infrared small target detection model according to the video frame sequence samples to obtain a trained multi-frame infrared small target detection model, and perform target detection using the trained multi-frame infrared small target detection model.

[0023] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0024] Obtain video frame sequence samples; the video frame sequence samples are obtained by preprocessing an infrared video;

[0025] Construct a multi-frame infrared small target detection model; the multi-frame infrared small target detection model includes an encoder, a temporal Transformer module, a decoder, and an output module; the encoder and the decoder are connected by a convolutional layer to form a U-shaped network; the temporal Transformer module includes a temporal channel cross Transformer sub-module and a temporal spatial cross Transformer sub-module;

[0026] Through the encoder, perform successive downsampling on the current video frame to obtain spatial features of multiple different scales, and output the spatial features output by the last-level downsampling to the decoder after convolution;

[0027] Through the temporal channel cross Transformer sub-module, map each spatial feature to the same size and then splice them according to the channel dimension to obtain a first spliced feature, perform dual-output ConvLSTM processing on the first spliced feature to obtain an updated channel-level memory, perform cross-attention processing on the temporal features included in the updated channel-level memory and the convolved first spliced feature to obtain a first fusion feature, and split the first fusion feature according to the channel dimension to obtain multiple channel-level fusion features;

[0028] Through the temporal spatial cross Transformer sub-module, splice each channel-level fusion feature according to the batch dimension to obtain a second spliced feature, perform dual-output ConvLSTM processing on the second spliced feature to obtain an updated spatial-level memory, perform cross-attention processing on the temporal features included in the updated spatial-level memory and the convolved second spliced feature to obtain a second fusion feature, split the second fusion feature according to the batch dimension to obtain multiple spatial-level fusion features, and reconstruct each spatial-level fusion feature into features of different scales to obtain corresponding spatio-temporal fusion features;

[0029] Through the decoder, perform successive upsampling on the received convolved spatial features and each spatio-temporal fusion feature, and output decoded features;

[0030] Through the output module, process the decoded features to obtain a target segmentation map;

[0031] Train the multi-frame infrared small target detection model according to the video frame sequence samples to obtain a trained multi-frame infrared small target detection model, and use the trained multi-frame infrared small target detection model for target detection.

[0032] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0033] Obtain a video frame sequence sample; the video frame sequence sample is obtained by preprocessing an infrared video;

[0034] Construct a multi-frame infrared small target detection model; the multi-frame infrared small target detection model includes an encoder, a temporal Transformer module, a decoder, and an output module; the encoder and the decoder are connected by a convolutional layer to form a U-shaped network; the temporal Transformer module includes a temporal channel cross Transformer sub-module and a temporal spatial cross Transformer sub-module;

[0035] The encoder performs progressive downsampling on the current video frame to obtain spatial features at multiple different scales, and convolves the spatial features output by the last stage of downsampling and outputs them to the decoder;

[0036] The temporal channel cross Transformer sub-module maps each spatial feature to the same size and then splices them according to the channel dimension to obtain a first spliced feature, performs dual-output ConvLSTM processing on the first spliced feature to obtain an updated channel-level memory, performs cross-attention processing on the temporal features included in the updated channel-level memory and the convolved first spliced feature to obtain a first fusion feature, and splits the first fusion feature according to the channel dimension to obtain multiple channel-level fusion features;

[0037] The temporal spatial cross Transformer sub-module splices each channel-level fusion feature according to the batch dimension to obtain a second spliced feature, performs dual-output ConvLSTM processing on the second spliced feature to obtain an updated spatial-level memory, performs cross-attention processing on the temporal features included in the updated spatial-level memory and the convolved second spliced feature to obtain a second fusion feature, splits the second fusion feature according to the batch dimension to obtain multiple spatial-level fusion features, and reconstructs each spatial-level fusion feature into features at different scales to obtain corresponding spatio-temporal fusion features;

[0038] The decoder performs progressive upsampling on the received convolved spatial features and each spatio-temporal fusion feature, and outputs decoded features;

[0039] The output module processes the decoded features to obtain a target segmentation map;

[0040] Train the multi-frame infrared small target detection model according to the video frame sequence samples to obtain a trained multi-frame infrared small target detection model, and use the trained multi-frame infrared small target detection model for target detection.

[0041] The above multi-frame infrared small target detection method, device, computer device and storage medium first use the encoder of the U-shaped network to process the current frame to obtain spatial features at different levels, and then use the temporal channel cross Transformer sub-module and the temporal spatial cross Transformer sub-module in the temporal Transformer module to achieve cross-fusion of the spatial features of the current frame and the temporal features in memory in the channel dimension and the spatial dimension respectively. At the same time, the dual-output ConvLSTM module is used to process the temporal features and update the memory, and the spatio-temporal features are stored and utilized through the memory mechanism, breaking through the limitation of the prior art that spatio-temporal features must be extracted in a fixed time window, and being able to extract richer spatio-temporal features from long video sequences. Finally, the decoder of the U-shaped network is used to process the fused spatio-temporal features to obtain the target segmentation map, so as to more accurately locate the target and estimate the target shape. The embodiments of the present invention significantly improve the performance of multi-frame infrared small target detection, and can improve the accuracy of small target detection and the accuracy of small target positioning. Description of the Drawings

[0042] Figure 1 It is a schematic flowchart of the multi-frame infrared small target detection method in an embodiment;

[0043] Figure 2 It is a schematic structural diagram of the multi-frame infrared small target detection model in an embodiment;

[0044] Figure 3 It is a schematic structural diagram of the temporal channel cross Transformer sub-module and the temporal spatial cross Transformer sub-module in an embodiment;

[0045] Figure 4 It is the target detection comparison result between the method of the present invention and the DTUM method in different scenarios in an embodiment, wherein, (1) the sequence image represents the sky scene, (2) the sequence image represents the mountain scene, (3) the sequence image represents the forest scene, (4) the sequence image represents the building scene, (5) the sequence image represents the vegetation scene;

[0046] Figure 5 It is a structural block diagram of the multi-frame infrared small target detection device in an embodiment;

[0047] Figure 6 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments

[0048] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0049] In one embodiment, as Figure 1 shown, a multi-frame infrared small target detection method is provided, including the following steps:

[0050] Step 102, obtain a video frame sequence sample.

[0051] The video frame sequence sample is obtained by preprocessing the infrared video. By preprocessing the infrared video to obtain a video frame sequence that is standardized and meets the size requirements. Among them, the calculation process of image standardization is:

[0052] ;

[0053] In the formula, is the unstandardized video frame at time , and are respectively the image mean and standard deviation statistically calculated in the dataset, is the standardized video frame (current frame) at time t .

[0054] The process of size adjustment is: add constant patches on the right side and the lower side of to make the width of W and the height H respectively meet the requirements of being divisible by .

[0055] Step 104, construct a multi-frame infrared small target detection model.

[0056] As Figure 2 shown, the structural schematic diagram of the multi-frame infrared small target detection model. The multi-frame infrared small target detection model includes an encoder, a temporal Transformer module, a decoder, and an output module. The encoder and the decoder are connected by convolutional layers to form a U-shaped network. The temporal Transformer module includes a temporal channel cross Transformer sub-module and a temporal spatial cross Transformer sub-module.

[0057] Input the video frame at the current moment into the encoder of the U-shaped network to obtain spatial features at different levels. Input the spatial features of the current frame and the temporal features included in the memory into the temporal Transformer module for feature cross-fusion, including channel dimension feature fusion of the temporal channel cross Transformer sub-module and spatial dimension feature fusion of the temporal spatial cross Transformer sub-module. The overall calculation method is:

[0058] ;

[0059] In the formula, represents the set of denotes memory, TTM represents the temporal Transformer module. After the residual connection with the hierarchical spatial features the result of spatio-temporal feature fusion at the is the i hierarchy. Among them, the dual-output ConvLSTM module is used to process the temporal features and update the memory. After fusion, the spatio-temporal features are input to the decoder of the U-shaped network to obtain the target segmentation map.

[0060] Step 106: The encoder downsamples the current video frame step by step to obtain spatial features at multiple different scales, and convolves the spatial features output by the last level of downsampling and outputs them to the decoder.

[0061] The encoder of the U-shaped network contains 4 downsampling modules. The downsampling method is max pooling, and the downsampling factor is 2, obtaining spatial features at 4 different hierarchies , ([[]] ) where is the channel dimension.

[0062] Step 108: The temporal channel cross Transformer sub-module maps each spatial feature to the same size and then concatenates them according to the channel dimension to obtain the first concatenated feature. The first concatenated feature is processed by the dual-output ConvLSTM to obtain the updated channel-level memory. Cross-attention processing is performed on the temporal features included in the updated channel-level memory and the convolved first concatenated feature to obtain the first fusion feature. The first fusion feature is split according to the channel dimension to obtain multiple channel-level fusion features.

[0063] Figure 2 Shown in

[0064] is the case of only containing 1 temporal channel cross Transformer sub-module and 1 temporal spatial cross Transformer sub-module. Taking this structure as an example, the data processing process is described as follows: Figure 3 As shown in the structural schematic diagram of the temporal channel cross Transformer sub-module and the temporal spatial cross Transformer sub-module, the temporal channel cross Transformer sub-module first maps to features of the same size c , is the channel dimension, ; then, the 4 different hierarchies of are concatenated according to the channel dimension to obtain ; Next, convolution is used Process Obtain , and process using a dual-output ConvLSTM (D-ConvLSTM) , Memory and that of the previous frame , to obtain the updated memory , and , calculated as follows:

[0065] , ;

[0066] Finally, perform spatio-temporal feature cross-fusion in the channel dimension. First, calculate and perform weighted processing according to the channel dimension attention weights, and then use a feed-forward fully connected network for mapping processing:

[0067] ;

[0068] ;

[0069] where is an optional scaling coefficient, FFN is a feed-forward fully connected network, is the first fusion feature, and after splitting it by dimension, the final result of the temporal-channel cross Transformer sub-module is obtained .

[0070] Step 110, use the spatio-temporal cross Transformer sub-module to perform batch dimension concatenation on each channel-level fusion feature to obtain a second concatenated feature, perform dual-output ConvLSTM processing on the second concatenated feature to obtain the updated spatial-level memory, perform cross-attention processing on the temporal features included in the updated spatial-level memory and the convolved second concatenated feature to obtain a second fusion feature, split the second fusion feature by batch dimension to obtain multiple spatial-level fusion features, and reconstruct each spatial-level fusion feature into different scale features to obtain the corresponding spatio-temporal fusion features.

[0071] As Figure 3 shown, the spatio-temporal cross Transformer sub-module first performs batch dimension concatenation on the input , ( ) to obtain ; Next, use convolution to process to obtain , and process using a dual-output ConvLSTM (D-ConvLSTM) , Memory and that of the previous frame , obtain the updated memory , and , and the calculation is as follows:

[0072] , ;

[0073] Finally, perform spatio-temporal feature cross-fusion in the spatial dimension. First, calculate and weight the attention weights in the spatial dimension, and then use a feed-forward fully connected network for mapping processing:

[0074] ;

[0075] ;

[0076] where is an optional scaling coefficient, is the first fusion feature, and the final result of the spatio-temporal cross Transformer sub-module is obtained after splitting it along the batch dimension .

[0077] In steps 108 and 110, a dual-output ConvLSTM module is used to process the temporal features and update the memory. This module takes , and as inputs, and and as outputs. represents the memory to store t the multi-frame spatio-temporal information before the (-1)th moment. D-ConvLSTM sets a forget gate , an input gate and two output gates , , and the calculation process is as follows:

[0078] ;

[0079] In the formula, represents the activation function, represents pixel-wise multiplication, represents channel-wise concatenation, Ch represents channel-wise splitting, represents convolution, represents the updated memory.

[0080] Step 112: The decoder performs progressive upsampling on the received convolved spatial features and each spatio-temporal fusion feature, and outputs decoded features.

[0081] Step 114: Process the decoded features through the output module to obtain the target segmentation map.

[0082] The decoder of the U-shaped network contains 4 upsampling modules. The upsampling method is bilinear interpolation, and the upsampling factor is 2. It sequentially combines the outputs of the temporal Transformer modules at different levels to finally obtain the output of the detection model .

[0083] Step 116: Train the multi-frame infrared small target detection model based on the video frame sequence samples to obtain the trained multi-frame infrared small target detection model, and use the trained multi-frame infrared small target detection model for target detection.

[0084] In the above multi-frame infrared small target detection method, first, the encoder of the U-shaped network is used to process the current frame to obtain spatial features at different levels. Then, the temporal channel cross Transformer sub-module and the temporal spatial cross Transformer sub-module in the temporal Transformer module are used to achieve the cross-fusion of the spatial features of the current frame and the temporal features in the memory in the channel dimension and the spatial dimension respectively. At the same time, the dual-output ConvLSTM module is used to process the temporal features and update the memory. The spatio-temporal features are stored and utilized through the memory mechanism, breaking through the limitation of the prior art that spatio-temporal features must be extracted in a fixed time window, and being able to extract richer spatio-temporal features from long video sequences. Finally, the decoder of the U-shaped network is used to process the fused spatio-temporal features to obtain the target segmentation map, so as to more accurately locate the target and estimate the target shape. The embodiments of the present invention significantly improve the performance of multi-frame infrared small target detection and can improve the accuracy of small target detection and the accuracy of small target positioning.

[0085] In one embodiment, the number of the temporal channel cross Transformer sub-module and the temporal spatial cross Transformer sub-module is at least one. Among them, when the number is two or more, the temporal channel cross Transformer sub-module and the temporal spatial cross Transformer sub-module are alternately connected in series.

[0086] In one embodiment, performing dual-output ConvLSTM processing on the first concatenated feature to obtain the updated channel-level memory includes: Performing dual-output ConvLSTM processing on the first concatenated feature to obtain the updated channel-level memory as:

[0087] ;

[0088] wherein, is the channel-level memory at time is Channel-level key at a moment is Channel-level value at a moment is the first splicing feature is the dual-output ConvLSTM processing is Channel-level memory at a moment is Channel-level key at a moment is Channel-level value at a moment

[0089] In one embodiment, cross-attention processing is performed on the temporal features included in the updated channel-level memory and the convolved first splicing feature to obtain a first fusion feature. Splitting the first fusion feature along the channel dimension to obtain multiple channel-level fusion features includes: performing cross-attention processing on the temporal features included in the updated channel-level memory and the convolved first splicing feature, and the obtained first fusion feature is:

[0090] ;

[0091] wherein is the first fusion feature, FFN is the feed-forward fully connected network , is the convolved first splicing feature is Channel-level key at a moment is Channel-level value at a moment is the activation function is the scaling coefficient

[0092] In one embodiment, performing dual-output ConvLSTM processing on the second splicing feature to obtain the updated spatial-level memory includes: performing dual-output ConvLSTM processing on the second splicing feature, and the obtained updated spatial-level memory is:

[0093] ;

[0094] wherein is Spatial-level memory at a moment is Spatial-level key at a moment is Spatial-level value at a moment is the second splicing feature is the dual-output ConvLSTM processing is Spatial-level memory at a moment is Spatial-level key at a moment is Spatial-level value at a moment

[0095] In one embodiment, cross-attention processing is performed on the temporal features and the convolved second concatenated features included in the updated spatial-level memory to obtain a second fusion feature. Splitting the second fusion feature along the batch dimension to obtain multiple spatial-level fusion features includes: performing cross-attention processing on the temporal features and the convolved second concatenated features included in the updated spatial-level memory, and the obtained second fusion feature is:

[0096] ;

[0097] wherein is the second fusion feature, FFN is a feed-forward fully connected network , is the convolved first concatenated feature is Spatial-level key at a moment is Spatial-level value at a moment is an activation function is a scaling factor

[0098] In one embodiment, the encoder includes a plurality of downsampling modules connected in sequence; the decoder includes a plurality of upsampling modules connected in sequence

[0099] In one embodiment Figure 4 shows the target detection comparison results of the proposed method and the DTUM method (Li R, An W, Xiao C, et al. Direction-coded Temporal U-shape Module for Multiframe Infrared Small Target Detection[J]. IEEE Transactions on Neural Networks and Learning Systems, 2025, PP(36): 555–568.) in different scenarios. Among them, (1) the sequence image represents the sky scene, (2) the sequence image represents the mountain scene, (3) the sequence image represents the forest scene, (4) the sequence image represents the building scene, (5) the sequence image represents the vegetation scene. In the mountain scene, the target moves slowly. In the forest scene, there are tree occlusion interferences. In the vegetation scene, the background moves fast. According to the output target segmentation map, the proposed method can more accurately locate the positions of small-sized and weak-feature targets and estimate the target shapes, avoiding Figure 4 (2) andFigure 4 The missed detection cases in the results of the DTUM method shown in (3), especially Figure 4 in the complex background shown in (5), avoid the serious false alarm cases in the results of the DTUM method, proving that the proposed method can make better use of the spatio-temporal information in the video sequence to obtain better detection performance.

[0100] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0101] In one embodiment, as Figure 5 shown, a multi-frame infrared small target detection device is provided, including:

[0102] A sample acquisition module 502, configured to acquire a video frame sequence sample; the video frame sequence sample is obtained by preprocessing an infrared video;

[0103] A model construction module 504, configured to construct a multi-frame infrared small target detection model; the multi-frame infrared small target detection model includes an encoder, a temporal Transformer module, a decoder, and an output module; the encoder and the decoder are connected by a convolutional layer to form a U-shaped network; the temporal Transformer module includes a temporal channel cross Transformer sub-module and a temporal spatial cross Transformer sub-module;

[0104] An encoding module 506, configured to perform step-by-step downsampling on the current video frame through the encoder to obtain spatial features of multiple different scales, and output the spatially convolved features of the last-level downsampling output to the decoder;

[0105] The channel feature fusion module 508 is used to map each spatial feature to the same size through the temporal-channel cross Transformer sub-module, splice them according to the channel dimension to obtain the first spliced feature, perform dual-output ConvLSTM processing on the first spliced feature to obtain the updated channel-level memory, perform cross-attention processing on the temporal features and the convolved first spliced feature included in the updated channel-level memory to obtain the first fusion feature, and split the first fusion feature according to the channel dimension to obtain multiple channel-level fusion features;

[0106] The spatial feature fusion module 510 is used to splice each channel-level fusion feature according to the batch dimension through the temporal-spatial cross Transformer sub-module to obtain the second spliced feature, perform dual-output ConvLSTM processing on the second spliced feature to obtain the updated spatial-level memory, perform cross-attention processing on the temporal features and the convolved second spliced feature included in the updated spatial-level memory to obtain the second fusion feature, split the second fusion feature according to the batch dimension to obtain multiple spatial-level fusion features, and reconstruct each spatial-level fusion feature into different scale features to obtain the corresponding spatio-temporal fusion features;

[0107] The decoding module 512 is used to perform progressive upsampling on the received convolved spatial features and each spatio-temporal fusion feature through the decoder to output the decoded features;

[0108] The result output module 514 is used to process the decoded features through the output module to obtain the target segmentation map;

[0109] The target detection module 516 is used to train a multi-frame infrared small target detection model according to the video frame sequence samples to obtain a trained multi-frame infrared small target detection model, and use the trained multi-frame infrared small target detection model to perform target detection.

[0110] For the specific limitations of the multi-frame infrared small target detection device, reference can be made to the limitations of the multi-frame infrared small target detection method in the above text, which will not be elaborated here. Each module in the above multi-frame infrared small target detection device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0111] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a multi-frame infrared small target detection method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0112] Those skilled in the art can understand that Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0113] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method in the above embodiment.

[0114] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps of the method in the above embodiment.

[0115] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0116] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0117] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A multi-frame infrared small target detection method, characterized in that: The method comprises: Acquire video frame sequence samples; the video frame sequence samples are obtained by preprocessing the infrared video; Construct a multi-frame infrared small target detection model; the multi-frame infrared small target detection model includes an encoder, a temporal Transformer module, a decoder and an output module; the encoder and the decoder are connected through a convolutional layer to form a U-shaped network; the temporal Transformer module includes a temporal channel cross Transformer submodule and a temporal space cross Transformer submodule; The encoder downsamples the current video frame step by step to obtain spatial features of multiple scales, and convolves the spatial features output by the last downsampling step and outputs them to the decoder; Mapping each spatial feature to the same size through the temporal channel cross Transformer submodule and then splicing it according to the channel dimension to obtain a first spliced ​​feature, performing dual-output ConvLSTM processing on the first spliced ​​feature to obtain an updated channel-level memory, performing cross-attention processing on the temporal features contained in the updated channel-level memory and the convolved first spliced ​​features to obtain a first fused feature, and splitting the first fused feature according to the channel dimension to obtain multiple channel-level fused features; Through the temporal-spatial cross Transformer submodule, each channel-level fusion feature is spliced ​​according to the batch dimension to obtain a second spliced ​​feature, the second spliced ​​feature is processed by dual-output ConvLSTM to obtain an updated spatial-level memory, the temporal features contained in the updated spatial-level memory and the convolved second spliced ​​features are cross-attention processed to obtain a second fusion feature, the second fusion feature is split according to the batch dimension to obtain multiple spatial-level fusion features, and each spatial-level fusion feature is reconstructed into features of different scales to obtain corresponding spatiotemporal fusion features; The decoder performs step-by-step upsampling on the received convolved spatial features and each spatiotemporal fusion feature, and outputs decoded features; Processing the decoded features through the output module to obtain a target segmentation map; The multi-frame infrared small target detection model is trained according to the video frame sequence samples to obtain a trained multi-frame infrared small target detection model, and the trained multi-frame infrared small target detection model is used to perform target detection.

2. The method according to claim 1, characterized in that The number of the temporal channel crossover Transformer submodule and the temporal space crossover Transformer submodule is at least one, wherein when the number is two or more, the temporal channel crossover Transformer submodule and the temporal space crossover Transformer submodule are alternately connected in series.

3. The method according to claim 1, characterized in that: The performing of dual-output ConvLSTM processing on the first concatenated feature to obtain an updated channel-level memory comprises: The first concatenated feature is processed by dual-output ConvLSTM to obtain the updated channel-level memory: ; in, for Channel-level memory of the moment, for The channel-level key at the moment, for The channel level value at the moment, is the first splicing feature, For dual-output ConvLSTM processing, for Channel-level memory of the moment, for The channel-level key at the moment, for The channel-level value at the moment.

4. The method according to claim 1, characterized in that: The time series features contained in the updated channel-level memory and the first concatenated features after convolution are cross-attended to obtain the first fusion features including: The time series features contained in the updated channel-level memory and the first concatenated features after convolution are cross-attended to obtain the first fused feature: ; in, is the first fusion feature, FFN is a feed-forward fully connected network, , is the first concatenated feature after convolution, for The channel-level key at the moment, for The channel level value at the moment, is the activation function, is the scaling factor.

5. The method according to claim 1, characterized in that The updated spatial level memory obtained by performing dual-output ConvLSTM processing on the second concatenated feature includes: The updated spatial level memory obtained by performing dual-output ConvLSTM processing on the second concatenated feature is: ; in, for The spatial memory of the moment, for The spatial key of the moment, for The spatial level value at the moment, is the second splicing feature, For dual-output ConvLSTM processing, for The spatial memory of the moment, for The spatial key of the moment, for The spatial level value at the moment.

6. The method according to claim 1, characterized in that The temporal features contained in the updated spatial-level memory and the second concatenated features after convolution are cross-attended to obtain the second fused features including: The temporal features contained in the updated spatial-level memory and the second concatenated features after convolution are cross-attended to obtain the second fused feature: ; in, is the second fusion feature, FFN is a feed-forward fully connected network, , is the first concatenated feature after convolution, for The spatial key of the moment, for The spatial level value at the moment, is the activation function, is the scaling factor.

7. The method according to claim 1, characterized in that The encoder includes a plurality of down-sampling modules connected in sequence; the decoder includes a plurality of up-sampling modules connected in sequence.

8. A multi-frame infrared small target detection device, characterized in that: The device comprises: A sample acquisition module is used to acquire video frame sequence samples; the video frame sequence samples are obtained by preprocessing the infrared video; A model building module, used to build a multi-frame infrared small target detection model; the multi-frame infrared small target detection model includes an encoder, a temporal Transformer module, a decoder and an output module; the encoder and the decoder are connected through a convolutional layer to form a U-shaped network; the temporal Transformer module includes a temporal channel cross Transformer submodule and a temporal space cross Transformer submodule; The encoding module is used to downsample the current video frame step by step through the encoder to obtain spatial features of multiple different scales, and convolute the spatial features output by the last downsampling step and output them to the decoder; A channel feature fusion module is used to map each spatial feature to the same size through the temporal channel cross Transformer submodule and then splice it according to the channel dimension to obtain a first spliced ​​feature, perform dual-output ConvLSTM processing on the first spliced ​​feature to obtain an updated channel-level memory, perform cross-attention processing on the temporal features contained in the updated channel-level memory and the convolved first spliced ​​features to obtain a first fused feature, and split the first fused feature according to the channel dimension to obtain multiple channel-level fused features; A spatial feature fusion module is used to splice each channel-level fusion feature according to the batch dimension through the temporal-spatial cross Transformer submodule to obtain a second spliced ​​feature, perform dual-output ConvLSTM processing on the second spliced ​​feature to obtain an updated spatial-level memory, perform cross-attention processing on the temporal features contained in the updated spatial-level memory and the convoluted second spliced ​​features to obtain a second fusion feature, split the second fusion feature according to the batch dimension to obtain multiple spatial-level fusion features, and reconstruct each spatial-level fusion feature into features of different scales to obtain corresponding spatiotemporal fusion features; A decoding module, used for upsampling the received convolved spatial features and each spatiotemporal fusion feature step by step through the decoder, and outputting decoded features; A result output module, used for processing the decoded features through the output module to obtain a target segmentation map; The target detection module is used to train the multi-frame infrared small target detection model according to the video frame sequence samples to obtain the trained multi-frame infrared small target detection model, and use the trained multi-frame infrared small target detection model to perform target detection.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Link prediction method fusing attention mechanism and graph contrast learning

    CN118210976A