A multi-frame infrared small target detection method and system based on time-space integrated network
Through the method based on the integrated space-time network, the airspace and motion characteristics of weak infrared targets are extracted, and the problem of insufficient detection capabilities of weak infrared targets in the prior art is solved, thereby achieving more efficient target detection.
Patent Information
- Application Number
- CN202411200418.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-08-29
AI Technical Summary
When detecting weak infrared targets with low signal-to-miss ratios, it is difficult to effectively distinguish the target from the background, resulting in insufficient detection capabilities.
Using a method based on a space-time integrated network, the spatial significance characteristics of the space are extracted frame by frame through the U-shaped network, combined with a fixed-weight multi-scale three-dimensional motion feature convolution kernel, and a spatial-temporal feature fusion module of 3D convolution is used to establish a mapping of features to infrared weak target probability likelihood maps.
It significantly improves the detection ability of low-information and miscellaneous targets that are weaker than infrared, can more effectively distinguish targets from background, enhances space-time significance, and suppresses clutter and noise.
Smart Images

Figure CN119169263B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of infrared small target detection and processing, and in particular relates to a multi-frame infrared small target detection method and system based on a time-space integrated network. Background Art
[0002] Infrared imaging has the advantages of good concealment, strong anti-interference, long imaging distance and low cost, and is widely used in infrared search and tracking, civil maritime surveillance and vehicle detection. As a key technology, infrared small target detection aims to accurately locate the target of interest (such as drone) in infrared images. Due to the influence of background clutter and detection distance, infrared small targets only have a few pixels to dozens of pixels in the image, and their shape features are limited and the texture is unclear, which makes them easily submerged in background clutter and noise. In view of the above characteristics, infrared small target detection is a very challenging problem task.
[0003] In order to detect infrared weak targets, many new methods have been proposed in recent years, which can be divided into two categories: traditional methods based on model-driven and intelligent methods based on data-driven. Traditional methods based on model-driven conduct in-depth analysis of the characteristics of infrared weak targets and their differences from background, clutter, and noise, and build hand-made feature models and prior knowledge to enhance the infrared weak target area and suppress background, clutter, and noise. Intelligent methods based on data-driven mainly refer to methods based on deep learning, which intelligently and automatically extract deep dimensional features for separating targets and non-targets from a large amount of labeled data. The strong representation ability makes the deep learning-based methods perform well. However, for infrared weak targets with low signal-to-clutter ratio, the targets are blurred and not spatially significant. Whether it is traditional methods or deep learning methods, they can only use the spatial domain method and cannot effectively detect infrared weak targets with low signal-to-clutter ratio.
[0004] The space-time integrated design method realizes the detection of infrared weak targets by superimposing and accumulating the significant features in the spatial domain to form a space-time feature tensor, and then uses a 3D convolution-based network to map the space-time feature tensor into the probability of infrared weak targets. This method can greatly improve the detection capability of infrared weak targets with low signal-to-clutter ratio. Summary of the invention
[0005] In order to overcome the shortcomings of existing methods in detecting infrared dim targets with low signal-to-clutter ratio, the present invention proposes a multi-frame detection method and system for infrared dim targets based on a spatiotemporal integrated network. The proposed network first uses a U-type network to extract spatial domain saliency features frame by frame, then uses the proposed fixed weight multi-scale three-dimensional motion feature convolution kernel to extract the motion features of the target, and finally uses a spatiotemporal feature fusion module based on 3D convolution to establish a mapping from features to infrared dim target probability likelihood maps. The network is trained on the established data set using the constructed loss function to obtain the optimal weight parameters of the model. The obtained weight parameters can be used to obtain the infrared dim target probability likelihood map output by the network, and the maximum value method is used to fuse the infrared dim target probability likelihood map output by the network. After adaptive threshold segmentation, the position of the infrared dim target is finally obtained. After testing and analysis, the proposed spatiotemporal integrated infrared dim target detection method has better detection capabilities for infrared dim targets with low signal-to-clutter ratio than the existing methods.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] A multi-frame infrared small target detection method based on a time-space integrated network comprises the following steps:
[0008] S1: Build a spatiotemporal integrated infrared small target detection network model;
[0009] S2: Obtain an infrared image sample data set and use the infrared image sample data set to train the spatiotemporal integrated infrared dim small target detection network model;
[0010] S3: Input the image sequence to be detected into the trained spatiotemporal integrated infrared dim small target detection network model to obtain the network output, use the maximum value method to perform result-level fusion to obtain the fusion result, perform threshold segmentation on the fusion result, and obtain the infrared dim small target.
[0011] Preferably, in S1, the spatiotemporal integrated infrared dim small target detection network model includes: a spatial domain saliency feature generation network, a motion saliency extraction layer and a spatiotemporal saliency extraction network.
[0012] Preferably, the method for building a spatial salient feature generation network includes:
[0013] The input of the spatial salient feature generation network is an infrared sequence image Among them, 5 represents the length of the input sequence, W and H represent the width and height of the image respectively;
[0014] For a single frame in a sequence A spatial saliency feature generation network with a U-shaped network as the backbone network is used to extract spatial saliency. All frame images in the sequence share the weight parameters of the spatial saliency feature generation network.
[0015] The spatial salient feature generation network consists of four neural network submodules: Stem, Layer, Up, and Head. The number of output channels is 32. A single frame image is passed through the spatial salient feature generation network to obtain the spatial salient feature tensor.
[0016] The spatial salient feature tensor S of the five frames of images 5×32×W×H Reorganize into space-time tensor ST according to the time dimension 32 ×5×W×H , used for subsequent extraction of temporal motion features.
[0017] Preferably, the motion saliency extraction layer is composed of 32 designed fixed-weight 3D convolution kernels, divided into 4 scales, each scale has 8 3D motion directions;
[0018] Among them, the scales include 4 sizes: 5×5×5, 5×7×7, 5×9×9, and 5×11×11; the depth direction of each convolution kernel is 5; the 8 three-dimensional motion directions include: left-right, upper left-lower right, up-down, upper right-lower left, right-left, lower right-upper left, down-up, and lower left-upper right, which describe the motion direction of infrared weak targets in the time domain.
[0019] Preferably, the method for constructing the motion saliency extraction layer includes:
[0020] The obtained 32 fixed-weight 3D convolution kernels are concatenated into the motion saliency extraction layer 3D-FWCK Di , which is expressed as follows:
[0021]
[0022] Where Di represents the size of the convolution kernel, cat[·] represents the concatenation operator, Represents the convolution kernel in the left-right direction corresponding to the convolution kernel size, Represents the convolution kernel in the upper left-lower right direction under the corresponding convolution kernel size, Represents the convolution kernel in the up-down direction corresponding to the convolution kernel size, Represents the convolution kernel in the upper right-lower left direction under the corresponding convolution kernel size, Represents the convolution kernel in the right-left direction under the corresponding convolution kernel size, Represents the convolution kernel in the lower right-upper left direction under the corresponding convolution kernel size, represents the convolution kernel in the down-up direction corresponding to the convolution kernel size, Represents the convolution kernel in the lower left-upper right direction under the corresponding convolution kernel size;
[0023] The space-time tensor ST 32×5×W×H With 3D-FWCK Di Perform convolution to obtain the spatiotemporal motion feature tensor ST M , which is expressed as follows:
[0024] ST M =ST 32×5×W×H *3D-FWCK Di ,
[0025] Specifically, the space-time tensor ST 32×5×W×H Convolution is performed with the convolution kernel of each scale to generate a tensor of size 8×5×W×H. The tensors at the four scales are concatenated to obtain the spatiotemporal motion feature tensor ST. M , with a scale of 32×5×W×H, and the space-time tensor ST 32×5×W×H Be consistent.
[0026] Preferably, the spatiotemporal saliency extraction network based on 3D convolution transforms the spatiotemporal motion feature tensor ST M Mapped to the probability likelihood graph of infrared dim targets, the spatiotemporal saliency extraction network consists of 5 parallel spatiotemporal saliency modules. Each module operates independently in parallel and does not share weights, as shown below:
[0027] P t =STSM t (ST M )t∈[1,2,3,4,5],
[0028] Among them, STSM t represents the spatiotemporal saliency module operation of the t-th frame, P t is the probability likelihood map of infrared dim targets in the tth frame;
[0029] The spatiotemporal saliency module adopts a U-shaped network as the backbone structure, which consists of 3D Stem, 3D Down, 3D Up, and 3D Head network layers. After 3D convolution, 3D pooling, and 3D normalization operations, the spatiotemporal motion feature tensor ST M Mapped into infrared small target probability likelihood map.
[0030] Preferably, in S2, the acquired infrared image sample data set is composed of a training set and a test set, each of which includes multiple groups of data, each group of data includes a three-dimensional matrix Image_mat composed of 5 original images and a Mask_mat composed of 5 corresponding labels; wherein the original image is an infrared image of 5 consecutive frames, and the target position is manually marked according to the infrared image and recorded as 1, and the non-target position pixel is set to 0 to obtain a label; the image and the label are superimposed according to the time dimension to form Image_mat and Mask_mat;
[0031] The method of using the infrared image sample data set to train the spatiotemporal integrated infrared dim small target detection network model includes:
[0032] Input the training set sample images into the spatiotemporal integrated infrared dim small target detection network model to obtain the network output results of each set of training sample images;
[0033] Based on the network output results, construct a loss function;
[0034] Calculate the loss value according to the loss function;
[0035] Based on the loss value, determine whether the spatiotemporal integrated infrared dim small target detection network model converges; if so, determine that the network model training is completed, save the network model parameters obtained by the training, otherwise continue training;
[0036] The loss function is constructed by calculating the focal loss of the corresponding label frame by frame for the probability likelihood graphs of 5 frames of infrared dim targets output by the network, and taking the sum of the focal losses of the 5 frames as the total network loss; it is expressed as follows:
[0037] Loss=∑FL t t∈[1,2,3,4,5],
[0038] Where: FL refers to focus loss.
[0039] Preferably, in S3, the method of inputting the image sequence to be detected into the trained spatiotemporal integrated infrared small target detection network model to obtain the network output includes:
[0040] Slide the sequence of input T frames along the time direction to obtain the 5 consecutive frames of images required by the network as input, where the sliding interval is 1 frame, and obtain T-4 image groups, which are recorded as:
[0041] {G p ,p∈{1,2,...,T-4}},
[0042] The trained spatiotemporal integrated infrared small target detection network model is used to process T-4 image groups, and the network processing results are continuously obtained, which are recorded as:
[0043] {IR p ,p∈{1,2,...,T-4}},
[0044] Among them: IR p The result includes 5 frames of images, namely IR p = {IR p_1 ,IR p_2 ,IR p_3 ,IR p_4 ,IR p_5}.
[0045] Preferably, in S3, the method of using the maximum value method to perform result-level fusion to obtain the fusion result includes:
[0046] R K (i,j)=Max[IR (K-4)_5 (i,j),IR (K-3)_4 (i,j),IR (K-2)_3 (i,j),IR (K-1)_2 (i,j),IR K_1 (i,j)]
[0047] 1≤i≤W,1≤j≤H,1≤K≤T
[0048] Among them, (i, j) is the pixel position, R K (i, j) is the probability of a weak infrared target at pixel (i, j) in the Kth frame, Max[·] is the maximum value operation, IR (K-4)_5 (i, j) is the processing result of the network, (K-4)_5 represents the 5th frame of the network processing result of the K-4th image group, IR (K-3)_4 (i, j) is the processing result of the network, (K-3)_4 represents the 4th frame of the network processing result of the K-3th image group, IR (K-2)_3 (i, j) is the processing result of the network, (K-2)_3 represents the third frame of the network processing result of the K-2th image group, IR (K-1)_2 (i, j) is the processing result of the network, (K-1)_2 represents the second frame of the network processing result of the K-1th image group, IR (K-1) (i, j) is the processing result of the network, (K-1) represents the first frame of the network processing result of the K-1th image group, W and H represent the width and height of the image respectively, and T represents the number of frames in the sequence.
[0049] The present invention also provides an infrared weak target multi-frame detection system based on a time-space integrated network, comprising: a construction module, a training module and a detection module;
[0050] The building module is used to build a time-space integrated infrared small target detection network model;
[0051] The training module is used to obtain an infrared image sample data set and use the infrared image sample data set to train the spatiotemporal integrated infrared dim small target detection network model;
[0052] The detection module is used to input the image sequence to be detected into the trained spatiotemporal integrated infrared dim small target detection network model to obtain the network output, use the maximum value method to perform result-level fusion to obtain the fusion result, perform threshold segmentation on the fusion result, and obtain the infrared dim small target.
[0053] Compared with the prior art, the present invention has the following beneficial effects:
[0054] (1) The present invention proposes a spatial-temporal integrated infrared dim small target multi-frame detection network, which integrates the spatial salient features of the target and the temporal motion features of the target into a unified network architecture based on 3D convolution, and realizes end-to-end infrared dim small target multi-frame detection.
[0055] (2) According to the consistency of the short-term motion direction of infrared dim targets, the present invention proposes a multi-scale three-dimensional motion feature convolution kernel with fixed weights to enhance the spatiotemporal significance of infrared dim targets, while suppressing randomly generated clutter and noise, and improving the detection capability of infrared dim targets with low signal-to-clutter ratio.
[0056] (3) The present invention designs a spatiotemporal feature fusion processing module, which utilizes the spatiotemporal perception capability of 3D convolution to effectively fuse spatiotemporal features. At the same time, it uses five parallel and independently optimized spatiotemporal feature fusion processing modules to map the spatiotemporal tensor into a multi-frame infrared weak target probability likelihood map.
[0057] (4) The present invention proposes a multi-frame network method for applying to sequence images, designs a maximum value method for result-level data fusion, makes full use of context information, and improves the detection effect of infrared weak small targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0059] Figure 1 It is a multi-frame detection flow chart of infrared weak small target based on time-space integrated network provided by an embodiment of the present invention;
[0060] Figure 2 It is a general structural block diagram of the time-space integrated network of an embodiment of the present invention;
[0061] Figure 3 It is a structural diagram of each module of the spatial significant feature generation network according to an embodiment of the present invention;
[0062] Figure 4 is a schematic diagram of eight movement directions of an embodiment of the present invention;
[0063] Figure 5 is a network structure diagram of spatiotemporal saliency extraction according to an embodiment of the present invention;
[0064] Figure 6 1 is a diagram of each network structure of the spatiotemporal saliency module according to an embodiment of the present invention;
[0065] Figure 7 is a graph showing the results of the algorithm of the embodiment of the present invention and the comparative algorithm on the test data;
[0066] Figure 8 is the ROC curve of the algorithm of the embodiment of the present invention and the comparative algorithm on the test data;
[0067] Fig. 9 The infrared small target multi-frame detection system is composed of the present invention. DETAILED DESCRIPTION
[0068] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0069] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0070] Embodiment 1
[0071] The image data in the embodiment of the present invention is a sequence of infrared small target images, with a total of 104 sequences, each of which contains 100 frames of images, 100 sequences as training data, 4 sequences as test data, and the signal-to-clutter ratio of infrared weak small targets on the test data is less than 3.
[0072] See also Figure 1 The embodiment of the present invention proposes an infrared dim small target detection method based on a time-space integrated network, comprising the following steps:
[0073] Step 1: Build a spatiotemporal integrated infrared small target detection network model;
[0074] Step 2: Obtain an infrared image sample data set and use the infrared image sample data set to train the spatiotemporal integrated infrared dim small target detection network model;
[0075] Step 3: Input the image sequence to be detected into the trained spatiotemporal integrated infrared dim small target detection network model to obtain the network output, use the maximum value method to perform result-level fusion to obtain the fusion result, perform threshold segmentation on the fusion result, and obtain the infrared dim small target.
[0076] In this embodiment, step 1: building a time-space integrated infrared dim small target detection network model.
[0077] Through the design of time-space integration, the network can autonomously learn the time-space characteristics of infrared weak targets to enhance the detection capability of weak targets with extremely low signal-to-clutter ratio. Figure 2 The spatiotemporal integrated infrared dim target detection network is mainly composed of three parts: spatial saliency feature generation network, motion saliency extraction layer, and spatiotemporal saliency extraction network.
[0078] Step 1.1: Build a spatial saliency feature generation network
[0079] The input of the network is a sequence of infrared images Where 5 represents the length of the input sequence, W and H represent the width and height of the image respectively. A spatial saliency feature generation network with a U-shaped network as the backbone network is used to extract spatial saliency. All frame images in the sequence share the weight parameters of the spatial saliency feature generation network. Figure 3 The spatial salient feature generation network consists of four neural network submodules: Stem, Layer, Up, and Head. The number of its output channels is 32. A single frame image can be passed through the spatial salient feature generation network to obtain the spatial salient feature tensor.
[0080] The feature tensor S of the 5 frames of images 5×32×W×H Reorganize into space-time tensor ST according to the time dimension 32×5×W×H , used for subsequent extraction of temporal motion features:
[0081]
[0082] Among them: cat[·] represents the concatenation operator, They respectively represent the tensors extracted by the spatial saliency feature generation network for the 1st to 5th frame images.
[0083] Step 1.2: Build the motion saliency extraction layer
[0084] The motion saliency extraction layer consists of 32 designed fixed-weight 3D convolution kernels, divided into 4 scales, see Figure 4 , there are 8 motion directions at each scale. The scales include 5×5×5, 5×7×7, 5×9×9, and 5×11×11. The depth direction of each convolution kernel is 5. The 8 three-dimensional motion directions include: left-right (LR), upper left-lower right (LU-RD), upper-lower (UD), upper right-lower left (RU-LD), right-left (RL), lower right-upper left (RD-LU), lower-up (DU), and lower left-upper right (LD-RU), which describe the motion direction of infrared weak targets in the time domain.
[0085] Step 1.2.1: Construct a fixed-weight 3D convolution kernel
[0086] The parameters of the fixed convolution kernel with a scale of 5×5×5 in 8 directions are set as follows:
[0087]
[0088] The sum of all elements in each two-dimensional matrix is 1, and the elements with non-zero values are arranged in time and space according to the corresponding movement direction.
[0089] The parameter settings of the fixed convolution kernels at the other three scales refer to the 5×5×5 fixed convolution kernel. The difference is that the assignment positions and values are different in different spatial dimensions, but the direction rule remains unchanged.
[0090] The convolution kernel with a scale of 5×7×7 and left-right (LR) direction is set as follows, and the other 7 directions can be obtained similarly:
[0091]
[0092] The convolution kernel with a scale of 5×9×9 and left-right (LR) direction is set as follows, and the other 7 directions can be obtained similarly:
[0093]
[0094] The convolution kernel with a scale of 5×11×11 in the left-right (LR) direction is set as follows, and the other 7 directions can be obtained similarly:
[0095]
[0096] Step 1.2.2: Generate motion saliency extraction layer
[0097] The obtained 32 fixed-weight 3D convolution kernels are concatenated into the motion saliency extraction layer 3D-FWCK Di , which is expressed as follows:
[0098]
[0099] Where: Di represents the size of the convolution kernel, cat[·] represents the concatenation operator, Represents the convolution kernel in the left-right direction corresponding to the convolution kernel size, Represents the convolution kernel in the upper left-lower right direction under the corresponding convolution kernel size, Represents the convolution kernel in the up-down direction corresponding to the convolution kernel size, Represents the convolution kernel in the upper right-lower left direction under the corresponding convolution kernel size, Represents the convolution kernel in the right-left direction under the corresponding convolution kernel size, Represents the convolution kernel in the lower right-upper left direction under the corresponding convolution kernel size, represents the convolution kernel in the down-up direction corresponding to the convolution kernel size, Represents the convolution kernel in the lower left-upper right direction under the corresponding convolution kernel size.
[0100] Space-time tensor ST 32×5×W×H With 3D-FWCK Di Convolution is performed to obtain the spatiotemporal motion feature tensor ST M :
[0101] ST M =ST 32×5×W×H *3D-FWCK Di
[0102] Space-time tensor ST 32×5×W×H Convolution with the convolution kernel of each scale can generate a tensor of 8×5×W×H. The spatiotemporal motion feature tensor ST is obtained by concatenating the tensors at the four scales. M The scale of is 32×5×W×H, which is consistent with the space-time tensor.
[0103] Step 1.3 Spatiotemporal Saliency Extraction Network
[0104] The spatiotemporal saliency extraction network based on 3D convolution maps the spatiotemporal motion feature tensor into a probability likelihood map of infrared dim targets, see Figure 5 , which consists of 5 parallel spatiotemporal saliency modules, each of which operates independently in parallel and does not share weights.
[0105] P t =STSM t (ST M )t∈[1,2,3,4,5]
[0106] Among them, STSM t represents the spatiotemporal saliency module operation of the t-th frame, P t is the probability likelihood map of infrared dim targets in the tth frame.
[0107] See also Figure 6 The spatiotemporal saliency module adopts a U-shaped network as the backbone structure, which is composed of 3D Stem, 3D Down, 3D Up, 3D Head and other network layers. After 3D convolution, 3D pooling, 3D normalization and other operations, the spatiotemporal motion feature tensor is mapped into an infrared weak target probability likelihood map.
[0108] In this embodiment, step 2: prepare a data set and use the prepared data to perform network training to obtain a model and its parameters.
[0109] Step 2.1: Prepare the dataset
[0110] The dataset consists of a training set and a test set. Both the training set and the test set include multiple sets of data. Each set of data includes a three-dimensional matrix Image_mat consisting of 5 original images and a Mask_mat consisting of 5 corresponding labels. The original image is an infrared image of 5 consecutive frames. The target position is manually marked according to the infrared image and recorded as 1, and the non-target position pixels are set to 0 to obtain the label. The image and the label are superimposed according to the time dimension to form Image_mat and Mask_mat.
[0111] Step 2.2: Construct the loss function
[0112] For the 5 frames of infrared small target probability likelihood maps output by the network, the focal loss between them and the corresponding labels is calculated frame by frame, and the sum of the focal losses of the 5 frames is taken as the total network loss.
[0113] Loss=∑FL t t∈[1,2,3,4,5]
[0114] Where: FL refers to focus loss.
[0115] Step 2.3: Network training
[0116] Input the training set sample images into the network structure of step 1, obtain the network output results of each group of training sample images, and calculate the loss value according to the loss function constructed in step 2.2; based on the loss value, determine whether the target detection network converges; if so, determine that the network training is completed and save the network model parameters obtained by training.
[0117] In this embodiment, step 3: input the image sequence to be detected into the network model trained in step S2 to obtain the output of the network, and finally use the maximum value method to perform result-level fusion to obtain the fusion result.
[0118] Step 3.1 Pipeline to obtain the output of the network
[0119] When the present invention is applied to a sequence of infrared images with a length of T frames, the sequence of input T frames is first slid along the time direction to obtain 5 consecutive frames of images required by the network as input, wherein the sliding interval is 1 frame, and T-4 image groups can be obtained, which are recorded as:
[0120] {G p ,p∈{1,2,...,T-4}}.
[0121] Then, the network model obtained in step 2 and the trained parameters are used to process the image group, and the network processing results can be continuously obtained, which are recorded as:
[0122] {IR p ,p∈{1,2,...,T-4}},
[0123] Among them: IR p Includes the results of 5 frames of images, namely IR p = {IR p_1 ,IR p_2 ,IR p_3 ,IR p_4 ,IR p_5}.
[0124] Step 3.2 Result-level fusion
[0125] When a sequence of length T frames is processed by a sliding method with an interval of 1 frame and a dimension of 5 frames, the obtained result data will overlap. For the overlapping result data, the present invention adopts the maximum value method to perform result-level fusion to obtain the output, and the calculation method is as follows:
[0126] R K (i,j)=Max[IR (K-4)_5 (i,j),IR (K-3)_4 (i,j),IR (K-2)_3 (i,j),IR (K-1)_2 (i,j),IR K_1 (i,j)]
[0127] 1≤i≤W,1≤j≤H,1≤K≤T
[0128] Where: (i, j) is the pixel position, R K (i, j) is the probability of a weak infrared target at pixel (i, j) in the Kth frame, Max[·] is the maximum value operation, IR (K-4)_5 (i, j) is the processing result of the network, (K-4)_5 represents the 5th frame of the network processing result of the K-4th image group, IR (K-3)_4 (i, j) is the processing result of the network, (K-3)_4 represents the 4th frame of the network processing result of the K-3th image group, IR (K-2)_3(i, j) is the processing result of the network, (K-2)_3 represents the third frame of the network processing result of the K-2th image group, IR (K-1)_2 (i, j) is the processing result of the network, (K-1)_2 represents the second frame of the network processing result of the K-1th image group, IR (K-1) (i, j) is the processing result of the network, (K-1) represents the first frame of the network processing result of the K-1th image group, W and H represent the width and height of the image respectively, and T represents the number of frames in the sequence.
[0129] Through the above method, the probability likelihood map of infrared weak targets in each frame image {R K ,K∈{1,2,...,T}}, and perform threshold segmentation on it: those greater than the threshold are considered as targets, and those less than the threshold are considered as non-targets. In this embodiment, the threshold is set to 0.6.
[0130] The infrared small target detection performance of the present invention is quantitatively compared with other algorithms on four sequences, and the comparison methods include MPCM (Multiscale Patch-based Contrast Measure), RLCM (Multiscale Relative Local Contrast Measure), TLLCM (Tri-Layer Local Contrast Measure), NIPPS (Non-negative Infrared Patch-image model), RIPT (Reweighted Infrared Patch-Tensor), STLCF (Spatial-Temporal Local Contrast Filter), ISTDU_Net (Infrared Small-Target Detection U-Net), RISTD_Net (Robust Infrared Small Target Detection Network), and DTUM_Net (Direction-coded Temporal U-shape Module).
[0131] See also Figure 7, shows the processing results of different methods on a randomly selected frame of the test sequence, where the dotted rectangle box is missed detection, the solid rectangle box is accurate detection, and the solid ellipse mark is a false alarm. The method proposed in the present invention successfully detects all targets without false alarms. Figure 8 , the ROC (Receiver Operating Characteristic) curves of different methods on 4 test sequences are plotted. The ROC curve of the method proposed in the present invention is better than that of other comparison methods.
[0132] The three indicators of AUC (Area Under Curve), Signal to Clutter Ratio Gain (SCRG), and Background Suppression Factor (BSF) were calculated. The results are shown in the following table. The proposed method achieved good performance.
[0133] Table 1 is a comparison table of the performance of each method
[0134]
[0135]
[0136] It can be seen that the method proposed in the present invention utilizes an end-to-end spatiotemporal integrated network to accumulate the spatial saliency of small infrared targets according to the prior direction of motion, thereby achieving effective detection of weak infrared targets with low signal-to-clutter ratio, and has better performance than existing methods.
[0137] Embodiment 2
[0138] The present invention also provides an infrared dim target multi-frame detection system based on a time-space integrated network, which is characterized by comprising: a construction module, a training module and a detection module, see Fig. 9 ;
[0139] The construction module is used to build a spatial-temporal integrated infrared small target detection network model;
[0140] The training module is used to obtain an infrared image sample data set and use the infrared image sample data set to train the spatiotemporal integrated infrared dim small target detection network model;
[0141] The detection module is used to input the image sequence to be detected into the trained spatiotemporal integrated infrared dim small target detection network model to obtain the network output, use the maximum value method to perform result-level fusion to obtain the fusion result, perform threshold segmentation on the fusion result, and obtain the infrared dim small target.
[0142] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. A multi-frame infrared small target detection method based on a time-space integrated network, characterized in that: The following steps are involved: S1: Build a spatiotemporal integrated infrared small target detection network model; S2: Obtain an infrared image sample data set and use the infrared image sample data set to train the spatiotemporal integrated infrared dim small target detection network model; S3: inputting the image sequence to be detected into the trained spatiotemporal integrated infrared dim small target detection network model to obtain the network output, using the maximum value method to perform result-level fusion to obtain the fusion result, performing threshold segmentation on the fusion result, and obtaining the infrared dim small target; In S1, the spatiotemporal integrated infrared dim small target detection network model includes: a spatial domain saliency feature generation network, a motion saliency extraction layer and a spatiotemporal saliency extraction network; Methods for building a spatial saliency feature generation network include: The input of the spatial domain salient feature generation network is an infrared sequence image , where 5 represents the length of the input sequence, W and H represent the width and height of the image respectively; For a single frame in a sequence A spatial saliency feature generation network with a U-shaped network as the backbone network is used to extract spatial saliency. All frame images in the sequence share the weight parameters of the spatial saliency feature generation network. The spatial salient feature generation network consists of four neural network submodules: Stem, Layer, Up, and Head. The number of output channels is 32. A single frame image is passed through the spatial salient feature generation network to obtain the spatial salient feature tensor. ; The spatial salient feature tensor of the five frames of images Reorganize into space-time tensor according to the time dimension , used for subsequent extraction of temporal motion features; The motion saliency extraction layer is composed of 32 designed fixed-weight 3D convolution kernels, divided into 4 scales, each with 8 3D motion directions; The method for constructing the motion saliency extraction layer includes: The obtained 32 fixed-weight 3D convolution kernels are concatenated into a motion saliency extraction layer ; The space-time tensor and Perform convolution to obtain the spatiotemporal motion feature tensor ; The spatiotemporal saliency extraction network based on 3D convolution transforms the spatiotemporal motion feature tensor Mapped to the probability likelihood graph of infrared dim targets, the spatiotemporal saliency extraction network consists of 5 parallel spatiotemporal saliency modules. Each module operates independently in parallel and does not share weights, as shown below: , in, represents the spatiotemporal saliency module operation at the t-th frame, is the probability likelihood map of infrared dim targets in the tth frame; The spatiotemporal saliency module adopts a U-shaped network as the backbone structure, which consists of 3D Stem, 3D Down, 3D Up, and 3D Head network layers. After 3D convolution, 3D pooling, and 3D normalization operations, the spatiotemporal motion feature tensor is transformed into Mapped into infrared small target probability likelihood map.
2. The infrared small target multi-frame detection method based on time-space integrated network according to claim 1 is characterized in that: The scales include 4 sizes: 5×5×5, 5×7×7, 5×9×9, and 5×11×11; the depth direction of each convolution kernel is 5; the 8 three-dimensional motion directions include: left-right, upper left-lower right, upper-lower right, right-left, lower right-upper left, lower-upper, and lower left-upper right, which describe the motion direction of infrared weak targets in the time domain.
3. The infrared small target multi-frame detection method based on time-space integrated network according to claim 2 is characterized in that: in, represents the size of the convolution kernel, represents the concatenation operator, Represents the convolution kernel in the left-right direction corresponding to the convolution kernel size, Represents the convolution kernel in the upper left-lower right direction under the corresponding convolution kernel size, Represents the convolution kernel in the up-down direction corresponding to the convolution kernel size, Represents the convolution kernel in the upper right-lower left direction under the corresponding convolution kernel size, Represents the convolution kernel in the right-left direction under the corresponding convolution kernel size, Represents the convolution kernel in the lower right-upper left direction under the corresponding convolution kernel size, represents the convolution kernel in the down-up direction corresponding to the convolution kernel size, Represents the convolution kernel in the lower left-upper right direction under the corresponding convolution kernel size; , Specifically, the space-time tensor Convolution with the convolution kernel of each scale generates a tensor of size 8×5×W×H, and the tensors at the four scales are concatenated to obtain the spatiotemporal motion feature tensor. , with a scale of 32×5×W×H, and the space-time tensor Be consistent.
4. The infrared small target multi-frame detection method based on time-space integrated network according to claim 1 is characterized in that: In S2, the acquired infrared image sample data set is composed of a training set and a test set, and the training set and the test set both include multiple groups of data, each group of data includes a three-dimensional matrix Image_mat composed of 5 original images and a Mask_mat composed of 5 corresponding labels; wherein the original image is an infrared image of 5 consecutive frames, and the target position is manually marked according to the infrared image and recorded as 1, and the non-target position pixel is set to 0 to obtain a label; the image and the label are superimposed according to the time dimension to form Image_mat and Mask_mat; The method of using the infrared image sample data set to train the spatiotemporal integrated infrared dim small target detection network model includes: Input the training set sample images into the spatiotemporal integrated infrared dim small target detection network model to obtain the network output results of each set of training sample images; Based on the network output results, construct a loss function; Calculate the loss value according to the loss function; Based on the loss value, determine whether the spatiotemporal integrated infrared dim small target detection network model converges; if so, determine that the network model training is completed, save the network model parameters obtained by the training, otherwise continue training; The loss function is constructed by calculating the focal loss of the corresponding label frame by frame for the probability likelihood graphs of 5 frames of infrared dim targets output by the network, and taking the sum of the focal losses of the 5 frames as the total network loss; it is expressed as follows: , in: Refers to focal loss.
5. The infrared dim small target multi-frame detection method based on time-space integrated network according to claim 1 is characterized in that: In S3, the method of inputting the image sequence to be detected into the trained spatiotemporal integrated infrared small target detection network model to obtain the network output includes: Slide the sequence of input T frames along the time direction to obtain the 5 consecutive frames of images required by the network as input, where the sliding interval is 1 frame, and obtain T-4 image groups, which are recorded as: , The trained spatiotemporal integrated infrared small target detection network model is used to process T-4 image groups, and the network processing results are continuously obtained, which are recorded as: , in: The results include 5 frames of images, namely .
6. The infrared dim small target multi-frame detection method based on time-space integrated network according to claim 1 is characterized in that: In S3, the method of using the maximum value method to perform result-level fusion to obtain the fusion result includes: in, is the pixel location, is the pixel of the Kth frame The probability of being a weak infrared target, For maximum value operation, is the processing result of the network, The 5th frame represents the network processing result of the K-4th image group, is the processing result of the network, The 4th frame represents the network processing result of the K-3th image group, is the processing result of the network, The third frame represents the network processing result of the K-2th image group, is the processing result of the network, The second frame represents the network processing result of the K-1th image group, is the processing result of the network, It represents the first frame of the network processing result of the K-1th image group, W and H represent the width and height of the image respectively, and T represents the number of frames in the sequence.
7. An infrared dim target multi-frame detection system based on a time-space integrated network, the system is used to implement the method described in any one of claims 1 to 6, characterized in that: include: Building module, training module and detection module; The building module is used to build a time-space integrated infrared small target detection network model; The training module is used to obtain an infrared image sample data set and use the infrared image sample data set to train the spatiotemporal integrated infrared dim small target detection network model; The detection module is used to input the image sequence to be detected into the trained spatiotemporal integrated infrared dim small target detection network model to obtain the network output, use the maximum value method to perform result-level fusion to obtain the fusion result, perform threshold segmentation on the fusion result, and obtain the infrared dim small target.
Citation Information
Patent Citations
An infrared weak small target detection method based on a context aggregation network
CN109902715A
Infrared weak and small target detection method based on hybrid spatial modulation feature convolutional neural network
CN116486102A