A Multi-frame Infrared Small Target Super-resolution Method Based on Deformable Convolution

By adopting a deformable convolution method in infrared multi-frame small target super-resolution, multi-scale features are extracted and spatial details are restored through jump connections, the morphological changes caused by target motion changes and energy dispersion are solved, and efficient infrared small target super-resolution processing is achieved.

CN118864269BActive Publication Date: 2025-05-27HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410872983.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-01
Publication Date
2025-05-27
Estimated Expiration
2044-07-01

AI Technical Summary

Technical Problem

In infrared multi-frame small target super resolution, there are morphological changes caused by motion changes and energy dispersion of the target. The existing methods are difficult to effectively deal with these changes, limiting the performance of infrared small target super-segment processing.

Method used

Using a multi-frame infrared small-objective super-segment method based on deformable convolution, a convolutional neural network with a U-shaped structure extracts multi-scale features and restores the spatial details of the image through jump connections. Deformable convolution learns the characteristics of offsets, implicitly aligning the target, autonomously learns the convolution kernel that matches the target morphology, and combines three-dimensional convolution to enhance the target energy, thereby enhancing the significance of weak targets.

Benefits of technology

Effectively handle the timing motion, morphology and energy changes of the target, improve the spatial significance of the dark and weak targets, and improve the super-resolution performance of small infrared targets, so that the super-resolution image and the feature expression of the original image are highly consistent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118864269B_ABST
    Figure CN118864269B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-frame infrared small target super-resolution method based on deformable convolution. The method proposes a DCUNet network suitable for the super-resolution rate of multi-frame infrared small targets, fuses the multi-scale information of multi-frame infrared small target images, and restores the spatial detail information of the targets. A multi-frame alignment TADCM module is proposed to implicitly align targets with complex inter-frame motion states, morphological and energy temporal variations, so as to make full use of inter-frame information for mutual complementation and enhance the spatial saliency of dim targets. A method of using feature supervision to guide deformable convolution learning is proposed, that is, target segmentation results are output at the feature layers of the last two layers of the network, and the ground truth of target segmentation with pixel-level labels is used as supervision to constrain the encoding and upsampling processes, improve the accuracy of deformable convolution, and enable the edges and morphology of the targets to be fully restored.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and relates to a multi-frame infrared small target super-resolution method, in particular to a multi-frame infrared small target super-resolution method based on deformable convolution. Background Art

[0002] Infrared imaging has the characteristics of all-weather operation and being less affected by the environment, and plays an important role in the field of remote sensing. Its applications include infrared guidance, scene monitoring, target tracking, etc. However, limited by long-distance imaging and complex scenes, the targets in the infrared imaging results always show weak and small characteristics, lacking texture and detail information, and the targets are not significant in a single-frame image. Therefore, tasks such as the detection and segmentation of single-frame infrared small targets are difficult to be practically applied. Multi-frame images increase the time dimension compared with single-frame images, and the target information contained extends from the spatial domain to the spatio-temporal domain, so the information they possess is more abundant. In practical applications, for example, the image segmentation task in the autonomous driving scenario pays more attention to the spatial details of the target. Therefore, how to perform super-resolution on multi-frame infrared small target images to enrich the spatial scale information of infrared small targets is a key research issue in the field of infrared imaging.

[0003] Infrared multi-frame small target super-resolution aims to use multi-frame temporal information to achieve the resolution improvement of the target. There are the following difficulties in practical applications: (1) Small target: The general scale of infrared small targets is less than 9×9, lacking spatial details and internal texture information of the target. And affected by the infrared thermal radiation imaging mechanism, there is diffusion at the target edge, and the coupling with the environmental background makes the target contour information not obvious. (2) Weak energy: Long-distance imaging results in low target energy received on the image plane. Therefore, small targets in infrared images often show the characteristics of low signal-to-noise ratio. And there are complex backgrounds such as high-reflection clouds and surface high-temperature radiation sources in the infrared radiation scene, further causing the target to be submerged in the background. (3) Fast change: Infrared small targets move temporally in the imaging scene, and the movement trajectory is difficult to predict. Their relative position with the detection platform is constantly changing, resulting in temporal changes in the imaging position, imaging scale, and detection energy, bringing difficulties and challenges to the mutual complementation of multi-frame image information.

[0004] In recent years, researchers have proposed many super-resolution algorithms. For example, using the local motion and contrast prior of the target to achieve the super-resolution of infrared small targets; using the Bayesian method to estimate the basic motion of the target, the blur kernel, and the noise level, and reconstructing the high-resolution video sequence. However, the above methods are difficult to handle the situation of inter-frame motion changes of the target and the changes in the target's own morphology, restricting the performance of infrared small target super-resolution processing. Summary of the Invention

[0005] Aiming at the problem of morphological changes caused by target motion changes and energy dispersion in infrared multi-frame small target super-resolution, the present invention provides a multi-frame infrared small target super-resolution method based on deformable convolution. The present invention uses a convolutional neural network with a U-shaped structure as the basic framework to extract multi-scale features of infrared multi-frame images, and restores the spatial details of the images during the decoding process through skip connections. Utilizing the characteristic that deformable convolution can learn offsets, the target of the reference frame is implicitly aligned to the current frame to cope with the complex and variable target motion states. At the same time, taking advantage of the fact that deformable convolution can autonomously learn the convolutional kernel that matches the target morphology, it deals with the temporal and spatial deformation of the target. In addition, three-dimensional convolution with a convolutional kernel size of 1×1×5 is used to achieve the energy accumulation of the target to enhance the spatial saliency of weak targets, ensuring that more infrared small target details can be restored in the image super-resolution result of the current frame. To guide the learning of deformable convolution, the present invention outputs the target segmentation result at the last two feature layers of the network, and uses the target segmentation ground truth with pixel-level labels as supervision to improve the learning accuracy of deformable convolution.

[0006] The object of the present invention is achieved through the following technical solutions:

[0007] A multi-frame infrared small target super-resolution method based on deformable convolution, comprising the following steps:

[0008] Step 1: Construct a DCUNet network improved based on UNet

[0009] The DCUNet network framework consists of an encoding structure, skip connections, a decoding structure, a super-resolution head, and a target segmentation head with feature supervision;

[0010] The encoding structure uses a UNet network, including six layers of encoding structures, where:

[0011] The first layer of the encoding structure: First, an image with a dimension of k×c×h×w passes through a two-dimensional convolutional layer with a convolutional kernel size of 7×7, then through a batch normalization layer and an activation layer, and outputs a feature map with a dimension of k×64×h×w. Finally, the feature map is processed by the TADCM module, and the encoded feature is concatenated with the upsampled decoded feature through skip connections;

[0012] The second layer of the encoding structure: Perform max pooling on the feature map with a dimension of k×64×h×w output by the first layer of the encoding structure to obtain a dimension of The downsampled features are used as the input of the second-layer encoding structure. Three cascaded bottleneck structures are utilized to mine the high-level semantic information representing the global features of the image. Finally, max-pooling downsampling and multi-frame feature enhancement are performed on the feature map output by the bottleneck structure. The output features of max-pooling downsampling are input into the third-layer encoding structure. For multi-frame feature enhancement, a two-dimensional convolutional layer with a convolution kernel of 3×3 is used to double the number of output channels, and then the TADCM module is used to aggregate the multi-frame features into a single-frame output;

[0013] The third-layer encoding structure: Max-pooling is performed on the feature map with a dimension of to obtain downsampled features with a dimension of The downsampled features are used as the input of the third-layer encoding structure. Three cascaded bottleneck structures are utilized to mine the high-level semantic information representing the global features of the image. Finally, max-pooling downsampling and multi-frame feature enhancement are performed on the feature map output by the bottleneck structure. The output features of max-pooling downsampling are input into the fourth-layer encoding structure. For multi-frame feature enhancement, a two-dimensional convolutional layer with a convolution kernel of 3×3 is used to double the number of output channels, and then the TADCM module is used to aggregate the multi-frame features into a single-frame output;

[0014] The fourth-layer encoding structure: Max-pooling is performed on the feature map with a dimension of to obtain downsampled features with a dimension of The downsampled features are used as the input of the fourth-layer encoding structure. Three cascaded bottleneck structures are utilized to mine the high-level semantic information representing the global features of the image. Finally, max-pooling downsampling and multi-frame feature enhancement are performed on the feature map output by the bottleneck structure. The output features of max-pooling downsampling are input into the fifth-layer encoding structure. For multi-frame feature enhancement, a two-dimensional convolutional layer with a convolution kernel of 3×3 is used to double the number of output channels, and then the TADCM module is used to aggregate the multi-frame features into a single-frame output;

[0015] The fifth-layer encoding structure: Max-pooling is performed on the feature map with a dimension of to obtain downsampled features with a dimension of The downsampled features are used as the input of the fifth-layer encoding structure. Two cascaded bottleneck structures are utilized to mine the high-level semantic information representing the global features of the image. A residual connection structure is added to each bottleneck structure. Finally, max-pooling downsampling and multi-frame feature enhancement are performed on the feature map output by the bottleneck structure. The output features of max-pooling downsampling are input into the sixth-layer encoding structure. For multi-frame feature enhancement, a two-dimensional convolutional layer with a convolution kernel of 3×3 is used to double the number of output channels, and then the TADCM module is used to aggregate the multi-frame features into a single-frame output;

[0016] The sixth-layer encoding structure: With a dimension of The feature map is input into the sixth-layer encoding structure. Through two 2D convolutional layers with a convolutional kernel of 3×3 and an activation function layer, deep features with a dimension of are obtained. Finally, multi-frame features are aggregated through the TADCM module;

[0017] The skip connection concatenates the encoded features of the same layer with the upsampled decoded features to enhance the target details in the decoding process, that is: perform skip connection on the encoded features of the i-th layer and the decoded features of the i+1-th layer. In the present invention, i = [1, 2, 3, 4, 5]. The features after skip connection are expressed as follows:

[0018]

[0019] In the formula, represents the encoded features of the i-th layer, represents the decoded features of the i+1-th layer, cat represents the channel dimension concatenation operation, and SP is the subpixel convolution algorithm;

[0020] The decoding structure consists of a batch normalization layer, a 2D convolutional layer with a convolutional kernel of 3×3, an activation function layer, and subpixel convolution upsampling. Concatenate the decoded features output by the i+1-th layer decoding structure and the features of the skip connection of the i-th layer encoding structure as the input of the i-th layer decoding structure. The number of channels of the i+1-th layer decoding structure is twice that of the i-th layer decoding structure;

[0021] Add a target segmentation head to the last two layers of the decoding structure, and the target segmentation heads in the two layers share weights;

[0022] The target segmentation head structure consists of a 2D convolutional layer with a convolutional kernel of 3×3, a batch normalization layer, and a 2D convolutional layer with a convolutional kernel of 1×1. Segment the intermediate features output by the last two layers of the decoding structure, and then calculate the DCUNet loss with the original high-resolution image and the target segmentation ground truth respectively, and then optimize the network weights;

[0023] The super-resolution head structure consists of a 2D convolutional layer with a convolutional kernel of 3×3, an activation function layer, and a 2D convolutional layer with a convolutional kernel of 1×1. Receive the output of the last layer of the decoding structure and output a super-resolution image, that is: map the feature map to the finally output high-resolution image;

[0024] Step 2: In the training stage, first load the original high-resolution multi-frame infrared small target images, downsample the images to generate low-resolution images, superimpose Gaussian random blur on the downsampled images, and finally perform upsampling by bicubic interpolation as the input of the DCUNet network; in the inference stage, directly perform upsampling on the original images by bicubic interpolation and use them as the input of the trained DCUNet network;

[0025] Step 3: Use the DCUNet network to fuse the multi-scale information of multi-frame infrared small target images and restore the spatial detail information of the target, and finally output a high-resolution image.

[0026] Compared with the prior art, the present invention has the following advantages:

[0027] 1. A DCUNet network applicable to multi-frame infrared small target super-resolution is proposed to fuse the multi-scale information of multi-frame infrared small target images and restore the spatial detail information of the target. Different from the traditional interpolation method that can only calculate the extra pixels after super-resolution through neighboring pixels, the DCUNet network can fully explore the internal feature correlation between each pixel of the image, effectively mine the high-level semantic features of the image while retaining the image details, making the feature expression of the super-resolution image highly consistent with that of the original image, thereby improving the performance of super-resolution.

[0028] 2. A deformable convolutional module for multi-frame alignment, the Temporal Alignment Deformable Convolutional Module (TADCM) module, is proposed to implicitly align the targets with complex inter-frame motion states, morphological and energy temporal changes, so as to make full use of the inter-frame information for mutual complementation and improve the spatial saliency of dim targets.

[0029] 3. A method of using feature supervision to guide the learning of deformable convolution is proposed, that is, the target segmentation results are output at the feature layers of the last two layers of the network, and the ground truth of the target segmentation with pixel-level labels is used as supervision to constrain the encoding and upsampling processes, improve the accuracy of the deformable convolution, and enable the edges and morphology of the target to be fully restored. Description of the Drawings

[0030] Figure 1 It is a schematic diagram of the overall architecture of the DCUNet network;

[0031] Figure 2 It is a schematic diagram of the TADCM module;

[0032] Figure 3 It is a schematic diagram of Subpixel convolution;

[0033] Figure 4 It is a diagram of experimental results. Detailed Embodiments

[0034] The technical solutions of the present invention will be further described below in conjunction with the drawings, but are not limited thereto. Any modifications or equivalent replacements of the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention shall be covered by the protection scope of the present invention.

[0035] The present invention provides a multi-frame infrared small target super-resolution method based on deformable convolution. The method designs a DCUNet network suitable for multi-frame infrared small target super-resolution, and proposes a Temporal Alignment Deformable Convolutional Module (TADCM) for aligning multi-frame features to solve the changes in the temporal motion, morphology, and energy of the target, enhance the spatial saliency of dim targets, and finally proposes a feature supervision method, that is, outputting the target segmentation result at the last two feature layers of the network, and using the target segmentation ground truth with pixel-level labels as supervision to improve the accuracy of deformable convolution. The specific steps are as follows:

[0036] Step 1: Construct a DCUNet network improved based on UNet

[0037] The DCUNet network framework consists of an encoding structure, skip connections, a decoding structure, a super-resolution head, and a feature-supervised target segmentation head;

[0038] The encoding structure contains a residual connection structure, a bottleneck structure, and a TADCM module. The encoding structure uses a UNet network, including six layers of encoding structure, where:

[0039] The first layer of the encoding structure: First, an image with dimensions of k×c×h×w passes through a two-dimensional convolutional layer with a convolutional kernel size of 7×7, then through a batch normalization layer and an activation layer, and outputs a feature map with dimensions of k×64×h×w. Finally, the feature map is processed by the TADCM module, and the encoded feature is concatenated with the upsampled decoded feature through skip connections.

[0040] As Figure 2 shown, the TADCM module consists of a many-to-many module and a many-to-one module. Figure 2 (a), the many-to-many module first uses a deformable convolutional layer (DCN) to align the input multi-frame features to each frame one by one, and then uses a three-dimensional convolutional layer with a convolutional kernel of 1×1×5 to enhance the target features of each frame. The formula is as follows:

[0041]

[0042] In the formula, K represents the maximum number of frames of the input multi-frame image, is the k-th frame feature in the input feature, refers to the multi-frame features aligned to the k-th frame, is the output result of the many-to-many module, DCN is the deformable convolution, conv is the 3D convolution, and cat is the channel dimension concatenation operation.

[0043] As Figure 2As shown in (b), the many-to-one module first uses a deformable convolutional layer (DCN) to extract features of other frames, and then uses a three-dimensional convolutional layer with a convolutional kernel of 1×1×5 to aggregate the reference frame features into the current frame, as the output of the TADCM module, to achieve full utilization of multi-frame information. The formula is as follows:

[0044]

[0045] In the formula, K represents the maximum number of frames of the input multi-frame images, is the feature of the k-th frame output by the many-to-many module, is the aligned multi-frame features, is the single-frame feature finally enhanced by the many-to-one module. DCN is the deformable convolution, conv is the 3D convolution, and cat is the channel dimension concatenation operation.

[0046] The second-layer encoding structure: Perform max pooling on the feature map with a dimension of k×64×h×w output by the first-layer encoding structure to obtain a downsampled feature with a dimension of . Use this downsampled feature as the input of the second-layer encoding structure, and utilize three cascaded bottleneck structures to mine the high-level semantic information representing the global features of the image. Finally, perform max pooling downsampling and multi-frame feature enhancement on the feature map output by the bottleneck structure. The output feature of the max pooling downsampling is input to the third-layer encoding structure. For multi-frame feature enhancement, use a two-dimensional convolutional layer with a convolutional kernel of 3×3 to double the number of output channels, and then use the TADCM module to aggregate the multi-frame features into one-frame output.

[0047] The bottleneck structure consists of a two-dimensional convolutional layer, a batch normalization layer, and an activation function. At the same time, to alleviate the problems of gradient disappearance and model degradation, the present invention adds a residual connection structure from before the first bottleneck structure to the output position of the third bottleneck structure, enabling the model to skip the three cascaded bottleneck structures, which are represented by black lines in the first three-layer encoding structures in the figure, specifically as Figure 1 shown.

[0048] The third-layer encoding structure: The third-layer encoding structure is basically the same as the second-layer encoding structure, except that the number of input and output channels of the third-layer encoding structure is twice that of the second-layer encoding structure, and the feature map resolution is one-fourth of the upper-layer network, that is: perform max pooling on the feature map with a dimension of output by the second-layer encoding structure to obtain a feature map with a dimension of The downsampled features are used as the input of the third-layer encoding structure. Three cascaded bottleneck structures are utilized to mine the high-level semantic information representing the global features of the image. Finally, max-pooling downsampling and multi-frame feature enhancement are performed on the feature map output by the bottleneck structure. The output features of max-pooling downsampling are input into the fourth-layer encoding structure. For multi-frame feature enhancement, a two-dimensional convolutional layer with a 3×3 convolutional kernel is used to double the number of output channels, and then the TADCM module is used to aggregate the multi-frame features into a single-frame output. For details of the network structure, see Figure 1 。

[0049] Fourth-layer encoding structure: The fourth-layer encoding structure is basically the same as the third-layer encoding structure, except that the number of input and output channels of the fourth-layer encoding structure is twice that of the third-layer encoding structure, and the feature map resolution is one-fourth of the upper-layer network, that is: max-pooling is performed on the feature map with a dimension of to obtain downsampled features with a dimension of The downsampled features are used as the input of the fourth-layer encoding structure. Three cascaded bottleneck structures are utilized to mine the high-level semantic information representing the global features of the image. Finally, max-pooling downsampling and multi-frame feature enhancement are performed on the feature map output by the bottleneck structure. The output features of max-pooling downsampling are input into the fifth-layer encoding structure. For multi-frame feature enhancement, a two-dimensional convolutional layer with a 3×3 convolutional kernel is used to double the number of output channels, and then the TADCM module is used to aggregate the multi-frame features into a single-frame output. For details of the network structure, see Figure 1 。

[0050] Fifth-layer encoding structure: In the fifth-layer encoding structure, the bottleneck structure is reduced to two, and a residual connection structure is added to each bottleneck structure, that is: max-pooling is performed on the feature map with a dimension of to obtain downsampled features with a dimension of The downsampled features are used as the input of the fifth-layer encoding structure. Two cascaded bottleneck structures are utilized to mine the high-level semantic information representing the global features of the image. A residual connection structure is added to each bottleneck structure. Finally, max-pooling downsampling and multi-frame feature enhancement are performed on the feature map output by the bottleneck structure. The output features of max-pooling downsampling are input into the sixth-layer encoding structure. For multi-frame feature enhancement, a two-dimensional convolutional layer with a 3×3 convolutional kernel is used to double the number of output channels, and then the TADCM module is used to aggregate the multi-frame features into a single-frame output. For details of the network structure, see Figure 1 。

[0051] Sixth-layer encoding structure: The feature map with a dimension of is input into the sixth-layer encoding structure. Through two two-dimensional convolutional layers with 3×3 convolutional kernels and an activation function layer, deep features with a dimension of are obtained. Finally, the multi-frame features are aggregated through the TADCM module.

[0052] The skip connection concatenates the encoded features of the same layer with the upsampled decoded features to enhance the object details during the decoding process, that is: a skip connection is performed on the encoded features of the i-th layer and the decoded features of the (i + 1)-th layer. In the present invention, i = [1, 2, 3, 4, 5]. The features after the skip connection can be expressed as follows:

[0053]

[0054] In the formula, represents the encoded features of the i-th layer, represents the decoded features of the (i + 1)-th layer. cat represents the channel dimension concatenation operation, and SP is the subpixel convolution algorithm. The schematic diagram of the subpixel convolution algorithm is shown in Figure 3 By rearranging the values of 4 channels, an upsampled image with 2 times the resolution is formed.

[0055] The decoding structure consists of a batch normalization layer, a two-dimensional convolutional layer with a convolution kernel of 3×3, an activation function layer, and subpixel convolution upsampling. The decoded features output by the (i + 1)-th decoding structure are concatenated with the features of the skip connection of the i-th encoding structure as the input of the i-th decoding structure. The number of channels of the (i + 1)-th decoding structure is 2 times that of the i-th decoding structure.

[0056] An object segmentation head is added to the last two layers of the decoding structure. The object segmentation heads in the two layers share weights. The object segmentation head consists of a two-dimensional convolutional layer with a convolution kernel size of 3×3, a batch normalization layer, and a two-dimensional convolutional layer with a convolution kernel size of 1×1 that outputs 1 channel. DCUNet uses real annotations with pixel-level labels to supervise the results of object segmentation, avoiding the deformable convolution from learning clutter structures similar to the object background.

[0057] The final output part of the DCUNet consists of a super-resolution head and a feature-supervised object segmentation head, where:

[0058] The super-resolution head is responsible for outputting a super-resolution image, that is: mapping the feature map to the final output high-resolution image. The super-resolution head is stacked by two two-dimensional convolutional layers with a convolution kernel size of 3×3 and an activation layer. Since the infrared image is a single-channel grayscale image, a single-channel image is finally output through a convolutional layer with a convolution kernel size of 1×1.

[0059] The object segmentation head is responsible for outputting the intermediate feature object segmentation result, and then calculating the DCUNet loss with the original high-resolution image and the object segmentation ground truth respectively, thereby optimizing the network weights.

[0060] Total loss function of DCUNet It consists of two parts: super-resolution reconstruction loss and target segmentation loss, where:

[0061] Super-resolution reconstruction loss It is calculated using pixel-wise mean squared error (MSE), and the expression is as follows:

[0062]

[0063] In the formula, K represents the maximum number of frames of the input multi-frame images, X and Y respectively represent the length and width of the output super-resolution image, represents the original high-resolution image of the k-th frame, represents the low-resolution image after downsampling the original k-th frame image, and x and y respectively represent the row and column coordinates of the pixel points in the image.

[0064] Target segmentation loss and Use soft-IOU loss to balance positive and negative samples, where is the target segmentation loss of the low resolution, is the target segmentation loss of the high resolution, and the expression is as follows:

[0065]

[0066]

[0067] In the formula, (x, y) are the pixel coordinates of the image, C l and C h respectively represent the confidence maps of the low resolution and the high resolution, G l and G h respectively represent the target segmentation labels of the low resolution and the high resolution.

[0068] Use the weight λ to balance the super-resolution loss and the target segmentation loss, and the total loss function is:

[0069]

[0070] In the formula, represents the total loss function of DCUNet, λ represents the loss adjustment factor, and in this invention, λ = 0.5.

[0071] Step 2: Training phase. First, load the original high-resolution multi-frame infrared small target images, downsample the images to generate low-resolution images, add Gaussian random blur to the downsampled images, and finally perform upsampling using bicubic interpolation as the input to the DCUNet network. In the inference phase, directly upsample the original low-resolution images using bicubic interpolation and then use them as the input to the trained DCUNet network. As Figure 1 shown, the specific steps are as follows:

[0072] Step 2-1: In the training phase, first load an original high-resolution multi-frame infrared small target image with dimensions of k×c×h×w, where k represents the number of image frames, c represents the number of image channels, for infrared images c = 1, h represents the image height, and w represents the image width.

[0073] Step 2-2: Downsample the original high-resolution image with dimensions of k×c×h×w to obtain a low-resolution image with halved width and height and dimensions of .

[0074] Step 2-3: Add Gaussian random blur to the low-resolution image with dimensions of to obtain the paired low-resolution and high-resolution images required for DCUNet network training. At the same time, to match the input size of the DCUNet network, perform bicubic interpolation upsampling on the low-resolution image with dimensions of to obtain an image with dimensions of k×c×h×w with random blur as the input to the DCUNet network.

[0075] Step 2-4: In the inference phase, perform bicubic interpolation upsampling on the low-resolution image with dimensions of to obtain an image with dimensions of k×c×h×w as the input to the trained DCUNet network.

[0076] Step 3: Use the DCUNet network to fuse the multi-scale information of multi-frame infrared small target images and restore the spatial detail information of the targets, and finally output high-resolution images.

[0077] Step 4: Use two public datasets to conduct multi-scenario test verification on the algorithm.

[0078] The present invention uses the publicly available IRDST and NUDT-MIRSDT datasets to conduct experiments. These two datasets are divided into training set, validation set, and test set in a ratio of 5:2:3. In all experiments, the batch size is set to 4, the number of frames of multi-frame infrared small target images is set to 5, the number of optimization rounds is set to 30, and the initial learning rate is set to 1×10 -6 , and the initial learning rate is decreased by 10 times at the 10th and 20th rounds respectively. Finally, the Adagrad optimizer is used to train the network. The experimental results are asFigure 4 As shown, it can be observed that the performance of the super-resolution reconstruction of the present invention is superior to that of bilinear interpolation and bicubic interpolation, and is closest to the original image, and can well adapt to the changes of the target and the background.

Claims

1. A multi-frame infrared small target super-resolution method based on deformable convolution, characterized by The method comprises the following steps: Step 1: Build a DCUNet network based on UNet improvement The DCUNet network framework consists of an encoding structure, skip connections, a decoding structure, a super-resolution head, and a feature-supervised target segmentation head; The coding structure adopts the UNet network, including a six-layer coding structure, wherein: The first layer encoding structure: First, the image with a dimension of k×c×h×w is passed through a two-dimensional convolution layer with a convolution kernel size of 7×7, and then through a batch normalization layer and an activation layer, and the feature map with a dimension of k×64×h×w is output. Finally, the feature map is processed by the TADCM module, and the encoded features are concatenated with the upsampled decoded features through jump connections; The second layer encoding structure: The feature map with the dimension of k×64×h×w output by the first layer encoding structure is max-pooled to obtain a dimension of The down-sampled features are used as the input of the second-layer encoding structure. The high-level semantic information representing the global features of the image is mined using three cascaded bottleneck structures. Finally, the feature map output by the bottleneck structure is subjected to maximum pooling down-sampling and multi-frame feature enhancement. The output features of the maximum pooling down-sampling are input to the third-layer encoding structure. The multi-frame feature enhancement uses a two-dimensional convolutional layer with a convolution kernel of 3×3 to double the number of output channels, and then uses the TADCM module to aggregate the multi-frame features into one frame output. The third-layer encoding structure: The dimension of the output of the second-layer encoding structure is The feature map of is max-pooled, and the dimension is The down-sampled features are used as the input of the third-layer encoding structure. The high-level semantic information representing the global features of the image is mined using three cascaded bottleneck structures. Finally, the feature map output by the bottleneck structure is subjected to maximum pooling down-sampling and multi-frame feature enhancement. The output features of the maximum pooling down-sampling are input to the fourth-layer encoding structure. The multi-frame feature enhancement uses a two-dimensional convolutional layer with a convolution kernel of 3×3 to double the number of output channels, and then uses the TADCM module to aggregate the multi-frame features into one frame output. The fourth layer encoding structure: The dimension of the output of the third layer encoding structure is The feature map of is max-pooled, and the dimension is The down-sampled features are used as the input of the fourth-layer encoding structure. The high-level semantic information representing the global features of the image is mined using three cascaded bottleneck structures. Finally, the feature map output by the bottleneck structure is subjected to maximum pooling down-sampling and multi-frame feature enhancement. The output features of the maximum pooling down-sampling are input to the fifth-layer encoding structure. The multi-frame feature enhancement uses a two-dimensional convolutional layer with a convolution kernel of 3×3 to double the number of output channels, and then uses the TADCM module to aggregate the multi-frame features into one frame output. The fifth layer encoding structure: The dimension of the output of the fourth layer encoding structure is The feature map of is max-pooled, and the dimension is The down-sampled features are used as the input of the fifth-layer encoding structure. The high-level semantic information representing the global features of the image is mined using two cascaded bottleneck structures. Finally, the feature map output by the bottleneck structure is subjected to maximum pooling down-sampling and multi-frame feature enhancement. The output features of the maximum pooling down-sampling are input to the sixth-layer encoding structure. The multi-frame feature enhancement uses a two-dimensional convolutional layer with a convolution kernel of 3×3 to double the number of output channels, and then uses the TADCM module to aggregate the multi-frame features into one frame output. The sixth level encoding structure: the dimension is The feature map is input into the sixth encoding structure, and through two two-dimensional convolution layers with a convolution kernel of 3×3 and an activation function layer, the dimension is The deep features of the image are finally aggregated through the TADCM module; The jump connection concatenates the encoding features of the same layer with the upsampled decoding features to enhance the target details in the decoding process, that is, the i-th layer encoding features and the i+1-th layer decoding features are jump-connected, i = [1, 2, 3, 4, 5], and the features after the jump connection are The statement is as follows: In the formula, represents the i-th layer encoding feature, represents the i+1th layer decoding feature, cat represents the channel dimension splicing operation, and SP is the subpixel convolution algorithm; The decoding structure concatenates the decoding features output by the i+1th layer decoding structure and the features of the jump connection of the ith layer encoding structure as the input of the ith layer decoding structure. The number of channels of the i+1th layer decoding structure is twice that of the ith layer decoding structure. The target segmentation head structure segments the intermediate features output by the last two layers of decoding structures, and then calculates the DCUNet loss with the original high-resolution image and the target segmentation truth value, thereby optimizing the network weights; The super-resolution head structure receives the output of the last layer of the decoding structure and outputs a super-resolution image, that is, mapping the feature map to a high-resolution image that is finally output; Step 2: In the training phase, first load the original high-resolution multi-frame infrared small target image, downsample the image to generate a low-resolution image, superimpose Gaussian random blur on the downsampled image, and finally upsample it using bicubic interpolation as the input of the DCUNet network; in the inference phase, directly upsample the original image using bicubic interpolation as the input of the trained DCUNet network; Step 3: Use the DCUNet network to fuse the multi-scale information of multiple frames of infrared small target images, restore the spatial detail information of the target, and finally output a high-resolution image.

2. The multi-frame infrared small target super-resolution method based on deformable convolution according to claim 1 is characterized in that The TADCM module consists of a many-to-many module and a many-to-one module, wherein: The many-to-many module first uses a deformable convolution layer to align the input multi-frame features to each frame one by one, and then uses a three-dimensional convolution layer with a convolution kernel of 1×1×5 to enhance the target features of each frame. The formula is as follows: In the formula, K represents the maximum number of frames of the input multi-frame image. is the k-th frame feature in the input feature, refers to the multi-frame features aligned to the kth frame, It is the output result of many-to-many module, DCN is deformable convolution, conv is 3D convolution, and cat is channel dimension concatenation operation; The many-to-one module first uses a deformable convolution layer to extract features from other frames, and then uses a 3D convolution layer with a convolution kernel of 1×1×5 to aggregate the reference frame features to the current frame as the output of the TADCM module, making full use of multi-frame information. The formula is as follows: In the formula, K represents the maximum number of frames of the input multi-frame image. is the k-th frame feature output by the many-to-many module, is the aligned multi-frame feature, It is the single-frame feature after the final enhancement of the many-to-one module, DCN is deformable convolution, conv is 3D convolution, and cat is the channel-dimensional splicing operation.

3. The multi-frame infrared small target super-resolution method based on deformable convolution according to claim 1 is characterized in that The bottleneck structure consists of a two-dimensional convolutional layer, a batch normalization layer, and an activation function In the three cascaded bottleneck structures, a residual connection structure is added from the position before the first bottleneck structure to the output of the third bottleneck structure; in the two cascaded bottleneck structures, a residual connection structure is added to each bottleneck structure.

4. The multi-frame infrared small target super-resolution method based on deformable convolution according to claim 1 is characterized in that The target segmentation head structure consists of a two-dimensional convolution layer with a convolution kernel of 3×3, a batch normalization layer, and a two-dimensional convolution layer with a convolution kernel of 1×1.

5. The multi-frame infrared small target super-resolution method based on deformable convolution according to claim 1 is characterized in that The super-resolution head structure consists of a two-dimensional convolution layer with a convolution kernel of 3×3, an activation function layer, and a two-dimensional convolution layer with a convolution kernel of 1×1.

6. The multi-frame infrared small target super-resolution method based on deformable convolution according to claim 1 is characterized in that The total loss function of DCUNet It consists of two parts: super-resolution reconstruction loss and target segmentation loss, where: Super-resolution reconstruction loss MSE is used for calculation, and the expression is as follows: Where K represents the maximum number of frames of the input multi-frame image, X and Y represent the length and width of the output super-resolution image, respectively. represents the original k-th high-resolution image, represents the low-resolution image after the original k-th frame image is downsampled, and x and y represent the row and column coordinates of the pixel points in the image respectively; Object segmentation loss and Soft-IOU loss is used to balance positive and negative samples, where is the low-resolution object segmentation loss, is the high-resolution object segmentation loss, expressed as follows: Where (x, y) is the pixel coordinate of the image, C l and C h Distinguish between low-resolution and high-resolution confidence maps, G l and G h Represent low-resolution and high-resolution object segmentation labels respectively; The weight λ is used to balance the super-resolution loss and the target segmentation loss. The total loss function is: In the formula, represents the total loss function of DCUNet, and λ represents the loss adjustment factor.

7. The multi-frame infrared small target super-resolution method based on deformable convolution according to claim 1 is characterized in that The specific steps of step 2 are as follows: Step 21: In the training phase, first load an original high-resolution multi-frame infrared small target image with dimensions of k×c×h×w, where k represents the number of image frames, c represents the number of image channels, c=1 for infrared images, h represents the image height, and w represents the image width; Step 22: Downsample the original high-resolution image with dimensions k×c×h×w to obtain a half-dimension image with dimensions k×c×h×w. Low-resolution images; Step 23: Dimension Gaussian random blur is superimposed on the low-resolution image of , and the paired low-resolution image and high-resolution image required for DCUNet network training are obtained; at the same time, in order to match the input size of the DCUNet network, the dimension is The low-resolution image is upsampled by bicubic interpolation to obtain an image with random blur of dimension k×c×h×w as the input of DCUNet network; Step 24: Inference phase, for dimensions The low-resolution image is upsampled by bicubic interpolation to obtain an image with a dimension of k×c×h×w as the input of the trained DCUNet network.

Citation Information

Patent Citations

  • Infrared small target detection method based on space-time stability

    CN113034533A

  • Multi-frame infrared image super-resolution method and system based on feature cycle fusion

    CN113538229A