A Pixel-Level Infrared Image Stitching Method Based on Deep Learning

Through the pixel-level infrared image stitching method based on deep learning, using pixel-level deformation and encoder-decoder structure, ghosting, artifacts and parallax problems in infrared image stitching in the prior art are solved, and a more natural and realistic stitching effect is achieved.

CN116152055BActive Publication Date: 2025-06-27CHANGCHUN UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211390108.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-08
Publication Date
2025-06-27
Estimated Expiration
2042-11-08

AI Technical Summary

Technical Problem

Existing infrared image stitching methods are prone to ghosting and artifacts when dealing with large parallax and complex scenes, and the stitching effect is unnatural and cannot completely solve the parallax problem.

Method used

The pixel-level infrared image stitching method based on deep learning is adopted to perform pixel-level transformation through the pixel-level deformation module, and a stitching image with high naturalness is generated using the encoder-decoder structure to eliminate ghosting and artifacts.

Benefits of technology

It effectively solves the problem of large parallax, enhances the naturalness and authenticity of infrared stitching images, eliminates ghosting and artifacts, and improves the stitching effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116152055B_ABST
    Figure CN116152055B_ABST
Patent Text Reader

Abstract

A pixel-level infrared image stitching method based on deep learning, belonging to the technical field of image stitching. To solve the problems of ghosting and artifacts in the images obtained by existing infrared image stitching methods, Step 1: Construct a network model; Step 2: Prepare a dataset: Select the KAIST dataset, adjust the size of each image in the dataset, and fix the size of the input image; Step 3: Train the network model: Input the dataset prepared in Step 2 into the network model constructed in Step 1 for training; Step 4: Select the minimized loss function and the optimal evaluation metric; Step 5: Fine-tune the model: Use the LTIR dataset to train and fine-tune the model to obtain stable and usable model parameters, ultimately making the model have a better effect on infrared image stitching; Step 6: Save the model: Solidify the finally determined model parameters. When infrared image stitching is required, directly input the image into the network to obtain the stitched image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image stitching, and particularly relates to a pixel-level infrared image stitching method based on deep learning. Background Art

[0002] Infrared image stitching has been widely used in different fields such as biology, medicine, surveillance videos, autonomous driving, virtual reality, etc. Existing methods use an estimated deformation function to deform the feature points in the overlapping region of two infrared images, and homography is the most commonly used deformation function. However, when the camera motion model includes not only displacement but also rotation and scaling degrees of freedom, especially when the distance between the captured scene and the camera is very close, for the surfaces at different depths or planes in different directions at the same depth of the captured scene, the method based on global or local homography estimation cannot completely solve this problem and is affected by parallax. In these cases, the "parallax problem" will occur, resulting in poor stitching effects, such as ghosting and artifacts in the stitched image.

[0003] The Chinese patent publication number is "CN109886878A", and the name is "An infrared image stitching method based on coarse-to-fine registration". This method first obtains the infrared images to be stitched and obtains the adjacent relationships between all images based on the overlapping region; then, calculates the homography matrix between every two adjacent images based on the feature point pairs between every two adjacent images; then, selects an infrared image as the reference image, calculates the homography matrix of each other infrared image relative to the reference image based on the homography matrix between adjacent images, and then calculates the coordinates of each infrared image in the reference image coordinate system to obtain the coarsely registered infrared image; finally, taking the reference image as the initial fine registration image sequence, calculates the homography matrix of each coarsely and finely registered infrared image corresponding to the current fine registration image sequence in turn. The stitched image obtained by this method has problems of ghosting and artifacts and does not conform to the human visual effect. Summary of the Invention

[0004] In order to solve the problems of ghosting and artifacts in the images obtained by the existing infrared image stitching methods, the present invention provides a pixel-level infrared image stitching method based on deep learning. Applying pixel-level deformation in the image overlapping region instead of the traditional homography transformation, and using pixel-level deformation to handle the large parallax problem. Designing an encoder-decoder structure for image stitching, the obtained infrared stitched image is more natural, and at the same time, ghosting and artifacts are eliminated, thereby improving the image stitching effect.

[0005] The solution of the present invention to solve the technical problem is:

[0006] A pixel-level infrared image stitching method based on deep learning, which includes the following steps:

[0007] Step 1, construct the network model: The entire network consists of two parts, a pixel-level deformation module and a stitched image generation module. The pixel-level deformation module uses an optical flow estimation model to obtain the pixel-level deformation from the target image to the reference image, and uses the obtained pixel-level deformation to reposition each pixel point in the target image to obtain the deformed target image. The stitched image generation module stitches the deformed target image and the reference image to obtain the infrared stitched image.

[0008] Step 2, prepare the dataset: Select the KAIST dataset, adjust the size of each image in the dataset, and fix the size of the input image.

[0009] Step 3, train the network model: Input the dataset prepared in Step 2 into the network model constructed in Step 1 for training.

[0010] Step 4, select the loss function to be minimized and the optimal evaluation metric: By minimizing the loss function between the network output image and the label until the number of training times reaches the set threshold or the value of the loss function reaches the set range, it can be considered that the model parameters have been pre-trained, and save the model parameters. At the same time, select the optimal evaluation metric to measure the accuracy of the algorithm and evaluate the performance of the system.

[0011] Step 5, fine-tune the model: Use the LTIR dataset to train and fine-tune the model to obtain stable and usable model parameters, and finally make the model perform better in infrared image stitching.

[0012] Step 6, save the model: Fix the finally determined model parameters. When infrared image stitching operations need to be performed later, directly input the image into the network to obtain the final stitched image.

[0013] In Step 1, the pixel-level deformation module first estimates the pixel-level transformation from the target image to the reference image to obtain the pixel-level deformation field, and then re-positions each pixel in the target image to the reference image through differentiable forward deformation to obtain the deformed target image. The stitched image generation module consists of an encoder-decoder. The encoder consists of Residual Block 1, Residual Block 2, Residual Block 3, and Residual Block 4, where each residual block consists of a fully connected layer, an activation function, and global average pooling. The decoder consists of Convolution Block 1, Convolution Block 2, Deconvolution Block 1, and Deconvolution Block 2. Each convolution block and deconvolution block consists of a skip connection, a convolutional layer, an activation function, and a normalization layer, and the size of the convolutional kernel is uniformly n×n. The size of all feature maps is the same as the size of the input image.

[0014] In step 4, during the training process, the composite loss function is selected as the loss function. The deformation loss is used for the pixel-level deformation field, and the reconstruction loss and adversarial loss are used for the stitching image generation module to eliminate the unnatural seams and misalignments in the overlapping part of the final stitched image. The choice of the loss function affects the quality of the model, can truly reflect the difference between the predicted value and the true value, and can correctly feedback the quality of the model.

[0015] In step 4, during the training process, the root mean square error and structural similarity are selected as appropriate evaluation metrics, which can effectively evaluate the quality of the infrared image stitching algorithm results and the distortion degree of the infrared stitched images, and measure the role of the stitching network.

[0016] The beneficial effects of the present invention are as follows:

[0017] 1. The infrared image stitching framework proposed by the present invention consists of two modules: the pixel-level deformation module and the stitching image generation module. The pixel-level deformation module is used to locate the pixel points in the overlapping area to be stitched, and the pixel-level transformation is used to replace the traditional homography transformation, which solves the large parallax problem in the infrared image stitching task and enhances the naturalness and authenticity of the infrared stitched images.

[0018] 2. In the encoder of the stitching image generation module of the present invention, residual blocks are adopted to extract features from the feature maps, which can enhance the feature representation of the feature maps and reduce the loss of feature information in the feature maps.

[0019] 3. The present invention proposes a composite loss function composed of deformation loss, adversarial loss, and reconstruction loss, which improves the quality of the infrared stitched images from two aspects of edge structure and hole perception, and at the same time makes the images more realistic.

[0020] 4. The addition of the residual structure and skip connections in the present invention helps to reduce the network parameters, makes the network shallower, and the number of network parameters is less. Finally, the entire network has a simple structure, improving the stitching efficiency and accuracy. Brief Description of the Drawings

[0021] Figure 1 It is a flow chart of a pixel-level infrared image stitching method based on deep learning according to the present invention.

[0022] Figure 2 It is a structural diagram of a network model of a pixel-level infrared image stitching method based on deep learning.

[0023] Figure 3 It is a schematic diagram of the encoder-decoder structure of the stitching image generation module described in the present invention.

[0024] Figure 4 It is a composition structure diagram of each residual block in the encoder described in the present invention.

[0025] Figure 5 This is the composition structure diagram of each convolutional block and transposed convolutional block in the decoder of the present invention. Detailed implementation manners

[0026] The present invention will be further described in detail below with reference to the accompanying drawings.

[0027] As Figure 1 shown, a pixel-level infrared image stitching method based on deep learning, the method specifically includes the following steps:

[0028] Step 1, construct a network model. The entire network consists of two parts: a pixel-level deformation module and a stitched image generation module. The pixel-level deformation module uses an optical flow estimation model to obtain the pixel-level deformation from the target image to the reference image, and uses the obtained deformation field to reposition the pixel points of the target image. The stitched image generation module deforms and stitches the target image and the reference image. Among them, the pixel-level deformation module first estimates the pixel-level transformation from the target image to the reference image, and reorients each pixel in the target image to the reference image through differentiable forward deformation to obtain the deformed target image. The stitched image generation module consists of an encoder-decoder. The encoder consists of Residual Block 1, Residual Block 2, Residual Block 3, and Residual Block 4. Each residual block consists of a fully connected layer, an activation function, and global average pooling. The decoder consists of Convolutional Block 1, Convolutional Block 2, Transposed Convolutional Block 1, and Transposed Convolutional Block 2. Each convolutional block and transposed convolutional block consists of a skip connection, a convolutional layer, an activation function, and a normalization layer. The size of the convolutional kernel is uniformly n×n. The size of all feature maps is the same as the size of the input image.

[0029] Step 2, prepare the dataset. Select the KAIST dataset, adjust the size of each image in the dataset, and fix the size of the input image.

[0030] Step 3, train the network model. Input the dataset prepared in Step 2 into the network model constructed in Step 1 for training.

[0031] Step 4, select the loss function to be minimized and the optimal evaluation metric. By minimizing the loss function between the network output image and the label until the number of training iterations reaches the set threshold or the value of the loss function falls within the set range, it can be considered that the model parameters have been pre-trained, and the model parameters are saved. At the same time, select the optimal evaluation metric to measure the accuracy of the algorithm and evaluate the performance of the system; during the training process, the composite loss function is selected for the loss function. The pixel-level deformation module uses the deformation loss, and the stitched image generation module uses the reconstruction loss and the adversarial loss to eliminate the unnatural seams and misalignments in the overlapping parts of the final stitched image; the selection of the loss function affects the quality of the model, can truly reflect the difference between the predicted value and the true value, and can correctly feedback the quality of the model; the appropriate evaluation metrics selected are the root mean square error and the structural similarity, which can effectively evaluate the quality of the infrared image stitching algorithm results and the degree of image distortion, and measure the role of the stitching network.

[0032] Step 5, fine-tune the model. Use the LTIR dataset to train and fine-tune the model to obtain stable and usable model parameters, and further improve the infrared image stitching ability of the model. Finally, make the model have a better effect on infrared image stitching.

[0033] Step 6, save the model. Solidify the finally determined model parameters. When infrared image stitching operations need to be performed later, directly input the infrared image into the network to obtain the final stitched image.

[0034] Embodiment:

[0035] As Figure 1 shown, a pixel-level infrared image stitching method based on deep learning, which specifically includes the following steps:

[0036] Step 1, construct a network model.

[0037] As Figure 2 shown, the entire network consists of two parts: a pixel-level deformation module and a stitched image generation module. The pixel-level deformation module uses an optical flow estimation model to obtain the pixel-level deformation from the target image to the reference image:

[0038]

[0039] Among them, is the forward deformation, I0 and I1 are the target image and the reference image, F 0→t and F 1→t are the optical flows of the target image and the reference image, ψ is the filter, and φ is the synthesis network.

[0040] Then estimate the pixel-level deformation field, and reposition each pixel in the target image to the reference image through differentiable forward deformation to obtain the deformed target image. The differentiable forward deformation function is defined as follows:

[0041]

[0042] Among them, I0 is the input image, F 0→t is the optical flow, Z is the mask, is the summation operation.

[0043] The encoder-decoder structure diagram of the spliced image generation module is as Figure 3 shown. A skip connection is used between the encoder and the decoder. The reference image and the deformed target image are input to share the encoder and are averaged and fused. The encoder consists of four residual blocks; among them, residual block one, residual block two, residual block three, and residual block four are composed of two fully connected layers, an activation function, and global average pooling, aiming to reduce the problem of edge feature loss and guide the network to capture more critical features. The specific structure of each residual block is as Figure 4 shown; in order to ensure that the network can retain more structural information and fully extract the features of the infrared image, the activation function used in the present invention is the S-shaped function, and the S-shaped function is defined as follows:

[0044]

[0045] The decoder consists of two convolutional blocks and two deconvolutional blocks; among them, convolutional block one and convolutional block two are composed of a skip connection, a convolutional layer, an activation function, and a normalization layer, and the size of the convolutional kernel is uniformly 1×1, and the stride is 1; deconvolutional block one and deconvolutional block two are composed of a skip connection, a convolutional layer, an activation function, and a normalization layer, and the size of the convolutional kernel is uniformly 1×1, and the stride is 1. Among them, the first 1×1 convolution is mainly used for dimensionality increase, deconvolutional block one and deconvolutional block two are mainly used for enhancing the resolution of the feature map, and the last 1×1 convolution is used for dimensionality reduction. The activation function uses the S-shaped function. The target image and the reference image pass through the encoding module and the decoding module, and the features are deeply fused, and finally the size of the feature map is consistent with the size of the input image. The specific structure of each convolutional block and deconvolutional block is as Figure 5 shown.

[0046] Step 2, prepare the dataset. The infrared image dataset uses the KAIST dataset for training and testing. Among them, the training set contains 50,187 pictures, and the test set contains 45,141 pictures. The dataset captured various conventional traffic scenes including campuses, streets, and the countryside during the day and at night. The picture size is 640×480.

[0047] Step 3, train the network model. Perform image enhancement on the pictures in the dataset, randomly perform diffraction transformation on the same picture, and crop it to the size of the input picture as the input of the entire network, and use the pictures with annotations in the dataset as labels. The random size and position can be achieved through software algorithms. Using the pictures with annotations in the dataset as labels is to enable the network to learn better feature extraction capabilities and ultimately achieve better stitching effects.

[0048] Step 4, select the minimized loss function and the optimal evaluation metric. Achieve better stitching effects by minimizing the loss function. During the training process, the composite loss function is selected for the loss function. The pixel-level deformation field uses the deformation loss, and the stitching image generation module uses the reconstruction loss and the adversarial loss to eliminate the unnatural seams and misalignments in the overlapping parts of the final stitched image;

[0049] The stitched image is obtained by fusing the overlapping regions of the reference image and the deformed target image. By applying the same loss function to all pixels in the overlapping region, the network can output a more accurate stitching result. Then the deformation loss is defined as:

[0050]

[0051] where W ov and W nov are the estimated values of the deformation fields in the overlapping region and the non-overlapping region respectively. and are the ground truth labels of W ov and W nov Here, in order to make the model mainly focus on the overlapping region while maintaining the minimum guidance for the non-overlapping region, α regularizes the non-overlapping deformation supervision.

[0052] Since when the pixel-level deformation field deforms the pixel points in the overlapping region, unnecessary holes will be generated in the target image. When using image inpainting technology to fill these holes, it is impossible to distinguish all these holes from the background, and misalignment and seams may occur. To reduce these artifacts, we design a hole-aware reconstruction loss and an adversarial loss. Therefore, when the adversarial loss eliminates the artifacts, the reconstruction loss is used to constrain the stitching image generation module and retain the input image. The reconstruction loss is defined as:

[0053]

[0054]

[0055]

[0056] where, I S is the stitched image, I R is the reference image, IWT is the target image, and ⊙ is pixel-level multiplication.

[0057] The adversarial loss is defined as:

[0058]

[0059] where D(·) is the discriminator, and I S is the stitched image.

[0060] The total loss of the stitched image generation module is defined as:

[0061] L SIGMo = λ r L recon + λ a L adv

[0062] where λ r and λ a are the weighted parameters of the reconstruction loss and the adversarial loss, respectively.

[0063] The pixel-level deformation field and the loss function of the stitched image generation module help the network learn clearer edges and more detailed textures, enabling the stitched image to eliminate gaps, be stitched more naturally, and have a better visual effect.

[0064] In step 4, the appropriate evaluation metrics selected are the root mean square error and the structural similarity. The root mean square error can effectively evaluate the accuracy of pixel point registration in the overlapping area; the structural similarity measures the image similarity from three aspects: brightness, contrast, and structure, and is an index used to measure the similarity degree between two digital images. The root mean square error and the structural similarity are defined as follows:

[0065]

[0066]

[0067] where, and are the matching point pairs between the reference image and the target image in the images to be stitched, T is the transformation model, θ is the model parameter, and ||·|| is the distance between two points. μ x , μ y represent the mean and variance of the reference image and the target image, respectively, and represent the standard deviation of the reference image and the target image, respectively, and σ xy represents the covariance of the reference image and the target image, and C1 and C2 are constants.

[0068] Set the number of training times to 200. The size of the number of images input into the network each time is about 8 - 16. The upper limit of the size of the number of images input into the network each time is mainly determined by the performance of the computer graphics processor. Generally, the larger the number of images input into the network each time, the better, making the network more stable. The learning rate during the training process is set to 0.0001, which can not only ensure the network to fit quickly but also prevent the network from overfitting. The network parameter optimizer selects the adaptive moment estimation algorithm. Its main advantage lies in that after bias correction, there is a definite range for the learning rate in each iteration, making the parameters relatively stable. The threshold value of the loss function value is set to about 0.0003. If it is less than 0.0003, it can be considered that the training of the entire network has been basically completed.

[0069] Step 5, fine-tune the model. Use the LTIR dataset to train and fine-tune the model. For the LTIR dataset, we use 500 images for training and 200 images for testing.

[0070] Step 6, save the model. When the network training is completed, all the parameters in the network need to be saved. Then, by inputting the infrared images to be stitched into the network, the stitched images can be obtained. This network has no requirements for the size of the input images, and any size is acceptable.

[0071] Among them, the implementations of convolution, activation function, average fusion, fully connected layer, global average pooling, etc. are algorithms well-known to those skilled in the art. The specific processes and methods can be found in the corresponding textbooks or technical literatures.

[0072] The present invention constructs a pixel-level infrared image stitching method based on deep learning, which can directly stitch two narrow-view infrared images into a super-wide-view infrared image. The obtained infrared stitched image is not affected by parallax, and at the same time, artifacts are eliminated, making it more in line with the human eye visual effect. Under the same conditions, by calculating the relevant indicators of the images obtained by the existing methods, the feasibility and superiority of this method are further verified. The comparison of the relevant indicators of the existing technology and the method proposed by the present invention is shown in Table 1:

[0073] Table 1 Comparison of relevant indicators of the existing technology and the method proposed by the present invention

[0074]

[0075] It can be seen from the table that the method proposed by the present invention has lower root mean square error, higher structural similarity, and higher signal-to-noise ratio than the existing methods. These indicators further illustrate that the method proposed by the present invention has better stitching quality and lower computational complexity.

Claims

1. A pixel-level infrared image stitching method based on deep learning, characterized in that, The method includes the following steps: Step 1, construct a network model: The whole network consists of two parts, a pixel-level deformation module and a stitched image generation module: The pixel-level deformation module uses an optical flow estimation model to obtain the pixel-level deformation from the target image to the reference image, and uses the obtained pixel-level deformation to reposition each pixel point in the target image to obtain a deformed target image; The stitched image generation module stitches the deformed target image and the reference image to obtain an infrared stitched image; Step 2, prepare a dataset: Select the KAIST dataset, adjust the size of each image in the dataset, and fix the size of the input image; Step 3, train the network model: Input the dataset prepared in Step 2 into the network model constructed in Step 1 for training; Step 4, select the minimized loss function and the optimal evaluation metric: By minimizing the loss function between the network output image and the label until the number of training times reaches the set threshold or the value of the loss function reaches the set range, it can be considered that the model parameters have been pre-trained, and the model parameters are saved; At the same time, select the optimal evaluation metric to measure the accuracy of the algorithm and evaluate the performance of the system; Step 5, fine-tune the model: Use the LTIR dataset to train and fine-tune the model to obtain stable and usable model parameters, and finally make the model have a better effect on infrared image stitching; Step 6, save the model: Solidify the finally determined model parameters. When infrared image stitching operations need to be performed later, directly input the image into the network to obtain the final stitched image; In the above step 1, the pixel-level deformation module first estimates the pixel-level transformation from the target image to the reference image to obtain the pixel-level deformation field, and then relocates each pixel in the target image to the reference image through differentiable forward deformation to obtain the deformed target image. The stitched image generation module consists of an encoder-decoder. The encoder is composed of residual block 1, residual block 2, residual block 3, and residual block 4, where each residual block consists of a fully connected layer, an activation function, and global average pooling. The decoder is composed of convolution block 1, convolution block 2, deconvolution block 1, and deconvolution block 2. Each convolution block and deconvolution block consists of a skip connection, a convolutional layer, an activation function, and a normalization layer, and the size of the convolution kernel is uniformly ; the size of all feature maps is consistent with the size of the input image.

2. The pixel-level infrared image stitching method based on deep learning according to claim 1, wherein, In Step 4, during the training process, the loss function is selected to use a composite loss function. The pixel-level deformation field uses a deformation loss, and the stitched image generation module uses a reconstruction loss and an adversarial loss to eliminate the unnatural seams and misalignments in the overlapping parts of the final stitched image; The choice of the loss function affects the quality of the model, can truly reflect the difference between the predicted value and the true value, and can correctly feedback the quality of the model.

3. A pixel-level infrared image stitching method based on deep learning according to claim 1, characterized in that, In Step 4, during the training process, the evaluation metrics are selected as the root mean square error and the structural similarity, which can effectively evaluate the quality of the results of the infrared image stitching algorithm and the distortion degree of the infrared stitched image, and measure the role of the stitching network.

Citation Information

Patent Citations

  • An infrared image splicing method based on coarse-to-fine registration

    CN109886878A

  • End-to-end infrared and visible light image fusion method

    CN113298744A

  • Multi-platform multi-view image splicing method in airport large-range environment

    CN115222595A