A method for infrared video distortion correction based on deep learning
By constructing an infrared distortion correction method based on convolutional neural network, the infrared camera image deformation problem is solved, efficient and accurate infrared video distortion correction is achieved, the correction process is simplified, and the speed and accuracy are improved.
Patent Information
- Application Number
- CN202111321639.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-11-09
AI Technical Summary
The infrared pictures taken by existing infrared cameras have deformation problems. Ordinary black and white chessboards have small distinction in infrared images, unclear edges, and difficult to effectively correct infrared video distortion.
A infrared distortion correction method based on convolutional neural network is constructed. By creating an infrared distortion image data set, an efficient convolutional neural network structure is designed, and the neural network is used to automatically extract infrared distortion image features and establish mapping relationships to realize automatic correction of infrared distortion images.
It improves the efficiency and accuracy of infrared video distortion correction, simplifies the correction process, improves the speed by 20%, improves the accuracy by 15%, and has good universality and efficient automatic correction capabilities.
Smart Images

Figure CN114463192B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image correction method in the fields of image processing and deep learning, and particularly to a method for correcting infrared video distortion based on deep learning. Background Art
[0002] With the development of society and economy, people pay more and more attention to the protection of personal and property safety. Installing video surveillance as the most common security measure is widely used, and infrared cameras are widely used because of their outstanding night surveillance effect. Infrared cameras have the advantages of high brightness, long night vision distance, stable performance, and uniform light. However, the infrared images captured by infrared cameras have distortion problems. In order to better utilize the images and videos captured by infrared cameras, it is necessary to correct these distorted infrared images and videos into images and videos that conform to people's visual habits.
[0003] For the correction of ordinary cameras, black and white checkerboards are generally used at present. By using the obvious edge differentiation of black and white grids under visible light, the distance between the distorted focal points can be detected, so as to achieve the purpose of correcting the camera. However, ordinary black and white checkerboards have problems such as small differentiation, unclear focal points, and unclear edges in infrared images. Therefore, other methods are needed for the correction of infrared video distortion.
[0004] In the past few years, deep learning has shown great advantages in image target recognition, speech recognition, and natural language processing. Among many types of neural networks, the convolutional neural network is the most widely used one. In the early stage, due to the lack of large-scale training sets and computer capacity, it was very difficult to train a convolutional neural network with good performance and no overfitting problems. With the emergence of large-scale labeled databases such as the ImageNet database, and the greatly improved GPU-assisted computing power in recent years, the convolutional neural network has returned to the eyes of researchers.
[0005] Based on the great advantages shown by the convolutional neural network in image information processing, and with the rapid development of deep learning technology, more and more researchers have begun to try image correction methods based on convolutional neural networks. Among them, there is a fish-eye image correction method based on convolutional neural networks. This method constructs a convolutional neural network for estimating the distortion coefficients of fish-eye images, and uses artificially synthesized fish-eye images as the training set to train this convolutional neural network. This method realizes the correction of fish-eye images by converting the fish-eye image distortion problem into a fish-eye image distortion coefficient classification problem. Similarly, for infrared images, a distortion correction method based on neural networks can also be designed. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to automatically correct infrared distorted images by establishing a convolutional neural network, implement image distortion correction based on the neural network, and improve efficiency. The method uses a convolutional neural network to extract the feature information of infrared distorted images and establish a mapping relationship between the distortion features and the correction parameters of infrared distorted images.
[0007] The technical solution adopted by the present invention is as follows:
[0008] Step 1: Create a dataset of infrared distorted images;
[0009] Step 2: Construct a convolutional neural network that meets the accuracy requirements and can make rapid inferences;
[0010] Step 3: Use the convolutional neural network to train the dataset of infrared distorted images;
[0011] Step 4: Input the infrared distorted images with unknown distortion parameters and correction parameters to be measured into the trained convolutional neural network to obtain the correction parameters and correct the infrared distorted images to be measured.
[0012] The said Step 1 includes the following steps:
[0013] Collect multiple infrared undistorted images containing different scenes. Each image is given N sets of randomly formed distortion parameters to obtain N distorted infrared distorted images, thereby forming all infrared distorted images. The corresponding correction parameters are obtained by inverse distortion transformation of the distortion parameters of each infrared distorted image. The dataset is composed of all infrared distorted images and their corresponding all correction parameters.
[0014] In the said Step 2, the convolutional neural network is mainly composed of 3 - 4 neural structure units connected in sequence; each neural structure unit includes four consecutive convolutional layers. The input image of the neural structure unit outputs the final feature map after passing through the four consecutive convolutional layers. After the first convolutional layer, a 2 - fold downsampling operation and a 4 - fold downsampling operation are respectively connected. After the 2 - fold downsampling operation, it is connected to the output of the second convolutional layer after an upsampling operation. After the 2 - fold downsampling operation and the 4 - fold downsampling operation, they are respectively connected to the addition fusion layer after a convolutional operation and a 2 - fold upsampling. The output of the addition fusion layer is then connected to the output of the third convolutional layer.
[0015] The input graph X0 of the neural structure unit is processed by the first convolutional layer to obtain the feature map X1. The feature map X1 undergoes a 2-fold downsampling operation to obtain the feature map X4. The feature map X1 is input into the second convolutional layer to obtain the feature map X2. The feature map X2 and the feature map X4 are added and fused, and then input into the third convolutional layer for processing to obtain the feature map X3. At the same time, the feature map X4 undergoes a convolutional operation to obtain the feature map X5. The feature map X1 undergoes a 4-fold downsampling operation to obtain the feature map X6. The result after the feature map X6 undergoes a 2-fold upsampling operation and the result after the feature map X4 undergoes a convolutional operation are added and fused to obtain the feature map X5. The feature map X5 undergoes an upsampling operation and is added and fused with the feature map X3, and then input into the fourth convolutional layer for processing to obtain the final feature map Y.
[0016] The convolutional neural network designed by the present invention has high operating efficiency and can retain the pixel information of the infrared image to the greatest extent, making the correction more accurate.
[0017] In the convolutional neural network designed by the present invention, the high-resolution (X1, X2, X3) and low-resolution (X4, X5, X6) networks are connected in parallel, rather than serially connected as in traditional neural networks. It can always maintain high resolution throughout the process, rather than restoring the resolution through a process from low to high.
[0018] The present invention replaces the downsampling pooling layer in the traditional network with a convolutional layer, retaining the information carried by each pixel in the image to the greatest extent. It can retain the information contained in each pixel and the mutual relationship between each pixel, and achieve the goal of always maintaining high resolution.
[0019] The model proposed therein fuses low-resolution feature maps of the same depth and similar levels to improve the representation effect of the high-resolution feature map, and performs repeated multi-scale fusion, so that small-scale distortions can also be noticed by the model, thereby improving the accuracy of correction.
[0020] The last layer of the neural network uses a fully connected method to fuse multi-dimensional features into the dimension corresponding to the correction parameters, thereby playing a supervisory role during training.
[0021] In step 3, before inputting the infrared distorted image dataset into the convolutional neural network for training, preprocessing is also performed. The preprocessing includes randomly adding random noise, adding random color transformation, and converting to tensor information in sequence. The random color transformation means shuffling the RGB channels of all pixels in the image. Converting to tensor information means converting the pixel information in the image into tensor information tensor and then using it as the input of the convolutional neural network, with the correction parameter as the supervision signal. Randomly adding random noise can prevent the network from overfitting. Adding random color transformation can make the model pay more attention to the morphological features in the image. Converting to tensor information can input the training data into the GPU to accelerate training.
[0022] In step 3 described above, the training process is as follows: First, the calibration parameters and the infrared distorted image dataset are input into the convolutional neural network structure for preliminary training. Then, by comparing the difference between the original undistorted image and the calibrated undistorted image, the convolutional neural network is fine-tuned again. The trained convolutional neural network structure can accurately predict the calibration parameters of the infrared distorted image, achieving a high-precision calibration of the infrared distorted image.
[0023] Step 4 described above is as follows: Use an infrared thermal imager to collect a batch of infrared distorted images to be measured. After converting the pixel information of the infrared distorted images into tensor information, input them into the trained convolutional neural network batch by batch to obtain the calibration parameters corresponding to each image. Use the calibration parameters as the parameters for the affine transformation of the infrared image, and obtain the calibrated image after calculation.
[0024] The beneficial effects of the present invention are:
[0025] Adopting the method for infrared video distortion correction based on neural network of the present invention simplifies the correction process, improves the correction efficiency, and enhances the correction effect. In terms of software, due to the use of the end-to-end training method, the entire process only requires a neural network, simplifying the steps that require a large amount of manual calibration in the traditional method.
[0026] Since the automatic correction method itself is purely based on neural network, the advancement of the algorithm therein improves the distortion correction effect.
[0027] The number of parameters of the neural network in the present invention has been strictly screened. Therefore, while ensuring the calibration accuracy, the operation efficiency is also ensured, enabling a large number of calibration tasks to be completed quickly in a short time. Experiments were carried out on a dedicated deep learning workstation, and the test results showed that the calibration speed reached 20 images per second, and the accuracy rate reached more than 98%. Compared with the conventional convolutional neural network, the speed of the present invention is increased by 20%, and the accuracy rate is increased by 15%.
[0028] This method can overcome the limitations of different cameras and their imaging models on the distortion image correction algorithm, and has good universality. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a schematic flowchart of the method for infrared video correction based on deep learning in the present invention.
[0030] Figure 2 is a schematic diagram of the structural unit of the convolutional neural network in the present invention.
[0031] Figure 3 is a schematic flowchart of the neural network training and prediction process in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0032] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0033] As Figure 1 shown, the embodiments include the following steps:
[0034] Step S1: Create an infrared distorted image dataset.
[0035] Collect 3000 infrared undistorted images containing different scenes, and assign 10 randomly formed distortion parameters to each image, thus forming 30000 infrared distorted images.
[0036] Each image corresponds to a corresponding correction parameter. The dataset is composed of 30000 infrared distorted images and the corresponding 30000 correction parameters.
[0037] Step S2: Design a convolutional neural network that can achieve the required accuracy and make quick inferences.
[0038] As Figure 2 shown, taking a feature map with a size of 256*256*3 as an example, it is input into the neural structure unit.
[0039] (1) First, input the graph X0 through a convolutional layer with a convolutional kernel size of 3*3 and a stride of 2, compress the feature map size to 128*128, and expand the feature channels to 64. At this time, the obtained feature map X1 has a size of 128*128*64.
[0040] (2) Perform 2-fold and 4-fold downsampling on the compressed feature map X1 respectively to obtain a feature map X4 of 64*64*64 and a feature map X6 of 32*32*64.
[0041] (3) Use transposed convolution to perform 2-fold upsampling on the 32*32*64 feature map X6 and add it to the result after convolution of the 64*64*64 feature map X4 in the previous step to obtain the feature map X5 through addition and fusion.
[0042] (4) Perform 2-fold upsampling on the feature map X4 and add it to the result obtained by processing the feature map X1 through a convolutional layer with a convolutional kernel of 3*3 and a stride of 1 to form the feature map X2;
[0043] Perform 2-fold upsampling on the feature map X5 and add it to the result obtained by processing the feature map X2 through a convolutional layer with a convolutional kernel of 3*3 and a stride of 1 to form the feature map X3;
[0044] The feature map X3 is further processed by a convolutional layer with a convolutional kernel of 3*3 and a stride of 1 to obtain the feature map Y.
[0045] (5)Finally, a feature map Y with a size of 128*128*64 is obtained and connected to the next neural structure unit. The same process is performed on the next neural structure unit.
[0046] Figure 2 The part in the first row is of high resolution, and the parts in the following two rows belong to low resolution.
[0047] In the design of a convolutional neural network, the number of the above-mentioned structure units is crucial. If the number is too small, the accuracy is difficult to meet the requirements; if the number is too large, it will affect the running efficiency of the model and even lead to overfitting, affecting the generalization ability. To obtain the optimal number of convolutional units, the following method can be adopted: First, randomly determine the number of convolutional units and adjust the number according to the training and test results. If the inference time is too long, then reduce the number of convolutional units; if the training error is too large, increase the number of convolutional units. Finally, in combination with the test accuracy, gradually adjust and increase the number of convolutional units, and then adjust the convolutional units according to whether the output meets various indicators, so as to determine the number of convolutional units when the network error is minimized. The last layer of the neural network uses a fully connected method to fuse multi-dimensional features into 10 dimensions to correspond to the calibration parameters, playing a supervisory role during training.
[0048] Step S3: The convolutional neural network trains the data set.
[0049] As Figure 3 shown, first, perform random color transformation on the image to prevent the model from paying more attention to the morphological information in the image. Secondly, add Gaussian noise with a random range to the image to enhance the robustness of the model. Finally, convert the pixel information of the image into tensor information as the input of the convolutional neural network, and use the calibration parameters of the image as the supervision signal, with MSE as the loss function. The training process is as follows: First, the calibration parameter estimation network is initially trained, and then the neural network is fine-tuned again by comparing the differences between the original undistorted image and the corrected undistorted image. The trained neural network can accurately predict the calibration parameters of the infrared distorted image, realizing the correction of the infrared distorted image with high accuracy.
[0050] Step S4: As Figure 3 shown, input the infrared distorted image into the neural network to obtain the calibration parameters and correct the image. Use an infrared thermal imager to collect a batch of infrared distorted images, convert the image information into tensor information, and then input it into the trained convolutional neural network batch by batch to obtain the calibration parameters corresponding to each image. Use the calibration parameters as the parameters for the affine transformation of the infrared image, and after calculation, obtain the corrected image.
[0051] During the operation of the above convolutional neural network, different convolutional units perform convolution on the image to extract features respectively, and then use a non-linear activation function for processing. After obtaining the response result, it is continuously input into the next convolutional unit to extract features. The convolutional units gradually extract the morphological and semantic information of the image from low to high. The last convolutional layer is connected to a fully connected layer to output a vector corresponding to the dimension of the calibration parameter. Since the convolutional neural network can automatically extract the features in the image and perform non-linear fitting between the features and the set calibration parameters. By using a large number of distorted images and calibration parameters to train the model, the model has the generalization ability to automatically obtain the calibration parameters after inputting the image, thus greatly reducing the workload of manual calibration.
[0052] It can be seen that on the premise of ensuring that the accuracy rate can reach 98%, the efficiency of using the neural network for autonomous calibration is 3 times that of manual calibration. The neural network proposed by the present invention has a 20% improvement in processing speed and a 15% improvement in accuracy compared with the traditional neural network.
Claims
1. A method for infrared video distortion correction based on deep learning, characterized in that, The method includes the following steps: Step 1: Create an infrared distorted image dataset; Step 2: Construct a convolutional neural network that meets the accuracy requirements and can perform rapid inference; In Step 2, the convolutional neural network is composed of 3 - 4 neural structure units connected in sequence; each neural structure unit includes four consecutive convolutional layers. The graph input to the neural structure unit outputs the final feature map after passing through the four consecutive convolutional layers. After the first convolutional layer, 2x downsampling operation and 4x downsampling operation are respectively connected. After the 2x downsampling operation, it is connected to the output of the second convolutional layer after upsampling operation. After the 2x downsampling operation and 4x downsampling operation, they are respectively connected to the addition fusion layer after convolutional operation and 2x upsampling. The output of the addition fusion layer is then connected to the output of the third convolutional layer; The graph X0 input to the neural structure unit is processed by the first convolutional layer to obtain the feature map X1. The feature map X1 is subjected to 2x downsampling operation to obtain the feature map X4. The feature map X1 is input to the second convolutional layer to obtain the feature map X2. The feature map X2 and the feature map X4 are added and fused and then input to the third convolutional layer for processing to obtain the feature map X3; At the same time, the feature map X4 is subjected to convolutional operation to obtain the feature map X5. The feature map X1 is subjected to 4x downsampling operation to obtain the feature map X6. The result after the 2x upsampling operation of the feature map X6 and the result after the convolutional operation of the feature map X4 are added and fused to obtain the feature map X5. The feature map X5 is upsampled and added and fused with the feature map X3 and then input to the fourth convolutional layer for processing to obtain the final feature map Y; Step 3: Use the convolutional neural network to train the infrared distorted image dataset; Step 4: Input the infrared distorted image to be measured into the trained convolutional neural network to obtain the correction parameters and correct the infrared distorted image to be measured.
2. The method for infrared video distortion correction based on deep learning according to claim 1, characterized in that: The Step 1 includes the following steps: Collect multiple infrared undistorted images containing different scenes. Each image is given N randomly formed distortion parameters to obtain N distorted infrared distorted images, thus forming all infrared distorted images. The corresponding correction parameters are obtained by inverse distortion conversion of the distortion parameters of each infrared distorted image. The dataset is composed of all infrared distorted images and their corresponding all correction parameters.
3. The method for infrared video distortion correction based on deep learning according to claim 1, characterized in that: In Step 3, before inputting the infrared distorted image dataset into the convolutional neural network for training, preprocessing is also performed. The preprocessing includes randomly adding random noise, adding random color transformation, and converting to tensor information in sequence; The random color transformation means shuffling the RGB channels of the pixels in the image. Converting to tensor information means converting the pixel information in the image to tensor information tensor and then using it as the input of the convolutional neural network.
4. The method for infrared video distortion correction based on deep learning according to claim 1, characterized in that, In Step 3, the training process is as follows: First, use the correction parameters and the infrared distorted image dataset to input into the convolutional neural network structure for preliminary training, and then fine-tune the convolutional neural network again by comparing the difference between the original undistorted image and the corrected undistorted image.
5. The method for infrared video distortion correction based on deep learning according to claim 1, characterized in that The step 4 is as follows: Use an infrared thermal imager to collect a batch of infrared distorted images to be measured. After converting the pixel information of the infrared distorted images into tensor information, input them into the trained convolutional neural network batch by batch to obtain the correction parameters corresponding to each image. Use the correction parameters as the parameters for the affine transformation of the infrared image, and obtain the corrected image after calculation.
Citation Information
Patent Citations
Data processing method, data training method, data identifying method and device, and storage medium
WO2021057810A1