A method for improving the resolution of infrared images
A deep learning-based neural network model enhances infrared image resolution on embedded devices by channel expansion and feature map processing, addressing computational and power challenges, enabling real-time high-resolution imaging with reduced power and cost.
Patent Information
- Application Number
- CN202110310821.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-03-23
AI Technical Summary
The prior art is difficult to achieve real-time high-resolution infrared image generation on embedded devices, traditional algorithms are poor in effect, deep learning algorithms have high computational volume, high power consumption, high cost, and difficult to deploy under the limitations of device size and power consumption.
A neural network model based on deep learning is adopted to design a convolutional neural network model through channel expansion, feature map splitting and stitching, and expansion operations. The image super-resolution processing is used for image super-resolution processing, including convolution, deconvolution and bicubital interpolation operations.
On a low-power embedded platform, it can achieve 2 to N times the resolution of infrared image, and has real-time imaging capabilities, low power consumption and low cost, meeting the actual needs of embedded devices.
Smart Images

Figure CN112927140B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image super-resolution, and particularly to a method for improving the resolution of infrared images. Background Art
[0002] Infrared imaging technology forms images based on the thermal energy radiation of objects. The information, feature distribution, signal-to-noise ratio, clarity, etc. contained in the images are different from those of visible light images. Its advantage is that it does not require visible light, is less affected by the environment, and can image 24 hours a day and all-weather. Therefore, this technology has been widely used in the fields of security, target recognition, etc. However, its disadvantage is that the image details are less than those of visible light images. Therefore, in order to obtain a higher target recognition rate in infrared images, the image needs to have a higher resolution.
[0003] Image super-resolution technology aims to use algorithms to improve image resolution so that the image has more details. Previously, people often used various interpolation algorithms (such as bilinear interpolation, bicubic interpolation) to improve image resolution, and the generated images were relatively blurred. In recent years, due to the outstanding effects of deep learning algorithms, the market has increasingly tended to use deep learning algorithms to achieve image super-resolution. However, since deep learning algorithms have more parameters and require a higher computing speed of the device, the general operating environment is a PC or industrial control computer with higher power consumption. If the computing speed of the device is low, the speed of generating high-resolution images is slow.
[0004] In order to meet the requirements of real-time performance, currently, the embedded devices that achieve high-resolution infrared imaging on the market mainly use high-resolution imaging hardware to achieve, rather than using deep learning-based image super-resolution algorithms. The disadvantages of using high-resolution imaging hardware to achieve are: high power consumption, high cost, low economic benefits, and it is not conducive to promotion. The difficulty in deploying super-resolution algorithms on embedded devices is: if traditional algorithms such as interpolation are directly used, the super-resolution effect of the obtained infrared images is obviously not satisfactory; if deep learning algorithms are used, since general deep learning models have more parameters and a large amount of calculations, high-power consumption embedded devices need to be used. However, in many scenarios, due to restrictions such as volume and cost, embedded devices are generally small in volume and low in power consumption platforms, and it is easy to have a slow running speed and unable to generate high-resolution images in real time; if the deep learning model parameters and the amount of calculations are blindly reduced, it is easy to reduce the effect of the algorithm. Summary of the Invention
[0005] In order to solve the above technical problems existing in the prior art, the present invention provides a method for improving the resolution of infrared images, which is applied to an embedded device and includes:
[0006] Step S1, inputting a low-resolution infrared image whose resolution needs to be improved into a trained neural network model;
[0007] Step S2: Use the neural network model to perform channel expansion on the low-resolution infrared image to form a multi-channel image;
[0008] Step S3: Use the neural network model to split the feature map of the multi-channel image according to the number of channels, and process the feature map of each channel separately and then splice them again to form the feature map of the processed multi-channel image;
[0009] Step S4: Use the neural network model to expand the feature map output in Step S3, and form and output a high-resolution infrared image according to the expanded feature map.
[0010] Preferably, in Step S1, use a pre-prepared training sample set to train the neural network model so that the loss function of the neural network model reaches the training conditions, thereby completing the training process of the neural network model.
[0011] Preferably, in Step S2, perform a first convolution operation on the low-resolution infrared image once to correspondingly expand the number of channels of the low-resolution infrared image according to the size of the convolution kernel to form a multi-channel image.
[0012] Preferably, in Step S3, the process of processing the feature map of one channel of the multi-channel image specifically includes:
[0013] Step S31: Perform the first convolution operation on the feature map of the one channel once to obtain a first feature map;
[0014] Step S32: Perform a second convolution operation on the first feature map once to obtain a second feature map;
[0015] Step S33: Perform the first convolution operation on the second feature map once to obtain a third feature map;
[0016] Step S34: Perform an upsampling operation on the third feature map once to obtain a fourth feature map;
[0017] Step S35: Splice and process the first feature map and the fourth feature map to obtain a fifth feature map;
[0018] Step S36: Perform the first convolution operation on the fifth feature map once to obtain a sixth feature map;
[0019] Step S37: Splice and process the fourth feature map and the sixth feature map to obtain a seventh feature map;
[0020] Step S38, perform the first convolution operation on the seventh feature map once to obtain an eighth feature map;
[0021] Step S39, splice the eighth feature map with the eighth feature maps of other channels to obtain the feature map of the multi-channel image.
[0022] Preferably, step S4 includes:
[0023] Step S41, perform a first expansion operation on the feature map of the multi-channel image multiple times to obtain an expanded first expanded feature map;
[0024] Step S42, perform a second expansion operation on the feature map of the low-resolution infrared image once to obtain a second expanded feature map that is expanded by the same multiple as the first feature map;
[0025] Step S43, after splicing the first expanded feature map and the second expanded feature map according to the number of channels, obtain a third expanded feature map;
[0026] Step S44, perform the first convolution operation on the third expanded feature map once to obtain the high-resolution infrared image.
[0027] Preferably, the convolution kernel size of the first convolution operation is 3*3, and the stride is 1.
[0028] Preferably, the convolution kernel size of the second convolution operation is 3*3, and the stride is 2.
[0029] Preferably, the convolution kernel size of the first expansion operation is 4*4, and the stride is 1.
[0030] Preferably, the second expansion operation is a bicubic interpolation upsampling operation.
[0031] Preferably, the training condition associated with the neural network model is to set a loss function for the neural network model;
[0032] The loss function is specifically:
[0033]
[0034] Wherein,
[0035] MSE is the loss value calculated by the loss function;
[0036] m is the total number of pixel points of a real infrared image;
[0037] is the value of each pixel point of a real infrared image;
[0038] is the value of each pixel point of the high-resolution infrared image.
[0039] The beneficial effects of the technical solution of the present invention are as follows: Design a method for improving the resolution of infrared images based on a deep learning convolutional neural network, which can expand the resolution of the input image by more than 2 times, can be deployed on a low-power embedded platform, and has a high derivation operation speed and good effects. Therefore, compared with traditional infrared imaging embedded devices, it has the characteristics of high image resolution, real-time imaging, low power consumption and low cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic flowchart of a method for improving the resolution of infrared images in a preferred embodiment of the present invention;
[0041] Figure 2 is a schematic flowchart of step S3 of a method for improving the resolution of infrared images in a preferred embodiment of the present invention;
[0042] Figure 3 is a schematic flowchart of step S4 of a method for improving the resolution of infrared images in a preferred embodiment of the present invention;
[0043] Figure 4 is a schematic structural diagram of a convolutional neural network model in a preferred embodiment of the present invention;
[0044] Figure 5 is a schematic structural diagram of a grouped convolution splicing module in a convolutional neural network model in a preferred embodiment of the present invention;
[0045] Figure 6 is a schematic structural diagram of the first convolution operation of the grouped convolution splicing module in a preferred embodiment of the present invention;
[0046] Figure 7 is a schematic structural diagram of the second convolution operation of the grouped convolution splicing module in a preferred embodiment of the present invention;
[0047] Figure 8 is a schematic structural diagram of the first expansion operation of the grouped convolution splicing module in a preferred embodiment of the present invention;
[0048] Figures 9a - 9c is a visual effect comparison diagram under 4-fold magnification in an embodiment of a method for improving the resolution of infrared images. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0051] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but it is not limited to the present invention.
[0052] In a preferred embodiment of the present invention, in view of the above problems existing in the prior art, a method for improving the resolution of infrared images is provided, which is applied to an embedded device, such as Figure 1 shown, including:
[0053] Step S1, input a low-resolution infrared image whose resolution needs to be improved into a trained neural network model;
[0054] Step S2, use the neural network model to perform channel expansion on the low-resolution infrared image to form a multi-channel image;
[0055] Step S3, use the neural network model to split the feature maps of the multi-channel image according to the number of channels, and process the feature maps of each channel respectively and then splice them again to form the feature maps of the processed multi-channel image;
[0056] Step S4, use the neural network model to expand the feature maps output in Step S3, and form and output a high-resolution infrared image according to the expanded feature maps.
[0057] All embodiments of this technical solution use Rockchip RK3399Pro as the embedded platform to be deployed and RKNN as the inference framework. It can make the most of the convolutional neural network processor of its embedded platform to process multimedia data such as infrared videos and images.
[0058] The convolutional neural network model included in this platform consists of a conv convolutional layer, a transpose conv deconvolution layer, a concatenate splicing operation, a bicubic interpolation operation, a split channel grouping operation, and a Pixel shuffle operation, as Figure 4 shown; the grouped convolutional splicing module included in this convolutional neural network model, as Figure 5 shown.
[0059] The method for enhancing the resolution of infrared images based on convolutional neural network models with different settings can enhance the resolution of infrared images by 2 to N times, not limited to higher multiples. For example, an image of 320*256 can be enhanced to an image of 640*512, or an image of 320*256 can be enhanced to an image of 1280*1024, etc. The convolutional neural network model included in this technical solution, such as Figure 4 shown, sets two deconvolution layers to perform 2 times of deconvolution upsampling operations, and the bicubic interpolation upsampling is set to 4 times, which can enhance the resolution of infrared images by 4 times. Further, by adjusting Figure 4 the number of deconvolution layers, the size of the deconvolution kernel, the stride size, and the bicubic interpolation upsampling ratio in
[0060] other magnification factors can be achieved. If three deconvolution layers are set to perform 3 times of deconvolution upsampling operations, and the bicubic interpolation upsampling is set to 8 times, the resolution of infrared images can be enhanced by 8 times.
[0060] In the following embodiments of this technical solution, taking enhancing the resolution of infrared images by 4 times as an example, a convolutional neural network model including a Grouping convolution&concatenate block is trained. The feature extraction of low-resolution infrared images is mainly realized in the Grouping convolution&concatenate block, and the upsampling magnification factor is realized through the transpose conv deconvolution layer and the bicubic interpolation operation. When the upsampling magnification factor is 4, the structure is as Figure 4 shown, Figure 5 is Figure 4 the structure diagram of the Grouping convolution&concatenate block in
[0061] . The input represents the low-resolution infrared image whose resolution needs to be enhanced, and the output represents the high-resolution infrared image generated by using the trained neural network model. Taking a 320*256 low-resolution infrared image as an example, the resolution of the generated high-resolution infrared image is 1280*1024.
[0061] Specifically, the feature extraction of low-resolution infrared images is mainly realized in the Grouping convolution&concatenate block. As Figure 5 shown, the convolutional neural network model is provided with a channel expansion parameter n and a feature extraction times parameter i. The low-resolution infrared image is channel-expanded according to the channel expansion parameter n, and the Grouping convolution&concatenate block is set to n sub-modules of the Grouping convolution&concatenate block according to the parameter n. Each sub-module of the Grouping convolution&concatenate sub-module correspondingly extracts features from the feature maps of the infrared images of each channel. The feature extraction times parameter i is set so that the feature maps of the low-resolution infrared images are subjected to i times of feature extraction.
[0062] In a preferred embodiment, in step S1, a pre-prepared training sample set is used to train the neural network model so that the loss function of the neural network model reaches the training conditions, thereby completing the training process of the neural network model.
[0063] In this embodiment, infrared videos with a resolution of 1080*1280 are collected in different scenarios. The frames of these videos are used as labels (GT). Each frame is downsampled by bicubic interpolation to a resolution of 320*256 as the low-resolution map (LR). A total of 1700 GT maps and corresponding LR maps are obtained as training samples; infrared videos with a resolution of 320*256 are collected in different scenarios, and the frames of these videos are used as test samples. The training samples and test samples are sorted and screened to obtain 1700 pairs of training images GT and LR, and 100 test images.
[0064] Set the channel expansion parameter n in this convolutional neural network model to 7, the feature extraction times parameter i to 5, and the upsampling magnification to 4. Use this method to train the network model on a PC with the training samples. When the loss drops to a reasonable range, use the test samples to test on the PC. If the test results meet the requirements, convert the model into an RKNN model, deploy it on an embedded device, and use the test samples to test again, and calculate the average value of the output NIQE and the running speed.
[0065] Algorithm Average NIQE Average derivation speed on RK3399Pro Bicubic interpolation 14.18 The effect is poor and not deployed This method 12.08 30 ms / frame
[0066] Table 1
[0067] Combined with Figures 9a - 9c As can be seen from Table 1, this method meets the requirements of both speed and effect and can meet the actual application.
[0068] In the following embodiments, the batch size b = 100, the image height h = 320, the image width w = 256, and the number of channels c = 3 are taken as examples for detailed description. The number of channels c is variable and is realized by the number of convolution kernels specifically set in each convolutional layer.
[0069] In a preferred embodiment, in step S2, the low-resolution infrared image is subjected to a first convolution operation once to correspondingly expand the number of channels of the low-resolution infrared image according to the convolution kernel size to form a multi-channel image.
[0070] In this embodiment, the dimension of the low-resolution infrared image is the batch size b, the image height h, the image width w, and the number of channels c, forming an input array corresponding to the dimension of the low-resolution infrared image of (b, h, w, c). After passing through the conv_re convolution operation once, the dimension becomes (b, h, w, n*c), and it enters the grouped convolution splicing module.
[0071] Specifically, the number of convolution kernels in the convolution layer in this embodiment is 7*3.
[0072] Specifically, as shown in Table 2:
[0073]
[0074] Table 2
[0075] As Figure 2 shown, in a preferred embodiment, in step S3, the process of processing the feature map of one channel of the multi-channel image specifically includes:
[0076] Step S31, performing a first convolution operation on the feature map of one channel once to obtain a first feature map;
[0077] Step S32, performing a second convolution operation on the first feature map once to obtain a second feature map;
[0078] Step S33, performing a first convolution operation on the second feature map once to obtain a third feature map;
[0079] Step S34, performing an upsampling operation on the third feature map once to obtain a fourth feature map;
[0080] Step S35, splicing the first feature map and the fourth feature map to obtain a fifth feature map;
[0081] Step S36, performing a first convolution operation on the fifth feature map once to obtain a sixth feature map;
[0082] Step S37, splicing the fourth feature map and the sixth feature map to obtain a seventh feature map;
[0083] Step S38, performing a first convolution operation on the seventh feature map once to obtain an eighth feature map;
[0084] Step S39, splicing the eighth feature map with the eighth feature maps of other channels to obtain the feature map of the multi-channel image.
[0085] In this embodiment, the grouped convolution splicing module splits the multi-channel image into n along the channel dimension, and the dimension of each feature map is (b, h, w, c). The following operations are performed on the feature map of each channel: First, perform a conv_re convolution operation once to obtain a first feature map with a dimension of (b, h, w, 64); the feature map obtained after performing a conv_down convolution operation on the first feature map is the second feature map with a dimension of (b, h / 2, w / 2, 256); then perform a conv_re convolution operation on the second feature map once to obtain a third feature map with a dimension of (b, h / 2, w / 2, 256); then perform a pixelshuffle upsampling operation on the third feature map once to obtain a fourth feature map with a dimension of (b, h, w, 64); splice the first feature map and the fourth feature map along the channel to obtain a fifth feature map with the dimension changed to (b, h, w, 128); then perform a conv_re convolution operation on the fifth feature map once to obtain a sixth feature map with a dimension of (b, h, w, 64); then splice the fourth feature map and the sixth feature map along the channel to obtain a seventh feature map with a dimension of (b, h, w, 128); then perform a conv_re convolution operation on the seventh feature map once to obtain an eighth feature map with a dimension of (b, h, w, c). Finally, splice the eighth feature map with the eighth feature maps of other channels to change back to the feature map of the multi-channel image with a dimension of (b, h, w, n*c).
[0086] Specifically, in this embodiment, the number of convolution kernels of the convolution layers for obtaining the first feature map, the fourth feature map, and the sixth feature map is 64; the number of convolution kernels of the convolution layers for obtaining the second feature map and the third feature map is 256; the number of convolution kernels of the convolution layer for obtaining the eighth feature map is 3.
[0087] Specifically, as shown in Table 3:
[0088]
[0089]
[0090] Table 3
[0091] As Figure 3 shown, in a preferred embodiment, step S4 includes:
[0092] Step S41, perform a first expansion operation on the feature map of the multi-channel image multiple times to obtain an expanded first expanded feature map;
[0093] Step S42, perform a second expansion operation on the feature map of the low-resolution infrared image once to obtain a second expanded feature map corresponding to the same multiple expansion as the first feature map;
[0094] Step S43: After concatenating the first enlarged feature map and the second enlarged feature map according to the number of channels, a third enlarged feature map is obtained.
[0095] Step S44: Perform the first convolution operation on the third enlarged feature map once to obtain a high-resolution infrared image.
[0096] In this embodiment, the feature map of the multi-channel image passes through the i-th grouped convolution splicing module, and then performs the Transpose conv_re deconvolution operation twice to enlarge the feature map of the multi-channel image by 4 times to obtain the first enlarged feature map with the dimension of (b, 4*h, 4*w, c); the low-resolution infrared image is upsampled by bicubic interpolation by 4 times to obtain the second enlarged feature map corresponding to the same multiple enlargement of the first feature map with the dimension of (b, 4*h, 4*w, 2*c). The first enlarged feature map and the second enlarged feature map are concatenated according to the number of channels to obtain the third enlarged feature map. Finally, the third enlarged feature map is subjected to the conv_re convolution operation once to obtain a high-resolution infrared image with the dimension of (b, 4*h, 4*w, c), and the height and width of the low-resolution infrared image are each enlarged by 4 times.
[0097] Specifically, in this embodiment, the number of convolution kernels of the deconvolution layer for obtaining the first enlarged feature map enlarged by two times is 12, the number of convolution kernels of the deconvolution layer for obtaining the first enlarged feature map enlarged by four times is 3, and the number of convolution kernels of the convolution layer for obtaining the high-resolution infrared image is 3.
[0098] Specifically, as shown in
[0099]
[0100] Table 4
[0101] In a preferred embodiment, the convolution kernel size of the first convolution operation is 3*3 and the stride is 1.
[0102] Specifically, as Figure 6 shown, the first convolution operation in this embodiment is the conv_re convolution operation, that is, first perform the convolution operation with a 3*3 convolution kernel with a stride of 1, and then pass through the leaky relu activation function.
[0103] In a preferred embodiment, the convolution kernel size of the second convolution operation is 3*3 and the stride is 2.
[0104] Specifically, as Figure 7As shown, the second convolution operation in this embodiment is the conv_down convolution operation, that is, first perform convolution operation with a 3*3 convolution kernel and a stride of 2, and then pass through the leaky relu activation function. The only difference from the conv_re convolution operation is that the stride is 2, so as to achieve the purpose of downsampling.
[0105] In a preferred embodiment, the convolution kernel size of the first expansion operation is 4*4 and the stride is 1.
[0106] Specifically, as Figure 8 shown, the first expansion operation in this embodiment is the Transpose conv_re deconvolution operation, that is, perform deconvolution operation with a 4*4 convolution kernel and a stride of 1, pad 0 to the feature map of the multi-channel image after feature extraction, so that the height and width of the output feature map are doubled, and then pass through the leaky relu activation function.
[0107] In a preferred embodiment, the second expansion operation is the bicubic interpolation upsampling operation.
[0108] Specifically, the magnification factor of the bicubic interpolation upsampling operation is the same as that of the Transpose conv_re deconvolution operation. In this embodiment, the Transpose conv_re deconvolution layer is set to 2, and the input image is enlarged by 4 times, then the bicubic interpolation upsampling operation is also set to be enlarged by 4 times.
[0109] In a preferred embodiment, the training condition associated with the neural network model is to set a loss function for the neural network model;
[0110] The loss function is specifically:
[0111]
[0112] Among them,
[0113] MSE is the loss value calculated by the loss function;
[0114] m is the total number of pixel points of a real infrared image;
[0115] is the value of each pixel point of a real infrared image;
[0116] is the value of each pixel point of the high-resolution infrared image.
[0117] In this embodiment, the loss function used to train this convolutional neural network model is the MSE mean square error loss, and its MSE loss value is used to evaluate the distortion degree between the generated high-resolution image and the original real image collected. The smaller the MSE loss value, the better the accuracy of this convolutional neural network model for the generated high-resolution image. Using this MES loss function to train this convolutional neural network model, after training is completed, call the api to convert the trained model into a model that can run on the embedded chip RK3399Pro platform.
[0118] In summary, a method for improving the resolution of infrared images based on a deep learning convolutional neural network designed by the present invention can expand the resolution of the input image by 2 to N times. This method is used to train a convolutional neural network model including a grouped convolution splicing module. The feature extraction of the input image is mainly implemented in the grouped convolution splicing module, and the upsampling of the input image is achieved through deconvolution operations and bicubic interpolation operations. Other upsampling multiples can be achieved by adjusting the number of deconvolution layers, the size of the deconvolution kernel, the stride size, and the bicubic interpolation upsampling ratio in this convolutional neural network model.
[0119] The above are only preferred embodiments of the present invention, and do not limit the implementation manners and protection scope of the present invention. For those skilled in the art, it should be able to realize that the solutions obtained by equivalent substitution and obvious changes made by using the description and illustrations of the present invention should all be included in the protection scope of the present invention.
Claims
1. A method for improving the resolution of infrared images, applied to an embedded device, characterized in that, Including: Step S1: Input a low-resolution infrared image whose resolution needs to be improved into a trained neural network model; Step S2: Use the neural network model to perform channel expansion on the low-resolution infrared image to form a multi-channel image; Step S3: Use the neural network model to split the feature maps of the multi-channel image according to the number of channels, and process the feature maps of each channel separately and then splice them again to form the feature maps of the processed multi-channel image; Step S4: Use the neural network model to expand the feature maps of the multi-channel image output in Step S3, and form a high-resolution infrared image based on the expanded feature maps and output it; The neural network model consists of a convolutional layer, a transposed convolutional layer, a splicing operation, a bicubic interpolation operation, a split channel grouping operation, and a Pixel shuffle operation. Among them, two transposed convolutional layers are set to perform 2 times of transposed convolutional upsampling operations, and the bicubic interpolation operation is set to 4 times; Step S4 includes: Step S41: Perform a first expansion operation on the feature maps of the multi-channel image multiple times to obtain an expanded first expanded feature map; Step S42: Perform bicubic interpolation upsampling on the low-resolution infrared image by 4 times to obtain a second expanded feature map that is expanded by the same multiple as the first expanded feature map; Step S43: Perform splicing processing on the first expanded feature map and the second expanded feature map according to the number of channels to obtain a third expanded feature map; Step S44: Perform a first convolution operation on the third expanded feature map once to obtain the high-resolution infrared image.
2. The method for enhancing the resolution of an infrared image according to claim 1, wherein In Step S1, a pre-prepared training sample set is used to train the neural network model so that the loss function of the neural network model reaches the training conditions, thereby completing the training process of the neural network model.
3. A method for improving the resolution of infrared images according to claim 1, characterized in that, In Step S2, perform a first convolution operation on the low-resolution infrared image once to correspondingly expand the number of channels of the low-resolution infrared image according to the size of the convolution kernel to form a multi-channel image.
4. A method for enhancing the resolution of an infrared image according to claim 3, characterized in that, In Step S3, the process of processing the feature maps of one channel of the multi-channel image specifically includes: Step S31: Perform the first convolution operation on the feature maps of the one channel once to obtain a first feature map; Step S32: Perform a second convolution operation on the first feature map once to obtain a second feature map; Step S33: Perform the first convolution operation on the second feature map once to obtain a third feature map; Step S34: Perform an upsampling operation on the third feature map once to obtain a fourth feature map; Step S35: Perform splicing processing on the first feature map and the fourth feature map to obtain a fifth feature map; Step S36: Perform the first convolution operation on the fifth feature map once to obtain a sixth feature map; Step S37: Perform splicing processing on the fourth feature map and the sixth feature map to obtain a seventh feature map; Step S38: Perform the first convolution operation on the seventh feature map once to obtain an eighth feature map; Step S39: Concatenate the eighth feature map with the eighth feature maps of other channels to obtain the feature map of the multi-channel image.
5. A method for enhancing the resolution of an infrared image according to claim 3, characterized in that The convolution kernel size of the first convolution operation is 3*3, and the stride is 1.
6. A method for enhancing the resolution of an infrared image according to claim 4, characterized in that, The convolution kernel size of the second convolution operation is 3*3, and the stride is 2.
7. A method for improving the resolution of an infrared image according to claim 1, characterized in that, The convolution kernel size of the first expansion operation is 4*4, and the stride is 1.
8. A method for improving the resolution of infrared images according to claim 1, characterized in that, The second expansion operation is a bicubic interpolation upsampling operation.
9. A method for enhancing the resolution of an infrared image according to claim 2, characterized in that, The training condition associated with the neural network model is to set a loss function for the neural network model; The loss function is specifically: ; Wherein, The loss value calculated for the loss function; m is the total number of pixel points of a real infrared image; is the value of each pixel point of a real infrared image; is the value of each pixel point of the high-resolution infrared image.
Citation Information
Patent Citations
A super-resolution reconstruction method based on feature fusion of dual-channel convolution network
CN109509149A