Thermal Image Reconstruction Method, Network Training Method, Device, Equipment and Medium
By generating and using adjusted thermal image samples, training thermal image reconstruction networks is solved, and the problem of limited performance of thermal image super-resolution models in the prior art is achieved, achieving more efficient thermal image reconstruction and detail enhancement.
Patent Information
- Application Number
- CN202210619200.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-06-01
AI Technical Summary
Existing thermal image super-resolution models have limited performance because the training data comes from different thermal imaging sensors.
By obtaining low-resolution and high-resolution thermal image samples, adjusting them separately to generate new samples for training the thermal image reconstruction network to be trained. The network is trained by comparing the losses of different sample thermal images and predicting the losses between thermal images and real thermal images.
Improves the adaptability of thermal image reconstruction networks to poorly registered data and enhances the performance of super-resolution thermal image reconstruction, including improving image resolution and detailed information.
Smart Images

Figure CN114936969B_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of computer vision technology, and in particular, to a thermal image reconstruction method, a network training method, a device, a device and a medium. Background Art
[0002] In recent years, thermal imaging super-resolution technology based on deep learning has gradually developed and made certain progress. Some of these progress are for low-resolution to high-resolution data of structures, and some are for real thermal imaging super-resolution scenarios. In real thermal imaging super-resolution scenarios, since the training data comes from thermal imaging sensors with different camera parameters, the performance of the thermal image super-resolution model is limited. Summary of the Invention
[0003] In view of this, at least one thermal image reconstruction method, network training method, device, device and medium are provided in the embodiments of this application.
[0004] The technical solutions in the embodiments of this application are implemented as follows:
[0005] On the one hand, the embodiments of this application provide a thermal image reconstruction method, and the method includes:
[0006] Obtain a first sample thermal image with a first resolution and a second sample thermal image with a second resolution corresponding to the first sample thermal image; wherein, the second resolution is higher than the first resolution;
[0007] Adjust the first sample thermal image and the second sample thermal image respectively to obtain a third sample thermal image and a fourth sample thermal image;
[0008] In a to-be-trained thermal image reconstruction network, determine a first loss for comparing different sample thermal images based on the first sample thermal image, the third sample thermal image and the fourth sample thermal image;
[0009] Determine a second loss based on the second sample thermal image and a predicted thermal image; wherein, the predicted thermal image is the thermal image of the first sample thermal image predicted by the to-be-trained thermal image reconstruction network at the second resolution;
[0010] Train the to-be-trained thermal image reconstruction network based on the first loss and the second loss, so that the loss output by the trained thermal image reconstruction network meets a preset convergence condition.
[0011] On the other hand, the embodiments of this application provide an image reconstruction method, and the method includes:
[0012] Obtain a first thermal image to be super-resolution reconstructed;
[0013] In the trained thermal image reconstruction network, adjust the dimension of the first thermal image to obtain an adjusted first thermal image; wherein, the trained thermal image reconstruction network is trained by the method described in the above-mentioned one aspect;
[0014] Extract features from the adjusted first thermal image to obtain image features to be reconstructed;
[0015] Based on the super-resolution in the trained thermal image reconstruction network, perform super-resolution reconstruction on the image features to be reconstructed to obtain a reconstructed super-resolution thermal image;
[0016] Based on the reconstructed super-resolution thermal image and the adjusted first thermal image, determine a second thermal image with a resolution higher than that of the first thermal image.
[0017] On the other hand, an embodiment of the present application provides a training device for a thermal image reconstruction network, including:
[0018] A first acquisition module, configured to acquire a first sample thermal image with a first resolution and a second sample thermal image with a second resolution corresponding to the first sample thermal image; wherein, the second resolution is higher than the first resolution;
[0019] A first adjustment module, configured to adjust the first sample thermal image and the second sample thermal image respectively to obtain a third sample thermal image and a fourth sample thermal image;
[0020] A first determination module, configured to determine a first loss for comparing different sample thermal images in a thermal image reconstruction network to be trained based on the first sample thermal image, the third sample thermal image, and the fourth sample thermal image;
[0021] A second determination module, configured to determine a second loss based on the second sample thermal image and a predicted thermal image; wherein, the predicted thermal image is a thermal image of the first sample thermal image predicted by the thermal image reconstruction network to be trained at the second resolution;
[0022] A first training module, configured to train the thermal image reconstruction network to be trained based on the first loss and the second loss, so that the loss output by the trained thermal image reconstruction network meets a preset convergence condition.
[0023] On the other hand, an embodiment of the present application provides a thermal image reconstruction device, and the device includes:
[0024] An acquisition module, configured to acquire a first thermal image to be super-resolution reconstructed;
[0025] An adjustment module, configured to adjust the dimension of the first thermal image in a trained thermal image reconstruction network to obtain an adjusted first thermal image; wherein, the trained thermal image reconstruction network is trained by the method described in one aspect above.
[0026] An extraction module, configured to extract features from the adjusted first thermal image to obtain image features to be reconstructed.
[0027] A reconstruction module, configured to perform super-resolution reconstruction on the image features to be reconstructed based on the super-resolution in the trained thermal image reconstruction network to obtain a reconstructed super-resolution thermal image.
[0028] A determination module, configured to determine a second thermal image with a resolution higher than that of the first thermal image based on the reconstructed super-resolution thermal image and the adjusted first thermal image.
[0029] In another aspect, an embodiment of the present application provides a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.
[0030] In yet another aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method.
[0031] In yet another aspect, an embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.
[0032] In yet another aspect, an embodiment of the present application provides a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above method.
[0033] In the embodiments of the present application, after obtaining a first sample thermal image with low resolution and a second sample thermal image with super resolution corresponding to the first sample thermal image, by adjusting the first sample thermal image and the second sample thermal image, a third sample thermal image and multiple frames of fourth sample thermal images are obtained; in this way, a third sample thermal image that is not pixel-level corresponding to the first sample thermal image and a fourth sample thermal image that is not pixel-level corresponding to the second sample thermal image can be introduced, so as to enrich the input samples. Then, the first sample thermal image, the second sample thermal image, the third sample thermal image, and the fourth sample thermal images are input into the thermal image reconstruction network to be trained, and a first loss for comparing different sample thermal images is determined; in this way, by comparing the first sample thermal image and the third sample thermal image as positive samples with the second sample thermal image and the fourth sample thermal images as negative samples, a first loss that can represent the difference between different thermal images is obtained. Moreover, a second loss is determined through the second sample thermal image and the predicted thermal image output by the thermal image reconstruction network to be trained; the predicted thermal image is obtained by performing super resolution prediction on the first sample thermal image, so that the second loss can represent the similarity between the real super resolution thermal image, that is, the second sample thermal image, and the predicted thermal image; finally, the thermal image reconstruction network to be trained is trained through the first loss and the second loss to obtain a trained thermal image reconstruction network; in this way, the thermal image reconstruction network to be trained is trained through the first loss for comparing sample thermal images with non-corresponding pixels, so that the trained thermal image reconstruction network can perform contrast learning on different sample thermal images, thereby improving the adaptability of the trained thermal image reconstruction network to poorly registered data; and the thermal image reconstruction network to be trained is trained through the second loss between the second sample thermal image and the predicted thermal image. Since the difference between the second sample thermal image and the predicted thermal image is considered in the second loss, more image features can be introduced during the training process, so that the thermal image reconstruction network to be trained learns more image features, and further improves the performance of the trained thermal image reconstruction network in performing super resolution thermal image reconstruction.
[0034] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to explain the technical solutions of the present application.
[0036] Figure 1 It is a schematic implementation flowchart of a method for training a thermal image reconstruction network provided by an embodiment of the present application;
[0037] Figure 2Another implementation process schematic diagram of a training method for a thermal image reconstruction network provided by an embodiment of the present application;
[0038] Figure 3A Another implementation process schematic diagram of a training method for a thermal image reconstruction network provided by an embodiment of the present application;
[0039] Figure 3B Implementation process schematic diagram of a thermal image reconstruction method provided by an embodiment of the present application;
[0040] Figure 4 Another implementation process schematic diagram of a training method for a thermal image reconstruction network provided by an embodiment of the present application;
[0041] Figure 5 Schematic diagram of the implementation framework of a thermal image reconstruction network provided by an embodiment of the present application;
[0042] Figure 6 Another schematic diagram of the implementation framework of a thermal image reconstruction network provided by an embodiment of the present application;
[0043] Figure 7 Schematic diagram of the composition structure of a training device for a thermal image reconstruction network provided by an embodiment of the present application;
[0044] Figure 8 Schematic diagram of the composition structure of a thermal image reconstruction device provided by an embodiment of the present application;
[0045] Figure 9 Schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0046] In order to make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be further elaborated in detail below in conjunction with the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0047] In the following descriptions, reference is made to "some embodiments" which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing this application and are not intended to limit this application.
[0049] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are explained first. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.
[0050] 1) Computer vision refers to using cameras and computers to replace human eyes for machine vision such as target recognition, tracking, and measurement, and further performing graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instruments for detection.
[0051] 2) Deep learning is a machine learning method. As an artificial neural network, it can independently construct (train) basic rules based on example data during the learning process. Especially in the field of machine vision, neural networks usually adopt the method of supervised learning for training, that is, training through example data and predefined results of the example data.
[0052] 3) Super-resolution is to improve the resolution of the original image by hardware or software methods. The process of obtaining a high-resolution image through a series of low-resolution images is super-resolution reconstruction. Super-resolution reconstruction: exchanging time bandwidth (acquiring a multi-frame image sequence of the same scene) for spatial resolution to achieve the conversion of time resolution to spatial resolution.
[0053] The embodiments of the present application provide a training method for a thermal image reconstruction network, and this method can be executed by the processor of a computer device. Among them, the computer device may refer to devices with data processing capabilities such as servers, laptop computers, tablet computers, desktop computers, smart TVs, and mobile devices (such as mobile phones, portable video players, personal digital assistants, dedicated messaging devices, portable game devices). Figure 1 It is a schematic diagram of the implementation process of a training method for a thermal image reconstruction network provided by the embodiments of the present application, as Figure 1 shown, this method includes the following steps S101 to step S104:
[0054] Step S101, obtain a first sample thermal image with a first resolution and a second sample thermal image with a second resolution corresponding to the first sample thermal image.
[0055] In some embodiments, the first resolution is less than the second resolution; if the first resolution is regarded as the low resolution, then the second resolution can be the high resolution obtained by super-resolution reconstruction of the first resolution. The first sample thermal image can be a thermal image with a relatively low resolution acquired for a preset scene, that is, the number of pixels per inch in the first sample thermal image is small. For example, the number of pixels per inch is less than or equal to 100, that is, the first resolution is less than or equal to 100 Dots Per Inch (DPI). For example, the first resolution of the first sample thermal image is 72 DPI. The preset scene in the first sample thermal image can be any scene. For example, in a road traffic scene, a thermal image of the road is acquired as the first sample thermal image; or, in a campus scene, a thermal image of a teaching building is acquired as the first sample thermal image; or, in a wild scene, a thermal image of a forest is acquired as the first sample thermal image, etc. In the process of obtaining the first sample thermal image, it can be a thermal image acquired by a sensor with the first resolution, or a thermal image with the first resolution received from other devices.
[0056] In some embodiments, the second resolution can be any resolution greater than the first resolution. The second resolution can be a relatively high resolution, that is, a resolution that makes the thermal image clearer. In some possible implementation manners, the second resolution can be a resolution greater than a certain threshold, and it can be set that the second resolution is much greater than the maximum value of the first resolution. For example, if the maximum value of the first resolution is 100 DPI, then the second resolution can be set to be greater than or equal to 200 DPI. It can also be set according to the image clarity. In some possible implementation manners, since the pixels cannot be distinguished by the naked eye when the image resolution is 300 DPI, the second resolution can be set to be greater than or equal to 300 DPI; in this way, the resolution of the second sample thermal image is sufficient to meet the user's requirement for clarity.
[0057] In some possible implementation manners, the first sample thermal image and the second sample thermal image are images acquired for the same preset scene, that is, the first sample thermal image and the second sample thermal image are thermal images with the same picture content but different resolutions. The first sample thermal image and the second sample thermal image can have the same size, or different sizes but the same picture content. The second sample thermal image can be acquired by a sensor with the second resolution, or a thermal image with the second resolution received from other devices. The second sample thermal image can be understood as the true value of the first sample thermal image at the second resolution. The second sample thermal image can also be understood as the true value thermal image of the super-resolution corresponding to the low-resolution first sample thermal image.
[0058] Step S102: Adjust the first sample thermal image and the second sample thermal image respectively to obtain a third sample thermal image and a fourth sample thermal image.
[0059] In some embodiments, randomly adjust the first sample thermal image to obtain a third sample thermal image, and randomly adjust the second sample thermal image to obtain a fourth sample thermal image; here, the adjustment method for the first sample thermal image is the same as that for the second sample thermal image.
[0060] In some possible implementation manners, the way to obtain the third sample thermal image and the fourth sample thermal image can be to randomly rotate or translate the thermal image. Taking the random rotation or translation of the first sample thermal image as an example: First, obtain a rotation matrix according to the angle to be rotated and the set rotation center; then, perform an affine transformation on the first sample thermal image according to the rotation matrix to obtain a third sample thermal image at any angle and any center; it is also possible to first define a translation matrix for the first sample thermal image, specify the translation amounts in the x direction and the y direction respectively, and then translate the first sample thermal image according to the translation amounts to obtain the third sample thermal image. In this way, the size and the content of the third sample thermal image are the same as those of the first sample thermal image, but the pixel positions in the third sample thermal image are different from those in the first sample thermal image, that is, the pixel levels of the first sample thermal image and the third sample thermal image do not correspond. Similarly, the size and the objects in the fourth sample thermal image are the same as those in the second sample thermal image, but the pixel levels of the fourth sample thermal image and the second sample thermal image do not correspond.
[0061] Step S103: In the thermal image reconstruction network to be trained, determine a first loss for comparing different sample thermal images based on the first sample thermal image, the third sample thermal image, and the fourth sample thermal image.
[0062] In some embodiments, after adjusting the first sample thermal image and the second sample thermal image, input the obtained third sample thermal image, the fourth sample thermal image, and the first sample thermal image into the thermal image reconstruction network to be trained. In the thermal image reconstruction network to be trained, input the first sample thermal image and the third sample thermal image into the encoder of the thermal image reconstruction network to be trained respectively, and input the second sample thermal image and the fourth sample thermal image into the auxiliary encoder for feature extraction. Based on the extracted image features, determine the similarity between the first sample thermal image and the third sample thermal image, and the similarity between the first sample thermal image and the fourth sample thermal image, and determine the first loss based on these two types of similarities. In this way, the first loss is used for contrastive learning of different sample thermal images, so that the thermal image reconstruction network to be trained can learn the image features of thermal images from different sensors.
[0063] In some possible implementation manners, since the multi-frame third-sample thermal images are obtained after adjusting the first-sample thermal image, the third-sample thermal images can be regarded as positive samples that are very similar to the first-sample thermal image. Therefore, it is desired that the image features of the first-sample thermal image and the image features of the third-sample thermal image are as similar as possible. Since the second-sample thermal image is the ground-truth image of the first-sample thermal image at the second resolution, that is, the real image under super-resolution. Since the fourth-sample thermal image is obtained after adjusting the second-sample thermal image, the fourth-sample thermal image is similar to the second-sample thermal image as the ground-truth. Therefore, the fourth-sample thermal image can also be regarded as coming from a different sensor from the first-sample thermal image. Thus, the fourth-sample thermal image can be used as the negative sample corresponding to the first-sample thermal image. Therefore, it is desired that the distance between the image features of the fourth-sample thermal image and the image features of the first-sample thermal image is as large as possible. Based on this, the first loss is optimized so that the distance between the image features of the fourth-sample thermal image and the image features of the first-sample thermal image is large, and the distance between the image features of the third-sample thermal image and the image features of the first-sample thermal image is small.
[0064] Step S104, determine a second loss based on the second-sample thermal image and the predicted thermal image.
[0065] In some embodiments, the predicted thermal image is the thermal image of the first-sample thermal image at the second resolution predicted by the to-be-trained thermal image reconstruction network. This second loss is used to characterize the similarity between the second-sample thermal image and the predicted thermal image. Since the second-sample thermal image is the ground-truth image of the first-sample thermal image at the second resolution, and this predicted thermal image is the thermal image of the first-sample thermal image at the second resolution predicted by the to-be-trained thermal image reconstruction network, in this way, by comparing the second-sample thermal image and the predicted thermal image, the second loss characterizing the difference between the ground-truth and the predicted value can be obtained.
[0066] In some possible implementation manners, this second loss may include the difference between the second-sample thermal image and the predicted thermal image, as well as the similarity between the image features of the second-sample thermal image and the image features of the predicted thermal image. In this way, by introducing the image features of the second-sample thermal image as the ground-truth and the predicted thermal image as the predicted value into the second loss, more detailed features can be involved in the second loss.
[0067] Step S105, train the to-be-trained thermal image reconstruction network based on the first loss and the second loss, so that the loss output by the trained thermal image reconstruction network meets a preset convergence condition.
[0068] In some embodiments, by combining the first loss and the second loss as the total loss, the network parameters of the to-be-trained thermal image reconstruction network are adjusted so that the loss output by the trained thermal image reconstruction network meets a preset convergence condition. The preset convergence condition can be set based on the requirements of the to-be-trained thermal image reconstruction network for the loss. For example, for the first loss, since it is desired that the distance between the image features of the fourth sample thermal image and the image features of the first sample thermal image in the first loss is large, and the distance between the image features of the third sample thermal image and the image features of the first sample thermal image is small, the preset convergence condition may include a first distance threshold and a second distance threshold, where the first distance threshold is greater than the second distance threshold. If the distance between the image features of the fourth sample thermal image and the image features of the first sample thermal image is greater than or equal to the first distance threshold, and the distance between the image features of the third sample thermal image and the image features of the first sample thermal image is less than the second distance threshold, then it is considered that the first loss meets the preset convergence condition. In some possible implementation manners, the first distance threshold and the second distance threshold can be set based on the requirements of the thermal image reconstruction network for performance. For example, the first distance threshold and the second distance threshold are set according to the accuracy of the thermal image reconstruction network.
[0069] In some embodiments, for the second loss, since the second loss represents the similarity between the true value and the predicted value, the preset convergence condition may include a similarity threshold. If the similarity in the second loss is greater than or equal to the similarity threshold, it indicates that the second loss meets the preset convergence condition. In some possible implementation manners, the similarity threshold can also be set according to the requirements of the thermal image reconstruction network for performance.
[0070] In some possible implementation manners, the network parameters of each module in the to-be-trained thermal image reconstruction network are adjusted by the first loss and the second loss together. Among them, the encoder in the to-be-trained thermal image reconstruction network is optimized by the first loss, and the network parameters of the entire to-be-trained thermal image reconstruction network are optimized by the second loss to obtain the trained thermal image reconstruction network.
[0071] In the embodiments of the present application, after obtaining the first sample thermal image with low resolution and the second sample thermal image with super resolution corresponding to the first sample thermal image, a third sample thermal image and a fourth sample thermal image are introduced. By being able to introduce a third sample thermal image that is not pixel-level corresponding to the first sample thermal image, and a fourth sample thermal image that is not pixel-level corresponding to the second sample thermal image, the input samples can be enriched. Then, by comparing the first sample thermal image and the third sample thermal image as positive samples with the second sample thermal image and the fourth sample thermal image as negative samples, a first loss that can represent the difference between different thermal images is obtained. And, a second loss is determined through the second sample thermal image and the predicted thermal image output by the thermal image reconstruction network to be trained; the predicted thermal image is obtained by performing super resolution prediction on the first sample thermal image. In this way, the second loss can represent the similarity between the true super resolution thermal image, that is, the second sample thermal image, and the predicted thermal image. Finally, the thermal image reconstruction network to be trained is trained through the first loss and the second loss to obtain a trained thermal image reconstruction network. In this way, by training the thermal image reconstruction network to be trained through the first loss for comparing different sample thermal images, the trained thermal image reconstruction network can perform contrastive learning on sample thermal images that are not pixel-level corresponding, thereby improving the adaptability of the trained thermal image reconstruction network to poorly registered data. Moreover, by training the thermal image reconstruction network to be trained through the second loss between the second sample thermal image and the predicted thermal image, since the difference between the second sample thermal image and the predicted thermal image is considered in the second loss, more image features can be introduced during the training process, enabling the thermal image reconstruction network to be trained to learn more image features, and further improving the performance of the trained thermal image reconstruction network for super resolution thermal image reconstruction.
[0072] In some embodiments, by rotating or translating the first sample thermal image and the second sample thermal image to obtain a third sample thermal image that is not pixel-level corresponding to the first sample thermal image, and a fourth sample thermal image that is not pixel-level corresponding to the second sample thermal image, the positive samples and negative samples in the training process can be enriched. That is, the above step S102 can be implemented through the following steps S121 and S122 (not shown in the figure):
[0073] Step S121, perform at least one rotation or translation on the first sample thermal image to obtain at least one frame of third sample thermal image.
[0074] In some embodiments, after collecting the first sample thermal image, in the first sample thermal image, the first sample thermal image is rotated or translated, or rotated first and then translated, etc., to obtain a changed third sample thermal image.
[0075] In some embodiments, in the first sample thermal image, by performing one rotation or translation on the first sample thermal image, a frame of the third sample thermal image is obtained. In this way, the resolution and size of the third sample thermal image are the same as those of the first sample thermal image, and the picture content in the third sample thermal image and the first sample thermal image is the same. However, the pixels in the first sample thermal image and the third sample thermal image do not correspond at the pixel level. During the training process, it is required that the distance between the image features of the third sample thermal image and the first sample thermal image in the first loss is as small as possible.
[0076] Step S122: Perform at least one rotation or translation on the second sample thermal image to obtain at least one frame of the fourth sample thermal image.
[0077] In some embodiments, after the second sample thermal image is collected, in the second sample thermal image, the second sample thermal image is rotated or translated, or rotated first and then translated, etc., to obtain multiple frames of the fourth sample thermal image after the change; the fourth sample thermal image can be three frames or more. Since the first sample thermal image and the second sample thermal image are thermal images with different resolutions collected for the same object, the object in the second sample thermal image is the same as the object in the first sample thermal image. By randomly rotating or translating the second sample thermal image, the simulation of pixel-level non-corresponding data in different images is realized, so as to obtain the fourth sample thermal image and the second sample thermal image with non-corresponding pixels. In this way, the resolution and size of the fourth sample thermal image are the same as those of the second sample thermal image, and the object in the fourth sample thermal image and the second sample thermal image is the same. However, the pixels in the fourth sample thermal image do not correspond to the pixels in the second sample thermal image. In this way, by translating or rotating the second sample thermal image, a fourth sample thermal image with poor registration with the first sample thermal image is obtained. In this way, the fourth sample thermal image can be used as a negative sample of the first sample thermal image. In some possible implementation manners, the number of frames of the fourth sample thermal image is greater than the number of frames of the third sample thermal image. In this way, the number of frames of the third sample thermal image as a positive sample is less than the number of frames of the fourth sample thermal image as a negative sample, so that more negative samples are used for comparison learning, which is more helpful for improving the model's ability to adapt to data with poor registration. During the training process, since the fourth sample thermal image is used as a negative sample of the first sample thermal image, it is required that the distance between the image features of the fourth sample thermal image and the first sample thermal image in the first loss is as large as possible.
[0078] In the above steps S121 and S122, by rotating or translating the object in the first sample thermal image, a third sample thermal image is obtained, thereby enriching the positive samples similar to the first sample thermal image; by rotating or translating the object in the second sample thermal image, a fourth sample thermal image is obtained, thereby enriching the negative samples different from the first sample thermal image. In this way, by rotating or translating the first sample thermal image and the second sample thermal image, at least one frame of the third sample thermal image that is pixel-level non-corresponding to the first sample thermal image and multiple frames of the fourth sample thermal image that are pixel-level non-corresponding to the second sample thermal image are obtained, simulating a scenario where two cameras cannot completely overlap during simultaneous shooting, resulting in slightly different shooting perspectives and pixel-level non-correspondence; thus, it is convenient for the network to perform contrastive learning on different samples during the subsequent training process, so that the network can be unaffected by poorly registered data.
[0079] In some embodiments, by using different encoders to extract features from different samples, the first loss can perform object learning on the extracted image features, that is, the above step S103 can be achieved through Figure 2 the steps shown as follows:
[0080] Step S201, use the first encoder in the thermal image reconstruction network to be trained to extract features from the first sample thermal image and the third sample thermal image, obtaining a first image feature and a third image feature.
[0081] In some embodiments, the first encoder is used to extract features from low-resolution thermal images, that is, the first encoder is used to extract features from thermal images of the first resolution. The first encoder includes a convolution module and a downsampling module. Among them, the convolution module may include at least one convolution layer (for example, one) for extracting features from the input thermal image, and the downsampling module is used to downsample the image features output by the convolution module and continue to extract features, thereby obtaining the final image features.
[0082] In some possible implementation manners, in the thermal image reconstruction network to be trained, the first encoder may be connected to the input module of the network. In this way, after the sample thermal images of the first resolution (for example, the first sample thermal image and the third sample thermal image) are input from the input module, they can enter the first encoder; before the first encoder, a bicubic interpolation function may also be connected. The bicubic interpolation function is used to perform bicubic interpolation on the sample thermal images input by the input module to increase the dimension of the sample thermal images, and then the sample thermal images with increased dimensions are input into the encoder for feature extraction.
[0083] Step S202, use a second encoder that does not share weights with the first encoder to extract features from the fourth sample thermal image, obtaining a second image feature and a fourth image feature.
[0084] In some embodiments, the second encoder has the same structure as the first encoder, but does not share weights with the first encoder. Thus, the second encoder can be regarded as an auxiliary encoder for feature extraction of the second sample thermal image and the fourth sample thermal image. The implementation process of the second encoder for feature extraction of the fourth sample thermal image is the same as that of the first encoder for feature extraction of the first sample thermal image and the third sample thermal image, that is, feature extraction is performed on the input thermal image through the convolutional module in the second encoder, and then, the image features output by the convolutional module are downsampled and further feature extraction is performed through the downsampling module to obtain the final image features.
[0085] Step S203: Determine the first loss based on the first image feature, the third image feature, and the fourth image feature.
[0086] In some embodiments, the first loss is obtained by determining the similarities between the first image feature and the third image feature and between the first image feature and the fourth image feature respectively, and then comparing the two similarities. That is, the similarity between the first image feature and the third image feature, and the similarities between the first image feature and multiple fourth image features are analyzed. By determining the difference between the two similarities, the first loss is obtained. In this way, the first loss can characterize the differences between sample thermal images from different sources, enabling the network to perform contrastive learning on sample thermal images from different sources.
[0087] In the embodiments of the present application, by using different encoders to perform feature extraction on the first sample thermal image, the third sample thermal image, and the fourth sample thermal image respectively, the fourth image feature is independent of the first image feature and the third image feature, reducing the interference between data from different sources; and, by combining the first image feature with the third image feature and the fourth image feature respectively to determine the first loss for contrastive learning, the first loss can enable the to-be-trained thermal image reconstruction network to perform contrastive learning on sample thermal images from different sources, improving the adaptability of the network.
[0088] In some embodiments, the first loss is constructed by comparing image features with different resolutions, that is, the above step S203 can be implemented through the following steps S231 to S233 (not shown in the figure):
[0089] Step S231: Determine the first similarity between the first image feature and the third image feature.
[0090] In some embodiments, the first similarity may be the sum of one or more similarities. If the third image feature is one, then the first similarity is one; if the third image feature is multiple, then the first similarity is the sum of the similarities between the first image feature and each third image feature.
[0091] In some possible implementation manners, the first similarity may be the feature distance between the first image feature and the third image feature. For example, the cosine distance between two feature vectors. When the third image feature is one, it may also be obtained by first multiplying the first image feature and the third image feature, using the multiplication result as an exponent, with e as the base, and using such a result as the first similarity between the two features. Since the first image feature is derived from the first sample thermal image, the third image feature is derived from the third sample thermal image, and the third sample thermal image is obtained by rotating or translating the first sample thermal image, it can be understood that the first image feature and the third image feature have similar sources, so the larger the first similarity, the better.
[0092] Step S232, determine the second similarity between the first image feature and the fourth image feature.
[0093] In some embodiments, the second similarity may be the sum of one or more similarities. If the fourth image feature is one, then the second similarity is one; if the fourth image feature is multiple, then the first similarity is the sum of the similarities between the first image feature and each fourth image feature.
[0094] In some possible implementation manners, the determination manner of the second similarity is the same as that of the first similarity. That is, the second similarity may be characterized by the feature distance between the first image feature and the fourth image feature. The larger this feature distance is, the smaller the similarity between the two features is. It may also be represented by the exponent of the first image feature and the fourth image feature to represent the second similarity. Since the fourth image feature is derived from the fourth sample thermal image, and the fourth sample thermal image is obtained by rotating or translating the second sample thermal image, and the second sample thermal image and the first sample thermal image have different sources, it can be understood that the first image feature and the fourth image feature have dissimilar sources, so the smaller the second similarity, the better.
[0095] Step S233, compare the first similarity and the second similarity to obtain the first loss.
[0096] In some embodiments, the first similarity and the second similarity are represented in the same form. For example, both the first similarity and the second similarity are represented by a feature distance. By determining the ratio of the two similarities, a first loss is obtained. If both the first similarity and the second similarity are represented by an exponent, then the logarithm of the ratio of the two similarities is taken as the first loss, so that the image features of different resolutions can be fully compared in the first loss.
[0097] In the above steps S231 to S233, the similarity between the first sample thermal image and the third sample thermal image with the same resolution, and the similarity between the first sample thermal image and the fourth sample thermal image with different resolutions are represented by determining the first similarity and the second similarity between the first image feature and the third image feature and the fourth image feature respectively; in this way, by comparing and learning the unregistered sample thermal images, the network can be made not to be affected by the poorly registered sample thermal images.
[0098] In some embodiments, before the encoder extracts features from the thermal image, the dimension of the thermal image is adjusted so that the dimension of the thermal image input to the encoder is higher, that is, before the above step S201, there is also a step S211 (not shown in the figure):
[0099] Step S211, adjust the dimensions of the first sample thermal image and the third sample thermal image respectively to obtain an adjusted first sample thermal image and an adjusted third sample thermal image.
[0100] In some embodiments, in the thermal image reconstruction network to be trained, after the first sample thermal image and the third sample thermal image are input from the input module, they enter the bicubic interpolation module. The bicubic interpolation module can be implemented by a bicubic interpolation function. The first sample thermal image and the third sample thermal image are bicubically interpolated by the bicubic interpolation function to increase the dimensions of the first sample thermal image and the third sample thermal image. In this way, the adjusted first sample thermal image and the adjusted third sample thermal image output by the bicubic interpolation module have the same size as the first sample thermal image, but the pixel dimension in the image is greater than the pixel dimension of the first sample thermal image.
[0101] After step S211, the encoder extracts features from the dimension-adjusted thermal image, which can be implemented by the following step S212 (not shown in the figure):
[0102] Step S212, use the first encoder to extract features from the adjusted first sample thermal image and the adjusted third sample thermal image respectively to obtain the first image feature and the third image feature.
[0103] In some embodiments, in the first encoder, feature extraction is first performed on the modulated first sample thermal image and the modulated third sample thermal image, and then the image features are downsampled and feature extraction is continued to obtain the first image features and the third image features. In this way, by first dimensionality increasing the input first sample thermal image and third sample thermal image, and using the first encoder to perform feature extraction on the dimensionality-increased modulated first sample thermal image and modulated third sample thermal image, the dimensionality of the extracted image features is higher, that is, the resolution of the image features is higher, which is beneficial to subsequent super-resolution reconstruction of the image features.
[0104] In some possible implementation manners, feature extraction of the input sample thermal image is implemented through the interaction of each module in the first encoder, that is, the above step S212 can be implemented through the following steps:
[0105] First step, use the convolutional module in the first encoder to perform feature extraction on the modulated first sample thermal image and the modulated third sample thermal image to obtain first candidate features and second candidate features.
[0106] In some embodiments, the convolutional module in the first encoder can be implemented by two convolutional layers (for example, two connected 3×3 convolutional layers). The modulated first sample thermal image and the modulated third sample thermal image are input into the first convolutional layer for feature extraction, and the extracted features are output to the second convolutional layer for feature extraction again to obtain the first candidate features and the second candidate features.
[0107] Second step, use the downsampling module in the first encoder to downsample the first candidate features and the second candidate features to obtain the first image features and the third image features.
[0108] In some embodiments, the downsampling module can be at least one. For example, the first encoder includes three connected downsampling modules; wherein, the structure of each downsampling module is the same and can be a convolutional network structure composed of residual blocks. The convolutional module inputs the first candidate features and the second candidate features into the first downsampling module for downsampling and feature extraction again. Then, the first downsampling module inputs the result into the second downsampling module for downsampling and feature extraction again. The second downsampling module inputs the result into the third downsampling module for downsampling and feature extraction again. The finally obtained image features are the first image features and the third image features. In this way, by performing feature extraction on the dimensionality-increased modulated first sample thermal image and modulated third sample thermal image through convolutional layers, and downsampling the extracted candidate features through the sampling module, the first image features and the third image features obtained have a higher feature dimension and a smaller feature size, realizing that the small-sized features include rich high-dimensional features.
[0109] In some embodiments, after the first encoder extracts features from the first sample thermal image, it is continuously input into the decoder of the network to achieve super-resolution reconstruction of the first sample thermal image, thereby obtaining a predicted super-resolution image, that is, a preset thermal image. This can be achieved through the following steps:
[0110] First step, in the decoder corresponding to the first encoder, perform super-resolution reconstruction on the first image features of the first sample thermal image to obtain a reconstructed thermal image.
[0111] In some embodiments, the decoder corresponding to the first encoder can be understood as a decoder with a structure symmetric to that of the first encoder, that is, a decoder used to decode the output result of the first encoder. According to the second resolution, the decoder first upsamples the input first image features, and then reconstructs the upsampled features through a convolution module to obtain a reconstructed thermal image.
[0112] Second step, fuse the reconstructed thermal image and the adjusted first sample thermal image corresponding to the first sample thermal image to obtain the predicted thermal image.
[0113] In some embodiments, based on the reconstructed thermal image, combine the original upsampled and adjusted first sample thermal image to obtain the predicted thermal image. For example, supplement the reconstructed thermal image on the basis of the upsampled and adjusted first sample thermal image to make the resolution in the predicted thermal image higher and as close as possible to the second resolution, that is, make the predicted thermal image as similar as possible to the second sample thermal image.
[0114] In the above first step and second step, through the decoder matched with the first encoder, perform super-resolution reconstruction on the image features output by the encoder, and combine the upsampled and adjusted first sample thermal image with increased dimensions to predict the super-resolved thermal image corresponding to the first sample thermal image, so that the image features in the original first sample thermal image can be fully considered in the predicted thermal image, improving the accuracy of the preset thermal image.
[0115] In some possible implementation manners, decoding of the image features input to the first encoder is achieved through the interaction of each module of the decoder, thereby obtaining a reconstructed thermal image. That is, "in the decoder corresponding to the first encoder, based on the second resolution, perform super-resolution reconstruction on the first image features of the first sample thermal image to obtain a reconstructed thermal image with the second resolution" in the above first step can be achieved through the following steps A and B:
[0116] Step A, use the upsampling module in the decoder to upsample the first image features to obtain upsampled features.
[0117] Here, the first image feature is upsampled by the upsampling module in the decoder, so that the size of the upsampled feature is restored to the size of the image feature output by the convolutional module of the encoder. The upsampling module corresponds to the downsampling module in the encoder. If there are three downsampling modules, then there are also three upsampling modules, which correspond to the downsampling modules one by one. The upsampling module is also a convolutional network structure composed of residual blocks. In this way, if in the encoder, the features extracted by the convolutional module are downsampled three times continuously by three downsampling modules, then in this decoder, upsampling is performed three times continuously according to the same size as the downsampling module, so that the upsampled feature is restored to the size of the image feature output by the convolutional module of the encoder.
[0118] Step B, the upsampled feature is reconstructed by the reconstruction module in the decoder to obtain the reconstructed thermal image.
[0119] Here, the reconstruction module in the decoder can be implemented by at least one convolutional layer. For example, this reconstruction module can also be implemented by two connected 3×3 convolutional layers. In this way, the vector representing the upsampled feature is convolved successively by two connected 3×3 convolutional layers to realize reconstructing the feature into a thermal image. In this way, by using a decoder symmetric to the encoder structure, first upsampling the input first image feature, and then reconstructing the upsampled feature, a super-resolution reconstructed thermal image can be obtained. In this way, since the first image feature is derived from the upsampled first sampled thermal image, the resolution of the first image feature is greater than the first resolution. In this way, by upsampling and reconstructing the high-resolution first image feature, a reconstructed thermal image with a resolution higher than that of the first sampled thermal image can be obtained, and further, the predicted thermal image obtained from this reconstructed thermal image is closer to the second sampled thermal image, so as to improve the accuracy of the network for super-resolution thermal image reconstruction.
[0120] In some embodiments, by analyzing the difference between the predicted thermal image and the second sampled thermal image as the ground truth, and introducing the similarity between the image features in the two types of thermal images to determine the second loss, so that the second loss can introduce more detailed features into the network, that is, the above step S104 can be Figure 3A implemented by the steps shown as follows:
[0121] Step S301, based on the second sampled thermal image and the predicted thermal image, determine the basic loss representing the difference between the images.
[0122] In some embodiments, first, matrices representing the second sample thermal image and the predicted thermal image are respectively determined; then, the difference between the two matrices is determined; in this way, the difference is used as the basic loss characterizing the difference between the images. Thus, the basic loss can represent the similarity between the predicted value and the true value. During the training process, it is desired that the basic loss be as small as possible. The smaller the basic loss, the smaller the difference between the true value and the predicted value, and thus the higher the accuracy of the super-resolution thermal image reconstruction by the thermal image reconstruction network.
[0123] Step S302: Based on the second image features of the second sample thermal image and the predicted features of the predicted thermal image, determine a feature loss characterizing the similarity between the features.
[0124] In some embodiments, first, vectors representing the second image features and vectors representing the predicted features of the predicted thermal image are respectively determined; then, the cosine distance between the two vectors is determined; finally, multiple cosine distances are obtained through multiple second image features and multiple predicted features of the predicted thermal images, and the multiple cosine distances are summed element by element to obtain the feature loss characterizing the similarity between the features. In this way, during the network training process, introducing the feature loss between the second image features as the true value and the predicted features of the predicted value can enhance the detailed features in the network.
[0125] Step S303: Determine the basic loss and the feature loss as the second loss.
[0126] In some embodiments, adding the feature loss on the basis of the basic loss can introduce detailed features into the second loss, thereby realizing the enhancement of the detailed features.
[0127] In the embodiments of the present application, the difference between the predicted value and the true value is represented by the basic loss, and on the basis of the basic loss, the feature loss between the features of the true value and the predicted value is added. Thus, the basic loss and the feature loss are combined to adjust the network parameters of the thermal image reconstruction network, so that the thermal image reconstruction network can learn more detailed features, and further the restored and reconstructed thermal image of the trained thermal image reconstruction network has richer detailed features.
[0128] In some embodiments, the encoder in the network is optimized through the first loss that realizes contrast learning, and the entire network parameters are optimized through the second loss that introduces the feature loss, so as to obtain the trained thermal image reconstruction network. That is, the above step S105 can be implemented through the following steps S151 and S152 (not shown in the figure):
[0129] Step S151: Using the first loss, adjust the structural parameters of the first encoder and the second encoder in the to-be-trained thermal image reconstruction network to obtain an intermediate thermal image reconstruction network.
[0130] In some embodiments, the first encoder and the second encoder have the same structure. The structural parameters include: the weights of the convolutional kernels, the weights of the downsampling modules, and the coding rates of the encoders, etc. By comparing and learning the third sample thermal images and the fourth sample thermal images from different sources, the first loss is obtained. Using this first loss for backpropagation to optimize the structural parameters of the first encoder and the second encoder, so that the first encoder and the second encoder can accurately encode data with poor sample and label registration, thereby enabling the network to adapt to data with poor registration.
[0131] Step S152: Using the second loss, adjust the network parameters of the intermediate thermal image reconstruction network to obtain the trained thermal image reconstruction network.
[0132] In some embodiments, the network parameters of the intermediate thermal image reconstruction network include: weights and learning rates, etc. By combining the basic loss and the feature loss to optimize the network parameters of the entire network. In this way, during the training process, through the backpropagation of the basic loss, the trained thermal image reconstruction network can reconstruct a more accurate super-resolution thermal image, and through the backpropagation of the feature loss, the trained thermal image reconstruction network can reconstruct more detailed information.
[0133] In the above steps S151 and S152, the first loss obtained by comparing and learning sample thermal images with different resolutions is used to enhance the encoders in the to-be-trained thermal image reconstruction network, so that the trained thermal image reconstruction network is not affected by data with poor registration; and a feature loss between the ground truth and the predicted value is introduced in the second loss, and the overall network parameters of the to-be-trained thermal image reconstruction network are optimized by increasing the feature loss, so that the thermal image reconstructed by the trained thermal image reconstruction network has richer details and higher accuracy.
[0134] In some possible implementation manners, by using different sensors to collect thermal images of the same preset object, a first sample thermal image and a second sample thermal image are obtained. That is, the above step S101 can be implemented by the following steps S111 and S112 (not shown in the figure):
[0135] Step S111: Use a first-resolution sensor to collect a thermal image of the preset object to obtain the first sample thermal image.
[0136] In some embodiments, the first-resolution sensor may be a thermal imaging camera with a first resolution. By using this thermal imaging camera to collect thermal images of a preset object, a first sample thermal image with the first resolution is obtained. The preset object may refer to a foreground pattern in the picture. For example, if the first sample thermal image is a thermal image collected in a road scene, then the preset object may be the road or vehicles on the road in the collected thermal image; if the first sample thermal image is a thermal image of a teaching building collected in a campus scene, then the object may be the teaching building in the image; if the first sample thermal image is a thermal image of a person collected, then the preset object may be the person in the image.
[0137] Step S111: Use a second-resolution sensor to collect a thermal image of the preset object to obtain the second sample thermal image.
[0138] In some embodiments, the second-resolution sensor may be a thermal imaging camera with a second resolution, that is, a super-resolution thermal imaging camera. By using this super-resolution thermal imaging camera to collect thermal images of a preset object, a second sample thermal image with the second resolution is obtained. In this way, the second sample thermal image can be used as the super-resolution thermal image reconstructed from the first sample thermal image, that is, the ground truth label of the first sample thermal image at the second resolution. Thus, the first sample thermal image and the second sample thermal image can be regarded as a pair of registered data. In this way, in the subsequent processing, by rotating or translating the objects in the first sample thermal image and the second sample thermal image, a third sample thermal image and a fourth sample thermal image are obtained. Therefore, the first sample thermal image and the fourth sample thermal image can be regarded as unregistered data.
[0139] In the above steps S111 and S112, by using sensors with different resolutions to collect thermal images of the same preset object, two sample thermal images with the same object in the picture but different resolutions can be obtained. In this way, the second sample thermal image can be used as the ground truth label of the first sample thermal image at the second resolution. Thus, the first sample thermal image and the second sample thermal image used as the ground truth label are made to match better.
[0140] An embodiment of the present application provides a thermal image reconstruction method, as Figure 3B shown, Figure 3B is a schematic flowchart of the implementation process of a thermal image reconstruction method provided by an embodiment of the present application. The method includes the following steps S31 to step S35:
[0141] Step S31: Obtain a first thermal image to be super-resolution reconstructed.
[0142] In some embodiments, the first thermal image is the thermal image that needs to be super-resolved. The first thermal image can be a thermal image captured for any scene, which can be a thermal image with a complex background or a thermal image with a simple background. The first thermal image can be captured by a thermal image acquisition device (such as a thermal image camera), or can also be a thermal image received from other devices.
[0143] Step S32: In the trained thermal image reconstruction network, adjust the dimension of the first thermal image to obtain the adjusted first thermal image.
[0144] In some embodiments, the trained thermal image reconstruction network is trained by the training method provided in the above embodiments. The trained thermal image reconstruction network includes: an input module, a module for dimension adjustment (such as a bicubic interpolation module), an encoder, a decoder, and an output module; wherein, the first thermal image to be super-resolved is input into the trained thermal image reconstruction network through the input module of the trained thermal image reconstruction network, and then output from the input module to the dimension adjustment module. The bicubic interpolation function in this module is used to increase the dimension of the first thermal image to obtain an adjusted thermal image with a dimension higher than that of the first thermal image, which further indicates that the resolution of the adjusted thermal image is also higher than that of the first thermal image.
[0145] Step S33: Extract features from the adjusted first thermal image to obtain the features of the image to be reconstructed.
[0146] In some embodiments, the dimension adjustment module of the trained thermal image reconstruction network outputs the adjusted first thermal image to the encoder, and in this encoder, features are extracted from the adjusted first thermal image to obtain the features of this image.
[0147] Step S34: Based on the super-resolution in the trained thermal image reconstruction network, perform super-resolution reconstruction on the features of the image to be reconstructed to obtain the reconstructed super-resolution thermal image.
[0148] In some embodiments, the encoder of the trained thermal image reconstruction network outputs the image features to the decoder, and in this decoder, the reconstructed features are decoded to obtain the reconstructed super-resolution image. In this decoder, upsampling and convolution operations are performed on the input features to be reconstructed, so as to realize the super-resolution reconstruction of the features of the image to be reconstructed and obtain the reconstructed super-resolution thermal image, that is, the reconstructed super-resolution thermal image; in this way, the resolution of the reconstructed super-resolution thermal image is the super-resolution corresponding to the trained thermal image reconstruction network, that is, the second resolution in the above embodiments.
[0149] Step S35: Based on the reconstructed super-resolution thermal image and the adjusted first thermal image, determine the second thermal image with a resolution higher than that of the first thermal image.
[0150] In some embodiments, the second thermal image is a super-resolution thermal image predicted by a trained thermal image reconstruction network for the first thermal image, that is, the objects in the second thermal image are the same as those in the first thermal image, but the resolution of the second thermal image is higher than that of the first thermal image. The decoder outputs the reconstructed super-resolution thermal image after super-resolution reconstruction to the output module, where the reconstructed super-resolution thermal image and the adjusted first thermal image are combined, so that the reconstructed super-resolution thermal image can be optimized by the adjusted first thermal image. For example, the details of the reconstructed super-resolution thermal image are supplemented by the picture content in the adjusted first thermal image, etc., so that the accuracy of the output second thermal image is higher.
[0151] In the embodiments of the present application, for any first thermal image obtained, a trained thermal image reconstruction network is used to perform super-resolution reconstruction on the first thermal image; since during the training process, the trained thermal image reconstruction network can adapt to poorly registered data through contrast learning and can reconstruct rich detail information by introducing feature loss; in this way, the trained thermal image reconstruction network performs super-resolution reconstruction on the features of the image to be reconstructed, and the obtained reconstructed super-resolution thermal image is combined with the adjusted first thermal image, so that the second thermal image reconstructed by super-resolution has more detail information and higher accuracy.
[0152] The following describes the application of the training method of the thermal image reconstruction network provided by the embodiments of the present application in an actual scenario, taking the reconstruction of super-resolution thermal images as an example for illustration.
[0153] The production materials of thermal imaging cameras are very expensive and the production process is very complicated. If you want to improve the resolution of thermal imaging cameras, the cost of improvement at the hardware level is very high. Therefore, thermal imaging super-resolution technology is the key to reducing the cost of high-quality thermal imaging and making it widely available for civilian use.
[0154] In a real thermal imaging super-resolution scenario, the training data usually used comes from different thermal imaging sensors. Due to different camera parameters, the pixels of these training data are not in one-to-one correspondence. Registering the training data is a key step in the training of the thermal image super-resolution model, and the quality of data registration directly affects the performance of the trained model. In this way, if the data cannot be fully registered, the poorly registered data will affect and limit the performance of the thermal image super-resolution model.
[0155] Based on this, an embodiment of the present application provides a training method for a thermal image reconstruction network, which uses a U-shaped super-resolution network (corresponding to the thermal image reconstruction network in the above embodiment) to solve the problem of performance degradation of the super-resolution in the thermal imaging network caused by inter-domain drift. By constructing positive and negative samples and optimizing the encoder part of the network using contrastive learning, the U-shaped super-resolution network has a better adaptation ability to mis-registered data, thereby solving the problem of network performance degradation caused by mis-registered data. And using feature loss to solve the problem of less detailed information in the super-resolution thermal image reconstructed by the network.
[0156] An embodiment of the present application provides a training method for a thermal image reconstruction network, which can be achieved through Figure 4 the steps shown below:
[0157] Step S401: Randomly rotate or translate the low-resolution thermal image to generate an adjusted thermal image.
[0158] Here, the low-resolution thermal image I LR can be the first sample thermal image in the above embodiment; the adjusted thermal image can be the adjusted first sample thermal image in the above embodiment.
[0159] Step S402: Input the low-resolution thermal image and the adjusted thermal image into the encoder of the U-shaped super-resolution network to obtain the encoded low-resolution feature and the adjusted low-resolution feature.
[0160] Here, the encoded low-resolution feature corresponds to the first image feature in the above embodiment and can be represented as F LR ; the adjusted low-resolution feature corresponds to the third image feature in the above embodiment, and the third image feature can be represented as
[0161] While performing steps S401 and S402, the high-resolution thermal image I HR is processed as shown in steps S403 and S404.
[0162] Step S403: Randomly rotate or translate the high-resolution thermal image to generate three frames of adjusted high-resolution thermal images.
[0163] Here, the high-resolution thermal image I HR can be the second sample thermal image in the above embodiment. The three frames of adjusted high-resolution thermal images correspond to the fourth sample thermal image in the above embodiment.
[0164] Step S404: Input the high-resolution thermal image and three adjusted high-resolution thermal images into the auxiliary encoder to obtain the auxiliary-encoded high-resolution image features and three adjusted high-resolution features.
[0165] Here, input the high-resolution thermal image I HR and three adjusted high-resolution thermal images into the auxiliary encoder to obtain the high-resolution image features and three adjusted high-resolution features, denoted as F HR ,
[0166] Step S405: Determine the comparison loss using the low-resolution features, adjusted low-resolution features, and three adjusted high-resolution features.
[0167] Here, use the low-resolution feature F LR , the adjusted low-resolution feature and three adjusted high-resolution features to determine the comparison loss L con .
[0168] Step S406: Backpropagate the comparison loss to adjust the encoder of the U-shaped super-resolution network.
[0169] After step S402, continue to perform decoding processing on the encoded features output by the encoder in the U-shaped super-resolution network, that is, enter step S407.
[0170] Step S407: Use the decoder to decode the low-resolution features to obtain the predicted thermal image output by the U-shaped super-resolution network.
[0171] Here, the predicted thermal image can be denoted as I SR .
[0172] Step S408: Determine the basic loss and feature loss based on the predicted thermal image and the high-resolution thermal image.
[0173] Here, based on I SR and I HR , determine L 1 and L CX .
[0174] Step S409: Backpropagate the basic loss and feature loss to adjust the network parameters of the U-shaped super-resolution network.
[0175] The above step S406 and step S409 can be executed alternately. Adjust the encoder of the U-shaped super-resolution network through step S406, and adjust the overall network parameters of the U-shaped super-resolution network through step S409.
[0176] In some embodiments, the architecture of the U-shaped super-resolution network is as follows Figure 5 shown. The network architecture includes: an encoder 501, a decoder 502, and a bicubic interpolation module 503. Among them, the encoder 501 includes: a 3×3 convolutional layer 51, a 1×1 convolutional layer 52, and three downsampling general modules (Ccommon Block) 53, 54, and 55. First, the bicubic interpolation module 503 performs bicubic interpolation on the input low-resolution thermal image 500 to obtain I LBIC , so as to increase the dimension of the low-resolution thermal image. Then, the low-resolution thermal image with increased dimension is input into the 3×3 convolutional layer 51 for feature extraction, and the extracted features are then input into the 1×1 convolutional layer 52 for further feature extraction. After that, the features extracted by the 1×1 convolutional layer 52 are input into the general module, and the general modules 53, 54, and 55 perform downsampling and feature extraction on the input features to obtain image features with a smaller image size but a higher dimension. And the image features are input into the decoder 502.
[0177] The decoder 502 includes: three upsampling general modules (Ccommon Block) 56, 57, and 58, and two 3×3 convolutional layers 504 and 505. Among them, the general modules 56, 57, and 58 perform upsampling and decoding on the input image features, and the upsampled features are input into the 3×3 convolutional layers 504 and 505. The 3×3 convolutional layers 504 and 505 perform high-resolution reconstruction on the input features to obtain the reconstructed thermal image. Then, the reconstructed thermal image and the low-resolution thermal image with increased dimension are fused to obtain the super-resolution thermal image 506 predicted by the U-shaped network.
[0178] In some embodiments, to enable the U-shaped super-resolution network to tolerate poorly registered data, the embodiments of the present application introduce contrastive learning to enhance the encoder in the network. As Figure 6 shown, it can be achieved through the following process:
[0179] For a training low-resolution thermal image data I LR 61 (that is, the collected low-resolution thermal image, for example, the sample thermal image of the first resolution), first, some image sets 601 (including at least one frame of image) that do not match the real label I HR pixels are generated by randomly rotating or translating 62; then, these image sets 601 are input into the encoder 63 of the network to obtain feature vectors At the same time, I LR is also input into the encoder to obtain a feature vector F LR ; F LR and together constitute the positive sample 64. Obviously, FLR and the more similar the better, because F LR and come from the same image, only with different degrees of pixel mismatch. To make F LR and more similar, the negative sample 65 is introduced in the embodiments of the present application, and the distance between the negative sample 65 and F LR is made as large as possible. Therefore, an auxiliary encoder 66 is introduced. The structure of the auxiliary encoder is the same as that of the encoder, but the weights are not shared. In this way, the true label I HR 67 (i.e., the collected high-resolution sample thermal image, for example, the sample thermal image of the second resolution) generates an image set 602 through rotation or translation 62 (as Figure 6 shown, three frames of thermal images can be obtained); the image set 602 is subjected to feature extraction through the auxiliary encoder 66 to obtain (i.e., the image features of the negative sample 65). Since I HR and I LR come from different sensors, the distance between F LR and is made as large as possible. Using F LR , and to determine and optimize the comparison loss 603, so that the encoder is not affected by poorly registered data and can obtain features beneficial to super-resolution reconstruction. The comparison loss L con is shown in formula (1):
[0180]
[0181] In formula (1), B represents the number of images in a batch during a training process, and τ represents a hyperparameter.
[0182] In some embodiments, in order to further enhance the detailed information of the reconstructed super-resolution image, a feature loss and a basic loss between the predicted super-resolution thermal image and the true super-resolution thermal image are introduced here. The super-resolution image and the true sample are subjected to feature extraction through a pre-trained convolutional neural network (for example, the Visual Geometry Group (VGG16)), and then the feature loss is calculated using the features of the super-resolution image and the features of the true sample, so as to enhance the detailed features. The feature loss L CX is shown in formula (2):
[0183]
[0184] In formula 2, cosine() is used to determine the cosine distance between vectors.
[0185] The basic loss between the predicted super-resolution thermal image and the real super-resolution thermal image is shown in Equation (3):
[0186]
[0187] In Equation 2, K represents the total number of frames of the sample images.
[0188] In the embodiments of the present application, a U-shaped super-resolution network is used to extract and encode low-resolution features, so as to obtain a super-resolution thermal image. At the same time, contrastive learning is used to enhance the encoder, so that the network is not affected by poorly registered data, and the feature loss is increased to optimize the network, so that the restored and reconstructed image has more detailed features.
[0189] Based on the foregoing embodiments, the embodiments of the present application provide a thermal image reconstruction device. The device includes each unit included and each module included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0190] Figure 7 It is a schematic structural diagram of the composition of a training device for a thermal image reconstruction network provided by the embodiments of the present application, as Figure 7 shown. The training device 700 of the thermal image reconstruction network includes:
[0191] A first acquisition module 701, configured to acquire a first sample thermal image with a first resolution and a second sample thermal image with a second resolution corresponding to the first sample thermal image; wherein, the second resolution is higher than the first resolution;
[0192] A first adjustment module 702, configured to adjust the first sample thermal image and the second sample thermal image respectively to obtain a third sample thermal image and a fourth sample thermal image;
[0193] A first determination module 703, configured to determine a first loss for comparing different sample thermal images in a thermal image reconstruction network to be trained based on the first sample thermal image, the third sample thermal image, and the fourth sample thermal image;
[0194] A second determination module 704, configured to determine a second loss based on the second sample thermal image and the predicted thermal image, where the predicted thermal image is a thermal image of the first sample thermal image predicted by the to-be-trained thermal image reconstruction network at the second resolution;
[0195] A first training module 705, configured to train the to-be-trained thermal image reconstruction network based on the first loss and the second loss, so that the loss output by the trained thermal image reconstruction network meets a preset convergence condition.
[0196] In some embodiments, the first adjustment module 702 includes: a first adjustment sub-module, configured to perform at least one rotation or translation on the first sample thermal image to obtain at least one frame of third sample thermal image; a second adjustment sub-module, configured to perform at least one rotation or translation on the second sample thermal image to obtain at least one frame of fourth sample thermal image.
[0197] In some embodiments, the first determination module 703 includes: a first extraction sub-module, configured to extract features from the first sample thermal image and the third sample thermal image by using a first encoder in the to-be-trained thermal image reconstruction network to obtain a first image feature and a third image feature; a second extraction sub-module, configured to extract features from the fourth sample thermal image by using a second encoder that does not share weights with the first encoder to obtain a fourth image feature; a first determination sub-module, configured to determine the first loss based on the first image feature, the third image feature, and the fourth image feature.
[0198] In some embodiments, the first determination sub-module includes: a first determination unit, configured to determine a first similarity between the first image feature and the third image feature; a second determination unit, configured to determine a second similarity between the first image feature and the fourth image feature; a first comparison unit, configured to compare the first similarity and the second similarity to obtain the first loss.
[0199] In some embodiments, the apparatus further includes: a second adjustment module, configured to adjust the dimensions of the first sample thermal image and the third sample thermal image respectively to obtain an adjusted first sample thermal image and an adjusted third sample thermal image; the first extraction sub-module is further configured to: extract features from the adjusted first sample thermal image and the adjusted third sample thermal image by using the first encoder to obtain the first image feature and the third image feature.
[0200] In some embodiments, the first extraction sub-module includes: a first extraction unit configured to perform feature extraction on the modulated first sample thermal image and the modulated third sample thermal image by using a convolutional module in the first encoder to obtain a first candidate feature and a second candidate feature; and a first downsampling unit configured to perform downsampling on the first candidate feature and the second candidate feature by using a downsampling module in the first encoder to obtain the first image feature and the third image feature.
[0201] In some embodiments, the apparatus further includes: a first reconstruction module configured to perform super-resolution reconstruction on the first image feature of the first sample thermal image in a decoder corresponding to the first encoder to obtain a reconstructed image; and a first fusion module configured to fuse the reconstructed thermal image and the modulated first sample thermal image corresponding to the first sample thermal image to obtain the predicted thermal image.
[0202] In some embodiments, the first reconstruction module includes: a first upsampling sub-module configured to perform upsampling on the first image feature by using an upsampling module in the decoder to obtain an upsampled feature; and a first reconstruction sub-module configured to perform reconstruction on the upsampled feature by using a reconstruction module in the decoder to obtain the reconstructed thermal image.
[0203] In some embodiments, the second determination module 704 includes: a second determination sub-module configured to determine a basic loss representing the difference between images based on the second sample thermal image and the predicted thermal image; a third determination sub-module configured to determine a feature loss representing the similarity between features based on the second image feature of the second sample thermal image and the predicted feature of the predicted thermal image; and a fourth determination sub-module configured to determine the basic loss and the feature loss as the second loss.
[0204] In some embodiments, the first training module 705 includes: a third adjustment sub-module configured to adjust structural parameters of a first encoder and a second encoder in the to-be-trained thermal image reconstruction network by using the first loss to obtain an intermediate thermal image reconstruction network; and a fourth adjustment sub-module configured to adjust network parameters of the intermediate thermal image reconstruction network by using the second loss to obtain the trained thermal image reconstruction network.
[0205] In some embodiments, the first acquisition module 701 includes: a first acquisition sub-module configured to acquire a thermal image of a preset object by using a first-resolution sensor to obtain the first sample thermal image; and a second acquisition sub-module configured to acquire a thermal image of the preset object by using a second-resolution sensor to obtain the second sample thermal image.
[0206] An embodiment of the present application provides an image reconstruction device, as Figure 8 shown, Figure 8 which is a schematic structural diagram of a thermal image reconstruction device provided by an embodiment of the present application. The thermal image reconstruction device 800 includes:
[0207] An acquisition module 801, configured to acquire a first thermal image to be super-resolution reconstructed;
[0208] An adjustment module 802, configured to adjust the dimension of the first thermal image in a trained thermal image reconstruction network to obtain an adjusted first thermal image; wherein, the trained thermal image reconstruction network is trained by the method described above;
[0209] An extraction module 803, configured to extract features of the adjusted first thermal image to obtain image features to be reconstructed;
[0210] A reconstruction module 804, configured to perform super-resolution reconstruction on the image features to be reconstructed based on the super-resolution in the trained thermal image reconstruction network to obtain a reconstructed super-resolution thermal image;
[0211] A determination module 805, configured to determine a second thermal image with a resolution higher than that of the first thermal image based on the reconstructed super-resolution thermal image and the adjusted first thermal image.
[0212] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects to the method embodiment. In some embodiments, the functions or modules included in the device provided by the embodiments of the present application can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0213] It should be noted that in the embodiments of the present application, if the training method of the above thermal image reconstruction network is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any arbitrary combination of hardware, software, and firmware.
[0214] An embodiment of the present application provides a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, some or all of the steps in the above method are implemented.
[0215] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.
[0216] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.
[0217] An embodiment of the present application provides a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be specifically implemented in a way of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0218] It should be noted here that: the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities or similarities can be referred to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0219] It should be noted that Figure 9 is a schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application. As Figure 9 shown, the hardware entity of the computer device 900 includes: a processor 901, a communication interface 902, and a memory 903, where:
[0220] The processor 901 generally controls the overall operation of the computer device 900.
[0221] The communication interface 902 can enable the computer device to communicate with other terminals or servers through a network.
[0222] The memory 903 is configured to store instructions and applications executable by the processor 901, and can also cache data to be processed or already processed by the processor 901 and each module in the computer device 900 (such as, image data, audio data, voice communication data, and video communication data), and can be implemented by a flash memory (FLASH) or a random access memory (RAM). Data transmission can be performed between the processor 901, the communication interface 902, and the memory 903 via a bus 904.
[0223] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence numbers of the above steps / processes does not mean the order of execution. The order of execution of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The sequence numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0224] It should be noted that, in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0225] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings, direct couplings, or communication connections between the components shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.
[0226] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0227] In addition, each functional unit in the embodiments of the present application may be all integrated in a processing unit, or each unit may be separately regarded as a unit, or two or more units may be integrated in one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of a combination of hardware and software functional units.
[0228] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical discs and other various media that can store program codes.
[0229] Alternatively, if the above-mentioned integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application essentially or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks, or optical discs and other various media that can store program codes.
[0230] The above is only the implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application.
Claims
1. A training method for a thermal image reconstruction network, characterized in that, the method includes: Obtain a first sample thermal image with a first resolution and a second sample thermal image with a second resolution corresponding to the first sample thermal image; wherein, the second resolution is higher than the first resolution; Perform at least one rotation or translation on the first sample thermal image to obtain at least one frame of third sample thermal images; Perform at least one rotation or translation on the second sample thermal image to obtain at least one frame of fourth sample thermal images; In the thermal image reconstruction network to be trained, based on the first sample thermal image, the third sample thermal images, and the fourth sample thermal images, determine a first loss for comparing different sample thermal images; Based on the second sample thermal image and the predicted thermal image, determine a second loss; wherein, the predicted thermal image is the thermal image of the first sample thermal image predicted by the thermal image reconstruction network to be trained at the second resolution; Based on the first loss and the second loss, train the thermal image reconstruction network to be trained so that the loss output by the trained thermal image reconstruction network meets a preset convergence condition; Wherein, in the thermal image reconstruction network to be trained, based on the first sample thermal image, the third sample thermal images, and the fourth sample thermal images, determining the first loss for comparing different sample thermal images includes: using a first encoder in the thermal image reconstruction network to be trained to extract features from the first sample thermal image and the third sample thermal images to obtain a first image feature and a third image feature; using a second encoder that does not share weights with the first encoder to extract features from the fourth sample thermal image to obtain a fourth image feature; based on the first image feature, the third image feature, and the fourth image feature, determine the first loss; Wherein, after obtaining the first image feature and the third image feature, the method further includes: in the decoder corresponding to the first encoder, perform super-resolution reconstruction on the first image feature of the first sample thermal image at the second resolution to obtain a reconstructed thermal image; fuse the reconstructed thermal image and the adjusted first sample thermal image corresponding to the first sample thermal image to obtain the predicted thermal image; Wherein, the determining the second loss based on the second sample thermal image and the predicted thermal image includes: based on the second sample thermal image and the predicted thermal image, determine a basic loss representing the difference between the images; based on the second image feature of the second sample thermal image and the predicted feature of the predicted thermal image, determine a feature loss representing the similarity between the features; determine the basic loss and the feature loss as the second loss.
2. The method according to claim 1, characterized in that, the determining the first loss based on the first image feature, the third image feature, and the fourth image feature includes: Determine a first similarity between the first image feature and the third image feature; Determine a second similarity between the first image feature and the fourth image feature; Compare the first similarity and the second similarity to obtain the first loss.
3. The method according to claim 1, wherein, before using the first encoder in the to-be-trained thermal image reconstruction network to extract features from the first sample thermal image and the third sample thermal image to obtain a first image feature and a third image feature, the method further includes: Adjust the dimensions of the first sample thermal image and the third sample thermal image respectively to obtain an adjusted first sample thermal image and an adjusted third sample thermal image; The using the first encoder in the to-be-trained thermal image reconstruction network to extract features from the first sample thermal image and the third sample thermal image to obtain a first image feature and a third image feature includes: Using the first encoder to extract features from the adjusted first sample thermal image and the adjusted third sample thermal image respectively to obtain the first image feature and the third image feature.
4. The method according to claim 3, wherein, the using the first encoder to extract features from the adjusted first sample thermal image and the adjusted third sample thermal image respectively to obtain the first image feature and the third image feature includes: Using a convolutional module in the first encoder to extract features from the adjusted first sample thermal image and the adjusted third sample thermal image to obtain a first candidate feature and a second candidate feature; Using a downsampling module in the first encoder to downsample the first candidate feature and the second candidate feature to obtain the first image feature and the third image feature.
5. The method according to claim 1, wherein, in the decoder corresponding to the first encoder, performing super-resolution reconstruction on the first image feature of the first sample thermal image to obtain a reconstructed thermal image, including: Using an upsampling module in the decoder to upsample the first image feature to obtain an upsampled feature; Using a reconstruction module in the decoder to reconstruct the upsampled feature to obtain the reconstructed thermal image.
6. The method according to any one of claims 1 to 5, wherein, the training the to-be-trained thermal image reconstruction network based on the first loss and the second loss such that the loss output by the trained thermal image reconstruction network meets a preset convergence condition includes: Using the first loss to adjust the structural parameters of the first encoder and the second encoder in the to-be-trained thermal image reconstruction network to obtain an intermediate thermal image reconstruction network; Using the second loss to adjust the network parameters of the intermediate thermal image reconstruction network to obtain the trained thermal image reconstruction network.
7. The method according to any one of claims 1 to 5, wherein, the obtaining a first sample thermal image with a first resolution and a second sample thermal image with a second resolution corresponding to the first sample thermal image includes: Using a first-resolution sensor to collect a thermal image of a preset object to obtain the first sample thermal image; Using a second-resolution sensor to collect a thermal image of the preset object to obtain the second sample thermal image.
8. A thermal image reconstruction method, characterized in that the method includes: obtaining a first thermal image to be super-resolution reconstructed; in a trained thermal image reconstruction network, adjusting the dimension of the first thermal image to obtain an adjusted first thermal image; wherein, the trained thermal image reconstruction network is trained by the method described in any one of the above claims 1 to 7; performing feature extraction on the adjusted first thermal image to obtain image features to be reconstructed; based on the super-resolution in the trained thermal image reconstruction network, performing super-resolution reconstruction on the image features to be reconstructed to obtain a reconstructed super-resolution thermal image; based on the reconstructed super-resolution thermal image and the adjusted first thermal image, determining a second thermal image with a resolution higher than that of the first thermal image.
9. A training device for a thermal image reconstruction network, characterized in that it includes: a first acquisition module, configured to acquire a first sample thermal image with a first resolution and a second sample thermal image with a second resolution corresponding to the first sample thermal image; wherein, the second resolution is higher than the first resolution; a first adjustment module, configured to perform at least one rotation or translation on the first sample thermal image to obtain at least one frame of third sample thermal image; perform at least one rotation or translation on the second sample thermal image to obtain at least one frame of fourth sample thermal image; a first determination module, configured to determine a first loss for comparing different sample thermal images in a thermal image reconstruction network to be trained based on the first sample thermal image, the third sample thermal image, and the fourth sample thermal image; a second determination module, configured to determine a second loss based on the second sample thermal image and a predicted thermal image; wherein, the predicted thermal image is a thermal image of the first sample thermal image predicted by the thermal image reconstruction network to be trained at the second resolution; a first training module, configured to train the thermal image reconstruction network to be trained based on the first loss and the second loss, so that the loss output by the trained thermal image reconstruction network meets a preset convergence condition; wherein, the first determination module includes: a first extraction sub-module, configured to use a first encoder in the thermal image reconstruction network to be trained to perform feature extraction on the first sample thermal image and the third sample thermal image to obtain a first image feature and a third image feature; a second extraction sub-module, configured to use a second encoder that does not share weights with the first encoder to perform feature extraction on the fourth sample thermal image to obtain a fourth image feature; a first determination sub-module, configured to determine the first loss based on the first image feature, the third image feature, and the fourth image feature; Wherein, the device further includes: a first reconstruction module, configured to perform super-resolution reconstruction on the first image feature of the first sample thermal image at a second resolution in the decoder corresponding to the first encoder, to obtain a reconstructed thermal image; a first fusion module, configured to fuse the reconstructed thermal image and the adjusted first sample thermal image corresponding to the first sample thermal image, to obtain the predicted thermal image; Wherein, the second determination module includes: a second determination sub-module, configured to determine a basic loss representing the difference between images based on the second sample thermal image and the predicted thermal image; a third determination sub-module, configured to determine a feature loss representing the similarity between features based on the second image feature of the second sample thermal image and the predicted feature of the predicted thermal image; a fourth determination sub-module, configured to determine the basic loss and the feature loss as the second loss.
10. A thermal image reconstruction device Characterized in that It includes An acquisition module, configured to acquire a first thermal image to be super-resolution reconstructed; An adjustment module, configured to adjust the dimension of the first thermal image in a trained thermal image reconstruction network, to obtain an adjusted first thermal image; wherein, the trained thermal image reconstruction network is trained by the method according to any one of claims 1 to 7 above; An extraction module, configured to extract features of the adjusted first thermal image, to obtain image features to be reconstructed; A reconstruction module, configured to perform super-resolution reconstruction on the image features to be reconstructed based on the super-resolution in the trained thermal image reconstruction network, to obtain a reconstructed super-resolution thermal image; A determination module, configured to determine a second thermal image with a resolution higher than that of the first thermal image based on the reconstructed super-resolution thermal image and the adjusted first thermal image.
11. A computer device, including a memory and a processor, the memory stores a computer program that can run on the processor, Characterized in that When the processor executes the program, it implements the steps in the method according to any one of claims 1 to 7, or, when the processor executes the program, it implements the steps in the method according to claim 8.
12. A computer-readable storage medium, on which a computer program is stored, Characterized in that When the computer program is executed by a processor, it implements the steps in the method according to any one of claims 1 to 7, or, when the computer program is executed by a processor, it implements the steps in the method according to claim 8.
Citation Information
Patent Citations
Training and reconstruction method of super-resolution reconstruction model of face image
CN112581370A
Model training method, super-resolution reconstruction method, device, equipment and medium
CN114494022A