An infrared image processing method, apparatus, device and medium

By fusing low-resolution IR images with high-resolution visible light images using dynamic weight determination, the method enhances IR image resolution and fidelity, addressing the limitations of low-resolution IR thermal imaging devices.

CN119624780BActive Publication Date: 2025-07-15SHANGHAI SHENQISHEN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510147217.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-15
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

Due to hardware cost and resolution limitations, existing infrared thermal imaging equipment has low infrared image resolution, making it difficult to meet the application needs of high-precision detection and detailed analysis.

Method used

By acquiring low-resolution infrared images and high-resolution visible light images, the weights are determined using similarity calculations, and image fusion is performed, super-resolution reconstruction is performed in combination with a guide filter and a convolutional neural network, and the weights are dynamically adjusted to preserve the temperature information of the infrared image and the high-frequency details of the visible light image.

Benefits of technology

Without increasing hardware costs, the resolution and reality of infrared images are improved, artifacts are prevented, key temperature and texture information are retained, and high-resolution, high-detailed infrared images are generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119624780B_ABST
    Figure CN119624780B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an infrared image processing method, apparatus, device and medium, relating to the technical field of image processing. The method includes: obtaining an initial infrared image with a first resolution and a visible light image with a second resolution, where the first resolution is lower than the second resolution; determining a first weight of the initial infrared image and a second weight of the visible light image based on the similarity between the initial infrared image and the visible light image, the similarity being positively correlated with the first weight, and the sum of the first weight and the second weight being a preset value; fusing the initial infrared image and the visible light image using the first weight and the second weight to obtain a first target infrared image with the second resolution. By applying the technical solution provided by the embodiment of the present application, the resolution of the infrared image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technologies, and in particular, to an infrared image processing method, apparatus, device, and medium. Background Art

[0002] Currently, infrared thermal imaging technology has been widely applied in the fields of security, medical treatment, industrial inspection, etc. However, limited by the hardware cost and resolution of infrared sensors, the resolution of infrared (IR) images (I) collected by existing infrared thermal imaging devices is relatively low, and the imaging quality is difficult to meet the application requirements of high-precision detection, detail analysis, etc. How to improve the resolution of infrared images without increasing the hardware cost is an urgent problem to be solved currently. Summary of the Invention

[0003] The purpose of the embodiments of the present application is to provide an infrared image processing method, apparatus, device, and medium to improve the resolution of infrared images. The specific technical solutions are as follows:

[0004] In a first aspect, the embodiments of the present application provide an infrared image processing method, and the method includes:

[0005] Obtain an initial infrared image with a first resolution and a visible light image with a second resolution, where the first resolution is lower than the second resolution;

[0006] Extract a first feature of each pixel point in the initial infrared image and a second feature of each pixel point in the visible light image;

[0007] Calculate the similarity between the initial infrared image and the visible light image at each pixel point by using the first feature and the second feature of each pixel point;

[0008] Determine a first weight of the initial infrared image and a second weight of the visible light image by using the similarity between the initial infrared image and the visible light image, where the similarity is positively correlated with the first weight, and the sum of the first weight and the second weight is a preset value;

[0009] Fuse the initial infrared image and the visible light image by using the first weight and the second weight to obtain a first target infrared image with the second resolution;

[0010] The step of determining the first weight of the initial infrared image and the second weight of the visible light image by using the similarity between the initial infrared image and the visible light image includes:

[0011] Determine the first weight of the initial infrared image and the second weight of the visible light image at each pixel point by using the similarity between the initial infrared image and the visible light image at each pixel point;

[0012] The step of fusing the initial infrared image and the visible light image by using the first weight and the second weight to obtain the first target infrared image with the second resolution includes:

[0013] Fuse each pixel point in the initial infrared image and the visible light image by using the first weight of the initial infrared image and the second weight of the visible light image at each pixel point to obtain the first target infrared image with the second resolution.

[0014] In some embodiments, the step of fusing each pixel point in the initial infrared image and the visible light image by using the first weight of the initial infrared image and the second weight of the visible light image at each pixel point to obtain the first target infrared image with the second resolution includes:

[0015] Perform weighted averaging on the first feature of each pixel point in the initial infrared image and the second feature of each pixel point in the visible light image by using the first weight and the second weight at each pixel point to obtain the first target infrared image with the second resolution.

[0016] In some embodiments, the step of determining the first weight of the initial infrared image and the second weight of the visible light image by using the similarity between the initial infrared image and the visible light image further includes:

[0017] Determine the first weight of the initial infrared image and the second weight of the visible light image at the multiple scales by using the similarity between the initial infrared image and the visible light image at the multiple scales;

[0018] The step of fusing the initial infrared image and the visible light image by using the first weight and the second weight to obtain the first target infrared image with the second resolution further includes:

[0019] Fuse the initial infrared image and the visible light image by using the first weight and the second weight at each scale to obtain the target sub-image at each scale;

[0020] Fuse the target sub-images at the multiple scales by using the weights corresponding to the multiple scales to obtain the first target infrared image with the second resolution.

[0021] In some embodiments, after obtaining the first target infrared image with the second resolution, the method further includes:

[0022] Perform a preset process on the first target infrared image to obtain a second target infrared image. The preset process includes at least one of the following: denoising using a residual block and detail enhancement using an adversarial network.

[0023] In some embodiments, the method further includes:

[0024] Obtain a first sample infrared image, a second sample infrared image, and a sample visible light image. The resolution of the first sample infrared image is a third resolution, and the resolutions of the second sample infrared image and the sample visible light image are a fourth resolution, where the third resolution is lower than the fourth resolution;

[0025] Input the first sample infrared image and the sample visible light image into a preset network model to obtain a reconstructed infrared image. The preset network model is used to determine a third weight of the first sample infrared image and a fourth weight of the sample visible light image based on the similarity between the first sample infrared image and the sample visible light image. The similarity is positively correlated with the third weight, and the sum of the third weight and the fourth weight is the preset value; use the third weight and the fourth weight to fuse the first sample infrared image and the sample visible light image, and output the reconstructed infrared image;

[0026] Obtain the reconstruction loss and the structural similarity loss between the reconstructed infrared image and the second sample infrared image;

[0027] Determine whether the preset network model converges according to the reconstruction loss and the structural similarity loss.

[0028] In some embodiments, the preset network model includes an adversarial network, and the method further includes:

[0029] Obtain the adversarial loss between the reconstructed infrared image and the second sample infrared image;

[0030] The step of determining whether the preset network model converges according to the reconstruction loss and the structural similarity loss includes:

[0031] Determine whether the preset network model converges according to the sum of the reconstruction loss, the structural similarity loss, and the adversarial loss.

[0032] In a second aspect, an embodiment of the present application provides an infrared image processing device, which includes:

[0033] An acquisition module, configured to acquire an initial infrared image with a first resolution and a visible light image with a second resolution, where the first resolution is lower than the second resolution;

[0034] A determination module, configured to extract the first feature of each pixel point in the initial infrared image and the second feature of each pixel point in the visible light image; calculate the similarity between the initial infrared image and the visible light image at each pixel point by using the first feature and the second feature of each pixel point; determine the first weight of the initial infrared image and the second weight of the visible light image by using the similarity between the initial infrared image and the visible light image, where the similarity is positively correlated with the first weight, and the sum of the first weight and the second weight is a preset value;

[0035] A fusion module, configured to fuse the initial infrared image and the visible light image by using the first weight and the second weight to obtain a first target infrared image with the second resolution;

[0036] The determination module is specifically configured to: determine the first weight of the initial infrared image and the second weight of the visible light image at each pixel point by using the similarity between the initial infrared image and the visible light image at each pixel point;

[0037] The fusion module is specifically configured to: fuse each pixel point in the initial infrared image and the visible light image by using the first weight of the initial infrared image and the second weight of the visible light image at each pixel point to obtain a first target infrared image with the second resolution.

[0038] In some embodiments, the fusion module is specifically configured to: perform weighted averaging on the first feature of each pixel point in the initial infrared image and the second feature of each pixel point in the visible light image by using the first weight and the second weight at each pixel point to obtain a first target infrared image with the second resolution.

[0039] In some embodiments, the determination module is specifically configured to: determine the first weight of the initial infrared image and the second weight of the visible light image at multiple scales by using the similarity between the initial infrared image and the visible light image at multiple scales;

[0040] The fusion module is specifically configured to: fuse the initial infrared image and the visible light image by using the first weight and the second weight at each scale to obtain a target sub-image at each scale; fuse the target sub-images at multiple scales by using the weights corresponding to the multiple scales to obtain a first target infrared image with the second resolution.

[0041] In some embodiments, the device further includes a processing module, configured to: after obtaining the first target infrared image with the second resolution, perform a preset process on the first target infrared image to obtain a second target infrared image, where the preset process includes at least one of the following: denoising using a residual block and detail enhancement using an adversarial network.

[0042] In some embodiments, the device further includes a training module, configured to:

[0043] Obtain a first sample infrared image, a second sample infrared image, and a sample visible light image, where the resolution of the first sample infrared image is a third resolution, and the resolutions of the second sample infrared image and the sample visible light image are a fourth resolution, and the third resolution is lower than the fourth resolution;

[0044] Input the first sample infrared image and the sample visible light image into a preset network model to obtain a reconstructed infrared image. The preset network model is configured to use the similarity between the first sample infrared image and the sample visible light image to determine a third weight of the first sample infrared image and a fourth weight of the sample visible light image. The similarity is positively correlated with the third weight, and the sum of the third weight and the fourth weight is the preset value; use the third weight and the fourth weight to fuse the first sample infrared image and the sample visible light image, and output the reconstructed infrared image;

[0045] Obtain a reconstruction loss and a structural similarity loss between the reconstructed infrared image and the second sample infrared image;

[0046] Determine whether the preset network model converges according to the reconstruction loss and the structural similarity loss.

[0047] In some embodiments, the preset network model includes an adversarial network, and the training module is further configured to: obtain an adversarial loss between the reconstructed infrared image and the second sample infrared image;

[0048] Specifically, the training module is configured to: determine whether the preset network model converges according to the sum of the reconstruction loss, the structural similarity loss, and the adversarial loss.

[0049] In a third aspect, an embodiment of the present application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0050] The memory is configured to store a computer program;

[0051] A processor, when executing a program stored in a memory, implements the method described in any one of the above first aspects.

[0052] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the above first aspects is implemented.

[0053] An embodiment of the present application also provides a computer program product containing instructions, which when running on a computer, causes the computer to execute the method described in any one of the above first aspects in the embodiments.

[0054] Advantages of the embodiments of the present application:

[0055] In the technical solution provided by the embodiment of the present application, a high-resolution visible light image is used to guide the super-resolution reconstruction of a low-resolution infrared image, which can ensure that the temperature information of the infrared image is consistent with the high-frequency details of the visible light image. In addition, using the similarity between the low-resolution infrared image and the high-resolution visible light image, the weights of the infrared image and the visible light image are dynamically determined when reconstructing the infrared image, and the infrared image and the visible light image are fused using the dynamically determined weights to achieve super-resolution reconstruction. When the similarity between the visible light image and the infrared image is relatively high, the weight tilts towards the infrared image, and more characteristics of the infrared image are retained. When the similarity is relatively low, the weight tilts towards the visible light image, which can effectively prevent artifacts from generating while retaining key temperature and texture information, ensuring that the generated infrared image has high resolution and high authenticity, and improving the resolution of the infrared image.

[0056] Of course, when implementing any product or method of the present application, it is not necessarily required to achieve all the above-mentioned advantages simultaneously. Description of the Drawings

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.

[0058] Figure 1 It is the first flow schematic diagram of the infrared image processing method provided by the embodiment of the present application;

[0059] Figure 2 It is the second flow schematic diagram of the infrared image processing method provided by the embodiment of the present application;

[0060] Figure 3 It is the third flow schematic diagram of the infrared image processing method provided by the embodiment of the present application;

[0061] Figure 4 It is a schematic flowchart of a training method provided by an embodiment of the present application;

[0062] Figure 5 It is a schematic structural diagram of an infrared image processing device provided by an embodiment of the present application;

[0063] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0064] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.

[0065] Currently, infrared thermal imaging technology has been widely used in the fields of security, medical treatment, industrial inspection, etc. Infrared thermal imaging devices detect the infrared radiation on the surface of an object and convert it into a visible image. Since infrared thermal imaging technology can work in a completely dark, smoky or occluded environment and can directly measure temperature information, it has irreplaceable advantages in harsh environments. However, limited by the hardware cost and resolution of infrared sensors, the resolution of infrared images collected by existing infrared thermal imaging devices is relatively low, and the imaging quality is difficult to meet the application requirements of high-precision detection, detail analysis, etc. How to improve the resolution of infrared images without increasing the hardware cost is an urgent problem to be solved currently.

[0066] In the prior art, super-resolution (SR) reconstruction (rec) can be performed based on methods such as interpolation, sparse representation, and deep learning to improve the image resolution. However, the method based on interpolation only relies on the low-resolution information of the image itself and cannot restore the high-frequency details in the image. The generated high-resolution (HR) images are usually blurred, especially in the edge and texture regions. The method based on sparse representation has a high computational complexity, a slow processing speed, and it is difficult to construct a high-quality dictionary in some complex scenarios. In addition, the method based on deep learning relies on large-scale high-quality data training and is unstable in new environments.

[0067] To solve the above problems, an embodiment of the present application provides an infrared image processing method, which can be applied to electronic devices such as computers, servers, and clusters. For the convenience of description, the electronic device will be used as the execution subject in the following, which does not play a limiting role.

[0068] See Figure 1 , Figure 1 which is the first schematic flowchart of the infrared image processing method provided by the embodiment of the present application. The infrared image processing method includes the following steps.

[0069] Step S11: Obtain an initial infrared image with a first resolution and a visible light image with a second resolution, where the first resolution is lower than the second resolution.

[0070] Step S12: Determine a first weight of the initial infrared image and a second weight of the visible light image by using the similarity between the initial infrared image and the visible light image. The similarity is positively correlated with the first weight, and the sum value of the first weight and the second weight is a preset value.

[0071] Step S13: Use the first weight and the second weight to fuse the initial infrared image and the visible light image to obtain a first target infrared image with the second resolution.

[0072] In the technical solution provided by the embodiment of the present application, the electronic device uses a high-resolution visible light image to guide the super-resolution reconstruction of a low-resolution infrared image, which can ensure that the temperature information of the infrared image is consistent with the high-frequency details of the visible light image. In addition, the electronic device uses the similarity between the low-resolution infrared image and the high-resolution visible light image to dynamically determine the weight of the infrared image and the weight of the visible light image when reconstructing the infrared image, and uses the dynamically determined weights to fuse the infrared image and the visible light image to achieve super-resolution reconstruction. When the similarity between the visible light image and the infrared image is relatively high, the weight tilts towards the infrared image, and more characteristics of the infrared image are retained. When the similarity is relatively low, the weight tilts towards the visible light image, which can effectively prevent artifacts from generating while retaining key temperature and texture information, ensuring that the generated infrared image has high resolution and high authenticity, and improving the resolution of the infrared image.

[0073] In the above step S11, the electronic device collects an infrared image and a visible light (VIS) image of the same scene through a collection device such as an infrared thermal imaging device or a camera. The collected infrared image is an image with a low resolution (LR), and the collected visible light image is an image with a high resolution. In the embodiment of the present application, the collection device can be a separate device or integrated on the electronic device, and the collection device is not limited herein.

[0074] The electronic device can obtain the infrared image and the visible light image collected by the collection device. In one example, the electronic device can directly use the collected visible light image as the visible light image to be processed , and use the collected infrared image as the initial infrared image to be processed . In another example, the electronic device may use the captured visible light image as the visible light image to be processed , and preprocess the captured infrared image, such as performing preliminary noise filtering and normalization processing on the captured infrared image, etc., and use the preprocessed infrared image as the initial infrared image to be processed . Here, the manner in which the electronic device obtains the initial infrared image and the visible light image is not limited.

[0075] In the embodiment of the present application, to implement super-resolution reconstruction of the initial infrared image, the electronic device may first resize the initial infrared image, and perform the subsequent steps S12 and S13 on the resized initial infrared image. The scale (scale, s) of the resized initial infrared image is the same as the scale of the visible light image, and the number of pixel points of the resized initial infrared image is the same as the number of pixel points of the visible light image, that is, the resolution of the resized initial infrared image is the same as the resolution of the visible light image. After the electronic device performs the resizing process, the obtained resized initial infrared image is relatively blurred and has low quality. Therefore, it is necessary to perform super-resolution reconstruction on the resized initial infrared image to improve clarity and texture details. The subsequent initial infrared image mentioned refers to the resized initial infrared image.

[0076] In the above step S12, the first weight is the weight of the initial infrared image, the second weight is the weight of the visible light image, and the sum value of the first weight and the second weight is a preset value. The preset value can be 1, 2, etc., and can be specifically set according to the actual situation. For the convenience of description, the subsequent description will take the preset value as 1 as an example, and is not limited here.

[0077] The similarity between the initial infrared image and the visible light image is positively correlated with the first weight and negatively correlated with the second weight. That is, the higher the similarity, the greater the first weight and the smaller the second weight; the lower the similarity, the smaller the first weight and the greater the second weight.

[0078] In one example, after the electronic device determines the similarity between the initial infrared image and the visible light image, it may directly use the determined similarity as the first weight of the initial infrared image, and calculate the difference between the preset value and the first weight as the second weight of the visible light image.

[0079] In another example, the electronic device may use the determined similarity as an independent variable, combine a preset positive correlation function, calculate the first weight, and calculate the difference between the preset value and the first weight to obtain the second weight. The preset positive correlation function may be a proportional function with a slope greater than 0, a square function, etc., and is not limited thereto.

[0080] Similarly, the electronic device can calculate the second weight by combining a preset negative correlation function, and calculate the difference between the preset value and the second weight to obtain the first weight. The preset negative correlation function can be a direct proportional function with a slope less than 0, an inverse proportional function, etc.

[0081] The electronic device can also calculate the first weight and the second weight by combining a preset positive correlation function and a negative correlation function respectively. In this case, the sum of the preset positive correlation function and the negative correlation function is the preset value. Here, the method for the electronic device to determine the first weight and the second weight is not limited.

[0082] In the above step S13, the electronic device uses the first weight of the initial infrared image and the second weight of the visible light image to fuse the initial infrared image and the visible light image (fusion). That is, taking the visible light image as a guiding image, the electronic device uses the high-resolution information of the visible light image to guide the initial infrared image for super-resolution reconstruction. After that, the electronic device obtains the fused infrared image, and this fused infrared image is the first target infrared image after super-resolution reconstruction. The resolution of this first target infrared image is the same as that of the visible light image, and it is a higher-quality and clearer infrared image.

[0083] In the embodiments of the present application, the electronic device can perform super-resolution reconstruction using a preset network model to obtain the first target infrared image. The preset network model is a pre-trained Guided Super Resolution (GuidedSR) network model. The training process of the preset network model will be described in detail later and will not be elaborated here for the time being.

[0084] The electronic device can input the initial infrared image and the visible light image into the preset network model. The preset network model determines the first weight and the second weight using the similarity between the initial infrared image and the visible light image, and uses the first weight and the second weight to fuse the initial infrared image and the visible light image to obtain and output the first target infrared image. The process of the preset network model for super-resolution reconstruction can be expressed as:

[0085] Formula (1)

[0086] Wherein, represents the first target infrared image; represents the preset network model; represents the initial infrared image; represents the visible light image.

[0087] See Figure 2 , Figure 2This is the second flowchart of the infrared image processing method provided by the embodiments of the present application. The infrared image processing method may include the following steps.

[0088] Step S21: Obtain an initial infrared image with a first resolution and a visible light image with a second resolution, where the first resolution is lower than the second resolution. This is the same as step S11 above.

[0089] Step S22: Extract the first feature of each pixel point in the initial infrared image and the second feature of each pixel point in the visible light image.

[0090] Step S23: Use the first feature and the second feature of each pixel point to calculate the similarity between the initial infrared image and the visible light image at each pixel point.

[0091] Step S24: Use the similarity between the initial infrared image and the visible light image at each pixel point to determine the first weight of the initial infrared image and the second weight of the visible light image at each pixel point.

[0092] Step S25: Use the first weight of the initial infrared image and the second weight of the visible light image at each pixel point to fuse each pixel point in the initial infrared image and the visible light image to obtain a first target infrared image with the second resolution.

[0093] In the technical solution provided by the embodiments of the present application, the electronic device can adaptively adjust the first weight and the second weight according to the first feature and the second feature of each pixel point of the low-resolution infrared image and the high-resolution visible light image, and then fuse the infrared image and the visible light image for each pixel point, improving the fusion accuracy, making the obtained fused pixel points more accurate, while enhancing the image and increasing the resolution, retaining the key information (such as temperature information) of the infrared image, and ensuring the consistency between image enhancement and temperature information.

[0094] In the above step S22, the first feature is the image feature of the initial infrared image, and the second feature is the image feature of the visible light image. The electronic device performs feature extraction on the initial infrared image and the visible light image respectively to obtain the image features of the initial infrared image and the visible light image, that is, obtain the first feature of each pixel point in the initial infrared image and the second feature of each pixel point in the visible light image.

[0095] In the embodiments of the present application, the number of pixel points in the initial infrared image and the visible light image is the same. Any pixel point can be represented as (x, y), where x represents the abscissa and y represents the ordinate.

[0096] An electronic device can extract features using a feature extraction network, which can be a convolutional neural network or the like. Here, the feature extraction network is not limited. The electronic device can input the initial infrared image and the visible light image into the corresponding feature extraction networks of the images respectively for feature extraction. The process of the feature extraction network for feature extraction can be expressed as:

[0097] Formula (2)

[0098] Formula (3)

[0099] Wherein, represents the second feature of the pixel point in the visible light image; represents the feature extraction network of the visible light image; represents the visible light image; represents the first feature of the pixel point in the initial infrared image; represents the feature extraction network of the initial infrared image; represents the initial infrared image.

[0100] In the above step S23, after the electronic device obtains the first feature and the second feature, for each pixel point, it calculates the similarity between the first feature and the second feature of the pixel point, and this similarity is the similarity between the initial infrared image and the visible light image at this pixel point.

[0101] In the embodiments of the present application, the electronic device can use distance algorithms such as Euclidean distance, Canberra distance, and Manhattan distance to calculate the distance between the first feature and the second feature, and use the distance to represent the similarity. The smaller the distance, the higher the similarity; the larger the distance, the lower the similarity. Here, the way for the electronic device to calculate the similarity and represent the similarity is not limited.

[0102] In the above step S24, for each pixel point, the electronic device determines the first weight and the second weight at this pixel point by using the similarity between the initial infrared image and the visible light image at this pixel point. In the embodiments of the present application, the specific way for the electronic device to determine the first weight and the second weight by using the similarity can refer to the relevant description in the above step S12 part, and will not be elaborated here.

[0103] In the above step S25, for each pixel point, the electronic device fuses the pixel point in the initial infrared image and the pixel point in the visible light image by using the first weight and the second weight at this pixel point to obtain the fused pixel point. The electronic device obtains all the fused pixel points, and further obtains the first target infrared image composed of all the pixel points.

[0104] In some embodiments, the above step S25 can be implemented through the following steps: Using the first weight and the second weight at each pixel, perform weighted averaging on the first feature of each pixel in the initial infrared image and the second feature of each pixel in the visible light image to obtain a first target infrared image with a second resolution.

[0105] For each pixel, the electronic device calculates the product of the first weight at this pixel and the first feature of this pixel in the initial infrared image (abbreviated as the first product), and the product of the second weight at this pixel and the second feature of this pixel in the visible light image (abbreviated as the second product), accumulates the two calculated products to obtain the fused feature value of this pixel. The electronic device obtains the fused features of all pixels, that is, obtains the features of the first target infrared image, and further obtains the first target infrared image. By means of weighted averaging, the electronic device can consider the relative importance of the initial infrared image and the visible light image and can more accurately obtain the fused features.

[0106] In the embodiments of the present application, the electronic device can use a Guided Filter (GF) to execute the above step S23 to step S25 for super-resolution reconstruction. The electronic device can input the first feature and the second feature of each pixel into the guided filter, and the guided filter determines the first weight and the second weight, and fuses the first feature and the second feature of each pixel to obtain the fused feature of each pixel, and further obtains the features of the first target infrared image, that is, obtains the first target infrared image. The process of the guided filter for super-resolution reconstruction can be expressed as:

[0107] Formula (4)

[0108] Where, represents the feature of the pixel in the first target infrared image, and can also represent the first target infrared image; represents the guided filter; represents the first feature of the pixel in the initial infrared image; represents the second feature of the pixel in the visible light image.

[0109] After the electronic device inputs the first feature and the second feature of each pixel into the guided filter, it calculates the similarity between the initial infrared image and the visible light image at each pixel. The similarity at the pixel (x, y) can be expressed as:

[0110] Formula (5)

[0111] Among them, d represents the distance between the first feature and the second feature at the pixel point (x, y), which is used to represent the similarity at the pixel point (x, y). The larger d is, the smaller the similarity is; the smaller d is, the larger the similarity is. (x, y) represents the second feature of the pixel point (x, y) in the visible light image; (x, y) represents the first feature of the pixel point (x, y) in the initial infrared image; represents a preset constant, which is a relatively small constant to prevent the denominator from being zero and can be specifically set according to the actual situation and is not limited here.

[0112] Furthermore, the electronic device can determine the first weight and the second weight at each pixel point by using the similarity at the pixel point (x, y). The second weight at the pixel point (x, y) can be expressed as:

[0113] Formula (6)

[0114] The first weight at the pixel point (x, y) can be expressed as:

[0115] Formula (7)

[0116] Among them, (x, y) represents the second weight of the visible light image at the pixel point (x, y); 1 - (x, y) represents the first weight of the initial infrared image at the pixel point (x, y); (x, y) represents the second feature of the pixel point (x, y) in the visible light image; (x, y) represents the first feature of the pixel point (x, y) in the initial infrared image; represents a preset constant.

[0117] The electronic device can perform weighted averaging on the first feature of the pixel point (x, y) in the initial infrared image and the second feature of the pixel point (x, y) in the visible light image by using the first weight and the second weight at the pixel point (x, y). The output of the guided filter at the pixel point (x, y) can be expressed as:

[0118] Formula (8)

[0119] Among them, (x, y) represents the feature of the pixel point (x, y) in the first target infrared image; (x, y) represents the second weight of the visible light image at the pixel point (x, y); (x, y) represents the second feature of the pixel point (x, y) in the visible light image; 1 - (x, y) represents the first weight of the initial infrared image at the pixel point (x, y). (x, y) represents the first feature of the pixel point (x, y) in the initial infrared image.

[0120] Formulas (6) to (8) indicate that, based on the differences in the local features (i.e., the first feature and the second feature) of the initial infrared image and the visible light image at the pixel point (x, y), the first weight and the second weight can be dynamically adjusted. When the similarity of the local features of the initial infrared image and the visible light image at the pixel point (x, y) is relatively high, the second weight is small and the first weight is large, and the weight tilts towards the initial infrared image; when the similarity of the local features of the initial infrared image and the visible light image at the pixel point (x, y) is relatively low, the second weight is large and the first weight is small, and the weight tilts towards the visible light image.

[0121] In the embodiments of the present application, the electronic device can also calculate, for each pixel point, the first product of the first weight at this pixel point and the pixel value of this pixel point in the initial infrared image, and the second product of the second weight at this pixel point and the pixel value of this pixel point in the visible light image, and accumulate the two calculated products to obtain the pixel value after fusion of this pixel point, and further obtain the first target infrared image. Here, the manner in which the electronic device uses the first weight and the second weight to fuse pixel points is not limited.

[0122] In the technical solution provided by the embodiments of the present application, for the initial infrared image and the visible light image at the same scale and the same resolution, each pixel point in the initial infrared image is relatively blurred and inaccurate. The electronic device uses the first weight and the second weight to fuse the features of the initial infrared image and the features of the visible light image for each pixel point, accurately calculates the feature value of each pixel point, and realizes super-resolution reconstruction.

[0123] See Figure 3 , Figure 3 FIG. is the third process schematic diagram of the infrared image processing method provided by the embodiments of the present application. The above infrared image processing method may include the following steps.

[0124] Step S31, obtain an initial infrared image with a first resolution and a visible light image with a second resolution, where the first resolution is lower than the second resolution. It is the same as the above step S11.

[0125] Step S32, use the similarities of the initial infrared image and the visible light image at multiple scales to determine the first weight of the initial infrared image and the second weight of the visible light image at multiple scales.

[0126] Step S33, use the first weight and the second weight at each scale to fuse the initial infrared image and the visible light image to obtain the target sub-image at each scale.

[0127] In step S34, using the weights corresponding to multiple scales, fuse the target sub-images at multiple scales to obtain a first target infrared image with a second resolution.

[0128] In the technical solution provided by the embodiment of the present application, the electronic device performs processing on the initial infrared image and the visible light image at multiple scales, and fuses the target sub-images at multiple scales in an adaptive weighted manner, ensuring the detail processing at different scales, effectively utilizing the initial infrared image information and visible light image information at different scales, enhancing the detail performance and resolution of the fused infrared image, and making the super-resolution reconstruction effect more stable.

[0129] In the above step S32, the electronic device converts the initial infrared image and the visible light image into images at multiple scales. For each scale, using the similarity between the initial infrared image and the visible light image at that scale, determine the first weight and the second weight at that scale. The specific manner in which the electronic device determines the first weight and the second weight using the similarity can refer to the relevant description in the above step S12.

[0130] In the embodiment of the present application, the electronic device can extract the image features of the initial infrared image and the visible light image at multiple scales, calculate the similarity between the first feature of the initial infrared image and the second feature of the visible light image at multiple scales, and use the calculated similarity as the similarity between the initial infrared image and the visible light image at multiple scales.

[0131] For example, the electronic device can, for the initial infrared image and the visible light image with a scale of m×n, respectively convert the scale of the images into m×n, (m / 2)×(n / 2), and (m / 4)×(n / 4) to obtain the initial infrared image and the visible light image at 3 scales, and extract the image features from the initial infrared image and the visible light image at these 3 scales to obtain low-frequency image features, medium-frequency image features, and high-frequency image features, and calculate the similarities at these 3 scales based on the extracted features.

[0132] In one example, the electronic device can extract the image features of the initial infrared image and the visible light image based on pixel points, calculate the similarity, and determine the first weight and the second weight.

[0133] After obtaining the initial infrared images and visible light images at multiple scales, for each scale, the electronic device can extract the first feature of each pixel in the initial infrared image at that scale and the second feature of each pixel in the visible light image. The electronic device can use the feature extraction network for multiple scales to input the initial infrared image and visible light image at each scale into the feature extraction network corresponding to the image and scale respectively for feature extraction. The image features extracted by the feature extraction network at scale s can be expressed as:

[0134] Equation (9)

[0135] Equation (10)

[0136] Among them, s represents any scale; represents the second feature of the pixel in the visible light image at scale s; represents the feature extraction network for the visible light image at scale s; represents the visible light image at scale s; represents the first feature of the pixel in the initial infrared image at scale s; represents the feature extraction network for the initial infrared image at scale s; represents the initial infrared image at scale s.

[0137] For each scale, the electronic device can use the first feature and the second feature of each pixel at that scale to calculate the similarity between the initial infrared image and the visible light image at each pixel at that scale. The specific way for the electronic device to extract the first feature and the second feature to calculate the similarity can be referred to the relevant description in the above Figure 2 part and will not be elaborated here.

[0138] For each scale, the electronic device can use the similarity between the initial infrared image and the visible light image at each pixel at that scale to determine the first weight and the second weight at that pixel at that scale. The specific way for the electronic device to determine the first weight and the second weight can be referred to the relevant description in the above Figure 2 part and will not be elaborated here.

[0139] In the above step S33, for each scale, the electronic device uses the first weight and the second weight at that scale to fuse the initial infrared image and the visible light image at that scale to obtain the fused image at that scale. The fused image is the target sub-image at that scale. The specific way for the electronic device to fuse the initial infrared image and the visible light image can be referred to the relevant description in the above step S13 part.

[0140] In one example, the electronic device may fuse the initial infrared image and the visible light image based on pixel points.

[0141] After obtaining the first weight and the second weight at each pixel point under multiple scales, for each scale, the electronic device may use the first weight and the second weight at each pixel point under that scale to fuse each pixel point in the initial infrared image and the visible light image, so as to obtain the fused image under that scale.

[0142] In the embodiments of the present application, the electronic device may use a guided filter to obtain target sub-images under multiple scales. The electronic device may input the first feature and the second feature of each pixel point under each scale into the guided filter corresponding to each scale. The guided filter corresponding to each scale determines the first weight and the second weight of each pixel point under that scale, and fuses each pixel point in the initial infrared image and the visible light image under that scale, so as to obtain the fused feature value of each pixel point under that scale, and obtain the feature of the target sub-image under that scale, that is, obtain the target sub-image under that scale. The process of the guided filter corresponding to scale s obtaining the target sub-image under scale s can be expressed as:

[0143] Formula (11)

[0144] Where s represents any scale; represents the feature of the pixel point in the target sub-image under scale s, and may also represent the target sub-image under scale s; represents the guided filter under scale s; represents the first feature of the pixel point in the initial infrared image under scale s; represents the second feature of the pixel point in the visible light image under scale s. For the specific manner in which the electronic device obtains the fused image, reference may be made to the relevant description in the above Figure 2 part, which will not be elaborated here.

[0145] In the above step S34, the weights corresponding to multiple scales may be determined during the training of the preset network model, or may be preset, and this is not limited. The electronic device uses the determined weights corresponding to multiple scales (that is, the weighting coefficients of multiple scales) to perform weighted summation on the target sub-images under multiple scales, so as to achieve multi-scale image fusion and obtain the first target infrared image. The process of fusing the target sub-images under different scales to obtain the first target infrared image can be expressed as:

[0146] Formula (12)

[0147] Where represents the feature of the pixel point in the first target infrared image, and may also represent the first target infrared image; s represents any scale; represents the weight corresponding to the scale s; represents the feature of the pixel points in the target sub-image at the scale s.

[0148] In some embodiments, after obtaining the first target infrared image, the electronic device may perform subsequent processing on the first target infrared image to obtain a second target infrared image that has undergone super-resolution reconstruction and subsequent processing. The above infrared image processing method may further include the following steps: performing preset processing on the first target infrared image to obtain a second target infrared image, and the preset processing includes at least one of the following: performing denoising processing using a residual block and performing detail enhancement processing using an adversarial network.

[0149] In the embodiments of the present application, according to the different processing methods included in the preset processing, the situation of the electronic device processing the first target infrared image can be divided into the following three types.

[0150] Situation 1, the preset processing includes performing denoising processing using a residual block.

[0151] In the embodiments of the present application, the electronic device may use a residual block in a pre-trained Residual Network (ResNet) to remove the noise in the first target infrared image. The electronic device may, according to the features of the obtained first target infrared image, input the features of the first target infrared image into the residual network and use the residual block to learn the noise distribution of the first target infrared image. The process of the residual block learning the noise distribution can be expressed as:

[0152] Formula (13)

[0153] Among them, represents the noise distribution in the first target infrared image; represents the residual block; represents the feature of the pixel points in the first target infrared image.

[0154] After obtaining the noise distribution of the first target infrared image, the electronic device may remove the noise component in the features of the first target infrared image to obtain the features of the denoised first target infrared image, that is, obtain the denoised first target infrared image. The process of removing the noise component can be expressed as:

[0155] Formula (14)

[0156] Among them, represents the features of the denoised first target infrared image, and can represent the denoised first target infrared image; represents the feature of the pixel points in the first target infrared image; represents the noise distribution in the first target infrared image. The denoised first target infrared image obtained by the electronic device is the second target infrared image.

[0157] In Case 2, the preset processing includes performing detail enhancement processing using an adversarial network.

[0158] In the embodiments of the present application, the electronic device can use the Generative Adversarial Networks (GAN) (abbreviated as the adversarial network) in the pre-trained detail enhancement module to enhance the details of the first target infrared image. The electronic device can, according to the features of the obtained first target infrared image, input the features of the first target infrared image into the detail enhancement module, and use the generator of the adversarial network in the detail enhancement module to generate the first target infrared image with enhanced details. The first target infrared image with enhanced details obtained by the electronic device is the second target infrared image.

[0159] In Case 3, the preset processing includes performing denoising processing using a residual block and performing detail enhancement processing using an adversarial network.

[0160] In the embodiments of the present application, in order to suppress noise and enhance details simultaneously, the electronic device can adopt a joint optimization strategy. First, use the residual block in the pre-trained residual network to perform denoising processing on the first target infrared image to obtain the denoised first target infrared image. For specific details, refer to the relevant description in Case 1. After the electronic device obtains the features of the denoised first target infrared image, then use the adversarial network in the pre-trained detail enhancement module to perform detail enhancement processing on the features of the denoised first target infrared image, and obtain the processed image as the second target infrared image. The process of the detail enhancement module performing detail enhancement processing can be expressed as:

[0161] Formula (15)

[0162] where, represents the second target infrared image; represents the detail enhancement module; represents the features of the denoised first target infrared image.

[0163] In the embodiments of the present application, the electronic device can also first use the adversarial network in the pre-trained detail enhancement module to perform detail enhancement processing on the first target infrared image, and then use the residual block in the pre-trained residual network to perform denoising processing on the first target infrared image with enhanced details to obtain the second target infrared image, and this is not limited.

[0164] In the technical solution provided by the embodiment of the present application, the electronic device uses a residual block for noise suppression, and by learning the noise distribution of the image, improves the visual quality of the reconstructed image; the detail enhancement module uses an adversarial network to enhance the detail information in the infrared image, further improving the edge sharpness of the image and the resolution of small target areas, ensuring that the infrared image has better details and quality visually, and effectively improving the clarity and reliability of the infrared image. In addition, the electronic device can also adopt a joint optimization strategy of noise suppression and detail enhancement, which reduces the noise while improving the visual quality of the image.

[0165] In the embodiment of the present application, the electronic device can perform preset processing on the first target infrared image by using a preset network model. That is to say, the preset network model can include a residual network and an adversarial network. The electronic device inputs the initial infrared image and the visible light image into the preset network model to obtain the first target infrared image, and uses the residual network and the adversarial network to perform denoising processing and detail enhancement processing on the first target infrared image respectively to obtain the second target infrared image.

[0166] In some embodiments, before processing the initial infrared image, the electronic device can also train the preset network model. Refer to Figure 4 , Figure 4 which is a schematic flowchart of a training method provided by the embodiment of the present application. The above training method can include the following steps.

[0167] Step S41, obtain a first sample infrared image, a second sample infrared image, and a sample visible light image. The resolution of the first sample infrared image is the third resolution, and the resolutions of the second sample infrared image and the sample visible light image are the fourth resolution, and the third resolution is lower than the fourth resolution.

[0168] Step S42, input the first sample infrared image and the sample visible light image into the preset network model to obtain a reconstructed infrared image. The preset network model is used to determine the third weight of the first sample infrared image and the fourth weight of the sample visible light image by using the similarity between the first sample infrared image and the sample visible light image. The similarity is positively correlated with the third weight, and the sum of the third weight and the fourth weight is a preset value; use the third weight and the fourth weight to fuse the first sample infrared image and the sample visible light image, and output the reconstructed infrared image.

[0169] Step S43, obtain the reconstruction loss and the structural similarity loss between the reconstructed infrared image and the second sample infrared image.

[0170] Step S44, determine whether the preset network model converges according to the reconstruction loss and the structural similarity loss. If so, execute step S45; if not, adjust the parameters of the preset network model, and return to execute step S42.

[0171] Step S45, output the preset network model.

[0172] In the technical solution provided by the embodiments of the present application, since the electronic device uses the visible light image for guidance during the super-resolution reconstruction of the infrared image, the amount of data of the infrared images required for training the preset network model is reduced. In addition, the electronic device uses a joint loss function including a reconstruction loss and a structural similarity loss to optimize the preset network model, ensuring that the reconstructed infrared image reaches the optimal level in terms of both structural similarity and overall visual effect. In addition, the coefficients of each loss term and the parameters of the preset network model can be dynamically adjusted through training to achieve the best effect.

[0173] In the above step S41, the electronic device acquires a sample infrared image and a sample visible light image in the same scene. For the convenience of description, the acquired sample infrared image and sample visible light image in the same scene will be referred to as a sample set hereinafter. The sample infrared image includes a low-resolution first sample infrared image and a high-resolution second sample infrared image, and the sample visible light image is a high-resolution visible light image.

[0174] In the embodiments of the present application, the electronic device may acquire the first sample infrared image, the second sample infrared image, and the sample visible light image in multiple scenes, that is, acquire multiple sample sets. For the convenience of description, hereinafter, the training using one sample set will be taken as an example for illustration, which is not restrictive.

[0175] In the embodiments of the present application, the electronic device may first scale the acquired first sample infrared image, and input the scaled infrared image into the preset network model. The scaled first sample infrared image is a blurred and low-quality infrared image. At this time, the scaled first sample infrared image has the same scale, the same number of pixel points, and the same resolution as the sample visible light image. Hereinafter, the first sample infrared image mentioned will be the scaled first sample infrared image.

[0176] In the embodiments of the present application, the first resolution and the third resolution may be the same or different; the second resolution and the fourth resolution may be the same or different, and this is not limited herein.

[0177] In the above step S42, the electronic device inputs the first sample infrared image and the sample visible light image into the preset network model. The preset network model uses the first sample infrared image and the sample visible light image, takes the sample visible light image as a guidance map, and uses the high-resolution information of the sample visible light image to guide the first sample infrared image for super-resolution reconstruction, and obtains and outputs the fused reconstructed infrared image.

[0178] In the embodiments of the present application, the preset network model can extract the third feature of each pixel in the first sample infrared image and the fourth feature of each pixel in the sample visible light image, calculate the similarity between the first sample infrared image and the sample visible light image at each pixel using the third feature and the fourth feature of each pixel, determine the third weight of the first sample infrared image and the fourth weight of the sample visible light image at each pixel using the similarity between the first sample infrared image and the sample visible light image at each pixel, and fuse each pixel in the first sample infrared image and the sample visible light image using the third weight of the first sample infrared image and the fourth weight of the sample visible light image at each pixel to obtain a reconstructed infrared image.

[0179] The process of the electronic device using the preset network model to perform super-resolution reconstruction on the first sample infrared image and the sample visible light image to obtain a reconstructed infrared image is the same as the process of performing super-resolution reconstruction on the initial infrared image and the visible light image to obtain the first target infrared image or the second target infrared image described above. For specific details, please refer to the relevant descriptions in the above Figures 1 to 3 and other parts for obtaining the first target infrared image or the second target infrared image.

[0180] In step S43 above, the electronic device uses the reconstructed infrared image and the second sample infrared image to obtain the loss (Loss, L) of the preset network model for performing super-resolution reconstruction on the first sample infrared image. The electronic device can calculate the reconstruction loss between the reconstructed infrared image and the second sample infrared image, as shown in formula (16).

[0181] Formula (16)

[0182] Where, represents the reconstruction loss; represents the reconstructed infrared image; represents the second sample infrared image.

[0183] The electronic device can calculate the structural similarity (Structural Similarity, SSIM) loss between the reconstructed infrared image and the second sample infrared image, as shown in formula (17).

[0184] Formula (17)

[0185] Where, represents the structural similarity loss; represents the reconstructed infrared image; represents the second sample infrared image; SSIM(·) represents calculating the structural similarity.

[0186] In the above step S44, after obtaining the reconstruction loss and the structural similarity loss, the electronic device can determine the size relationship between the reconstruction loss and the structural similarity loss and a threshold value to determine whether the preset network model converges.

[0187] In one example, the electronic device can use preset reconstruction weights and structural similarity weights to perform weighted summation on the reconstruction loss and the structural similarity loss, calculate the total loss, and determine whether the total loss is less than a preset first threshold value to determine whether the preset network model converges.

[0188] When the total loss is less than the first threshold value, the electronic device can determine that the preset network model converges, execute step S45, output the preset network model, and perform infrared image processing based on the output preset network model; when the total loss is greater than or equal to the first threshold value, the electronic device can determine that the preset network model does not converge, then adjust the parameters in the preset network model, such as the weights corresponding to multiple scales, etc., and return to execute step S42 to continue training the adjusted preset network model using other obtained sample sets. The electronic device loops through steps S42 to S44 until it determines that the preset network model converges.

[0189] In another example, the electronic device can determine whether the reconstruction loss is less than a preset second threshold value and determine whether the structural similarity loss is less than a preset third threshold value to determine whether the preset network model converges.

[0190] When the reconstruction loss is less than the second threshold value and the structural similarity loss is less than the third threshold value, the electronic device can determine that the preset network model converges; when the reconstruction loss is greater than or equal to the second threshold value, or the structural similarity loss is greater than or equal to the third threshold value, the electronic device can determine that the preset network model does not converge. For the specific process, reference can be made to the relevant descriptions in the above examples.

[0191] In another example, the electronic device can combine the reconstruction loss, the structural similarity loss, and the total loss to determine whether the preset network model converges.

[0192] When the total loss is less than the first threshold value, the reconstruction loss is less than the second threshold value, and the structural similarity loss is less than the third threshold value, the electronic device can determine that the preset network model converges; when the total loss is greater than or equal to the first threshold value, the reconstruction loss is greater than or equal to the second threshold value, or the structural similarity loss is greater than or equal to the third threshold value, the electronic device can determine that the preset network model does not converge. For the specific process, reference can be made to the relevant descriptions in the above examples.

[0193] No limitations are imposed on the values of the first threshold value, the second threshold value, and the third threshold value here.

[0194] In some embodiments, the preset network model may include an adversarial network, and the electronic device may further obtain the adversarial loss between the reconstructed infrared image and the second sample infrared image, as shown in formula (18).

[0195] Formula (18)

[0196] Wherein, represents the adversarial loss; represents the second sample infrared image; represents the reconstructed infrared image, represents the expectation.

[0197] In one example, the above step S44 may be implemented by the following steps: determining whether the preset network model converges according to the sum of the reconstruction loss, the structural similarity loss, and the adversarial loss.

[0198] The electronic device may use the preset reconstruction weight, structural similarity weight, and adversarial weight to perform weighted summation on the reconstruction loss, the structural similarity loss, and the adversarial loss, and calculate the total loss, as shown in formula (19).

[0199] Formula (19)

[0200] Wherein, represents the total loss; represents the reconstruction weight; represents the reconstruction loss; represents the adversarial weight; represents the adversarial loss; represents the structural similarity weight; represents the structural similarity loss.

[0201] Furthermore, the electronic device determines whether the total loss is less than a preset fourth threshold to determine whether the preset network model converges.

[0202] When the total loss is less than the fourth threshold, the electronic device may determine that the preset network model converges, execute step S45, output the preset network model, and perform infrared image processing based on the output preset network model; when the total loss is greater than or equal to the fourth threshold, the electronic device may determine that the preset network model does not converge, then adjust the parameters in the preset network model, such as the weights corresponding to multiple scales and the parameters in the adversarial network, etc., and return to execute step S42, and continue to train the adjusted preset network model using the other obtained sample sets. The electronic device repeatedly executes steps S42 to S44 until it determines that the preset network model converges.

[0203] In another example, the electronic device can determine whether the reconstruction loss is less than a preset fifth threshold, determine whether the structural similarity loss is less than a preset sixth threshold, and determine whether the adversarial loss is less than a preset seventh threshold to determine whether the preset network model converges.

[0204] When the reconstruction loss is less than the fifth threshold, the structural similarity loss is less than the sixth threshold, and the adversarial loss is less than the seventh threshold, the electronic device can determine that the preset network model has converged; when the reconstruction loss is greater than or equal to the fifth threshold, the structural similarity loss is greater than or equal to the sixth threshold, or the adversarial loss is greater than or equal to the seventh threshold, the electronic device can determine that the preset network model has not converged. The specific process can be found in the relevant description in the above example.

[0205] In another example, the electronic device may determine whether a preset network model converges by combining reconstruction loss, structural similarity loss, adversarial loss, and total loss.

[0206] When the total loss is less than the fourth threshold, the reconstruction loss is less than the fifth threshold, the structural similarity loss is less than the sixth threshold and the adversarial loss is less than the seventh threshold, the electronic device can determine that the preset network model has converged; when the total loss is greater than or equal to the fourth threshold, the reconstruction loss is greater than or equal to the fifth threshold, the structural similarity loss is greater than or equal to the sixth threshold, or the adversarial loss is greater than or equal to the seventh threshold, the electronic device can determine that the preset network model has not converged. The specific process can be found in the relevant description in the above example.

[0207] The values of the fourth threshold, the fifth threshold, the sixth threshold and the seventh threshold are not limited here.

[0208] By applying the technical solution provided in the embodiments of the present application, the electronic device uses reconstruction loss, adversarial loss and structural similarity loss for joint optimization, so that the reconstructed infrared image reaches the optimal level in terms of structural similarity, edge clarity and overall visual effect.

[0209] In the technical solution provided in the embodiment of the present application, a visible light image is used as a guide, and an adaptive guided filter and a convolutional neural network are used to perform super-resolution reconstruction on a low-resolution infrared image. That is, the infrared image is scaled, and the characteristic value of each pixel of the scaled infrared image is determined to achieve super-resolution reconstruction of the infrared image. In addition, the characteristic values of pixels at multiple scales are determined to achieve multi-modal guided super-resolution reconstruction, ensuring that the temperature information of the infrared image is consistent with the high-frequency details of the visible light. Without increasing the hardware cost, the processed infrared image has high resolution and high detail expression, and solves the problems of high-frequency detail loss and noise interference in the prior art.

[0210] In addition, during the fusion process of infrared images and visible light images by the guided filter, the local feature similarity between the infrared image and the visible light image is dynamically weighted and adjusted to control the fusion ratio, effectively avoiding the problem of introducing artifacts during the guidance process of the visible light image. At the same time, key temperature and texture information are retained to ensure that the generated image has high resolution and high authenticity.

[0211] Meanwhile, during the process of super-resolution reconstruction of infrared images using the technical solution provided in the embodiments of the present application, the computational complexity is low, the processing speed is fast, and it is more convenient to perform super-resolution reconstruction on infrared images.

[0212] Corresponding to the above embodiments of the infrared image processing method, the embodiments of the present application also provide an infrared image processing device. Refer to Figure 5 , which is a schematic structural diagram of the infrared image processing device provided in the embodiments of the present application. The infrared image processing device includes:

[0213] An acquisition module 51, configured to acquire an initial infrared image with a first resolution and a visible light image with a second resolution, where the first resolution is lower than the second resolution;

[0214] A determination module 52, configured to extract a first feature of each pixel point in the initial infrared image and a second feature of each pixel point in the visible light image; use the first feature and the second feature of each pixel point to calculate the similarity between the initial infrared image and the visible light image at each pixel point; use the similarity between the initial infrared image and the visible light image to determine a first weight of the initial infrared image and a second weight of the visible light image, where the similarity is positively correlated with the first weight, and the sum value of the first weight and the second weight is a preset value;

[0215] A fusion module 53, configured to fuse the initial infrared image and the visible light image using the first weight and the second weight to obtain a first target infrared image with the second resolution;

[0216] The above determination module 52 is specifically configured to: use the similarity between the initial infrared image and the visible light image at each pixel point to determine the first weight of the initial infrared image and the second weight of the visible light image at each pixel point;

[0217] The above fusion module 53 is specifically configured to: use the first weight of the initial infrared image and the second weight of the visible light image at each pixel point to fuse each pixel point in the initial infrared image and the visible light image to obtain a first target infrared image with the second resolution.

[0218] In the technical solution provided by the embodiment of the present application, by using a high-resolution visible light image to guide the super-resolution reconstruction of a low-resolution infrared image, it can be ensured that the temperature information of the infrared image is consistent with the high-frequency details of the visible light image. In addition, by using the similarity between the low-resolution infrared image and the high-resolution visible light image, the weights of the infrared image and the visible light image are dynamically determined when reconstructing the infrared image, and the infrared image and the visible light image are fused using the dynamically determined weights to achieve super-resolution reconstruction. When the similarity between the visible light image and the infrared image is high, the weight tilts towards the infrared image, and more characteristics of the infrared image are retained. When the similarity is low, the weight tilts towards the visible light image, which can effectively prevent the generation of artifacts while retaining key temperature and texture information, ensuring that the generated infrared image has high resolution and high authenticity, and improving the resolution of the infrared image.

[0219] In some embodiments, the above-mentioned fusion module 53 may specifically be used to: perform weighted averaging on the first feature of each pixel point in the initial infrared image and the second feature of each pixel point in the visible light image by using the first weight and the second weight at each pixel point to obtain a first target infrared image with a second resolution.

[0220] In some embodiments, the above-mentioned determination module 52 may specifically be used to: determine the first weight of the initial infrared image and the second weight of the visible light image at multiple scales by using the similarity between the initial infrared image and the visible light image at multiple scales;

[0221] The above-mentioned fusion module 53 may specifically be used to: fuse the initial infrared image and the visible light image by using the first weight and the second weight at each scale to obtain a target sub-image at each scale; fuse the target sub-images at multiple scales by using the weights corresponding to multiple scales to obtain a first target infrared image with a second resolution.

[0222] In some embodiments, the above-mentioned infrared image processing device may further include a processing module, and the processing module is used to: after obtaining the first target infrared image with a second resolution, perform a preset process on the first target infrared image to obtain a second target infrared image, and the preset process includes at least one of the following: denoising processing by using a residual block and detail enhancement processing by using an adversarial network.

[0223] In some embodiments, the above-mentioned infrared image processing device may further include a training module, and the training module is used to:

[0224] Obtain a first sample infrared image, a second sample infrared image, and a sample visible light image, where the resolution of the first sample infrared image is a third resolution, and the resolutions of the second sample infrared image and the sample visible light image are a fourth resolution, and the third resolution is lower than the fourth resolution;

[0225] Input the first sample infrared image and the sample visible light image into a preset network model to obtain a reconstructed infrared image. The preset network model is used to determine the third weight of the first sample infrared image and the fourth weight of the sample visible light image by using the similarity between the first sample infrared image and the sample visible light image. The similarity is positively correlated with the third weight, and the sum of the third weight and the fourth weight is a preset value. Then, fuse the first sample infrared image and the sample visible light image by using the third weight and the fourth weight, and output the reconstructed infrared image.

[0226] Obtain the reconstruction loss and the structural similarity loss between the reconstructed infrared image and the second sample infrared image.

[0227] Determine whether the preset network model converges according to the reconstruction loss and the structural similarity loss.

[0228] In some embodiments, the preset network model includes an adversarial network. The above training module can also be used to: obtain the adversarial loss between the reconstructed infrared image and the second sample infrared image.

[0229] Specifically, the above training module can be used to: determine whether the preset network model converges according to the sum of the reconstruction loss, the structural similarity loss, and the adversarial loss.

[0230] An embodiment of the present application also provides an electronic device, as Figure 6 shown, including a processor 61, a communication interface 62, a memory 63, and a communication bus 64. Among them, the processor 61, the communication interface 62, and the memory 63 complete communication with each other through the communication bus 64.

[0231] The memory 63 is used to store a computer program.

[0232] When the processor 61 is used to execute the program stored in the memory 63, it implements the steps of any of the above infrared image processing methods.

[0233] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0234] The communication interface is used for communication between the above electronic device and other devices.

[0235] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0236] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0237] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above infrared image processing methods are implemented.

[0238] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the infrared image processing methods in the above embodiments.

[0239] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0240] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0241] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, electronic device, computer-readable storage medium, and computer program product, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments.

[0242] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.

Claims

1. An infrared image processing method, characterized in that, The method includes: Obtaining an initial infrared image with a first resolution and a visible light image with a second resolution, where the first resolution is lower than the second resolution; Extracting a first feature of each pixel point in the initial infrared image and a second feature of each pixel point in the visible light image; Calculating the similarity between the initial infrared image and the visible light image at each pixel point by using the first feature and the second feature of each pixel point; Determining a first weight of the initial infrared image and a second weight of the visible light image by using the similarity between the initial infrared image and the visible light image, where the similarity is positively correlated with the first weight, and the sum of the first weight and the second weight is a preset value; Fusing the initial infrared image and the visible light image by using the first weight and the second weight to obtain a first target infrared image with the second resolution; The step of determining the first weight of the initial infrared image and the second weight of the visible light image by using the similarity between the initial infrared image and the visible light image includes: Determining the first weight of the initial infrared image and the second weight of the visible light image at each pixel point by using the similarity between the initial infrared image and the visible light image at each pixel point; The step of fusing the initial infrared image and the visible light image by using the first weight and the second weight to obtain a first target infrared image with the second resolution includes: Fusing each pixel point in the initial infrared image and the visible light image by using the first weight of the initial infrared image and the second weight of the visible light image at each pixel point to obtain a first target infrared image with the second resolution.

2. The method according to claim 1, characterized in that, The step of fusing each pixel point in the initial infrared image and the visible light image by using the first weight of the initial infrared image and the second weight of the visible light image at each pixel point to obtain a first target infrared image with the second resolution includes: Performing weighted averaging on the first feature of each pixel point in the initial infrared image and the second feature of each pixel point in the visible light image by using the first weight and the second weight at each pixel point to obtain a first target infrared image with the second resolution.

3. The method according to claim 1, wherein The step of determining the first weight of the initial infrared image and the second weight of the visible light image by using the similarity between the initial infrared image and the visible light image further includes: Determining the first weight of the initial infrared image and the second weight of the visible light image at multiple scales by using the similarity between the initial infrared image and the visible light image at multiple scales; The step of fusing the initial infrared image and the visible light image by using the first weight and the second weight to obtain a first target infrared image with the second resolution further includes: Fusing the initial infrared image and the visible light image by using the first weight and the second weight at each scale to obtain a target sub-image at each scale; Using the weights corresponding to the multiple scales, fuse the target sub-images at the multiple scales to obtain the first target infrared image with the second resolution.

4. The method according to claim 1, wherein After obtaining the first target infrared image with the second resolution, the method further includes: Performing a preset process on the first target infrared image to obtain a second target infrared image, where the preset process includes at least one of the following: denoising using a residual block and detail enhancement using an adversarial network.

5. The method according to claim 1, wherein The method further includes: Obtaining a first sample infrared image, a second sample infrared image, and a sample visible light image, where the resolution of the first sample infrared image is a third resolution, and the resolutions of the second sample infrared image and the sample visible light image are a fourth resolution, and the third resolution is lower than the fourth resolution; Inputting the first sample infrared image and the sample visible light image into a preset network model to obtain a reconstructed infrared image. The preset network model is used to determine a third weight of the first sample infrared image and a fourth weight of the sample visible light image using the similarity between the first sample infrared image and the sample visible light image. The similarity is positively correlated with the third weight, and the sum of the third weight and the fourth weight is the preset value; using the third weight and the fourth weight, fuse the first sample infrared image and the sample visible light image, and output the reconstructed infrared image; Obtaining a reconstruction loss and a structural similarity loss between the reconstructed infrared image and the second sample infrared image; Determining whether the preset network model converges according to the reconstruction loss and the structural similarity loss.

6. The method according to claim 5, wherein The preset network model includes an adversarial network, and the method further includes: Obtaining an adversarial loss between the reconstructed infrared image and the second sample infrared image; The step of determining whether the preset network model converges according to the reconstruction loss and the structural similarity loss includes: Determining whether the preset network model converges according to the sum of the reconstruction loss, the structural similarity loss, and the adversarial loss.

7. An infrared image processing apparatus, characterized in that, The apparatus includes: An acquisition module, configured to acquire an initial infrared image with a first resolution and a visible light image with a second resolution, where the first resolution is lower than the second resolution; A determination module, configured to extract a first feature of each pixel point in the initial infrared image and a second feature of each pixel point in the visible light image; calculate the similarity between the initial infrared image and the visible light image at each pixel point using the first feature and the second feature of each pixel point; determine a first weight of the initial infrared image and a second weight of the visible light image using the similarity between the initial infrared image and the visible light image, where the similarity is positively correlated with the first weight, and the sum of the first weight and the second weight is the preset value; A fusion module, configured to fuse the initial infrared image and the visible light image using the first weight and the second weight to obtain the first target infrared image with the second resolution; The determining module is specifically configured to: determine a first weight of the initial infrared image and a second weight of the visible light image at each pixel point by using the similarity between the initial infrared image and the visible light image at each pixel point; The fusion module is specifically configured to: fuse each pixel point in the initial infrared image and the visible light image by using the first weight of the initial infrared image and the second weight of the visible light image at each pixel point, so as to obtain a first target infrared image with the second resolution.

8. The device according to claim 7, characterized in that, The fusion module is specifically configured to: perform weighted averaging on the first feature of each pixel point in the initial infrared image and the second feature of each pixel point in the visible light image by using the first weight and the second weight at each pixel point, so as to obtain a first target infrared image with the second resolution.

9. The device according to claim 7, characterized in that, The determining module is specifically configured to: determine the first weight of the initial infrared image and the second weight of the visible light image at multiple scales by using the similarity between the initial infrared image and the visible light image at multiple scales; The fusion module is specifically configured to: fuse the initial infrared image and the visible light image by using the first weight and the second weight at each scale to obtain a target sub-image at each scale; fuse the target sub-images at the multiple scales by using the weights corresponding to the multiple scales, so as to obtain a first target infrared image with the second resolution.

10. The device according to claim 7, characterized in that, The device further includes a processing module, and the processing module is configured to: after obtaining the first target infrared image with the second resolution, perform a preset process on the first target infrared image to obtain a second target infrared image, where the preset process includes at least one of the following: denoising by using a residual block and detail enhancement by using an adversarial network.

11. The device according to claim 7, characterized in that, The device further includes a training module, and the training module is configured to: obtain a first sample infrared image, a second sample infrared image, and a sample visible light image, where the resolution of the first sample infrared image is a third resolution, and the resolutions of the second sample infrared image and the sample visible light image are a fourth resolution, and the third resolution is lower than the fourth resolution; input the first sample infrared image and the sample visible light image into a preset network model to obtain a reconstructed infrared image, where the preset network model is configured to determine a third weight of the first sample infrared image and a fourth weight of the sample visible light image by using the similarity between the first sample infrared image and the sample visible light image, the similarity is positively correlated with the third weight, and the sum of the third weight and the fourth weight is the preset value; fuse the first sample infrared image and the sample visible light image by using the third weight and the fourth weight, and output the reconstructed infrared image; obtain a reconstruction loss and a structural similarity loss between the reconstructed infrared image and the second sample infrared image; determine whether the preset network model converges according to the reconstruction loss and the structural similarity loss.

12. The device according to claim 11, characterized in that, The preset network model includes a confrontation network, and the training module is further configured to: obtain the confrontation loss between the reconstructed infrared image and the second sample infrared image; Specifically, the training module is configured to: determine whether the preset network model converges according to the sum of the reconstruction loss, the structural similarity loss, and the confrontation loss.

13. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used for storing a computer program; The processor is configured to implement the method according to any one of claims 1-6 when executing the program stored on the memory.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Infrared and visible light image fusion method based on multi-scale generative adversarial network

    CN111145131A

  • Image resolution reconstruction method and device, storage medium and electronic equipment

    CN116630152A