High-precision infrared image temperature expression method based on U-Net architecture

By constructing a multi-head self-attention enhanced U-Net architecture, combining moving window feature extraction and multi-level residual connections, and designing a composite network loss function, the problem of insufficient accuracy in temperature detection of infrared images is solved, and high-precision temperature expression and image quality improvement are achieved.

CN120599432BActive Publication Date: 2025-10-10SHANDONG UNIV OF FINANCE & ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511094713.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-10-10
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

The temperature detection accuracy of existing infrared imaging technology is significantly affected by image quality, especially by the uncertainty of object emissivity, instrument errors and environmental complexity, which leads to inaccurate temperature measurement and limits its widespread application.

Method used

A multi-head self-attention enhanced U-Net architecture is adopted, combined with moving window feature extraction and multi-level residual connections, a composite network loss function is designed, infrared image and visual image features are fused, and content-enhanced infrared images are generated through a dense residual convolution module.

Benefits of technology

It significantly improves the accuracy and quality of temperature expression of infrared images and promotes their application in more fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599432B_ABST
    Figure CN120599432B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of infrared images, and particularly provides a high-precision infrared image temperature expression method based on a U-Net architecture. The method comprises the following steps: constructing a multi-head self-attention enhanced U-Net architecture, adopting a moving window feature extraction strategy to capture the correlation of local pixels of an image; constructing a dense residual convolution module based on multi-level residual connection to generate a content enhanced infrared image; and designing a composite network loss function to fuse the features of the infrared image and a corresponding visual image. The method improves the quality of the infrared image, ensures that temperature expression is more accurate, and promotes wider application of the infrared image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of infrared image technology, and in particular to a high-precision infrared image temperature expression method based on a U-Net architecture. Background Art

[0002] Infrared thermal imaging technology is widely used in IoT applications in various fields such as agriculture, industry, construction, and medical services. Due to its advantages such as non-contact, non-destructive testing, wide temperature range, and low power consumption, this technology is gaining increasing attention.

[0003] Infrared thermal imaging technology has made significant progress in improving temperature detection accuracy. Existing technologies can deduce object temperature from infrared images, but the accuracy of the results is significantly affected by image quality. Thermal image formation is highly sensitive to various interference factors, including uncertainty in object emissivity, instrument errors, and environmental complexity. These factors often lead to significant inaccuracies in temperature measurements, hindering the widespread application of infrared imaging. In recent years, deep learning technology has been widely used in various fields, including object detection, image classification, and image segmentation. Although many deep learning-based solutions have been used to achieve temperature measurement and fault detection using infrared images, these advanced methods still have significant room for improvement in terms of accuracy and efficiency. Summary of the Invention

[0004] In view of this, the present invention provides a high-precision infrared image temperature expression method based on the U-Net architecture, which is used to improve the quality of infrared images, ensure more accurate temperature expression, and thus promote its wider application.

[0005] In a first aspect, the present invention provides a high-precision infrared image temperature expression method based on a U-Net architecture, the method comprising:

[0006] Step 1: Build a multi-head self-attention enhanced U-Net architecture and adopt a moving window feature extraction strategy to capture the correlation of local pixels in the image;

[0007] Step 2: Based on step 1, a dense residual convolution module based on multi-level residual connections is constructed to generate content-enhanced infrared images;

[0008] Step 3: Based on step 2, design a composite network loss function to fuse the features of the infrared image and the corresponding visual image.

[0009] Optionally, step 1 includes:

[0010] The multi-head self-attention enhanced U-Net architecture consists of three parts: the encoding network, the decoding network, and the temperature mapping network;

[0011] In the encoding network, the original image is first convolved and normalized twice to generate a global feature matrix. Subsequently, the global feature matrix is ​​input into the moving window enhanced multi-head self-attention SW-MHSA module for downsampling. In the SW-MHSA module, the global feature matrix undergoes a three-stage multi-head self-attention feature extraction process; after each stage, an average pooling operation is performed.

[0012] In the decoding network, the compressed feature matrix is ​​first subjected to the SW-MHSA operation. Then, the output feature matrix is ​​connected to the feature matrix of the same size in the encoding network through skip connections. The connected feature matrix is ​​then subjected to dimensionality reduction and input into the next SW-MHSA module, where it is then expanded through upsampling. After four rounds of SW-MHSA operations, the feature matrix with the same size as the original image is obtained in the decoding stage.

[0013] In the temperature mapping network, a gradual dimensionality reduction process is performed to retain the fine-grained features in the generated feature matrix; a dense residual convolutional network is used for feature fusion to generate infrared images. At the same time, a local full convolutional network is used to map the image content with the calibrated infrared temperature annotation matrix to optimize the pixel value distribution of the infrared image.

[0014] Optionally include:

[0015] Based on the multi-head self-attention MHSA algorithm, a moving window enhanced multi-head self-attention SW-MHSA is designed. After the global feature matrix is ​​input into the SW-MHSA module, it is first normalized at the layer level, and then the internal features are extracted through the SW-MHSA operation to generate a feature matrix of the same size as the input. The generated feature matrix is ​​then normalized at the layer level and enhanced by dense residual convolution through a feedforward network. It is then input again into the next SW-MHSA module to learn the relationship between adjacent windows and extract deep features from the input feature matrix. The output result is again normalized at the layer level and enhanced by dense residual convolution through a feedforward network. Finally, it is connected to the input features through a skip network to output the window image feature representation matrix.

[0016] The structure of the SW-MHSA module, for each image window, the input feature matrix First, it is processed by depth convolution and layer normalization, and then the query is constructed. ,key Sum matrix 、 and Perform a linear transformation, then transform the matrix K Transpose and combine with the matrix Q Multiply to generate the feature matrix ; Then, the feature matrix Apply the softmax function and linear normalization, and then add the matrix Multiply, the generated matrix is ​​connected to the feature map of the original input through the residual connection X Connect to get the output matrix ;

[0017] Each attention head is calculated using the following formula:

[0018] ;

[0019] in, represents transpose; Represents the dimension of the word vector;

[0020] By applying the SW-MHSA module on n blocks, the final result matrix is ​​as follows:

[0021] ;

[0022] in, Represents the output projection matrix.

[0023] Optionally, step 2 includes:

[0024] In the construction of the U-Net architecture, each convolutional layer can directly access the output of all previous layers; by using the 1×1 convolutional layer to further integrate features, the activation function Leaky ReLU and batch normalization Batch Norm are used to implement nonlinear transformations. The output of the layer is:

[0025] ;

[0026] in, represents a 3×3 convolution operation, It means connecting the output feature matrices of all previous layers.

[0027] Optionally, step 3 includes:

[0028] Design three different loss functions, namely mean square error MSE loss , structural similarity SSIM loss and Temperature Mapping TM Loss ; Total loss function Expressed as:

[0029] ;

[0030] in, are hyperparameters used to control the contribution of each loss component;

[0031] Mean Squared Error (MSE) loss It is used to quantify the difference between two similar images at the pixel level as an indicator for evaluating the fidelity of the generated image. Its calculation formula is:

[0032] ;

[0033] in, Indicates the total number of pixels, represents the true value of the pixel, Represents the predicted value of the pixel;

[0034] Structural Similarity SSIM Loss It represents the structural similarity of two images and its calculation formula is:

[0035] ;

[0036] in, and Represents two images; Representing an image The mean of Representing an image The mean of and represents the variance of the image; represents the covariance of the image; and Represents two very small constants, used to avoid division by zero;

[0037] Temperature Mapping™ Loss It is set to the average difference between the reconstructed image and the thermal temperature matrix, which is calculated as:

[0038] ;

[0039] in, and Represent the midpoints of the thermal temperature matrix and the corresponding values ​​in the reconstructed image; Represents a constant used to map pixel values ​​to their corresponding temperatures;

[0040] Finally, a comprehensive optimized loss function is developed, which combines the above three network losses and optimizes the temperature expression ability of reconstructed infrared images according to the error back propagation mechanism of deep neural networks.

[0041] In a second aspect, an embodiment of the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the high-precision infrared image temperature expression method based on the U-Net architecture in the first aspect or any possible implementation of the first aspect.

[0042] In a third aspect, an embodiment of the present invention provides an electronic device comprising: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the device, enable the device to execute the high-precision infrared image temperature expression method based on the U-Net architecture in the first aspect or any possible implementation of the first aspect.

[0043] In the technical solution provided by the present invention, the method includes constructing a multi-head self-attention enhanced U-Net architecture, adopting a moving window feature extraction strategy to capture the correlation of local pixels in the image; constructing a dense residual convolution module based on multi-level residual connections to generate content-enhanced infrared images; designing a composite network loss function to fuse the features of the infrared image with the corresponding visual image. This method improves the quality of the infrared image, ensures more accurate temperature expression, and thus promotes its wider application. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 A flowchart of a high-precision infrared image temperature expression method based on a U-Net architecture provided by an embodiment of the present invention;

[0046] Figure 2 A schematic diagram of a multi-head self-attention enhanced U-Net architecture provided by an embodiment of the present invention;

[0047] Figure 3 A schematic diagram of multi-head self-attention enhanced by a moving window provided in an embodiment of the present invention;

[0048] Figure 4 A schematic diagram of the movement of multi-head self-attention enhanced by a moving window provided in an embodiment of the present invention;

[0049] Figure 5 A schematic diagram of the SW-MHSA module structure provided in an embodiment of the present invention;

[0050] Figure 6 A schematic diagram of a dense residual convolution module provided by an embodiment of the present invention;

[0051] Figure 7 A schematic diagram of the temperature expression accuracy of reconstructed infrared images provided by an embodiment of the present invention;

[0052] Figure 8 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0054] It should be understood that the embodiments described are only a portion of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by persons of ordinary skill in the art without creative work are within the scope of protection of the present invention.

[0055] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings.

[0056] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. Furthermore, the character " / " in this document generally indicates an "or" relationship between the associated objects.

[0057] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0058] Principle of infrared thermal imaging technology: Infrared thermal imaging technology converts the emissivity of natural objects into electrical signals, and then converts these signals into infrared images, detecting and visualizing the distribution of infrared radiation energy density based on the temperature difference between the object and the surrounding environment.

[0059] Typically, an infrared thermal imaging system consists of three components: an optical sensing unit, an infrared detector, and a video signal amplifier. In this system, infrared radiation is first attenuated by the atmosphere and then focused onto the optical sensing unit's infrared detector, which converts the radiation into electrical signals. Because these electrical signals are weak, they are then amplified by a signal amplifier and recorded as an infrared image temperature matrix.

[0060] To visually display temperature variations in infrared images, the temperature matrix is ​​usually converted into a three-channel color image. However, due to the inherent limitations of infrared thermal imaging quality and external interference, the original image often lacks edge and texture details, resulting in low resolution and a lower signal-to-noise ratio than natural images.

[0061] Principle of infrared temperature measurement: According to Stefan Boltzmann's law, when the black body temperature is T, its radiation intensity is Expressed as:

[0062] ;

[0063] in, Indicates the radiation intensity, the proportionality coefficient represents the Stefan-Boltzmann constant.

[0064] Assume that the temperature of an object is , its radiation intensity is:

[0065] ;

[0066] The emissivity of the object As the temperature changes, the temperature of the object Calculated by the following formula:

[0067] ;

[0068] Since the emissivity of an object is less than or equal to 1.0, its temperature is always higher than that of a black body. According to Planck's radiation law, the temperature of the surface of an object is expressed as:

[0069] ;

[0070] in, represents the emissivity of the object, and represent the emissivity and transmittance of the atmosphere respectively; Indicates the surface temperature of an object. 、 represent the ambient temperature and atmospheric temperature respectively; Indicates the temperature measured by an infrared thermal imaging device; the exact object temperature is calculated as:

[0071] ;

[0072] in, Represents a constant, which is related to the operating band.

[0073] According to the above formula, the performance of the infrared thermal imaging system is significantly affected by factors such as atmospheric temperature, emissivity, transmittance and ambient temperature. Therefore, infrared images usually need to be enhanced to achieve accurate temperature expression capabilities.

[0074] Figure 1 The flowchart of the high-precision infrared image temperature expression method based on the U-Net architecture provided by the embodiment of the present invention is as follows: Figure 1 As shown, the method includes:

[0075] Step 1: Construct a multi-head self-attention enhanced U-Net architecture and adopt a moving window feature extraction strategy to capture the correlation of local pixels in the image.

[0076] In an embodiment of the present invention, considering the low quality of the original image, an efficient U-Net architecture combining the SW-MHSA module and jump connections is proposed to enhance the temperature expression capability of the infrared image; the U-Net architecture is used to fully reconstruct the original image. In addition, a moving window-enhanced multi-head self-attention network is developed to improve feature extraction capabilities and transfer detailed information from the original image to multiple scales of the generated feature matrix through jump connections. At the same time, a gradual dimensionality reduction process is designed to maintain the fine-grained features of the generated feature matrix. As a result, the reconstructed infrared image not only has better visual effects, but also has significantly improved temperature expression capabilities.

[0077] In the embodiment of the present invention, step 1 includes:

[0078] The U-Net architecture of the present invention with multi-head self-attention enhancement consists of three parts: Figure 2 As shown, the three parts are encoding network, decoding network and temperature mapping network;

[0079] In the encoding network, the original image is first convolved twice and normalized to generate a global feature matrix. This global feature matrix is ​​then input into the moving window enhanced multi-head self-attention (SW-MHSA) module for downsampling. In the SW-MHSA module, the global feature matrix undergoes a three-stage multi-head self-attention feature extraction process; after each stage, an average pooling operation is performed to reduce the size of the matrix and increase its dimensionality, thereby facilitating the extraction of features from both details and global regions of the original image.

[0080] Accordingly, in the decoding network, the compressed feature matrix is ​​first subjected to SW-MHSA operation to further explore its internal relationship characteristics. Next, the output feature matrix is ​​connected to the feature matrix of the same size in the encoding network through a skip connection to enrich its feature information; then, the connected feature matrix is ​​subjected to dimensionality reduction operation and input into the next SW-MHSA module, and then its size is expanded through upsampling operation to generate a detailed enhanced content matrix; after four rounds of SW-MHSA operation, a series of feature matrices of the same size as the original image are obtained in the decoding stage. It is worth noting that the present invention adopts a skip connection structure to transfer the features of matrices of the same size from the encoding stage to the decoding stage, thereby enhancing the detail information of the final matrix. Specifically, in order to reduce the influence of the original image noise, the skip connection is deliberately omitted in the last layer of the decoding network, so as to better maintain the quality of the generated matrix.

[0081] In the temperature mapping network, a gradual dimensionality reduction process is used to preserve the fine-grained features in the generated feature matrix. Furthermore, a dense residual convolutional network is employed to fuse features and generate infrared images, thereby improving the quality of the reconstructed infrared images. Simultaneously, during this stage, a local fully convolutional network is used to map the image content to a calibrated infrared temperature annotation matrix, optimizing the pixel value distribution of the infrared image. Ultimately, a reconstructed infrared image with accurate temperature representation is obtained.

[0082] The attention mechanism mimics the human ability to selectively focus on key information, enabling neural networks to dynamically adjust their attention weights, thereby enhancing the processing of important information and reducing unimportant content. Its core idea is to calculate the correlation between different parts of the input data and the task objectives and assign different weights based on these correlations. As a result, the network can more effectively extract key information from images, significantly improving model performance.

[0083] In the embodiment of the present invention, based on the multi-head self-attention MHSA algorithm, a moving window enhanced multi-head self-attention SW-MHSA is designed. Compared with the MHSA module, the method of the present invention reduces the memory requirement by adopting the window convolution technology, thereby significantly reducing the computational complexity; Figure 3As shown in the figure, after the global feature matrix is ​​input into the SW-MHSA module, it is first normalized at the layer level, and then the internal features are extracted through the SW-MHSA operation to generate a feature matrix of the same size as the input; the generated feature matrix is ​​then normalized at the layer level, enhanced by dense residual convolution through the feedforward network, and then input into the next SW-MHSA module again to learn the relationship between adjacent windows and extract deep features from the input feature matrix. The output result is again normalized at the layer level, enhanced by dense residual convolution through the feedforward network, and finally connected to the input feature through the jump network to output the window image feature representation matrix. Figure 4 As shown in the figure, since the moving window enhanced multi-head self-attention can extract more associations between related windows while reducing the size of the input matrix, it significantly improves the performance of the MHSA-based feature extraction network.

[0084] The structure of the SW-MHSA module is as follows: Figure 5 As shown, R represents reshape, + and × represent element-wise addition and multiplication, respectively. For each image window, the input feature matrix First, it is processed by Deep Convolutional (Dconv) and Layer Normalization (LN), and then the query is constructed ,key Sum matrix 、 and Perform a linear transformation, then transform the matrix K Transpose and combine with the matrix Q Multiply to generate the feature matrix ; Then, the feature matrix Apply the softmax function and linear normalization (Line Normalization, LN), and then with the matrix Multiply, the generated matrix is ​​connected to the feature map of the original input through the residual connection X Connect to get the output matrix ;

[0085] Each attention head is calculated using the following formula:

[0086] ;

[0087] in, represents transpose; Represents the dimension of the word vector;

[0088] Each head is able to focus on different aspects of the input data, thereby obtaining more image features than a simple weighted average. By applying the SW-MHSA module on n blocks, the final result matrix is ​​as follows:

[0089] ;

[0090] in, Represents the output projection matrix.

[0091] Step 2: Based on step 1, a dense residual convolution module based on multi-level residual connections is constructed to generate content-enhanced infrared images.

[0092] In the embodiment of the present invention, step 2 includes:

[0093] In order to reduce the loss of features in the dimensionality reduction process, a dense residual convolution module based on multi-level residual connections is constructed to generate content-enhanced infrared images. Figure 6 As shown in Figure 2, in the construction of the U-Net architecture, each convolutional layer can directly access the output of all previous layers, thereby promoting the transfer of features and extracting fine features from the input matrix. In addition, by using 1×1 convolutional layers to further integrate features, activation functions (Leaky ReLU, LReLU) and batch normalization (Batch Norm, BN) are used to implement nonlinear transformations. The output of the layer is:

[0094] ;

[0095] in, represents a 3×3 convolution operation, Represents the concatenation of the output feature matrices of all previous layers. Therefore, by using densely connected convolutional blocks, fine-grained features are effectively preserved and content-enhanced infrared images are obtained.

[0096] To further enhance the temperature representation of reconstructed infrared images, this paper employs a fixed-coefficient local convolutional network to map image pixels to temperature values. Specifically, the pixel values ​​of each window are averaged and compared with the average value of correspondingly sized windows in the thermal temperature matrix. Backpropagation is performed using the total average temperature error between the reconstructed infrared image and the thermal temperature matrix as the network loss, thereby optimizing the temperature representation of the reconstructed infrared image.

[0097] Step 3: Based on step 2, design a composite network loss function to fuse the features of the infrared image and the corresponding visual image.

[0098] In this embodiment of the present invention, step 3 includes:

[0099] Design three different loss functions, namely mean square error MSE loss , structural similarity SSIM loss and Temperature Mapping TM Loss ; Total loss function Expressed as:

[0100] ;

[0101] in, are hyperparameters used to control the contribution of each loss component;

[0102] Mean Squared Error (MSE) loss It is used to quantify the difference between two similar images at the pixel level as an indicator for evaluating the fidelity of the generated image. Its calculation formula is:

[0103] ;

[0104] in, Indicates the total number of pixels, represents the true value of the pixel, Represents the predicted value of the pixel;

[0105] Structural Similarity SSIM Loss It represents the structural similarity of two images and its calculation formula is:

[0106] ;

[0107] in, and Represents two images; Representing an image The mean of Representing an image The mean of and represents the variance of the image; represents the covariance of the image; and Represents two very small constants, used to avoid division by zero;

[0108] Temperature Mapping™ Loss It is set to the average difference between the reconstructed image and the thermal temperature matrix, which is calculated as:

[0109] ;

[0110] in, and Represent the midpoints of the thermal temperature matrix and the corresponding values ​​in the reconstructed image; Represents a constant used to map pixel values ​​to their corresponding temperatures. By using the infrared image temperature mapping loss, the temperature expression capability of the reconstructed infrared image is further improved.

[0111] Finally, a comprehensive optimized loss function is developed, which combines the above three network losses and optimizes the temperature expression ability of reconstructed infrared images according to the error back propagation mechanism of deep neural networks.

[0112] I. Experimental setup

[0113] In the experiments, an efficient U-Net architecture combining SW-MHSA modules and skip connections was constructed to evaluate its performance in temperature representation. The infrared image dataset contains 3,000 infrared images of heating pipes with a resolution of 640 × 512 pixels and JPEG format, covering the temperature range from -25°C to 135°C. Furthermore, the infrared image database is divided into three parts: a training part containing 2,400 images, and a validation part and a test part containing 300 images each. To speed up image processing, all infrared images are resized to 320 × 256 pixels. To improve the temperature representation of the reconstructed infrared images, the average temperature values ​​of each 16 × 16 pixel block (a total of 320 blocks) are annotated. These values ​​are used to construct a temperature matrix to optimize the reconstructed infrared images, thereby improving the accuracy of temperature representation. The experiments were conducted on a Dell server equipped with an i7-7700K CPU and 32GB of memory, using an Nvidia RTX 3090i (24GB) graphics card. The training parameters are shown in Table 1.

[0114] Table 1 Training parameters

[0115] ;

[0116] II. Impact of image content enhancement:

[0117] Due to the influence of camera electromagnetic interference and external factors, the initially captured infrared image usually contains a large number of noise points. These noise pixels not only reduce the quality of the infrared image, but also weaken its temperature expression ability. Therefore, it is essential to develop an effective content enhancement method to improve the temperature expression ability of the infrared image. The present application constructs a high-efficiency U-Net architecture combined with a SW-MHSA module and a skip connection, aiming to reduce the noise in the infrared image and enable the reconstructed infrared image to better express the temperature. The experimental results are shown in Table 2, and it can be seen that after the content of the infrared image is enhanced by the method of the present application, the reconstructed infrared image has a significant improvement in temperature expression performance. Specifically, after the content of the infrared image is enhanced, the proportion of blocks with a temperature error less than 1℃ in each image increases by 1.66%; while the proportion of blocks with a temperature error less than 5℃ reaches 99.42%, which is significantly better than the unenhanced image. In addition, the U-Net architecture image reconstruction method proposed in the present application adopts a skip connection technology to transfer the detail information of the original image to the reconstructed image. It is worth noting that in order to reduce the propagation of image noise, the final skip connection is omitted in this method. Therefore, the infrared image noise will be reduced overall, and the correspondence between pixel value and temperature is well established, thereby significantly improving the temperature estimation accuracy of the reconstructed infrared image.

[0118] Table 2 Influence of infrared image temperature expression accuracy

[0119] ;

[0120] III. Temperature expression ability of reconstructed infrared image:

[0121] In order to further evaluate the performance of the proposed scheme, about 50 infrared images were used to verify its temperature expression ability. The results show that more than 91.62% of the blocks in each infrared image achieve an average temperature expression error of less than 1℃, as shown in Table 2. Figure 7 A large number of experimental results show that the infrared image temperature expression scheme based on the U-Net architecture can make the reconstructed infrared image achieve very high temperature expression accuracy. Since this method enriches the texture and edges, it significantly improves the content quality of the reconstructed infrared image. In addition, through multiple network losses, the optimized infrared image minimizes the error between the thermal temperature matrix and the reconstructed infrared image, achieving excellent temperature expression ability. Therefore, the reconstructed infrared image not only has excellent visual effect, but also has very excellent temperature expression ability.

[0122] IV. Conclusion:

[0123] In view of the low resolution and high noise interference of the original image, this paper proposes an efficient U-Net architecture to enhance its temperature expression ability by combining SW-MHSA modules and skip connections. A moving window enhanced multi-head self-attention module is developed to explore the intrinsic characteristics of the original image, and skip connections are used in the proposed U-Net architecture to reduce the noise of the original image. In addition, a composite network loss function is designed to improve the visual quality and temperature expression ability of the reconstructed infrared image. The effects of different SW-MHSA module combinations and composite network loss functions are deeply analyzed and verified in the experiment. Experimental results on a large number of infrared images show that the infrared image reconstruction scheme based on the U-Net architecture not only achieves excellent visual quality in the reconstructed infrared images, but also demonstrates a strong high temperature representation capability.

[0124] To enhance the temperature representation of infrared images for the Internet of Things (IoT), this paper constructs an efficient U-Net architecture that combines a moving window-enhanced multi-head self-attention (SW-MHSA) module with skip connections. The performance of the multi-head self-attention module is enhanced by employing a moving window strategy to capture local pixel correlations. Meanwhile, skip connections are used to preserve details of the original image at multiple scales. Specifically, by removing the last skip connection between the initial feature matrix and the generated feature matrix, noise in the original image is effectively reduced and content features are enhanced, ultimately resulting in a high-quality image. Furthermore, a dense residual convolution module based on multi-level residual connections was designed to preserve more detailed information in the generated feature matrix. Each convolutional layer in this module has direct access to the outputs of all previous layers, facilitating the transfer of key information during dimensionality reduction. This not only promotes the transfer and reuse of features at all scales, but also further enhances the temperature representation capability of the reconstructed infrared image. To preserve the detailed information of the infrared image and thus enhance its temperature representation capability, a composite network loss function was developed, consisting of mean squared error (MSE) loss, structural similarity (SSIM) loss, and temperature mapping (TM) loss. The MSE and SSIM losses are used to fuse the features of the infrared image with the corresponding visual image. The temperature matrix (TM) loss establishes a correlation between pixel values ​​and their corresponding temperatures, significantly improving the texture information, visual quality, and temperature representation capability of the reconstructed infrared image. Compared with existing state-of-the-art methods, the proposed method performs well in terms of temperature representation accuracy in infrared images.

[0125] In the technical solution provided by the present invention, the method includes constructing a multi-head self-attention enhanced U-Net architecture, adopting a moving window feature extraction strategy to capture the correlation of local pixels in the image; constructing a dense residual convolution module based on multi-level residual connections to generate content-enhanced infrared images; designing a composite network loss function to fuse the features of the infrared image with the corresponding visual image. This method improves the quality of the infrared image, ensures more accurate temperature expression, and thus promotes its wider application.

[0126] Each step of the embodiment of the present invention may be performed by an electronic device, including but not limited to a tablet computer, a portable PC, a desktop computer, etc.

[0127] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program. When the program is run, the electronic device where the computer-readable storage medium is located is controlled to execute the above-mentioned embodiment of the high-precision infrared image temperature expression method based on the U-Net architecture.

[0128] Figure 8 A schematic diagram of an electronic device provided by an embodiment of the present invention is shown in FIG. Figure 8 As shown, the electronic device 21 includes: a processor 211, a memory 212, and a computer program 213 stored in the memory 212 and executable on the processor 211. When the computer program 213 is executed by the processor 211, the high-precision infrared image temperature expression method based on the U-Net architecture in the embodiment is implemented. To avoid repetition, they are not described here one by one.

[0129] The electronic device 21 includes, but is not limited to, a processor 211 and a memory 212. Those skilled in the art will understand that Figure 8 It is only an example of the electronic device 21 and does not constitute a limitation of the electronic device 21. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0130] The processor 211 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0131] The memory 212 can be an internal storage unit of the electronic device 21, such as the hard drive or memory of the electronic device 21. The memory 212 can also be an external storage device of the electronic device 21, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 21. Furthermore, the memory 212 can include both the internal storage unit of the electronic device 21 and an external storage device. The memory 212 is used to store computer programs and other programs and data required by the network device. The memory 212 can also be used to temporarily store data that has been output or is about to be output.

[0132] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0133] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for expressing infrared image temperature based on U-Net architecture, characterized in that: The method comprises: Step 1: Build a multi-head self-attention enhanced U-Net architecture and adopt a moving window feature extraction strategy to capture the correlation of local pixels in the image; Step 2: Based on step 1, a dense residual convolution module based on multi-level residual connections is constructed to generate content-enhanced infrared images; Step 3: Based on step 2, design a composite network loss function to fuse the features of the infrared image and the corresponding visual image; The step 1 comprises: The multi-head self-attention enhanced U-Net architecture consists of three parts: the encoding network, the decoding network, and the temperature mapping network; In the encoding network, the original image is first convolved and normalized twice to generate a global feature matrix. Subsequently, the global feature matrix is ​​input into the moving window enhanced multi-head self-attention SW-MHSA module for downsampling. In the SW-MHSA module, the global feature matrix undergoes a three-stage multi-head self-attention feature extraction process; after each stage, an average pooling operation is performed. In the decoding network, the compressed feature matrix is ​​first subjected to the SW-MHSA operation. Then, the output feature matrix is ​​connected to the feature matrix of the same size in the encoding network through skip connections. The connected feature matrix is ​​then subjected to dimensionality reduction and input into the next SW-MHSA module, where it is then expanded through upsampling. After four rounds of SW-MHSA operations, the feature matrix with the same size as the original image is obtained in the decoding stage. In the temperature mapping network, a gradual dimensionality reduction process is used to preserve the fine-grained features in the generated feature matrix. A dense residual convolutional network is used for feature fusion to generate infrared images. At the same time, a local full convolutional network is used to map the image content with the calibrated infrared temperature annotation matrix to optimize the pixel value distribution of the infrared image. The step 3 includes: Design three different loss functions, namely mean square error MSE loss , structural similarity SSIM loss and Temperature Mapping TM Loss ; Total loss function Expressed as: ; in, are hyperparameters used to control the contribution of each loss component; Mean Squared Error (MSE) loss It is used to quantify the difference between two similar images at the pixel level as an indicator for evaluating the fidelity of the generated image. Its calculation formula is: ; in, Indicates the total number of pixels, represents the true value of the pixel, Represents the predicted value of the pixel; Structural Similarity SSIM Loss It represents the structural similarity of two images and its calculation formula is: ; in, and Represents two images; Representing an image The mean of Representing an image The mean of and represents the variance of the image; represents the covariance of the image; and Represents two very small constants, used to avoid division by zero; Temperature Mapping™ Loss It is set to the average difference between the reconstructed image and the thermal temperature matrix, which is calculated as: ; in, and Represent the midpoints of the thermal temperature matrix and the corresponding values ​​in the reconstructed image; Represents a constant used to map pixel values ​​to their corresponding temperatures; Finally, a comprehensive optimized loss function is developed, which combines the above three network losses and optimizes the temperature expression ability of reconstructed infrared images according to the error back propagation mechanism of deep neural networks.

2. The method according to claim 1, characterized in that include: Based on the multi-head self-attention MHSA algorithm, a moving window enhanced multi-head self-attention SW-MHSA is designed; After the global feature matrix is ​​input into the SW-MHSA module, it is first normalized at the layer level, and then the internal features are extracted through the SW-MHSA operation to generate a feature matrix of the same size as the input. The generated feature matrix is ​​then normalized at the layer level and enhanced by dense residual convolution through a feedforward network before being input into the next SW-MHSA module to learn the relationship between adjacent windows and extract deep features from the input feature matrix. The output result is again normalized at the layer level and enhanced by dense residual convolution through a feedforward network. Finally, it is connected to the input features through a skip network to output the window image feature representation matrix. The structure of the SW-MHSA module, for each image window, the input feature matrix First, it is processed by depth convolution and layer normalization, and then the query is constructed. ,key Sum matrix 、 and Perform a linear transformation, then transform the matrix K Transpose and combine with the matrix Q Multiply to generate the feature matrix ; Then, the feature matrix Apply the softmax function and linear normalization, and then add the matrix Multiply, the generated matrix is ​​connected to the feature map of the original input through the residual connection X Connect to get the output matrix ; Each attention head is calculated using the following formula: ; in, represents transpose; Represents the dimension of the word vector; By applying the SW-MHSA module on n blocks, the final result matrix is ​​as follows: ; in, Represents the output projection matrix.

3. The method according to claim 1, characterized in that The step 2 includes: In the construction of the U-Net architecture, each convolutional layer can directly access the output of all previous layers; by using the 1×1 convolutional layer to further integrate features, the activation function Leaky ReLU and batch normalization Batch Norm are used to implement nonlinear transformations. The output of the layer is: ; in, represents a 3×3 convolution operation, It means connecting the output feature matrices of all previous layers.

4. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the infrared image temperature expression method based on the U-Net architecture according to any one of claims 1 to 3.

5. An electronic device, characterized in that: include: one or more processors; Memory; And one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, enable the device to execute the infrared image temperature expression method based on the U-Net architecture according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Method for performing super-resolution reconstruction on infrared image of power equipment

    CN118864250A

  • Shield tunnel thermal disaster diagnosis method and system based on deep learning

    CN119862450A