Single-frame infrared thermal image enhancement method and system
By simultaneously acquiring infrared thermal images and visible light images, calculating virtual diffusion coefficients and generating Gaussian diffusion kernels, performing convolution and Fourier transforms, and constructing a conditional generative adversarial network, the problems of noise, artifacts, and defocusing in single-frame infrared image enhancement technology are solved, achieving high-precision thermal flux assessment and image enhancement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN UNIV
- Filing Date
- 2026-04-29
- Publication Date
- 2026-06-02
AI Technical Summary
Existing single-frame infrared image enhancement technologies struggle to achieve high-precision thermal flux assessment under varying hardware performance, environmental interference, and scene conditions. They suffer from severe noise and artifact interference, making it difficult to balance physical models and data-driven methods. Furthermore, they fail to meet the needs of image defocus restoration and multispectral detection, thus failing to satisfy the detection requirements of complex application scenarios.
By simultaneously acquiring infrared thermal images and visible light images, calculating virtual diffusion coefficients, generating Gaussian diffusion kernels and convolution kernels, performing convolution operations and three-dimensional Fourier transforms, constructing a conditional generative adversarial network, incorporating multispectral scene conditional information, training the network with a dual loss function, and outputting a target-enhanced thermal flow image.
It improves image quality, enhances the physical rationality and positioning accuracy of heat flow inversion, suppresses noise and halo phenomena, ensures that the enhancement results conform to thermophysical laws, and improves the reliability and comprehensiveness of the detection results.
Smart Images

Figure CN122134588A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of multispectral infrared image processing, and in particular to a method and system for enhancing single-frame infrared thermal images. Background Technology
[0002] Infrared thermal imaging technology, with its core advantages of being non-contact, non-destructive, and capable of penetrating some obstructions, has become an indispensable key technology in fields such as industrial defect detection, power equipment fault diagnosis, aerospace component quality inspection, and electronic component reliability assessment. This technology converts the thermal radiation signal of a target into a visual image, enabling efficient monitoring of equipment status and product quality. However, in practical engineering applications, detection schemes relying on single-frame infrared images often suffer from image quality limitations imposed by hardware performance, environmental interference, and scene conditions, making it difficult to meet the ever-increasing demands for high-precision diagnosis. This leads to missed defects and increased errors in thermal flow assessment. Current technology mainly faces the following bottlenecks: 1. Single-frame images have limited information dimensions, making accurate thermal physics inversion difficult. A single-frame infrared image can only capture the static heat distribution of the target surface at a specific moment, completely lacking dynamic information about heat conduction over time. This limitation makes it extremely difficult to invert the heat flow transfer process inside the target from a single-frame image. For example, when assessing corrosion defects on the inner wall of a pipe, it is impossible to determine the depth and extent of the defect based on the trend of heat flow changes. At the same time, image quality is severely limited by hardware: the camera pixel width restricts spatial resolution, causing micron-level defect features to be easily lost; while the integration time setting faces a dilemma—too short a time results in insufficient signal and blurred image, while too long a time will miss transient thermal processes.
[0003] 2. Severe noise and artifact interference presents inherent contradictions in traditional image processing methods. Infrared imaging is susceptible to complex noise interference, including random electromagnetic noise in industrial environments, inherent fixed-pattern noise of sensors, and low-frequency noise introduced by ambient temperature fluctuations. This noise significantly reduces the signal-to-noise ratio of the image. Existing traditional filtering and enhancement methods based on the spatial or frequency domains have limited effectiveness in dealing with such interference: methods such as Gaussian filtering and median filtering, while suppressing noise, often smooth out key details or disrupt the physical continuity of thermal distribution. More significantly, when using frequency domain transformation for processing, the "halo" artifact phenomenon caused by factors such as signal truncation can create false bright or dark edges at defect edges, leading to serious deviations in defect location and size assessment.
[0004] 3. The disconnect between physical models and data-driven methods makes it difficult to balance physical realism and scene adaptability. Existing single-frame infrared image enhancement techniques mainly follow two paths, each with significant drawbacks. First, enhancement methods based on physical models such as pseudospectral methods and finite element methods, while ensuring the physical validity of the results based on the heat conduction equation, typically have fixed model parameters, making it difficult to adapt to complex real-world scenes and diverse noise types. Second, data-driven methods based on deep learning such as conditional generative adversarial networks and U-Net, while demonstrating strong learning capabilities in noise suppression and detail restoration, may produce results that deviate from physical laws, such as generating temperature values that violate the laws of thermal radiation or incorrect heat flow directions. This results in low reliability of thermal evaluation results in scenarios not covered by training data.
[0005] 4. The mismatch between single-frame image defocus restoration and the requirements of multispectral detection. In multispectral infrared detection, switching between different wavelength filters can easily cause focal length shift, resulting in image defocusing and blurred boundaries. In scenarios where only single-frame images can be acquired (such as high-speed moving part detection and high-risk charged object detection), existing single-frame refocusing techniques often only pursue improved visual clarity while neglecting the authenticity of the thermophysical properties of the restored image, failing to meet the accuracy requirements of quantitative heat flow analysis.
[0006] In summary, as industries such as aerospace, power, and electronics place increasingly stringent demands on detection accuracy, temperature resolution, and result reliability, traditional single-frame infrared image enhancement methods, due to the aforementioned limitations, are no longer sufficient to meet the needs of complex application scenarios. Therefore, the industry urgently requires a novel technical solution that integrates the rigor of physical models with data-driven adaptability, effectively suppressing noise and eliminating halo artifacts while ensuring that the enhancement results conform to thermophysical laws, thereby achieving the goal of high-quality enhancement and accurate thermophysical inversion from single-frame infrared images. Summary of the Invention
[0007] This application proposes a method and system for enhancing single-frame infrared thermal images to address the shortcomings of the prior art.
[0008] According to a first aspect of the embodiments of this application, a method for enhancing a single-frame infrared thermal image is provided, comprising: S1: Simultaneously acquire the original single-frame infrared thermal image and the corresponding visible light image, and calculate the virtual diffusion coefficient based on the integration time of the original single-frame infrared thermal image, the camera pixel width, and the thermal radiation characteristics of the selected multispectral filter. S2: Based on the virtual diffusion coefficient, a Gaussian diffusion kernel and a sequence convolution kernel are generated, and a convolution kernel sequence adapted to the thermal diffusion characteristics of multispectral bands is formed. S3: Add white noise to the original single-frame infrared thermal image to generate a noisy image sequence, and perform a convolution operation between the noisy image sequence and the convolution kernel sequence to obtain a smooth image sequence; S4: Perform residual calculation, augmentation processing, and three-dimensional Fourier transform on the smoothed image sequence, and use dot product kernels to correct the frequency domain data; S5: The corrected frequency domain data is inversely transformed back to the spatiotemporal domain, the effective image data is filtered and the process is repeated multiple times. The halo effect of the image is eliminated by summing multiple sets of data to obtain an intermediate image that meets the target quality. S6: Construct a conditional generative adversarial network with a U-shaped network as the generator, incorporate multispectral scene condition information, and use dual loss function constraints to train the conditional generative adversarial network. Input the intermediate image into the conditional generative adversarial network after network training. S7: The target enhanced thermal flow image is output by generating an adversarial network based on the conditions, and the original single-frame infrared thermal image and the corresponding visible light image are compared, displayed and stored with the target enhanced thermal flow image. The visible light image is used to assist in the analysis of scene texture and environmental factors.
[0009] In some implementations, the simultaneous acquisition of the original single-frame infrared thermal image and the corresponding visible light image includes: The image acquisition system simultaneously acquires the original single-frame infrared thermal image and the corresponding visible light image. The image acquisition system uses a microcontroller as the core control unit and is equipped with an infrared camera, a motor-driven filter wheel, and a visible light camera. It is used to control the rotation of the filter wheel through a preset program, ensure that the rotation of the filter wheel is synchronized with the timing of the camera acquisition, and acquire image sequences containing slide-free data and multispectral data.
[0010] In some implementations, calculating the virtual diffusion coefficient based on the integration time of the original single-frame infrared thermal image, the camera pixel width, and the thermal radiation characteristics of the selected multispectral filter includes: The initial data for the virtual diffusion coefficient are calculated based on the following formula:
[0011] in, Represents the virtual diffusion coefficient; This indicates the width of the camera pixels. This represents the integration time of the original single-frame infrared thermal image; The initial data of the virtual diffusion coefficient are fine-tuned based on the thermal diffusion rate corresponding to the average wavelength of the selected multispectral filter.
[0012] In some implementations, the Gaussian diffusion core is calculated based on the following expression:
[0013] The sequence convolution kernel is calculated and generated based on the following expression:
[0014] in, It is an exponential function with the natural constant e as its base; Indicates radial distance; Indicates the diffusion coefficient; Indicates the fundamental time constant; Represents a continuously changing time variable. The value of is Discretize into units. The range of values is ,2 , ..., (N+1) N is an integer determined based on the differences in thermal response across multiple spectral bands.
[0015] In some embodiments, the method further includes: The standard deviation of the added white noise is dynamically determined based on the average residual between the original single-frame infrared thermal image and the first image in the smoothed image sequence. The standard deviation of the white noise to be added is adaptively adjusted according to the type of filter selected. When the filter type is a short-pass filter, the standard deviation of the white noise is greater than the standard deviation of the white noise when the filter type is a long-pass filter or a band-pass filter.
[0016] In some implementations, the steps of performing residual calculation, augmentation, and three-dimensional Fourier transform on the smoothed image sequence, and correcting the frequency domain data using a dot product kernel, include: The smoothed image sequence is subjected to residual calculation to obtain the residual image sequence; The residual image sequence is expanded by 20% in both the row and column directions of the image coordinates; A three-dimensional Fourier transform is performed on the expanded residual map sequence to convert the spatiotemporal domain signal into a frequency domain signal; The transformed frequency domain signal is corrected using a dot product kernel, the expression of which is shown below:
[0017] in, This represents the length frequency corresponding to the Fourier transform. This represents the width frequency corresponding to the Fourier transform. The value is represented as The imaginary unit, The dot product core represents the time frequency corresponding to the Fourier transform. It is used to capture the spatial, temporal, and band correlation features in the multispectral frequency domain and to specifically correct the differences in frequency domain energy distribution of different filters.
[0018] In some implementations, the step of inversely transforming the corrected frequency domain data back to the spatiotemporal domain, filtering valid image data and repeating the process multiple times, and eliminating image halo phenomena through multiple sets of data summation operations to obtain an intermediate image that meets the target quality includes: The corrected frequency domain data is restored to the spatiotemporal domain through inverse three-dimensional Fourier transform, and the transformation result containing the data in the spatiotemporal domain is obtained. Filter and retain the first 5 valid image data from the transformation results; Repeat steps S4 to S6 a total of m times, where the specific value of m is determined by the number of filters and the image noise level. The m×5 images obtained from m repeated processing are summed along the direction of the number of processing steps m, and the band-specific halos of different filters are canceled out by statistical averaging effect to obtain an intermediate image that meets the target quality.
[0019] In some implementations, the construction of a conditional generative adversarial network (GAN) using a U-shaped network as a generator, incorporating multispectral scene conditional information and employing dual loss function constraints to train the GAN, and inputting the intermediate image into the trained GAN, includes: The conditional generative adversarial network is constructed, wherein the generator adopts a U-shaped network structure with symmetrical encoding and decoding, and is equipped with multi-scale feature fusion skip connections; When training the conditional generative adversarial network, multispectral scene condition information is incorporated, including wavelength parameters of eight filters, heat flux labels corresponding to infrared images, camera integration time, and pixel width. The conditional generative adversarial network is trained using a dual loss function constraint, and the intermediate image is input into the conditional generative adversarial network after the network training is completed. The dual loss function includes defining the GAN loss and defining... The loss, defined as the GAN loss, is calculated using the following expression:
[0020] The definition The loss is calculated using the following expression:
[0021] Among them, the This represents calculating the mathematical expectation over a large number of samples; the... This indicates the conditional information of the input generator; the Represents a real image; the This indicates that the generator is based on conditional information. The generated image; the This indicates that the discriminator recognizes the real image. Corresponding condition information The discrimination probability; This indicates that the discriminator evaluates the generated image. Corresponding condition information The discrimination probability; Represents generator of Loss value.
[0022] In some implementations, the stored processing parameters include virtual diffusion coefficient, white noise standard deviation, GAN loss value, etc. The loss values and corresponding filter parameters are compared and displayed, including the original single-frame infrared thermal images acquired with different filters and the enhanced and deblurred single-frame infrared thermal images.
[0023] According to a second aspect of this application, a single-frame infrared thermal image enhancement system is provided, comprising: The original image acquisition module is used to simultaneously acquire the original single-frame infrared thermal image and the corresponding visible light image, and calculate the virtual diffusion coefficient based on the integration time of the original single-frame infrared thermal image, the camera pixel width, and the thermal radiation characteristics of the selected multispectral filter. The convolution kernel sequence generation module is used to generate Gaussian diffusion kernels and sequence convolution kernels based on the virtual diffusion coefficient, and form a convolution kernel sequence adapted to the thermal diffusion characteristics of multispectral bands. The smooth image sequence generation module is used to add white noise to the original single-frame infrared thermal image to generate a noisy image sequence, and to perform a convolution operation between the noisy image sequence and the convolution kernel sequence to obtain a smooth image sequence. The smoothed image sequence transformation module is used to perform residual calculation, augmentation processing, and three-dimensional Fourier transform on the smoothed image sequence, and to correct the frequency domain data using the dot product kernel; The intermediate image generation module is used to inversely transform the corrected frequency domain data back to the spatiotemporal domain, filter effective image data and repeat the process multiple times, and eliminate the image halo phenomenon through the summation operation of multiple sets of data to obtain an intermediate image that meets the target quality. The Conditional Generative Adversarial Network (CGAN) processing module is used to construct a CGAN with a U-shaped network as the generator, incorporate multispectral scene condition information, and use dual loss function constraints to train the CGAN. The intermediate image is input into the CGAN trained by the network. The target-enhanced thermal flux image generation module is used to output a target-enhanced thermal flux image through the conditional generative adversarial network, and to compare, display and store the original single-frame infrared thermal image and the corresponding visible light image with the target-enhanced thermal flux image. The visible light image is used to assist in the analysis of scene texture and environmental factors.
[0024] The beneficial effects of the single-frame infrared thermal image enhancement method and system of this application embodiments include at least the following: This application embodiment ensures the temporal consistency of data acquisition by simultaneously acquiring original single-frame infrared thermal images and visible light images, providing complete input for multispectral analysis. The virtual diffusion coefficient is calculated based on integration time, pixel width, and the thermal radiation characteristics of the filter, making the heat conduction model conform to actual hardware and optical conditions, thus improving the accuracy of subsequent physical modeling. This application embodiment generates Gaussian diffusion kernels and sequential convolution kernels to form a convolution kernel sequence adapted to multispectral bands, capable of simulating the thermal diffusion characteristics of different bands. Time-series expansion compensates for the lack of dynamic temporal information in single-frame images, enhancing the physical rationality of heat flow inversion. This application embodiment generates noisy image sequences by adding white noise, improving the model's robustness to real noise; convolution operations smooth the image sequences, effectively suppressing random noise while preserving core heat flow features, providing a high-quality data foundation for frequency domain processing. This application embodiment highlights image differences and avoids boundary effects through residual calculation and augmentation processing; three-dimensional Fourier transform combined with dot product kernels corrects frequency domain data, accurately capturing spatial-temporal-band correlations, optimizing the separation and enhancement of heat flow features. This application embodiment eliminates halo phenomena by inversely transforming back to the spatiotemporal domain and filtering effective data, and by performing multiple rounds of summation operations. It utilizes statistical averaging to cancel band-specific artifacts, obtaining high-quality intermediate images and improving the clarity and positioning accuracy of defect edges. This application embodiment constructs a conditional generative adversarial network with a U-shaped network as the generator, incorporating multispectral scene condition information to ensure the enhancement process conforms to thermophysical laws. Dual loss function constraints are used in training to balance visual realism and pixel-level accuracy, outputting enhanced images with rich detail and physical realism. This application embodiment outputs enhanced thermal flow images of the target and compares them with the original infrared and visible light images for intuitive evaluation of the enhancement effect. The visible light image assists in analyzing scene texture and environmental factors, improving the reliability and comprehensiveness of the detection results and supporting subsequent defect identification decisions. Attached Figure Description
[0025] Figure 1This is a flowchart illustrating the single-frame infrared thermal image enhancement method according to an embodiment of this application. Figure 2 This is a flowchart of the synchronization triggering control program according to an embodiment of this application; Figure 3 This is a structural diagram of the UNet network according to an embodiment of this application; Figure 4 This is a circuit diagram showing the improved treaty connection according to an embodiment of this application; Figure 5 A structural diagram of the conditional generation adversarial network for embodiments of this application; Figure 6 This is a flowchart illustrating a specific implementation of the single-frame infrared thermal image enhancement method according to an embodiment of this application. Figure 7 This is a hardware structure diagram illustrating the single-frame infrared thermal image enhancement system according to an embodiment of this application; Figure 8(a) is a schematic diagram of a visible light image of an embodiment of this application; Figure 8(b) is a schematic diagram of infrared thermal imaging of an embodiment of this application; Figure 9(a) is the original infrared grayscale image of an embodiment of this application; Figure 9(b) is a schematic diagram of the data-enhanced image according to an embodiment of this application. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the sulfur-containing polymer electrolyte and its preparation method will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0027] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed embodiments of the present application, but merely to illustrate selected embodiments of the present application. Other embodiments obtained by those skilled in the art based on the embodiments of the present application without inventive effort are all within the scope of protection of the embodiments of the present application.
[0028] It can be noted that similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it will not be further defined and explained in subsequent figures according to the embodiments of this application.
[0029] This application discloses a method and system for enhancing single-frame infrared thermal images. The method is implemented based on an ultra-low frequency polarity adaptive circuit structure system. The purpose is to specifically solve the key problems of single-frame infrared images under the self-built multispectral rotary infrared data acquisition system, such as blurred details, noise interference, halo phenomenon, insufficient physical authenticity, and poor defocusing repair effect caused by filter switching.
[0030] See attached document Figure 1 As shown, this single-frame infrared thermal image enhancement method includes the following steps S1-S7.
[0031] S1: Simultaneously acquire the original single-frame infrared thermal image (which can also be understood as the original captured image) and the corresponding visible light image, and calculate the virtual diffusion coefficient based on the integration time of the original single-frame infrared thermal image, the camera pixel width, and the thermal radiation characteristics of the selected multispectral filter. .
[0032] Wherein, the virtual diffusion coefficient It is the core bridge connecting camera hardware parameters and the physical laws of heat conduction. The pixel width w determines the image spatial resolution, and the integration time... The virtual diffusion coefficient reflects the time-based acquisition characteristics of the thermal signal. The wavelength differences in the multispectral filter lead to varying thermal radiation energy distributions, thus affecting the calculated virtual diffusion coefficient. The virtual diffusion coefficient needs to be adapted to both hardware acquisition characteristics and multispectral thermal radiation characteristics, and is applicable to eight different types of filters. It will automatically fine-tune according to the thermal diffusion rate corresponding to the average wavelength of the filter to ensure the physical rationality of subsequent heat conduction simulation.
[0033] In some implementations, calculating the virtual diffusion coefficient based on the integration time of the original single-frame infrared thermal image, the camera pixel width, and the thermal radiation characteristics of the selected multispectral filter includes: The initial data for the virtual diffusion coefficient are calculated based on the following formula:
[0034] in, Represents the virtual diffusion coefficient; This indicates the width of the camera pixels. This represents the integration time of the original single-frame infrared thermal image. The initial data of the virtual diffusion coefficient is fine-tuned based on the thermal diffusion rate corresponding to the average wavelength of the selected multispectral filter. For example, during the calculation process, the coefficient needs to be adaptively fine-tuned based on the thermal diffusion rate corresponding to the average wavelength of eight filters (including bandpass, long pass, and short pass types) to adapt to the differences in thermal radiation characteristics of different bands.
[0035] In one exemplary embodiment, the synchronous acquisition of the original single-frame infrared thermal image and the corresponding visible light image includes: building a multispectral infrared thermal imaging data acquisition system, controlling the multispectral infrared thermal imaging data acquisition system to complete the synchronous acquisition of single-frame infrared thermal images and visible light images at multiple time points, and acquiring the original infrared thermal image without a glass slide, the multi-band infrared thermal image, and the corresponding visible light scene image.
[0036] In some implementations, the simultaneous acquisition of the original single-frame infrared thermal image and the corresponding visible light image includes: simultaneously acquiring the original single-frame infrared thermal image and the corresponding visible light image based on the image acquisition system.
[0037] The image acquisition system (which can also be understood as a multispectral infrared thermal imaging data acquisition system) uses a microcontroller as the core control unit and is equipped with an infrared camera, a motor-driven filter wheel, and a visible light camera. It is used to control the rotation of the filter wheel through a preset program, and to ensure that the rotation of the filter wheel is synchronized with the timing of the camera acquisition, and to acquire image sequences containing slide-free data and multispectral data.
[0038] For example, the image acquisition system includes an infrared camera (e.g., MAG62), a development board (e.g., Arduino), a motor-driven filter wheel, and eight different types of filters.
[0039] For example, the image acquisition system uses an Arduino UNO development board as the core control unit, equipped with a MAG62 infrared camera, a motor-driven filter wheel (containing 8 different types of filters), and a visible light camera. The direction pin, pulse pin, camera trigger pin, and status indicator pin of the development board are respectively connected to the motor driver direction control terminal, pulse signal input terminal, camera external trigger interface, and system status indicator. The system uses a falling edge trigger mode to control camera acquisition. Through program design, the timing of filter wheel rotation and camera acquisition is synchronized to ensure that 10 images, including 2 images without a glass slide and 8 images of multispectral data, are accurately acquired with one rotation of the turntable.
[0040] For example, when controlling the multispectral infrared thermal imaging data acquisition system to perform multi-time point acquisition, the acquisition time sequence covers early morning, noon, and evening, and the acquisition months cover February-March and April-June. Among them, the acquisition is divided into segments according to specific time intervals from 6:00 to 21:45, and each segment is continuously photographed 3 times to study the impact of sunlight excitation on infrared defect detection under different seasons and times.
[0041] Among these, the synchronous control of the multispectral filter rotary table motor and the MAG62 infrared camera during data acquisition is particularly important. To ensure that the filter wheel completes one rotation in less than 3 seconds, and that data acquisition occurs simultaneously when the filter center is aligned with the camera lens, this embodiment of the application incorporates a synchronous trigger control program design for this process. Specifically, this embodiment achieves timing synchronization between the filter wheel rotation control and the infrared camera through hardware interconnection and algorithmic design. This ensures that for every rotation of the rotary table, the infrared camera accurately acquires 10 images, including 2 images without a slide and 8 multispectral images. (See attached document.) Figure 2 As shown, the exemplary synchronous trigger control program design includes: after power-on, the control program performs initialization, that is, first sets the motor stepper pulse pin, direction pin, camera trigger pin, and LED pin as outputs, fixes the direction as HIGH, sets the camera trigger to high by default, and turns off the LED, and prints "System Ready" via the serial port. Then, it starts rotation, enters the main loop, and executes the main logic only once on the first run. Then it cycles through 10 segments. For each segment, it first sends 320 step pulses at 625µs / step (each step first pulls the step pin high for 2µs and then low). The motor rotates about 36°, which takes about 200ms. Then it pauses for 1000ms, pulls the camera trigger pin low for 10ms and then high to complete one shot, pauses again for 1000ms, and prints the current segment number. When the 10 segments are completed, that is, the turntable rotates one revolution (i.e., 360°), it prints a completion message and lights up the LED. Since it has been marked as running, no further actions are performed in subsequent loops.
[0042] S2: Based on the virtual diffusion coefficient, a Gaussian diffusion kernel and a sequence convolution kernel are generated, and a convolution kernel sequence adapted to the thermal diffusion characteristics of multispectral bands is formed.
[0043] This step, based on the virtual diffusion coefficient obtained in step S1, first generates a basic Gaussian diffusion kernel to characterize the instantaneous thermal diffusion state at the moment of single-frame acquisition, accurately matching the temporal dimension characteristics of multispectral single-frame images. Then, the maximum integer multiple of the acquisition time and the image integration time is taken to generate a sequence convolution kernel to ensure complete coverage of the thermal diffusion state under different multispectral bands, providing support for generating continuous thermal imaging sequences.
[0044] In some implementations, the Gaussian diffusion core is calculated based on the following expression:
[0045] The sequence convolution kernel is calculated and generated based on the following expression:
[0046] in, It is an exponential function with the natural constant e as its base; Indicates radial distance; Represents the virtual diffusion coefficient; Indicates the fundamental time constant; Represents a continuously changing time variable. The value of is Discretize into units. The range of values is ,2 , ..., (N+1) N is an integer determined based on the differences in thermal response across multiple spectral bands, for example, in the embodiments of this application. The value range of N is determined by combining the thermal conduction time scale and the differences in thermal response of multispectral bands. For example, the thermal response of the long-wave infrared band is slower, so the value of N is larger, while the thermal response of the short-wave infrared band is faster, so N can be appropriately reduced to ensure that the sequence convolution kernel can completely cover the thermal diffusion state under different multispectral bands, thus providing support for generating continuous thermal imaging sequences.
[0047] S3: Add white noise to the original single-frame infrared thermal image to generate a noisy image sequence, and perform a convolution operation between the noisy image sequence and the convolution kernel sequence to obtain a smooth image sequence.
[0048] In some embodiments, the method further includes: dynamically determining the standard deviation of the added white noise based on the average residual between the original single-frame infrared thermal image and the first image in the smoothed image sequence; adaptively adjusting the value of the standard deviation of the white noise to be added according to the selected filter type, wherein the standard deviation of the white noise is greater when the filter type is a short-pass filter than when the filter type is a long-pass filter or a band-pass filter.
[0049] For example, the number of smooth image sequences output after each convolution operation increases by 1 (i.e., N+1).
[0050] For example, to improve the anti-interference capability of the embodiments of this application in multispectral scenarios and to simulate the noise characteristics under different filters, the embodiments of this application randomly add m standard deviations to the original single-frame infrared image. The white noise is used to generate m infrared images containing white noise, where, This is the average residual between the original acquired image and the first image in the Gaussian smoothed array, and it needs to be dynamically adjusted according to the filter type. For example, images acquired using a short-pass filter have higher noise levels. The values are increased accordingly to ensure that the noise addition intensity matches the actual noise distribution of the multispectral image. Then, each infrared thermogram in the image sequence is multiplied by a Gaussian sequence convolution to obtain N+1 smooth image sequences. This step can effectively suppress band-specific noise in the multispectral image while preserving the core features of heat flow, laying a high-quality data foundation for subsequent frequency domain transformation.
[0051] S4: Perform residual calculation, augmentation processing, and three-dimensional Fourier transform on the smoothed image sequence, and use dot product kernels to correct the frequency domain data.
[0052] In some embodiments, the process of performing residual calculation, augmentation, and three-dimensional Fourier transform on the smoothed image sequence, and correcting the frequency domain data using a dot product kernel, includes: performing residual calculation on the smoothed image sequence to obtain a residual map sequence, which reflects image differences; augmenting the residual map sequence by 20% along both the row and column directions of the image coordinates to ensure that the thermal radiation characteristics of different filter edges are not lost during the transformation; performing a three-dimensional Fourier transform on the augmented residual map sequence to convert the spatiotemporal domain signal into a frequency domain signal; and correcting the transformed frequency domain signal using a dot product kernel, the expression of which is shown below:
[0053] in, This represents the length frequency corresponding to the Fourier transform. This represents the width frequency corresponding to the Fourier transform. The value is represented as The imaginary unit, This represents the time frequency corresponding to the Fourier transform. The dot product kernel is used to capture the spatial, temporal, and band correlation characteristics in the multispectral frequency domain and to specifically correct the differences in frequency domain energy distribution among different filters. Based on the above formula, the correlation characteristics of spatial and temporal bands in the multispectral frequency domain can be accurately captured, and the differences in frequency domain energy distribution among different filters can be specifically corrected. S5: The corrected frequency domain data is inversely transformed back to the spatiotemporal domain, the effective image data is filtered and the process is repeated multiple times. The halo effect of the image is eliminated by summing multiple sets of data to obtain an intermediate image that meets the target quality.
[0054] In some implementations, the step of inversely transforming the corrected frequency domain data back to the spatiotemporal domain, filtering valid image data, and repeating the process multiple times to eliminate image halo phenomena through multiple sets of data summation operations to obtain an intermediate image that meets the target quality includes: restoring the corrected frequency domain data to the spatiotemporal domain through an inverse three-dimensional Fourier transform and obtaining a transformation result containing data from the spatiotemporal domain; filtering and retaining the first 5 valid image data from the transformation result; repeating steps S4 to S6 a total of m times, where the specific value of m is determined by the number of filters and the image noise level; summing the m×5 images obtained from the m repetitions along the direction of the number of processing times m, and using the statistical averaging effect to cancel the band-specific halo of different filters to obtain an intermediate image that meets the target quality.
[0055] For example, the process of inversely transforming the corrected frequency domain data back to the spatiotemporal domain, filtering effective image data, and repeating the process multiple times, and eliminating image halo phenomena through multiple sets of data summation to obtain an intermediate image that meets the target quality, includes: restoring the frequency domain data after dot product processing to spatiotemporal space through inverse three-dimensional Fourier transform; considering the effectiveness and redundancy of the data, retaining the first 5 data images of the array, which is the optimal choice verified by a large number of experiments and can most accurately extract the core features of heat flow in multispectral scenes; repeating steps 4-6 a total of m times, where the value of m needs to be determined in combination with the number of filters and the noise level. Next, halo elimination processing is performed by summing the m×5 images obtained after m repetitions of processing along the m direction. In multispectral scenes or infrared detection systems, the halo performance of different filters varies. For example, the halo edge of a bandpass filter is sharper, while the halo range of a long-pass filter is wider. By superimposing multiple sets of data and utilizing the statistical averaging effect, these band-specific halos can be naturally eliminated, while preserving image details and heat flow distribution information, ultimately obtaining a high-quality intermediate image without halo interference.
[0056] S6: Construct a Conditional Generative Adversarial Network (CGAN) with a U-shaped network as the generator, incorporate multispectral scene condition information, and train the CGAN using a dual loss function constraint. Input the intermediate image into the trained CGAN.
[0057] In some implementations, constructing a conditional generative adversarial network (GAN) with a U-shaped network (Unet) as the generator, incorporating multispectral scene conditional information, and training the GAN using a double loss function constraint, and inputting the intermediate image into the trained GAN, includes: constructing the GAN, wherein the generator adopts a U-shaped network structure with symmetrical encoding and decoding, and is equipped with multi-scale feature fusion skip connections; incorporating multispectral scene conditional information during the training of the GAN, the conditional information including wavelength parameters of eight filters, heat flux labels corresponding to infrared images, camera integration time, and pixel width; training the GAN using a double loss function constraint, and inputting the intermediate image into the trained GAN.
[0058] See attached document Figure 3 As shown, the network structure of the U-shaped network is illustrated. UNet, as the core generator for multispectral infrared image enhancement, possesses a typical encoder-decoder symmetrical U-shaped structure. Its core consists of an encoder (downsampling path) responsible for feature compression and a decoder (upsampling path) responsible for resolution restoration. The encoder progressively extracts and compresses features from the input multispectral infrared image through multiple convolution operations and max pooling operations. Each downsampling operation halves the feature map size and doubles the number of channels, thereby gradually extracting high-level heat flow distribution patterns and defects from low-level filter band textures and image edge details. The encoder extracts core semantic information such as structural features; the decoder gradually restores the spatial resolution of the image through upsampling operations such as transpose convolution, during which the number of channels is halved. The multi-scale feature fusion skip connection structure specially adopted in this study can accurately fuse the shallow detail features of the corresponding layer of the encoder with the current deep semantic features of the decoder at the key nodes of the decoder's resolution restoration. This not only makes up for the inevitable loss of details during the upsampling process, but also retains the band-specific features of the images acquired by different filters. This perfectly meets the dual requirements of multispectral images, which need to retain band individuality and restore the physical laws of thermal flow.
[0059] In the augmented network built based on the U-shaped network in this application embodiment, the skip connection of the UNet network belongs to a gated dynamic weighted feature fusion structure, and its structure diagram is as follows. Figure 4 As shown, Figure 4The diagram showcases the improved skip connection architecture, employing a gated dynamic weighted feature fusion mechanism. It generates weight vectors through convolution and sigmoid activation, precisely selecting encoder and decoder features to improve detail preservation and noise suppression in infrared image enhancement. Its core is the precise fusion of encoder and decoder features through adaptive weight allocation. This precise fusion specifically involves using features g from a certain layer of the encoder and features x from the current layer of the decoder. l As input, the channels and spatial dimensions of the two features are first aligned using 1×1×1 convolutions. Then, the aligned features are summed element-wise, and after introducing non-linearity through ReLU activation, another 1×1×1 convolution is used to extract the core information of the fused features. Finally, a weight vector in the range of 0 to 1 is generated using Sigmoid activation. Each element of this vector corresponds to the contribution of a feature at a specific "spatial-channel-depth" position. Finally, this weight is combined with the decoder feature x. l Element-wise multiplication enables "on-demand transfer"—weights close to 1 retain effective details from the encoder (such as defect textures in infrared images), while weights close to 0 suppress redundant interference (such as background noise), achieving dynamic feature fusion. In contrast, the original UNet's skip connection is an indiscriminate channel stitching mechanism, directly stacking the encoder's corresponding layer features and the decoder's upsampled features along the channel dimension. This lacks feature selection and non-linear interaction, easily introducing redundant information and doubling the number of channels, leading to a dramatic increase in computation. This gated structure, however, is an actively weighted non-linear fusion, capturing the correlation between encoder and decoder features (such as the correspondence between heat flow features and defect textures) through intermediate non-linear processing, while maintaining feature dimension conservation. It is more suitable for tasks like infrared image enhancement that require distinguishing between effective details and interference (such as halo artifacts), selectively preserving heat flow details and suppressing redundant edges, improving fusion efficiency while enhancing the accuracy of feature representation.
[0060] See attached document Figure 5 As shown, the structure of the CGAN network is presented, consisting of a UNet generator and a discriminator. It incorporates multispectral conditional information (such as filter wavelengths and camera parameters) and achieves accurate temperature data reconstruction from heat flow images through collaborative optimization of GAN loss and L1 loss functions. The Conditional Generative Adversarial Network (CGAN) is composed of this optimized UNet generator and a discriminator specifically adapted for multispectral scenes. Both deeply integrate conditional information unique to multispectral scenes, including wavelength parameters of eight filters, heat flow labels corresponding to infrared images, camera integration time, and pixel width, among other hardware acquisition parameters. The introduction of this conditional information frees the image generation process from the randomness of traditional GANs, giving it greater controllability and specificity, and enabling precise adaptation to the output requirements of multispectral rotary infrared data acquisition systems.
[0061] For example, the dual loss function includes defining the GAN loss (Generative Adversarial Network loss) and defining... loss.
[0062] For example, the GAN loss is defined as being calculated using the following expression:
[0063] The definition The loss is calculated using the following expression:
[0064] Among them, the This represents calculating the mathematical expectation over a large number of samples; the... This indicates the conditional information of the input generator; the Represents a real image; the This indicates that the generator is based on conditional information. The generated image; the This indicates that the discriminator recognizes the real image. Corresponding condition information The discrimination probability; This indicates that the discriminator evaluates the generated image. Corresponding condition information The discrimination probability; Represents generator of Loss value.
[0065] The GAN loss function originates from the continuous adversarial training game between the generator and the discriminator. This involves the generator taking random noise and the aforementioned multispectral conditional information as input, and generating realistic enhanced images through a series of sequential operations such as convolution, normalization, and activation. Its core objective is to continuously improve the realism of the generated images to "deceive" the discriminator. The discriminator, on the other hand, takes real multispectral infrared images and the fake images generated by the generator as core inputs, and combines them with corresponding conditional information. After extracting features through multiple convolutions, it outputs a probability value in the range of 0 to 1, used to determine whether the input image is a real sample or a generated fake sample. Its core objective is to continuously improve its ability to distinguish between real and fake samples. During iterative training, the generator and discriminator mutually promote and continuously optimize each other, ultimately enabling the generator to stably generate images that visually closely resemble the realism of multispectral scenes.
[0066] in, Loss or ( The loss function focuses on constraining the pixel-level differences between the generated image and the original real image, ensuring that the generated image is highly consistent with the original scene in core physical information such as heat flux intensity and temperature values. This avoids deviating from actual physical laws due to an excessive pursuit of visual realism, which is crucial for subsequent professional applications of infrared images such as defect detection and heat flux assessment, achieving a dual guarantee of visual effect and physical accuracy. To further improve the network's generalization ability in multispectral scenes, the training set specifically covers full-scene infrared image data collected by eight different filters. The high-quality intermediate image after dehaling processing in this step is input into the trained CGAN network, and the optimized UNet generator outputs the final enhanced heat flux image or temperature image.
[0067] S7: The target enhanced thermal flow image is output by generating an adversarial network based on the conditions, and the original single-frame infrared thermal image and the corresponding visible light image are compared, displayed and stored with the target enhanced thermal flow image. The visible light image is used to assist in the analysis of scene texture and environmental factors.
[0068] For example, the target enhanced heat flux image can also be a target temperature image. The stored processing parameters include virtual diffusion coefficient, white noise standard deviation, GAN loss value, etc. The loss values and corresponding filter parameters are compared and displayed, including the original single-frame infrared thermal images acquired with different filters and the enhanced and deblurred single-frame infrared thermal images.
[0069] See attached document Figure 6 The illustrated embodiment of this application, based on the above S1-S7 steps, demonstrates the overall flowchart of the image enhancement algorithm. It summarizes the complete process from multispectral data acquisition, virtual diffusion coefficient calculation, noisy sequence generation, frequency domain transformation to CGAN restoration, showcasing an integrated processing chain combining physical modeling and deep learning. This not only completely eliminates noise, halos, and other interference factors in the original image, accurately restoring key details such as minute defects and edge contours, but also ensures, through the strict constraints of the L1 loss function, that core physical data such as heat flow distribution and temperature values are highly consistent with the original scene, without any distortion issues deviating from reality. In multispectral scenarios, cross-band fusion processing can be performed on the enhancement results corresponding to different filters, integrating the advantages of each band, such as utilizing the detail advantages of short-pass filter images and the thermal flow stability advantages of long-pass filter images, to form multi-dimensional, comprehensive enhanced images. This provides more comprehensive and reliable data support for subsequent applications such as industrial component defect identification and power equipment thermal flow assessment, significantly improving the accuracy and efficiency of subsequent detection tasks.
[0070] This application embodiment ensures the temporal consistency of data acquisition by simultaneously acquiring original single-frame infrared thermal images and visible light images, providing complete input for multispectral analysis. The virtual diffusion coefficient is calculated based on integration time, pixel width, and the thermal radiation characteristics of the filter, making the heat conduction model conform to actual hardware and optical conditions, thus improving the accuracy of subsequent physical modeling. This application embodiment generates Gaussian diffusion kernels and sequential convolution kernels to form a convolution kernel sequence adapted to multispectral bands, capable of simulating the thermal diffusion characteristics of different bands. Time-series expansion compensates for the lack of dynamic temporal information in single-frame images, enhancing the physical rationality of heat flow inversion. This application embodiment generates noisy image sequences by adding white noise, improving the model's robustness to real noise; convolution operations smooth the image sequences, effectively suppressing random noise while preserving core heat flow features, providing a high-quality data foundation for frequency domain processing. This application embodiment highlights image differences and avoids boundary effects through residual calculation and augmentation processing; three-dimensional Fourier transform combined with dot product kernels corrects frequency domain data, accurately capturing spatial-temporal-band correlations, optimizing the separation and enhancement of heat flow features. This application embodiment eliminates halo phenomena by inversely transforming back to the spatiotemporal domain and filtering effective data, and by performing multiple rounds of summation operations. It utilizes statistical averaging to cancel band-specific artifacts, obtaining high-quality intermediate images and improving the clarity and positioning accuracy of defect edges. This application embodiment constructs a conditional generative adversarial network with a U-shaped network as the generator, incorporating multispectral scene condition information to ensure the enhancement process conforms to thermophysical laws. Dual loss function constraints are used in training to balance visual realism and pixel-level accuracy, outputting enhanced images with rich detail and physical realism. This application embodiment outputs enhanced thermal flow images of the target and compares them with the original infrared and visible light images for intuitive evaluation of the enhancement effect. The visible light image assists in analyzing scene texture and environmental factors, improving the reliability and comprehensiveness of the detection results and supporting subsequent defect identification decisions.
[0071] This application also discloses a single-frame infrared thermal image enhancement system, including: an original image acquisition module, a convolution kernel sequence generation module, a smooth image sequence generation module, a smooth image sequence transformation module, an intermediate image generation module, a conditional generative adversarial network processing module, and a target-enhanced thermal flow image generation module.
[0072] For example, the original image acquisition module is used to simultaneously acquire the original single-frame infrared thermal image and the corresponding visible light image, and calculate the virtual diffusion coefficient based on the integration time of the original single-frame infrared thermal image, the camera pixel width, and the thermal radiation characteristics of the selected multispectral filter.
[0073] For example, the convolution kernel sequence generation module is used to generate Gaussian diffusion kernels and sequence convolution kernels based on the virtual diffusion coefficient, and form a convolution kernel sequence adapted to the thermal diffusion characteristics of multispectral bands.
[0074] For example, a smooth image sequence generation module is used to add white noise to the original single-frame infrared thermal image to generate a noisy image sequence, and to perform a convolution operation between the noisy image sequence and the convolution kernel sequence to obtain a smooth image sequence.
[0075] For example, the smoothed image sequence transformation module is used to perform residual calculation, augmentation processing, and three-dimensional Fourier transform on the smoothed image sequence, and to correct the frequency domain data using a dot product kernel.
[0076] For example, the intermediate image generation module is used to inversely transform the corrected frequency domain data back to the spatiotemporal domain, filter valid image data and repeat the process multiple times, and eliminate the image halo phenomenon by adding multiple sets of data to obtain an intermediate image that meets the target quality.
[0077] For example, the conditional generative adversarial network processing module is used to construct a conditional generative adversarial network with a U-shaped network as the generator, incorporate multispectral scene condition information, and use dual loss function constraints to train the conditional generative adversarial network, and input the intermediate image into the conditional generative adversarial network after network training.
[0078] For example, the target enhanced thermal flow image generation module is used to output a target enhanced thermal flow image through the conditional generative adversarial network, and to compare, display and store the original single-frame infrared thermal image and the corresponding visible light image with the target enhanced thermal flow image, wherein the visible light image is used to assist in the analysis of scene texture and environmental factors.
[0079] In some implementations, the original single-frame infrared thermal image and the corresponding visible light image are acquired simultaneously, based on an image acquisition system. This image acquisition system (which can also be understood as a multispectral infrared thermal imaging data acquisition system) uses a microcontroller as its core control unit and is equipped with an infrared camera, a motor-driven filter wheel, and a visible light camera. It controls the rotation of the filter wheel through a preset program, ensures that the rotation of the filter wheel is synchronized with the timing of camera acquisition, and acquires image sequences containing slide-less data and multispectral data. For example, the image acquisition system includes an infrared camera (e.g., MAG62), a development board (e.g., Arduino), a motor-driven filter wheel, and eight different types of filters.
[0080] For example, the image acquisition system uses an Arduino UNO development board as the core control unit, equipped with a MAG6 infrared camera, a motor-driven filter wheel (containing 8 different types of filters), and a visible light camera. The direction pin, pulse pin, camera trigger pin, and status indicator pin of the development board are respectively connected to the motor driver direction control terminal, pulse signal input terminal, camera external trigger interface, and system status indicator. The system uses a falling edge trigger mode to control camera acquisition. Through program design, the timing of filter wheel rotation and camera acquisition is synchronized to ensure that 10 images, including 2 images without a glass slide and 8 images of multispectral data, are accurately acquired with one rotation of the turntable.
[0081] For example, when controlling the multispectral infrared thermal imaging data acquisition system to perform multi-time point acquisition, the acquisition time sequence covers early morning, noon, and evening, and the acquisition months cover February-March and April-June. Among them, the acquisition is divided into segments according to specific time intervals from 6:00 to 21:45, and each segment is continuously photographed 3 times to study the impact of sunlight excitation on infrared defect detection under different seasons and times.
[0082] See attached document Figure 7The diagram illustrates the hardware structure flowchart of a multispectral infrared thermal imaging data acquisition system. The core control unit of this system utilizes the Arduino UNO development board, which is equipped with an ATmega328P microcontroller. This board features 14 digital input / output pins and 6 analog input pins, supporting both USB power supply and external power supply modes. Its programming architecture employs the standard C / C++ language framework, relying on the Arduino integrated development environment for program design. With its mature and stable hardware architecture, low development cost, and comprehensive open-source ecosystem, the Arduino UNO can efficiently complete core functions such as data acquisition, complex logic operations, and communication with external devices, providing reliable control assurance for system operation. In the hardware system, the functions of each pin include: Direction pin (dirPin, pin 7), connected to the direction control terminal (DIR) of the stepper motor driver, outputting high and low levels to control the forward and reverse rotation of the motor; Pulse pin (pulsePin, pin 8), connected to the pulse signal input terminal (PUL) of the stepper motor driver, precisely controlling the number of steps and angle of motor rotation through the output pulse sequence; Camera trigger pin (triggerPin, pin 9), serving as an external trigger interface for the camera, controlling camera image acquisition through the falling edge of the output periodic pulse voltage; and Status indicator pin (ledPin, pin 13), connected to the system status indicator. The system features an LED display indicating its operating status; a DC12V interface for powering the camera via a dedicated power cord; a LAN interface for connecting to a computer or network device via a network cable, enabling infrared image data transmission and control command interaction between the camera and the computer; an IN interface connected to pin 9 of the Arduino development board to receive the periodic pulse voltage output from pin 9. When the filter wheel rotates to be horizontal with the camera lens, the falling edge signal triggers the camera's external control acquisition, ensuring synchronization between camera acquisition and filter wheel rotation; and a GND interface connected to the ground terminal of the Arduino development board to unify the reference potential, stabilize signal transmission, and ensure stable system operation.
[0083] The connection of the infrared camera is crucial for its normal operation and coordinated operation with the system. The camera's external trigger interface is responsible for receiving trigger signals sent by the external control system to achieve precise synchronous data acquisition. This system adopts a falling edge trigger mode, meaning the camera initiates image acquisition the instant the trigger signal switches from high to low. To ensure reliable triggering, the timing parameters of the trigger signal must maintain a stable high-low level for at least 5ms.
[0084] In some implementations, this single-frame infrared thermal image enhancement system also includes an image acquisition module. Its core function is to simultaneously acquire a single-frame infrared thermal image (combined with eight filters on a multispectral rotary table, covering bandpass, long-pass, and short-pass types; specific parameters are shown in Table 1 below) and a visible light image, providing raw data for subsequent processing. This module uses a MAG62 infrared camera as its core, coupled with an adjustable integration time hardware configuration. It can set an appropriate integration time τ according to the requirements of the detection scenario, while accurately recording hardware parameters such as the camera pixel width w. Furthermore, this module achieves multispectral infrared image acquisition by switching the filter wheel via an Arduino development board and a motor. It also has data preprocessing capabilities, standardizing the format of the acquired raw images to ensure the integrity and consistency of the image data, providing reliable input for subsequent steps such as virtual diffusion coefficient calculation and noise addition.
[0085] Table 1: Filter Parameter Table
[0086] Referring to Table 1, filter number 6 represents a blank filter.
[0087] In some implementations, this single-frame infrared thermal image enhancement system further includes a virtual heat flow generation module, which is responsible for generating a virtual heat flow image based on the single-frame infrared thermal image and is a key link connecting physical modeling and image enhancement. This module has a built-in heat conduction control equation solver, which can automatically calculate the virtual diffusion coefficient α based on τ, w, and filter parameters (such as the thermal radiation characteristics corresponding to the average wavelength) input from the image acquisition module, and generate the corresponding Gaussian diffusion core sequence. Simultaneously, the module has noise addition and convolution operation functions, capable of generating infrared image sequences containing white noise according to a set number of iterations, and performing convolution operations between the image and the Gaussian sequence to output a smooth image sequence, laying the foundation for subsequent frequency domain transformation.
[0088] In some embodiments, this single-frame infrared thermal image enhancement system further includes a frequency domain processing module, which focuses on performing spatiotemporal frequency domain transformation, Bessel core inversion, and inverse transformation operations to achieve halo elimination and thermal flow feature enhancement. This module integrates a three-dimensional Fourier transform algorithm and a Bessel core inversion algorithm, enabling precise transformation of the residual-expanded image. It corrects the frequency domain data using a specific dot product core and then inversely transforms the processed data back to the spatiotemporal domain. Furthermore, this module has data filtering and overlay functions, retaining valid image data and performing multiple rounds of data summation operations to efficiently eliminate halo phenomena in the image, outputting high-quality intermediate image data that adapts to the image characteristics of different multispectral bands.
[0089] In some implementations, this single-frame infrared thermal image enhancement system further includes a CGAN restoration module, which employs a CGAN network structure composed of a UNet generator and a discriminator to accurately restore temperature data from the image. This module incorporates an optimized UNet generator whose skip connection structure effectively preserves image detail features. Simultaneously, through the synergistic effect of the GAN loss function and the L1 loss function, it ensures that the generated temperature data is both visually realistic and closely approximates the original data. This module can iteratively optimize the frequency-domain processed multispectral image data, outputting the final enhanced heat flow image or temperature image, thus completing the entire image enhancement process.
[0090] In some implementations, this single-frame infrared thermal image enhancement system also includes a result output module, which performs the functions of result display and data storage. It can compare and display the original thermal image (acquired using different filters) with the enhanced and deblurred thermal image, allowing users to intuitively observe the enhancement effect. Simultaneously, it supports storing the enhanced heat flow image, temperature image, and various processing parameters (such as diffusion coefficient, noise standard deviation, loss function value, filter parameters, etc.) for subsequent traceability and analysis.
[0091] This single-frame infrared thermal image enhancement system is based on a combination of pseudospectral physical modeling, spatiotemporal frequency domain optimization, and CGAN deep learning. It constructs a physical model of heat conduction using pseudospectral methods and, combined with the thermal radiation characteristics of multispectral filters, clarifies the heat flow transfer patterns in the spatiotemporal domain of a single-frame image, ensuring the authenticity of the heat flow distribution from a physical perspective. Utilizing the adaptive learning capabilities of CGAN, it accurately repairs image noise and detail blurring in different multispectral bands, achieving detail enhancement with limited single-frame data. Through frequency domain transformation and superposition operations, it eliminates defocusing and halo interference caused by filter switching. Simultaneously, it integrates multi-module processing flows, forming an automated system from "multispectral single-frame data acquisition → physical model heat flow expansion → frequency domain de-haloing and noise reduction → CGAN detail calibration → enhanced image output," achieving high-quality enhancement of single-frame infrared images without manual intervention. This provides a reliable data foundation for defect identification and heat flow assessment in multispectral infrared detection, significantly improving detection efficiency and accuracy.
[0092] This application embodiment utilizes a self-built multispectral infrared thermal imaging acquisition system to perform full-cycle, multi-time-point data acquisition. The core components of this system include a MAG62 infrared camera, an Arduino development board, and a motor-driven filter wheel. Its hardware connection and synchronization control process are as follows: Figure 1 and Figure 2As shown, the device completes one round of acquisition by rotating once under the drive of the motor, simultaneously acquiring the original infrared thermal image without a glass slide (as a reference) and multi-band infrared thermal images (covering the spectral ranges of bandpass, long-pass, and short-pass). At the same time, the visible light camera acquires scene texture and weather information to help distinguish non-defective structures (such as tile joints) and record environmental variables. By comparing the visible light and infrared data, as shown in Figures 8(a) and 8(b), the comparison results of the visible light image in Figure 8(a) and the infrared thermal imaging data in Figure 8(b) are shown. From the comparison, it can be clearly seen that the visible light image mainly records the surface texture features of the scene (such as tile joints, decorative lines, etc.) and real-time weather conditions, while the infrared thermal imaging data is based on the thermal radiation difference of the target, which can effectively penetrate surface visual interference and accurately capture the abnormal heat distribution caused by internal defects in the wall (such as hollows and cracks). This comparison vividly highlights the core advantage of infrared thermal imaging technology in internal defect detection: its imaging results are not affected by factors such as ambient light intensity and surface texture blurring, highlighting the advantages of infrared thermal imaging in penetrating surface interference and capturing internal defects (such as hollow areas) and abnormal heat distribution.
[0093] This application employs a refined multi-time-point data acquisition sequence: starting daily at 6:00 AM, data is collected in segments at 15-minute intervals (covering key nodes such as 6:00, 9:00, 12:00, 15:00, 18:00, and 21:00), with each segment consisting of three consecutive images, concluding at 21:45. This design leverages the slow temperature change of walls, and the 15-minute intervals balance data accuracy and redundancy, fully covering the entire cycle of sunlight excitation from weak to strong, providing temporal support for studying the impact of sunlight on defect detection at different times. Furthermore, the data acquisition covers February to March (high humidity environment in spring) and April to June (strong sunlight environment in spring and summer). By comparing seasonal differences, the impact of humidity and sunlight intensity on heat transfer differences is analyzed, enriching the data diversity.
[0094] In actual data acquisition, image defocusing is mainly caused by camera resolution limitations and optical focal length shifts due to filter switching. This application addresses this issue through a physical modeling enhancement scheme: First, a dynamic thermal imaging sequence is generated based on a single-frame infrared thermal image and the Green's function of the heat conduction equation to compensate for missing temporal information. Second, the sequence undergoes spatiotemporal frequency domain transformation and Bessel core inversion to separate and correct defocusing and noise interference (Bessel core adapts to the radial diffusion characteristics of heat flow). Finally, a virtual heat flow image is generated through inverse transformation, and Planck's law is used to convert thermal radiation intensity into temperature data, outputting an enhanced heat flow or temperature image. This method effectively eliminates halo artifacts in the original image. Referring to Figures 9(a) and 9(b), a comparison is shown between the original infrared grayscale image in Figure 9(a) and the image enhanced by the method described in this application in Figure 9(b). It can be clearly seen from Figures 9(a) and 9(b) that the enhanced image has significantly improved detail information, and key features such as the thermal flow difference between defective and non-defective areas and edge contours are accurately restored. At the same time, this method effectively solves the defocusing problem of the original image and eliminates halo artifacts (i.e., false bright or dark edges of the image) that are easily caused by the RPHF algorithm, providing a more reliable visual basis for defect identification and quantitative analysis, thereby achieving the unity of detail enhancement and physical authenticity, and providing a reliable data foundation for defect identification.
[0095] This application constructs a core technology system from three dimensions: physical modeling, frequency domain optimization, and deep learning-based restoration. Through the synergistic effect of key steps such as virtual heat flow generation, spatiotemporal frequency domain transformation, and GAN network calibration, a complete process covering data input, processing, and output is formed. This system balances physical accuracy and image enhancement effects in multispectral scenarios, adapting to the actual needs of infrared detection scenarios involving multispectral filter switching. This application constructs a heat conduction physical model using a pseudospectral method, enabling the expansion and frequency domain optimization from single-frame images to continuous thermal imaging sequences. This ensures the physical rationality of heat flow data in multispectral scenarios. Utilizing the generator and discriminator mechanisms of the CGAN network, temperature data in the image is accurately restored, improving the realism and detail of the enhanced image. Multi-step noise suppression and halo elimination further optimize image quality. Simultaneously, this application implements a system deeply compatible with the multispectral rotary infrared data acquisition system. It eliminates the need for multi-frame image input, outputting enhanced heat flow or temperature images from a single frame of infrared thermal image, significantly improving the efficiency and accuracy of multispectral infrared detection and providing high-quality data support for subsequent defect identification and heat flow assessment. This method organically combines physical modeling and deep learning techniques to address issues such as noise interference, detail blurring, and halo phenomena in single-frame infrared images. It achieves precise enhancement of heat flow and temperature images and is applicable to infrared thermal imaging detection systems and multispectral infrared thermal imaging detection systems. It has significant application value, especially in scenarios such as surface scratch detection of industrial parts, identification of internal material defects, and defocused image repair under channel spectral filtering. It provides technical support for the precision and efficiency of infrared detection.
[0096] It is understood that the above embodiments are merely exemplary implementations used to illustrate the principles of this application, and this application is not limited thereto. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and substance of this application, and these modifications and improvements are also considered to represent the scope of protection of this application.
Claims
1. A method for enhancing a single-frame infrared thermal image, characterized in that, include: S1: Simultaneously acquire the original single-frame infrared thermal image and the corresponding visible light image, and calculate the virtual diffusion coefficient based on the integration time of the original single-frame infrared thermal image, the camera pixel width, and the thermal radiation characteristics of the selected multispectral filter. S2: Based on the virtual diffusion coefficient, a Gaussian diffusion kernel and a sequence convolution kernel are generated, and a convolution kernel sequence adapted to the thermal diffusion characteristics of multispectral bands is formed. S3: Add white noise to the original single-frame infrared thermal image to generate a noisy image sequence, and perform a convolution operation between the noisy image sequence and the convolution kernel sequence to obtain a smooth image sequence; S4: Perform residual calculation, augmentation processing, and three-dimensional Fourier transform on the smoothed image sequence, and use dot product kernels to correct the frequency domain data; S5: The corrected frequency domain data is inversely transformed back to the spatiotemporal domain, the effective image data is filtered and the process is repeated multiple times. The halo effect of the image is eliminated by summing multiple sets of data to obtain an intermediate image that meets the target quality. S6: Construct a conditional generative adversarial network with a U-shaped network as the generator, incorporate multispectral scene condition information, and use dual loss function constraints to train the conditional generative adversarial network. Input the intermediate image into the conditional generative adversarial network after network training. S7: The target enhanced thermal flow image is output by generating an adversarial network based on the conditions, and the original single-frame infrared thermal image and the corresponding visible light image are compared, displayed and stored with the target enhanced thermal flow image. The visible light image is used to assist in the analysis of scene texture and environmental factors.
2. The method according to claim 1, characterized in that, The synchronous acquisition of the original single-frame infrared thermal image and the corresponding visible light image includes: The image acquisition system simultaneously acquires the original single-frame infrared thermal image and the corresponding visible light image. The image acquisition system uses a microcontroller as the core control unit and is equipped with an infrared camera, a motor-driven filter wheel, and a visible light camera. It is used to control the rotation of the filter wheel through a preset program, ensure that the rotation of the filter wheel is synchronized with the timing of the camera acquisition, and acquire image sequences containing slide-free data and multispectral data.
3. The method according to claim 1, characterized in that, The calculation of the virtual diffusion coefficient based on the integration time of the original single-frame infrared thermal image, the camera pixel width, and the thermal radiation characteristics of the selected multispectral filter includes: The initial data for the virtual diffusion coefficient are calculated based on the following formula: in, Represents the virtual diffusion coefficient; This indicates the width of the camera pixels. This represents the integration time of the original single-frame infrared thermal image; The initial data of the virtual diffusion coefficient are fine-tuned based on the thermal diffusion rate corresponding to the average wavelength of the selected multispectral filter.
4. The method according to claim 3, characterized in that, The Gaussian diffusion core is calculated and generated based on the following expression: The sequence convolution kernel is calculated and generated based on the following expression: in, It is an exponential function with the natural constant e as its base; Indicates radial distance; Represents the virtual diffusion coefficient; Indicates the fundamental time constant; Represents a continuously changing time variable. The value of is Discretize into units. The range of values is ,2 , ..., (N+1) N is an integer determined based on the differences in thermal response across multiple spectral bands.
5. The method according to claim 1, characterized in that, The method further includes: The standard deviation of the added white noise is dynamically determined based on the average residual between the original single-frame infrared thermal image and the first image in the smoothed image sequence. The standard deviation of the white noise to be added is adaptively adjusted according to the type of filter selected. When the filter type is a short-pass filter, the standard deviation of the white noise is greater than the standard deviation of the white noise when the filter type is a long-pass filter or a band-pass filter.
6. The method according to claim 4, characterized in that, The process of performing residual calculation, augmentation, and three-dimensional Fourier transform on the smoothed image sequence, and correcting the frequency domain data using a dot product kernel, includes: The smoothed image sequence is subjected to residual calculation to obtain the residual image sequence; The residual image sequence is expanded by 20% in both the row and column directions of the image coordinates; A three-dimensional Fourier transform is performed on the expanded residual map sequence to convert the spatiotemporal domain signal into a frequency domain signal; The transformed frequency domain signal is corrected using a dot product kernel, the expression of which is shown below: in, This represents the length frequency corresponding to the Fourier transform. This represents the width frequency corresponding to the Fourier transform. The value is represented as The imaginary unit, The dot product core represents the time frequency corresponding to the Fourier transform. It is used to capture the spatial, temporal, and band correlation features in the multispectral frequency domain and to specifically correct the differences in frequency domain energy distribution of different filters.
7. The method according to claim 1, characterized in that, The process of inversely transforming the corrected frequency domain data back to the spatiotemporal domain, filtering valid image data and repeating the process multiple times, and eliminating image halo phenomena through multiple sets of data summation operations to obtain an intermediate image that meets the target quality includes: The corrected frequency domain data is restored to the spatiotemporal domain through inverse three-dimensional Fourier transform, and the transformation result containing the data in the spatiotemporal domain is obtained. Filter and retain the first 5 valid image data from the transformation results; Repeat steps S4 to S6 a total of m times, where the specific value of m is determined by the number of filters and the image noise level. The m×5 images obtained from m repeated processing are summed along the direction of the number of processing steps m, and the band-specific halos of different filters are canceled out by statistical averaging effect to obtain an intermediate image that meets the target quality.
8. The method according to claim 1, characterized in that, The construction of a conditional generative adversarial network (GAN) using a U-shaped network as a generator, incorporating multispectral scene conditional information and employing dual loss function constraints for network training, and inputting the intermediate image into the trained GAN includes: The conditional generative adversarial network is constructed, wherein the generator adopts a U-shaped network structure with symmetrical encoding and decoding, and is equipped with multi-scale feature fusion skip connections; When training the conditional generative adversarial network, multispectral scene condition information is incorporated, including wavelength parameters of eight filters, heat flux labels corresponding to infrared images, camera integration time, and pixel width. The conditional generative adversarial network is trained using a dual loss function constraint, and the intermediate image is input into the conditional generative adversarial network after the network training is completed. The dual loss function includes defining the GAN loss and defining... The loss, defined as the GAN loss, is calculated using the following expression: The definition The loss is calculated using the following expression: Among them, the This represents calculating the mathematical expectation over a large number of samples; the... This indicates the conditional information of the input generator; the Represents a real image; the This indicates that the generator is based on conditional information. The generated image; the This indicates that the discriminator recognizes the real image. Corresponding condition information The discrimination probability; This indicates that the discriminator evaluates the generated image. Corresponding condition information The discrimination probability; Represents generator of Loss value.
9. The method according to claim 1, characterized in that, The stored processing parameters include virtual diffusion coefficient, white noise standard deviation, GAN loss value, The loss values and corresponding filter parameters are compared and displayed, including the original single-frame infrared thermal images acquired with different filters and the enhanced and deblurred single-frame infrared thermal images.
10. A single-frame infrared thermal image enhancement system, characterized in that, include: The original image acquisition module is used to simultaneously acquire the original single-frame infrared thermal image and the corresponding visible light image, and calculate the virtual diffusion coefficient based on the integration time of the original single-frame infrared thermal image, the camera pixel width, and the thermal radiation characteristics of the selected multispectral filter. The convolution kernel sequence generation module is used to generate Gaussian diffusion kernels and sequence convolution kernels based on the virtual diffusion coefficient, and form a convolution kernel sequence adapted to the thermal diffusion characteristics of multispectral bands. The smooth image sequence generation module is used to add white noise to the original single-frame infrared thermal image to generate a noisy image sequence, and to perform a convolution operation between the noisy image sequence and the convolution kernel sequence to obtain a smooth image sequence. The smoothed image sequence transformation module is used to perform residual calculation, augmentation processing, and three-dimensional Fourier transform on the smoothed image sequence, and to correct the frequency domain data using the dot product kernel; The intermediate image generation module is used to inversely transform the corrected frequency domain data back to the spatiotemporal domain, filter effective image data and repeat the process multiple times, and eliminate the image halo phenomenon through the summation operation of multiple sets of data to obtain an intermediate image that meets the target quality. The Conditional Generative Adversarial Network (CGAN) processing module is used to construct a CGAN with a U-shaped network as the generator, incorporate multispectral scene condition information, and use dual loss function constraints to train the CGAN. The intermediate image is input into the CGAN trained by the network. The target-enhanced thermal flux image generation module is used to output a target-enhanced thermal flux image through the conditional generative adversarial network, and to compare, display and store the original single-frame infrared thermal image and the corresponding visible light image with the target-enhanced thermal flux image. The visible light image is used to assist in the analysis of scene texture and environmental factors.