Infrared thermal imaging denoising and imaging optimization method based on deep learning
Patent Information
- Application Number
- CN202610677237.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-05-18
AI Technical Summary
[0004]为了解决上述技术问题,本申请提供基于深度学习的红外热成像去噪及成像优化方法,以解决现有的问题
本申请针对固定模式噪声与条纹噪声在空间、频域及方向域上的特异性解耦分析,构建了第一噪声系数与第二噪声系数,成功将复杂耦合噪声的物理表象转化为可量化的数学特征,该静态特征提取机制能够使模型准确识别当前图像中主导噪声的类型与强度,为后续在损失函数中实施精准的差异化约束奠定了数据基础,有效避免了传统方法将混合噪声统一均化处理而导致的去噪不足或边缘过度平滑问题;
Smart Images

Figure CN122199318B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of infrared thermal imaging denoising technology, specifically to infrared thermal imaging denoising and imaging optimization methods based on deep learning. Background Technology
[0002] Infrared thermal imaging technology, with its advantages of being non-contact and all-weather, has been widely used in industrial inspection, medical, and security fields. However, due to the limitations of the detector hardware's physical characteristics, the acquired raw images are often accompanied by severe noise interference, directly reducing the image signal-to-noise ratio and the accuracy of subsequent tasks such as target identification and temperature measurement. Therefore, how to effectively suppress noise and optimize imaging quality has always been a core requirement in this field.
[0003] In real-world scenarios, infrared images are often subject to complex coupling interference from fixed-pattern noise, stripe noise, and random noise. Traditional methods are mostly static processing flows targeting single types of noise, making it difficult to handle the superposition of multiple noises. While existing deep learning methods attempt to handle mixed noise, their loss functions typically impose equal penalty weights on all types of noise, essentially remaining limited by "uniform processing." This lack of differentiated constraints in the optimization mechanism causes the network to become directionally dispersed when facing coupled noise, easily leading to incomplete denoising or excessive smoothing of details, making it difficult to meet the dual requirements of high-precision imaging for purity and detail preservation. Summary of the Invention
[0004] To address the aforementioned technical problems, this application provides a deep learning-based infrared thermal imaging denoising and imaging optimization method to solve the existing issues.
[0005] The infrared thermal imaging denoising and imaging optimization method based on deep learning in this application adopts the following technical solution: One embodiment of this application provides a method for infrared thermal imaging denoising and imaging optimization based on deep learning, the method comprising the following steps: Real-time acquisition of infrared images of the target object; Each frame of infrared image is divided into multiple image blocks. The average pixel value of all pixels in each image block is calculated. The average pixel value of all image blocks in each preset direction is fitted. The difference between the average pixel value of each image block and the corresponding fitted value in each direction is analyzed. The distance of each image block to the nearest edge in the infrared image and the degree of dispersion of the pixel value of all pixels in each image block are also analyzed. The first noise coefficient of each frame of infrared image is determined and used to quantify the significance of fixed pattern noise. The energy spectrum of each frame of infrared image in polar coordinates is obtained. By analyzing the energy distribution in each polar angle direction and the changing trend of the peak value in the average pixel value of all image blocks in each direction, the second noise coefficient of each frame of infrared image is determined, which is used to quantify the intensity of stripe noise. The first and second noise coefficients are smoothed, and the differences in the smoothed values of the first and second noise coefficients between adjacent frames of infrared images are analyzed. The first noise drift index and the second noise drift index are constructed respectively, and the dynamic noise index is determined by combining the differences between the first and second noise coefficients and their corresponding smoothed values. Based on the first and second noise figures and the dynamic noise index, noise weights are determined, and the loss function is optimized to train the infrared thermal imaging denoising model.
[0006] Preferably, the method for determining the first noise figure of each frame of infrared image is as follows: The mean difference between the average pixel value of each image patch and its corresponding fitted value in all directions is calculated and denoted as the mean pixel difference. The first noise coefficient of each frame of infrared image is positively correlated with the dispersion of pixel values of all pixels in each image block, and negatively correlated with the mean pixel difference and the distance of each image block to the nearest edge in the infrared image.
[0007] Preferably, the method for determining the second noise figure of each frame of infrared image is as follows: Calculate the total energy in each polar angle direction of the energy spectrum of each frame of infrared image in polar coordinate system, and take the result of dividing the maximum value of the total energy in all polar angle directions by the average value of the total energy in all polar angle directions as the energy feature value of each frame of infrared image. A fitted straight line is obtained by fitting the peak value of the average pixel value of all image blocks in each direction. The absolute value of the slope of the fitted straight line is recorded as the pixel change trend value in each direction. The mean of the pixel change trend values in all directions is recorded as the pixel change feature value of each frame of infrared image. Based on the energy characteristic value and the pixel change characteristic value, the second noise coefficient of each frame of infrared image is determined.
[0008] Preferably, the expression for the second noise coefficient of each frame of infrared image is: In the formula, This represents the second noise coefficient of the i-th frame of the infrared image; This represents the energy feature value of the i-th frame of the infrared image; Let represent the pixel change feature value of the i-th frame of the infrared image; min() represents the minimum value function; This indicates a constant that is pre-defined as being greater than 0.
[0009] Preferably, the construction of the first noise drift index and the second noise drift index respectively includes: Calculate the mean of the difference between the first noise coefficient smoothing value and the mean of the difference between the second noise coefficient smoothing value among all adjacent infrared images in a preset number of consecutive infrared images, and denot them as the first noise drift index and the second noise drift index, respectively.
[0010] Preferably, the expression for the dynamic noise index is: In the formula, Represents the dynamic noise index of the i-th frame of the infrared image; , Let represent the first noise drift index and the second noise drift index of the i-th frame infrared image, respectively; , represents the degree of dispersion of the difference between the first noise coefficient and its smoothed value in a preset number of infrared images, and the degree of dispersion of the difference between the second noise coefficient and its smoothed value, respectively; min() represents the minimum value function.
[0011] Preferably, the method for determining the noise weight is as follows: The noise weight can take three values: The first type of noise weight is the proportion of the first noise figure to the sum of the first noise figure, the second noise figure, and the dynamic noise index. The second type of noise weight is the proportion of the second noise figure in the sum of the first noise figure, the second noise figure, and the dynamic noise index. The third type of noise weight is the proportion of the dynamic noise index in the sum of the first noise figure, the second noise figure, and the dynamic noise index.
[0012] Preferably, the optimization of the loss function includes: The optimized loss function value corresponding to the s-th frame of the infrared image The expression is: ; This represents the denoised image obtained by processing the s-th frame of the infrared image in the preset training set of the infrared thermal imaging denoising model using the U-Net model. This represents the reference image of the s-th frame infrared image extracted by the multi-frame adaptive averaging fusion method based on temporal registration; This represents the noise weight of the s-th frame of the infrared image; This represents the total number of noise weight values for the s-th frame of the infrared image; The residual value is obtained by considering the noise weight based on the m-th value and the residual map between the denoised infrared image of the s-th frame and its corresponding reference image; || represents the L2 norm.
[0013] Preferably, the residual value obtained based on the noise weight of the m-th value and the residual map between the denoised infrared image of the s-th frame and its corresponding reference image includes: Under the first noise weight, the residual map between the denoised infrared image of the s-th frame and its corresponding reference image is used as the input of the low-pass filtering algorithm, and the filtered residual map is output. The average of the squares of all pixel values in the filtered residual map is used as the residual value of the denoised infrared image of the s-th frame. Under the second noise weight, the sum of the gradient magnitudes of all pixels in the residual image between the denoised infrared image of frame s and its corresponding reference image is calculated, and the sum of the gradient magnitudes in the horizontal direction and the vertical direction are recorded as the horizontal gradient sum and the vertical gradient sum, respectively. The mean of the horizontal gradient sum and the vertical gradient sum is used as the residual value of the denoised infrared image of frame s. Under the third noise weight value, the variance of all pixel values in the residual map between the denoised infrared image of frame s and its corresponding reference image is calculated and used as the residual value of the denoised infrared image of frame s.
[0014] Preferably, the trained infrared thermal imaging denoising model includes: All infrared images in the preset training set are used as input to the neural network, the corresponding reference images are used as target labels, the optimized loss function is used as the loss function in the neural network, the neural network is trained, and the trained neural network model is output, which is denoised as the infrared thermal imaging denoising model.
[0015] This application has at least the following beneficial effects: This application focuses on the specific decoupling analysis of fixed pattern noise and stripe noise in the spatial, frequency, and directional domains. It constructs a first noise figure and a second noise figure, successfully transforming the physical appearance of complex coupled noise into quantifiable mathematical features. This static feature extraction mechanism enables the model to accurately identify the type and intensity of the dominant noise in the current image, laying a data foundation for implementing precise differential constraints in the loss function. It effectively avoids the problems of insufficient denoising or excessive edge smoothing caused by the uniform homogenization of mixed noise in traditional methods. Furthermore, this application introduces inter-frame temporal evolution features and constructs a dynamic noise index. By using sliding trend separation and residual variance measurement, it separates the disordered jumps of random noise from the slow drift of structured noise in the temporal dimension, making up for the limitations of single-frame static feature extraction. This enables comprehensive and accurate quantification of fixed pattern noise, stripe noise and random noise in images in both spatial and temporal dimensions. Finally, based on the first and second noise figures and the dynamic noise index, this application determines the noise weights and optimizes the loss function to train the infrared thermal imaging denoising model. Attached Figure Description
[0016] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart illustrating the steps of a deep learning-based infrared thermal imaging denoising and imaging optimization method provided in one embodiment of this application; Figure 2 This is a schematic diagram of the first and second noise extraction processes provided in one embodiment of this application. Detailed Implementation
[0018] To further illustrate the technical means and effects adopted by this application to achieve the intended inventive objective, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of the deep learning-based infrared thermal imaging denoising and imaging optimization method proposed in this application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0020] The following description, in conjunction with the accompanying drawings, details the specific scheme of the infrared thermal imaging denoising and imaging optimization method based on deep learning provided in this application.
[0021] This application provides an embodiment of a deep learning-based infrared thermal imaging denoising and imaging optimization method. Specifically, the following deep learning-based infrared thermal imaging denoising and imaging optimization method is provided. Please refer to [link to relevant documentation]. Figure 1 The method includes the following steps: Step S1: Acquire infrared images of the target object in real time.
[0022] In this embodiment, the infrared thermal imaging system comprises an infrared lens, an infrared focal plane array detector, a signal acquisition and processing circuit, and an embedded platform for control and display. The infrared lens focuses the infrared radiation emitted by the target object onto the detector's focal plane. The infrared focal plane array detector receives the infrared energy radiated by the target and converts it into an analog electrical signal. The signal acquisition and processing circuit conditions, quantizes, and preprocesses the analog signal output by the detector. The embedded platform serves as the system's main control and interaction unit, performing functions such as parameter configuration, image enhancement, analysis, and display. During data acquisition, a preset number of infrared images of the target object are continuously acquired at a preset frame rate to form time-series image data. The frame rate is 25Hz, and the number of consecutive acquisition frames is set to 30 frames to meet the analysis requirements of the temporal variation characteristics of the images.
[0023] Step S2: Determine the first noise coefficient of the quantized fixed pattern noise by analyzing the statistical characteristics and spatial fitting differences of image patches in the infrared image, and determine the second noise coefficient of the quantized stripe noise by combining the polar coordinate energy spectrum distribution and the directional peak trend.
[0024] Infrared thermal imaging is susceptible to interference from complex noise, leading to a decrease in signal-to-noise ratio. Denoising is a crucial step in ensuring the quality of subsequent detection. Existing technologies mostly filter out single types of noise; even though some deep learning methods attempt to handle mixed noise, their loss functions still treat multiple coupled noises as a whole and apply uniform equalization. Since various noises overlap and their proportions dynamically change in real-world scenarios, this "one-size-fits-all" penalty mechanism cannot implement differentiated constraints, easily leading to a contradiction between insufficient denoising of dominant noise and excessive smoothing of subtle textures. Therefore, overcoming the limitations of uniform processing and achieving refined decoupling of the influence of multiple noises at the model optimization level is key to improving denoising quality.
[0025] Under normal circumstances, an ideal infrared thermal imaging image exhibits continuous and smooth characteristics in the spatial domain, with strong correlation between adjacent pixels and local gradient abrupt changes only at the target edges. In the frequency domain, it is dominated by low-frequency components, with continuous and uniform high-frequency energy distribution. However, when subjected to noise interference from different physical mechanisms, these characteristics undergo specific changes: For fixed-pattern noise, due to the spatial non-uniformity of the gain and bias of the detector pixel response, this noise exhibits a large-scale, slow brightness shift trend on a macroscopic scale, and is accompanied by significant local random fluctuations on a microscopic scale. For stripe noise, the shared readout circuit of pixels in the same row or column causes errors to accumulate consistently along a specific direction, resulting in a continuously extending striped structure in the spatial domain, correspondingly exciting discrete energy peaks in a specific direction in the frequency domain. Furthermore, due to the overlapping characteristics of different noises in the spatial, frequency, and directional domains, severe cross-interference occurs: for example, the low-frequency shift of fixed-pattern noise can distort the stripe structure.
[0026] Based on the above analysis of the physical mechanisms and characteristic variations of different noises, it is clear that to achieve accurate decoupling, differentiated feature extraction paths must be constructed to address their specific characteristics. For fixed-pattern noise, the contradiction between its macroscopic low-frequency shift trend and microscopic local fluctuations needs to be resolved, and quantification is achieved by measuring the degree to which local spatial statistics deviate from the global spatial fitting baseline. For stripe noise, distortion interference from the low-frequency background needs to be avoided, and its discrete energy peaks and directional aggregation characteristics excited in the frequency domain need to be directly anchored. Therefore, this embodiment determines the first noise coefficient of quantized fixed-pattern noise by analyzing the statistical characteristics of image patches and the spatial fitting differences in infrared images, and determines the second noise coefficient of quantized stripe noise by combining the polar coordinate energy spectrum distribution and directional peak trend. The schematic diagrams of the first and second noise extraction processes provided in this embodiment are shown below. Figure 2 As shown, the specific process for obtaining the first and second noise figures is as follows: S2.1: Divide each frame of infrared image into multiple image blocks, calculate the average pixel value of all pixels in each image block, fit the average pixel value of all image blocks in each preset direction, analyze the difference between the average pixel value of each image block and its corresponding fitted value in each direction, the distance of each image block to the nearest edge in the infrared image, and the degree of dispersion of the pixel value of all pixels in each image block, determine the first noise coefficient of each frame of infrared image, which is used to quantify the significance of fixed pattern noise.
[0027] First, each frame of infrared image is divided into multiple image blocks, each with a size of 9×9. In practical applications, as other implementation methods, implementers can also set the image block size according to specific circumstances. This embodiment does not impose any special restrictions.
[0028] Furthermore, this embodiment calculates the average pixel value of all pixels within each image block, fits the average pixel value of all image blocks in each preset direction, analyzes the difference between the average pixel value of each image block and its corresponding fitted value in each direction, the distance of each image block to the nearest edge in the infrared image, and the dispersion of the pixel values of all pixels within each image block, and determines the first noise coefficient of each frame of infrared image. Specifically: In this embodiment, the average pixel value of all pixels in each image block is calculated, and a curve fitting algorithm based on the least squares method is used to fit the average pixel value of all image blocks in each preset direction to obtain the fitting curve in each direction. In practical applications, as other implementation methods, implementers may also use other fitting algorithms to fit the data according to specific circumstances. This embodiment does not impose any special restrictions.
[0029] It should be noted that the preset directions in this embodiment include the horizontal direction, the vertical direction, and the 45° direction.
[0030] The process of fitting data using a curve fitting algorithm based on the least squares method is a well-known technique and will not be described in detail here.
[0031] Furthermore, this embodiment analyzes the difference between the average pixel value of each image block and its corresponding fitted value in each direction, the distance of each image block to the nearest edge in the infrared image, and the dispersion of pixel values of all pixels within each image block to determine the first noise coefficient of each frame of infrared image, specifically: In this embodiment, the mean of the difference between the average pixel value of each image block and its corresponding fitted value in all directions is calculated and denoted as the mean pixel difference. In this embodiment, the absolute difference between the average pixel value of each image block and its corresponding fitted value in all directions is taken as the difference between the average pixel value of each image block and its corresponding fitted value in all directions. In practical applications, as other implementation methods, implementers may also use other methods to measure the difference between data, such as the square or ratio of the difference, depending on the specific circumstances. This embodiment does not impose any special restrictions on the selection of methods to measure the difference between data.
[0032] The first noise coefficient of each frame of infrared image is positively correlated with the dispersion of pixel values of all pixels in each image block, and negatively correlated with the mean pixel difference and the distance of each image block to the nearest edge in the infrared image.
[0033] It should be noted that there are many methods to measure the dispersion of data. In this embodiment, the variance of the pixel values of all pixels in each image block is calculated as the dispersion of the pixel values of all pixels in each image block. In addition, before calculating the distance of each image block to the nearest edge in the infrared image, the Canny edge detection algorithm is used to extract the edges in the infrared image.
[0034] The process of extracting edges from an image using the Canny edge detection algorithm is a well-known technique and will not be elaborated further.
[0035] It should be understood that a positive correlation means that the dependent variable increases as the independent variable increases, and the dependent variable decreases as the independent variable decreases. The specific relationship can be additive or multiplicative, etc., and is determined by the actual application. This application does not impose any special restrictions.
[0036] Preferably, as one implementation, the first noise figure of the i-th frame infrared image in this embodiment... The expression is: In the formula, The mean of the dispersion of all image patches in the i-th frame of the infrared image; This represents the mean of the normalized pixel difference values of image block j in the i-th frame of the infrared image; This represents the normalized distance from image block j in the i-th frame of the infrared image to the nearest edge in the infrared image. This represents the number of all image blocks in the i-th frame of the infrared image.
[0037] It should be noted that there are many commonly used normalization methods. In this embodiment, the pixel difference values of all image blocks and the distances of all image blocks to the corresponding nearest edges in the infrared image are normalized using the maximum-minimum value normalization method. In practical applications, as other implementation methods, implementers may also use other normalization methods according to specific circumstances. This embodiment does not impose any special restrictions on the selection of normalization methods.
[0038] The process of normalizing data using the maximum-minimum normalization method is a well-known technique and will not be elaborated further.
[0039] Based on the first noise coefficient, it can be understood that the first noise coefficient is used to characterize the spatial significance of fixed-pattern noise in a single frame of infrared image. It reflects the severity of the combined influence of macroscopic low-frequency brightness shift and micro-local non-uniformity of distribution on the infrared image. Due to the normalization process introduced in the calculation, the first noise coefficient is a dimensionless parameter. The calculation of the first noise coefficient is directly affected by three factors: the dispersion of pixel values within the image patch, the mean of pixel difference, and the distance from the image patch to the nearest edge. When the dispersion of pixel values within the image patch is greater, it indicates that the local random fluctuations are more severe, and the first noise coefficient increases positively. This reflects the enhanced micro-inconsistency of the detector pixel response, which will cause the model to impose stronger low-frequency smoothing constraints on this region during training. Conversely, when the dispersion of pixel values is smaller, it indicates that the local pixel distribution is more uniform, and the first noise coefficient decreases. This reflects that the micro-perturbation of fixed-pattern noise is weaker, and the model will correspondingly reduce the smoothing pressure to protect texture details. Meanwhile, when the mean pixel difference and the distance from the image patch to the edge are larger, the first noise coefficient will decrease significantly. This is because the real target edge itself has a high gradient and does not conform to the low-frequency shift trend. The larger difference and edge distance convey to the model the signal that "this is a valid edge rather than fixed-pattern noise", thus effectively avoiding the denoising algorithm from misjudging the target edge as a noise structure and over-erasing it. Conversely, when the mean pixel difference and the distance from the image patch to the edge are smaller, the first noise coefficient will increase. This is because a smaller pixel difference means that the local statistics of the corresponding region are highly consistent with the global low-frequency fitting baseline, and a smaller edge distance indicates that the region is far from the real target structure. Both of these convey to the model the signal that "there is no valid edge here, it is just a slow brightness shift caused by fixed-pattern noise", thus prompting the model to apply stronger low-frequency smoothing constraints to the corresponding region during training, achieving complete filtering of this type of background noise.
[0040] S2.2: Obtain the energy spectrum of each frame of infrared image in polar coordinates. By analyzing the energy distribution in each polar angle direction and the changing trend of the peak value in the average pixel value of all image blocks in each direction, determine the second noise coefficient of each frame of infrared image, which is used to quantify the intensity of stripe noise.
[0041] First, in this embodiment, each frame of infrared image is used as input to a two-dimensional Fourier transform algorithm, outputting a two-dimensional spectrum with horizontal spatial frequency as the horizontal axis and vertical spatial frequency as the vertical axis. The spectrum is then transformed into a polar coordinate system, where the polar radius dimension... The absolute magnitude of the frequency, polar dimension The spatial direction of frequency is characterized, thereby constructing a polar coordinate energy spectrum with angle as the horizontal axis and frequency magnitude as the vertical axis. In the polar coordinate energy spectrum, the discrete energy peaks excited by the stripe structure with a specific direction in the image are converged and aligned to the same polar angle column, thereby decoupling the directional features in the spatial domain into the energy peak distribution in the polar angle dimension.
[0042] The process of obtaining the two-dimensional spectrum using the two-dimensional Fourier transform algorithm, and the process of converting the rectangular coordinate system to the polar coordinate system are well-known techniques, and will not be described in detail here.
[0043] Furthermore, the total energy in each polar angle direction of the energy spectrum of each frame of infrared image is calculated in the polar coordinate system. The maximum value of the total energy in all polar angle directions is divided by the average value of the total energy in all polar angle directions, and the result is used as the energy feature value of each frame of infrared image. A fitted straight line is obtained by fitting the peak value of the average pixel value of all image blocks in each direction. The absolute value of the slope of the fitted straight line is recorded as the pixel change trend value in each direction. The mean value of the pixel change trend values in all directions is recorded as the pixel change feature value of each frame of infrared image. The peak extraction method adopts an automatic multi-scale peak search algorithm, and the fitting algorithm used in the fitting process is the least squares method. In practical applications, as other implementation methods, implementers may also adopt other fitting methods according to specific circumstances. This embodiment does not impose special restrictions.
[0044] It should be noted that the process of extracting peaks using the automatic multi-scale peak finding algorithm and the process of linearly fitting the data using the least squares method are both well-known techniques, and the specific process will not be described in detail.
[0045] Furthermore, in this embodiment, based on the energy feature value and the pixel change feature value, a second noise coefficient for each frame of infrared image is determined, specifically: In this embodiment, the expression for the second noise coefficient of each frame of infrared image is: In the formula, This represents the second noise coefficient of the i-th frame of the infrared image; This represents the energy feature value of the i-th frame of the infrared image; Let represent the pixel change feature value of the i-th frame of the infrared image; min() represents the minimum value function; This represents a preset constant greater than 0, used to prevent the denominator from being 0. In this embodiment... The value of is 0.01. Provided that the denominator is not zero and does not excessively affect the calculation result, the implementer may also set it according to the specific situation. This embodiment does not impose any special restrictions.
[0046] The second noise figure can be understood as a characterizing of the structured intensity of stripe noise in a single frame of infrared image. It reflects the degree of coupling between the specific directional stripe structure caused by readout circuit errors in the spatial and frequency domains. It is also a dimensionless parameter due to normalization. The calculation of the second noise figure is influenced by both energy eigenvalues and pixel change eigenvalues: the larger the energy eigenvalues and pixel change eigenvalues, the more significantly the second noise figure increases. This reflects a high concentration of energy in a specific direction in the frequency domain and a strong and regular stripe grayscale abrupt change in the spatial domain, indicating that stripe noise dominates image degradation. In this case, the model will be driven to penalize directional gradient errors to forcefully stripe the stripe structure. Conversely, the smaller the energy eigenvalues and pixel change eigenvalues, the smaller the second noise figure decreases. This reflects a lack of regular stripe extension structures in the image, with stripe noise having a weak impact or being severely distorted by other noise. The model will automatically reduce the optimization weights for directional features to prevent artifacts or damage to the true horizontal or vertical texture due to incorrect penalty.
[0047] Thus, this embodiment has conducted a specific decoupling analysis of fixed-pattern noise and stripe noise in the spatial, frequency, and directional domains, constructing a first noise figure and a second noise figure. It successfully transforms the physical manifestation of complex coupled noise into quantifiable mathematical features. This static feature extraction mechanism enables the model to accurately identify the type and intensity of the dominant noise in the current image, laying a data foundation for implementing precise differential constraints in the loss function. It effectively avoids the problems of insufficient denoising or excessive edge smoothing caused by the uniform homogenization of mixed noise in traditional methods.
[0048] Step S3: Smooth the first and second noise coefficients, analyze the differences in the smoothed values of the first and second noise coefficients between adjacent infrared images, construct the first noise drift index and the second noise drift index respectively, and determine the dynamic noise index by combining the differences between the first and second noise coefficients and their corresponding smoothed values.
[0049] The first and second noise figures mentioned above quantify the degree of noise coupling from a static spatial dimension. However, relying solely on static features is insufficient to achieve extremely precise constraints; it is necessary to further introduce inter-frame temporal features for dynamic analysis. Due to the different physical causes of various noises, their evolution patterns in the temporal domain differ significantly: fixed-pattern noise originates from the inherent response bias of pixels and exhibits extremely strong temporal stability between consecutive frames; stripe noise, although subject to temporal fluctuations due to interference from the readout circuit with the input signal, still strictly retains its directional structural characteristics; while random noise originates from independent physical random events and exhibits completely disordered jumps between frames. This significant difference in temporal stability provides a crucial basis for further stripping and quantifying noise from a dynamic perspective.
[0050] Based on the differences in the aforementioned temporal evolution characteristics, this embodiment smooths the first and second noise figures, analyzes the differences in the smoothed values of the first and second noise figures between adjacent infrared images, constructs the first noise drift index and the second noise drift index respectively, and determines the dynamic noise index by combining the differences between the first and second noise figures and their corresponding smoothed values. The specific process is as follows: In this embodiment, firstly, by smoothing the first and second noise coefficients, the differences in the smoothed values of the first and second noise coefficients between adjacent infrared images are analyzed, and a first noise drift index and a second noise drift index are constructed respectively. Specifically: In this embodiment, the first noise coefficient of each frame of infrared image is used as the input of the exponential smoothing algorithm, wherein the step size is set to 1 frame and the sliding window size is set to 30 frames, and the smoothed value of the first noise coefficient of each frame of infrared image is output; the method for obtaining the smoothed value of the second noise coefficient is the same as the method for obtaining the smoothed value of the first noise coefficient; wherein, the process of smoothing the data using the exponential smoothing algorithm is a well-known technique and will not be described in detail.
[0051] Furthermore, the mean of the difference between the first noise coefficient smoothing value and the mean of the difference between the second noise coefficient smoothing value among all adjacent infrared images in a preset number of consecutive infrared images are calculated and denoted as the first noise drift index and the second noise drift index, respectively.
[0052] It should be noted that in this embodiment, the absolute difference of the first noise figure smoothing value between all adjacent infrared frames is taken as the difference of the first noise figure smoothing value between all adjacent infrared frames. In actual application, as other implementation methods, implementers may also use other methods such as the square or ratio of the difference to measure the data difference, depending on the specific situation. This embodiment does not impose any special restrictions.
[0053] Based on the first and second noise drift indices, it can be understood that the first and second noise drift indices are used to characterize the degree of temporal fluctuation of fixed pattern noise and stripe noise between consecutive time frames, respectively. They reflect the absolute amplitude of the dynamic fluctuation of structured noise with the strength of the input signal. Since they are the average of the absolute differences of the normalized smoothing coefficients of adjacent frames, the first and second noise drift indices are also dimensionless parameters. The larger the values of the first and second noise drift indices, the more directly it indicates that the corresponding structured noise has undergone a drastic temporal drift in a short period of time, indicating that hardware factors such as readout circuit deviation are extremely unstable due to changes in scene thermal radiation. Conversely, the smaller the first and second noise drift indices, the more it indicates that the corresponding noise has extremely high static stability in time and is almost unaffected by changes in scene content.
[0054] Furthermore, in this embodiment, the dynamic noise index is determined based on the first noise drift index and the second noise drift index, combined with the differences between the first noise figure and the second noise figure and their corresponding smoothing values. Specifically: The expression for the dynamic noise figure is: In the formula, Represents the dynamic noise index of the i-th frame of the infrared image; , Let represent the first noise drift index and the second noise drift index of the i-th frame infrared image, respectively; , represents the degree of dispersion of the difference between the first noise coefficient and its smoothed value in a preset number of infrared images, and the degree of dispersion of the difference between the second noise coefficient and its smoothed value, respectively; min() represents the minimum value function.
[0055] It should be noted that, in this embodiment, the absolute difference between the first noise coefficient and its smoothed value of a preset number of infrared images is taken as the difference between the first noise coefficient and its smoothed value of the preset number of infrared images; the absolute difference between the second noise coefficient and its smoothed value of the preset number of infrared images is taken as the difference between the second noise coefficient and its smoothed value of the preset number of infrared images; and the variance of the difference between the first noise coefficient and its smoothed value and the variance of the difference between the second noise coefficient and its smoothed value of the preset number of infrared images are taken as the degree of dispersion of the difference between the first noise coefficient and its smoothed value and the degree of dispersion of the difference between the second noise coefficient and its smoothed value of the preset number of infrared images, respectively.
[0056] Based on the dynamic noise index, it can be understood that the dynamic noise index is used to characterize the independent perturbation intensity of pure high-frequency random noise in a time-series image sequence. It reflects the degree of disordered jumps of residual noise between frames after removing the slow drift trend of structured noise. Since it is calculated from dimensionless parameters, it is itself a dimensionless parameter. The calculation of the dynamic noise index is jointly affected by the minimum value of the two types of noise drift indices and the dispersion of the residuals of the first and second noise coefficients: the smaller the minimum value of the first and second noise drift indices, the lower the baseline term of the dynamic noise index, which reflects that the temporal fluctuation of the structured noise floor is extremely suppressed. This provides a clean baseline for purifying random noise features. However, the greater the dispersion of the residuals of the first and second noise coefficients, the greater the dynamic noise index becomes. This reflects that there are a large number of rapid jumps in the original temporal features that cannot be explained by the smoothed baseline, i.e., the dynamic random noise is extremely strong. At this time, the model will be forced to significantly increase the suppression weight of global variance during training. Conversely, the smaller the dispersion of the residuals of the first and second noise coefficients, the smaller the dynamic noise index becomes. This reflects that the inter-frame pixel fluctuations are smooth and the proportion of random noise is extremely low. The model will automatically relax the global smoothing constraint to preserve the high-frequency micro-details of the image to the greatest extent.
[0057] Thus, this embodiment, by introducing inter-frame temporal evolution features and constructing a dynamic noise index, utilizes sliding trend separation and residual variance measurement to separate the disordered jumps of random noise from the slow drift of structured noise in the temporal dimension, making up for the limitations of single-frame static feature extraction. This achieves comprehensive and accurate quantification of fixed-pattern noise, stripe noise, and random noise in images in both spatial and temporal dimensions.
[0058] Step S4: Based on the first and second noise coefficients and the dynamic noise index, determine the noise weights and optimize the loss function to train the infrared thermal imaging denoising model.
[0059] Following the data acquisition method in step S1, γ sample data from different scenarios are collected. In this embodiment, γ is set to 2000. The first noise coefficient, second noise coefficient, and dynamic noise index for each sample are calculated according to the method in step S2. Further, each sample, along with its corresponding first and second noise coefficients and dynamic noise index, forms a dataset. 75% of the dataset is randomly selected from all samples as the training set, and the remaining 25% is used as the test set. The U-Net model is trained using the training set. The U-Net model adopts an encoder-decoder structure, with 5 encoder layers, a 3×3 kernel size, ReLU activation, Adam optimizer, a learning rate of 0.0001, a batch size of 8, and 100 training epochs. To enhance the model's ability to filter different types of noise, this embodiment determines noise weights based on the first and second noise coefficients and the dynamic noise index, and optimizes the loss function to train the infrared thermal imaging denoising model. The specific process is as follows: The optimized loss function value corresponding to the s-th frame of the infrared image The expression is: ; This represents the denoised image obtained by processing the s-th frame of the infrared image in the preset training set of the infrared thermal imaging denoising model using the U-Net model. This represents the reference image of the s-th frame infrared image extracted by the multi-frame adaptive averaging fusion method based on temporal registration; This represents the noise weight for the m-th value of the s-th frame of the infrared image. This represents the total number of noise weight values for the s-th frame of the infrared image; The residual value is obtained by considering the noise weight based on the m-th value and the residual map between the denoised infrared image of the s-th frame and its corresponding reference image; || represents the L2 norm.
[0060] It should be noted that the noise weight can take three values: The first type of noise weight is the proportion of the first noise figure to the sum of the first noise figure, the second noise figure, and the dynamic noise index. The second type of noise weight is the proportion of the second noise figure in the sum of the first noise figure, the second noise figure, and the dynamic noise index. The third type of noise weight is the proportion of the dynamic noise index in the sum of the first noise figure, the second noise figure, and the dynamic noise index.
[0061] It should be noted that before calculating the proportion, a preset constant greater than 0 needs to be added to the sum of the first noise figure, the second noise figure, and the dynamic noise index to prevent the denominator from being 0. In this embodiment, the preset constant greater than 0 is set to 0.01. Under the premise of ensuring that the denominator is not 0 and does not excessively affect the calculation result, the implementer can also set it according to the specific situation. This embodiment does not impose any special restrictions.
[0062] Based on the noise weight, it can be understood that the noise weight, as a bridge connecting physical noise quantization and gradient allocation in deep networks, is a dimensionless proportional parameter. Its three different values represent the proportion of penalty gradients applied by the model to fixed-pattern noise, stripe noise, and random noise in a single iteration: The first value reflects that low-frequency spatial offset is absolutely dominant in infrared images, and the model focuses its optimization efforts on suppressing the mean square error of low-frequency components; the second value reflects that the strip structure error in a specific direction is the most prominent, and the model is forced to increase the gradient penalty in the horizontal and vertical directions; the third value reflects that high-frequency disturbances that are not correlated between frames are the core contradiction restricting image quality, and the model will correspondingly increase the optimization pressure on the global pixel variance of the residual map. Through this adaptive weight allocation based on the proportion of real-time noise, the limitation of the traditional loss function's "one-size-fits-all" equalization penalty is completely broken.
[0063] To further explain, the specific process of acquiring the denoised image is as follows: In the preset training set, a single frame of noisy infrared image is input into the constructed U-Net denoising model. First, it is processed by an encoder network containing five downsampling layers to extract deep spatial features, mapping the image to a high-dimensional latent space and capturing multi-scale noise and texture features. Subsequently, the feature map is upsampled and stitched layer by layer through the decoder network to gradually restore the spatial resolution. During this process, the model, guided by the loss function with differentiated noise weights constructed in step S4, continuously optimizes the network parameters through backpropagation. Finally, the decoder outputs a denoised image with the same size as the input image. This image retains the details of the infrared target to the maximum extent while effectively suppressing coupled noise. The extraction process of the reference image is as follows: The extraction of the reference image is based on the multi-frame adaptive averaging fusion method of temporal registration. For continuously acquired infrared images, the motion vector between adjacent frames is first calculated using a feature registration algorithm with sub-pixel accuracy to eliminate spatial position deviations caused by the detector micro-scanning mechanism or small relative motion of the scene to achieve strict alignment. After completing the temporal registration, the aligned multi-frame image sequence is fused by point-by-point weighted averaging in the pixel dimension. This process utilizes the zero correlation characteristic of random noise in the temporal domain (i.e., the principle of non-correlation accumulation cancellation) to significantly reduce dynamic random noise. At the same time, the number of fused frames is adjusted through an adaptive weighting mechanism to balance the suppression effect of fixed pattern noise and edge ghosting phenomenon, and finally a static reference image with high signal-to-noise ratio and low distortion is generated.
[0064] The process of obtaining the denoised image using the U-Net model, the process of extracting the reference image using the multi-frame adaptive averaging fusion method based on temporal registration, and the training process of the U-Net model are all well-known techniques and will not be elaborated further.
[0065] Thus, this embodiment transforms the three types of noise quantification indicators extracted by decoupling into differentiated noise weights and injects them into the loss function, giving the denoising model the ability to adaptively adjust the optimization direction according to the real-time noise ratio. This eliminates the defects of traditional methods in uniformly equalizing mixed noise from the underlying logic of model training, improves the accuracy of filtering out complex coupled noise, and preserves faint infrared details with high fidelity.
[0066] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments of this specification have been described above. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0067] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0068] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; modifications to the technical solutions described in the foregoing embodiments, or equivalent substitutions of some of the technical features, do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for denoising and optimizing infrared thermal imaging based on deep learning, characterized in that, The method includes the following steps: Real-time acquisition of infrared images of the target object; Each frame of infrared image is divided into multiple image blocks. The average pixel value of all pixels in each image block is calculated. The average pixel value of all image blocks in each preset direction is fitted. The difference between the average pixel value of each image block and the corresponding fitted value in each direction is analyzed. The distance of each image block to the nearest edge in the infrared image and the degree of dispersion of the pixel value of all pixels in each image block are also analyzed. The first noise coefficient of each frame of infrared image is determined and used to quantify the significance of fixed pattern noise. The energy spectrum of each frame of infrared image in polar coordinates is obtained. By analyzing the energy distribution in each polar angle direction and the changing trend of the peak value in the average pixel value of all image blocks in each direction, the second noise coefficient of each frame of infrared image is determined, which is used to quantify the intensity of stripe noise. The first and second noise coefficients are smoothed, and the differences in the smoothed values of the first and second noise coefficients between adjacent infrared images are analyzed. A first noise drift index and a second noise drift index are constructed, including: calculating the mean of the differences in the smoothed values of the first and second noise coefficients between all adjacent infrared images in a preset number of consecutive frames, denoted as the first noise drift index and the second noise drift index, respectively; and determining the dynamic noise index by combining the differences between the first and second noise coefficients and their corresponding smoothed values. The expression for the dynamic noise index is: In the formula, Represents the dynamic noise index of the i-th frame of the infrared image; , Let represent the first noise drift index and the second noise drift index of the i-th frame infrared image, respectively; , These represent the degree of dispersion of the difference between the first noise coefficient and its smoothed value, and the degree of dispersion of the difference between the second noise coefficient and its smoothed value, respectively, for a preset number of infrared image frames; min() represents the minimum value function; Based on the first and second noise figures and the dynamic noise index, noise weights are determined, and the loss function is optimized for training the infrared thermal imaging denoising model. The optimization of the loss function includes: The optimized loss function value corresponding to the s-th frame of the infrared image The expression is: ; This represents the denoised image obtained by processing the s-th frame of the infrared image in the preset training set of the infrared thermal imaging denoising model using the U-Net model. This represents the reference image of the s-th frame infrared image extracted by the multi-frame adaptive averaging fusion method based on temporal registration; This represents the noise weight for the m-th value of the s-th frame of the infrared image. This represents the total number of noise weight values for the s-th frame of the infrared image; The residual value is obtained by considering the noise weight based on the m-th value and the residual map between the denoised infrared image of the s-th frame and its corresponding reference image; || represents the L2 norm.
2. The infrared thermal imaging denoising and imaging optimization method based on deep learning as described in claim 1, characterized in that, The method for determining the first noise coefficient of each frame of infrared image is as follows: The mean difference between the average pixel value of each image patch and its corresponding fitted value in all directions is calculated and denoted as the mean pixel difference. The first noise coefficient of each frame of infrared image is positively correlated with the dispersion of pixel values of all pixels in each image block, and negatively correlated with the mean pixel difference and the distance of each image block to the nearest edge in the infrared image.
3. The infrared thermal imaging denoising and imaging optimization method based on deep learning as described in claim 1, characterized in that, The method for determining the second noise coefficient of each frame of infrared image is as follows: Calculate the total energy in each polar angle direction of the energy spectrum of each frame of infrared image in polar coordinate system, and take the result of dividing the maximum value of the total energy in all polar angle directions by the average value of the total energy in all polar angle directions as the energy feature value of each frame of infrared image. A fitted line is obtained by fitting the peak value of the average pixel value of all image blocks in each direction. The absolute value of the slope of the fitted line is recorded as the pixel change trend value in each direction. The average of the pixel change trend values in all directions is recorded as the pixel change feature value of each frame of infrared image; Based on the energy characteristic value and the pixel change characteristic value, the second noise coefficient of each frame of infrared image is determined.
4. The infrared thermal imaging denoising and imaging optimization method based on deep learning as described in claim 3, characterized in that, The expression for the second noise coefficient of each frame of infrared image is: In the formula, This represents the second noise coefficient of the i-th frame of the infrared image; This represents the energy feature value of the i-th frame of the infrared image; Let represent the pixel change feature value of the i-th frame of the infrared image; min() represents the minimum value function; This indicates a constant that is pre-defined as being greater than 0.
5. The infrared thermal imaging denoising and imaging optimization method based on deep learning as described in claim 1, characterized in that, The method for determining the noise weight is as follows: The noise weight can take three values: The first type of noise weight is the proportion of the first noise figure to the sum of the first noise figure, the second noise figure, and the dynamic noise index. The second type of noise weight is the proportion of the second noise figure in the sum of the first noise figure, the second noise figure, and the dynamic noise index. The third type of noise weight is the proportion of the dynamic noise index in the sum of the first noise figure, the second noise figure, and the dynamic noise index.
6. The infrared thermal imaging denoising and imaging optimization method based on deep learning as described in claim 1, characterized in that, The residual value obtained based on the noise weight of the m-th value and the residual map between the denoised infrared image of the s-th frame and its corresponding reference image includes: Under the first noise weight, the residual map between the denoised infrared image of the s-th frame and its corresponding reference image is used as the input of the low-pass filtering algorithm, and the filtered residual map is output. The average of the squares of all pixel values in the filtered residual map is used as the residual value of the denoised infrared image of the s-th frame. Under the second noise weight, the sum of the gradient magnitudes of all pixels in the residual image between the denoised infrared image of frame s and its corresponding reference image is calculated, and the sum of the gradient magnitudes in the horizontal direction and the vertical direction are recorded as the horizontal gradient sum and the vertical gradient sum, respectively. The mean of the horizontal gradient sum and the vertical gradient sum is used as the residual value of the denoised infrared image of frame s. Under the third noise weight value, the variance of all pixel values in the residual map between the denoised infrared image of frame s and its corresponding reference image is calculated and used as the residual value of the denoised infrared image of frame s.
7. The infrared thermal imaging denoising and imaging optimization method based on deep learning as described in claim 1, characterized in that, The trained infrared thermal imaging denoising model includes: All infrared images in the preset training set are used as input to the neural network, the corresponding reference images are used as target labels, the optimized loss function is used as the loss function in the neural network, the neural network is trained, and the trained neural network model is output, which is denoised as the infrared thermal imaging denoising model.
Citation Information
Patent Citations
A variance gradient constraint method infrared image edge preserving and denoising method
CN109559286A
Multi-source noise removal method and system for infrared imaging system
CN121724863A