Infrared image heterogeneity joint correction method based on histogram and neural network

The combination of histogram matching and neural networks effectively addresses non-uniformity and noise issues in infrared imaging, enhancing image quality and object identification.

CN120318084APending Publication Date: 2025-07-15XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510359107.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

There are problems of thermal noise, detector inhomogeneity and time drift in the existing infrared image technology, which affects the effect of image recognition and tracking tasks.

Method used

The combined correction method based on histogram and neural network is adopted to reconstruct infrared video sequences through histogram correction and trained neural network model to reduce background noise and improve image contrast, and image reconstruction is performed using CNN to reduce residual inhomogeneity.

Benefits of technology

The quality of infrared images is significantly improved, background noise and mean square error are reduced, and the distinction of objects in the image is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318084A_ABST
    Figure CN120318084A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared image heterogeneity joint correction method based on a histogram and a neural network. The infrared image heterogeneity joint correction method comprises the following steps: acquiring a to-be-corrected infrared video sequence; wherein the infrared video sequence comprises L video frames; performing histogram correction on the ith video frame by adopting the first video frame to obtain the ith video frame after initial correction; the value of i is 2-L; performing non-uniformity measurement value estimation on the first video frame to obtain a non-uniformity correction measurement value matrix; based on the non-uniformity measurement value matrix and the trained neural network model, reconstructing the ith initially corrected video frame to obtain an ith corrected video frame; and taking a video sequence formed by the first video frame and the obtained L-1 corrected video frames according to a time sequence as a corrected infrared video sequence. According to the invention, thermal noise can be eliminated, the non-uniformity of the detector is reduced, and the overall quality of the image sequence is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of infrared image correction, and particularly relates to a method for jointly correcting non-uniformity of infrared images based on histogram and neural network. Background Technique

[0002] Infrared imaging technology has been widely used in many fields such as military, scientific research and commerce due to its excellent anti-interference ability, including military reconnaissance, night driving, disease detection, forest fire monitoring, weather forecasting and abnormal event identification. However, due to deficiencies in device manufacturing and theoretical basis, there is still much room for improvement in infrared focal plane arrays and imaging technology. In the context of increasingly fierce competition, it is particularly important to perform a series of restoration and enhancement on the infrared images generated by the focal plane. During the acquisition process of infrared image sequences, complex problems are faced. In addition to the inherent low contrast, other noises may also appear. As the temperature of the camera increases, this noise will increase, resulting in more background noise, thus masking the target object and affecting the recognition and tracking tasks. In addition, the difference in light response between each individual detector and adjacent detectors in the array, as well as fluctuations in temperature, unstable bias voltage and scene illumination, will all cause time drift, manifested as slow changes in the image. Summary of the Invention

[0003] In order to solve the above problems existing in the prior art, the present invention provides a method for jointly correcting non-uniformity of infrared images based on histogram and neural network, aiming to eliminate thermal noise, reduce the non-uniformity of detectors, and improve the overall quality of image sequences.

[0004] The technical problems to be solved by the present invention are realized through the following technical solutions:

[0005] The present invention provides a method for jointly correcting non-uniformity of infrared images based on histogram and neural network, including:

[0006] Obtain an infrared video sequence to be corrected; wherein, the infrared video sequence includes L video frames;

[0007] Perform histogram correction on the i-th video frame using the first video frame to obtain the i-th initially corrected video frame; the value of i ranges from 2 to L; L is a positive integer greater than 1;

[0008] Estimate the non-uniformity measurement value of the first video frame to obtain a non-uniformity correction measurement value matrix;

[0009] Based on the non-uniformity measurement value matrix and the trained neural network model, reconstruct the i-th initially corrected video frame to obtain the i-th corrected video frame;

[0010] The video sequence formed by arranging the first video frame and the obtained L - 1 corrected video frames in chronological order is used as the corrected infrared video sequence.

[0011] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0012] First, through histogram framework matching, the present invention reduces the average gradient, edge intensity, and entropy of infrared images, improves the contrast, reduces background noise, and makes objects easier to distinguish in infrared images. Then, a neural network based on CNN is used to reconstruct the images, greatly reducing the residual non - uniformity and mean square error of the images. Thus, through the implementation of joint non - uniformity correction of histogram and neural network, the present invention significantly improves the quality of the obtained infrared images.

[0013] The following will further elaborate on the present invention in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings

[0014] Figure 1 is a flowchart of a method for joint non - uniformity correction of infrared images based on histogram and neural network provided by an embodiment of the present invention;

[0015] Figure 2 is another flowchart of a method for joint non - uniformity correction of infrared images based on histogram and neural network provided by an embodiment of the present invention;

[0016] Figure 3 is a flowchart for histogram correction of a video frame provided by an embodiment of the present invention;

[0017] Figure 4 is a structural schematic diagram of a neural network model provided by an embodiment of the present invention;

[0018] Figure 5 is a flowchart of the ECA module executing the attention mechanism provided by an embodiment of the present invention. Detailed Embodiments

[0019] The following further describes the present invention in detail with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0020] The present invention provides a method for joint non - uniformity correction of infrared images based on histogram and neural network, aiming to solve the deficiencies of existing models in simulating temperature changes and dependence on scene motion, eliminate thermal noise, reduce the non - uniformity of detectors, and improve the overall quality of image sequences.

[0021] Figure 1 is a flowchart of a method for joint non - uniformity correction of infrared images based on histogram and neural network provided by an embodiment of the present invention;Figure 2 Another flowchart of the infrared image non-uniformity joint correction method based on histogram and neural network provided by the embodiment of the present invention; combined with Figure 1 and Figure 2 as shown, the method includes:

[0022] S101. Obtain an infrared video sequence to be corrected; wherein, the infrared video sequence includes L video frames.

[0023] S102. Perform histogram correction on the i-th video frame using the first video frame to obtain the i-th initially corrected video frame; the value of i ranges from 2 to L; L is a positive integer greater than 1.

[0024] S103. Estimate the non-uniformity correction measurement value for the first video frame to obtain a non-uniformity correction measurement value matrix.

[0025] S104. Reconstruct the i-th initially corrected video frame based on the non-uniformity correction measurement value matrix and the trained neural network model to obtain the i-th corrected video frame.

[0026] S105. Use the first video frame and the L - 1 corrected video frames obtained to form a video sequence in chronological order as the corrected infrared video sequence.

[0027] Exemplarily, Figure 3 is a flowchart for performing histogram correction on a video frame; combined with Figure 3 , in the present invention, the above S102 can be implemented through the following steps:

[0028] S1021. Calculate the average value of the pixel values of the first video frame g0(x, y) and the average value of the pixel values of the i-th video frame g i (x, y)

[0029] Specifically, and are respectively expressed as follows:

[0030]

[0031]

[0032] wherein, X is the number of pixel rows of a video frame, and Y is the number of pixel columns in a video frame.

[0033] S1022. Calculate the standard deviation σ0 of the pixel values of the first video frame using the average value of the pixel values of the first video frame and calculate the standard deviation of the pixel values of the i-th video frame using the average value of the pixel values of the i-th video frame Calculate the standard deviation σ of the pixel values of the i-th video frame i .

[0034] Specifically, the expressions for σ0 and σ i are as follows:

[0035]

[0036]

[0037] S1023. Use the standard deviations σ0 and σ i to calculate the multiplicative correction factor K of the i-th video frame i .

[0038] Specifically, the expression for K i is:

[0039] S1024. Use the average value and the multiplicative correction factor K of the i-th video frame i to calculate the additional correction factor of the i-th video frame

[0040] Specifically, the expression for

[0041] S1025. Use the multiplicative correction factor K i and the additional correction factor to perform histogram correction on the i-th video frame to obtain the i-th initially corrected video frame.

[0042] Specifically, the expression for the i-th initially corrected video frame g ic (x, y) is as follows:

[0043] Through histogram matching correction, the present invention successfully reduces the background noise of infrared images, enhances the distinguishability of objects in the infrared image sequence, and improves the overall image quality. Specifically, this process effectively eliminates the influence of background noise by reducing the average gradient, edge intensity, and entropy across frames, thereby improving the quality of the obtained infrared image sequence.

[0044] In the present invention, the above S103 can be implemented through the following steps:

[0045] S1031. Use a small overlapping window with a size of S×S and slide it from the first pixel of the first video frame to the last pixel of the first video frame. Each time, slide one pixel. After each slide, calculate the variance corresponding to the first pixel within the current window, so that the variance corresponding to each pixel in the first video frame can be obtained. S is a positive integer greater than 1.

[0046] The small overlapping window is a sliding window technique widely used in data sequences or image processing. Among them, the size of each window is relatively small, and there is a certain overlapping area between adjacent windows. The degree of overlap can be adjusted according to specific application requirements. It should be noted that the value of S can be set according to actual needs.

[0047] Specifically, the calculation formula for the variance V(x, y) corresponding to each pixel (x, y) in a video frame is: where, when the pixel (x, y) is the first pixel within a window, Ω is the small neighborhood around the pixel (x, y), that is, Ω is the remaining pixels within the window, I(u, v) represents each pixel in Ω, |Ω| represents the number of elements in Ω, and μ x,y is the average value of the pixel values in the neighborhood Ω.

[0048] S1032. After obtaining the variance corresponding to each pixel in the first video frame, use a small overlapping window with a size of S×S to continue sliding from the first pixel of the first video frame to the last pixel of the first video frame. Each time, slide one pixel. After each slide, select a variance from the variances corresponding to the pixels within the current window as the non-uniformity measurement value η corresponding to the first pixel within the current window. The matrix composed of the non-uniformity measurement values η corresponding to all pixels in the first video frame is used as the non-uniformity measurement value matrix.

[0049] Specifically, after each slide, select a median variance from the variances corresponding to all pixels within the current window as the non-uniformity measurement value η corresponding to the first pixel within the current window. For example, when there are 4 variances, sort these 4 variances in order of size to obtain a variance sequence, and select the median of the variance sequence as the non-uniformity measurement value η. Specifically, it is expressed as

[0050] The present invention calculates the variance of the pixel values of video frames and combines it with the non-uniformity measurement value. This non-uniformity correction measure aims to capture the non-uniformity fluctuations. The currently common solution is to calculate the variance between the responses of each detector element under a constant-temperature uniform blackbody surface to address the non-uniformity characteristics caused by the differences in the response rates of detector elements. However, this method has poor effects in non-uniform scenes, mainly because it cannot effectively distinguish the non-uniformity of the focal plane array from the non-uniformity of the scene itself. However, the method proposed by the present invention effectively avoids the influence of neighborhood outliers on the non-uniformity correction measurement method proposed by the present invention. In a relatively harsh usage environment, for example, when the infrared focal plane is exposed to a dynamic moving scene, this method can still ensure the correlation between non-uniformity measurement and temperature, indicating that the measures proposed by the present invention are independent of the scene content. This method utilizes the characteristics of a small neighborhood and can significantly weaken the influence of scene non-uniformity on the calculation results.

[0051] Feature extraction is the basis of the attention mechanism. Feature extraction can achieve dimensionality reduction, denoising, and information compression of images, thereby extracting useful information. Both feature extraction and the attention mechanism rely on convolutional operations. In the present invention, the trained neural network model includes: a downsampling layer, three parallel convolutional layers, a feature fusion layer, an efficient channel attention module, and an upsampling layer. Specifically, the input of the downsampling layer is used as the input of the neural network model, the output of the downsampling layer is connected to the input of the three parallel convolutional layers, the output of the three parallel convolutional layers is connected to the input of the feature fusion layer, the output of the feature fusion layer is connected to the input of the efficient channel attention module, the output of the efficient channel attention module is connected to the input of the upsampling layer, and the output of the upsampling layer is used as the output of the neural network model. In the present invention, the dimensionality of image features can be correspondingly shrunk and expanded through downsampling and upsampling. Based on the above network structure, the above S104 can be implemented through the following steps:

[0052] S1041. Stitch the non-uniformity correction measurement value matrix and the i-th initially corrected video frame to obtain a stitched feature.

[0053] S1042. Use the stitched feature as the input of the trained neural network model. The trained neural network model performs a downsampling operation on the stitched feature to obtain a downsampled feature, performs multi-scale feature extraction on the downsampled feature, and performs feature fusion on the extracted multi-scale features to obtain a multi-scale fused feature. Perform efficient channel attention processing on the multi-scale fused feature to obtain a weighted feature map, and perform an upsampling operation on the weighted feature map to obtain the i-th corrected video frame.

[0054] In the present invention, the original image is fed together with the non-uniformity measurement into a trained neural network model, and through processing using the neural network model, a non-uniformity corrected image is finally obtained. This process ensures that the quality of the image is significantly improved, making the effect of non-uniformity correction more remarkable. By generating a weighted feature map, different weights can be assigned to each position of the feature map.

[0055] In some embodiments, after each of the three parallel convolutional layers, there is also a Leaky ReLU activation function to introduce non-linearity.

[0056] In the present invention, through upsampling, the reduced image features are gradually restored to the original image scale for deconvolution decoding. Downsampling can use a 2×2 max-pooling layer with a stride of 2, which can reduce the size of the input data, enabling the network to capture more useful features. Upsampling uses a deconvolution layer to expand the input image scale to twice the original.

[0057] In the present invention, the convolutional layer serves as a multi-scale feature extraction unit, and the sizes of the convolutional kernels used in the three parallel convolutional layers and the feature fusion layer are all different. Exemplarily, Figure 4 is a schematic structural diagram of the neural network model. As Figure 4 shown, the convolutional layer serves as a multi-scale feature extraction unit. Through the multi-scale feature extraction unit, not only the adaptability to targets of different scales is improved, but also the expression ability of the network is enhanced without increasing the network depth, effectively avoiding the gradient dispersion phenomenon. Specifically, among these three parallel convolutional layers, the first convolutional layer is composed of a cascade of 32 1×5 asymmetric convolutional kernels and 32 5×1 asymmetric convolutional kernels. Among them, the cascade of 1×5 and 5×1 asymmetric convolutional kernels is equivalent to forming a 5×5 convolutional kernel in the receptive field. The second convolutional layer is composed of 64 3×3 convolutional kernels, and the third convolutional layer is composed of 32 1×1 convolutional kernels. The feature fusion layer is composed of 64 1×1 convolutional kernels. Figure 4 The attention mechanism in out (x,y) = ∑ m ∑ n K(m,n)·I(x+m,y+n) + b′, where F out (x,y) is the pixel value of the output feature map at position (x,y), K(m,n) is the weight of the convolutional kernel (also known as the filter or feature detector) at position (m,n). m and n are the indices on the width and height of the convolutional kernel, I(x+m,y+n) is the pixel value of the input image at position (x+m,y+n), and b′ is the bias term of the convolutional kernel.

[0058] In the present invention, when performing convolution and deconvolution operations, a method of zero-padding at the image edge is adopted, which can keep the size of the infrared image unchanged. This method effectively suppresses the degradation phenomenon at the infrared image edge. Zero-padding is an edge-padding strategy that ensures the consistency of the image size after convolution and deconvolution operations by adding a circle of zero-valued pixels at the image edge.

[0059] In the present invention, the core idea of the ECA module is to capture the inter-channel dependencies through one-dimensional convolution. Compared with traditional attention mechanisms, the ECA module avoids the complex processes of dimensionality reduction and dimensionality increase, thus achieving the characteristics of high efficiency and light weight. At the same time, the ECA module does not require operations of dimensionality reduction and dimensionality increase, so it can retain the information integrity of the original channel features. The ECA module adaptively calculates the kernel size k of one-dimensional convolution according to the number of channels. After obtaining the kernel size k, the ECA module applies one-dimensional convolution to the input features to learn the importance of each channel relative to other channels. Compared with traditional attention mechanisms, the ECA module does not need to perform operations of dimensionality reduction and dimensionality increase, thus retaining the information integrity of the original channel features. This design helps the model to better utilize the inter-channel dependencies and improve the feature representation ability. The ECA module can adaptively adjust the kernel size of one-dimensional convolution according to the number of channels, enabling it to flexibly capture channel dependencies within different ranges. This adaptive mechanism enables the ECA module to work effectively in networks of different scales and at different depths of the hierarchy, thus improving the overall performance and efficiency of the model. Exemplarily, Figure 5 is a flowchart for the ECA module to execute the attention mechanism. As Figure 5 shown, the efficient channel attention module is used to perform global average pooling on the features output by the feature fusion layer to obtain pooled features. Through global average pooling, it can adapt to different input sizes and prevent overfitting. Then, the channel weights w of the pooled features are calculated, the kernel size k of one-dimensional convolution is calculated according to the number of channels of the pooled features, and one-dimensional convolution operation is performed on the pooled features according to the channel weights and the kernel size k of one-dimensional convolution to obtain a spatial attention map. The spatial attention map is multiplied by the features output by the feature fusion layer to obtain a weighted feature map. It should be noted that calculating the channel weights w of the pooled features is a conventional technique, and the present invention does not elaborate on this. Specifically, the calculation formula for the kernel size k is as follows: where C is the number of channels of the input features, γ and b are hyperparameters, and to ensure that the kernel size k is an odd number, the absolute value is taken and rounded down to the nearest odd number. After obtaining the kernel size k, the ECA module applies one-dimensional convolution to the input features to learn the importance of each channel relative to other channels. Specifically, converting the input feature in to the output feature out through a one-dimensional convolution operation with a kernel size of k can be represented by the following formula: out = Conv1D k (in), Conv1Dk Denotes a one-dimensional convolution operation with a kernel size of k.

[0060] In the present invention, the dataset used for training the neural network model can be existing publicly available infrared image datasets. For example, the infrared dataset of MIT and the remote sensing image library of NASA, etc. These datasets are usually labeled and suitable for model training and testing. Thus, the present invention obtains rich data resources, meeting the dataset requirements for accurate image processing and analysis. The loss function used for training the neural network model is the sum of squared errors E of the above three parallel convolutional layers. During training, this loss function needs to be minimized to optimize the model. Specifically, the expression of this loss function is as follows:

[0061]

[0062] where f represents the function of the proposed neural network model, and I n (x, y) is the image input into this neural network model, w represents the weights of the network, T represents the transpose, N is the total number of weight updates performed in one epoch, and the baseline J n is obtained by taking the average of the pixel values of the nth frame image. To avoid overfitting, the present invention adds an L2 regularization term to the loss function, where v is the regularization factor. The network weights w are updated using the Stochastic Gradient Descent (SGD) optimizer. The SGD weight update rule uses an adaptive learning rate α, and the weight update formula is: where n is the update batch index, that is, it represents the nth update. At the same time, the present invention uses an adaptive learning rate strategy, and its learning rate α0 gradually decreases within one epoch:

[0063] The present invention processes the original image using the neural network model and finally obtains a non-uniformity correction image. This process ensures that the quality of the image is significantly improved, making the non-uniformity correction effect more remarkable.

[0064] It should be noted that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0065] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0066] In the specification, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude the case of a plurality. Certain measures are recited in different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0067] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A joint non-uniformity correction method for infrared images based on histogram and neural network, characterized in that Including: Obtain an infrared video sequence to be corrected; wherein, the infrared video sequence includes L video frames; Perform histogram correction on the i-th video frame using the first video frame to obtain the i-th initially corrected video frame; the value range of i is 2 to L; L is a positive integer greater than 1; Estimate the non-uniformity measurement values for the first video frame to obtain a non-uniformity correction measurement value matrix; Based on the non-uniformity measurement value matrix and the trained neural network model, reconstruct the i-th initially corrected video frame to obtain the i-th corrected video frame; Use the first video frame and the L - 1 corrected video frames obtained in chronological order to form a video sequence as the corrected infrared video sequence.

2. The infrared image non-uniformity joint correction method based on histogram and neural network according to claim 1, characterized in that, The step of performing histogram correction on the i-th video frame using the first video frame to obtain the i-th initially corrected video frame includes: Calculate the average value of the pixel values of the first video frame respectively and the average value of the pixel values of the i-th video frame Adopt the average value of the pixel values of the first video frame Calculate the standard deviation σ0 of the pixel values of the first video frame, and adopt the average value of the pixel values of the i-th video frame Calculate the standard deviation σ of the pixel values of the i-th video frame i ; Using the standard deviations σ0 and σ i Calculate the multiplication correction factor K for the i-th video frame i ; Using the mean value and the multiplication correction factor K of the i-th video frame i to calculate the additional correction factor of the i-th video frame Using the multiplication correction factor K i and the additional correction factor perform histogram correction on the i-th video frame to obtain the i-th initially corrected video frame.

3. The infrared image non-uniformity joint correction method based on histogram and neural network according to claim 2, characterized in that K i and The expressions are as follows:

4. The infrared image non-uniformity joint correction method based on histogram and neural network according to claim 2, characterized in that The expression of the i-th initially corrected video frame is as follows: where g ic (x, y) is the i-th initially corrected video frame, and g i (x, y) is the i-th video frame.

5. The infrared image non-uniformity joint correction method based on histogram and neural network according to claim 1, characterized in that The step of estimating the non-uniformity measurement values for the first video frame to obtain a non-uniformity measurement value matrix includes: Use a small overlapping window with a size of S×S to slide from the first pixel of the first video frame to the last pixel of the first video frame, where each time it slides one pixel, and calculate the variance corresponding to the first pixel within the current window after each slide; S is a positive integer greater than 1; After obtaining the variance corresponding to each pixel in the first video frame, continue to use the small overlapping window with a size of S×S to slide from the first pixel of the first video frame to the last pixel of the first video frame, where each time it slides one pixel, and select one variance from the variances corresponding to the pixels within the current window as the non-uniformity measurement value corresponding to the first pixel within the current window, and use the matrix composed of the non-uniformity measurement values corresponding to all pixels in the first video frame as the non-uniformity measurement value matrix.

6. The non-uniformity joint correction method for infrared images based on histogram and neural network according to claim 1, characterized in that The trained neural network model includes: a downsampling layer, three parallel convolutional layers, a feature fusion layer, an efficient channel attention module, and an upsampling layer; The output of the downsampling layer is connected to the input of the three parallel convolutional layers, the output of the three parallel convolutional layers is connected to the input of the feature fusion layer, the output of the feature fusion layer is connected to the input of the efficient channel attention module, and the output of the efficient channel attention module is connected to the input of the upsampling layer.

7. The infrared image non-uniformity joint correction method based on histogram and neural network according to claim 6, characterized in that The efficient channel attention module is used to perform global average pooling on the features output by the feature fusion layer to obtain pooled features, and then calculate the channel weights of the pooled features, calculate the kernel size of the one-dimensional convolution according to the number of channels of the pooled features, perform one-dimensional convolution operation on the pooled features according to the channel weights and the kernel size of the one-dimensional convolution to obtain a spatial attention map, and multiply the spatial attention map and the features output by the feature fusion layer to obtain a weighted feature map.

8. The non-uniformity joint correction method for infrared images based on histogram and neural network according to claim 6, characterized in that, The sizes of the convolution kernels used in the three parallel convolutional layers and the feature fusion layer are all different.

9. The non-uniformity joint correction method for infrared images based on histogram and neural network according to claim 6, characterized in that Reconstructing the i-th initially corrected video frame based on the non-uniformity measurement value matrix and the trained neural network model to obtain the i-th corrected video frame, including: Concatenating the non-uniformity correction measurement value matrix and the i-th initially corrected video frame to obtain a concatenated feature; Using the concatenated feature as the input of the trained neural network model, the trained neural network model performs a downsampling operation on the concatenated feature to obtain a downsampled feature, performs multi-scale feature extraction on the downsampled feature, and performs feature fusion on the extracted multi-scale features to obtain a multi-scale fused feature, performs efficient channel attention processing on the multi-scale fused feature to obtain a weighted feature map, and performs an upsampling operation on the weighted feature map to obtain the i-th corrected video frame.

10. The non-uniformity joint correction method for infrared images based on histogram and neural network according to claim 5, characterized in that, After each sliding, selecting a variance from the variances corresponding to the pixels within the current window as the non-uniformity measurement value corresponding to the first pixel within the current window, including: After each sliding, selecting a variance median from the variances corresponding to all the pixels within the current window as the non-uniformity measurement value corresponding to the first pixel within the current window.