Infrared video denoising method based on zero sample learning
Through the infrared video denoising method with zero sample learning, the symmetric loss function is constructed using adjacent frame training, which solves the problem of infrared video denoising dependent labeled data, realizes efficient and accurate infrared video denoising, improves imaging quality, and is suitable for security monitoring and military reconnaissance and other fields.
Patent Information
- Application Number
- CN202510363987.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-22
AI Technical Summary
The existing infrared video denoising methods rely on a large amount of labeled data training, and their generalization capabilities are limited, making it difficult to adapt to noise changes in complex environments, affecting imaging quality.
The zero-sample learning method is adopted to generate noise data, construct input data pairs, and build symmetric loss functions, and train the front and back frames of adjacent frames to realize infrared video denoising and avoid relying on labeled data.
Without relying on labeled data, the imaging quality of infrared video in complex environments is significantly improved, and is suitable for industrial scenarios with high real-time requirements.
Smart Images

Figure FT_1 
Figure SMS_1 
Figure FDA0005329479840000011
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing and video enhancement. A method based on zero-shot learning is designed for infrared video denoising, aiming to improve the imaging quality of infrared videos in complex environments and is applicable to fields such as security monitoring and military reconnaissance. Background Art
[0002] Infrared imaging technology is widely used in fields such as security monitoring, military reconnaissance, and autonomous driving because it can work at night or in complex meteorological conditions. However, infrared videos are easily disturbed by various noises during the acquisition process, such as thermal noise, dark current noise, etc., resulting in a decline in image quality and affecting subsequent tasks such as target detection, recognition, and tracking.
[0003] Traditional infrared video denoising methods mainly rely on manually designed filters or models, such as median filtering, wavelet transform, etc. Although these methods can remove noise to a certain extent, they often rely on prior assumptions about the noise characteristics and are difficult to adapt to noise changes in different scenarios. In recent years, denoising methods based on deep learning have made significant progress, but most methods require a large number of noisy and clean image pairs as training data, which is often difficult to obtain in practical applications. In addition, existing deep learning methods usually need to be trained for specific types of noise, and their generalization ability is limited.
[0004] Zero-Shot Learning (ZSL), as an emerging machine learning paradigm, aims to use limited labeled data and rich semantic information to learn a model, so as to achieve the recognition and processing of unseen categories. Its core idea is to enable the model to reason and predict unseen categories by learning the semantic associations between known and unknown categories. However, the application of zero-shot learning in the field of infrared video denoising is still relatively rare at present, mainly because the noise characteristics of infrared videos are complex and diverse, and there is a lack of effective semantic description methods.
[0005] To address the above problems, the present invention proposes an infrared video denoising method based on zero-shot learning. This method uses the previous frame and the subsequent frame of the video as the input image and the target image of the denoising network respectively, and is trained using a loss function to form an efficient infrared video denoising method. Through this method, the present invention realizes effective denoising of infrared videos without relying on any labeled data, significantly improves the imaging quality of infrared videos in complex environments, and provides strong support for the wide application of infrared videos. Summary of the Invention
[0006] The present invention designs an infrared video denoising method based on zero - sample learning. It is achieved through three steps: generating noise data, constructing input data pairs, and building a symmetric loss function. Without relying on any labeled data, it can efficiently and accurately denoise infrared videos, significantly improving the imaging quality of infrared videos in complex environments.
[0007] The present invention is realized through the following technical solutions, including the following steps:
[0008] The first step: Generate noise data;
[0009] The second step: Construct input data pairs;
[0010] The third step: Build a symmetric loss function.
[0011] The creativity of the present invention is mainly reflected in:
[0012] Different from the vast majority of existing video denoising methods that are trained through large - scale data sets, the present invention newly designs a zero - sample - based infrared video denoising method. Specifically, the present invention only needs to input an infrared video into the denoising network model. By using the previous frame and the subsequent frame of adjacent frames of the video as the input image and the target image of the denoising network model respectively, and then training with a symmetric loss function, the function of denoising infrared videos can be realized. This method does not need to rely on any labeled data during training. Compared with video denoising methods that rely on large - scale data sets for training, this method can greatly reduce the training time and is more suitable for industrial scenarios with strict real - time requirements. Brief Description of the Drawings
[0013] Figure 1 : Schematic diagram of the input data flow in the infrared video denoising method based on zero - sample learning Detailed Embodiment
[0014] The following details the embodiments of the present invention. These embodiments are implemented on the premise of the technical solutions of the present invention, and detailed implementation methods and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.
[0015] Embodiment:
[0016] The first step: Generate noise data;
[0017] First, select an infrared video with clear content and no any distortion interference. Secondly, randomly add Gaussian noise with a standard deviation in the range of [0, 30] to each frame of the infrared video to generate noise data. The reason for randomly adding noise is as follows: on the one hand, it increases the difference between frames; on the other hand, it improves the robustness and generalization ability of the denoising network.
[0018] The second step: Construct input data pairs;
[0019] Inspired by the recently proposed famous Blind Spot Network (BSN) which can generate a good denoiser from a single noisy image, the present invention first utilizes the characteristics that there is strong similarity and correlation between adjacent frames in a video, and forms a data pair from the previous frame F i and the subsequent frame F i+1 of adjacent frames of a noisy infrared video. In this data pair, the previous frame F i serves as the input image of the denoising network, and the subsequent frame F i+1 serves as the target image of the denoising network, as shown in Figure 1 . For example, if the duration of an infrared video is 20 seconds and each second contains 15 frames of images, then the 300 frames of images in this video can be constructed into 299 groups of input data pairs: {(F1,F2),(F2,F3),(F3,F4),…,(F i ,F i+1 ),…,(F 299 ,F 300 ). In the present invention, the denoising network is constructed by n layers of common square convolutions. In order to make the infrared video denoising method proposed by the present invention efficient and fast, in the present invention, n = 3.
[0020] Step 3: Construct a symmetric loss function;
[0021] In order to more effectively optimize the training process and prompt the denoising network to efficiently utilize the input data, the present invention uses a symmetric loss to train the denoising network. Specifically, in the backpropagation process, not only the mean squared error (MSE) loss of "the previous frame F i as the input image and the subsequent frame F i+1 as the target image" is calculated, but also the MSE loss of "the subsequent frame F i+1 as the input image and the previous frame F i as the target image" is calculated, and it is considered that the two have the same importance. The symmetric loss function in the backpropagation process is calculated as follows:
[0022]
[0023] where θ is the network parameter of the denoising network f θ (·), which is updated as the symmetric loss Loss(·) is updated; is the mean squared error operation, which calculates the mean of the sum of the squares of the errors of each pixel point between the denoised image and the target image.
Claims
1. Generate noisy data: First, select an infrared video with clear content and no distortion interference. Second, randomly add Gaussian noise with a standard deviation in the range of [0, 30] to each frame of the infrared video to generate noisy data. The reason for randomly adding noise is that on the one hand, it increases the difference between frames, and on the other hand, it improves the robustness and generalization ability of the denoising network.
2. Construct input data pairs: First, the present invention utilizes the characteristics that there is strong similarity and correlation between adjacent frames in a video, and forms a group of data pairs from the previous frame F i and the subsequent frame F i+1 of adjacent frames of the noisy infrared video; in this data pair, the previous frame F i serves as the input image of the denoising network, and the subsequent frame F i+1 serves as the target image of the denoising network. For example, if the duration of an infrared video is 20 seconds and each second contains 15 frames of images, then the 300 frames of images in this video can be constructed into 299 groups of input data pairs: {(F1,F2),(F2,F3),(F3,F4),…,(F i ,F i+1 ),…,(F 299 ,F 300 )}; in the present invention, the denoising network is constructed by n layers of common square convolutions. In order to make the infrared video denoising method proposed by the present invention efficient and fast, n = 3 in the present invention.
3. Construct a symmetric loss function: To more effectively optimize the training process and prompt the denoising network to efficiently utilize the input data, the present invention uses a symmetric loss to train the denoising network; specifically, during the backpropagation process, not only calculate the mean squared error (MSE) loss of "the previous frame F i as the input image, and the subsequent frame F i+1 as the target image", but also calculate the MSE loss of "the subsequent frame F i+1 as the input image, and the previous frame F i as the target image", and consider the two to have the same importance; the symmetric loss function calculation during the backpropagation process is as follows: Among them, θ is the network parameter of the denoising network f θ (·), which is updated as the symmetric loss Loss(·) is updated; is the mean square error operation, which calculates the mean of the sum of the squares of the errors of each pixel point between the denoised image and the target image.