Two-frame wide dynamic short frame denoising system and method, model system and training method thereof

By training a two-frame wide dynamic range short-frame denoising model and utilizing noise modeling and luminance mapping techniques, the problems of high noise and color deviation in short-frame denoising are solved, thereby improving the quality of wide dynamic range image synthesis and motion region identification.

CN122289046APending Publication Date: 2026-06-26HEFEI JUNZHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI JUNZHENG TECH CO LTD
Filing Date
2024-12-25
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing wide dynamic range image synthesis methods, short frame denoising techniques suffer from problems such as high noise, color deviation, horizontal stripe noise, and false textures, which affect image quality and motion region identification.

Method used

By training a two-frame wide dynamic range short-frame denoising model, noise modeling methods are used to simulate the consistency of noise intensity between long and short frames. Brightness mapping and exposure ratio adjustment are employed to generate training samples and optimize the model, including downsampling, feature extraction, residual calculation and upsampling processes.

Benefits of technology

It improves the quality of wide dynamic range synthesized images, maintains consistent noise intensity between long and short frames, reduces false textures and detail blur, and enhances image realism and clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289046A_ABST
    Figure CN122289046A_ABST
Patent Text Reader

Abstract

This application discloses a two-frame wide dynamic range (WDR) short-frame denoising system and method, as well as its model system and training method. The method utilizes noise modeling to simulate long and short frames, which is more conducive to the use of the short-frame denoising model on real data. Furthermore, it does not require the collection of a large amount of real sample data. A machine learning model uses information from the long frame to guide the denoising of the short frame, maintaining consistency in noise intensity between the long and short frames, which helps improve the quality of the output image in subsequent WDR synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and in particular to a two-frame wide dynamic range short-frame denoising system and method, as well as its model system and training method. Background Technology

[0002] Wide Dynamic Range (WDR) is a technique used in image processing to improve image quality, especially in environments with vastly different lighting conditions. It allows an image to clearly display details in both bright and dark areas within the same scene. WDR is typically achieved by combining images with different exposure times (e.g., a short frame and a long frame). Due to hardware limitations such as bandwidth, memory, and computational load, two-frame WDR is the mainstream technique for real-time WDR. A short frame is the frame with the shorter exposure time among the two frames output by the image sensor at one time. In scenes with a large dynamic range, the shorter exposure time of a short frame, controlled by automatic exposure, more accurately presents details in bright areas. However, due to the short exposure time, the image noise is relatively higher. A long frame is the frame with the longer exposure time among the two frames output by the image sensor at one time.

[0003] Image synthesis with different exposure times can be affected by noise, especially short frames with shorter exposure times, which can impact the quality of the synthesized image. Common solutions include weighted fusion and short-frame denoising.

[0004] In the weighted fusion method, long frames are referenced for extremely dark areas with low signal-to-noise ratio, short frames are referenced for extremely bright areas with overexposure, and different weights are assigned to long and short frames for other areas. This method has the problem of reduced signal-to-noise ratio in overexposure and non-overexposure areas, and may affect the judgment of motion areas when the noise is high.

[0005] While short-frame denoising methods can avoid the problem of decreased signal-to-noise ratio in overexposed and unexposed areas and help in identifying moving areas, they still have some drawbacks: due to the short exposure time, the colors in short frames are prone to deviation, and some short frames acquired by sensors are prone to complex situations such as horizontal stripe noise; due to the low signal-to-noise ratio in short frames, traditional denoising methods are prone to producing problems such as pseudo-textures and blurred details. Summary of the Invention

[0006] To improve the quality of wide dynamic range synthesized images, this application provides a training method for a two-frame wide dynamic range short-frame denoising model and a short-frame denoising method.

[0007] In a first aspect, embodiments of this application provide a training method for a two-frame wide dynamic range short-frame denoising model. The training method includes: acquiring multiple pairs of denoised long and short frames, wherein each pair of long and short frames includes one short frame and one long frame. The method further includes, for each of the multiple pairs of long and short frames, performing the following operations: generating initial short frame noise based on the short frame in the pair, and adding the initial short frame noise to the short frame to obtain a noisy short frame; generating long frame noise based on the long frame in the pair, and adding the long frame noise to the long frame to obtain a noisy long frame; determining the noise coefficient corresponding to the pair of long and short frames; converting the initial short frame noise according to the noise coefficient to obtain target short frame noise with the same intensity as the long frame noise; adding the target short frame noise to the short frame to obtain a label image; determining the exposure ratio corresponding to the pair of long and short frames, and increasing the brightness of the noisy short frame according to the exposure ratio to obtain a noisy and brightened short frame; obtaining a training sample based on the noisy short frame, the noisy and brightened short frame, the noisy long frame, and the label image. The method further includes: training an initial model based on multiple training samples to obtain the short frame denoising model.

[0008] In some embodiments, the denoised short frame is obtained by: 1) acquiring multiple denoised images; 2) for each of the multiple denoised images, performing brightness mapping on the denoised image according to a preset brightness sampling strategy to obtain the denoised short frame.

[0009] In some embodiments, step 1) acquiring multiple denoised images includes: A) acquiring multiple original images collected in an environment where the illuminance exceeds a set illuminance threshold; B) normalizing the multiple original images to obtain multiple normalized images; and C) denoising the multiple normalized images to obtain the multiple denoised images.

[0010] In some embodiments, step 2) performing brightness mapping on the denoised image based on a preset brightness sampling strategy to obtain a denoised short frame includes: a) randomly selecting a brightness interval as a target brightness interval from the multiple brightness intervals according to the probability allocation of the preset multiple brightness intervals; b) performing uniform sampling within the target brightness interval to obtain a target average brightness value; c) performing brightness mapping on the denoised image according to the ratio of the target average brightness value to the average brightness value of the denoised image to obtain a denoised short frame.

[0011] In some embodiments, the longer frame in each pair of long and short frames is generated based on the shorter frame in that pair.

[0012] The longer frame in each pair of long and short frames is generated as follows: S1, random sampling is performed within a preset exposure ratio range to obtain a candidate exposure ratio; S2, the shorter frame in the pair of long and short frames is brightened according to the candidate exposure ratio to obtain a candidate longer frame; S3, it is determined whether the overexposure percentage of the candidate longer frame is lower than a preset ratio, and whether the candidate exposure ratio is the left endpoint of the exposure closed range; if the overexposure percentage of the candidate longer frame is not lower than the preset ratio, and the candidate exposure ratio is not the left endpoint of the exposure ratio range, then S1 to S3 are repeated; otherwise, S4 is executed; S4, the candidate longer frame is determined as the longer frame in the pair of long and short frames.

[0013] In some embodiments, for each of the plurality of pairs of long and short frames, the noise coefficient is determined as follows: the long frame in the pair of long and short frames is binarized to obtain a binarized image, such that: the region with a brightness of 0 in the binarized image corresponds to the region in the long frame with a brightness less than a set brightness threshold, and the region with a brightness of 1 in the binarized image corresponds to the region in the long frame with a brightness not less than the brightness threshold; the noise coefficient is determined according to the following formula:

[0014]

[0015] Where nr represents the noise figure, N s N represents the initial short frame noise. l The long frame noise is represented by er, and the exposure ratio is represented by M. op The binary image is represented by var, where var represents the variance.

[0016] In some embodiments, for each of the plurality of pairs of long and short frames, the noisy short frame, the noisy and brightened short frame, the noisy long frame, and the label image are determined according to the following calculation formula:

[0017] I ins =clip(I cs +N s +d bl ,0,1);

[0018] I inser =clip((I cs +N s +d bl )·er,0,er);

[0019] I inl =clip(I cl +N l ,0,1);

[0020] I odns =clip(Ics +N ss ,0,1).

[0021] Among them, I ins I represents the short frame after noise addition. inser I represents the short frame after noise addition and brightening. inl I represents the long frame after noise addition. odns Represents the label image; I cs N represents the shorter frame in the pair of long and short frames. s N represents the initial short frame noise. ss Indicates the target short frame noise; I cl N represents the longer frame in the pair of long and short frames. l The long frame noise is represented by d. bl This represents the short-frame black level perturbation obtained by uniform sampling within a preset interval; nr represents the noise coefficient, er represents the exposure ratio; clip(X,a,b) represents the truncation function, used to limit the brightness of each pixel in the input image X to a specified range [a,b].

[0022] In some embodiments, the short-frame denoising model includes a downsampling module, a first-scale feature extraction module, a second-scale feature extraction module, a residual calculation module, a deconvolution module, an upsampling module, and a post-processing module. The downsampling module is used to: for each pair of long and short frames, concatenate the noisy short frame, the noisy and brightened short frame, and the noisy long frame to obtain a first concatenation result; and perform a pixel unshuffle operation on the first concatenation result to obtain a downsampling result. The first-scale feature extraction module is used to: extract features from the downsampling result to obtain a first-scale feature map. The second-scale feature extraction module is used to: extract features from the first-scale feature map to obtain a second-scale feature map. The residual calculation module is used to: perform residual calculation based on the second-scale feature map to obtain a residual feature map. The deconvolution module is used to: perform a deconvolution operation on the residual feature map to obtain a deconvolution result. The upsampling module is used to: concatenate the first-scale feature map and the residual feature map to obtain a second concatenation result; perform a second convolution operation on the second concatenation result to obtain a second convolution result; and perform a pixel shuffle operation on the second convolution result to obtain predicted short frame noise. The post-processing module is used to: determine the predicted short frame denoising result based on the denoised short frame and the predicted short frame noise.

[0023] In some embodiments, the number of input channels corresponding to the Pixel Unshuffle operation is 12. The first-scale feature extraction module includes a first convolutional layer, a second convolutional layer, and a ReLU activation function layer connected in sequence. The kernel size of both the first and second convolutional layers is 3×3, the stride is 2, and the padding is 1. The second-scale feature extraction module includes a third convolutional layer, a fourth convolutional layer, and a ReLU activation function layer connected in sequence. The kernel size of the third convolutional layer is 3×3, the stride is 2, and the padding is 1. The kernel size of the fourth convolutional layer is 3×3, the stride is 1, and the padding is 1. The deconvolution module includes a deconvolutional layer and a ReLU activation function layer. The kernel size of the deconvolutional layer is 2×2, the stride is 2, and the padding is 0. The kernel size corresponding to the second convolution operation is 3x3, the stride is 1, and the padding is 1. The number of output channels corresponding to the Pixel Shuffle operation is 4.

[0024] Secondly, embodiments of this application provide a two-frame wide dynamic range short frame denoising method. The short frame denoising method includes: acquiring a target short frame and a target long frame; determining a target exposure ratio; increasing the brightness of the target short frame according to the target exposure ratio to obtain a brightened target short frame; and processing the target short frame, the brightened target short frame, and the target long frame using a short frame denoising model to obtain a denoised target short frame. The short frame denoising model is obtained through the training method of the two-frame wide dynamic range short frame denoising model as described in the first aspect.

[0025] Thirdly, embodiments of this application provide a training system for a two-frame wide dynamic range short-frame denoising model, including a sample generation module and a training module. The sample generation module is used to acquire multiple pairs of denoised long and short frames, wherein each pair of long and short frames includes one short frame and one long frame. The sample creation module is further configured to perform the following operations for each of the multiple pairs of long and short frames: generate initial short frame noise based on the short frame in the pair of long and short frames, and add the initial short frame noise to the short frame to obtain a noisy short frame; generate long frame noise based on the long frame in the pair of long and short frames, and add the long frame noise to the long frame to obtain a noisy long frame; determine the noise coefficient corresponding to the pair of long and short frames; convert the initial short frame noise according to the noise coefficient to obtain a target short frame noise with the same intensity as the long frame noise; add the target short frame noise to the short frame to obtain a label image; determine the exposure ratio corresponding to the pair of long and short frames, and increase the brightness of the noisy short frame according to the exposure ratio to obtain a noisy and brightened short frame; and obtain a training sample based on the noisy short frame, the noisy and brightened short frame, the noisy long frame, and the label image. The training module is configured to: train the initial model based on multiple training samples to obtain the short frame denoising model.

[0026] Fourthly, embodiments of this application provide a two-frame wide dynamic range short-frame denoising system, including an input module and a denoising module. The input module is used to: acquire a target short frame and a target long frame; and determine the exposure ratio corresponding to the pair of short and long frames, and increase the brightness of the target short frame according to the exposure ratio to obtain a brightened target short frame. The denoising module is used to: process the target short frame, the brightened target short frame, and the target long frame using a short-frame denoising model to obtain a denoised target short frame, wherein the short-frame denoising model is obtained through the training method of the two-frame wide dynamic range short-frame denoising model as described in the first aspect.

[0027] Fifthly, embodiments of this application provide a training device, including a processor and a memory. The memory stores a computer program, and when the processor executes the computer program, it implements the training method as described in any one of the first aspects.

[0028] In a sixth aspect, embodiments of this application provide an inference device, including a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the short frame denoising method as described in the second aspect.

[0029] The beneficial effects of the embodiments of this application include, but are not limited to:

[0030] (1) Using noise modeling to simulate long and short frames is more conducive to the use of short frame denoising models on real data. At the same time, it does not require the collection of a large amount of real sample data. The machine learning model uses long frame information to guide the denoising of short frames, keeping the noise intensity of long and short frames consistent, which helps to improve the quality of the output image of subsequent wide dynamic range synthesis.

[0031] (2) Brightness mapping can obtain more realistic short frame brightness, which is beneficial for the use of short frame denoising models on real data. Attached Figure Description

[0032] Figure 1 This is an exemplary flowchart of the training method for the two-frame wide dynamic range short frame denoising model provided in the embodiments of this application.

[0033] Figure 2 This is an exemplary flowchart of the process for generating clean short frames provided in the embodiments of this application.

[0034] Figure 3 This is an exemplary flowchart of the process for generating a clean long frame provided in the embodiments of this application.

[0035] Figure 4 This is an exemplary flowchart of the two-frame wide dynamic range short frame denoising method provided in the embodiments of this application.

[0036] Figure 5 This is a schematic diagram of the structure of an exemplary short frame denoising model provided in the embodiments of this application.

[0037] Figure 6 This is an exemplary block diagram of the training system for the two-frame wide dynamic range short frame denoising model provided in the embodiments of this application.

[0038] Figure 7 This is an exemplary block diagram of a two-frame wide dynamic range short-frame denoising system provided in the embodiments of this application.

[0039] Figure 8 This is a schematic diagram of the structure of an exemplary training / inference device provided in an embodiment of this application. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be described in detail below with reference to the accompanying drawings.

[0041] Figure 1 This is an exemplary flowchart of the training method for a two-frame wide dynamic range short-frame denoising model provided in this application embodiment. Flow 100 can be executed by a training device. Figure 1 As shown, the training process 100 includes the following steps.

[0042] Step 110: Obtain multiple pairs of long and short frames after denoising.

[0043] Each pair of long and short frames consists of one short frame and one long frame. For ease of description, the denoised image will be referred to as the clean image below. Similarly, the denoised short frame is also called the clean short frame, and the denoised long frame is also called the clean long frame. That is, each pair of (denoised) long and short frames consists of one clean short frame and one clean long frame.

[0044] In some embodiments, clean short frames can be simulated using luminance mapping. Specifically, the training device can acquire multiple denoised images (clean images). Then, for each of the multiple denoised images, the training device can perform luminance mapping on the denoised image according to a preset luminance sampling strategy to obtain a denoised short frame (i.e., a clean short frame). More details about luminance mapping can be found in [reference needed]. Figure 2 And its related descriptions.

[0045] In some embodiments, multiple raw images with low noise and high signal-to-noise ratio can be acquired in a well-lit environment to prepare clean images. Specifically, the training device can acquire multiple raw images in an environment where the illuminance exceeds a set illuminance threshold; for example, 2000 raw images can be acquired using an image sensor in an indoor / outdoor scene where the illuminance exceeds 10 lux. Next, the training device can normalize each of the raw images to obtain multiple normalized images. Finally, the training device can denoise each of the normalized images to obtain multiple denoised images (i.e., multiple clean images). For example, the process of acquiring clean images can be represented as follows:

[0046]

[0047] Among them, I o Let represent the original image, dn represent the traditional denoising algorithm BM3D, the input is a normalized image, and blc represent the black level of the image sensor. It should be understood that the arithmetic operations between an image and a single numerical value are a simplified representation of performing the same operation on each pixel in the image. For example, calculating I in equation (1) o -blc indicates that the original image I o The brightness of each pixel in the image is minus the black level (blc).

[0048] In some embodiments, the clean long frame in each pair of long and short frames is generated based on the clean short frame in that pair. For more details on the generation of clean long frames, please refer to [link / reference needed]. Figure 3 And its related descriptions.

[0049] Step 120: For each of the multiple pairs of long and short frames, execute steps 121 to 127.

[0050] Through step 120, the training device can generate multiple training samples.

[0051] Step 121: Generate initial short frame noise based on clean short frames, and add initial short frame noise to clean short frames to obtain noisy short frames.

[0052] Step 122: Generate long frame noise based on the clean long frame, and add long frame noise to the clean long frame to obtain the noisy long frame.

[0053] Existing noise modeling methods can be used to model noise in image sensors, and then noise can be randomly generated based on the noise generation parameters (i.e., model parameters), as shown below:

[0054] N = G(I) c ,p) (2).

[0055] Where N represents the generated noise, Ic G represents a clean image (clean short frame or clean long frame), G represents the noise generation function (i.e., noise model), and p represents the noise generation parameters.

[0056] Step 123: Determine the noise figure corresponding to the pair of long and short frames.

[0057] The noise figure reflects the intensity relationship between the initial short frame noise and the long frame noise.

[0058] In some embodiments, for each of the plurality of pairs of long and short frames, the training device may binarize the (clean) long frame in the pair to obtain a binarized image, such that: the region with a brightness of 0 in the binarized image corresponds to the region in the clean long frame with a brightness less than a set brightness threshold, and the region with a brightness of 1 in the binarized image corresponds to the region in the clean long frame with a brightness not less than the brightness threshold. The purpose of binarization is to remove the noise variance in overexposed regions and ensure that the noise intensity of the long and short frames in non-overexposed regions is consistent. For example, for the normalized image, the brightness threshold can be 0.9. Furthermore, the training device may determine the noise coefficient according to the following calculation formula:

[0059]

[0060] Where nr represents the noise figure, N s N represents the initial short frame noise. l This represents long frame noise, er represents the exposure ratio (see step 126), and M... op This represents a binary image, where var represents the variance (each pixel can be considered as a data point).

[0061] Step 124: Convert the initial short frame noise according to the noise coefficient to obtain the target short frame noise with the same intensity as the long frame noise.

[0062] The inventors discovered that real data often exhibits inconsistent noise levels between long and short frames. This inconsistency in noise levels at the boundary between long and short frame synthesis results in a poor-looking synthesized image. Therefore, the key to improving the quality of the synthesized image lies in how to create noisy short frame data (i.e., the label image in step 125) with noise levels consistent with those of the long frames.

[0063] Taking the calculation formula (3) as an example, the training device can perform noise conversion in the following manner:

[0064]

[0065] Where, N ss This represents the target short frame noise.

[0066] Step 125: Add target short frame noise to the clean short frame to obtain the label image.

[0067] Step 126: Determine the exposure ratio corresponding to the pair of long and short frames, and increase the brightness of the short frame after adding noise according to the exposure ratio to obtain the short frame after adding noise and brightening.

[0068] For composite images, the exposure ratio typically refers to the brightness ratio between the brightest and darkest parts of the same image. In a composite image, extremely dark areas are referenced to longer frames, and extremely bright areas are referenced to shorter frames; therefore, each pair of long and short frames corresponds to an exposure ratio.

[0069] Step 127: Based on the noisy short frame, the noisy and brightened short frame, the noisy long frame, and the label image, obtain a training sample.

[0070] Each training sample may include a short frame with noise, a short frame with noise and brightness, a long frame with noise, and a label image. Correspondingly, the model input includes a short frame with noise, a short frame with noise and brightness, and a long frame with noise.

[0071] In some embodiments, for each of the plurality of pairs of long and short frames, the portions of the training samples can be determined as follows:

[0072] I ins =clip(I cs +N s +d bl ,0,1) (5);

[0073] I inser =clip((I cs +N s +d bl )·er,0,er) (6);

[0074] I inl =clip(I cl +N l ,0,1) (7);

[0075] I odns =clip(I cs +N ss ,0,1) (8).

[0076] Among them, I ins I represents the short frame after noise addition. inser Indicates a short frame after adding noise and brightening, I inl I represents a long frame after noise has been added. odns Represents a label image; I cs N represents a clean, short frame. s N represents the initial short frame noise. ss Indicates target short frame noise; I cl N represents a clean long frame.l Indicates long frame noise; d bl This represents the black level perturbation of a short frame obtained by uniform sampling within a preset interval; nr represents the noise figure, and er represents the exposure ratio. Furthermore, clip(X,a,b) represents a truncation function used to limit the brightness of each pixel in the input image X to a specified range [a,b] to ensure the validity of the output image brightness values. For example, when the input image X is a normalized image, the range [a,b] is [0,1].

[0077] Step 130: Train the initial model based on multiple training samples to obtain a short-frame denoising model.

[0078] In some embodiments, the short-frame denoising model can be designed as a neural network, specifically, refer to Figure 5 And its related descriptions.

[0079] The training objective is to reduce the difference between the predicted denoised image and the labeled image (also known as the prediction error). Based on this, the loss function can include the prediction error of multiple training samples. For example, the Adam optimizer can be used to optimize the model parameters with an initial learning rate of 0.0001, 60 training epochs, and the learning rate decreasing by a factor of 0.5 every 10 training epochs. The loss function includes an L1 loss term.

[0080] Figure 2 This is an exemplary flowchart illustrating the process of generating clean short frames according to embodiments of this application. Figure 2 As shown, the clean short frame generation process 200 includes the following steps.

[0081] Step 210: Based on the probability allocation of multiple preset brightness intervals, randomly select one brightness interval from the multiple brightness intervals as the target brightness interval.

[0082] For example, five brightness intervals can be pre-divided as [0.001, 0.008], [0.008, 0.02], [0.02, 0.05], [0.05, 0.1], and [0.1, 0.3], with corresponding probabilities of 0.5, 0.3, 0.1, 0.05, and 0.05, respectively. The probability of any interval being selected is its corresponding probability; for example, the probability of selecting the interval [0.001, 0.008] is 0.5.

[0083] Step 220: Perform uniform sampling within the target brightness range to obtain the average target brightness value.

[0084] Step 230: Based on the ratio of the target brightness mean to the brightness mean of the clean image, perform brightness mapping on the clean image to obtain a clean short frame.

[0085] The training device can perform brightness mapping in the following manner:

[0086]

[0087] Among them, I cs Indicates a clean, short frame, I c Represents a clean image, mean(I) c ) represents a clean image I c The average brightness of the clean image I c The average brightness of all pixels in the image.

[0088] Figure 3 This is an exemplary flowchart of the process for generating a clean long frame according to an embodiment of this application. Figure 3 As shown, the clean long frame generation process 300 includes the following steps.

[0089] Step 310: Random sampling is performed within the preset exposure ratio range to obtain candidate exposure ratios.

[0090] As an example only, the exposure ratio range can be set to [2, 64]. The lower limit of 2 is to take into account that the long frame exposure in bright scenes is already relatively small. For example, the long frame exposure line (relative time unit) in bright scenes is 4, and the short frame exposure line is 2 (the minimum exposure line of the image sensor OS04A10 is 2). The upper limit of 64 is to take into account that when the exposure ratio exceeds 64 in scenes with a large dynamic range, the short frame signal-to-noise ratio is very low and many anomalies are likely to occur.

[0091] Step 320: Magnify the brightness of clean short frames according to the candidate exposure ratio to obtain candidate long frames.

[0092] For example, the training device can perform luminance amplification on clean short frames in the following manner:

[0093] I cl =clip(I cs ·er,0,1) (10).

[0094] Among them, I cl Indicates a candidate long frame, I cs This indicates a clean, short frame, and er indicates the candidate exposure ratio.

[0095] Step 330: Determine whether the overexposure percentage of the candidate long frame is lower than a preset percentage, and whether the candidate exposure ratio is the left endpoint of the exposure ratio range.

[0096] For a normalized image, pixels with a brightness of 1 can be defined as overexposed points. The corresponding percentages of overexposed points are shown below:

[0097]

[0098] oer represents the percentage of overexposure points, sum(I) cl ==1) indicates candidate long frame I cl The number of pixels with a brightness of 1 (i.e., overexposed pixels), where size represents the candidate long frame I. cl The total number of pixels.

[0099] For example, the preset ratio can be 0.5 or greater. This is just an example. Figure 3 The preset ratio is set to 0.5, and the left endpoint of the exposure ratio range is set to 2.

[0100] If the proportion of overexposed points in the candidate long frame is not less than the preset proportion, and the candidate exposure ratio is not the left endpoint of the exposure closed interval, then repeat steps 310 to 330. Otherwise, proceed to step 340.

[0101] Step 340: Determine the candidate long frames as clean long frames.

[0102] When a candidate exposure ratio is at the left endpoint of the exposure ratio interval (e.g., er = 2), and the overexposure percentage of the candidate long frame is still not lower than a preset ratio (e.g., oer ≥ 0.5), the candidate image generated according to the left endpoint of the exposure ratio interval (e.g., er = 2) is determined as a clean long frame. When a candidate exposure ratio is not at the left endpoint of the exposure ratio interval (e.g., er > 2), as long as the overexposure percentage of the candidate long frame is not lower than a preset ratio (e.g., oer ≥ 0.5), it indicates that the candidate exposure ratio is unreasonable. The candidate exposure ratio is then resampled and candidate long frames are generated until the overexposure percentage of the newly generated candidate long frame is lower than a preset ratio (e.g., oer < 0.5). The training device then determines the newly generated candidate long frame as a clean long frame. Of course, if the overexposure percentage of the initially generated candidate long frame is lower than a preset ratio (e.g., oer < 0.5), the candidate long frame is directly determined as a clean long frame.

[0103] Figure 4 This is an exemplary flowchart of a two-frame wide dynamic range short-frame denoising method provided in this application embodiment. Flow 400 can be executed by an inference device. Figure 4 As shown, the short frame denoising process 400 includes the following steps.

[0104] Step 410: Obtain the target short frame and the target long frame.

[0105] The target short frame and the target long frame are a pair of long and short frames used to synthesize a wide dynamic range image. Before image synthesis, the target short frame needs to be denoised.

[0106] Step 420: Determine the target exposure ratio, and increase the brightness of the target short frame according to the target exposure ratio to obtain the brightened target short frame.

[0107] The target exposure ratio can be obtained from the camera. For brightness amplification, please refer to the previously introduced formulas (6) and (10).

[0108] Step 430: Process the target short frame, the brightened target short frame, and the target long frame using the short frame denoising model to obtain the denoised target short frame. The short frame denoising model can be obtained through training process 100.

[0109] For more details on the short frame denoising model, please refer to the aforementioned embodiments.

[0110] It should be noted that the above description of the process is for illustrative purposes only and does not limit the scope of this application. Those skilled in the art can make various modifications and changes to the process under the guidance of this application.

[0111] Figure 5 This is a schematic diagram of the structure of an exemplary short frame denoising model provided in an embodiment of this application. The short frame denoising model 500 includes a downsampling module 510, a first feature extraction module 520, a second feature extraction module 530, a residual calculation module 540, a deconvolution module 550, an upsampling module 560, and a post-processing module 570.

[0112] The downsampling module 510 is used to: for each pair of multiple pairs of long and short frames, splice together a noisy short frame (denoted as I). ins ), and the short frame after adding noise and brightening (denoted as I) inser ) and the long frame after adding noise (denoted as I) inl The first stitching result is obtained; a pixel unshuffle operation is performed on the first stitching result to obtain a downsampling result. In some embodiments, the number of input channels corresponding to the pixel unshuffle operation is 12.

[0113] The first-scale feature extraction module 520 is used to extract features from the downsampling results to obtain a first-scale feature map. In some embodiments, the first-scale feature extraction module 520 includes a first convolutional layer, a second convolutional layer, and a ReLU activation function layer connected in sequence. The kernel size of both the first and second convolutional layers is 3×3, the stride is 2, and the padding is 1.

[0114] The second-scale feature extraction module 530 is used to extract features from the first-scale feature map to obtain a second-scale feature map. In some embodiments, the second-scale feature extraction module 530 includes a third convolutional layer, a fourth convolutional layer, and a ReLU activation function layer connected in sequence. The third convolutional layer has a kernel size of 3×3, a stride of 2, and padding of 1. The fourth convolutional layer has a kernel size of 3×3, a stride of 1, and padding of 1.

[0115] The residual calculation module 540 is used to: perform residual calculation based on the second-scale feature map to obtain the residual feature map.

[0116] The deconvolution module 550 is used to perform a deconvolution operation on the residual feature map to obtain a deconvolution result. In some embodiments, the deconvolution module includes a deconvolution layer and a ReLU activation function layer. The deconvolution layer has a kernel size of 2×2, a stride of 2, and padding of 0.

[0117] The upsampling module 560 is used to: concatenate the first-scale feature map and the residual feature map to obtain a second concatenation result; perform a second convolution operation on the second concatenation result to obtain a second convolution result; and perform a pixelShuffle operation on the second convolution result to obtain predicted short-frame noise. In some embodiments, the kernel size corresponding to the second convolution operation is 3x3, the stride is 1, and the padding is 1. The number of output channels corresponding to the pixel shuffle operation is 4.

[0118] Post-processing module 570 is used for: based on the noisy short frame (I ins ) and predicted short frame noise (denoted as N) ps ), determine the predicted short frame denoising result (denoted as I) ps The post-processing module 570 can calculate the predicted short frame denoising results as follows:

[0119] I ps =clip(I ins -N ps ,0,1) (12).

[0120] Figure 6 This is an exemplary block diagram of a training system for a two-frame wide dynamic range short-frame denoising model provided in this application embodiment. This system can be implemented as at least part of a training device through software, hardware, or a combination of both. Figure 6 As shown, the training system 600 includes a sample creation module 610 and a training module 620.

[0121] The sample creation module 610 is used to acquire multiple pairs of denoised long and short frames, wherein each pair of long and short frames includes one short frame and one long frame. The sample creation module 610 is further used to perform the following operations for each pair of long and short frames: generate initial short frame noise based on the short frame in the pair, and add the initial short frame noise to the short frame to obtain a denoised short frame; generate long frame noise based on the long frame in the pair, and add the long frame noise to the long frame to obtain a denoised long frame; determine the noise coefficient corresponding to the pair of long and short frames; convert the initial short frame noise according to the noise coefficient to obtain target short frame noise with the same intensity as the long frame noise; add the target short frame noise to the short frame to obtain a label image; determine the exposure ratio corresponding to the pair of long and short frames, and increase the brightness of the denoised short frame according to the exposure ratio to obtain a denoised and brightened short frame; and obtain a training sample based on the denoised short frame, the denoised and brightened short frame, the denoised long frame, and the label image.

[0122] The training module 620 is used to: train the initial model based on multiple training samples to obtain the short frame denoising model.

[0123] The training system and training method provided in this application are based on the same concept. For more details regarding the system and its modules, please refer to... Figure 1 The details and related descriptions will not be repeated here.

[0124] Figure 7 This is an exemplary block diagram of a two-frame wide dynamic range short-frame denoising system provided in an embodiment of this application. This system can be implemented as at least a part of an inference device through software, hardware, or a combination of both. Figure 7 As shown, the short frame denoising system 700 includes an input module 710 and a denoising module 720.

[0125] The input module 710 is used to: acquire a target short frame and a target long frame; and determine the exposure ratio corresponding to the pair of short and long frames, and increase the brightness of the target short frame according to the exposure ratio to obtain a brightened target short frame.

[0126] The denoising module 720 is used to process the target short frame, the brightened target short frame, and the target long frame using a short frame denoising model to obtain a denoised target short frame. The short frame denoising model can be obtained through training process 100.

[0127] The short-frame denoising system and method provided in this application belong to the same concept. For more details about the system and its modules, please refer to... Figure 4 The details and related descriptions will not be repeated here.

[0128] It should be noted that the above division of functional modules is only an example. For those skilled in the art, after understanding the system principle, they can arbitrarily combine, split, or replace the modules, as well as add or omit one or more modules, without violating the system principle.

[0129] This application also provides a training device. (See reference...) Figure 8 The training device 800 includes a processor 810 and a memory 820. The memory 820 stores a computer program for model training. When the processor 820 executes the computer program, it implements the training method provided in this embodiment. More details about the training method and its steps can be found in... Figure 1 The relevant descriptions can be found here, so I will not repeat them here.

[0130] This application also provides an inference device. (See reference...) Figure 8 The inference device 800 includes a processor 810 and a memory 820. The memory 820 stores a computer program for short-frame denoising. When the processor 820 executes the computer program, it implements the short-frame denoising method provided in this embodiment. More details about the short-frame denoising method and its steps can be found in... Figure 4 The relevant descriptions can be found here, so I will not repeat them here.

[0131] It should be noted that the training device and the inference device can be the same device or different devices.

[0132] Unless otherwise defined, the terms used herein should have the ordinary meaning as understood by those skilled in the art. Unless otherwise specified, the terms "first," "second," "third," and similar words used herein do not indicate any order, quantity, or importance, but are merely used to distinguish different entities. The terms "comprising," "including," or any variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements may include not only the elements expressly listed, but also other elements not expressly listed, such as elements inherent to such process, method, product, or apparatus.

[0133] Finally, it should be noted that the above embodiments are only some embodiments of this application and not all embodiments; in the various embodiments of this application, unless otherwise specified or logically conflicting, the terms and / or descriptions between different embodiments are consistent and can be referenced by each other, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship; although the technical solutions of this application have been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features, and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the various embodiments of this application.

Claims

1. A training method for a two-frame wide dynamic range short-frame denoising model, characterized in that, include: Step 110: Obtain multiple pairs of denoised long and short frames, wherein each pair of long and short frames includes one short frame and one long frame; wherein the denoised short frame is obtained in the following way: 1) Acquire multiple denoised images; including: A) Acquire multiple raw images in an environment where the illuminance exceeds a set illuminance threshold; B) Normalize each of the original images to obtain multiple normalized images; C) Denoise the multiple normalized images respectively to obtain the multiple denoised images; 2) For each of the multiple denoised images, perform luminance mapping on the denoised image according to a preset luminance sampling strategy to obtain a denoised short frame; including: a) Based on the probability allocation of multiple preset brightness intervals, randomly select one brightness interval from the multiple brightness intervals as the target brightness interval; b) Perform uniform sampling within the target brightness range to obtain the average target brightness value; c) Based on the ratio of the target brightness mean to the brightness mean of the denoised image, perform brightness mapping on the denoised image to obtain the denoised short frame. Step 120: For each of the multiple pairs of long and short frames, perform the following steps 121-127: Step 121: Generate initial short frame noise based on the short frame in the pair of long and short frames, and add the initial short frame noise to the short frame to obtain a noisy short frame; Step 122: Generate long frame noise based on the long frame in the pair of long and short frames, and add the long frame noise to the long frame to obtain a noisy long frame; Step 123: Determine the noise figure corresponding to the pair of long and short frames; Step 124: Convert the initial short frame noise according to the noise coefficient to obtain a target short frame noise with the same intensity as the long frame noise; Step 125: Add the target short frame noise to the short frame to obtain the label image; Step 126: Determine the exposure ratio corresponding to the pair of long and short frames, and increase the brightness of the noise-added short frame according to the exposure ratio to obtain the noise-added and brightened short frame. Step 127: Based on the noisy short frame, the noisy and brightened short frame, the noisy long frame, and the label image, obtain a training sample; Step 130: Train the initial model based on multiple training samples to obtain the short frame denoising model.

2. The training method for the two-frame wide dynamic range short-frame denoising model as described in claim 1, characterized in that, The longer frame in each pair of long and short frames is generated based on the shorter frame in that pair.

3. The training method for the two-frame wide dynamic range short-frame denoising model as described in claim 2, characterized in that, The longer frame in each pair of long and short frames is generated in the following way: S1, Random sampling is performed within the preset exposure ratio range to obtain candidate exposure ratios; S2, according to the candidate exposure ratio, the brightness of the short frame in the pair of long and short frames is increased to obtain the candidate long frame; S3, determine whether the overexposure point ratio of the candidate long frame is lower than a preset ratio, and whether the candidate exposure ratio is the left endpoint of the exposure closed interval; If the overexposure percentage of the candidate long frame is not lower than the preset percentage, and the candidate exposure ratio is not the left endpoint of the exposure ratio range, then S1 to S3 are executed repeatedly. Otherwise, execute S4; S4, the candidate long frame is determined as the long frame in the pair of long and short frames.

4. The training method for the two-frame wide dynamic range short-frame denoising model as described in claim 1, characterized in that, For each of the plurality of pairs of long and short frames, the noise figure is determined as follows: The long frame in the pair of long and short frames is binarized to obtain a binarized image, such that: the region with a brightness of 0 in the binarized image corresponds to the region in the long frame with a brightness less than a set brightness threshold, and the region with a brightness of 1 in the binarized image corresponds to the region in the long frame with a brightness not less than the brightness threshold. The noise figure is determined using the following formula: Where nr represents the noise figure, N s N represents the initial short frame noise. l The long frame noise is represented by er, and the exposure ratio is represented by M. op The binary image is represented by var, where var represents the variance.

5. The training method for the two-frame wide dynamic range short-frame denoising model as described in claim 1, characterized in that, For each of the multiple pairs of long and short frames, the noisy short frame, the noisy and brightened short frame, the noisy long frame, and the label image are determined according to the following calculation formula: I ins =clip(I cs +N s +d bl ,0,1); IN inser =clip((I cs +N s +d bl )·is,0,is); I inl =clip(I cl +N l ,0,1); I odns =clip(I cs +N ss ,0,1); Among them, I ins I represents the short frame after noise addition. inser I represents the short frame after noise addition and brightening. inl I represents the long frame after noise addition. odns Represents the label image; I cs N represents the shorter frame in the pair of long and short frames. s N represents the initial short frame noise. ss Indicates the target short frame noise; I cl N represents the longer frame in the pair of long and short frames. l The long frame noise is represented by d. bl This represents the short-frame black level perturbation obtained by uniform sampling within a preset interval; nr represents the noise coefficient, er represents the exposure ratio; clip(X,a,b) represents the truncation function, used to limit the brightness of each pixel in the input image X to a specified range [a,b].

6. The training method for the two-frame wide dynamic range short-frame denoising model as described in claim 1, characterized in that, The short frame denoising model includes a downsampling module, a first-scale feature extraction module, a second-scale feature extraction module, a residual calculation module, a deconvolution module, an upsampling module, and a post-processing module. The downsampling module is used to: for each of the multiple pairs of long and short frames, splice the noisy short frame, the noisy and brightened short frame, and the noisy long frame to obtain a first splicing result; and perform a Pixel Unshuffle operation on the first splicing result to obtain a downsampling result. The first-scale feature extraction module is used to: extract features from the downsampling results to obtain a first-scale feature map; The second-scale feature extraction module is used to: extract features from the first-scale feature map to obtain a second-scale feature map; The residual calculation module is used to: perform residual calculation based on the second scale feature map to obtain a residual feature map; The deconvolution module is used to: perform a deconvolution operation on the residual feature map to obtain a deconvolution result; the upsampling module is used to: concatenate the first scale feature map and the residual feature map to obtain a second concatenation result; perform a second convolution operation on the second concatenation result to obtain a second convolution result; and perform a PixelShuffle operation on the second convolution result to obtain predicted short frame noise; The post-processing module is used to: determine the denoising result of the predicted short frame based on the noise-added short frame and the noise of the predicted short frame.

7. The training method for the two-frame wide dynamic range short-frame denoising model as described in claim 6, characterized in that, The number of input channels corresponding to the PixelUnshuffle operation is 12; The first scale feature extraction module includes a first convolutional layer, a second convolutional layer, and a ReLU activation function layer connected in sequence; the kernel size of the first convolutional layer and the second convolutional layer is 3×3, the stride is 2, and the padding is 1. The second-scale feature extraction module includes a third convolutional layer, a fourth convolutional layer, and a ReLU activation function layer connected in sequence; the third convolutional layer has a kernel size of 3×3, a stride of 2, and padding of 1; the fourth convolutional layer has a kernel size of 3×3, a stride of 1, and padding of 1. The deconvolution module includes a deconvolution layer and a ReLU activation function layer; the kernel size of the deconvolution layer is 2×2, the stride is 2, and the padding is 0. The second convolution operation corresponds to a kernel size of 3x3, a stride of 1, and padding of 1; the pixel shuffle operation corresponds to an output channel number of 4.

8. A two-frame wide dynamic range short-frame noise reduction method, characterized in that, include: Acquire the target short frame and the target long frame; Determine the target exposure ratio, and then increase the brightness of the target short frame according to the target exposure ratio to obtain the brightened target short frame. The target short frame, the brightened target short frame, and the target long frame are processed using a short frame denoising model to obtain a denoised target short frame. The short frame denoising model is obtained by training a two-frame wide dynamic range short frame denoising model as described in any one of claims 1 to 7.

9. A training system for a two-frame wide dynamic range short-frame denoising model, characterized in that, The system is applicable to any of the methods described in claims 1-8 above, and the system includes a sample creation module and a training module; The sample production module is used to obtain multiple pairs of long and short frames after denoising, wherein each pair of long and short frames includes one short frame and one long frame. The sample production module is further configured to perform the following operations for each of the multiple pairs of long and short frames: generate initial short frame noise based on the short frame in the pair of long and short frames, and add the initial short frame noise to the short frame to obtain a noisy short frame; generate long frame noise based on the long frame in the pair of long and short frames, and add the long frame noise to the long frame to obtain a noisy long frame. Determine the noise coefficients corresponding to the pair of long and short frames; convert the initial short frame noise according to the noise coefficients to obtain target short frame noise with the same intensity as the long frame noise; add the target short frame noise to the short frame to obtain the label image; Determine the exposure ratio corresponding to the pair of long and short frames, and then increase the brightness of the noise-added short frame according to the exposure ratio to obtain the noise-added and brightened short frame. A training sample is obtained based on the noisy short frame, the noisy and brightened short frame, the noisy long frame, and the label image; The training module is used to: train the initial model based on multiple training samples to obtain the short frame denoising model.

10. A two-frame wide dynamic range short-frame denoising system, characterized in that, Includes an input module and a noise reduction module; The input module is used to: acquire a target short frame and a target long frame; and determine the exposure ratio corresponding to the pair of short and long frames, and increase the brightness of the target short frame according to the exposure ratio to obtain the brightened target short frame. The denoising module is used to: process the target short frame, the brightened target short frame, and the target long frame using a short frame denoising model to obtain a denoised target short frame. The short frame denoising model is obtained by training the two-frame wide dynamic range short frame denoising model as described in any one of claims 1 to 8.