Spatial domain denoising method adaptive to time domain denoising under extremely dark light

The neural network predicts noise and combines the time-domain denoising and divides the region into noise data, which solves the problem of uneven noise removal in extremely dark light scenes, and achieves more uniform airspace denoising and higher image quality.

CN120013791APending Publication Date: 2025-05-16HEFEI JUNZHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311519475.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-05-16

Smart Images

  • Figure CN120013791A_ABST
    Figure CN120013791A_ABST
Patent Text Reader

Abstract

The invention provides a space domain denoising method adaptive to time domain denoising under extremely dark light. The method comprises the following steps: S1, acquiring noise modeling data; s2, calibrating noise parameters; s3, preparing clean data; s4, a random sampling strategy is involved; s5, static region noise generation; s6, generating dynamic region noise; and S7, generating pairing data. According to an existing noise modeling method, in combination with time domain denoising, the static region and the dynamic region are divided, small noise is generated for the static region, large noise is generated for the dynamic region, real noise data obtained after time domain denoising are simulated, time domain denoising is better adapted, and the situation that space domain denoising is not uniform or noise remains is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent monitoring video processing, and in particular relates to a spatial domain denoising method for adaptive temporal domain denoising under extremely dark light. Background Art

[0002] In the prior art, image sensors face serious noise problems due to the limited number of photons in extremely dark light scenes. Usually, only time domain denoising or only spatial domain denoising cannot effectively solve the noise problem. The usual solution is to combine time domain denoising with spatial domain denoising. Time domain denoising can effectively reduce the noise in static areas, but the noise in dynamic areas is difficult to eliminate. In addition to solving the residual noise in static areas, spatial domain denoising is more important to solve the noise problem in dynamic areas.

[0003] At present, traditional spatial domain denoising methods include various filtering based on pixel features such as Gaussian filtering, bilateral filtering, non-local mean filtering, etc., methods based on transform domain such as wavelet shrinkage method, and methods combining features and transforms such as BM3d algorithm.

[0004] At present, the spatial domain denoising algorithms based on neural networks use paired data to directly predict noise, or model the noise formation process, including photon shot noise, line noise, readout noise, etc., and then use neural networks to predict noise.

[0005] However, the existing traditional spatial denoising methods are relatively ineffective in extremely dark light scenes and are prone to produce pseudo textures and residual noise; the neural network-based methods are relatively effective, but the use of paired data for training requires a sufficient amount of paired data, which is difficult to obtain in many cases; the noise modeling method can simulate the generation of a large amount of paired data, but the existing method denoises the entire image, and cannot effectively model the situation where there is less noise in the static area and more noise in the dynamic area after time domain denoising, ultimately resulting in uneven denoising or residual noise.

[0006] In addition, the commonly used terms in the prior art include:

[0007] Image noise: Image noise mainly refers to the rough part of the image produced in the process of CCD or CMOS receiving light as a signal and outputting it. It also refers to foreign pixels that should not appear in the image, usually caused by electronic interference.

[0008] Temporal denoising: Denoising is performed by combining the previous and next frame images in the temporal domain.

[0009] Spatial denoising: denoising a single frame image in the spatial domain. Summary of the invention

[0010] In order to solve the above problems, the purpose of this application is to use a neural network to predict noise based on the existing noise modeling method, taking into account the different noise intensities in static and dynamic areas, so that the denoising in the spatial domain is uniform and the residual noise is small.

[0011] Specifically, the present invention provides a spatial domain denoising method for adapting temporal domain denoising under extremely dim light, the method comprising the following steps:

[0012] S1, collect the data required for noise modeling, including black frame data and flat frame data:

[0013] S1.1, collect black frame data, use black tape to completely cover the lens and place it in a completely dark room without light, and collect 100 frames of raw images at ISO of 100, 200, 400, 800, 1600, and 3200 respectively;

[0014] S1.2, collect flat frame data, fix the image sensor and point the lens at uniform white paper under uniform ambient light, respectively at ISO 100, 200, 400, 800, 1600, 3200, take 1 as the starting point of exposure line, take the exposure line with no overexposure, i.e. the maximum value of 12-bit raw image data does not exceed 4095 as the end point, sample 20 groups of data at average intervals, each group of data includes 2 frames of raw images;

[0015] S2, calibration noise modeling parameters:

[0016] S2.1, calibrate the system gain k parameter, select step S1.2 to collect each group of flat frames 256x256 area and record it as I f1 , I f2 , calculate the mean m and variance v using formula (1), where Mean represents the mean and Var represents the variance.

[0017] m=Mean((I f1 +I f2 ) / 2)

[0018] v=Var(I f1 -I f2 ) Formula (1)

[0019] For each iso, 20 sets of data are respectively subjected to formula (1) to obtain 20 sets of parameters, and linear regression is performed with m and v as the horizontal axis and vertical axis respectively, as shown in formula (2), where k is the system gain under the current iso;

[0020] y=kx+b Formula (2)

[0021] S2.2, calibrated read noise N γeadParameters, select step S1.1 to collect black frames under the same iso, subtract the black level and the line noise that obeys the standard normal distribution, and then use the Tukeylambda distribution to fit to get tlshape and σ t1 , each iso can get the corresponding tlshape and σ t1 , with each iso corresponding to log(k), log(σ t1 ) is the horizontal axis and the vertical axis for linear regression to obtain the slope a tl and the intercept b tl ;

[0022] S2.3, determine the sampling range of system gain k, the maximum system gain k max is the maximum and minimum gain k of each iso system gain min Set to 0.05; determine the tlshape sampling range, maximum tlshape max Set to 0.1, minimum tlshape min is the minimum value obtained by each iso calibration;

[0023] S3 is prepared to contain 2,500 clean raw data of different scenes;

[0024] S4, design random sampling strategy:

[0025] S4.1, the system gain sampling is as shown in formula (2), where log(k s ) is the sampling system gain parameter, U is uniform distribution, k min and k max are the minimum and maximum values ​​of the system gain k, respectively.

[0026] log(k s )~U(log(k min ), log( xmax )) Formula (2)

[0027] S4.2, Tukeylambda distribution shape parameter sampling is as shown in formula (3), where tlshape s is the shape of the sampled Tukeylambda distribution, Y is a uniform distribution, k min and k max The minimum and maximum system gain,

[0028] tlshape~U(tlshape min , tlshape max ) Formula (3)

[0029] S4.3, Tukeylambda distribution standard deviation parameter sampling is as shown in formula (4), where N is the normal distribution sampling, σ tlis the standard deviation parameter of the normal distribution, σ tls is the standard deviation of the sampled Tukeylambda distribution, a tl and b tl The parameters obtained by calibration in step S2.2, log(k s ) is the system gain parameter sampled in step S4.1,

[0030] log(σ tls )~N(a tl ·log(k s )+b tl , σ tl ) Formula (4);

[0031] S5, static region noise generation:

[0032] S5.1, each time noise is generated, a multi-frame weighted average frame number MFM is randomly generated between 1 and 200 with uniform distribution, and the system gain log(k s ), generate Tukeylambda distribution shape tlshape by random sampling according to formula (3) s , generate Tukeylambda distribution with standard deviation σ by random sampling according to formula (4) tls ;

[0033] S5.2, static region Poisson distribution noise N ps It is generated as follows (5), where P is the Poisson distribution and I is the static area of ​​the complete image;

[0034]

[0035] S5.3, read noise N γs It is generated as follows (6), where TL is the Tukeylambda distribution;

[0036]

[0037] S6, dynamic area noise generation:

[0038] S6.1, randomly generate a dynamic area mask, generate a ratio of the full image length to the dynamic area length between 1.5 and 3 with uniform distribution, generate a ratio of the full image width to the dynamic area width between 1.5 and 3 with uniform distribution, generate a dynamic area in the full image, randomly generate a center point in the dynamic area, the center point is within 10x10 in the middle of the dynamic area, and ensure that the center point is generated within 10x10 pixels in the center of the dynamic area, split the dynamic area into 4 sub-dynamic areas according to the coordinates of the center point, specifically, determine the 4 sub-dynamic areas with the center coordinates as the vertex and the 4 vertices of the dynamic area as another vertex;

[0039] S6.2, for the sub-dynamic region, regenerate the 3-sample system gain log(k sm ), plus the maximum system gain log(k max ) There are 4 different system gains in total, and the 4 different system gains are randomly assigned to 4 sub-dynamic areas;

[0040] log(k sm )~U(log(k s ), log(k max )) Formula (7)

[0041] S6.3 Similarly, step S5 generates the Poisson distribution noise N of the four sub-dynamic regions according to equations (5) and (6). pm and read noise N γm ;

[0042] S7, generate paired data, add the noise generated by the static area and the dynamic area to the clean data according to formula (8), where I c For clean data, I N For the final generated noise data, form paired data,

[0043] I N =I C +(1-mask)·(N ps +N γs )+mask·(N pm +N γm ) formula (8).

[0044] The method also includes step S8: training the neural network model, generating paired data online according to step S7 during the training, using Adam as the optimizer, with a learning rate of 0.0001, a training cycle of 120, a learning rate reduced by 0.1 times every 20 cycles, and a training loss of L1loss.

[0045] The method uses the imx327 image sensor.

[0046] Therefore, the advantages of this application are: according to the existing noise modeling method, combined with time domain denoising to divide the static area and the dynamic area, generate less noise for the static area, generate more noise for the dynamic area, simulate the noise data after real time domain denoising, better adapt to time domain denoising, and effectively reduce the situation of uneven spatial denoising or residual noise. This method is applicable to Beijing Junzheng's chip T41 AIISP, the field is image signal processing, improve the noise level in extremely dark light scenes, and improve image quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention.

[0048] Figure 1 This is a schematic diagram of the application process.

[0049] Figure 2 It is a flowchart of the present application method including the training model step. DETAILED DESCRIPTION

[0050] In order to more clearly understand the technical content and advantages of the present invention, the present invention is now further described in detail in conjunction with the accompanying drawings.

[0051] like Figure 1 As shown, this application proposes a spatial denoising method for adapting temporal denoising under extremely dark light, and the image sensor used is imx327, wherein the main implementation steps of the method are as follows:

[0052] Step S1 collects the data required for noise modeling, including black frame data and flat frame data:

[0053] S1.1 collects black frame data. Use black tape to completely cover the lens and place it in a completely dark room. Collect 100 frames of raw images at ISO 100, 200, 400, 800, 1600, and 3200 respectively.

[0054] S1.2, collect flat frame data, fix the image sensor and keep the lens facing a uniform white paper under uniform ambient light, respectively at ISO 100, 200, 400, 800, 1600, 3200, take 1 as the starting point of the exposure line, take the exposure line with no overexposure, i.e. the maximum value of the 12-bit raw image data does not exceed 4095 as the end point, sample 20 groups of data at an average interval, each group of data includes 2 frames of raw images; the overexposure, i.e. the 12-bit data range is 0-4095, overexposure means that the data is saturated to 4095, and the collected data is not necessarily exactly underexposed. Generally, the human eye is used to preliminarily observe whether there is overexposure, and if there is overexposed data, the overexposed data is discarded;

[0055] Step S2 calibrates the noise modeling parameters:

[0056] S2.1, calibrate the system gain k parameter, select step S1.2 to collect each group of flat frames 256x256 area and record it as I f1 , I f2 , calculate the mean m and variance v using formula (1), where Mean represents the mean and Var represents the variance.

[0057] m=Mean((I f1 +I f2 ) / 2)

[0058] v=Var(I f1 -I f2 ) Formula (1)

[0059] For each iso, 20 sets of data are respectively subjected to formula (1) to obtain 20 sets of parameters, and linear regression is performed with m and v as the horizontal axis and vertical axis respectively, as shown in formula (2), where k is the system gain under the current iso;

[0060] y=kx+b Formula (2)

[0061] S2.2, calibrated read noise N γead Parameters, select step S1.1 to collect black frames under the same iso, subtract the black level and the line noise that obeys the standard normal distribution, and then use the Tukeylambda distribution to fit to get tlshape and σ tl , each iso can get the corresponding tlshape and σ tl , with each iso corresponding to log(k), log(σ tl ) is the horizontal axis and the vertical axis for linear regression to obtain the slope a tl and the intercept b tl ; The Tukeylambda distribution has shape parameters, standard deviation parameters, etc., and formula (3) only represents the shape parameters;

[0062] S2.1 obtains the system gain k for each iso, and S2.2 obtains the Tukey lambda related parameter σ for each iso tl ;

[0063] S2.3 Determine the sampling range of system gain k and the maximum system gain k max is the maximum and minimum gain k of each iso system gain min Set to 0.05; determine the tlshape sampling range, maximum tlshape max Set to 0.1, minimum tlshape min is the minimum value obtained by each iso calibration;

[0064] Step S3 prepares clean raw data containing 2500 images of different scenes;

[0065] Step S4, design random sampling strategy:

[0066] S4.1, the system gain sampling is as shown in formula (2), where log(k s ) is the sampling system gain parameter, U is uniform distribution, k min and k max are the minimum and maximum values ​​of the system gain k, respectively.

[0067] log(k s )~U(log(k min ), log(k max )) Formula (2)

[0068] Among them, ~ indicates that it conforms to a certain distribution;

[0069] S4.2, Tukeylambda distribution shape parameter sampling is as shown in formula (3), where tlshape s is the shape of the sampled Tukeylambda distribution, U is the uniform distribution, kx in and k max The minimum and maximum system gain,

[0070] tlshape~U(tlshape min , tlshape max ) Formula (3)

[0071] S4.3, Tukeylambda distribution standard deviation parameter sampling is as shown in formula (4), where N is the normal distribution sampling, σ tl is the standard deviation parameter of the normal distribution, σ tls is the standard deviation of the sampled Tukeylambda distribution, a tl and b tl The parameters obtained by calibration in step S2.2, log(k s ) is the system gain parameter sampled in step S4.1,

[0072] log(σ tls )~N(a tl ·log(k s )+b tl , σ tl ) Formula (4);

[0073] Step S5, static area noise generation:

[0074] S5.1, each time noise is generated, a multi-frame weighted average frame number MFM is randomly generated between 1 and 200 with uniform distribution, and the system gain log(k s ), generate Tukeylambda distribution shape tlshape by random sampling according to formula (3) s , generate Tukeylambda distribution with standard deviation σ by random sampling according to formula (4) tls ;

[0075] S5.2 Static region Poisson distribution noise N ps It is generated as follows (5), where P is the Poisson distribution and I is the static area of ​​the complete image;

[0076]

[0077] S5.3 Readout noise N γs It is generated as follows (6), where TL is the Tukeylambda distribution;

[0078]

[0079] Step S6: Dynamic area noise generation:

[0080] S6.1, randomly generate a dynamic area mask, generate a ratio of the length of the complete image to the length of the dynamic area with a uniform distribution between 1.5 and 3, and generate a ratio of the width of the complete image to the width of the dynamic area with a uniform distribution between 1.5 and 3. Here, the width and height of the dynamic area are mainly generated. The ratio is the width of the complete image to the width of the dynamic area. The height is similar. For example, if the width of the complete image is 1080 and the ratio is 2, the width of the dynamic area is 540. Generate a dynamic area in the complete image, randomly generate a center point in the dynamic area, and the center point is within 10x10 in the middle of the dynamic area. At the same time, ensure that the center point is generated within 10x10 pixels in the center of the dynamic area. Split the dynamic area into 4 sub-dynamic areas according to the coordinates of the center point. Specifically, the 4 sub-dynamic areas are determined with the center coordinates as the vertex and the 4 vertices of the dynamic area as the other vertex;

[0081] S6.2, for the sub-dynamic region, regenerate the 3-sample system gain log(k sm ), plus the maximum system gain log(k max ) There are 4 different system gains in total, and the 4 different system gains are randomly assigned to 4 sub-dynamic areas; 3 are generated because there is a maximum system gain behind, and there are 4 in total, and it is necessary to ensure that at least one maximum system gain can be covered;

[0082] log(k sm )~U(log(k s ), log(k max )) Formula (7)

[0083] S6.3 Similarly, step S5 generates the Poisson distribution noise N of the four sub-dynamic regions according to equations (5) and (6). pm and read noise N γm The noise in the dynamic area is larger than that in the static area. The noise size is mainly controlled by the system gain. S6.1 and S6.2 simulate the dynamic area position and the noise size in the dynamic area respectively. The noise generation method is the same.

[0084] Step S7 generates paired data, and adds the noise generated by the static area and the dynamic area to the clean data according to formula (8), where I c For clean data, IN The noise data is finally generated to form paired data.

[0085] I N =I C +(1-mask)·(N ps +N γs )mask·(N pm +N γm ) Formula (8)

[0086] like Figure 2 As shown, the method may further include step S8, training the neural network model, generating paired data online according to step S7 during training, using Adam as the optimizer, with a learning rate of 0.0001, a training cycle of 120, a learning rate reduced by 0.1 times every 20 cycles, and a training loss of L1loss.

[0087] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A spatial denoising method for adapting temporal denoising under extremely dim light, characterized in that: The method comprises the following steps: S1, collect the data required for noise modeling, including black frame data and flat frame data: S1.1, collect black frame data, use black tape to completely cover the lens and place it in a completely dark room without light, and collect 100 frames of raw images at ISO of 100, 200, 400, 800, 1600, and 3200 respectively; S1.2, collect flat frame data, fix the image sensor and point the lens at uniform white paper under uniform ambient light, respectively at ISO 100, 200, 400, 800, 1600, 3200, take 1 as the starting point of exposure line, take the exposure line with no overexposure, i.e. the maximum value of 12-bit raw image data does not exceed 4095 as the end point, sample 20 groups of data at average intervals, each group of data includes 2 frames of raw images; S2, calibration noise modeling parameters: S2.1, calibrate the system gain k parameter, select step S1.2 to collect each group of flat frames 256x256 area and record it as I f1 , I f2 , calculate the mean m and variance v using formula (1), where Mean represents the mean and Var represents the variance. m=Mean((I f1 +I f2 ) / 2) v = Var(I f1 - I f2 ) Equation (1) For each iso, 20 sets of data are respectively subjected to formula (1) to obtain 20 sets of parameters, and linear regression is performed with m and v as the horizontal axis and vertical axis respectively, as shown in formula (2), where k is the system gain under the current iso; y=kx+b Formula (2) S2.2, calibrated read noise N read Parameters, select step S1.1 to collect black frames under the same iso, subtract the black level and the line noise that obeys the standard normal distribution, and then use the Tukeylambda distribution to fit to get tlshape and σ t1 , each iso can get the corresponding tlshape and σ t1 , with each iso corresponding to log(k), log(σ t1 ) is the horizontal axis and the vertical axis for linear regression to obtain the slope a tl and the intercept b tl ; S2.3, determine the sampling range of system gain k, the maximum system gain k max is the maximum and minimum gain k of each iso system gain min Set to 0.05; determine the tlshape sampling range, maximum tlshape max Set to 0.1, minimum tlshape min is the minimum value obtained by each iso calibration; S3 is prepared to contain 2,500 clean raw data of different scenes; S4, design random sampling strategy: S4.1, the system gain sampling is as shown in formula (2), where log(k s ) is the sampling system gain parameter, U is uniform distribution, k min and k max are the minimum and maximum values ​​of the system gain k, respectively. log(k s )~U(log(k min ), log(k max )) Equation (2) S4.2, Tukeylambda distribution shape parameter sampling is as shown in formula (3), where tlshape s is the shape of the sampled Tukeylambda distribution, U is the uniform distribution, k min and k max The minimum and maximum system gain, tlshape~U(tlshape min , tlshape max ) Formula (3) S4.3, Tukeylambda distribution standard deviation parameter sampling is as shown in formula (4), where N is the normal distribution sampling, σ tl is the standard deviation parameter of the normal distribution, σ tls is the standard deviation of the sampled Tukeylambda distribution, a tl and b tl The parameters obtained by calibration in step S2.2, log(k s ) is the system gain parameter sampled in step S4.1, log(σ tls )~N(a tl ·log(k s )+b tl ,s tl ) expression(4); S5, static region noise generation: S5.1, each time noise is generated, a multi-frame weighted average frame number MFM is randomly generated between 1 and 200 with uniform distribution, and the system gain log(k s ), generate Tukeylambda distribution shape tlshape by random sampling according to formula (3) s , generate Tukeylambda distribution with standard deviation σ by random sampling according to formula (4) tls ; S5.2, static region Poisson distribution noise N ps It is generated as follows (5), where P is the Poisson distribution and I is the static area of ​​the complete image; S5.3, read noise N rs It is generated as follows (6), where TL is the Tukeylambda distribution; S6, dynamic area noise generation: S6.1, randomly generate a dynamic area mask, generate a ratio of the full image length to the dynamic area length between 1.5 and 3 with uniform distribution, generate a ratio of the full image width to the dynamic area width between 1.5 and 3 with uniform distribution, generate a dynamic area in the full image, randomly generate a center point in the dynamic area, the center point is within 10x10 in the middle of the dynamic area, and ensure that the center point is generated within 10x10 pixels in the center of the dynamic area, split the dynamic area into 4 sub-dynamic areas according to the coordinates of the center point, specifically, determine the 4 sub-dynamic areas with the center coordinates as the vertex and the 4 vertices of the dynamic area as another vertex; S6.2, for the sub-dynamic region, regenerate the 3-sample system gain log(k sm ), plus the maximum system gain log(k max ) There are 4 different system gains in total, and the 4 different system gains are randomly assigned to 4 sub-dynamic areas; log(k sm ) ~ U(log(k s ), log(k max )) Equation (7) S6.3 Similarly, step S5 generates the Poisson distribution noise N of the four sub-dynamic regions according to equations (5) and (6). pm and read noise N rm ; S7, generate paired data, add the noise generated by the static area and the dynamic area to the clean data according to formula (8), where I c For clean data, I N For the final generated noise data, form paired data, I N = I C + (1 - mask)·(N ps + N rs ) + mask·(N pm + N rm ) Equation (8).

2. The spatial domain denoising method for adaptive temporal domain denoising under extremely dark light according to claim 1, characterized in that: The method also includes step S8: training the neural network model, generating paired data online according to step S7 during the training, using Adam as the optimizer, with a learning rate of 0.0001, a training cycle of 120, a learning rate reduced by 0.1 times every 20 cycles, and a training loss of L1loss.

3. The spatial domain denoising method for adaptive temporal domain denoising under extremely dark light according to claim 1, characterized in that: The method uses the imx327 image sensor.