Method for synthesizing 3D denoising task training data set

By collecting and synthesizing background images in multiple scenes of different brightness and combining foreground data to synthesize images, the problem of difficulty in constructing high-quality Raw domain video training data sets in the prior art is solved, and the effective 3D denoising effect of the model in low-light scenes is achieved to reduce noise flicker.

CN120031735APending Publication Date: 2025-05-23HEFEI JUNZHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311574747.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively build high-quality Raw domain video training data sets for 3D denoising tasks, especially in low-light scenarios, resulting in poor denoising effect of the model in dynamic areas and serious noise flickering.

Method used

By acquiring and synthesizing background images in multiple different brightness scenes, weighted averaging and further BM3D denoising processing are performed to simulate the 3D denoising effect. Then, the image is synthesized based on the foreground data, and the time domain and airspace noise of the synthesis are judged based on the displacement area, and a data set that meets the training requirements are generated.

Benefits of technology

In low-light scenarios, the model has a strong 3D denoising effect in static areas, and can also lightly 3D denoising in dynamic areas, reducing noise flickering and generating cleaner images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031735A_ABST
    Figure CN120031735A_ABST
Patent Text Reader

Abstract

The invention provides a method for synthesizing a 3D denoising task training data set, and the method comprises the steps: S1, collecting background images, and collecting data in a plurality of different brightness scenes; s2, carrying out background image denoising processing, carrying out weighted averaging on a plurality of background images, and simulating 3D denoising; s3, further de-noising the background image, and further de-noising by using BM3D; s4, acquiring foreground data; s5, synthesizing the image, synthesizing the foreground and the background, and generating an image simulating real displacement; and S6, synthesizing the data set, synthesizing noise on the image, and synthesizing the time domain and the final Ground Truth according to the judgment of the displacement region. Through the method, the collected and synthesized data set can well fit the difference of time domain and space domain denoising tasks in a 3D denoising task, and through the method, the model can have a very strong 3D denoising effect in a static region and also have a certain 3D effect in a dynamic region, so that the model has a good denoising effect. The phenomenon that in a low-light scene, airspace denoising cannot well complete a denoising task on a motion area is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video denoising processing, and in particular relates to a method for synthesizing a 3D denoising task training data set. Background Art

[0002] With the rapid development of image tasks in recent years, people's demand for high-quality images and videos has become increasingly prominent. However, in the existing technology, noise will have a negative impact on the visual quality of the image. Noise reduction and denoising can reduce or eliminate the noise in the image, making it clearer, easier to observe and understand, which is very important for many application fields such as photography, medical imaging, and surveillance video. In many image analysis and recognition tasks, noise may interfere with the extraction and accuracy of key information. Through noise reduction and denoising processing, the signal quality of the image can be improved and the interference of noise can be reduced, thereby improving the performance of tasks such as image analysis, target detection, and image recognition.

[0003] It is difficult to ensure high-quality imaging under many extreme conditions. Currently, various solutions often set a very high ISO under low-light conditions in order to increase image brightness, but this will directly lead to a large amount of noise in the generated image video, and the image quality will deteriorate, which is far from meeting the requirements of various tasks. Therefore, the denoising task is very important. In the Raw domain video directly collected by the camera, the noise generally follows a close to simple Poisson-Gaussian distribution, but in the RGB image generated after processing by the device ISP, the noise distribution becomes very complex because its distribution has changed after conversion, which makes the denoising task very difficult. Therefore, Raw domain denoising has more advantages than RGB domain denoising. Therefore, removing Raw domain noise is of great significance for generating higher quality images.

[0004] Deep learning methods based on convolutional neural networks have been widely used in various ISP tasks, but such methods often require a large amount of paired data. Currently, common public datasets are basically based on image data pairs in the sRGB domain, and a small amount of Raw domain datasets are usually collected in a way similar to stop-motion animation.

[0005] Therefore, the current dataset construction method is far from meeting the higher requirements of denoising tasks. The method of collecting Raw domain datasets is relatively difficult, which will increase the difficulty of obtaining datasets. The clean GroundTruth collected by stop-motion animation is also difficult to align the model with the moving area during training, and the noise in the dynamic area is difficult to remove. It is impossible to perform different supervision on the time domain and spatial domain respectively, resulting in the model not being able to converge well.

[0006] In addition, the commonly used terms in the prior art include:

[0007] Ground Truth: refers to the label data in the data set, which is also the expected result of the model in model training. AEGain: refers to the gain of electrical signals and digital signals in the automatic exposure algorithm of digital image processing. The gain after light is converted into current signals through the lens is called analog gain, and the gain after the current signal is converted into digital signals is called digital gain.

[0008] AWB: Automatic white balance. The original RAW image captured by the camera will always have color deviations after the interpolation algorithm. The R, G, and B channels of the camera need to be multiplied by different gain values ​​to restore the normal color tone.

[0009] Bayer format: Bayer format images are derived from the Bayer array, which is one of the main technologies for CCD or CMOS sensors to capture color images. The Bayer array was invented by Bryce Bayer, a scientist at Eastman Kodak, and is widely used in digital images. For color images, each pixel can be represented by three colors, RGB. The simplest sampling method is to use three filters on each pixel. The red filter transmits red wavelengths, the green filter transmits green wavelengths, and the blue filter transmits blue wavelengths. In this way, in order to collect the three basic colors of RGB, three filters are required for each point. This method is expensive, and because the three filters must be aligned to the same point, it is not easy to manufacture. The Bayer format can solve this problem well. Each pixel uses only one color filter. In addition, by analyzing the human eye's perception of color, it is found that the human eye is more sensitive to green, so there is more green in the Bayer format image, and the green pixel is the sum of the R and B pixels. Summary of the invention

[0010] In order to solve the above problems, the purpose of the present application is: through this method, the collected and synthesized data set can well fit the difference between the time domain and spatial domain denoising tasks in the 3D denoising task. Through this method, the model can have a strong 3D denoising effect in the static area and a certain 3D effect in the dynamic area, and solve the problem that the spatial domain denoising cannot complete the denoising task of the moving area well in the low-light scene, and the resulting noise flickering phenomenon in the moving area. Compared with the prior art, the method of synthesizing a 3D denoising task training data set of the present invention can be used to construct a large amount of "clean-noise" paired Raw domain video sequence image data, supporting Raw domain video denoising work.

[0011] Specifically, the present invention provides a method for synthesizing a 3D denoising task training data set, the method comprising the following steps:

[0012] S1. Background image collection, collecting data under multiple scenes with different brightness;

[0013] S2. Background image denoising, weighted averaging of multiple images, and simulated 3D denoising;

[0014] S3. The background image is further denoised using BM3D;

[0015] S4. Acquisition of prospect data;

[0016] S5. Synthesize the image, synthesize the foreground and background, and generate an image simulating real displacement;

[0017] S6. Synthetic dataset, synthesize noise on the image, synthesize the time domain and the final Ground Truth based on the judgment of the displacement area.

[0018] The method further comprises:

[0019] S1. Background image collection:

[0020] The data set covers a suitable interval and collects data of low-light scenes with different brightness, including between 0.1lux and 2lux. The image is brightened by increasing the gain value of the analog gain, and the maximum value of the analog gain of the current sensor is set to the maximum value of the gain. The digital gain is not used as the standard. Different scenes contain different scene brightness. With the same background and the same shooting parameters, n images are collected. For different brightness scenes, the value of n is also different. The darker the scene, the larger the value of n. M scenes are collected as needed, and the larger the m, the better.

[0021] S2. Background image denoising:

[0022] In the previous step S1, for each scene, there are n images taken with the same camera parameters, assuming that T 1 , T 2 , ...T n , synthesized by adding and averaging, as follows

[0023]

[0024] This can simulate the effect of 3D denoising and obtain images with lower noise. Where n represents the number of images taken with the same camera parameters, the numerator of the formula is a summation formula, which represents the addition of n images, and T represents the result after weighted averaging.

[0025] S3. Further denoising of background image:

[0026] For the m images processed by step S2, although the noise is controlled to a certain extent, there is still some residual noise, which cannot be directly used as the ground truth of the label image after synthesis; therefore, further denoising is performed, and the denoising degree is controlled without losing details, and most of the residual noise is removed, so that the flat area is smooth enough, and artificial screening is used, which shows that the flat area is smooth and noise-free, and this is used as the final ground truth of the label image;

[0027] S4. Foreground acquisition:

[0028] The foreground can directly obtain a large amount of different foreground data from the clean RAW image using the existing cutout algorithm. The cutout material image data is dual-channel, the first channel is the foreground RAW image, and the second channel is the mask information for the image position. The image size is 1080p, and the mask value in the foreground area is 1, and in the non-foreground area is 0;

[0029] S5. Image synthesis:

[0030] Randomly take a background image after step S3, count the AWB information and brightness mean Lb of the background image, and obtain the deviations on the three channels of R, G, and B; randomly take a foreground image, and count the brightness mean Li, adjust the foreground brightness according to Lb and Li to make it close to the background brightness, and adjust the foreground and background color deviation according to the AWB information to make the foreground and background consistent in color;

[0031] Randomly composite the foreground onto the background, and then generate a series of displacement, affine, scaling, and rotation operations to simulate the movement of the foreground on the background, thereby obtaining a series of continuous frames;

[0032] At the same time, the displacement information of the foreground between frames can be obtained according to the image position mask. Here, a clean and noise-free continuous frame image is synthesized as F 1 , F 2 , ..., F k , a total of k pieces;

[0033] S6. Dataset synthesis:

[0034] In the synthesis of different scenes F 1 , F 2 , ..., F k The upper distribution synthesizes Poisson noise and Gaussian noise to obtain I 1 , I 2 ,...I k , as the input of the data set; I 1 , I 2 ,...I k The non-motion area is extracted by mask and F is used1 , F 2 , ..., F k The non-moving part of the ground truth is replaced, that is, the ground truth of the static area is completely noise-free, so that only the moving area has noise, which is called fugt. 1 ′, fugt 2 ′,...fugt k '; The overlapping parts of the dynamic regions of consecutive frames are also subjected to 3D denoising, which is expressed as follows when synthesizing the data set:

[0035] Sequence fugt 1 ′, fugt 2 ′,...fugt k The noise of the overlapping part of the dynamic area is reduced by changing the additive noise noise in this area to noise / x, where x is the sequence number in the implementation process, that is, fugt 1 ′, fugt 2 ′,...fugt k ′ subscript;

[0036] The resulting sequence is called Fugt 1 , Fugt 2 ,...Fugt k , as the Ground Truth for time domain denoising, and F 1 , F 2 , ..., F k It can be directly used as the final Ground Truth after denoising in the time domain and spatial domain.

[0037] In the step S1, in practical applications, n is 200 at 2 lux and n is 1000 at 0.1 lux.

[0038] In the step S3, in actual application, the pre-processed label image is further denoised using the conventional denoising algorithm BM3D in the prior art. The BM3D algorithm includes: A. Basic estimation:

[0039] 1) For each target tile, find the most hyperparameter MAXN1 similar tiles nearby. To avoid the influence of noise, the tiles are transformed into 2D and then the similarity is measured by Euclidean distance. After sorting the tiles from small to large in distance, take the first MAXN1 and stack them into a three-dimensional array:

[0040]

[0041] 2) For the third dimension of the 3D array, that is, the array of pixels at the same position of each tile after the tiles are stacked, after DCT transformation, a hard threshold is used to reduce the pixel value that is less than the hyperparameter λ. 3D The components are set to 0; at the same time, the number of non-zero components is counted as a reference for subsequent weights; and then the third dimension is inversely transformed;

[0042]

[0043] 3) After inverse transformation, these blocks are put back to their original positions, and the weights are superimposed using the number of non-zero components. Finally, the superimposed image is divided by the weight of each point to obtain the basic estimated image. At this time, the noise points of the image are greatly removed;

[0044]

[0045] B. Final Estimate:

[0046] 1) For each target block of the noisy original image, the Euclidean distance of the corresponding basic estimated block can be directly used to measure the similarity; sort the blocks by distance from small to large and take the top MAXN1 blocks; stack the basic estimated blocks and the noisy original image blocks into two three-dimensional arrays respectively;

[0047]

[0048] 2) Perform DCT transformation on the third dimension of the 3D array containing the basic estimate, that is, the array consisting of the pixels at the same position of each block after the blocks are stacked, and obtain the coefficients using the following formula;

[0049]

[0050] 3) Multiply the coefficients by the noisy 3D image block and put it back to its original position, and finally make a weighted average adjustment to get the final estimated image;

[0051]

[0052] In step S5, k is set to 25 in practical applications.

[0053] In step S6, x is set to 3.

[0054] Therefore, the advantages of this application are:

[0055] 1. 3D denoising datasets have always been difficult to collect. Some synthetic datasets face the problem of not being able to fit the real scene time domain and spatial domain denoising well. Real datasets face the problem of difficult to obtain label data and too single noise distribution. This method uses a series of fixed steps to obtain a large number of data that meet the training requirements, reducing costs, and can be synthesized online, thereby greatly expanding the amount of data and enhancing the robustness of the model. During the training process, the model can be adaptively 3D denoising in static areas, and mild 3D denoising in dynamic areas, and then 2D denoising, so as to achieve the effect of no flickering and cleaner noise in dynamic areas in low-light scenes;

[0056] 2. Many-to-one data, different brightness scenes, and different ISO levels can enable model training to better adapt to 3D denoising tasks under different conditions, thereby improving the model's generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention.

[0058] Figure 1 It is a schematic diagram of the 3D denoising dataset process in this application.

[0059] Figure 2 It is a schematic diagram of the synthetic data set in this application. DETAILED DESCRIPTION

[0060] In order to more clearly understand the technical content and advantages of the present invention, the present invention is now further described in detail in conjunction with the accompanying drawings.

[0061] The present invention belongs to the technical field of 3D denoising in low-level visual tasks of deep learning neural networks, and relates to a method for synthesizing a training data set for a 3D denoising task.

[0062] like Figure 1 As shown, this application proposes a method for synthesizing a 3D denoising task training dataset, such as Figure 1 As shown, the process of synthesizing a 3D denoising dataset includes:

[0063] S1. Background image collection, collecting data under multiple scenes with different brightness;

[0064] S2. Background image denoising, weighted averaging of multiple images, and simulated 3D denoising;

[0065] S3. The background image is further denoised using BM3D;

[0066] S4. Acquisition of prospect data;

[0067] S5. Synthesize the image, synthesize the foreground and background, and generate an image simulating real displacement;

[0068] S6. Synthetic dataset, synthesize noise on the image, synthesize the time domain and the final Ground Truth based on the judgment of the displacement area.

[0069] The specific implementation steps of this method are as follows:

[0070] Step S1. Background image collection

[0071] The data set needs to cover a suitable range. Since it mainly solves problems in low-light scenes, it collects data of low-light scenes with different brightness, including between 0.1lux and 2lux. Traditional ISPs often brighten images by increasing the gain value, but too large a gain value will always bring about unfavorable factors such as high noise and overexposure of bright areas. Gain includes analog gain and digital gain, where analog gain is an electrical signal amplifier and digital gain is a digital signal amplifier. It can be seen that digital gain improves visual effects and does not add new information. Analog gain can maximize the existence of information. Therefore, in practical applications, the maximum value of the analog gain of the current sensor is set to the maximum value of the gain, and digital gain is not used as the standard.

[0072] Different scenes have different scene brightness. For the same background and shooting parameters, n images are collected. For different brightness scenes, the value of n is different. The darker the scene, the larger n is. For example, in actual applications, n is 200 at 2lux and 1000 at 0.1lux. m scenes are collected as needed, and the larger m is, the better. In the specific implementation, more than 200 scenes are collected.

[0073] Step S2: Background image denoising

[0074] In the previous step S1, for each scene, there are n images taken with the same camera parameters, such as (T 1 , T 2 , ...T n ), synthesized by adding and averaging, as follows

[0075]

[0076] This can simulate the effect of 3D denoising and obtain images with lower noise. Where n represents the number of images taken with the same camera parameters, the numerator is the summation formula, which represents the addition of n images, and T represents the result after weighted averaging.

[0077] Step S3: Background image further denoising

[0078] For the m images processed by step S2, although the noise is controlled to a certain extent, there is still some residual noise, and it is not appropriate to use them directly as the ground truth of the label image after synthesis. Therefore, further denoising is required. In practical application, the traditional denoising algorithm BM3D is used to further denoise the preprocessed label image. The denoising degree is controlled without losing details, and most of the residual noise is removed to make the flat area smooth enough. Artificial screening is used to show that the flat area is smooth and noise-free, which is used as the final ground truth of the label image.

[0079] In practical applications, the traditional denoising algorithms used include:

[0080] A. Basic Estimates:

[0081] 1) For each target tile, find the most hyperparameter MAXN1 similar tiles nearby. To avoid the influence of noise, the tiles are transformed into 2D and then the similarity is measured by Euclidean distance. After sorting the tiles from small to large in distance, take the first MAXN1 and stack them into a three-dimensional array:

[0082]

[0083] 2) For the third dimension of the 3D array, that is, the array of pixels at the same position of each tile after the tiles are stacked, after DCT transformation, a hard threshold is used to reduce the pixel value that is less than the hyperparameter λ. 3D The components are set to 0; at the same time, the number of non-zero components is counted as a reference for subsequent weights; and then the third dimension is inversely transformed;

[0084]

[0085] 3) After inverse transformation, these blocks are put back to their original positions, and the weights are superimposed using the number of non-zero components. Finally, the superimposed image is divided by the weight of each point to obtain the basic estimated image. At this time, the noise points of the image are greatly removed;

[0086]

[0087] B. Final Estimate:

[0088] 1) For each target block of the noisy original image, the Euclidean distance of the corresponding basic estimated block can be directly used to measure the similarity; sort the blocks by distance from small to large and take the top MAXN1 blocks; stack the basic estimated blocks and the noisy original image blocks into two three-dimensional arrays respectively;

[0089]

[0090] 2) Perform DCT transformation on the third dimension of the 3D array containing the basic estimate, that is, the array consisting of the pixels at the same position of each block after the blocks are stacked, and obtain the coefficients using the following formula;

[0091]

[0092] 3) Multiply the coefficients by the noisy 3D image block and put it back to its original position, and finally make a weighted average adjustment to get the final estimated image;

[0093]

[0094] Step S4. Foreground acquisition

[0095] The foreground can directly use the cutout algorithm on the clean RAW image to obtain a large amount of different foreground data. The cutout material image data is dual-channel, the first channel is the foreground RAW image, and the second channel is the mask information for the image position. The image size is 1080p, and the mask value in the foreground area is 1, and in the non-foreground area is 0. The cutout algorithm can be cut using an open source cutout tool or tools such as Photoshop.

[0096] Step S5: Image synthesis

[0097] Randomly take a background image after step S3, count the AWB information and brightness mean Lb of the background image, and get the deviations on the three channels of R, G, and B. Randomly take a foreground image, and count the brightness mean Li, adjust the foreground brightness according to Lb and Li to make it close to the background brightness, and adjust the foreground and background color deviation according to the AWB information to make the foreground and background consistent in color.

[0098] Randomly synthesize the foreground on the background, and then generate a series of displacement, affine, scaling, and rotation operations to simulate the movement of the foreground on the background, thereby obtaining a series of continuous frames. At the same time, the displacement information of the foreground between frames can be obtained according to the image position mask map. Here, a clean and noise-free continuous frame image synthesized is assumed to be F 1 , F 2 , ..., F k , there are k pictures in total, and k is 25 in practical applications.

[0099] Step S6. Dataset synthesis

[0100] In the synthesis of different scenes F 1 , F 2 , ..., F k The upper distribution synthesizes Poisson noise, Gaussian noise, etc. to obtain I 1 , I 2 ,...I k , as the input of the data set; I1 , I 2 ,...I k The non-motion area is extracted by mask and F is used 1 , F 2 , ..., F k The non-moving part of the ground truth is replaced, that is, the ground truth of the static area is completely noise-free, so that only the moving area has noise. The sequence is called fugt 1 ′, fugt 2 ′,...fugt k ′.

[0101] In actual use, traditional 3D denoising uses a small degree of 3D denoising on dynamic areas instead of using 2D algorithms. So referring to this idea, the overlapping parts of the dynamic areas of consecutive frames are also subjected to 3D denoising, which is shown in the synthetic data set as follows:

[0102] Sequence fugt 1 ′, fugt 2 ′,...fugt k The noise of the overlapping part of the dynamic area is reduced by changing the additive noise noise in this area to noise / x, where x is the sequence number in the implementation process, that is, fugt 1 ′, fugt 2 ′,...fugt k ′, and x is 3 in the experiment.

[0103] The resulting sequence is called Fugt 1 , Fugt 2 ,...Fugt k , as the Ground Truth for time domain denoising, and F 1 , F 2 , ..., F k It can be directly used as the final Ground Truth after denoising in the time domain and spatial domain.

[0104] like Figure 2 As shown, it is a specific diagram of step S5, that is, a schematic diagram of the synthetic data set: from I1, I2, ..., Ik; the gray part is noisy, such as I1, I2, Ik; from Fugt1, Fugt2, ..., Fugtk; the white part is noise-free, such as Fugt1; among them, the noise in the overlapping part of the two circles of Fugt2 is smaller. If the gray is noise, the overlapping part is noise / 3.0, and this difference is not shown in the image.

[0105] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for synthesizing a training dataset for 3D denoising tasks, It is characterized in that The method comprises the following steps: S1. Background image collection, collecting data under multiple scenes with different brightness; S2. Background image denoising, weighted averaging of multiple images, and simulated 3D denoising; S3. The background image is further denoised using BM3D; S4. Acquisition of prospect data; S5. Synthesize the image, synthesize the foreground and background, and generate an image simulating real displacement; S6. Synthetic dataset, synthesize noise on the image, synthesize the time domain and the final GroundTruth based on the judgment of the displacement area.

2. A method for synthesizing a 3D denoising task training data set according to claim 1, It is characterized in that The method further comprises: S1. Background image collection: The data set covers a suitable interval and collects data of low-light scenes with different brightness, including between 0.1lux and 2lux. The image is brightened by increasing the gain value of the analog gain, and the maximum value of the analog gain of the current sensor is set to the maximum value of the gain. The digital gain is not used as the standard. Different scenes contain different scene brightness. With the same background and the same shooting parameters, n images are collected. For different brightness scenes, the value of n is also different. The darker the scene, the larger the value of n. M scenes are collected as needed, and the larger the m, the better. S2. Background image denoising: In the previous step S1, for each scene, there are n images taken with the same camera parameters, assuming that T 1 , T 2 , ...T n , synthesized by adding and averaging, as follows This can simulate the effect of 3D denoising and obtain images with lower noise. Where n represents the number of images taken with the same camera parameters, the numerator in the formula is a summation formula, which represents the addition of n images, and T represents the result after weighted averaging. S3. Further denoising of background image: For the m images processed by step S2, although the noise is controlled to a certain extent, there is still some residual noise, which cannot be directly used as the ground truth of the label image after synthesis; therefore, further denoising is performed, and the denoising degree is controlled without losing details, and most of the residual noise is removed, so that the flat area is smooth enough, and artificial screening is used, which shows that the flat area is smooth and noise-free, and this is used as the final ground truth of the label image; S4. Foreground acquisition: The foreground can directly obtain a large amount of different foreground data from the clean RAW image using the existing cutout algorithm. The cutout material image data is dual-channel, the first channel is the foreground RAW image, and the second channel is the mask information for the image position. The image size is 1080p, and the mask value in the foreground area is 1, and in the non-foreground area is 0; S5. Image synthesis: Randomly take a background image after step S3, count the AWB information and brightness mean Lb of the background image, and obtain the deviations on the three channels of R, G, and B; randomly take a foreground image, and count the brightness mean Li, adjust the foreground brightness according to Lb and Li to make it close to the background brightness, and adjust the foreground and background color deviation according to the AWB information to make the foreground and background consistent in color; Randomly composite the foreground onto the background, and then generate a series of displacement, affine, scaling, and rotation operations to simulate the movement of the foreground on the background, thereby obtaining a series of continuous frames; At the same time, the displacement information of the foreground between frames can be obtained according to the image position mask. Here, a clean and noise-free continuous frame image is synthesized as F 1 , F 2 , ..., F k , a total of k pieces; S6. Dataset synthesis: In the synthesis of different scenes F 1 , F 2 , ..., F k The upper distribution synthesizes Poisson noise and Gaussian noise to obtain I 1 , I 2 ,...I k , as the input of the data set; I 1 , I 2 ,...I k The non-motion area is extracted by mask and F is used 1 , F 2 , ..., F k The non-moving part of the ground truth is replaced, that is, the ground truth of the static area is completely noise-free, so that only the moving area has noise, which is called fugt. 1 ′, fugt 2 ′,...fugt k '; The overlapping parts of the dynamic regions of consecutive frames are also subjected to 3D denoising, which is expressed as follows when synthesizing the data set: Sequence fugt 1 ′, fugt 2 ′,...fugt k The noise of the overlapping part of the dynamic area is reduced by changing the additive noise noise in this area to noise / x, where x is the sequence number in the implementation process, that is, fugt 1 ′, fugt 2 ′,...fugt k The subscript of ; The resulting sequence is called Fugt 1 , Fugt 2 ,...Fugt k , as the Ground Truth for time domain denoising, and F 1 , F 2 , ..., F k It can be directly used as the final Ground Truth after denoising in the time domain and spatial domain.

3. A method for synthesizing a 3D denoising task training data set according to claim 1, It is characterized in that In the step S1, in practical applications, n is 200 at 2 lux and n is 1000 at 0.1 lux.

4. A method for synthesizing a 3D denoising task training data set according to claim 1, It is characterized in that In step S3, in actual application, the pre-processed label image is further denoised using the conventional denoising algorithm BM3D in the prior art. The BM3D algorithm includes: A. Basic Estimates: 1) For each target tile, find the most hyperparameter MAXN1 similar tiles nearby. To avoid the influence of noise, the tiles are transformed into 2D and then the similarity is measured by Euclidean distance. After sorting the tiles from small to large in distance, take the first MAXN1 and stack them into a three-dimensional array: 2) For the third dimension of the 3D array, that is, the array of pixels at the same position of each tile after the tiles are stacked, after DCT transformation, a hard threshold is used to reduce the pixel value that is less than the hyperparameter λ. 3D The components are set to 0; at the same time, the number of non-zero components is counted as a reference for subsequent weights; and then the third dimension is inversely transformed; 3) After inverse transformation, these blocks are put back to their original positions, and the weights are superimposed using the number of non-zero components. Finally, the superimposed image is divided by the weight of each point to obtain the basic estimated image. At this time, the noise points of the image are greatly removed; B. Final Estimate: 1) For each target block of the noisy original image, the Euclidean distance of the corresponding basic estimated block can be directly used to measure the similarity; sort the blocks by distance from small to large and take the top MAXN1 blocks; stack the basic estimated blocks and the noisy original image blocks into two three-dimensional arrays respectively; 2) Perform DCT transformation on the third dimension of the 3D array containing the basic estimate, that is, the array consisting of the pixels at the same position of each block after the blocks are stacked, and obtain the coefficients using the following formula; 3) Multiply the coefficients by the noisy 3D image block and put it back to its original position, and finally make a weighted average adjustment to get the final estimated image; 5. A method for synthesizing a 3D denoising task training data set according to claim 2, It is characterized in that In step S5, k is set to 25 in practical applications.

6. A method for synthesizing a 3D denoising task training data set according to claim 2, It is characterized in that In step S6, x is set to 3.