Model training method and model training apparatus

By generating multiple noisy videos and training an initial network model, the problem of high denoising difficulty in microscopic videos was solved, and efficient denoising effect of microscopic videos was achieved.

CN114331901BActive Publication Date: 2025-11-04BEIJING CHAOWEIJING BIOLOGICAL TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111663390.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-11-04
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

The signal-to-noise ratio of microscopic videos is much lower than that of natural images, making denoising difficult and resulting in poor denoising performance.

Method used

By generating multiple noisy videos and training an initial network model using a noise model, a loss function is constructed to generate a video denoising model. This model learns the noise information of the video to be denoised, thus achieving effective denoising of the video.

Benefits of technology

It improves the denoising effect of microscopic videos, generating clearer denoised videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114331901B_ABST
    Figure CN114331901B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image processing, in particular to a model training method and device, a computer readable storage medium and an electronic device, and solves the problems of high difficulty and poor denoising effect of microscopic video denoising. The model training method generates a corresponding noise-increasing video of a to-be-denoised video based on the to-be-denoised video, and trains an initial network model based on the noise-increasing video and the to-be-denoised video to generate a video denoising model, so that the initial network model can learn the noise information of the to-be-denoised video in the training, and the generated video denoising model can accurately denoise the to-be-denoised video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a model training method and a model training device, and a computer readable storage medium and an electronic device. BACKGROUND

[0002] Microscopic video refers to a video generated by recording a dynamic experimental process observed under an optical microscope by using a camera and then processing the video by using related software. For example, a calcium imaging video is a kind of microscopic video. Fluorescent proteins are used to specifically mark neuron cells, and fluorescent proteins marking calcium ions can emit different intensities of fluorescence according to the changes in the concentration of calcium ions in neuron cells. The changes in the fluorescent signals can be recorded in the form of a video. The video for realizing imaging according to the changes in the concentration of calcium ions can be simply referred to as a calcium imaging video.

[0003] In the process of producing a calcium imaging video, in order to reduce the phototoxicity to organisms, the number of photons participating in imaging is much smaller than that in natural imaging. Therefore, the signal-to-noise ratio of various microscopic videos including the calcium imaging video is much smaller than that of a video of natural imaging, resulting in high difficulty in denoising the microscopic video and poor denoising effect. SUMMARY

[0004] Therefore, an embodiment of the present application provides a model training method and a model training device, and a computer readable storage medium and an electronic device, which solve the problems of high difficulty in denoising a microscopic video and poor denoising effect.

[0005] In a first aspect, an embodiment of the present application provides a model training method, which comprises: generating a plurality of noise-increasing videos corresponding to a to-be-denoised video based on the to-be-denoised video; establishing an initial network model and training the initial network model based on the plurality of noise-increasing videos and the to-be-denoised video to generate a video denoising model, wherein the video denoising model is used to denoise the to-be-denoised video to generate a noise-reduced video corresponding to the to-be-denoised video.

[0006] In combination with the first aspect of the present application, in some embodiments, the generating of the plurality of noise-increasing videos corresponding to the to-be-denoised video based on the to-be-denoised video comprises: generating the plurality of noise-increasing videos corresponding to the to-be-denoised video based on the to-be-denoised video by using a noise model, wherein the plurality of noise-increasing videos contain different noise intensities, and the noise model comprises any one of the following models: an additive noise prior model, a multiplicative noise prior model and an additive-multiplicative compound prior model.

[0007] With reference to the first aspect of the present application, in some embodiments, the generating, based on the to-be-denoised video, a plurality of noise-increased videos corresponding to the to-be-denoised video comprises: determining noise intensity estimation information of the to-be-denoised video based on the to-be-denoised video; performing N times of noise adding operations on the to-be-denoised video based on the noise intensity estimation information to determine N noise-increased videos corresponding to the to-be-denoised video, wherein N is a positive integer.

[0008] With reference to the first aspect of the present application, in some embodiments, the determining, based on the to-be-denoised video, noise intensity estimation information of the to-be-denoised video comprises: determining at least one frame of video image corresponding to the to-be-denoised video based on the to-be-denoised video; generating a plurality of image regions corresponding to the at least one frame of video image based on the at least one frame of video image; calculating a gray mean value and a gray variance corresponding to each of the plurality of image regions; and determining the noise intensity estimation information based on the gray mean value and the gray variance corresponding to each of the plurality of image regions.

[0009] With reference to the first aspect of the present application, in some embodiments, the determining, based on the gray mean value and the gray variance corresponding to each of the plurality of image regions, the noise intensity estimation information comprises: performing a linear regression operation based on the gray mean value and the gray variance corresponding to each of the plurality of image regions to determine the noise intensity estimation information of the to-be-denoised video.

[0010] With reference to the first aspect of the present application, in some embodiments, the performing, based on the gray mean value and the gray variance corresponding to each of the plurality of image regions, the linear regression operation to determine the noise intensity estimation information of the to-be-denoised video comprises: establishing a data point graph of the gray mean value and the gray variance based on the gray mean value and the gray variance corresponding to each of the plurality of image regions, the data point graph comprising a plurality of data points; performing a linear regression operation on the plurality of data points to determine a regression straight line function; and determining the noise intensity estimation information of the to-be-denoised video based on the regression straight line function, wherein the noise intensity estimation information of the to-be-denoised video comprises a Poisson noise intensity, a Gaussian noise mean value, and a Gaussian noise variance.

[0011] With reference to the first aspect of the present application, in some embodiments, the training, based on the plurality of noise-increased videos and the to-be-denoised video, an initial network model to generate a video denoising model comprises: constructing a loss function of the initial network model, wherein the loss function comprises a time domain smoothing kernel norm regular term and a spatial domain smoothing entropy regular term; and training the initial network model based on the plurality of noise-increased videos, the to-be-denoised video, and the loss function to generate the video denoising model.

[0012] With reference to the first aspect of the present application, in some embodiments, the to-be-denoised video comprises a to-be-denoised microscopic video.

[0013] In a second aspect, an embodiment of the present application provides a model training apparatus, comprising: a generation module configured to generate a plurality of noise-increasing videos corresponding to a to-be-denoised video based on the to-be-denoised video; and a training module configured to establish an initial network model and train the initial network model based on the plurality of noise-increasing videos and the to-be-denoised video to generate a video denoising model, wherein the video denoising model is configured to denoise the to-be-denoised video to generate a noise-reduced video corresponding to the to-be-denoised video.

[0014] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores instructions, when the instructions are executed by a processor of an electronic device, enable the electronic device to perform the model training method of the first aspect and / or the video denoising method of the second aspect.

[0015] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising: a processor; a memory for storing computer executable instructions; and the processor is configured to execute the computer executable instructions to implement the model training method of the first aspect and / or the video denoising method of the second aspect.

[0016] The to-be-denoised video is a video containing internal noise, and there is no noise-free ideal denoising gold standard as a label in the training sample to train the initial network model. Therefore, the model training method and the model training apparatus provided by the embodiment of the present application, and the computer-readable storage medium and the electronic device, generate the noise-increasing video corresponding to the to-be-denoised video based on the to-be-denoised video, and train the initial network model based on the noise-increasing video and the to-be-denoised video to generate the video denoising model, so that the to-be-denoised video is used as the label of the noise-increasing video to train the initial network model, so that the initial network model can learn the noise information of the to-be-denoised video in the training, so that the generated video denoising model can accurately denoise the to-be-denoised video. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 Fig. 1 shows an application scenario schematic diagram of the model training method provided by an embodiment of the present application.

[0018] Figure 2 Fig. 2 shows a flowchart of the model training method provided by an embodiment of the present application.

[0019] Figure 3 Fig. 3 shows a flowchart of the model training method provided by another embodiment of the present application.

[0020] Figure 4 Fig. 4 shows a flowchart of the model training method provided by another embodiment of the present application.

[0021] Figure 5 Fig. 5 shows a flowchart of the model training method provided by another embodiment of the present application.

[0022] Figure 6 Fig. 2 shows a flowchart of a model training method according to another embodiment of the present application.

[0023] Figure 6a Fig. 3 shows a video image according to an embodiment of the present application.

[0024] Figure 7 Fig. 4 shows a flowchart of a model training method according to another embodiment of the present application.

[0025] Figure 8 Fig. 5 shows a flowchart of a video denoising method according to an embodiment of the present application.

[0026] Figure 9 Fig. 6 shows a flowchart of a video denoising method according to another embodiment of the present application.

[0027] Figure 10 Fig. 7 shows a structural diagram of a model training device according to an embodiment of the present application.

[0028] Figure 11 Fig. 8 shows a structural diagram of a model training device according to another embodiment of the present application.

[0029] Figure 12 Fig. 9 shows a structural diagram of a model training device according to another embodiment of the present application.

[0030] Figure 13 Fig. 10 shows a structural diagram of a model training device according to another embodiment of the present application.

[0031] Figure 14 Fig. 11 shows a structural diagram of a model training device according to another embodiment of the present application.

[0032] Figure 15 Fig. 12 shows a structural diagram of a model training device according to another embodiment of the present application.

[0033] Figure 16 Fig. 13 shows a structural diagram of a video denoising device according to an embodiment of the present application.

[0034] Figure 17 Fig. 14 shows a structural diagram of a video denoising device according to another embodiment of the present application.

[0035] Figure 18 Fig. 15 shows a structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0036] With reference to the drawings and the embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0037] Figure 1 Fig. 1 shows an application scenario of the model training method provided by an embodiment of the present application. Figure 1 The scenario shown includes a server 110 and a video shooting device 120 in communication connection with the server 110. Specifically, the server 110 is configured to generate a noise-increasing video corresponding to a to-be-de-noised video based on the to-be-de-noised video, establish an initial network model, and train the initial network model based on the noise-increasing video and the to-be-de-noised video to generate a video de-noising model, wherein the video de-noising model is configured to de-noise the to-be-de-noised video to generate a de-noised video corresponding to the to-be-de-noised video.

[0038] Exemplarily, in actual application, the video shooting device 120 is configured to shoot a to-be-de-noised video and send the acquired to-be-de-noised video to the server 110, and the server 110 is configured to determine a de-noised video corresponding to the to-be-de-noised video based on the received to-be-de-noised video, and can display the de-noised video corresponding to the to-be-de-noised video to a user.

[0039] Exemplary method

[0040] Figure 2 Fig. 2 shows a flowchart of the model training method provided by an embodiment of the present application. As shown in Fig. 2, the model training method provided by the embodiment of the present application includes the following steps. Figure 2 As shown in Fig. 2, the model training method provided by the embodiment of the present application includes the following steps.

[0041] Step 210: generating a noise-increasing video corresponding to a to-be-de-noised video based on the to-be-de-noised video.

[0042] Specifically, the to-be-de-noised video is a video containing internal noise. The noise-increasing video can be a video containing double noise. The double noise includes internal noise and external noise.

[0043] In an embodiment of the present application, the to-be-de-noised video can be a to-be-de-noised microscopic video. For example, the to-be-de-noised video can be a calcium imaging video.

[0044] In an embodiment of the present application, a plurality of noise-increasing videos corresponding to the video to be denoised can be generated based on the video to be denoised by using a noise model. The plurality of noise-increasing videos contain different noise intensities. The noise model can be an additive noise prior model, a multiplicative noise prior model, or an additive-multiplicative compound prior model. The multiplicative noise prior model can be a Poisson noise model, the additive noise prior model can be a Gaussian noise model, and the additive-multiplicative compound prior model can be a compound model of Poisson noise and Gaussian noise.

[0045] At step 220, an initial network model is established, and the initial network model is trained based on the noise-increasing videos and the video to be denoised to generate a video denoising model.

[0046] For example, the video denoising model is used to denoise the video to be denoised to generate a noise-reduced video corresponding to the video to be denoised. The noise-increasing videos use y * to represent, and the video to be denoised uses y to represent, thereby forming a training set {y * , y}. The training set {y * , y} is used to train the initial network model to generate the video denoising model.

[0047] Specifically, the video to be denoised can be a video containing internal noise, and the noise-increasing video can be a video containing double noise. The double noise includes internal noise and external noise. The noise-free ideal denoising gold standard refers to a video completely free of noise. Since there is no corresponding noise-free ideal denoising gold standard as a label in the training sample to train the initial network model, the present application generates noise-increasing videos corresponding to the video to be denoised based on the video to be denoised, and trains the initial network model based on the noise-increasing videos and the video to be denoised to generate a video denoising model, so that the video to be denoised is used as a label of the noise-increasing video to train the initial network model, so that the initial network model can learn the noise information of the video to be denoised in the training, so that the generated video denoising model can accurately denoise the video to be denoised.

[0048] Figure 3 Fig. 2 shows a flowchart of a model training method provided by another embodiment of the present application. Based on the embodiment shown in Fig. 1, the embodiment shown in Fig. 2 extends the embodiment shown in Fig. 1. Figure 2 Fig. 3 shows a flowchart of a model training method provided by another embodiment of the present application. Based on the embodiment shown in Fig. 1, the embodiment shown in Fig. 3 extends the embodiment shown in Fig. 1. Figure 3 Fig. 4 shows a flowchart of a model training method provided by another embodiment of the present application. Based on the embodiment shown in Fig. 1, the embodiment shown in Fig. 4 extends the embodiment shown in Fig. 1. Figure 3 Fig. 5 shows a flowchart of a model training method provided by another embodiment of the present application. Based on the embodiment shown in Fig. 1, the embodiment shown in Fig. 5 extends the embodiment shown in Fig. 1. Figure 2 Fig. 6 shows a flowchart of a model training method provided by another embodiment of the present application. Based on the embodiment shown in Fig. 1, the embodiment shown in Fig. 6 extends the embodiment shown in Fig. 1.

[0049] As shown in Fig. 1, in the embodiment of the present application, the step of generating a plurality of noise-increasing videos corresponding to the video to be denoised based on the video to be denoised includes the following steps. Figure 3

[0050] ​At step 310, noise intensity estimation information of the to-be-denoised video is determined based on the to-be-denoised video.

[0051] The noise intensity estimation information is estimation information obtained by estimating internal noise contained in the to-be-denoised video. Specifically, the to-be-denoised video and the noise intensity estimation information can be fused to add external noise to the to-be-denoised video, so as to generate a noise-increased video corresponding to the to-be-denoised video. That is, the noise-increased video is obtained by adding external noise to the to-be-denoised video according to the noise intensity estimation information.

[0052] At step 320, N times of noise adding operations are performed on the to-be-denoised video based on the noise intensity estimation information, so as to determine N noise-increased videos corresponding to the to-be-denoised video.

[0053] Specifically, each time the noise adding operation is performed, the to-be-denoised video is added with noise intensity estimation information once. N is a positive integer. When N = 1, the to-be-denoised video is added with noise once to determine the first noise-increased video. When N = 2, the to-be-denoised video is added with noise twice to determine the second noise-increased video. In this way, a set of noise-increased videos {y *} is obtained. The set of noise-increased videos {y *} includes N noise-increased videos.

[0054] By performing the noise adding operation on the to-be-denoised video multiple times to obtain multiple noise-increased videos, the multiple noise-increased videos obtained all contain different amounts of noise intensity estimation information, which increases the richness of the set of noise-increased videos, thereby increasing the richness of the training set {y * , y} and providing a rich data basis for training the initial network model.

[0055] Figure 4 Fig. 6 shows a flowchart of a model training method provided by another embodiment of the present application. The embodiment shown in Fig. 6 is extended on the basis of the embodiment shown in Fig. 1. Figure 2 The embodiment shown in Fig. 6 is extended on the basis of the embodiment shown in Fig. 1. Figure 4 The embodiment shown in Fig. 6 is extended on the basis of the embodiment shown in Fig. 1. Figure 4 The embodiment shown in Fig. 6 is extended on the basis of the embodiment shown in Fig. 1. Figure 2 The embodiment shown in Fig. 6 is extended on the basis of the embodiment shown in Fig. 1.

[0056] As shown in Fig. 6, in the embodiment of the present application, the step of determining the noise intensity estimation information of the to-be-denoised video based on the to-be-denoised video includes the following steps. Figure 4 At step 410, at least one video image corresponding to the to-be-denoised video is determined based on the to-be-denoised video.

[0057]

[0058] ​Specifically, the to-be-denoised video includes a plurality of video images. At least one video image corresponding to the to-be-denoised video is determined based on the to-be-denoised video. The at least one video image can be any one or more video images selected from the plurality of video images included in the to-be-denoised video.

[0059] At step 420, a plurality of image regions corresponding to the at least one video image are generated based on the at least one video image.

[0060] Specifically, the at least one video image can be one video image or a plurality of video images. The plurality of image regions corresponding to the at least one video image can be generated by cropping the one video image into the plurality of image regions or by cropping the plurality of video images into the plurality of image regions. For example, the one video image can be cropped into k image regions, and the plurality of video images can also be cropped into k image regions. The video image can be randomly cropped into the plurality of image regions, i.e., the plurality of image regions can overlap or not overlap. The image regions can be represented by Q k When k = 1, Q k = Q1 represents the first image region, when k = 2, Q k = Q2 represents the second image region, and so on.

[0061] At step 430, the gray mean and the gray variance corresponding to each of the plurality of image regions are calculated.

[0062] For example, the gray mean is represented by E[Q k ], and the gray variance is represented by σ 2 [Q k ]. When k = 1, E[Q k ] = E[Q1] represents the gray mean corresponding to the first image region, and σ 2 [Q k ] = σ 2 [Q1] represents the gray variance corresponding to the first image region; when k = 2, E[Q k ] = E[Q2] represents the gray mean corresponding to the second image region, and σ 2 [Q k ] = σ 2 [Q2] represents the gray variance corresponding to the second image region, and so on.

[0063] At step 440, noise intensity estimation information is determined based on the gray mean and the gray variance corresponding to each of the plurality of image regions.

[0064] Specifically, a relationship function of the gray mean and the gray variance corresponding to each of the plurality of image regions can be constructed to calculate the noise intensity estimation information.

[0065] In embodiments of this application, at least one video frame corresponding to the video to be denoised is determined based on the video to be denoised, and grayscale analysis is performed on the at least one video frame to determine noise intensity estimation information. Since video images mainly represent noise in grayscale, determining noise intensity estimation information by performing grayscale analysis on video images improves the accuracy of noise intensity estimation information.

[0066] Figure 5 The diagram shown is a flowchart illustrating a model training method provided in another embodiment of this application. Figure 4 This application extends from the embodiments shown. Figure 5 The illustrated embodiment will be described in detail below. Figure 5 The illustrated embodiments and Figure 4 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.

[0067] like Figure 5 As shown in the embodiments of this application, the step of determining noise intensity estimation information based on the gray-level mean and gray-level variance corresponding to each of multiple image regions includes the following steps.

[0068] Step 510: Perform linear regression based on the mean and variance of gray levels of each of the multiple image regions to determine the noise intensity estimation information of the video to be denoised.

[0069] Specifically, linear regression is a statistical analysis method that uses regression analysis in mathematical statistics to determine the quantitative relationship of interdependence between two or more variables.

[0070] By performing linear regression, the linear relationship between the gray-level mean and the gray-level variance can be determined, thereby identifying noise intensity estimation information within this linear relationship. The method is simple, reliable, and efficient.

[0071] Figure 6 The diagram shown is a flowchart illustrating a model training method provided in another embodiment of this application. Figure 5 This application extends from the embodiments shown. Figure 6 The illustrated embodiment will be described in detail below. Figure 6 The illustrated embodiments and Figure 5 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.

[0072] like Figure 6 As shown in the embodiments of this application, the step of determining the noise intensity estimation information of the video to be denoised by performing linear regression operation based on the gray mean and gray variance corresponding to each of multiple image regions includes the following steps.

[0073] At step 610, a data point diagram of the gray value mean and the gray value variance is established based on the gray value mean and the gray value variance of each of the plurality of image regions, and the data point diagram includes a plurality of data points.

[0074] Specifically, a coordinate system can be established with the gray value mean E[Q k ] as the horizontal axis and the gray value variance σ 2 [Q k ] as the vertical axis. Then, the plurality of data points are determined according to the gray value mean and the gray value variance of each of the plurality of image regions, so as to establish the data point diagram of the gray value mean and the gray value variance. Each data point represents the gray value mean and the gray value variance of a corresponding image region.

[0075] At step 620, a linear regression operation is performed on the plurality of data points to determine a regression straight line function.

[0076] Exemplarily, the linear regression operation performed on the plurality of data points can be a linear regression fitting on the plurality of data points, so as to obtain the regression straight line function.

[0077] At step 630, the noise intensity estimation information of the to-be-de-noised video is determined based on the regression straight line function.

[0078] Exemplarily, the noise intensity estimation information of the to-be-de-noised video includes a Poisson noise intensity, a Gaussian noise mean and a Gaussian noise variance. The Poisson noise intensity can be represented by , the Gaussian noise mean can be represented by , and the Gaussian noise variance can be represented by .

[0079] The regression straight line function is as follows:

[0080]

[0081] As can be seen from the above formula, the slope a of the regression straight line function is The intercept d of the regression straight line function is The linear regression operation performed on the plurality of data points to determine the regression straight line function can determine the slope a and the intercept d of the regression straight line function. Therefore, the determination of the regression straight line function can directly obtain the Poisson noise intensity and can determine the relationship between the Poisson noise intensity , the Gaussian noise mean and the Gaussian noise variance . That is, after the determination of the regression straight line function, the relationship between the Poisson noise intensity , the Gaussian noise mean and the Gaussian noise variance is known, and the relationship between the Gaussian noise mean and the Gaussian noise variance and Gaussian noise variance unknown.

[0082] Gaussian noise mean It can be calculated based on the grayscale of the video image. For example... Figure 6a As shown, the video image includes white highlight areas with low grayscale values ​​and black areas with high grayscale values. (Gaussian noise mean) This could be the average grayscale value of a black region with high grayscale values. The mean Gaussian noise can be obtained by calculating the average grayscale value of the black region with high grayscale values. Then, according to the formula The variance of Gaussian noise can then be calculated.

[0083] This embodiment first establishes a data point map of gray-level mean and variance based on the gray-level mean and variance corresponding to multiple image regions. Then, a linear regression operation is performed on multiple data points to determine the regression line function. Finally, based on the regression line function, the Poisson noise intensity, Gaussian noise mean, and Gaussian noise variance of the video to be denoised are determined. Since the internal noise contained in the video to be denoised is mainly a mixture of Poisson and Gaussian noise, and this mixture is characterized by Poisson noise intensity, Gaussian noise mean, and Gaussian noise variance, determining these parameters identifies the main internal noise of the video, thus providing accurate training data for the initial network model training.

[0084] Figure 7 The diagram shown is a flowchart illustrating a model training method provided in another embodiment of this application. Figure 2 This application extends from the embodiments shown. Figure 7 The illustrated embodiment will be described in detail below. Figure 7 The illustrated embodiments and Figure 2 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.

[0085] like Figure 7 As shown in the embodiments of this application, the step of training an initial network model based on multiple noisy videos and videos to be denoised to generate a video denoising model includes the following steps.

[0086] Step 710: Construct the loss function for the initial network model.

[0087] For example, the loss function includes a temporal smoothing kernel norm regularization term and a spatial smoothing entropy regularization term. The loss function is as follows:

[0088] L = L rec +λ1L nuclear +λ2L entropy

[0089] wherein,

[0090] In the above formula, L is a loss value, L rec is a fidelity term, L nuclear is a time domain smooth kernel norm regularization term, L entropy is a spatial domain smooth entropy regularization term, λ1 is a weight of the time domain smooth kernel norm regularization term, λ2 is a weight of the spatial domain smooth entropy regularization term, is an output video of the initial network model in the process of training the initial network model, y is a video to be denoised, and ε is an infinitesimal, is a gray value of the output video of the initial network model . is a Hessian matrix of the output video of the initial network model . is a Hessian matrix of the output video of the initial network model In a, b, c three different directions, the second derivative is calculated, The Hessian matrix of the output video of the initial network model

[0091]

[0092] Step 720, training the initial network model based on the plurality of noise-increasing videos, the video to be denoised, and the loss function to generate a video denoising model.

[0093] Specifically, the noise-increasing video y * is taken as the input of the initial network model, and the video to be denoised y is taken as the label of the noise-increasing video y * to train the initial network model. In the training process, the weight λ1 of the time domain smooth kernel norm regularization term and the weight λ2 of the spatial domain smooth entropy regularization term are constantly adjusted until the function value L of the loss function converges, thereby obtaining the video denoising model.

[0094] The video to be denoised can be a calcium imaging video of a neuron cell. In the calcium imaging process, the location of the cell body of the neuron cell does not change over time, but the gray value of the cell body changes, that is, the calcium imaging video contains both time and space information, thereby indicating that the calcium imaging video contains a spatio-temporal coupling structure. Most microscopic videos contain spatio-temporal coupling structures.

[0095] The time domain smoothing kernel norm regular term and the space domain smoothing entropy regular term are set in the loss function, low rank information in the space-time coupling structure can be extracted and noise signals changing too fast in the time domain can be eliminated. Therefore, by constructing the loss function, and training the initial network model based on the noise-increasing video, the video to be denoised and the loss function to generate the video denoising model, the initial network model can learn the space-time coupling structure and noise in the video to be denoised, so that the obtained video denoising model can more accurately retain the space-time coupling structure in the video to be denoised and remove the noise in the video to be denoised.

[0096] Figure 8 Fig. 1 is a flowchart of a video denoising method according to an embodiment of the present application. As shown in Fig. 1, the video denoising method according to the embodiment of the present application includes the following steps. Figure 8

[0097] Step 810, obtaining a video denoising model.

[0098] Exemplarily, the video denoising model is generated based on the model training method in any of the above embodiments.

[0099] Step 820, generating a noise-reduced video corresponding to the video to be denoised based on the video to be denoised by using the video denoising model.

[0100] By using the video denoising model in the above embodiments, a noise-reduced video corresponding to the video to be denoised is generated based on the video to be denoised, which can accurately remove the noise of the video to be denoised and obtain a clearer noise-reduced video.

[0101] Figure 9 Fig. 2 is a flowchart of a video denoising method according to another embodiment of the present application. Based on the embodiment of the present application Figure 8 The embodiment of the present application extends from the embodiment of the present application Figure 9 The embodiment of the present application extends from the embodiment of the present application Figure 9 The embodiment of the present application extends from the embodiment of the present application Figure 8 The embodiment of the present application extends from the embodiment of the present application

[0102] As shown in Fig. 2, the video denoising method according to the embodiment of the present application includes the following steps. Figure 9 As shown in Fig. 2, the video denoising method according to the embodiment of the present application includes the following steps.

[0103] Step 910, generating a plurality of video segments to be denoised corresponding to the video to be denoised based on the video to be denoised.

[0104] Exemplarily, the video segment to be denoised can include 8 video images, or can include 16 video images. The skilled in the art can select the number of frames included in the video segment to be denoised according to actual needs, which is not limited in the present application.

[0105] ​At step 920, the video denoising model is used to denoise the plurality of to-be-denoised video segments respectively to obtain a plurality of denoised video segments corresponding to the plurality of to-be-denoised video segments respectively.

[0106] Specifically, the plurality of to-be-denoised video segments can be sequentially input into the video denoising model to output the plurality of denoised video segments.

[0107] At step 930, the plurality of denoised video segments corresponding to the plurality of to-be-denoised video segments are spliced to generate a denoised video corresponding to the to-be-denoised video.

[0108] For example, the to-be-denoised video is a 20-frame video segment. The 1st-16th frame of the to-be-denoised video can be divided into a to-be-denoised video segment, the 2nd-17th frame can be divided into a to-be-denoised video segment, the 3rd-18th frame can be divided into a to-be-denoised video segment, the 4th-19th frame can be divided into a to-be-denoised video segment, and the 5th-20th frame can be divided into a to-be-denoised video segment, thereby obtaining five to-be-denoised video segments. Then, the five to-be-denoised video segments are input into the video denoising model respectively to obtain five denoised video segments. Finally, the five denoised video segments are spliced together to generate a denoised video corresponding to the to-be-denoised video. For the repeated video images in the five denoised video segments, the average gray value of the repeated video images can be taken as the video image in the spliced denoised video.

[0109] By dividing the to-be-denoised video into a plurality of to-be-denoised video segments, denoising the to-be-denoised video segments respectively, and then splicing the denoised video segments into a denoised video, the memory requirement of the hardware device in the video denoising process can be reduced, and the video denoising efficiency can be improved.

[0110] The method embodiments of the present application are described in detail above Figures 1-9 , and the device embodiments of the present application are described in detail below Figures 10-17 . It should be understood that the description of the method embodiments corresponds to the description of the device embodiments, and therefore, the parts not described in detail can be referred to the foregoing method embodiments.

[0111] Exemplary apparatus

[0112] Figure 10 FIG. 1 shows a structure schematic diagram of a model training device provided by an embodiment of the present application. As shown in FIG. 1, the model training device 1000 includes: Figure 10

[0113] The generating module 1010 is configured to generate a plurality of noise-increased video segments corresponding to the to-be-denoised video based on the to-be-denoised video.

[0114] ​Training module 1020 is configured to establish an initial network model and train the initial network model based on multiple noisy videos and videos to be denoised in order to generate a video denoising model. The video denoising model is used to denoise the videos to be denoised in order to generate a denoised video corresponding to the videos to be denoised.

[0115] Figure 11 The diagram shown is a structural schematic of a model training device provided in another embodiment of this application. Figure 10 This application extends from the embodiments shown. Figure 11 The illustrated embodiment will be described in detail below. Figure 11 The illustrated embodiments and Figure 10 The differences between the illustrated embodiments are not repeated here, and the similarities are not. Figure 11 As shown, the generation module 1010 includes:

[0116] The determining unit 1011 determines the noise intensity estimation information of the video to be denoised based on the video to be denoised;

[0117] The noise addition unit 1012 is configured to perform N noise addition operations on the video to be denoised based on noise intensity estimation information, so as to determine N different noise-added videos corresponding to the video to be denoised, where N is a positive integer.

[0118] In one embodiment of this application, the generation module 1010 is further configured to use a noise model to generate multiple noise-enhanced videos corresponding to the video to be denoised, wherein the multiple noise-enhanced videos contain different noise intensities, and the noise model includes any of the following models: additive noise prior model, multiplicative noise prior model, and additive multiplicative composite prior model.

[0119] Figure 12 The diagram shown is a structural schematic of a model training device provided in another embodiment of this application. Figure 10 This application extends from the embodiments shown. Figure 12 The illustrated embodiment will be described in detail below. Figure 12 The illustrated embodiments and Figure 10 The differences between the illustrated embodiments are not repeated here, and the similarities are not. Figure 12 As shown, the determining unit 1011 includes:

[0120] The image determination subunit 1021 is configured to determine at least one frame of video image corresponding to the video to be denoised based on the video to be denoised.

[0121] The region determination subunit 1022 is configured to generate multiple image regions corresponding to at least one frame of video image based on at least one frame of video image;

[0122] The calculation subunit 1023 is configured to calculate the gray-level mean and gray-level variance of each of the multiple image regions.

[0123] The information determining sub-unit 1024 is configured to determine the noise intensity estimation information based on the respective gray mean and gray variance of the plurality of image regions.

[0124] Figure 13 Fig. 8 shows a structural schematic diagram of a model training apparatus provided by another embodiment of the present application. The model training apparatus shown in Fig. 8 is extended on the basis of the model training apparatus shown in Fig. 7. Figure 12 Fig. 8 shows a structural schematic diagram of a model training apparatus provided by another embodiment of the present application. The model training apparatus shown in Fig. 8 is extended on the basis of the model training apparatus shown in Fig. 7. Figure 13 The differences between the embodiment shown in Fig. 8 and the embodiment shown in Fig. 7 will be described below, and the same parts will not be described again. As shown in Fig. 8, the model training apparatus includes: Figure 13 The differences between the embodiment shown in Fig. 8 and the embodiment shown in Fig. 7 will be described below, and the same parts will not be described again. As shown in Fig. 8, the model training apparatus includes: Figure 12 The differences between the embodiment shown in Fig. 8 and the embodiment shown in Fig. 7 will be described below, and the same parts will not be described again. As shown in Fig. 8, the model training apparatus includes: Figure 13 The linear regression sub-unit 1310 is configured to perform linear regression operation based on the respective gray mean and gray variance of the plurality of image regions, and determine the noise intensity estimation information of the to-be-de-noised video.

[0125]

[0126] Fig. 8 shows a structural schematic diagram of a model training apparatus provided by another embodiment of the present application. The model training apparatus shown in Fig. 8 is extended on the basis of the model training apparatus shown in Fig. 7. Figure 14 Fig. 8 shows a structural schematic diagram of a model training apparatus provided by another embodiment of the present application. The model training apparatus shown in Fig. 8 is extended on the basis of the model training apparatus shown in Fig. 7. Figure 13 The differences between the embodiment shown in Fig. 8 and the embodiment shown in Fig. 7 will be described below, and the same parts will not be described again. As shown in Fig. 8, the model training apparatus includes: Figure 14 The differences between the embodiment shown in Fig. 8 and the embodiment shown in Fig. 7 will be described below, and the same parts will not be described again. As shown in Fig. 8, the model training apparatus includes: Figure 14 The differences between the embodiment shown in Fig. 8 and the embodiment shown in Fig. 7 will be described below, and the same parts will not be described again. As shown in Fig. 8, the model training apparatus includes: Figure 13 The differences between the embodiment shown in Fig. 8 and the embodiment shown in Fig. 7 will be described below, and the same parts will not be described again. As shown in Fig. 8, the model training apparatus includes: Figure 14 The point graph establishing sub-unit 1311 is configured to establish a data point graph of the gray mean and the gray variance based on the respective gray mean and gray variance of the plurality of image regions, and the data point graph includes a plurality of data points.

[0127] The function determining sub-unit 1312 is configured to perform linear regression operation on the plurality of data points, and determine a regression straight line function.

[0128] The estimation information determining sub-unit 1313 is configured to determine the noise intensity estimation information of the to-be-de-noised video based on the regression straight line function, and the noise intensity estimation information of the to-be-de-noised video includes Poisson noise intensity, Gaussian noise mean and Gaussian noise variance.

[0129]

[0130] Fig. 8 shows a structural schematic diagram of a model training apparatus provided by another embodiment of the present application. The model training apparatus shown in Fig. 8 is extended on the basis of the model training apparatus shown in Fig. 7. Figure 15 Fig. 8 shows a structural schematic diagram of a model training apparatus provided by another embodiment of the present application. The model training apparatus shown in Fig. 8 is extended on the basis of the model training apparatus shown in Fig. 7. Figure 10 The differences between the embodiment shown in Fig. 8 and the embodiment shown in Fig. 7 will be described below, and the same parts will not be described again. As shown in Fig. 8, the model training apparatus includes: Figure 15 The differences between the embodiment shown in Fig. 8 and the embodiment shown in Fig. 7 will be described below, and the same parts will not be described again. As shown in Fig. 8, the model training apparatus includes: Figure 15 The differences between the embodiment shown in Fig. 8 and the embodiment shown in Fig. 7 will be described below, and the same parts will not be described again. As shown in Fig. 8, the model training apparatus includes: Figure 10The differences between the embodiments shown will not be described again. As Figure 15 As shown, the training module 1020 includes:

[0131] The function construction unit 1021 is configured to construct a loss function of the initial network model, where the loss function includes a time domain smoothing kernel norm regular term and a spatial domain smoothing entropy regular term.

[0132] The model generation unit 1022 is configured to train the initial network model based on the plurality of noise-increasing videos, the video to be denoised, and the loss function to generate the video denoising model.

[0133] Figure 16 As shown, an embodiment of the present application provides a structure schematic diagram of a video denoising device. As Figure 16 As shown, the video denoising device 1600 includes:

[0134] The acquisition module 1610 is configured to acquire a video denoising model, where the video denoising model is generated based on the model training method in the above embodiments.

[0135] The denoising module 1620 is configured to generate a noise-reduced video corresponding to the video to be denoised based on the video to be denoised by using the video denoising model.

[0136] Figure 17 As shown, another embodiment of the present application provides a structure schematic diagram of a video denoising device. In the present application Figure 16 Based on the embodiments shown in the present application Figure 17 The embodiments shown in the present application will be described below Figure 17 The embodiments shown in the present application are different from Figure 16 The differences between the embodiments shown will not be described again. As Figure 17 As shown, the denoising module 1620 includes:

[0137] The video segment generation unit 1621 is configured to generate a plurality of video segments to be denoised corresponding to the video to be denoised based on the video to be denoised.

[0138] The segment denoising unit 1622 is configured to respectively denoise the plurality of video segments to be denoised by using the video denoising model to obtain a plurality of noise-reduced video segments corresponding to the plurality of video segments to be denoised respectively.

[0139] The splicing unit 1623 is configured to splice the plurality of noise-reduced video segments corresponding to the plurality of video segments to be denoised respectively to generate a noise-reduced video corresponding to the video to be denoised.

[0140] Exemplary electronic device

[0141] Figure 18 As shown, an embodiment of the present application provides a structure schematic diagram of an electronic device. As Figure 18As shown, the electronic device 180 includes one or more processors 1801 and a memory 1802, and computer program instructions stored in the memory 1802 which, when executed by the processor 1801, cause the processor 1801 to perform the model training method and / or the video denoising method of any one of the above embodiments.

[0142] The processor 1801 can be a central processing unit (CPU) or other form of processing unit having data processing and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions.

[0143] The memory 1802 can include one or more computer program products, which can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory, for example, can include random access memory (RAM), cache memory, and / or the like. The non-volatile memory, for example, can include read-only memory (ROM), hard disk, flash memory, and / or the like. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 1801 can execute the program instructions to implement the steps in the model training method and / or the video denoising method of various embodiments of the present application described above and / or other desired functions.

[0144] In one example, the electronic device 180 can further include an input device 1803 and an output device 1804, which are interconnected through a bus system and / or other forms of connection mechanisms (not shown in the figure). Figure 18

[0145] In addition, the input device 1803 can further include, for example, a keyboard, a mouse, a microphone, and / or the like.

[0146] The output device 1804 can output various information to the outside, which can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and / or the like.

[0147] Of course, in order to simplify, Figure 18 In the figure, only some of the components in the electronic device 180 related to the present application are shown, and components such as buses, input / output interfaces, and / or the like are omitted. In addition to this, the electronic device 180 can further include any other appropriate components according to specific application cases.

[0148] Exemplary computer-readable storage medium

[0149] ​In addition to the above method and device, the embodiments of the present application can also be a computer program product, comprising computer program instructions, which, when executed by a processor, cause the processor to perform the steps in the model training method and / or the video denoising method of any of the above embodiments.

[0150] The computer program product can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++, etc., and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server.

[0151] In addition, the embodiments of the present application can also be a computer readable storage medium, which stores computer program instructions, which, when executed by a processor, cause the processor to perform the steps in the model training method and / or the video denoising method according to various embodiments of the present application described in the above "Exemplary Methods" section of the specification.

[0152] The computer readable storage medium can take the form of one or more combinations of any of the following: a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, an electrical, a magnetic, an optical, an electromagnetic, an infrared, or a semiconductor system, device or apparatus, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0153] The basic principles of the present application are described above in combination with specific embodiments, but it should be noted that the advantages, advantages, effects, etc. mentioned in the present application are only examples and are not limiting, and these advantages, advantages, effects, etc. cannot be considered as the necessary possession of each embodiment of the present application. In addition, the above specific details are only for the purpose of example and for the purpose of understanding, and are not limited to the above specific details, and the above specific details do not limit the present application to be necessarily implemented with the above specific details.

[0154] The block diagrams of the devices, apparatuses, equipment, systems referred to in this application are only illustrative examples and are not intended to require or imply that the connection, arrangement, configuration must be as shown in the block diagrams. These devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner as will be appreciated by those skilled in the art. Words such as "include," "contain," "have," etc. are open-ended words that are to be interpreted to mean "including but not limited to," and are to be taken in a non-limiting sense. The words "or" and "and" as used herein are to be interpreted as the word "and / or," and are to be taken in a non-limiting sense, unless the context clearly indicates otherwise. The word "such as" as used herein is to be interpreted as the phrase "such as but not limited to," and is to be taken in a non-limiting sense.

[0155] It is also to be noted that in the devices, apparatuses and methods of the present application, the various components or steps can be decomposed and / or recombined. These decompositions and / or recombinations are to be considered as equivalents of the present application.

[0156] The above description of disclosed aspects is given for illustrative purposes only and is not intended to limit the scope of the application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects without departing from the scope of the application. Thus, the present application is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0157] The above description has been given for illustrative and descriptive purposes. In addition, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.

[0158] The above description is only preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement and the like made within the spirit and principle of the present application should be included in the scope of the present application.

Claims

1. A model training method, characterized in that, The method comprises: generating a plurality of noise-enhanced videos corresponding to the to-be-denoised video based on the to-be-denoised video; establishing an initial network model and training the initial network model based on the plurality of noise-enhanced videos and the to-be-denoised video to generate a video denoising model, wherein the video denoising model is used to denoise the to-be-denoised video to generate a noise-reduced video corresponding to the to-be-denoised video; The method comprises: determining noise intensity estimation information of the to-be-denoised video based on the to-be-denoised video; performing N times of noise addition operations on the to-be-denoised video based on the noise intensity estimation information to determine N noise-enhanced videos corresponding to the to-be-denoised video, wherein N is a positive integer; The method comprises: determining at least one video image corresponding to the to-be-denoised video based on the to-be-denoised video; generating a plurality of image regions corresponding to the at least one video image based on the at least one video image; calculating the gray mean and the gray variance of each of the plurality of image regions; determining the noise intensity estimation information based on the gray mean and the gray variance of each of the plurality of image regions.

2. The model training method of claim 1, wherein, The method comprises: generating a plurality of noise-enhanced videos corresponding to the to-be-denoised video based on the to-be-denoised video using a noise model, wherein the plurality of noise-enhanced videos contain different noise intensities, and the noise model comprises any one of the following models: an additive noise prior model, a multiplicative noise prior model, and an additive-multiplicative compound prior model.

3. The model training method of claim 1, wherein, The method comprises: performing a linear regression operation based on the gray mean and the gray variance of each of the plurality of image regions to determine the noise intensity estimation information of the to-be-denoised video.

4. The model training method of claim 3, wherein, The method comprises: establishing a data point graph of the gray mean and the gray variance based on the gray mean and the gray variance of each of the plurality of image regions, wherein the data point graph comprises a plurality of data points; determining a regression straight line function by performing the linear regression operation on the plurality of data points; determining the noise intensity estimation information of the to-be-denoised video based on the regression straight line function, wherein the noise intensity estimation information of the to-be-denoised video comprises a Poisson noise intensity, a Gaussian noise mean, and a Gaussian noise variance.

5. The model training method of claim 1, wherein, The method comprises: constructing a loss function of the initial network model, wherein the loss function comprises a time domain smoothing kernel norm regular term and a spatial domain smoothing entropy regular term; training the initial network model based on the plurality of noise-enhanced videos, the to-be-denoised video, and the loss function to generate a video denoising model. 6.The model training method of any one of claims 1-5, wherein, The to-be-denoised video comprises a to-be-denoised microscopic video.

7. A model training apparatus characterized by comprising: The method comprises: The generating module is configured to generate a plurality of noise-increasing videos corresponding to the to-be-denoised video based on the to-be-denoised video; The training module is configured to establish an initial network model, and train the initial network model based on the plurality of noise-increasing videos and the to-be-denoised video to generate a video denoising model, wherein the video denoising model is used to denoise the to-be-denoised video to generate a noise-reduced video corresponding to the to-be-denoised video; The generating module is configured to generate a plurality of noise-increasing videos corresponding to the to-be-denoised video based on the to-be-denoised video; The generating module is configured to generate a plurality of noise-increasing videos corresponding to the to-be-denoised video based on the to-be-denoised video; The generating module is configured to generate a plurality of noise-increasing videos corresponding to the to-be-denoised video based on the to-be-denoised video; The generating module is configured to generate a plurality of noise-increasing videos corresponding to the to-be-denoised video based on the to-be-denoised video; The generating module is configured to generate a plurality of noise-increasing videos corresponding to the to-be-denoised video based on the to-be-denoised video; The storage medium stores instructions, and when the instructions are executed by the processor of the electronic device, the electronic device can execute the model training method in any one of claims 1 to 6. The electronic device comprises: a processor; 8. A computer-readable storage medium, characterized in that, a memory for storing computer executable instructions; 9. An electronic device, comprising: the processor is configured to execute the computer executable instructions to implement the model training method in any one of claims 1 to 6. ​ ​ ​

Citation Information

Patent Citations

  • Training method, denoising method and device for neural network

    CN107689034A

  • Image demosaicing

    US20150215590A1