Video frame processing method and device, video frame denoising method and device

By decomposing the probability diffusion model to decompose the noise into basis and residual noise, the video frames are denoised, which solves the problem of video quality degradation and achieves the restoration of high-quality video frames.

CN116208807BActive Publication Date: 2025-09-23ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310101110.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-28
Publication Date
2025-09-23
Estimated Expiration
2043-01-28

AI Technical Summary

Technical Problem

In the prior art, video is disturbed by noise, resulting in degradation of video quality, which affects visual effects and further processing.

Method used

The noise is decomposed into basic noise shared by consecutive video frames and independent residual noise by decomposing the probability diffusion model. The diffusion model is used to add noise to the video frames, and then the high-quality video frames are restored through the denoising network.

Benefits of technology

This simplifies the denoising process, improves the quality of video frames, and makes the subsequently generated videos higher quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116208807B_ABST
    Figure CN116208807B_ABST
Patent Text Reader

Abstract

Embodiments of this specification provide a video frame processing method and apparatus, and a video frame denoising method and apparatus. The video frame processing method includes determining an initial set of video frames of a target video; determining, based on the at least two video frames and a diffusion time step, a first target noise and a second target noise for the at least two video frames at the diffusion time step; and denoising the at least two video frames based on the at least two video frames, the diffusion time step, a diffusion parameter, a decomposition parameter, and the first target noise and second target noise for the at least two video frames at the diffusion time step to obtain a target video frame set. In the process of denoising consecutive video frames, this method decomposes the noise into a shared basic noise and independent residual noise, thereby achieving a continuous set of video frames that have been denoised to have a shared noise component, making it easier to restore the continuous video frames in the subsequent denoising process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a video frame processing method and apparatus, a video frame denoising method and apparatus, a computing device, and a computer-readable storage medium. Background Art

[0002] In existing technologies, due to limitations in shooting conditions and the influence of sending, transmitting, and receiving devices, videos are often affected by noise, which degrades the video quality, affects the visual effect, and hinders further processing of the video. Therefore, video denoising is needed to improve video quality. Summary of the Invention

[0003] In view of this, embodiments of this specification provide a video frame processing method. One or more embodiments of this specification also relate to a video frame processing device, a video frame denoising method, a video frame denoising device, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.

[0004] According to a first aspect of the embodiments of this specification, a video frame processing method is provided, including:

[0005] Determining an initial video frame set of a target video, wherein the initial video frame set includes at least two video frames;

[0006] Determining, based on the at least two video frames and the diffusion time step, a first target noise and a second target noise of the at least two video frames at the diffusion time step, wherein the first target noise of the at least two video frames at the diffusion time step is the same and the second target noise is different;

[0007] Noise is added to the at least two video frames based on the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set, wherein the target video frame set includes at least two noisy video frames.

[0008] According to a second aspect of the embodiments of this specification, a video frame processing apparatus is provided, including:

[0009] a video frame determination module, configured to determine an initial video frame set of a target video, wherein the initial video frame set includes at least two video frames;

[0010] a noise determination module configured to determine, based on the at least two video frames and the diffusion time step, a first target noise and a second target noise of the at least two video frames at the diffusion time step, wherein the first target noise of the at least two video frames at the diffusion time step is the same and the second target noise is different;

[0011] The noise adding module is configured to add noise to the at least two video frames based on the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set, wherein the target video frame set includes at least two noisy video frames.

[0012] According to a third aspect of the embodiments of this specification, a video frame denoising method is provided, comprising:

[0013] Determining a set of video frames to be processed, wherein the set of video frames to be processed includes at least two noisy video frames to be processed;

[0014] Inputting the at least two noisy video frames to be processed into a denoising network of a diffusion model to obtain at least two denoised video frames to be processed,

[0015] The denoising network of the diffusion model is the denoising network of the diffusion model in the above-mentioned video frame processing method.

[0016] According to a fourth aspect of the embodiments of this specification, a video frame denoising apparatus is provided, comprising:

[0017] a noisy video frame determination module, configured to determine a set of video frames to be processed, wherein the set of video frames to be processed includes at least two noisy video frames to be processed;

[0018] The video frame denoising module is configured to input the at least two noisy video frames to be processed into the denoising network of the diffusion model to obtain at least two denoised video frames to be processed,

[0019] The denoising network of the diffusion model is the denoising network of the diffusion model in the above-mentioned video frame processing method.

[0020] According to a fifth aspect of the embodiments of this specification, there is provided a computing device, including:

[0021] memory and processor;

[0022] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned video frame processing method or video frame denoising method are implemented.

[0023] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned video frame processing method or video frame denoising method.

[0024] According to a seventh aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned video frame processing method or video frame denoising method.

[0025] One embodiment of the present specification implements a video frame processing method, including determining an initial video frame set of a target video, wherein the initial video frame set includes at least two video frames; determining, based on the at least two video frames and a diffusion time step, a first target noise and a second target noise of the at least two video frames at the diffusion time step, wherein the first target noise of the at least two video frames at the diffusion time step is the same and the second target noise is different; and adding noise to the at least two video frames based on the at least two video frames, the diffusion time step, a diffusion parameter, a decomposition parameter, the first target noise and the second target noise of the at least two video frames at the diffusion time step to obtain a target video frame set, wherein the target video frame set includes at least two noisy video frames.

[0026] Specifically, in the process of adding noise to continuous video frames in an initial video frame set, the video frame processing method decomposes the noise into a first target noise shared by the continuous video frames and an independent second target noise, so that the continuous video frames are noisy into continuous video frames with shared noise components during the diffusion process (i.e., the noise adding process), making it easier to restore the continuous video frames in the subsequent denoising process, and more likely to generate higher quality continuous video frames, thereby generating high-quality videos based on the higher quality continuous video frames. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a schematic diagram of a specific implementation scenario of a decomposition probability diffusion model training provided by an embodiment of this specification;

[0028] Figure 2 This is a flowchart of a video frame processing method provided by one embodiment of this specification;

[0029] Figure 3 This is a schematic diagram of a noise adding process of a diffusion model in a video frame processing method provided by one embodiment of this specification;

[0030] Figure 4 This is a schematic diagram of a denoising process of a diffusion model in a video frame processing method provided by one embodiment of this specification;

[0031] Figure 5 This is a specific processing process of adding noise and removing noise using a diffusion model in a video frame processing method provided by an embodiment of this specification;

[0032] Figure 6 This is a flowchart of a video frame denoising method provided by one embodiment of this specification;

[0033] Figure 7 This is a structural diagram of a video frame processing device provided by an embodiment of this specification;

[0034] Figure 8 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION

[0035] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0036] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0037] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0038] First, the terms involved in one or more embodiments of this specification are explained.

[0039] DPM: English full name, Diffusion Probabilistic Model, Chinese full name, Probabilistic Diffusion Model.

[0040] Diffusion process: The forward process of DPM; in this process, DPM adds random noise to data (such as images, video frames, etc.) through a Markov chain, and ultimately converts the data samples into Gaussian noise (such as noisy images, noisy video frames, etc.).

[0041] Denoising process: The reverse process of DPM. In this process, DPM models data generation as a denoising process, and converts Gaussian noise into data samples through repeated denoising.

[0042] DecDPM: English full name, Decomposed DPM, Chinese full name, decomposed probability diffusion model.

[0043] Base generator: Basic generator.

[0044] Res i dua l generator: residual generator.

[0045] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0046] This specification provides a video frame processing method. One or more embodiments of this specification also relate to a video frame processing apparatus, a video frame denoising method, a video frame denoising device, a computing device, a computer-readable storage medium, and a computer program, each of which is described in detail in the following embodiments.

[0047] See also Figure 1 , Figure 1 A schematic diagram of a specific implementation scenario of a decomposition probability diffusion model training provided according to an embodiment of this specification is shown.

[0048] Figure 1 The cloud-side device 102 and the terminal-side device 104 are included, wherein the cloud-side device 102 can be understood as a cloud server. Of course, in another possible implementation scheme, the cloud-side device 102 can also be replaced by a physical server; the terminal-side device 104 includes but is not limited to a desktop computer, a laptop computer, etc. For ease of understanding, in the embodiments of this specification, the cloud-side device 102 is a cloud server and the terminal-side device 104 is a laptop computer as an example for detailed introduction.

[0049] In specific implementation, the decomposition probability diffusion model can be trained on the cloud side device 102. Figure 1 As shown, Figure 1 The decomposition probability diffusion model includes a noise adding network and a denoising network. The specific model training is as follows:

[0050] Determine consecutive initial video frames, such as 16 consecutive initial video frames: initial video frame 1, initial video frame 2...initial video frame 16; determine the diffusion time steps of the decomposition probability diffusion model, such as 1000 diffusion time steps: diffusion time step 1, diffusion time step 1...diffusion time step 1000; and determine the true noise (i.e., basic noise, residual noise), diffusion coefficient, and decomposition coefficient of each initial video frame at the corresponding diffusion time step, wherein the basic noise is shared in the consecutive initial video frames, and the residual noise is independent in the consecutive initial video frames, that is, the residual noise of each initial video frame is different.

[0051] Specifically, the real noise, diffusion coefficient, and decomposition coefficient of the continuous initial video frames and each initial video frame in the continuous initial video frames at the corresponding diffusion time step are input into the denoising network of the decomposition probability diffusion model, and the denoising process of the decomposition probability diffusion model is implemented in the denoising network to obtain continuous, noisy target video frames.

[0052] Then, the target video frame is input into the denoising network of the decomposition probability diffusion model, and the denoising process of the decomposition probability diffusion model is implemented in the denoising network to obtain the predicted noise corresponding to the continuous, noisy target video frame; finally, based on the actual noise and predicted noise of the target video frame, the noise loss function is calculated, and the decomposition probability diffusion model is trained and adjusted to obtain the trained decomposition probability diffusion model.

[0053] When the end-side device 104 needs to use the decomposition probability diffusion model, it can call the decomposition probability diffusion model trained by the cloud-side device 102 for functional use. In addition, if the computing resources and computing power of the end-side device 104 are sufficient, the decomposition probability diffusion model trained in the cloud-side device 102 can also be deployed on the end-side device 104. The specific deployment and implementation depends on the actual application and is not limited here.

[0054] The decomposition probability diffusion model provided in the embodiments of this specification decomposes the noise added to continuous initial video frames during the standard diffusion process into two parts: basic noise and residual noise, wherein the basic noise is shared between video frames; and realizes that during the diffusion process of continuous video frames, the continuous video frames are not noisily added as independent noise sequences, but rather continuous video frames with noise of shared components; and the denoising network of the decomposition probability diffusion model is trained and adjusted by combining the continuous video frames with noise of shared components with real noise, so that the decomposition probability diffusion model is easier to restore the continuous video frames in the subsequent denoising process, and is more likely to generate higher quality video.

[0055] See also Figure 2 , Figure 2 A flowchart of a video frame processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0056] Step 202: Determine an initial video frame set of the target video.

[0057] The initial video frame set includes at least two video frames.

[0058] Specifically, the target video can be understood as a video of any length, any type, and any format, such as a sports video in MPEG (Moving Picture Experts Group) format with a playback time of 30 minutes, a variety show video in AVI (Audio Video Interleaved) format with a playback time of 60 minutes, etc.

[0059] Then, the initial video frame set of the target video can be understood as the initial video frame set containing all video frames of the target video; in actual applications, the initial video frame set includes at least two video frames. For example, in the above embodiment, the initial video frame set can include 16 video frames.

[0060] Step 204: Determine a first target noise and a second target noise of the at least two video frames at the diffusion time step according to the at least two video frames and the diffusion time step.

[0061] The at least two video frames have the same first target noise and different second target noise at the diffusion time step.

[0062] Specifically, the diffusion time step can be understood as the time step in the Markov chain of the constructed diffusion model.

[0063] For example, taking at least two video frames as 16 video frames, after determining the 16 video frames and the diffusion time step, the first target noise and the second target noise for each of the 16 video frames at the diffusion time step can be determined. The first target noise can be understood as the base noise in the above embodiment, and the second target noise can be understood as the residual noise in the above embodiment. In actual applications, both the base noise and the residual noise are random Gaussian white noise, except that the base noise corresponding to different video frames is the same.

[0064] In practical applications, the diffusion model requires a sufficiently large number of diffusion steps to completely destroy the video frame signal in order to achieve a good denoising effect. Therefore, in order to achieve a better denoising effect, this can be achieved by increasing the number of diffusion time steps, for example, setting the diffusion time step to at least two or more. As described in the above embodiment, the diffusion time step can be set to 1000, etc. Then, when the diffusion time step is at least two, it is necessary to determine the basic noise and residual noise of at least two video frames at each diffusion time step. The specific implementation method is as follows:

[0065] The diffusion time steps include at least two;

[0066] Accordingly, determining the first target noise and the second target noise of the at least two video frames at the diffusion time step according to the at least two video frames and the diffusion time step includes:

[0067] Determining a first target noise and a second target noise of the at least two video frames at a target diffusion time step according to the at least two video frames and the at least two diffusion time steps,

[0068] The target diffusion time step is any one of the at least two diffusion time steps; for example, if the at least two diffusion time steps include diffusion time step 1 and diffusion time step 2, the target diffusion time step may be diffusion time step 1 or diffusion time step 2.

[0069] Continuing with the above example, if the at least two video frames include video frame 1 and video frame 2, and the diffusion time steps include diffusion time step 1 and diffusion time step 2, determining the first target noise and the second target noise of the at least two video frames at the target diffusion time step based on the at least two video frames and the at least two diffusion time steps, it can be understood that, based on video frame 1, video frame 2, diffusion time step 1, and diffusion time step 2, determining the basic noise and residual noise of video frame 1 and video frame 2 at diffusion time step 1, or determining the basic noise and residual noise of video frame 1 and video frame 2 at diffusion time step 2.

[0070] The video frame processing method provided in the embodiments of this specification increases the number of diffusion steps of the diffusion model and sets multiple diffusion time steps to add noise to the video frames, thereby greatly improving the noise addition effect of the video frames.

[0071] Step 206: Noise the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set.

[0072] The target video frame set includes at least two noisy video frames; using the above example, such as noisy video frame 1 and noisy video frame 2.

[0073] Specifically, the diffusion parameter can be understood as the diffusion coefficient. In practical applications, the diffusion coefficient in the diffusion process (i.e., the denoising process) is predetermined. For example, the diffusion coefficient can be set to (0, 1). For example, the cosine strategy can be used to pre-calculate the diffusion coefficient corresponding to each diffusion time step. That is, the diffusion coefficient for each diffusion time step can be pre-calculated according to actual needs through a fixed formula or method. The decomposition parameter can be understood as the decomposition coefficient. In practical applications, the decomposition coefficient in the diffusion process can also be understood as predetermined. For example, the decomposition coefficient can be set to [0, 1]. The larger the decomposition coefficient, the greater the proportion of basic noise in the diffusion process, that is, the more noise components are shared between video frames. Therefore, when consecutive video frames are similar, that is, the difference is small, the decomposition coefficient will be selected to be larger. If the difference between consecutive video frames is large, the decomposition coefficient will be selected to be smaller.

[0074] In a specific implementation, after determining at least two video frames to be loaded, the number of diffusion steps: diffusion time step, diffusion coefficient corresponding to each diffusion time step, and decomposition coefficient, the at least two video frames can be denoised according to the above parameters to obtain a target video frame set consisting of the at least two noisy video frames.

[0075] In order to improve the video frame noise addition effect, when the diffusion time step is set to at least two, the specific implementation method of adding noise to the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion coefficient corresponding to each diffusion time step, and the decomposition coefficient to obtain the target video frame set is as follows:

[0076] The step of adding noise to the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set includes:

[0077] Noising the target video frame according to the target video frame, the target diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the target video frame at the target diffusion time step, and the second target noise to obtain a target video frame set.

[0078] The target video frame is any one of the at least two video frames.

[0079] Continuing with the above example, when the at least two video frames are video frame 1 and video frame 2, the target video frame can be understood as video frame 1 or video frame 2.

[0080] For example, the target video frame is video frame 1 or video frame 2, and the target diffusion time step is diffusion time step 1 or diffusion time step 2; then, according to the target video frame, the target diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the target video frame at the target diffusion time step, and the second target noise, the target video frame is denoised to obtain a target video frame set. This can be understood as denoising video frame 1 according to video frame 1, diffusion time step 1, the diffusion coefficient of diffusion time step 1, the decomposition coefficient of diffusion time step 1, the basic noise of video frame 1 at diffusion time step 1, and the residual noise to obtain an initial noisy video frame 1, and further denoising the initial noisy video frame 1 according to the initial noisy video frame 1, diffusion time step 2, the diffusion coefficient of diffusion time step 2, the decomposition coefficient of diffusion time step 2, the basic noise of video frame 1 at diffusion time step 2, and the residual noise to obtain a target noisy video frame 1.

[0081] Similarly, according to the target video frame, the target diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the target video frame at the target diffusion time step, and the second target noise, the target video frame is denoised to obtain a target video frame set. This can be understood as denoising video frame 2 according to video frame 2, diffusion time step 1, the diffusion coefficient of diffusion time step 1, the decomposition coefficient of diffusion time step 1, the basic noise of video frame 2 at diffusion time step 1, and the residual noise to obtain an initial noisy video frame 2, and further denoising the initial noisy video frame 2 according to the initial noisy video frame 2, diffusion time step 2, the diffusion coefficient of diffusion time step 2, the decomposition coefficient of diffusion time step 2, the basic noise of video frame 2 at diffusion time step 2, and the residual noise to obtain a target noisy video frame 2.

[0082] Finally, a target video frame set is formed according to the target noisy video frame 1 and the target noisy video frame 2.

[0083] In practical applications, in the subsequent DecDPM denoising process, it is necessary to estimate the noise of each noisy video frame; if the decomposition coefficient of the intermediate video frame is equal to 1, it means that the residual noise of the intermediate video frame is 0, then the basic noise can be directly estimated from the intermediate video frame; therefore, in order to simplify the subsequent denoising process, the decomposition coefficient of the intermediate video frame can be set to 1; the decomposition coefficients of other video frames can be sqrt(2) / 2.

[0084] Then, when at least two video frames are divided into an intermediate video frame and other video frames, a specific implementation method of adding noise to the at least two video frames to obtain a target video frame set is as follows:

[0085] The step of adding noise to the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set includes:

[0086] Determining a middle video frame of the at least two video frames and a first decomposition parameter corresponding to the middle video frame;

[0087] Noising the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion parameter, the first decomposition parameter, the second decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set;

[0088] The second decomposition parameter is a decomposition parameter corresponding to other video frames in the at least two video frames except the middle video frame.

[0089] For example, if at least two consecutive video frames are 16 video frames, then the middle video frame can be understood as the 8th video frame; similarly, if at least two consecutive video frames are 20 video frames, then the middle video frame can be understood as the 10th video frame.

[0090] In practical applications, since the decomposition coefficients corresponding to the intermediate video frame and other video frames at different diffusion time steps are different, the diffusion process of at least two video frames is as follows:

[0091] First, an intermediate video frame, a diffusion time step, a first decomposition coefficient and a diffusion coefficient corresponding to the intermediate video frame at each diffusion time step, a basic noise and a residual noise of the intermediate video frame at each diffusion time step are determined among at least two video frames, and the intermediate video frame is denoised to obtain a noisy intermediate video frame. Similarly, other video frames, other video frames other than the intermediate video frame, the diffusion time step, the second decomposition coefficient and the diffusion coefficient corresponding to the other video frames at each diffusion time step, the basic noise and the residual noise of the other video frames at each diffusion time step are determined among at least two video frames, and the other video frames are denoised to obtain noisy other video frames. Then, a target video frame set is obtained based on the noisy intermediate video frame and the noisy other video frames.

[0092] In a specific implementation, it can be considered as inputting at least two video frames into the denoising network of the diffusion model. In the denoising network of the diffusion model, a diffusion process is implemented on the at least two video frames according to the above parameters, thereby achieving a fast and accurate denoising process for the at least two video frames. The specific implementation method is as follows:

[0093] The step of adding noise to the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set includes:

[0094] Inputting the at least two video frames into a denoising network of a diffusion model;

[0095] In the noise adding network, noise is added to the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set.

[0096] The diffusion model may be understood as the decomposed probability diffusion model of the above embodiment, and the model may include a noise adding network to implement the diffusion process and a noise removing network to implement the noise removing process.

[0097] See also Figure 3 , Figure 3 A schematic diagram of a noise adding process of a diffusion model in a video frame processing method provided by an embodiment of this specification is shown.

[0098] Figure 3 The x in the example can be understood as a continuous video frame, such as at least two video frames in the above embodiment; Z1 to Z T It can be understood as an intermediate variable in the DecDPM diffusion process, b represents the basic noise, and r represents the residual noise.

[0099] Assume x={xi |i=1, 2, ..., N) are consecutive video frames, is the intermediate variable in the DecDPM diffusion process, where N is the number of consecutive video frames, such as 16 or 18, and t = 0, 1, 2, ..., T is the number of diffusion steps, that is, the diffusion time step.

[0100] Then, after determining the above parameters, the diffusion process of the noise adding network of the diffusion model on the continuous video frames is described as formula 1:

[0101]

[0102] in,

[0103] α t ∈(0,1) is the diffusion coefficient; is the basic noise, shared in consecutive frames; is the residual noise, which is different in each frame; i ∈[0, 1] is the decomposition coefficient. In actual use, the diffusion process includes 1000 steps of noise addition, that is, T = 1000.

[0104] like Figure 3 As described above, the continuous video frame x passes through the entire diffusion process of the noise adding network of the diffusion model to obtain Z T The noisy consecutive video frames shown above.

[0105] Furthermore, in order to simplify the denoising process of the subsequent diffusion model denoising network, the decomposition coefficient of the middle video frame of the continuous video frames is set to 1 in the example of this specification, that is, in, is a Gaussian rounding function. At this time, the diffusion process of the noise-added network of the diffusion model can be expressed as the following formula 2:

[0106]

[0107] Then, after obtaining the target video frame set, the denoising network of the diffusion model can be trained based on the noisy video frames in the target video frame set. When the trained denoising network of the diffusion model is applied to the video frame denoising process, it is easier to restore continuous video frames and is more likely to generate higher quality videos. The specific implementation method is as follows:

[0108] After obtaining the target video frame set, the method further includes:

[0109] Determine a target noisy video frame from the at least two noisy video frames, and determine a first target noise and a second target noise of the target noisy video frame at the diffusion time step;

[0110] Inputting the target noisy video frame into a denoising network of a diffusion model to obtain a first predicted noise and a second predicted noise of the target noisy video frame at the diffusion time step;

[0111] The denoising network is trained according to the first target noise, the second target noise, the first predicted noise and the second predicted noise of the target noisy video frame at the diffusion time step.

[0112] The target video frame set includes at least two noisy video frames, and the target noisy video frame can be understood as any one of the at least two noisy video frames. For example, the at least two noisy video frames include noisy video frame 1 and noisy video frame 2, and the target noisy video frame can be understood as noisy video frame 1 or noisy video frame 2.

[0113] The first target noise and the second target noise of the target noisy video frame at the diffusion time step can be understood as the basic noise and the residual noise of the target noisy video frame at each diffusion time step.

[0114] Specifically, the specific training process of the denoising network of the diffusion model is as follows:

[0115] First, a target noisy video frame is determined from at least two noisy video frames, and a first target noise and a second target noise of the target noisy video frame at each diffusion time step, i.e., a real base noise and a residual noise, are determined; then, the target noisy video frame is input into a denoising network of a diffusion model for denoising, and a first predicted noise and a second predicted noise of the target noisy video at each diffusion time step, i.e., a predicted base noise and a residual noise, are obtained as output by the denoising network; finally, the denoising network of the diffusion model is trained according to the real base noise and residual noise, the predicted base noise and the residual noise of the target noisy video frame at each diffusion time step.

[0116] In practical applications, since the target noisy video frame is denoised using both base noise and residual noise, in order to achieve better denoising results, two denoising networks, the base generator and the residual generator, are used to predict the base noise and residual noise in the target noisy video frame, respectively. The specific implementation is as follows:

[0117] Inputting the target noisy video frame into the denoising network of the diffusion model to obtain the first predicted noise and the second predicted noise of the target noisy video frame at the diffusion time step includes:

[0118] Inputting the target noisy video frame into a first denoising network of a diffusion model to obtain a first predicted noise of the target noisy video frame at the diffusion time step; and

[0119] The target noisy video frame is input into the second denoising network of the diffusion model to obtain a second predicted noise of the target noisy video frame at the diffusion time step.

[0120] Among them, when the first predicted noise is understood as the predicted basic noise, the first denoising network can be understood as the basic generator; when the second predicted noise is understood as the predicted residual noise, the second denoising network can be understood as the residual generator.

[0121] In the case where the first denoising network is used as a base generator and the second denoising network is used as a residual generator, the target noisy video frame is input into the denoising network of the diffusion model to obtain the first predicted noise and the second predicted noise of the target noisy video frame at the diffusion time step. The specific implementation method is as follows:

[0122] The target noisy video frame is input into the basic generator of the diffusion model to obtain the predicted basic noise of the target noisy video frame at each diffusion time step; at the same time, the target noisy video frame is input into the residual generator of the diffusion model to obtain the predicted residual noise of the target noisy video frame at each diffusion time step.

[0123] The specific implementation in practical applications can be understood as inputting the target noisy video frame into the diffusion model, in which the base generator and the residual generator simultaneously output the predicted base noise and predicted residual noise of the target noisy video frame at each diffusion time step; subsequently, the diffusion model can be adjusted and trained based on the predicted base noise, residual noise and the actual base noise, residual noise.

[0124] Then, according to the predicted basic noise, residual noise and the actual basic noise, residual noise, the specific implementation method of adjusting and training the diffusion model is as follows:

[0125] The training of the denoising network according to the first target noise, the second target noise, the first predicted noise and the second predicted noise of the target noisy video frame at the diffusion time step comprises:

[0126] A target loss function is calculated according to the first target noise, the second target noise, the first predicted noise and the second predicted noise of the target noisy video frame at the diffusion time step, and the denoising network is trained according to the target loss function.

[0127] Specifically, after obtaining the first target noise, second target noise, first predicted noise and second predicted noise of the target noisy video frame at the diffusion time step, the corresponding target loss function can be calculated according to the first target noise, second target noise, first predicted noise and second predicted noise of the target noisy video frame at each diffusion time step, and the denoising network of the diffusion model is trained according to the target loss function to improve the subsequent video denoising effect of the diffusion model.

[0128] In a specific embodiment, the calculating of the target loss function according to the first target noise, the second target noise, the first predicted noise, and the second predicted noise of the target noisy video frame at the diffusion time step includes:

[0129] Calculating a first loss function according to a first target noise and a first predicted noise of the target noisy video frame at the diffusion time step;

[0130] Calculating a second loss function according to a second target noise and a second predicted noise of the target noisy video frame at the diffusion time step;

[0131] A target loss function is determined based on the first loss function and the second loss function.

[0132] See also Figure 4 , Figure 4 A schematic diagram of a denoising process of a diffusion model in a video frame processing method provided by an embodiment of this specification is shown.

[0133] Figure 4 The diffusion model in

[15] uses two denoising networks to predict the basic noise b t and residual noise Specific denoising process Figure 4 As shown, based on Figure 4 By performing the denoising process in , we can obtain the denoised continuous video frames.

[0134] Specifically, the diffusion model uses two denoising networks: a base generator and a residual generator, which are used to predict the base noise and residual noise respectively. The noise predicted each time can be expressed by the following formula 2:

[0135]

[0136] in, represents the mapping function of the base generator (such as the first loss function of the above embodiment), represents the mapping function of the residual generator (such as the second loss function of the above embodiment), and

[0137] Then the objective loss function of the denoising network of the diffusion model can be expressed by the following formula 3:

[0138]

[0139] The video frame processing method provided in the embodiments of this specification calculates the target loss function of the diffusion model through the mapping function of the base generator and the mapping function of the residual generator, adjusts the network parameters of the denoising network of the diffusion model through the target loss function, and obtains the denoising network of the diffusion model. This makes it easier for the denoising network of the diffusion model to restore continuous video frames in the subsequent denoising process of the noisy video frames, and is more likely to generate higher quality videos.

[0140] Specifically, the video frame processing method provided in the embodiments of this specification decomposes the noise into a first target noise shared by the continuous video frames and an independent second target noise during the process of adding noise to the continuous video frames in the initial video frame set, so that the continuous video frames are noisy into continuous video frames with shared noise components during the diffusion process (i.e., the noise adding process), making it easier to restore the continuous video frames in the subsequent denoising process, and more likely to generate higher quality continuous video frames, thereby generating high-quality videos based on the higher quality continuous video frames.

[0141] See also Figure 5 , Figure 5 The specific processing process of adding noise and removing noise of the diffusion model in a video frame processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0142] Step 502: Determine an initial video frame set of a target video, wherein the initial video frame set includes at least two video frames.

[0143] Step 504: Determine, based on the at least two video frames and the diffusion time step, a first target noise and a second target noise of the at least two video frames at the diffusion time step, wherein the first target noise of the at least two video frames at the diffusion time step is the same, and the second target noise is different.

[0144] Step 506: Noise the at least two video frames based on the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set, wherein the target video frame set includes at least two noisy video frames.

[0145] Step 508: Determine a target noisy video frame from the at least two noisy video frames, and determine a first target noise and a second target noise of the target noisy video frame at the diffusion time step.

[0146] Step 510: Input the target noisy video frame into the denoising network of the diffusion model to obtain the first predicted noise and the second predicted noise of the target noisy video frame at the diffusion time step.

[0147] Step 512: training the denoising network according to the first target noise, the second target noise, the first predicted noise and the second predicted noise of the target noisy video frame at the diffusion time step.

[0148] During specific implementation, the specific implementation process of step 502 to step 512 refers to the above embodiment, and this specification does not impose any limitation on this.

[0149] The video frame processing method provided in the embodiments of this specification decomposes the noise into a first target noise shared by the continuous video frames and an independent second target noise during the process of adding noise to the continuous video frames in the initial video frame set, so that the continuous video frames are noisy into continuous video frames with shared noise components during the diffusion process (i.e., the noise adding process), making it easier to restore the continuous video frames in the subsequent denoising process and more likely to generate higher quality continuous video frames, thereby generating high-quality video based on the higher quality continuous video frames.

[0150] In addition, a new decomposition probability diffusion model based on video frames is trained in the video frame processing method provided in one embodiment of the present specification. The decomposition probability diffusion model has a new diffusion process, which can decompose the noise in the standard video frame diffusion process into two parts: basic noise and residual noise, wherein the basic noise is shared between consecutive video frames; the decomposition probability diffusion model also provides a new denoising framework, which uses two generators (i.e., a basic generator and a residual generator) to estimate the basic noise and the residual noise respectively; so that the subsequent use of the decomposition probability diffusion model in the process of video frame denoising can make good use of the temporal correlation and redundancy of consecutive video frames, thereby greatly improving the quality and efficiency of video generation.

[0151] See also Figure 6 , Figure 6 A flowchart of a video frame denoising method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0152] Step 602: Determine a set of video frames to be processed, wherein the set of video frames to be processed includes at least two noisy video frames to be processed.

[0153] Step 604: Input the at least two noisy video frames to be processed into a denoising network of a diffusion model to obtain at least two denoised video frames to be processed.

[0154] The denoising network of the diffusion model is the denoising network of the diffusion model in the above-mentioned video frame processing method.

[0155] Specifically, the set of video frames to be processed can be understood as a set consisting of two or more consecutive noisy video frames to be processed. The video frames to be processed can also be understood as video frames generated by text. For example, a text encoder is first used to encode a text into a vector, and then the video frames to be processed are generated based on the vector. Of course, in actual applications, the generation form, format, and content of the video frames to be processed are not limited in any way in the embodiments of this specification and can be set according to actual needs.

[0156] In a specific implementation, after determining the set of video frames to be processed, at least two noisy video frames included in the set can be input into the diffusion model denoising network for denoising, thereby obtaining at least two denoised video frames to be processed. The diffusion model denoising network can be understood as the diffusion model denoising network in the video frame processing method of the above embodiment, and therefore, no further explanation of the diffusion model denoising network is provided.

[0157] The video frame denoising method provided in the embodiments of this specification implements denoising of at least two noisy video frames to be processed included in a set of video frames to be processed based on the denoising network of the diffusion model in the video frame processing method of the above embodiments. It can make good use of the temporal correlation and redundancy of consecutive video frames, and greatly improve the quality and efficiency of video generation.

[0158] Corresponding to the above method embodiment, this specification also provides an embodiment of a video frame processing device, Figure 7 FIG. 1 shows a schematic diagram of the structure of a video frame processing device provided by an embodiment of this specification. Figure 7 As shown, the device includes:

[0159] The video frame determination module 702 is configured to determine an initial video frame set of a target video, wherein the initial video frame set includes at least two video frames;

[0160] The noise determination module 704 is configured to determine, based on the at least two video frames and the diffusion time step, a first target noise and a second target noise of the at least two video frames at the diffusion time step, wherein the first target noise of the at least two video frames at the diffusion time step is the same and the second target noise is different;

[0161] The noise adding module 706 is configured to add noise to the at least two video frames based on the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise and the second target noise of the at least two video frames at the diffusion time step, to obtain a target video frame set, wherein the target video frame set includes at least two noisy video frames.

[0162] Optionally, the noise determination module 704 is further configured to:

[0163] Determining a middle video frame of the at least two video frames and a first decomposition parameter corresponding to the middle video frame;

[0164] Noising the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion parameter, the first decomposition parameter, the second decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set;

[0165] The second decomposition parameter is a decomposition parameter corresponding to other video frames in the at least two video frames except the middle video frame.

[0166] Optionally, the noise adding module 706 is further configured to:

[0167] Inputting the at least two video frames into a denoising network of a diffusion model;

[0168] In the noise adding network, noise is added to the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set.

[0169] Optionally, the device further includes:

[0170] The denoising module is configured as follows:

[0171] Determine a target video frame from the at least two noisy video frames, and determine a first target noise and a second target noise of the target video frame at the diffusion time step;

[0172] Inputting the target video frame into a denoising network of a diffusion model to obtain a first predicted noise and a second predicted noise of the target video frame at the diffusion time step;

[0173] The denoising network is trained according to the first target noise, the second target noise, the first predicted noise and the second predicted noise of the target video frame at the diffusion time step.

[0174] Optionally, the denoising module is further configured to:

[0175] Inputting the target video frame into a first denoising network of a diffusion model to obtain a first predicted noise of the target video frame at the diffusion time step; and

[0176] The target video frame is input into a second denoising network of a diffusion model to obtain a second predicted noise of the target video frame at the diffusion time step.

[0177] Optionally, the denoising module is further configured to:

[0178] A target loss function is calculated according to the first target noise, the second target noise, the first predicted noise and the second predicted noise of the target video frame at the diffusion time step, and the denoising network is trained according to the target loss function.

[0179] Optionally, the denoising module is further configured to:

[0180] Calculating a first loss function according to a first target noise and a first predicted noise of the target video frame at the diffusion time step;

[0181] Calculating a second loss function according to a second target noise and a second predicted noise of the target video frame at the diffusion time step;

[0182] A target loss function is determined based on the first loss function and the second loss function.

[0183] Optionally, the diffusion time step includes at least two;

[0184] Accordingly, the noise determination module 704 is further configured to:

[0185] Determining a first target noise and a second target noise of the at least two video frames at a target diffusion time step according to the at least two video frames and the at least two diffusion time steps,

[0186] The target diffusion time step is any diffusion time step of the at least two diffusion time steps.

[0187] Optionally, the noise determination module 704 is further configured to:

[0188] Noising the target video frame according to the target video frame, the target diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the target video frame at the target diffusion time step, and the second target noise to obtain a target video frame set.

[0189] The target video frame is any one of the at least two video frames.

[0190] The video frame processing device provided by the embodiment of this specification decomposes the noise into a first target noise shared by the continuous video frames and an independent second target noise during the process of adding noise to the continuous video frames in the initial video frame set, so that the continuous video frames are noisy into continuous video frames with shared noise components during the diffusion process (i.e., the noise adding process), making it easier to restore the continuous video frames in the subsequent denoising process and more likely to generate higher quality continuous video frames, thereby generating high-quality video based on the higher quality continuous video frames.

[0191] The above is a schematic diagram of a video frame processing device according to this embodiment. It should be noted that the technical solution of this video frame processing device and the technical solution of the aforementioned video frame processing method are based on the same concept. For details not described in detail in the technical solution of the video frame processing device, please refer to the description of the technical solution of the aforementioned video frame processing method.

[0192] Corresponding to the above method embodiment, this specification also provides an embodiment of a video frame denoising device, the device comprising:

[0193] a noisy video frame determination module, configured to determine a set of video frames to be processed, wherein the set of video frames to be processed includes at least two noisy video frames to be processed;

[0194] The video frame denoising module is configured to input the at least two noisy video frames to be processed into the denoising network of the diffusion model to obtain at least two denoised video frames to be processed,

[0195] The denoising network of the diffusion model is the denoising network of the diffusion model in the above-mentioned video frame processing method.

[0196] The video frame denoising method provided in the embodiments of this specification implements denoising of at least two noisy video frames to be processed included in a set of video frames to be processed based on the denoising network of the diffusion model in the video frame processing method of the above embodiments. It can make good use of the temporal correlation and redundancy of consecutive video frames, and greatly improve the quality and efficiency of video generation.

[0197] The above is a schematic diagram of a video frame denoising device according to this embodiment. It should be noted that the technical solution of this video frame denoising device and the technical solution of the aforementioned video frame denoising method are based on the same concept. For details not described in detail in the technical solution of the video frame denoising device, please refer to the description of the technical solution of the aforementioned video frame denoising method.

[0198] Figure 88 shows a block diagram of a computing device 800 according to one embodiment of the present disclosure. Components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and a database 850 is used to store data.

[0199] The computing device 800 also includes an access device 840 that enables the computing device 800 to communicate via one or more networks 860. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0200] In one embodiment of the present specification, the above components of the computing device 800 and Figure 8 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 8 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0201] Computing device 800 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 800 may also be a mobile or stationary server.

[0202] The processor 820 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned video frame processing method or video frame denoising method. The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above-mentioned video frame processing method or video frame denoising method belong to the same concept. For details not described in detail in the technical scheme of the computing device, please refer to the description of the technical scheme of the above-mentioned video frame processing method or video frame denoising method.

[0203] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned video frame processing method or video frame denoising method.

[0204] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solution of the video frame processing method or the video frame denoising method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the video frame processing method or the video frame denoising method described above.

[0205] An embodiment of the present specification further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned video frame processing method or video frame denoising method.

[0206] The above is an illustrative embodiment of a computer program. It should be noted that the technical solution of this computer program is based on the same concept as the technical solution of the aforementioned video frame processing method or video frame denoising method. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the aforementioned video frame processing method or video frame denoising method.

[0207] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0208] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0209] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0210] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0211] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A video frame processing method, comprising: Determining an initial video frame set of a target video, wherein the initial video frame set includes at least two video frames; Determining, based on the at least two video frames and the diffusion time step, a first target noise and a second target noise of the at least two video frames at the diffusion time step, wherein the first target noise of the at least two video frames at the diffusion time step is the same and the second target noise is different; Noise is added to the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise and the second target noise of the at least two video frames at the diffusion time step to obtain a target video frame set, wherein the target video frame set includes at least two noisy video frames, the diffusion parameter is a diffusion coefficient, the decomposition parameter is a decomposition coefficient, and the decomposition coefficient is used to control the proportion of the first target noise.

2. The video frame processing method according to claim 1, wherein the step of adding noise to the at least two video frames based on the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise of the at least two video frames to obtain the target video frame set comprises: Determining a middle video frame of the at least two video frames and a first decomposition parameter corresponding to the middle video frame; Noising the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion parameter, the first decomposition parameter, the second decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set; The second decomposition parameter is a decomposition parameter corresponding to other video frames in the at least two video frames except the middle video frame.

3. The video frame processing method according to claim 1, wherein the step of adding noise to the at least two video frames based on the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise of the at least two video frames to obtain the target video frame set comprises: Inputting the at least two video frames into a denoising network of a diffusion model; In the noise adding network, noise is added to the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set.

4. The video frame processing method according to any one of claims 1 to 3, further comprising: after obtaining the target video frame set; Determine a target noisy video frame from the at least two noisy video frames, and determine a first target noise and a second target noise of the target noisy video frame at the diffusion time step; Inputting the target noisy video frame into a denoising network of a diffusion model to obtain a first predicted noise and a second predicted noise of the target noisy video frame at the diffusion time step; The denoising network is trained according to the first target noise, the second target noise, the first predicted noise and the second predicted noise of the target noisy video frame at the diffusion time step.

5. The video frame processing method according to claim 4, wherein inputting the target noisy video frame into a denoising network of a diffusion model to obtain a first predicted noise and a second predicted noise of the target noisy video frame at the diffusion time step comprises: Inputting the target noisy video frame into a first denoising network of a diffusion model to obtain a first predicted noise of the target noisy video frame at the diffusion time step; as well as The target noisy video frame is input into the second denoising network of the diffusion model to obtain a second predicted noise of the target noisy video frame at the diffusion time step.

6. The video frame processing method according to claim 4, wherein the training of the denoising network based on the first target noise, the second target noise, the first predicted noise, and the second predicted noise of the target noisy video frame at the diffusion time step comprises: A target loss function is calculated according to the first target noise, the second target noise, the first predicted noise and the second predicted noise of the target noisy video frame at the diffusion time step, and the denoising network is trained according to the target loss function.

7. The video frame processing method according to claim 6, wherein the calculating the target loss function based on the first target noise, the second target noise, the first predicted noise, and the second predicted noise of the target noisy video frame at the diffusion time step comprises: Calculating a first loss function according to a first target noise and a first predicted noise of the target noisy video frame at the diffusion time step; Calculating a second loss function according to a second target noise and a second predicted noise of the target noisy video frame at the diffusion time step; A target loss function is determined based on the first loss function and the second loss function.

8. The video frame processing method according to claim 1 or 2, wherein the diffusion time step comprises at least two; Accordingly, determining the first target noise and the second target noise of the at least two video frames at the diffusion time step according to the at least two video frames and the diffusion time step includes: Determining a first target noise and a second target noise of the at least two video frames at a target diffusion time step according to the at least two video frames and the at least two diffusion time steps, The target diffusion time step is any diffusion time step of the at least two diffusion time steps.

9. The video frame processing method according to claim 8, wherein the step of adding noise to the at least two video frames based on the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise of the at least two video frames to obtain the target video frame set comprises: Noising the target video frame according to the target video frame, the target diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the target video frame at the target diffusion time step, and the second target noise to obtain a target video frame set. The target video frame is any one of the at least two video frames.

10. A video frame processing device, comprising: a video frame determination module, configured to determine an initial video frame set of a target video, wherein the initial video frame set includes at least two video frames; a noise determination module configured to determine, based on the at least two video frames and the diffusion time step, a first target noise and a second target noise of the at least two video frames at the diffusion time step, wherein the first target noise of the at least two video frames at the diffusion time step is the same and the second target noise is different; The noise adding module is configured to add noise to the at least two video frames according to the at least two video frames, the diffusion time step, the diffusion parameter, the decomposition parameter, the first target noise of the at least two video frames at the diffusion time step, and the second target noise to obtain a target video frame set, wherein the target video frame set includes at least two noisy video frames, the diffusion parameter is a diffusion coefficient, the decomposition parameter is a decomposition coefficient, and the decomposition coefficient is used to control the proportion of the first target noise.

11. A video frame denoising method, comprising: Determining a set of video frames to be processed, wherein the set of video frames to be processed includes at least two noisy video frames to be processed; Inputting the at least two noisy video frames to be processed into a denoising network of a diffusion model to obtain at least two denoised video frames to be processed, The denoising network of the diffusion model is the denoising network of the diffusion model in the video frame processing method according to any one of claims 5 to 7.

12. A video frame denoising device, comprising: a noisy video frame determination module, configured to determine a set of video frames to be processed, wherein the set of video frames to be processed includes at least two noisy video frames to be processed; The video frame denoising module is configured to input the at least two noisy video frames to be processed into the denoising network of the diffusion model to obtain at least two denoised video frames to be processed, The denoising network of the diffusion model is the denoising network of the diffusion model in the video frame processing method according to any one of claims 5 to 7.

13. A computing device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the video frame processing method according to any one of claims 1 to 9 or the video frame denoising method according to claim 11 are implemented.

14. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the video frame processing method according to any one of claims 1 to 9 or the video frame denoising method according to claim 11.

Citation Information

Patent Citations

  • Image processing apparatus, image processing method, and program

    CN101729760A

  • Temporal compressive sensing systems

    CN108474755A