Method, device, equipment and product for adjusting parameters of neural network model
By determining the discard ratio of video blocks during the denoising stage of the diffusion model and adjusting the neural network model, the problem of high computational complexity and difficulty in convergence in the video generation task is solved, and the quality of the generated video and the convergence speed of the model are improved.
Patent Information
- Application Number
- CN202510125076.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-06-17
AI Technical Summary
In video or image generation tasks, the diffusion model has a lower video generation quality due to high computational complexity, difficulty in convergence and limited motion amplitude of the generated results.
By determining the discarding ratio of the video block during the denoising phase of the diffusion model and discarding the video block through this proportion, the neural network model is adjusted to capture global spatiotemporal features from the finite information, reducing the amount of training data, thereby improving the convergence speed of the model and the quality of the generated video.
This method discards more video blocks during the denoising and high noise phase of the diffusion model, enhances the learning ability of the overall structure and motion mode of the video, and improves the convergence speed of the model and the quality of the generated video.
Smart Images

Figure CN120163199A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this specification generally relate to the field of computers, and more particularly to a method, apparatus, computing device, and computer program product for adjusting parameters of a neural network model. Background Art
[0002] The task of generating videos or images refers to the process of automatically creating video or image content from existing data through computer models and algorithms. Taking the video generation task as an example, it usually uses machine learning techniques to generate corresponding video sequences based on input text descriptions, images, or other content.
[0003] With the development and application of artificial intelligence, models with a large number of parameters, such as Diffusion Model, etc., are increasingly widely used in the task of generating videos or images. Since a large amount of data and complex operations need to be processed by a large model during the process of generating high-quality images or videos, high requirements are put forward for the performance and optimization of the model. Summary of the Invention
[0004] Embodiments of this specification provide a method, apparatus, computing device, and computer program product for adjusting parameters of a neural network model.
[0005] In a first aspect of this specification, a method for adjusting parameters of a neural network model is provided. The method includes: determining a first discard ratio of video blocks in a first video; discarding video blocks in the first video according to the first discard ratio to adjust the first video; determining noise in the adjusted first video through a neural network model to denoise the adjusted first video to form a second video; adjusting parameters in the neural network model based on the noise in the adjusted first video; determining a second discard ratio of video blocks in the second video, where the second discard ratio is lower than the first discard ratio; discarding video blocks in the second video according to the second discard ratio to adjust the second video; determining noise in the adjusted second video through a neural network model to denoise the adjusted second video to form a third video for the next adjustment of parameters of the neural network model; and adjusting parameters in the neural network model based on the noise in the adjusted second video.
[0006] In a second aspect of the present specification, an apparatus for adjusting parameters of a neural network model is provided. The apparatus includes: a first dropout ratio determination module configured to determine a first dropout ratio of video blocks in a first video; a first video adjustment module configured to adjust the first video by dropping video blocks in the first video according to the first dropout ratio; a first noise determination module configured to determine noise in the adjusted first video through the neural network model to denoise the adjusted first video to form a second video; a first parameter adjustment module configured to adjust parameters in the neural network model based on the noise in the adjusted first video; a second dropout ratio determination module configured to determine a second dropout ratio of video blocks in the second video, where the second dropout ratio is lower than the first dropout ratio; a second video adjustment module configured to adjust the second video by dropping video blocks in the second video according to the second dropout ratio; a second noise determination module configured to determine noise in the adjusted second video through the neural network model to denoise the adjusted second video to form a third video for adjusting parameters of the neural network model next time; and a second parameter adjustment module configured to adjust parameters in the neural network model based on the noise in the adjusted second video.
[0007] In a third aspect of the present specification, a computing device is provided. The computing device includes: a processor; and a memory coupled to the processor, the memory having instructions stored therein, which when executed by the processor, cause the computing device to execute the method according to the first aspect of the present specification.
[0008] In a fourth aspect of the present specification, a computer program product is provided, including a computer program, which when executed by a processor, implements the method according to the first aspect of the present specification.
[0009] It should be understood that the content described in the summary of the invention section is not intended to limit the key or important features of the embodiments of the present specification, nor to limit the scope of the present disclosure. Other features of the present specification will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In combination with the drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present specification will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0011] Figure 1 A schematic diagram of an example environment in which some embodiments of the present specification can be implemented is shown;
[0012] Figure 2 A schematic diagram of a process for training a neural network model according to some embodiments of the present specification is shown;
[0013] Figure 3 A schematic diagram showing the process of adding noise to a training video in some embodiments of this specification;
[0014] Figure 4 A flowchart showing a method for adjusting the parameters of a neural network model in some embodiments of this specification; and
[0015] Figure 5 A schematic block diagram of a computing device in some embodiments of this specification.
[0016] In all the drawings, the same or similar reference numerals denote the same or similar elements. Detailed Description of the Embodiments
[0017] Embodiments of this specification will be described in more detail below with reference to the accompanying drawings. Although some embodiments of this specification are shown in the drawings, it should be understood that this specification can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand this specification. It should be understood that the drawings and embodiments of this specification are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0018] In the description of the embodiments of this specification, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0019] As mentioned above, the performance and optimization of the model are crucial for video generation tasks. Models, especially generative models, have made significant progress in the computer field, particularly in the generation of images and videos. As a generative modeling method with powerful performance, the diffusion model can be applied to the generation of images and videos with high quality by defining a forward noise-adding stage and a reverse denoising stage. However, when applying the diffusion model to video generation, due to the high-dimensionality and temporal complexity of videos, the training and inference of the diffusion model face many difficulties. Among them, high computational complexity, difficulty in the convergence of the diffusion model, and limited motion amplitude of the generated results are the main problems.
[0020] To solve the above problems, improvements are made in the related art from aspects such as the diffusion model architecture, training strategy, and optimization algorithm. For example, a more massive diffusion model, a model architecture with more comprehensive attention, and the introduction of prior guidance to direct the generation process are utilized. However, while the methods in the related art improve the performance of the diffusion model, they often come with higher computational costs and more complex training processes. As a result, it is difficult for the diffusion model in the related art to converge during training, and the quality of the generated videos is low.
[0021] For this reason, embodiments of this specification provide a method for adjusting the parameters of a neural network model. In the current iteration of the denoising stage of the diffusion model, first determine the discard ratio of video blocks of the video, where this discard ratio is lower than the discard ratio applied in the previous iteration. Then, discard some video blocks in the video according to this discard ratio to adjust the video. Next, input the adjusted video into the neural network model, and predict the noise in it through the neural network model, so as to denoise the adjusted video to generate a video for the next iteration. Further, adjust the parameters in the neural network model according to the predicted noise. In this way, through multiple iterations of the above process, the training of the neural network model in the diffusion model is completed.
[0022] In this way, more video blocks can be discarded in the high-noise stage of diffusion model denoising, enabling the neural network model to capture global spatio-temporal features from limited information, enhancing the learning ability of the overall structure and motion pattern of the video. Moreover, by discarding some video blocks, the amount of data for training the neural network model can be reduced. Therefore, the embodiments of this specification can improve the convergence speed of the neural network model and the quality of the generated videos.
[0023] Figure 1 A schematic diagram of an example environment 100 in which some embodiments of this specification can be implemented is shown. Refer to Figure 1 , in environment 100, there is a computing device 102, where the computing device 102 includes but is not limited to computing chips, computing systems, single servers, distributed servers, or cloud-based servers, etc. In some embodiments, the computing device 102 can be used to perform the training and inference of the diffusion model.
[0024] In some embodiments, in the noise addition stage of diffusion model training, the computing device 102 adds noise to the sample videos in the training set multiple times to obtain the final noisy video. Then, in the denoising stage of the diffusion model, the computing device 102 denoises the final noisy video multiple times through the neural network model 108 to obtain the final denoised video.
[0025] During one iteration process, the computing device 102 can determine the discard ratio 106 (which can be used as the first discard ratio) of video blocks in the denoised video 104 (which can be used as the first video). Then, the computing device 102 discards some video blocks in the denoised video 104 according to this discard ratio 106, so as to obtain an adjusted denoised video. Further, the computing device 102 inputs the adjusted denoised video into the neural network model 108 in the diffusion model, predicts the noise in the adjusted denoised video through the neural network model 108, then generates a denoised denoised video 110 (which can be used as the second video) based on the predicted noise, and adjusts the parameters in the neural network model 108 based on the prediction result. It should be understood that the denoised video 104 can be either the final noisy video obtained after adding noise multiple times (i.e., the first iteration at the beginning of the denoising stage at this time), or the denoised video obtained after several iterations in the denoising stage (i.e., a part of the iterations in the denoising stage have been carried out at this time).
[0026] In the next iteration process, the computing device 110 can calculate the discard ratio (which can be used as the second discard ratio) of the denoised video 110. Then, the computing device 102 discards some video blocks in the denoised video 110 according to this discard ratio, so as to obtain an adjusted denoised video. Further, the computing device 102 inputs the adjusted denoised video into the neural network model 108 in the diffusion model, predicts the noise in the adjusted denoised video through the neural network model 108, then generates a denoised denoised video (which can be used as the third video) based on the predicted noise, and adjusts the parameters in the neural network model 108 based on the prediction result. Among them, the denoised video generated in the next iteration process will continue to be iterated in the next next iteration process.
[0027] In the above two iterations, the discard ratio 106 of the denoised video 104 in the previous iteration is higher than the discard ratio of the denoised video 110 in the next iteration. In this way, a higher proportion of video blocks can be discarded in the high-noise stage of the denoising stage of the diffusion model, and as the number of iterations increases, the discard ratio of video blocks is gradually reduced. In this way, more video blocks can be discarded in the high-noise stage of the diffusion model denoising, so that the neural network model 108 can capture global spatio-temporal features from limited information, enhancing the learning ability of the overall video structure and motion pattern. And by discarding some video blocks, the amount of data for training the neural network model 108 can be reduced. Therefore, the embodiments of this specification can improve the convergence speed of the neural network model 108 and improve the quality of the generated video.
[0028] It should be understood that the architecture and functions in the exemplary environment 100 are described only for exemplary purposes, without implying any limitation on the scope of this specification. Embodiments of this specification can also be applied to other environments with different structures and / or functions.
[0029] The following will be combined with Figures 2 to 5 to describe in detail the process according to an embodiment of this specification. For ease of understanding, the specific data mentioned in the following description are all exemplary and are not used to limit the protection scope of the present disclosure. It should be understood that the embodiments described below may also include additional actions not shown and / or actions shown may be omitted, and the scope of the present disclosure is not limited in this regard.
[0030] Figure 2 A schematic diagram of a process 200 for training a neural network model according to some embodiments of this specification is shown. In some embodiments, in the exemplary environment 100 shown in Figure 1 , the process 200 can be executed by the computing device 102. It should be understood that although the following content is described with the computing device 102 as the execution subject, the process 200 can also be executed by other devices. The process 200 may also include additional actions not shown and / or actions shown may be omitted, and the scope of the present disclosure is not limited in this regard.
[0031] In some embodiments, the noise addition stage of the diffusion model can be defined as q(x t |x t-1 ), where t represents the time step of iteration, 0 ≤ t ≤ T, and x t is the noisy video in the noise addition stage. The computing device 102 can obtain the training video x0 from the training set and gradually add noise to the training video x0 to obtain a noisy video x T containing random noise approximately following a Gaussian distribution. In some embodiments, the formula for the noise addition stage is shown in Equation (1):
[0032]
[0033] where N represents the Gaussian distribution, βt represents the noise scheduling parameter at time step t, and I is the identity matrix.
[0034] Figure 3 A schematic diagram of a process 300 for adding noise to a training video according to some embodiments of this specification is shown. In Figure 3Among them, the grid represents video blocks of a video frame, which may include one or more pixels. The circles represent the weights pre-calculated before adding noise (such as the weights representing contour information and the weights representing non-contour information). The shading of the grid represents the mask corresponding to the grid. Process 300 involves schematic diagrams of adding noise to video frames at different time steps, including video frame 302 of the training video, noisy video frame 304, noisy video frame 306, and noisy video frame 308. It should be understood that as the time step increases, the noise in the video frame also increases. For example, the noisy video frame 304 may be the result of adding noise after 250 time steps, the noisy video frame 306 may be the result of adding noise after 500 time steps, and the noisy video frame 308 may be the result of adding noise after 750 time steps.
[0035] Return to reference Figure 2 , in some embodiments, the denoising phase of the diffusion model can be defined as p θ (x t-1 |x t ), where θ represents the parameters of the neural network model 108 applied in the denoising phase. The noisy video x T After going through the denoising phase of T time steps, the noise in it can be restored to the original data distribution. It should be understood that the order of the time steps in the noise addition phase (from time step 0 to time step T) is opposite to that in the denoising phase (from time step T to time step 0). For example, in the noise addition phase, it is from time step (t - 1) to time step t, while in the denoising phase, it is from time step t to time step (t - 1). In some embodiments, the formula for the denoising phase is shown in Equation (2):
[0036] p θ (x t-1 |x t ) = N(x t-1 ; μ θ (x t , t), ∑ θ (x t , t)) (2)
[0037] Among them, μ θ (x t , t) and ∑ θ (x t , t) are the mean and covariance parameterized for the neural network model 108 respectively, and can be obtained by training the neural network model 108.
[0038] In some embodiments, the training phase of the diffusion model generally includes the above-mentioned noise addition phase and denoising phase. Since the noise addition phase is a parameter-free operation, the training of the diffusion model generally refers to adjusting the parameters in the neural network model 108 during the denoising phase. During the training process, the parameters of the neural network model 108 can be adjusted based on the loss between the noise predicted by the neural network model 108 and the real noise (i.e., the noise added at the corresponding time step during the noise addition phase). For example, the mean squared error between the predicted noise and the real noise is calculated and minimized to adjust the parameters in the neural network model 108. In some embodiments, the loss of the neural network model 108 can be as shown in Equation (3):
[0039]
[0040] where ε is the noise added during the noise addition phase, which follows the standard normal distribution sampling rule, and ε θ is the noise predicted by the neural network model 108.
[0041] In some embodiments, during the inference process of the diffusion model, multiple time steps of the denoising phase are required. During the training process of the diffusion model, it includes a noise addition phase with multiple time steps (corresponding to multiple iterations) and a denoising phase with multiple time steps, and during the training process, noise addition and denoising are performed globally on the video. However, at high time steps in the denoising phase (i.e., when t is close to T, at this time the denoised video x t has high noise), there will be more local noise in the denoised video x t leading to overfitting of the neural network model 108, and at low time steps in the denoising phase (i.e., when t is close to 0, at this time the denoised video x t has low noise), the neural network model 108 cannot focus on the recovery of the detailed information of the denoised video x t and may continue to modify the global structure. Therefore, this fixed noise scheduling strategy is difficult to balance the reconstruction of the video in terms of global structure and local details, resulting in unstable convergence of the neural network model 108 and insufficient detail fidelity during inference. The following embodiments will describe how to improve the first efficiency of the neural network model 108 and improve the detail fidelity during inference.
[0042] In some embodiments, in a video generation task, the training videos in the dataset are usually represented in the form of tokens (Tokens), and each token corresponds to a video block in the training video, such as a spatial position or an image block in the training video, etc. In the denoising phase of the diffusion model, for high time steps (i.e., when t is close to T, at this time the denoised video x t has high noise), a certain proportion of tokens can be randomly discarded. As the time step t decreases (the denoised video x tThe noise in also gradually decreases), and the discard ratio of the tokens is gradually reduced until no tokens are discarded at time step t = 0. The specific implementation of the above process of discarding tokens is shown in the following embodiments.
[0043] In some embodiments, at time step t in the denoising phase of the diffusion model, the computing device 102 may first calculate the denoised video x t of the video blocks with a discard ratio p t . It should be understood that the discard ratio p t decreases as the time step t decreases (i.e., as the number of iterations increases). Therefore, in some embodiments, the discard ratio p t can be calculated by Equation (4):
[0044]
[0045] where p max is the maximum discard ratio preset in the computing device 102, such as 10%, and γ is a hyperparameter that controls the decreasing speed of the discard ratio and can be adjusted according to the dataset and the training task.
[0046] In some embodiments, after determining the discard ratio p t at time step t, the computing device 102 may set the token mask vector M t corresponding to the denoised video x t . Among them, the elements in the token mask vector M t indicate whether the tokens at the corresponding positions in the denoised video x t are discarded. For example, if a token at a certain position is discarded, the corresponding element in the token mask vector M t can be represented as 0, and if the token is retained, the corresponding element can be represented as 1. If the discard ratio p t = 5%, then the element 1 in the token mask vector M t occupies 95%, and the element 0 occupies 5%. Then, the computing device 102 multiplies the token vector of the denoised video x t by this token mask vector M t to discard some tokens in the denoised video x t to obtain the adjusted denoised video x' t . The specific calculation is shown in Equation (5): t x'
[0047] = M t ⊙ x t (5) t where ⊙ represents element-wise multiplication.
[0048]
[0049] In some embodiments, after obtaining the denoised video x' t the computing device 102 may input the denoised video x' t into the neural network model 108 to predict the noise in the denoised video x' t through forward propagation of the neural network model 108, so as to denoise the denoised video x' t to obtain the denoised video x t-1 . Among them, the denoised video x t-1 is used for training the neural network model 108 at time step (t-1).
[0050] In some embodiments, after the computing device 102 predicts the noise in the denoised video x' t through the neural network model 108, the loss of the neural network model 108 may be calculated according to the predicted noise and the actual noise (i.e., the noise added in the corresponding time step during the noise addition stage). Since a part of the tokens are discarded at time step t, in order to prevent the training signal of the neural network model 108 from being too weak, in the embodiments of this specification, the computing device 102 may adaptively adjust the loss L in formula (3), so as to dynamically adjust the neural network according to the importance of the tokens and the token discard strategy, as described in the following embodiments.
[0051] In some embodiments, the device w in the computing device 102 it is the loss weight of the i-th token at time step t, which can be expressed as shown in formula (6):
[0052]
[0053] where m it is the element corresponding to the i-th token in the token mask vector M t . It should be understood that if the i-th token is discarded, then m it =0, so the corresponding loss weight w it =0. If the i-th token is retained, then m it =1, so the corresponding loss weight At this time, the loss corresponding to the i-th token is amplified. In this way, the loss of the tokens not discarded in the denoised video x t is amplified by times, thus making up for the loss of loss information caused by the discarded tokens.
[0054] Based on this, an adjusted adaptive loss function can be obtained, as shown in formula (7):
[0055]
[0056] where ε iis the actual noise corresponding to the i-th token, that is, the noise added to the i-th token during the noise addition stage, ε θ,i is the noise in the i-th token predicted by the neural network model 108. By converging the loss function L′, the parameters of the neural network model 108 can be adjusted. It should be understood that by introducing the adaptive loss function, it can be ensured that the sum of the loss weights is consistent with the original loss function, avoiding drastic fluctuations in the loss value. In this way, the computing device 102 can ensure that the tokens not discarded play a greater role in the training task by adjusting the loss weights, prevent insufficient training signals caused by token discarding, and ensure the stable convergence of the model.
[0057] In this way, in the high-noise stage of the denoising phase, due to the lower information content in the denoised video x t the computing device 102 can discard some tokens to prevent the neural network model 108 from overfitting to the noise, enabling the neural network model 108 to focus on the global structure and important features, which helps to learn a more generalizable parameter representation. At the same time, by reducing the number of tokens, the computational amount of the neural network model 108 can also be reduced, improving the training speed.
[0058] Figure 4 shows a flowchart of a method for adjusting the parameters of a neural network model according to some embodiments of the present specification. In some embodiments, in the Figure 1 example environment 100 shown, the method 400 can be executed by the computing device 102. It should be understood that although the following content is described with the computing device 102 as the execution subject, the method 400 can also be executed by other devices. The process 200 can also include additional actions not shown and / or can omit the shown actions, and the scope of the present disclosure is not limited in this regard.
[0059] At 402, determine a first discard ratio of video blocks in the first video. In some embodiments, referring to Figure 1 and Figure 2 , the first video can be the denoised video x t in the denoising phase, and the first discard ratio can be the discard ratio p t , and the computing device 102 can calculate the ratio of video blocks to be discarded in the denoised video x t , that is, the discard ratio p t . In some embodiments, the computing device 102 can calculate the discard ratio p t according to the current time step t. Alternatively or additionally, the computing device 102 can calculate the discard ratio p t according to the noise level in the denoised video x t . It should be understood that the calculation method of the discard ratio p t is not limited in this embodiment.
[0060] At 404, video blocks in the first video are discarded by a first discard ratio to adjust the first video. In some embodiments, referring to Figure 1 and Figure 2 , after determining the discard ratio p t , computing device 102 may randomly discard video blocks of the discard ratio p t in the denoised video x t , for example, randomly discard 5% of the video blocks. In some embodiments, control device 102 may directly perform a discard process on the video blocks in the denoised video x t . Alternatively or additionally, control device 102 may perform a discard process on the tokens corresponding to the video blocks in the denoised video x t .
[0061] At 406, the noise in the adjusted first video is determined by a neural network model to denoise the adjusted first video to form a second video, and the neural network model is applied in a diffusion model. In some embodiments, referring to Figure 1 and Figure 2 , in the diffusion model, denoising is performed by neural network model 108, and the second video may be the denoised video x t-1 in the denoising stage. Computing device 102 may input the denoised video x t into neural network model 108 to predict the noise in the denoised video x t by neural network model 108, and denoise the denoised video x t based on the predicted noise to form the denoised video x t-1 .
[0062] At 408, based on the noise in the adjusted first video, the parameters in the neural network model are adjusted. In some embodiments, referring to Figure 1 and Figure 2 , after predicting the noise in the denoised video x t by neural network model 108, computing device 108 adjusts the parameters in neural network model 108 based on the predicted noise. For example, computing device 108 may adjust the parameters in neural network model 108 based on the loss between the predicted noise and the actual noise to make the loss converge. It should be understood that 402 - 408 in method 400 is one iteration of training neural network model 108, and 410 - 416 is the next iteration of training neural network model 108.
[0063] At 410, a second discard ratio of the video blocks in the second video is determined, and the second discard ratio is lower than the first discard ratio. In some embodiments, referring to Figure 1 and Figure 2, the second video can be the denoised video x in the denoising stage t-1 , the second discard ratio can be the discard ratio p t-1 , the computing device 102 can compute the ratio of video blocks to be discarded in the denoised video x, that is, the discard ratio p t-1 . Among them, the discard ratio p t-1 . t-1 is lower than the discard ratio p t . In some embodiments, the computing device 102 can compute the discard ratio p according to the current time step (t - 1) t-1 . Alternatively or additionally, the computing device 102 can compute the discard ratio p according to the noise level in the denoised video x t-1 . It should be understood that there is no limitation on the computing method of the discard ratio p in this embodiment t-1 . t-1 .
[0064] At 412, the video blocks in the second video are discarded through the second discard ratio to adjust the second video. In some embodiments, refer to Figure 1 and Figure 2 , after determining the discard ratio p t-1 , the computing device 102 can randomly discard the video blocks with the discard ratio p in the denoised video x t-1 , for example, randomly discard 4% of the video blocks. In some embodiments, the control device 102 can directly perform the discard process on the video blocks in the denoised video x t-1 . Alternatively or additionally, the control device 102 can perform the discard process on the tokens corresponding to the video blocks in the denoised video x t-1 . t-1 .
[0065] At 414, the noise in the adjusted second video is determined through the neural network model to denoise the adjusted second video to form a third video for adjusting the parameters of the neural network model in the next time. In some embodiments, refer to Figure 1 and Figure 2 , the third video can be the denoised video x in the denoising stage t-2 , the denoised video x t-2 is used to adjust the neural network model 108 at the next iteration, that is, the time step (t - 2). The computing device 102 can input the denoised video x t-1 into the neural network model 108 to predict the noise in the denoised video x through the neural network model 108 t-1 , and denoise the denoised video x based on the predicted noise t-1 to form the denoised video x t-2 .
[0066] At 416, based on the noise in the adjusted second video, the parameters in the neural network model are adjusted. In some embodiments, referring to Figure 1 and Figure 2 , after the noise in the denoised video x t-1 is predicted by the neural network model 108, the computing device 108 adjusts the parameters in the neural network model 108 based on the predicted noise. For example, the computing device 108 can adjust the parameters in the neural network model 108 based on the loss between the predicted noise and the actual noise, so that the loss converges.
[0067] In this way, more video blocks can be discarded in the high-noise stage of diffusion model denoising, so that the neural network model can capture global spatio-temporal features from limited information, enhancing the learning ability of the overall structure and motion pattern of the video. And by discarding some video blocks, the amount of data for training the neural network model can be reduced. Therefore, the embodiments of this specification can improve the convergence speed of the neural network model and improve the quality of the generated video.
[0068] In some embodiments, the computing device 102 can obtain the first noise to be added to the first noisy video; add the first noise to the first noisy video to form a second noisy video; obtain the second noise to be added to the second noisy video; and add the second noise to the second noisy video to form a third noisy video as the first video.
[0069] Referring to Figure 2 , the first noisy video can be, for example, the noisy video x t-1 at the time step (t - 1) of the noise addition stage, the second noisy video can be, for example, the noisy video x t , and the third noisy video can be, for example, the noisy video x t+1 . At time step t. The computing device 102 can first determine the noise (which can be called the first noise) to be added to the noisy video x t-1 , and add the determined noise to the noisy video x t-1 , thereby forming the noisy video x t . Then, at the next time step t, the computing device 102 can first determine the noise (which can be called the second noise) to be added to the noisy video x t , and add the determined noise to the noisy video x t , thereby forming the noisy video x t+1 , and the noisy video x t+1 is used to continue adding noise at the next time step, i.e., time step (t + 1). In this way, the control device 102 can gradually add noise to the training video x0, thus ensuring the accuracy of the noise distribution.
[0070] In some embodiments, the computing device 102 may determine the time steps among a plurality of time steps that are related to the second video, where the plurality of time steps are applied to adjust the parameters in the neural network model multiple times; and determine a second discard ratio of video blocks in the second video based on the time steps.
[0071] Reference Figure 2 , in each time step of the denoising phase of the diffusion model, the computing device 102 may determine the discard ratio of the denoised video in the current time step according to the current time step (or the number of iterations that have been performed), that is, calculate the discard ratio p of the video blocks of the denoised video x according to the time step t t of the denoised video x t , for example, the discard ratio p t and the time step t may be in a proportional relationship or a quadratic function relationship. In this embodiment, for the second video, it may be the denoised video x t-1 , the computing device 102 may first determine its time step (t - 1), and then the computing device 102 may calculate the discard ratio p of the denoised video x according to the time step (t - 1) t-1 of the denoised video x t-1 (which may be referred to as the second discard ratio). In this way, the discard ratio of the denoised video can be quickly determined based on the time step, thereby improving the training efficiency of the diffusion model.
[0072] In some embodiments, the computing device 102 may obtain a preset maximum discard ratio and a hyperparameter, where the hyperparameter indicates the degree of influence of the time step on the second discard ratio; determine the ratio of the time step to the maximum time step among the plurality of time steps; and determine the second discard ratio based on the maximum discard ratio, the hyperparameter, and the ratio.
[0073] Regarding the maximum discard ratio, for example, it may be p in formula (4) max , regarding the hyperparameter, for example, it may be the hyperparameter γ in formula (4), where the hyperparameter γ can be used to adjust the degree of influence of the time step on the discard ratio, that is, the decreasing speed of the discard ratio during the process of the time step from the maximum time step T to the minimum time step 0. Regarding the ratio, for example, it may be in each time step, after determining the above maximum discard ratio p max , hyperparameter γ, ratio , the computing device 102 may calculate the discard ratio p according to the determined values t . For the second discard ratio in this embodiment, it may correspond to the time step (t - 1), so substituting the time step (t - 1) into formula (4) can obtain the discard ratio p t-1 that is, the second discard ratio.
[0074] In some embodiments, the controller 102 may generate a token mask vector associated with a plurality of tokens based on a second discard ratio, where the elements in the token mask vector indicate whether the corresponding tokens are discarded; and multiply the token mask vector by the vector of the plurality of tokens to discard tokens at the second discard ratio.
[0075] As described above, in a video generation task, the training videos in a dataset are usually represented in the form of tokens, and each token corresponds to a video patch in the training video, such as a spatial position or an image patch in the training video, etc. Therefore, in this embodiment, referring to Figure 2 , the computing device 102 may determine the denoised video x t at each time step t in the denoising phase, and determine the discard ratio p t of the denoised video x t . Then, according to the discard ratio p t , the computing device 102 sets the corresponding token mask vector M t for the denoised video x t . Wherein, the elements in the token mask vector M t indicate whether the tokens at the corresponding positions in the denoised video x t are discarded. Then, the computing device 102 multiplies the token vector of the denoised video x t by the token mask vector M t to discard some tokens in the denoised video x t to obtain an adjusted denoised video x' t-1 . The specific calculation method can be as shown in Equation (5). In this embodiment, for the second discard ratio, it may be the discard ratio p t-1 of the denoised video x . In this way, the data volume can be reduced based on the form of tokens, thereby improving the training efficiency of the neural network model 108.
[0076]
[0077] In some embodiments, the computing device 102 may obtain the actual noise associated with the second video; determine the loss of the neural network model based on the actual noise and the noise in the determined adjusted second video; and adjust the parameters in the neural network model based on the loss. Figure 2 Referring to t , at each time step t in the denoising phase, after discarding some video patches from the denoised video x t to obtain the denoised video x', the computing device 102 may input the denoised video x' t into the neural network model 108 to predict the noise in the denoised video x' t through forward propagation of the neural network model 108, so as to denoise the denoised video x' t to obtain the denoised video x t-1Among them, the denoised video x t-1 is used for the training of the neural network model 108 in the time step (t-1). Then, the computing device 102 can calculate the loss of the neural network model 108 based on the predicted noise and the actual noise (i.e., the noise added in the corresponding time step during the noise addition stage). Further, the computing device 102 adjusts the parameters in the neural network model 108 to make the loss converge. For the time step (t-1), the computing device 102 can input the denoised video x t-1 (i.e., the second video) into the neural network model 108, so as to repeat the process of training the neural network model 108 as described above.
[0078] In some embodiments, the actual noise includes the actual local noise of the video blocks of the second video, and the computing device 102 can determine the local loss of the video blocks of the second video based on the actual local noise and the local noise; determine the weights of the video blocks not discarded in the second video based on the second discard ratio; and determine the loss of the neural network model based on the local losses and weights corresponding to the non-discarded video blocks.
[0079] Reference Figure 2 , in each time step t of the denoising stage, the video block i corresponds to the token i, and the actual local noise of the video block can be, for example, the actual noise ε corresponding to the i-th token in Equation (7) i , that is, the noise added to the i-th token during the noise addition stage. The local noise can be the noise ε predicted by the neural network model 108 in the i-th token θ,i . After the computing device 102 determines the actual noise ε i corresponding to the i-th token and the noise ε θ,i , it can calculate the local loss corresponding to the video block of the token i according to the actual noise ε i and the noise ε θ,i .
[0080] Then, the computing device 102 can calculate the weight of the local loss of the video block according to the discard ratio p t of the video block in the denoised video x t in the time step i, such as the loss weight w it in Equation (7). Further, the computing device 102 can calculate the loss of the neural network model 108 according to the sum of the products of the local losses of all video blocks and the corresponding weights. For example, the computing device 102 can calculate the loss of the neural network model 108 in the manner shown in Equation (7). In this way, the computing device 102 can determine the loss weight of the video block according to the discard ratio of the video block, and then adjust the loss of the neural network model 108 to ensure that the sum of the loss weights is consistent with the original loss function, avoid drastic fluctuations in the loss value, and thus ensure the stable convergence of the model.
[0081] In some embodiments, the computing device 102 may determine the proportion of video blocks in the second video that are not discarded based on the second discard ratio; and determine the weight based on the reciprocal of the proportion of the non-discarded video blocks. Refer to Figure 2 , at each time step t in the denoising phase, the denoised video x t has a discard ratio of p t , then the proportion of the non-discarded video blocks is (1 - p t ). In this way, the computing device 102 may determine the weight corresponding to the video block based on the reciprocal of the proportion of the non-discarded video blocks (1 - p t ) according to the method shown in Equation (6). In this way, the loss of the non-discarded video blocks in the denoised video x t is amplified by times, thus compensating for the loss of information caused by the discarded tokens.
[0082] In some embodiments, the computing device 102 may determine the noise in the first inference video through a diffusion model to denoise the first inference video to form a second inference video; and determine the noise in the second inference video through the diffusion model to denoise the second inference video to form a third inference video for the next denoising. Refer to Figure 2 , after T time steps, the diffusion model completes the training process. At this time, the neural network model 108 may be directly used in the reverse phase of the diffusion model to perform video generation tasks. After the initial video is input into the diffusion model, the initial video will also go through the denoising phase of T time steps, thereby gradually denoising to form the final generated video. Among them, the first inference video, the second inference video, and the third inference video are all denoised videos generated in the denoising phase, which are similar to the above-mentioned denoised videos and will not be elaborated in this specification.
[0083] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments may be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the apparatus embodiments, the computing device embodiments, and the computer program product embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts may refer to the partial description of the method embodiments.
[0084] Figure 5 FIG. schematically shows a block diagram of a computing device 500 suitable for implementing the embodiments of the present invention. The computing device 500 may be used to implement the computing device 102. As Figure 5As shown, computing device 500 includes a processing unit (CPU) 501, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 502 or computer program instructions loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the computing device 500 can also be stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0085] A plurality of components in the computing device 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, and a storage unit 508. The processing unit 501 executes the various methods and processes described above, such as executing method 400. For example, in some embodiments, the various processes or operations described above can be implemented as a computer software program, which is stored in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the computing device 500 via the ROM 502 and / or a communication unit 509. When the computer program is loaded into the RAM 503 and executed by the CPU 501, the various methods and processes described above can be executed, such as one or more operations of executing method 400. Alternatively, in other embodiments, the CPU 501 can be configured to execute the various methods and processes described above, such as one or more actions of executing method 400, in any other suitable manner (e.g., by means of firmware).
[0086] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0087] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages - such as Smalltalk, C++, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., using an Internet service provider to connect through the Internet). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit may execute the computer-readable program instructions to implement various aspects of the present invention.
[0088] These computer-readable program instructions may be provided to the processing unit of a processor, general purpose computer, special purpose computer, or other programmable data processing apparatus in a voice interaction device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing apparatus, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing apparatus, and / or other devices to operate in a specific manner.
[0089] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0090] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
[0091] The above are only alternative embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for adjusting parameters of a neural network model, comprising: Determining a first discarding ratio of video blocks in the first video; Adjusting the first video by discarding video blocks in the first video according to the first discard ratio; Determining noise in the adjusted first video by using the neural network model to denoise the adjusted first video to form a second video, wherein the neural network model is applied to the diffusion model; Adjusting parameters in the neural network model based on the adjusted noise in the first video; Determine a second discarding ratio of video blocks in the second video, wherein the second discarding ratio is lower than the first discarding ratio; Adjusting the second video by discarding video blocks in the second video according to the second discard ratio; Determine the noise in the adjusted second video by using the neural network model to denoise the adjusted second video to form a third video for adjusting the parameters of the neural network model next time; as well as Based on the noise in the adjusted second video, parameters in the neural network model are adjusted.
2. The method according to claim 1, further comprising: Obtaining a first noise for adding to a first noisy video; Adding the first noise to the first noisy video to form a second noisy video; Acquire a second noise for adding to the second noisy video; as well as The second noise is added to the second noisy video to form a third noisy video as the first video.
3. The method according to claim 1 or 2, wherein determining the second discard ratio comprises: Determining a time step associated with the second video from a plurality of time steps, wherein the plurality of time steps are used to adjust parameters in the neural network model multiple times; as well as Based on the time step, the second discard ratio of video blocks in the second video is determined.
4. The method according to claim 3, wherein determining the second discard ratio comprises: Obtaining a preset maximum drop ratio and a hyperparameter, wherein the hyperparameter indicates the degree of influence of the time step on the second drop ratio; determining a ratio of the time step to a maximum time step among the plurality of time steps; as well as The second discard ratio is determined based on the maximum discard ratio, the hyperparameter, and the ratio.
5. The method of claim 1, wherein the plurality of video blocks in the second video correspond to a plurality of tokens, and discarding the video blocks in the second video comprises: Based on the second discard ratio, generating a token mask vector associated with the plurality of tokens, an element in the token mask vector indicating whether a corresponding token is discarded; as well as The token mask vector is multiplied by the vector of the plurality of tokens to discard the second discard ratio of tokens.
6. The method of claim 1, wherein adjusting the parameters of the neural network model comprises: Obtaining actual noise associated with the second video; as well as Determining a loss of the neural network model based on the actual noise and the determined noise in the adjusted second video; as well as Based on the loss, parameters in the neural network model are adjusted.
7. The method of claim 6, wherein the actual noise comprises actual local noise of a video block of the second video, the determined noise in the adjusted second video comprises local noise of the video block, and determining the loss comprises: determining a local loss of a video block of the second video based on the actual local noise and the local noise; Based on the second discard ratio, determining the weight of the video blocks that are not discarded in the second video; as well as The loss of the neural network model is determined based on the local losses and weights corresponding to the non-discarded video blocks.
8. The method of claim 7, wherein determining the weight comprises: Based on the second discard ratio, determining a ratio of video blocks in the second video that are not discarded; as well as The weight is determined based on the inverse of the proportion of video blocks that were not discarded.
9. The method according to claim 1, wherein the neural network model after adjusting parameters is used in the reverse phase of the diffusion model, and the method further comprises: Determining noise in the first inference video by using the diffusion model to denoise the first inference video to form a second inference video; as well as The noise in the second inference video is determined by the diffusion model to denoise the second inference video to form a third inference video for next denoising.
10. A device for adjusting parameters of a neural network model, comprising: A first discarding ratio determining module, configured to determine a first discarding ratio of video blocks in a first video; A first video adjustment module is configured to discard video blocks in the first video according to the first discard ratio to adjust the first video; A first noise determination module is configured to determine the noise in the adjusted first video through the neural network model to denoise the adjusted first video to form a second video; A first parameter adjustment module is configured to adjust parameters in the neural network model based on the noise in the adjusted first video; A second discarding ratio determining module is configured to determine a second discarding ratio of video blocks in the second video, wherein the second discarding ratio is lower than the first discarding ratio; A second video adjustment module is configured to discard video blocks in the second video by using the second discard ratio to adjust the second video; A second noise determination module is configured to determine the noise in the adjusted second video through the neural network model to denoise the adjusted second video to form a third video for adjusting the parameters of the neural network model next time; as well as The second parameter adjustment module is configured to adjust the parameters in the neural network model based on the noise in the adjusted second video.
11. A computing device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the computing device to perform the method according to any one of claims 1 to 9.
12. A computer program product comprising a computer program, the computer program being executed by a processor to implement the method according to any one of claims 1 to 9.