Video generation method and apparatus, device and storage medium

By training the video generation model, using the quality and motion consistency of the predicted video to determine the reward score, and combining the diffusion probability model and attention mechanism, the problems of insufficient video generation quality and motion consistency in the existing technology are solved, and high-quality and consistent video generation is achieved.

WO2025201082A1PCT designated stage Publication Date: 2025-10-02BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2025/082436
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2025-03-13
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing text-to-video generation models face challenges in generating high-quality and motion-consistent videos, especially in maintaining the continuity of the video, background consistency, and motion regularity.

Method used

By training the video generation model, the reward score is determined by using the quality of the predicted video and the motion consistency between the predicted video and the sample video. The model is optimized using reward feedback learning, combining the diffusion probability model, attention mechanism and motion prior information to improve the quality and consistency of the generated video.

Benefits of technology

The subjective and objective quality of the generated video is improved, the motion consistency of the video is enhanced, and the smoothness of the generated video and the consistency of the background are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025082436_02102025_PF_FP_ABST
    Figure CN2025082436_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a video generation method and apparatus, a device and a storage medium. The method comprises: obtaining input information for a trained video generation model, the input information at least comprising a text description of a target video to be generated; and using the trained video generation model to generate the target video on the basis of the input information, the video generation model being trained by: respectively determining a reward score for a predicted video on the basis of at least one of the quality of the predicted video and the motion consistency between the predicted video and a sample video, wherein the predicted video is generated by the video generation model under training at least on the basis of a sample text description matching the sample video; and training the video generation model on the basis of a predetermined training objective, wherein the training objective is configured to increase or maximize the reward score.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation method, device, equipment and storage medium

[0001] This application claims priority to the Chinese invention patent application entitled “Video Generation Method, Apparatus, Device and Storage Medium” filed on March 25, 2024, with application number 202410346959.4, the entire contents of which are incorporated herein by reference. Technical Field

[0002] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for video generation. Background Art

[0003] In recent years, the field of text-based visual content generation has rapidly developed. Text-to-video generation allows users to input text, and the model can automatically generate a video corresponding to the text or edit existing videos. Current methods face significant challenges in video generation continuity (consistency between the subject and background, and the ability to preserve motion patterns) and video quality (both subjective and objective quality). Summary of the Invention

[0004] In a first aspect of the present disclosure, a method for video generation is provided. The method includes: obtaining input information for a trained video generation model, the input information including at least a text description of a target video to be generated; and generating the target video based on the input information using the trained video generation model, wherein the video generation model is trained by: determining a reward score for the predicted video based on at least one of quality of the predicted video and motion consistency between the predicted video and a sample video, the predicted video being generated by the trained video generation model based on at least a sample text description matching the sample video, and training the video generation model based on a predetermined training objective, wherein the training objective is configured to increase or maximize the reward score.

[0005] In a second aspect of the present disclosure, a video generation apparatus is provided. The apparatus includes: an input information acquisition module configured to obtain input information for a trained video generation model, the input information including at least a text description of a target video to be generated; and a target video generation module configured to generate a target video based on the input information using the trained video generation model, wherein the video generation model is trained by: determining a reward score for the predicted video based on at least one of quality of the predicted video and motion consistency between the predicted video and a sample video, the predicted video being generated by the trained video generation model based at least on a sample text description matching the sample video, and training the video generation model based on a predetermined training objective, wherein the training objective is configured to increase or maximize the reward score.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the medium, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0008] In a fifth aspect of the present disclosure, a computer program product is provided, which includes a computer program, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0009] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0011] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] FIG2 shows a schematic diagram of the input and architecture of a video generation model according to some embodiments of the present disclosure;

[0013] FIG3 shows a schematic diagram of a training video generation model according to some embodiments of the present disclosure;

[0014] FIG4 shows a reward feedback learning algorithm according to some embodiments of the present disclosure;

[0015] FIG5 shows a schematic diagram of an environment in which embodiments of the present disclosure can be implemented;

[0016] FIG6 shows a schematic diagram of a process for model training according to some embodiments of the present disclosure;

[0017] FIG7 shows a block diagram of an apparatus for video generation according to some embodiments of the present disclosure; and

[0018] FIG8 illustrates a block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0019] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0020] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.

[0021] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0022] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0023] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to the electronic device, application, server or storage medium and other software or hardware that performs the operation of the technical solution of the present disclosure based on the prompt message.

[0024] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0026] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0027] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs. It typically includes an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes the input from the previous layer.

[0028] Generally speaking, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values ​​are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association between input and output (also called input-to-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. The testing stage can sometimes be integrated into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter values ​​obtained through training to determine the corresponding model output.

[0029] Reinforcement learning, also known as reinforcement learning, evaluation learning, or enhanced learning, is a machine learning technique used to describe and solve the problem of an intelligent agent learning strategies to maximize rewards or achieve specific goals during its interaction with the environment. Reinforcement learning focuses on the interaction between the agent and the environment, and its goal is generally to maximize rewards. In other words, reinforcement learning is a learning mechanism that learns how to map states to behaviors in order to maximize rewards. Such an intelligent agent needs to continuously experiment in the environment, continuously optimizing the state-behavior relationship through feedback (rewards) provided by the environment.

[0030] Reinforcement learning systems generally involve four elements: policy, reward, value, and environment or model. The following sections introduce these four elements separately.

[0031] A policy defines the actions a model should take in a given state, essentially mapping states to actions. A state refers to the state perceived by the model. Typically, the policy is the core of a reinforcement learning system, as it determines the actions to take in each state. Depending on the configuration, the policy itself can be a specific mapping or a random distribution.

[0032] Rewards define the objective of a reinforcement learning problem. At each time step, the environment sends a scalar value to the reinforcement learning system. Rewards determine how well a model performs. Therefore, reward signals are the primary factor influencing the policy. The model's task is to maximize the total reward accumulated over a period of time.

[0033] Value, or the value function, is a crucial concept in reinforcement learning. Unlike immediate rewards, a value function measures long-term benefits. It evaluates the benefits of a current action from a long-term perspective, rather than focusing solely on the immediate reward. Calculating the value function requires analyzing transitions between states.

[0034] The environment, also known as the model, is used to predict the next state and corresponding reward after a state and action are given.

[0035] Assume S is a finite state space, A is the action space for each state s∈S; p is the action from the state s in the tth step t To the t+1th step state s t+1 The state transition probability of , R is the immediate reward value obtained after the action a∈A is executed. The main goal of the model is to interact with the environment at each time step (taking the state as input) to find the optimal policy π to reach the goal while maximizing the cumulative reward (expected return) over the entire time period. The model takes the state s as input and returns the action a to be performed. At a specific time step t, the expected return Rt is the sum of the rewards from the current time step to the last time step t. When taking an action, the model chooses between relying on previous experience (exploitation) and collecting new experience (exploration) to make better decisions in the future.

[0036] Based on reinforcement learning, model training through reinforcement learning (RL) and feedback is also proposed. Such a training scheme is also called reinforcement learning from human feedback (RLHF). In the model training system of RLHF, the target model to be trained is also called the action model (actor model), which is used to map the state s in the reinforcement learning environment to the action a. The state s in reinforcement learning corresponds to the model input, and the action a corresponds to the model output. The output of the target model is the action logic, which includes the score determined by the target model for each potential action based on the input, and the action with the highest score is determined as the final selection of the target model. The reward model is configured to determine the reward score of the output estimated by the target model for the input. The reward model is used to evaluate the action a and state s of the target model, and determine the reward score based on the quality of the model output. In RLHF, it is expected to maximize the reward score.

[0037] FIG1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In environment 100, an electronic device 110 can utilize a video generation model 120 to perform a video generation task. In some implementations, the electronic device 110 can utilize the video generation model 120 to generate a target video 112 based on input information 102.

[0038] In Figure 1, electronic device 110 can be any type of device with computing capabilities, including terminal devices or server devices. The terminal device can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, media computer, multimedia tablet, personal communication system (PCS) device, personal navigation device, personal digital assistant (PDA), audio / video player, digital camera / camcorder, positioning device, television receiver, radio broadcast receiver, e-book device, gaming device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The server device can include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like.

[0039] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0040] Current text-to-image (T2I) models, trained on large-scale image-text pairs, demonstrate the ability to generate high-quality images guided by user-provided text prompts. Building on these pre-trained T2I models, personalized generation and conditional generation provide more fine-grained control over the generated images. The success achieved in image generation has been seamlessly extended to video generation, where text-to-video (T2V) models excel in generating coherent videos driven by text prompts.

[0041] However, T2V models still face the challenge of producing high-quality and motion-consistent videos. Video quality can include both technical quality and aesthetic quality. Technical quality is characterized by fewer artifacts and reduced blur, while aesthetic quality is measured subjectively by human perception of visual appeal. Motion consistency includes object consistency (the subject and background remain unchanged between frames) and motion smoothness (the motion follows physical principles).

[0042] In order to improve the quality of video generation, in an embodiment of the present disclosure, an improved video generation scheme is proposed. Specifically, input information for a trained video generation model is obtained, and the input information at least includes a text description of a target video to be generated. Using the trained video generation model, a target video is generated based on the input information, and the video generation model is trained by: determining a reward score for the predicted video based on at least one of the quality of the predicted video and the motion consistency between the predicted video and the sample video, respectively, the predicted video is generated by the video generation model being trained based on at least a sample text description that matches the sample video, and training the video generation model based on a predetermined training target, and the training target is configured to increase or maximize the reward score.

[0043] According to the solution disclosed in the present invention, the reward score of the predicted video is determined based on the quality of the predicted video and the motion consistency between the predicted video and the motion video, and the optimization direction of the video generation model is guided by reward feedback learning, thereby improving the quality and motion consistency of the generated predicted video.

[0044] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0045] FIG2 shows a schematic diagram 200 of the input and architecture of the video generation model 120 according to some embodiments of the present disclosure.

[0046] 2 , in order for the video generation model 120 to generate a video, it is necessary to obtain input information for the trained video generation model 120. The input information at least includes a text description 205 of the target video 112 to be generated.

[0047] In some embodiments, input information 102 also includes a reference video 210 used as a control condition for video generation. Reference video 210 may include multiple frames. Furthermore, a control map extractor 215 may be used to extract reference control maps from reference video 210, such as edge maps and depth maps 220. The extracted edge maps or depth maps 220 provide rich information about the three-dimensional structure of the scene, highlighting structural features and simplifying image information, thereby reducing the complexity of subsequent processing and analysis and improving computational efficiency.

[0048] In some embodiments, the video generation model 120 includes a diffusion probability model, and the input information 102 also includes noise information sampled from a noise distribution. For better understanding, the diffusion probability model will be briefly introduced below.

[0049] The diffusion probability model is a type of generative model, but its data generation process is based on a pair of Markov processes, namely the forward diffusion process and the backward denoising process. The forward diffusion process (expressed as: is the step-by-step interference data x( 0) ~q(x (0) ), through T stepwise noise addition steps x (1:T) =x1,…,x (t-1) ,x (t) ,…,x (T) , and obtain the static noise distribution x (T) ~q noise Through model training, the learned backward denoising process (expressed as: Perform the reverse process and gradually denoise the samples towards the data distribution to obtain data x (0) ~q(x (0) ). It can be seen that the backward denoising process can correspond to the desired data modeling process and finally obtain the desired data.

[0050] In some implementations, to represent the model (denoted as: p θ (x (0) )) fits to the data distribution q(x (0) ), the learning of the backward denoising process is usually achieved by optimizing the variational constraint of the log-likelihood, which can be expressed as follows:

[0051] After learning is completed, the model performing the backward denoising process is able to first noise (x (T) ) starts sampling and by using p θ (x (t-1) |x (t) ) is iterated to denoise until the desired data is obtained.

[0052] To inherit knowledge from the image domain, the first frame of the reference video 210 or the edge map or depth map 220 extracted from the reference video 210 can be introduced as a content prior to help generate more general videos. Specifically, noise information sampled from the noise distribution is added to each frame except the first frame, and the video generation model 120 uses the content prior to learn denoising. Exemplarily, the video generation model 120 can be trained using the following loss function:

[0053] Where ε is the true value noise, the diffusion probability model θ is used at time step t based on the conditional text description c p , control chart c f and the first frame v 1 The input x t Based on this, the video generation model 120 does not need to memorize the video content in the training set, but instead focuses on learning to reconstruct motion, which can achieve better generation effects with fewer training resources.

[0054] In some embodiments, the video generation model also includes an attention-based module. In the architecture 225 of the trained video generation model 120, a one-dimensional temporal attention layer / convolution layer 235 is added for each two-dimensional spatial attention layer / convolution layer 230. To further facilitate frame modeling, a spatiotemporal self-attention mechanism 240 can be employed, in which spatial and temporal relationships are jointly modeled to capture dependencies across frames. Each frame is processed by a two-dimensional spatial attention layer / convolution layer 230, and then these frame-level features are passed together to a trainable one-dimensional temporal attention layer / convolution layer 235 for frame modeling. In addition, to achieve fine-grained modeling, the spatial self-attention mechanism can be adjusted by merging spatiotemporal self-attention across frames, which can be formulated as:

[0055] in represents the token sequence of frame i, In formula (4), the features K and V of N frames are connected so that each position has the global perception of all video frames and tends to generate more consistent results.

[0056] In some embodiments, the input information 102 may also include a motion prior 245. The motion prior 245 may include residuals. To maintain consistent noise in static regions and introduce varying noise in dynamic regions, a residual-based noise prior can be used. Specifically, this can be accomplished by calculating pixel residuals between consecutive frames and then initializing the noise distribution accordingly after downsampling. This ensures that static regions exhibit consistent noise, while dynamic regions exhibit different noise patterns. Furthermore, a threshold can be used to distinguish between static and dynamic regions, thereby providing control over the smoothness of the generated video.

[0057] Alternatively or additionally, motion prior 245 can include optical flow information. To align the generated video stream with the motion depicted in the frames, a noise prior based on optical flow can be introduced. Specifically, optical flow can be calculated between consecutive frames in pixel space and then downsampled to the latent space. This can improve the consistency of the video generated by the model.

[0058] After obtaining the input information 102, the trained video generation model 120 can be used to generate a target video 112 based on the input information 102. The video generation model 120 is trained in the following manner: based on at least one of the quality of the predicted video and the motion consistency between the predicted video and the sample video, a reward score of the predicted video is determined. The predicted video is generated by the video generation model 120 being trained based on at least a sample text description that matches the sample video. The video generation model is trained based on a predetermined training objective, and the training objective is configured to increase or maximize the reward score. The sample video here can be a true value video used to guide the training process of the video generation model 120. Under a training strategy with the optimization goal of improving or maximizing the reward score, the video generation model 120 can learn the sample video and its corresponding text description, so that when given a new, unseen input, it can generate a predicted video close to the sample video.

[0059] In some embodiments, a quality reward model can be used to determine a quality reward score for a predicted video based on the quality of the predicted video. The quality of the predicted video here can include subjective quality and objective quality. The following describes how to determine a reward score for a predicted video using a reward model with reference to FIG3 , which shows a schematic diagram 300 of a training video generation model 120 according to some embodiments of the present disclosure.

[0060] As shown in FIG3 , in some embodiments, a subjective quality reward model 310 can be used to determine a first quality reward score for the predicted video 305. Subjective quality reward model 310 can be trained in a supervised manner using a training set of images or videos, including data such as the ground-truth quality reward scores for each image or video in the training set. The ground-truth quality reward scores are provided by users and are subjectively measured based on human perception of visual appeal. Subjective quality reward model 310, also known as an aesthetic scoring model, is trained using human rating data on images, with higher scores associated with higher aesthetic appeal.

[0061] In some embodiments, an objective quality reward model 315 is used to determine a second quality reward score for the predicted video 305. The objective quality reward model is configured to determine the quality reward score based on the presence of at least one quality-influencing factor in the input video. For example, quality-influencing factors may include image artifacts, noise, blur, etc. If a quality-influencing factor is present in the predicted video 305, the video score is low; otherwise, the score is high. It should be understood that more, fewer, or different objective quality-influencing factors may be provided as needed to objectively assess video quality.

[0062] In some embodiments, the loss function for quality can be expressed as a weighted sum of a subjective quality reward score and an objective quality reward score:

[0063] L quality =λ qt ReLU(b qt -R qt (v))+,λ qa ·R,L(b qa -R qa (v′)). (5)

[0064] where R qt represents the objective quality reward score of each frame, R qa represents the subjective quality reward score of each frame, b qt represents the upper bound of the objective quality reward model 315, b qa represents the upper bound of the subjective quality reward model 310, λ qt and λ qa Represent weights respectively.

[0065] Based on this, the quality reward model can guide the learning direction of the video generation model 120 and improve the training efficiency, thereby generating a predicted video 305 with better quality.

[0066] In some embodiments, a motion reward score for the predicted video 305 can be determined based on the motion consistency between the predicted video 305 and the sample video. The training objective is configured to increase or maximize the quality reward score and increase or maximize the motion reward score.

[0067] In some embodiments, sample optical flow information of the sample video and predicted optical flow information of the predicted video 305 can be determined, and a first motion reward score for the predicted video 305 can be determined based on the difference between the predicted optical flow information and the sample optical flow information. The optical flow between each frame of the training input video (i.e., the sample video before noise addition in the pre-training phase) can be used as the true value, and the optical flow between each frame of the output video (i.e., the predicted video 305) can be used as the predicted motion information. The difference between the output optical flow and the input optical flow is calculated and used as the reward score. The smaller the difference between the output optical flow and the input optical flow, the higher the reward score.

[0068] In some embodiments, a motion reward model 320 may be used to determine a first motion reward score for the predicted video 305. The motion reward model 320 is configured to determine the first motion reward score based on a difference between the predicted optical flow information and the sample optical flow information.

[0069] Based on this, by reducing or minimizing the difference between the predicted optical flow information and the sample optical flow information, a more consistent predicted video 305 can be generated.

[0070] In some embodiments, sample residual information of the sample video can be determined based on pixel differences between consecutive video frames of the sample video. Prediction residual information of the predicted video 305 is determined based on pixel differences between consecutive video frames of the predicted video 305. A second motion reward score for the predicted video 305 is determined based on the difference between the prediction residual information and the sample residual information. The residual information here can be the difference in pixel values ​​at the same location in the video frame. The difference between the prediction residual information and the sample residual information can be used as a reward score; the smaller the difference, the higher the reward score.

[0071] In some embodiments, the loss function for motion consistency may be expressed as a weighted sum of the first motion reward score and the second motion reward score:

[0072] L motion =-λ mr ·R mr (v, v′)-λ mf ·R mf (v,v′), (6)

[0073] Where v represents the sample video, v′=Decoder(x′0) represents the predicted video 305, and R mf and R mrThey represent the motion reward score based on optical flow (first motion reward score) and the motion reward score based on optical flow (second motion reward score), respectively, mf and λ mr Represent weights respectively.

[0074] Based on this, by reducing or minimizing the difference between the prediction residual information and the sample residual information, a more consistent prediction video 305 can be generated.

[0075] In some embodiments, after obtaining the reward score and calculating the loss, the total spatiotemporal reward loss can be obtained. The total spatiotemporal reward loss can be expressed as the sum of the above motion loss and mass:

[0076] FIG4 shows a reward feedback learning algorithm 400 according to some embodiments of the present disclosure. In some embodiments, the reward feedback learning algorithm shown in FIG4 can be used to optimize the video generation model 120. Continuing to refer to FIG3, specifically, a noise X can be randomly sampled. T 325, using the video generation model 120 to perform iterative reasoning to obtain X t+1 330, in this reasoning, the reasoning gradient is not back-propagated. Next, continue the reasoning to get X t 335, in this reasoning, has a gradient. Then, from X t 335 infers X0 340 and decodes X0 340 into predicted video 305. Predicted video 305 is input into the reward model to obtain a reward score. A loss is then calculated and used to optimize the video generation model 120. In this way, after optimization, the subjective aesthetic quality, objective quality, and continuity of the video are significantly improved, as shown in the predicted video display area 345 in Figure 3.

[0077] It should be understood that, in addition to the example algorithm shown in FIG4 , the video generation model 120 may be trained using any other training and gradient backpropagation algorithms. In some embodiments, during the training of the video generation model 120, the parameters of the various reward models, such as the subjective quality reward model 310, the objective quality reward model 315, and the motion reward model 320, remain unchanged. In other words, these reward models may be pre-trained or configured to support the training of the video generation model 120.

[0078] Figure 5 shows a schematic diagram of an environment 500 in which embodiments of the present disclosure can be implemented. In the environment 500 of Figure 5 , the model is generally shown to involve different stages, including a training stage 502 and an application stage 506. After the training stage is completed, there may also be a testing stage, which is not shown in the figure.

[0079] In the training phase 502, the model training system 510 is configured to perform training of the model 505 using the training data set 512. The model 505 may be, for example, the video generation model 120 in FIG1. ​​At the beginning of the training, the model may have initial parameter values. The training process is to update the parameter values ​​of the model 505 to the desired values ​​based on the training data.

[0080] In the application phase 506, the obtained model 505 has trained parameter values ​​and can be provided to the model application system 530 for use. In the application phase 506, the model 505 can be used to process the corresponding target input 532 in the actual scene and provide the corresponding target output 534. The model application system 530 can be configured to implement the electronic device 110 of Figure 1.

[0081] In Figure 5 , the model training system 510 and the model application system 530 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices may include any type of mobile, fixed, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframe computers, edge computing nodes, computing devices in cloud environments, etc.

[0082] It should be understood that the components and arrangements in the environment 500 shown in FIG5 are merely examples, and a computing system suitable for implementing the exemplary implementations described in the present disclosure may include one or more different components, other components, and / or different arrangements. For example, although shown as separate, the model training system 510 and the model application system 530 may be integrated into the same system or device. Implementations of the present disclosure are not limited in this respect.

[0083] FIG6 shows a schematic diagram of a process 600 for video generation according to some embodiments of the present disclosure. The process 600 may be implemented at the electronic device 110 of FIG1 .

[0084] In block 610 , the electronic device 110 obtains input information for a trained video generation model, where the input information includes at least a text description of a target video to be generated.

[0085] In block 620, the electronic device 110 generates a target video based on the input information using a trained video generation model, wherein the video generation model is trained by determining a reward score for the predicted video based on at least one of the quality of the predicted video and the consistency of motion between the predicted video and the sample video, wherein the predicted video is generated by the trained video generation model based on at least a sample text description that matches the sample video, and training the video generation model based on a predetermined training objective, wherein the training objective is configured to increase or maximize the reward score. The training of the video generation model can be implemented locally on the electronic device 110 or can be implemented remotely. In the case of remote implementation, the electronic device 110 can obtain the trained video generation model from the remote device for use.

[0086] In some embodiments, determining a reward score for a predicted video includes at least one of the following: determining a quality reward score for the predicted video based on the quality of the predicted video using a quality reward model; and determining a motion reward score for the predicted video based on motion consistency between the predicted video and a sample video; and wherein the training objective is configured to increase or maximize the quality reward score and increase or maximize the motion reward score.

[0087] In some embodiments, determining a quality reward score for a predicted video includes at least one of the following: determining a first quality reward score for the predicted video using a subjective quality reward model, wherein the subjective quality reward model is supervisedly trained using the following training data: a training image set or a training video set, and a true quality reward score for each image in the training image set or each video in the training video set, wherein the true quality reward score is provided by a user; and determining a second quality reward score for the predicted video using an objective quality reward model, wherein the objective quality reward model is configured to determine the quality reward score based on the presence of at least one quality influencing factor in the input video.

[0088] In some embodiments, determining the motion reward score of the predicted video includes: determining sample optical flow information of the sample video and predicted optical flow information of the predicted video; and determining a first motion reward score of the predicted video based on the difference between the predicted optical flow information and the sample optical flow information.

[0089] In some embodiments, determining the motion reward score of the predicted video includes: determining sample residual information of the sample video based on pixel differences between consecutive video frames of the sample video; determining prediction residual information of the predicted video based on pixel differences between consecutive video frames of the predicted video; and determining a second motion reward score of the predicted video based on the difference between the prediction residual information and the sample residual information.

[0090] In some embodiments, the input information further includes a reference video used as a control condition for video generation, or includes at least one of an edge map and a depth map extracted from the reference video.

[0091] In some embodiments, the video generation model also includes an attention-based module.

[0092] Figure 7 shows a block diagram of an apparatus 700 for video generation according to some embodiments of the present disclosure. Apparatus 700 may be implemented as or included in electronic device 110 of Figure 1. Each module / component in apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.

[0093] As shown in the figure, the device 700 includes an input information acquisition module 710, which is configured to obtain input information for a trained video generation model, wherein the input information includes at least a text description of a target video to be generated. The device 700 also includes a target video generation module 720, which is configured to generate a target video based on the input information using the trained video generation model, wherein the video generation model is trained by: determining a reward score for the predicted video based on at least one of the quality of the predicted video and the motion consistency between the predicted video and a sample video, wherein the predicted video is generated by the video generation model being trained based on at least a sample text description that matches the sample video, and training the video generation model based on a predetermined training objective, wherein the training objective is configured to increase or maximize the reward score.

[0094] In some embodiments, the target video generation module 720 includes a reward score determination module configured to determine a quality reward score of the predicted video based on the quality of the predicted video using a quality reward model; and to determine a motion reward score of the predicted video based on the motion consistency between the predicted video and the sample video; and wherein the training objective is configured to increase or maximize the quality reward score and increase or maximize the motion reward score.

[0095] In some embodiments, the reward score determination module includes a quality reward score determination module, which is configured to determine a first quality reward score for the predicted video using a subjective quality reward model, wherein the subjective quality reward model is supervisedly trained using the following training data: a training image set or a training video set, and a true quality reward score for each image in the training image set or each video in the training video set, wherein the true quality reward score is provided by a user; and determine a second quality reward score for the predicted video using an objective quality reward model, wherein the objective quality reward model is configured to determine the quality reward score based on the presence of at least one quality influencing factor in the input video.

[0096] In some embodiments, the reward score determination module includes a first motion reward score determination module, which is configured to determine sample optical flow information of the sample video and predicted optical flow information of the predicted video; and determine a first motion reward score of the predicted video based on the difference between the predicted optical flow information and the sample optical flow information.

[0097] In some embodiments, the reward score determination module includes a second motion reward score determination module, which is configured to determine sample residual information of the sample video based on pixel differences between consecutive video frames of the sample video; determine prediction residual information of the predicted video based on pixel differences between consecutive video frames of the predicted video; and determine a second motion reward score of the predicted video based on the difference between the prediction residual information and the sample residual information.

[0098] In some embodiments, the input information further includes a reference video used as a control condition for video generation, or includes at least one of an edge map and a depth map extracted from the reference video.

[0099] In some embodiments, the video generation model comprises a diffusion probability model, and the input information further comprises noise information sampled from a noise distribution.

[0100] In some embodiments, the video generation model also includes an attention-based module.

[0101] FIG8 shows a block diagram of an electronic device 800 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 800 shown in FIG8 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 800 shown in FIG8 may be used to implement the electronic device 110 of FIG1 or the apparatus 700 of FIG7.

[0102] As shown in FIG8 , electronic device 800 is in the form of a general-purpose computing device. Components of electronic device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processing unit 810 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 800.

[0103] The electronic device 800 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 800.

[0104] The electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG8 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0105] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 800 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 800 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0106] The input device 850 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 860 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 800 may also communicate with one or more external devices (not shown) via the communication unit 840 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 800, or with any device that allows the electronic device 800 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0107] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0108] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0109] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0110] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0111] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0112] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating a video, comprising: Obtaining input information for a trained video generation model, the input information comprising at least a text description of a target video to be generated; as well as Generate the target video based on the input information using the trained video generation model, wherein the video generation model is trained by: determining a reward score for each predicted video based on at least one of a quality of the predicted video and a motion consistency between the predicted video and a sample video, wherein the predicted video is generated by the video generation model being trained based on at least a sample text description matching the sample video, and The video generation model is trained based on a predetermined training objective, wherein the training objective is configured to increase or maximize the reward score.

2. The method of claim 1, wherein determining the reward score of the predicted video comprises at least one of the following: Determining a quality reward score for the predicted video based on the quality of the predicted video using a quality reward model; and Determining a motion reward score for the predicted video based on motion consistency between the predicted video and the sample video; and The training goal is configured to increase or maximize the quality bonus score and increase or maximize the exercise bonus score.

3. The method of claim 2, wherein determining the quality reward score of the predicted video comprises at least one of the following: Determining a first quality reward score for the predicted video using a subjective quality reward model, The subjective quality reward model is supervisedly trained using the following training data: a training image set or a training video set, and a true quality reward score for each image in the training image set or each video in the training video set, wherein the true quality reward score is provided by a user; and Determining a second quality reward score for the predicted video using an objective quality reward model, The objective quality reward model is configured to determine a quality reward score based on the presence of at least one quality influencing factor in the input video.

4. The method of claim 2, wherein determining the motion reward score of the predicted video comprises: Determining sample optical flow information of the sample video and predicted optical flow information of the predicted video; as well as A first motion reward score of the predicted video is determined based on a difference between the predicted optical flow information and the sample optical flow information.

5. The method of claim 2, wherein determining the motion reward score of the predicted video comprises: determining sample residual information of the sample video based on pixel differences between consecutive video frames of the sample video; determining prediction residual information of the predicted video based on pixel differences between consecutive video frames of the predicted video; as well as A second motion reward score for the predicted video is determined based on a difference between the prediction residual information and the sample residual information. 6 . The method according to claim 1 , wherein the input information further includes a reference video used as a control condition for video generation, or includes at least one of an edge map and a depth map extracted from the reference video.

7. The method of claim 1, wherein the video generation model comprises a diffusion probability model, and wherein the input information further comprises noise information sampled from a noise distribution.

8. The method of claim 7, wherein the video generation model further comprises an attention-based module.

9. A video generation device, comprising: An input information obtaining module is configured to obtain input information for a trained video generation model, wherein the input information at least includes a text description of a target video to be generated; as well as A target video generation module is configured to generate the target video based on the input information using the trained video generation model, wherein the video generation model is trained by: determining a reward score for each predicted video based on at least one of a quality of the predicted video and a motion consistency between the predicted video and a sample video, wherein the predicted video is generated by the video generation model being trained based on at least a sample text description matching the sample video, and The video generation model is trained based on a predetermined training objective, wherein the training objective is configured to increase or maximize the reward score.

10. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the apparatus to perform the method according to any one of claims 1 to 8 when executed by the at least one processing unit.

11. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

12. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Video generation method and device, electronic equipment and storage medium

    CN114598926A

  • Model training method and device based on video and text

    CN115240103A

  • Video generation model training method and device, equipment and storage medium

    CN117499711A

  • Video generation method, electronic equipment and computer readable storage medium

    CN117668297A

  • Generating videos using sequences of generative neural networks

    US11908180B1

Cited By

  • Model training method, video generation method, electronic equipment and storage medium

    CN120953453A