Training method and device of graph-based video model, equipment and storage medium
By acquiring image frames and various types of motion intensity data from the image-generated video model for calculation and training loss adjustment, the problem of insufficient motion modeling capability in image-generated video technology is solved, and high-quality video generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2026-03-20
AI Technical Summary
Existing image-to-video technology lacks the ability to perform motion modeling and accurately control motion intensity, resulting in low-quality videos. Furthermore, motion evaluation methods are easily affected by interference factors, making it difficult to accurately capture subtle changes in movement and limiting the realism and appeal of the videos.
By acquiring image frames and various types of motion intensity data from sample videos, a pre-defined image-generated video model is used for calculation. The training loss is determined based on the generated video, and the model parameters are adjusted to distinguish and decouple different types of motion intensity data for training.
It improves the accuracy of the image-generated video model in learning motion-related knowledge, ensuring the ability to control the motion intensity of the generated video and producing high-quality videos.
Smart Images

Figure CN119629426B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this application relate to the field of artificial intelligence technology, and in particular to a method, apparatus, device and storage medium for training a graph-generated video model. Background Technology
[0002] Image-to-video (IPV) refers to the use of artificial intelligence to transform static images into dynamic video content. It combines the latest advancements in image processing, computer vision, and deep learning, making it possible to create short videos with coherent motion from a single image. This technology is typically based on deep learning models that learn patterns in images and infer appropriate motion patterns to animate the images.
[0003] Image-to-video technology transforms static images into dynamic and engaging content by automatically adding animation effects and transitions, making it widely used in entertainment, advertising, and personal creation. For example, it can be used to create animations, game cutscenes, or movie special effects; businesses can quickly create product demonstration videos to enhance content appeal; and users can easily convert personal photos into fun video clips for social media sharing. However, in practical applications, achieving image-to-video while ensuring the quality of the generated videos has become a significant concern. Summary of the Invention
[0004] One or more embodiments of this application provide the following technical solutions:
[0005] This application provides a method for training a graph-generated video model, the method comprising:
[0006] Acquire a first sample video and extract image frames from the first sample video;
[0007] The motion intensity of the first sample video is evaluated by the trained motion estimation model, resulting in various types of motion intensity data for the first sample video.
[0008] The image frames and the various types of motion intensity data are input into a preset image-generated video model, which then calculates based on the image frames and the various types of motion intensity data to generate the corresponding video.
[0009] The training loss is determined based on the generated video, and the training of the image-generated video model is considered complete after the model parameters for the image-generated video model are adjusted according to the loss.
[0010] The application further provides a device for training a picture-to-video model, the device comprising:
[0011] a first obtaining module configured to obtain a first sample video and extract image frames from the first sample video;
[0012] a second obtaining module configured to obtain, by a trained motion estimation model, motion intensity data of multiple types of the first sample video obtained by performing motion intensity evaluation on the first sample video;
[0013] a generating module configured to input the image frames and the motion intensity data of multiple types into a preset picture-to-video model, and generate a corresponding video by the picture-to-video model based on the image frames and the motion intensity data of multiple types;
[0014] a training module configured to determine a training loss based on the generated video, and determine that the training of the picture-to-video model is completed after adjusting model parameters of the picture-to-video model according to the loss.
[0015] The application further provides an electronic device comprising:
[0016] a processor;
[0017] a memory for storing processor-executable instructions;
[0018] The processor implements the steps of the method according to any one of the above embodiments by running the executable instructions.
[0019] The application further provides a computer-readable storage medium having computer instructions stored thereon, the instructions being executed by a processor to implement the steps of the method according to any one of the above embodiments.
[0020] In the above technical solution, first, a sample video can be obtained, and image frames can be extracted from the sample video. Motion intensity data of multiple types of the first sample video obtained by performing motion intensity evaluation on the sample video by a trained motion estimation model can also be obtained. Then, the image frames and the motion intensity data of multiple types can be input into a preset picture-to-video model, and a corresponding video can be generated by the picture-to-video model based on the image frames and the motion intensity data of multiple types. Finally, a training loss can be determined based on the generated video, and the training of the picture-to-video model can be determined to be completed after adjusting model parameters of the picture-to-video model according to the loss.
[0021] In this way, on the one hand, different types of motion in the video can be distinguished, and the decoupled motion intensity data of multiple types can be injected into the video generation model for training, so that the learning effect of the video generation model based on the decoupled motion intensity data can be improved, and the accuracy of the motion-related knowledge learned by the video generation model can be ensured. On the other hand, the additional motion estimation model is used to evaluate the motion intensity of the video, which can ensure the accuracy of the various types of motion intensity data of the video obtained by evaluation, and help the video generation model better understand and generate the motion in the video. In this way, when the trained video generation model is used for video generation, the motion intensity of the generated video can be better controlled, and a high-quality video can be generated. BRIEF DESCRIPTION OF DRAWINGS
[0022] The drawings needed to be used in the description of the exemplary embodiments will be described below.
[0023] Figure 1 is a schematic diagram of a video generation system according to an exemplary embodiment of the present application;
[0024] Figure 2 is a schematic diagram of a training and use process of a video generation model according to an exemplary embodiment of the present application;
[0025] Figure 3 is a flowchart of a training method of a video generation model according to an exemplary embodiment of the present application;
[0026] Figure 4 is a flowchart of a training and use method of a motion estimation model according to an exemplary embodiment of the present application;
[0027] Figure 5 is a structural schematic diagram of a device according to an exemplary embodiment of the present application;
[0028] Figure 6 is a block diagram of a training device of a video generation model according to an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0029] The exemplary embodiments will be described in detail below with reference to the accompanying drawings. In the following description, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of the present application. Rather, they are merely examples consistent with some aspects of one or more embodiments of the present application.
[0030] It is to be noted that the steps of the respective methods are not necessarily performed in the order shown and described in this application in other embodiments. In some other embodiments, the steps included in the methods thereof can be more or less than described in this application. In addition, a single step described in this application can be broken down into multiple steps for description in other embodiments; and multiple steps described in this application can be combined into a single step for description in other embodiments.
[0031] Video from image refers to the use of artificial intelligence technology to convert static pictures into dynamic video content. It combines the latest advances in image processing, computer vision and deep learning, making it possible to create short videos with coherent actions from a single picture. This technology is usually based on deep learning models, which learn patterns in images and infer appropriate motion patterns to animate the images.
[0032] In the field of image-to-video, motion modeling capability is an important factor affecting the quality of generated videos. Motion modeling capability refers to the ability to understand and generate natural and smooth object or scene motion in the process of converting static images into dynamic videos. It involves learning the movement patterns of objects in image sequences and how to predict or generate new frames based on these patterns to create coherent and realistic video content.
[0033] Specifically, motion modeling capability usually includes motion estimation, motion prediction, deformation and animation, scene consistency, spatio-temporal consistency, etc. Among them, motion estimation refers to identifying the position changes of objects between two or more frames. Motion prediction refers to predicting the content of the next frame based on existing several frames of images, which requires the model to understand the motion trend of the object. Deformation and animation refer to non-rigid objects (such as human faces, animals), which also need to consider the changes in their shapes, and make still pictures "come to life" through appropriate algorithms. Scene consistency refers to ensuring that not only individual elements in the generated video look natural, but also the entire scene maintains consistency and logic. Spatio-temporal consistency refers to ensuring that the generated video has smooth motion on the time axis and reasonable relationships between objects in space.
[0034] However, in the current field of image-to-video, the lack of motion modeling capability is still a key problem that needs to be solved. The lack of motion modeling capability mainly manifests in the following aspects:
[0035] Motion Intensity Control: Motion intensity generally refers to the speed of movement, the degree of movement, etc. Existing video generation techniques have limitations in controlling the motion intensity of objects or scenes in generated videos. Ideally, it is desirable to automatically generate natural and smooth animations that meet expectations based on input content. However, current methods often struggle to accurately adjust the degree of such motion, resulting in generated videos that are either too stiff and lack dynamicity, or overly exaggerated and do not conform to real-world physics.
[0036] Motion Score Calculation: To evaluate and optimize the performance of motion in videos, a "motion score" is often defined to measure the intensity of motion in videos. Current methods infer motion intensity by analyzing the differences between adjacent frames in a video. However, this approach is susceptible to non-motion factors such as changes in lighting conditions, changes in background colors, or updates to the content itself. These external factors can lead to misjudgments about actual motion, affecting the quality of the final output video.
[0037] Interference Factor Processing: Although the frame difference-based approach is simple and intuitive, it faces many challenges in practical applications. For example, when the image sequence contains rapidly changing colors or patterns, these visual changes may be mistakenly interpreted as important motion features. Similarly, in cases of scene transitions or camera movements, it is difficult to distinguish between changes in viewpoint caused by camera movement and true object displacement.
[0038] Realistic Perception Matching: Humans are very sensitive to the realism of motion in videos and can easily identify unnatural or discontinuous movements. In contrast, current artificial intelligence systems have a lot of room for improvement in this area. They may not accurately capture subtle motion changes or handle complex multi-object interaction scenarios well, which limits the realism and appeal of the generated videos.
[0039] In existing technologies, for a model used for video generation (which can be referred to as a video generation model), it can learn the motion in the video during the model training process using a first sample video, and apply the learned motion-related knowledge to control the motion intensity of the generated video during subsequent model use.
[0040] However, on the one hand, since the motion in the video is relatively complex in actual application, for example, the motion in a video includes both object motion and camera motion, and the graph-based video model cannot decouple different types of motion in the learning process, which will result in low accuracy of the motion-related knowledge learned by the graph-based video model; on the other hand, since the way of calculating the motion score of the video based on the inter-frame difference has errors, it will also affect the understanding and generation of the motion in the video by the graph-based video model. This makes the motion intensity control ability of the graph-based video model for the generated video poor, and the quality of the generated video is difficult to guarantee, and the phenomenon of motion mixing is also easy to occur.
[0041] The one or more embodiments of the present application provide a training method of a graph-based video model. In the technical solution, first, a sample video can be obtained, and image frames can be extracted from the sample video. In addition, motion intensity evaluation of the sample video can be performed by a trained motion estimation model to obtain motion intensity data of the first sample video of multiple types. Then, the image frames and the motion intensity data of multiple types can be input into a preset graph-based video model. The graph-based video model can generate a corresponding video based on the image frames and the motion intensity data of multiple types. Finally, a training loss can be determined based on the generated video, and the training of the graph-based video model can be completed after adjusting the model parameters of the graph-based video model according to the loss.
[0042] In the above manner, on the one hand, different types of motion in the video can be distinguished, and decoupled motion intensity data of multiple types can be injected into the graph-based video model for training, so as to improve the learning effect of the graph-based video model based on the decoupled motion intensity data, and guarantee the accuracy of the motion-related knowledge learned by the graph-based video model. On the other hand, an additional motion estimation model is used to evaluate the motion intensity of the video, which can guarantee the accuracy of the motion intensity data of various types of the video obtained by evaluation, and help the graph-based video model to better understand and generate the motion in the video. In this way, when the trained graph-based video model is used for graph-based video generation, the motion intensity of the generated video can be well controlled, and a high-quality video can be generated.
[0043] Please refer to Figure 1 , Figure 1 is a schematic diagram of a graph-based video system according to an example embodiment of the present application.
[0044] As shown in Figure 1 , the above graph-based video system can include a server and at least one client accessing the server through any type of wired or wireless network.
[0045] The server can correspond to a server including a single independent physical host, or a server cluster including multiple independent physical hosts. Alternatively, the server can correspond to a virtual server or a cloud server carried by a host cluster.
[0046] The client can correspond to a terminal device such as a smartphone, a tablet computer, a notebook computer, a desktop computer, a PC (Personal Computer), a PDA (Personal Digital Assistant), a wearable device (for example, smart glasses, a smart watch, etc.), a smart in-vehicle device, or a game console.
[0047] The user can use the graph-to-video service provided by the graph-to-video system through the client. The client and the server can implement the graph-to-video service for the user through data interaction between them.
[0048] For example, the client can output a corresponding user interface to the user, so that the user can perform operations such as inputting an image for generating a video, inputting a description text and various types of motion intensity data (for example, a score reflecting the motion intensity) for assisting the graph-to-video, and selecting a graph-to-video service in the user interface, to use the graph-to-video service provided by the graph-to-video system. The client can send the image, the description text, and the various types of motion intensity data input by the user to the server. The server can generate a video based on the image under the guidance of the description text and the various types of motion intensity data as given conditions, and output the generated video to the user, that is, return the generated video to the client. The client can display the generated video to the user through the user interface for the user to view, thereby implementing the graph-to-video service for the user.
[0049] Specifically, the server can be deployed with a graph-to-video model, and the graph-to-video system can be based on the graph-to-video model to perform specific steps in the graph-to-video task.
[0050] Correspondingly, the server can also be deployed with a motion estimation model that can be used to train the graph-to-video model. The motion estimation model is used to evaluate various types of motion intensity data of the video. When training the graph-to-video model, the graph-to-video model can be trained based on a sample video, a description text corresponding to the sample video, and various types of motion intensity data of the sample video evaluated by the motion estimation model.
[0051] Alternatively, the server can also be deployed with a sample video database, each record in the sample video database including a sample video, description text corresponding to the sample video, and various types of motion intensity data of the sample video evaluated by the motion estimation model. When training the video generation model, the video generation model can be trained based on the corresponding sample videos, description texts, and various types of motion intensity data in the sample database together.
[0052] In addition, the server can also be deployed with other functional components or functional subsystems, such as a feature extraction component, a database retrieval component, a prompt text generation component, etc. These components or subsystems can work together with the video generation model deployed on the server to realize the video generation service.
[0053] It should be noted that the training of the video generation model described above can be directly completed on the server; alternatively, the training of the video generation model described above can be completed on other electronic devices with certain computing capabilities, and the trained video generation model can be deployed on the server; the present application does not make special limitations on this. The motion estimation model and the video generation model can be deployed on the same electronic device or on different electronic devices, and the present application does not make special limitations on this.
[0054] Please refer to Figure 2 , Figure 2 is a schematic diagram of a training and use process of a video generation model according to an example embodiment of the present application.
[0055] As shown in Figure 2 , when training the video generation model, a video as a sample (referred to as a first sample video) can be obtained first, and an image frame can be extracted from the first sample video.
[0056] The motion intensity of the first sample video can be evaluated by the trained motion estimation model, and various types of motion intensity data of the first sample video can be obtained. In this case, the various types of motion intensity data of the first sample video output by the trained motion estimation model can be obtained.
[0057] Having extracted the aforementioned image frames from the first sample video and obtained various types of motion intensity data from the first sample video, the image frames and these various types of motion intensity data can be input into the aforementioned image-generated video model. The image-generated video model then performs calculations based on the image frames and these various types of motion intensity data to generate a video corresponding to the image frames and these various types of motion intensity data. That is, the generated video is a video generated based on the image frames, and the motion intensity in the video matches these various types of motion intensity data.
[0058] Given the video generated by the aforementioned graph-to-video model, the training loss can be determined based on the generated video. After adjusting the model parameters for the graph-to-video model according to the loss, the training of the graph-to-video model is considered complete.
[0059] After training the aforementioned image-to-video model, the trained model can be used to perform image-to-video tasks. When using the trained image-to-video model, an image (which can be called the target image) and preset motion intensity data can be obtained. These target image and motion intensity data are then input into the trained image-to-video model, which calculates based on the target image and motion intensity data to generate the corresponding video.
[0060] Alternatively, when training the above-mentioned image-generated video model, text describing the motion in the first sample video (which can be called sample text) can also be obtained. Accordingly, the image frame, the sample text, and these various types of motion intensity data can be input into the image-generated video model, which will then calculate and generate the corresponding video based on the image frame, the sample text, and these various types of motion intensity data.
[0061] When using the trained graph-to-video model, the image used to generate the video (referred to as the target image), the text used to describe the motion in the video to be generated (referred to as the target text), and the preset motion intensity data can be obtained. The target image, the target text, and the motion intensity data are then input into the trained graph-to-video model, which calculates based on the target image, the target text, and the motion intensity data to generate the corresponding video.
[0062] Please refer to Figure 3 , Figure 3 This is a flowchart illustrating a training method for a graph-generated video model, as shown in an exemplary embodiment of this application.
[0063] like Figure 3 As shown, the training method for the above-mentioned image-generated video model may include the following steps:
[0064] Step 302: Obtain a first sample video, and extract image frames from the first sample video.
[0065] In this embodiment, the video generation model can be pre-set. The video generation model can be a deep learning-based generative model, such as a variational autoencoder (VAE), a generative adversarial network (GAN), a diffusion model, etc.
[0066] The generative model aims to learn the distribution of data so as to be able to generate new data similar to the training data. Therefore, using the generative model, a video sequence containing multiple images similar to an input image can be generated based on the input image, and the video sequence can be regarded as a video.
[0067] In some embodiments, the above-mentioned video generation model can be a diffusion model.
[0068] The diffusion model is a generative model, and its basic idea is to convert data into noise through a process of gradually adding noise, and then learn how to reverse this process to recover the original data from pure noise. The working principle of the diffusion model includes a forward process and a reverse process; in the forward process, starting from the original data, Gaussian noise is gradually added, and at each step, the data becomes closer to completely random noise; in the reverse process, the goal is to start from pure noise and gradually recover the original data through a series of denoising steps. The diffusion model can be used to generate new data, usually by drawing data from a noise distribution and converting it into new data that looks like it comes from the original data set through the reverse process.
[0069] When training the above-mentioned video generation model, first, a video as a sample (which can be referred to as a first sample video) can be obtained, and image frames can be extracted from the first sample video.
[0070] It should be noted that when extracting image frames from the above-mentioned first sample video, one or more image frames can be randomly extracted, or image frames can be extracted according to certain extraction rules (for example: extract every N frames), and the present application does not make special limitations thereto.
[0071] Step 304: Obtain the motion intensity data of the first sample video of multiple types by performing motion intensity evaluation on the first sample video by the trained motion estimation model.
[0072] In this embodiment, the motion estimation model trained can perform motion intensity evaluation on the first sample video, and obtain multiple types of motion intensity data of the first sample video. In this case, the multiple types of motion intensity data of the first sample video output by the trained motion estimation model can be obtained.
[0073] In actual applications, the motion in a video usually includes object motion and scene motion. For example, in a road video captured by a fixed camera, the scene does not move, but the cars, pedestrians and the like as objects move; in a road video captured by a moving camera, not only the cars, pedestrians and the like as objects move, but also the scene moves. Based on this, in some embodiments, the motion intensity data can be divided into two types, i.e., object motion intensity data and shot motion intensity data. Of course, more other types of motion intensity data can be divided according to actual conditions and requirements, which are not specially limited in this application.
[0074] In the above case, the multiple types of motion intensity data of the first sample video can include object motion intensity data and shot motion intensity data of the first sample video. That is, after the motion estimation model performs motion intensity evaluation on the first sample video, the object motion intensity and the shot motion intensity of the first sample video can be distinguished, and the object motion intensity data and the shot motion intensity data of the first sample video can be output.
[0075] In some embodiments, in order to reduce the complexity of data processing to a certain extent, the motion intensity data can be a score reflecting the motion intensity. For example, the object motion intensity data can be an object motion intensity score in the range of [0, 10], which focuses on the motion of the shooting subject in the video, such as the motion of a car or a windmill. The shot motion intensity data can be a shot motion intensity score in the range of [0, 10], which focuses on the operation of the camera when shooting the video, such as left-right panning, push-pull zooming and the like.
[0076] In some embodiments, inputting the first sample video into the trained motion estimation model, and performing motion intensity evaluation on the first sample video by the motion estimation model to obtain multiple types of motion intensity data of the first sample video can be performed online or offline.
[0077] Specifically, after the first sample video is obtained, the first sample video can be input into the trained motion estimation model in real time, and motion intensity evaluation on the first sample video can be performed by the motion estimation model to obtain multiple types of motion intensity data of the first sample video.
[0078] Alternatively, in order to accelerate the speed of obtaining the multiple types of motion intensity data of the first sample video, thereby improving the training efficiency of the video generation model, the multiple sample videos including the first sample video can be input into the trained motion estimation model respectively in advance, and the motion estimation model can perform motion intensity evaluation on each sample video to obtain the multiple types of motion intensity data of each sample video. In addition, each sample video and its multiple types of motion intensity data can be stored together to facilitate the obtaining of the multiple types of motion intensity data of each sample video.
[0079] Step 306: inputting the image frame and the multiple types of motion intensity data into a preset video generation model, and generating a corresponding video based on the image frame and the multiple types of motion intensity data by the video generation model.
[0080] In the embodiment, in the case that the image frame is extracted from the first sample video and the multiple types of motion intensity data of the first sample video are obtained, the image frame and the multiple types of motion intensity data can be input into the video generation model, and a video corresponding to the image frame and the multiple types of motion intensity data can be generated by the video generation model based on the image frame and the multiple types of motion intensity data. That is, the generated video is a video generated based on the image frame, and the motion intensity in the video matches the multiple types of motion intensity data.
[0081] In some embodiments, when the image frame and the multiple types of motion intensity data are input into the video generation model, and a corresponding video is generated by the video generation model based on the image frame and the multiple types of motion intensity data, the image frame and the multiple types of motion intensity data can be input into the video generation model, and the video generation model can generate a feature vector corresponding to the image frame and generate feature vectors corresponding to the multiple types of motion intensity data respectively. Subsequently, the video generation model can splice the feature vectors and further calculate based on the spliced feature vectors to generate a video corresponding to the image frame and the multiple types of motion intensity data.
[0082] In actual application, the model structure of the video generation model can include a linear layer (Linear Layer) for performing feature extraction tasks. That is, after the image frame and the multiple types of motion intensity data are input into the video generation model, the linear layer in the video generation model can first map the image frame into a corresponding feature vector and map the multiple types of motion intensity data into corresponding feature vectors respectively.
[0083] It should be noted that when the above graph video model is a diffusion model, since an integer T needs to be randomly generated as one of the training conditions when training the diffusion model, the feature vector corresponding to the integer T can be added to the above feature vector, and further calculation is performed based on the added feature vector.
[0084] By concatenating the feature vectors corresponding to the multiple types of motion intensity data to other feature vectors, the graph video model can further calculate based on the concatenated feature vectors to generate videos, which can better decouple different types of motion intensity data and further improve the learning effect of the graph video model based on the decoupled motion intensity data.
[0085] Step 308: Determine the training loss based on the generated video, and after adjusting the model parameters of the graph video model according to the loss, determine that the training of the graph video model is completed.
[0086] In this embodiment, in the case where the video generated by the above graph video model is obtained, the training loss can be determined based on the generated video, and after adjusting the model parameters of the graph video model according to the loss, it is determined that the training of the graph video model is completed.
[0087] It should be noted that the selection and calculation method of the above training loss will vary depending on the specific task target, model architecture and type of data used. For the above graph video model, the following training losses can be selected:
[0088] Reconstruction Loss: Reconstruction loss is used to measure the difference between the generated video frame and the target real video frame. For example, the sum of squares of pixel value differences (i.e. mean square error) between the generated frame and the target frame can be calculated, or the sum of absolute values of pixel value differences (i.e. absolute error) between the generated frame and the target frame can be calculated, or a pre-trained deep network (such as VGG-16) can be used to extract features, and the differences between the generated frame and the target frame in these feature spaces (i.e. perceptual loss) can be compared.
[0089] Motion Consistency Loss: In order to ensure the motion consistency between the generated video frames, motion consistency loss can be introduced. This can be achieved by calculating the optical flow field between adjacent frames or directly comparing the frame difference. For example, L1 or L2 norm can be used to measure whether the changes between the two frames meet the expected value.
[0090] Temporal Smoothness Loss: Encourages the generated video to be smoother in the temporal dimension, avoiding sudden changes.
[0091] Structural Similarity Index (SSIM) Loss: Used to evaluate the structural similarity between two images.
[0092] Of course, a variety of loss functions can also be combined to optimize the performance of the model. For example, a complete loss function can be a weighted sum of the above several losses, where the weight coefficients of each loss term can be adjusted according to actual conditions and requirements.
[0093] In practical applications, the model parameters of the above image-to-video model can be automatically adjusted based on the above losses according to pre-set model parameter adjustment rules.
[0094] Alternatively, the above losses can be output to the user, and the user can adjust the model parameters of the above image-to-video model based on the losses.
[0095] It should be noted that after the training of the above image-to-video model is completed, the trained image-to-video model can be used to perform image-to-video tasks.
[0096] Specifically, in some embodiments, when using the trained image-to-video model, an image used to generate a video (which can be referred to as a target image) and pre-set motion intensity data can be obtained, and the target image and the motion intensity data can be input into the trained image-to-video model. The image-to-video model calculates based on the target image and the motion intensity data to generate a corresponding video.
[0097] In practical applications, in order to make the generated video better meet the user's needs and improve the user's experience of using the image-to-video service, the user can be allowed to control the content, style, and even emotional tone of the generated video through a textual description. In this case, the image used to generate the video, the description text used to describe the motion in the video to be generated, and various types of motion intensity data can be input into the above image-to-video model, and the image-to-video model can generate a video based on the image under the guidance of the description text and the motion intensity data as given conditions.
[0098] Specifically, in some embodiments, when training the above-mentioned image-to-video model, text (which can be referred to as sample text) describing the motion in the above-mentioned first sample video can also be obtained. Accordingly, when inputting the above-mentioned image frame and the above-mentioned multiple types of motion intensity data into the image-to-video model, and generating a corresponding video based on the image frame and the multiple types of motion intensity data by the image-to-video model, the image frame, the sample text and the multiple types of motion intensity data can be specifically input into the image-to-video model, and a corresponding video can be generated by the image-to-video model based on the image frame, the sample text and the multiple types of motion intensity data.
[0099] In some embodiments, when using the trained image-to-video model, an image (which can be referred to as a target image) for generating a video, text (which can be referred to as target text) describing the motion in the video to be generated at this time, and preset motion intensity data can be obtained, and the target image, the target text and the motion intensity data can be input into the trained image-to-video model, and a corresponding video can be generated by the image-to-video model based on the target image, the target text and the motion intensity data.
[0100] In the above technical solution, first, a sample video can be obtained, and an image frame can be extracted from the sample video. In addition, multiple types of motion intensity data of the first sample video can be obtained by performing motion intensity evaluation on the sample video by a trained motion estimation model. Then, the image frame and the multiple types of motion intensity data can be input into a preset image-to-video model, and a corresponding video can be generated by the image-to-video model based on the image frame and the multiple types of motion intensity data. Finally, a training loss can be determined based on the generated video, and after adjusting the model parameters of the image-to-video model according to the loss, the training of the image-to-video model can be completed.
[0101] In the above manner, on the one hand, different types of motion in the video can be distinguished, and decoupled multiple types of motion intensity data can be injected into the image-to-video model for training, so that the learning effect of the image-to-video model based on the decoupled motion intensity data can be improved, and the accuracy of the motion-related knowledge learned by the image-to-video model can be ensured. On the other hand, using an additional motion estimation model to evaluate the motion intensity of the video can ensure the accuracy of the various types of motion intensity data of the video obtained by evaluation, and help the image-to-video model better understand and generate the motion in the video. In this way, when using the trained image-to-video model to generate image-to-video, the motion intensity of the generated video can be better controlled, and a high-quality video can be generated.
[0102] Please refer to Figure 4 ,Figure 4 is a flowchart of a method of training and using a motion estimation model according to an example embodiment of the present application.
[0103] As shown in Figure 4 the method of training and using the motion estimation model can include the following steps:
[0104] Step 402: Obtain a second sample video; wherein the second sample video is labeled with the multiple types of motion intensity data.
[0105] Step 404: Based on the second sample video, perform supervised training on the motion estimation model.
[0106] In this embodiment, in order to use the above-mentioned motion estimation model to evaluate the multiple types of motion intensity data of a video, it is also necessary to first train the motion estimation model.
[0107] Specifically, first, a video as a sample (which can be referred to as a second sample video) can be obtained. For any one of the second sample videos, the second sample video is labeled with the multiple types of motion intensity data of the second sample video as multiple labels of the second sample video.
[0108] Subsequently, based on the above-mentioned second sample video with labels, supervised training can be performed on the above-mentioned motion estimation model.
[0109] In some embodiments, in order to improve the model effect of the above-mentioned motion estimation model, a multi-layer dynamic convolution structure (Multi-layer Dynamic Convolution Structure) and a multi-layer perceptron (Multi-layer Perceptron, MLP) can be used; and the multi-layer dynamic convolution structure can be pre-trained on an ActivityNet dataset first, and then the second sample data is used to train the motion estimation model as a whole.
[0110] The ActivityNet dataset contains a large number of YouTube videos, which cover various daily activities, sports events, entertainment activities, etc., and each video is finely labeled, including the starting time, ending time and corresponding activity category of the action. By pre-training the above-mentioned multi-layer dynamic convolution structure, the speed of subsequently completing the training of the above-mentioned motion estimation model can be accelerated, and the efficiency of the training of the above-mentioned motion estimation model can be improved.
[0111] Step 406: inputting the first sample video into the trained motion estimation model, performing motion intensity evaluation on the first sample video by the motion estimation model, and obtaining the motion intensity data of the first sample video.
[0112] In this embodiment, after the training of the motion estimation model is completed, the trained motion estimation model can be used to perform the motion intensity evaluation task. That is, the first sample video can be input into the trained motion estimation model, and the motion intensity evaluation on the first sample video can be performed by the motion estimation model to obtain the motion intensity data of the first sample video.
[0113] Specifically, in some embodiments, the first sample video can be input into the trained motion estimation model, the motion trajectory in the first sample video can be identified by the motion estimation model, and the motion intensity evaluation on the first sample video can be performed according to the motion trajectory to obtain the motion intensity data of the first sample video.
[0114] For example, if the motion trajectory of an object in the video is identified, the motion intensity data of the object in the video can be obtained according to the change speed and intensity of the motion trajectory. For the motion trajectory of a background point in the video, the camera motion intensity data of the video can be obtained according to the change speed and intensity of the motion trajectory.
[0115] In actual applications, the motion estimation model can also be used in other ways to perform motion intensity evaluation on the video, and the present application does not make special limitations in this regard.
[0116] Corresponding to the foregoing embodiments of the training method of the video generation model, the present application also provides embodiments of a training device of the video generation model.
[0117] Reference is made to Figure 5 , Figure 5 is a structural schematic diagram of a device according to an example embodiment of the present application. At the hardware level, the device includes a processor 502, an internal bus 504, a network interface 506, a memory 508, and a non-volatile memory 510, and of course can also include other required hardware. One or more embodiments of the present application can be implemented in a software manner, such as reading a corresponding computer program from the non-volatile memory 510 into the memory 508 by the processor 502 and then running. Of course, in addition to the software implementation, one or more embodiments of the present application do not exclude other implementation manners, such as a logic device or a combination of software and hardware, and the like, that is, the execution subject of the following processing flow is not limited to each logical module, but can also be hardware or a logic device.
[0118] Referring to Figure 6 , Figure 6 is a block diagram of a device for training a graph-based video model according to an example embodiment of the present application.
[0119] The device for training a graph-based video model can be applied to Figure 5 the device shown in FIG. 8 to implement the technical solutions of the present application.
[0120] The device for training a graph-based video model can include:
[0121] A first obtaining module 602 obtains a first sample video and extracts image frames from the first sample video.
[0122] A second obtaining module 604 obtains a plurality of types of motion intensity data of the first sample video obtained by performing motion intensity evaluation on the first sample video by a trained motion estimation model.
[0123] A generating module 606 inputs the image frames and the plurality of types of motion intensity data into a preset graph-based video model, and generates a corresponding video based on the image frames and the plurality of types of motion intensity data by the graph-based video model.
[0124] A training module 608 determines a training loss based on the generated video, and determines that the training of the graph-based video model is completed after adjusting the model parameters of the graph-based video model according to the loss.
[0125] In some embodiments, the first sample video is input into the trained motion estimation model, and a plurality of types of motion intensity data of the first sample video are obtained by performing motion intensity evaluation on the first sample video by the motion estimation model, including:
[0126] The first sample video is input into the trained motion estimation model, and a plurality of types of motion intensity data of the first sample video are obtained by identifying motion trajectories in the first sample video by the motion estimation model and performing motion intensity evaluation on the first sample video according to the motion trajectories.
[0127] In some embodiments, the motion estimation model includes a multi-layer dynamic convolution structure and a multi-layer perceptron.
[0128] In some embodiments, the device further includes:
[0129] A third obtaining module obtains a second sample video, wherein the second sample video is labeled with the plurality of types of motion intensity data.
[0130] The second training module performs supervised training on the motion estimation model based on the second sample video.
[0131] In some embodiments, the inputting the image frame and the motion intensity data of the plurality of types into the preset image-to-video model, the generating, by the image-to-video model, a feature vector corresponding to the image frame, and the generating, by the image-to-video model, a feature vector corresponding to each of the motion intensity data of the plurality of types, and the further computing based on the feature vectors after the splicing to generate a corresponding video, comprise:
[0132] The inputting the image frame and the motion intensity data of the plurality of types into the preset image-to-video model, the generating, by the image-to-video model, a feature vector corresponding to the image frame, and the generating, by the image-to-video model, a feature vector corresponding to each of the motion intensity data of the plurality of types, and the further computing based on the feature vectors after the splicing to generate a corresponding video.
[0133] In some embodiments, the motion intensity data of the plurality of types comprises object motion intensity data and shot motion intensity data.
[0134] In some embodiments, the image-to-video model is a diffusion model.
[0135] In some embodiments, the apparatus further comprises:
[0136] The fourth obtaining module obtains a target image to be generated and preset motion intensity data;
[0137] The generating module is further configured to input the target image and the motion intensity data into the trained image-to-video model, and generate a corresponding video based on the image and the motion intensity data by the image-to-video model.
[0138] In some embodiments, the apparatus further comprises:
[0139] The fifth obtaining module obtains sample text describing the motion in the first sample video;
[0140] The inputting the image frame and the motion intensity data of the plurality of types into the preset image-to-video model, the generating, by the image-to-video model, a feature vector corresponding to the image frame, and the generating, by the image-to-video model, a feature vector corresponding to each of the motion intensity data of the plurality of types, and the further computing based on the feature vectors after the splicing to generate a corresponding video, comprise:
[0141] The inputting the image frame and the motion intensity data of the plurality of types into the preset image-to-video model, the generating, by the image-to-video model, a feature vector corresponding to the image frame, and the generating, by the image-to-video model, a feature vector corresponding to each of the motion intensity data of the plurality of types, and the further computing based on the feature vectors after the splicing to generate a corresponding video.
[0142] In some embodiments, the apparatus further comprises:
[0143] a sixth obtaining module, configured to obtain a target image to be generated, target text used for describing motion in a generated video, and preset motion intensity data;
[0144] The generation module is further configured to input the target image, the target text, and the motion intensity data into the trained image-to-video model, and generate a corresponding video based on calculation of the image-to-video model on the target image, the target text, and the motion intensity data.
[0145] For the device embodiment, it basically corresponds to the method embodiment, and thus the related parts can be referred to the part of the method embodiment. The device embodiments described above are merely illustrative, wherein the modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical modules, that is, they can be located in one place or distributed on multiple network modules. According to actual needs, some or all of the modules can be selected to achieve the purpose of the technical solution of the present application.
[0146] The system, device, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or an entity, or by a product with certain function. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0147] In a typical configuration, a computer includes one or more processors (CPU), input / output interface, network interface, and memory.
[0148] The memory can include a non-persistent memory in a computer readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer readable medium.
[0149] Computer-readable media includes permanent and non-permanent, moveable and non- moveable media that can be implemented by any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disks (DVDs) or other optical storage, magnetic cassettes, magnetic disks storage, quantum memory, graphene-based storage media, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definitions provided herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.
[0150] It should be noted that the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0151] The above detailed description has described only certain embodiments of the application. Other embodiments are within the scope of the application. In some instances, acts or steps can be performed in a different order from the order described in the application and still achieve desirable results. Additionally, the process depicted in the accompanying figures can not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.
[0152] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The term "and / or" includes any and all combinations of one or more of the associated listed items.
[0153] The description of the term "one embodiment", "some embodiments", "an example", "a specific example" or "an implementation" used in one or more embodiments of the present application means that the particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. The illustrative description of these terms does not necessarily refer to the same embodiment. Moreover, the described particular features or characteristics can be combined in any suitable manner in one or more embodiments of the present application. In addition, different embodiments and particular features or characteristics in different embodiments can be combined under appropriate circumstances.
[0154] It should be understood that although the terms first, second, third, etc. can be used herein to describe various information, the information should not be limited to these terms. These terms are only used to distinguish one piece of information from another piece of information. For example, without departing from the scope of one or more embodiments of the present application, first information can also be referred to as second information, and similarly, second information can also be referred to as first information. Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon" or "in response to determining".
[0155] The above description is only a preferred embodiment of one or more embodiments of the present application and is not intended to limit one or more embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of the present application should be included in the scope of protection of one or more embodiments of the present application.
[0156] The user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
Claims
1. A method for training a graph-generated video model, the method comprising: Acquire a first sample video and extract image frames from the first sample video; The motion estimation model, after being trained, is used to evaluate the motion intensity of the first sample video, resulting in various types of motion intensity data for the first sample video; wherein, the motion estimation model is used to evaluate the motion intensity of the input video to obtain various types of motion intensity data for the video. The image frames and the various types of motion intensity data are input into a preset image-generated video model, which then calculates based on the image frames and the various types of motion intensity data to generate the corresponding video. The training loss is determined based on the generated video, and the training of the image-generated video model is considered complete after the model parameters for the image-generated video model are adjusted according to the loss.
2. The method according to claim 1, wherein the first sample video is input into the trained motion estimation model, and the motion estimation model performs motion intensity evaluation on the first sample video to obtain multiple types of motion intensity data of the first sample video, including: The first sample video is input into the trained motion estimation model, which identifies the motion trajectory in the first sample video and evaluates the motion intensity of the first sample video based on the motion trajectory, thereby obtaining various types of motion intensity data of the first sample video.
3. The method according to claim 1, wherein the motion estimation model comprises a multi-layer dynamic convolutional structure and a multi-layer perceptron.
4. The method according to claim 1, further comprising: Obtain a second sample video; wherein the second sample video is labeled with the various types of motion intensity data; Based on the second sample video, the motion estimation model is trained in a supervised manner.
5. The method according to claim 1, wherein inputting the image frame and the various types of motion intensity data into a preset image-generated video model, and having the image-generated video model calculate based on the image frame and the various types of motion intensity data to generate a corresponding video, includes: The image frame and the various types of motion intensity data are input into a preset image-generated video model. The image-generated video model generates feature vectors corresponding to the image frame and feature vectors corresponding to the various types of motion intensity data. The feature vectors are then concatenated, and further calculations are performed based on the concatenated feature vectors to generate the corresponding video.
6. The method according to claim 1, wherein the multiple types of motion intensity data include object motion intensity data and lens motion intensity data.
7. The method according to claim 1, wherein the image-generated video model is a diffusion model.
8. The method according to claim 1, further comprising: Acquire the target image to be generated and the preset motion intensity data; The target image and the motion intensity data are input into the trained image-generated video model, which then calculates and generates a corresponding video based on the image and the motion intensity data.
9. The method according to claim 1, further comprising: Obtain sample text used to describe the motion in the first sample video; The step of inputting the image frame and the various types of motion intensity data into a preset image-generated video model, and having the image-generated video model calculate and generate a corresponding video based on the image frame and the various types of motion intensity data, includes: The image frame, the sample text, and the various types of motion intensity data are input into a preset image-generated video model. The image-generated video model then performs calculations based on the image frame, the sample text, and the various types of motion intensity data to generate the corresponding video.
10. The method according to claim 9, further comprising: Acquire the target image to be generated, the target text used to describe the motion in the generated video, and the preset motion intensity data; The target image, the target text, and the motion intensity data are input into the trained image-generated video model, which then performs calculations based on the target image, the target text, and the motion intensity data to generate the corresponding video.
11. A training apparatus for a graph-generated video model, the apparatus comprising: The first acquisition module acquires a first sample video and extracts image frames from the first sample video. The second acquisition module acquires various types of motion intensity data of the first sample video obtained by the trained motion estimation model performing motion intensity assessment on the first sample video; wherein, the motion estimation model is used to perform motion intensity assessment on the input video to obtain various types of motion intensity data of the video. The generation module inputs the image frames and the various types of motion intensity data into a preset image-generated video model, which then calculates based on the image frames and the various types of motion intensity data to generate the corresponding video. The training module determines the training loss based on the generated video, and after adjusting the model parameters of the image-generated video model according to the loss, determines that the training of the image-generated video model is complete.
12. An electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor implements the method as described in any one of claims 1 to 10 by executing the executable instructions.
13. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Self-adaption motion estimation method and module thereof
CN104995917A
Video motion estimation method and device and storage medium
CN110213591A