Model training method for video generation and video generation method and device

By extracting one-dimensional motion features and image features from driving videos and combining them with natural language text, the motion control signals and camera perspectives are decoupled, solving the problem of insufficient flexibility in existing video generation methods and achieving high-performance video generation that matches user expectations.

CN122053936APending Publication Date: 2026-05-15BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2026-02-02
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, motion-driven character generation video generation methods have poor camera angle flexibility and cannot achieve high-performance video generation effects that closely match user expectations.

Method used

One-dimensional motion features are extracted from the driving video by a motion encoder, and the model is trained by combining image features and natural language text features to decouple the motion control signal from the camera viewpoint. A video generation network is used for denoising and the parameters of the motion encoder and the video generation network are updated.

Benefits of technology

It achieves high-quality video generation that closely matches user expectations, with free control over camera angles and automatic motion alignment capabilities, thus enhancing the flexibility and accuracy of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122053936A_ABST
    Figure CN122053936A_ABST
Patent Text Reader

Abstract

The invention provides a model training method for video generation and a video generation method and device, and belongs to the technical field of computer vision. In the present disclosure, a model for video generation includes a motion encoder and a video generation network. In the training stage, one-dimensional motion features are extracted from a driving video through a motion encoder and serve as motion control signals. The one-dimensional feature structure can effectively filter the two-dimensional space layout information, so that the motion control signal only represents the motion state of the object and is completely decoupled from the lens visual angle of the driving video. Based on the decoupling design, the scheme realizes the separation of the action behavior in the driving video and the original lens view angle, so that when the trained model executes the video generation task in the reasoning stage, the lens view angle of the generated video is not limited to the lens view angle of the driving video any more, and the video generation efficiency is improved. Therefore, the free control of the lens view angle of the generated video becomes possible, and the video generation flexibility is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer vision technology, and in particular to a model training method, video generation method, and apparatus for video generation. Background Technology

[0002] With the widespread adoption of smart devices, the demand for video content is constantly growing. Correspondingly, video generation has also experienced rapid development and has become a research and application hotspot in the field of computer vision technology.

[0003] One popular video generation task is motion-driven character generation. This involves using motion control signals extracted from the driving video to control the actions of a character in a reference image, thereby generating a video in which the character mimics the actions in the driving video. Related techniques typically extract two-dimensional image features (such as a skeleton map) from the driving video as motion control signals to drive the character in the reference image to perform the motion mimicry.

[0004] However, this generation method has obvious limitations. The camera angle of the generated video can only replicate the angle of the driving video, which is not very flexible and cannot achieve high expressiveness and video generation effect that closely matches the user's expectations. Summary of the Invention

[0005] This disclosure provides a model training method, a video generation method, and an apparatus for video generation. The technical solution of this disclosure is shown below.

[0006] According to a first aspect of the present disclosure, a method for training a model for video generation is provided, the model including a motion encoder and a video generation network; the method includes: Acquire a driving video, a first target video, and a first reference image; wherein the first target video uses a first object in the first reference image as the subject of the shot and is used to replicate the action of a second object in the driving video; Image features are extracted from the first reference image, and first video features are extracted from the first target video. Noise is added to the first video features to obtain noisy video features. The motion encoder extracts one-dimensional motion features from the driving video, and the one-dimensional motion features are used to characterize the motion pattern of the second object. Under given control conditions, the noisy video features are denoised by the video generation network to obtain second video features; wherein, the control conditions include the one-dimensional motion features and the image features; Based on the first video features and the second video features, the parameters of the motion encoder and the video generation network are updated.

[0007] In some embodiments, extracting one-dimensional motion features from the driving video using the motion encoder includes: Data enhancement processing is performed on the video frames included in the driving video to obtain the processed video frames; Extract local motion features and global motion features from the processed video frames; The local motion features are used to characterize the motion pattern of a single part of the second object; the global motion features are used to indicate the motion patterns of multiple parts of the second object.

[0008] In other embodiments, the motion encoder includes a first motion encoder and a second motion encoder; the extraction of local motion features and global motion features from the processed video frames includes: The processed video frames are mapped to the implicit space using the first motion encoder to obtain the local motion features; The processed video frames are mapped to the implicit space using the second motion encoder to obtain the global motion features.

[0009] In other embodiments, the data enhancement processing of the video frames included in the driving video to obtain processed video frames includes at least one of the following: A random perspective transformation is performed on the video frames included in the driving video to obtain the processed time-frequency frame; Color dithering is performed on the video frames included in the driving video to obtain the processed time-frequency frame; wherein, color dithering includes adjusting brightness, contrast, saturation and hue.

[0010] In other embodiments, the control conditions further include: text features extracted from a first natural language text; wherein the first natural language text relates to at least one of the following: The camera angle of the first target video; The video style of the first target video.

[0011] In other embodiments, the first reference image is the first frame of the first target video.

[0012] In other embodiments, the method further includes any one of the following: Based on the cross-attention mechanism, the control conditions are used as keys and values ​​and input into the reverse denoising process of the video generation network. The noisy video features and the control conditions are spliced ​​together to obtain a first splicing result, and the first splicing result is input into the reverse denoising process; Based on the cross-attention mechanism, the first part of the control conditions is used as a key and a value and input into the reverse denoising process; and the noisy video features and the second part of the control conditions are concatenated to obtain a second concatenation result, which is then input into the reverse denoising process.

[0013] In other embodiments, the step of using the first part of the control conditions as a key and value, based on the cross-attention mechanism, and inputting it into the reverse denoising process; and concatenating the noisy video features and the second part of the control conditions to obtain a second concatenation result, and inputting the second concatenation result into the reverse denoising process, includes: Based on the cross-attention mechanism, the one-dimensional motion features and the text features are used as keys and values, respectively, and input into the reverse denoising process; and, The noisy video features and the image features are concatenated to obtain the second concatenation result, and the second concatenation result is input into the reverse denoising process.

[0014] In other embodiments, the camera angle of the first reference image is consistent with the camera angle of the driving video; the first target video uses the first object as the subject to replicate the camera angle and the action of the second object; updating the parameters of the motion encoder and the video generation network based on the first video features and the second video features includes: Based on the first video features and the second video features, the motion encoder and the video generation network are updated in the first stage.

[0015] In other embodiments, after the parameter update in the first phase, the method further includes: Acquire a second target video and a second reference image; wherein the camera angle of the second reference image is inconsistent with the camera angle of the driving video; the second target video takes the first object as the subject of the shot and is used to replicate the camera angle of the second reference image and the action of the second object; Based on the driving video, the second target video, and the second reference image, the motion encoder and the video generation network are updated in the second stage.

[0016] In other embodiments, the model further includes a motion decoder; updating the parameters of the motion encoder and the video generation network based on the first video features and the second video features includes: Based on the first video features and the second video features, a first loss is obtained; First human pose parameters are extracted from the first target video, and the one-dimensional motion features are mapped to second human pose parameters through the motion decoder; a second loss is obtained based on the first human pose parameters and the second human pose parameters. Based on the weight coefficients corresponding to the second loss, the weighted sum of the first loss and the second loss is obtained to obtain the total loss; wherein, the value of the weight coefficient gradually decays to zero after the start of training; Based on the total loss, update the parameters of the motion encoder and the video generation network.

[0017] According to a second aspect of the present disclosure, a video generation method is provided, comprising: Acquire driving video, user-input natural language text, and reference images; A model for video generation is invoked to generate a target video based on the driving video, the natural language text, and the reference image; wherein the model for video generation is trained using the model training method described above; the target video uses the object in the reference image as the subject and is used to replicate the orientation of the object included in the reference image and the action of the object included in the driving video; the camera angle of the target video is consistent with the camera angle of the reference image, or the camera angle of the target video is consistent with the camera angle indicated by the natural language text.

[0018] According to a third aspect of the present disclosure, a model training apparatus for video generation is provided, the model including a motion encoder and a video generation network; the apparatus includes: The first acquisition module is configured to acquire a driving video, a first target video, and a first reference image; wherein the first target video uses a first object in the first reference image as the subject of the shot and is used to replicate the action of a second object in the driving video. The first encoding module is configured to extract image features from the first reference image and extract first video features from the first target video; The noise-adding module is configured to perform noise-adding processing on the first video features to obtain noisy video features; The second encoding module is configured to extract one-dimensional motion features from the driving video through the motion encoder, the one-dimensional motion features being used to characterize the motion pattern of the second object; The first generation module is configured to perform denoising processing on the noisy video features through the video generation network under given control conditions to obtain the second video features; wherein the control conditions include the one-dimensional motion features and the image features; The training module is configured to update the parameters of the motion encoder and the video generation network based on the first video features and the second video features.

[0019] In some embodiments, the second encoding module is configured to: Data enhancement processing is performed on the video frames included in the driving video to obtain the processed video frames; Extract local motion features and global motion features from the processed video frames; The local motion features are used to characterize the motion pattern of a single part of the second object; the global motion features are used to indicate the motion patterns of multiple parts of the second object.

[0020] In other embodiments, the motion encoder includes a first motion encoder and a second motion encoder; the second encoding module is configured to: The processed video frames are mapped to the implicit space using the first motion encoder to obtain the local motion features; The processed video frames are mapped to the implicit space using the second motion encoder to obtain the global motion features.

[0021] In other embodiments, the second encoding module is configured to perform at least one of the following: A random perspective transformation is performed on the video frames included in the driving video to obtain the processed time-frequency frame; Color dithering is performed on the video frames included in the driving video to obtain the processed time-frequency frame; wherein, color dithering includes adjusting brightness, contrast, saturation and hue.

[0022] In other embodiments, the control conditions further include: text features extracted from a first natural language text; wherein the first natural language text relates to at least one of the following: The camera angle of the first target video; The video style of the first target video.

[0023] In other embodiments, the first reference image is the first frame of the first target video.

[0024] In other embodiments, the apparatus further includes a condition injection module; the condition injection module is configured to perform any of the following: Based on the cross-attention mechanism, the control conditions are used as keys and values ​​and input into the reverse denoising process of the video generation network. The noisy video features and the control conditions are spliced ​​together to obtain a first splicing result, and the first splicing result is input into the reverse denoising process; Based on the cross-attention mechanism, the first part of the control conditions is used as a key and a value and input into the reverse denoising process; and the noisy video features and the second part of the control conditions are concatenated to obtain a second concatenation result, which is then input into the reverse denoising process.

[0025] In other embodiments, the conditional injection module is configured to: Based on the cross-attention mechanism, the one-dimensional motion features and the text features are used as keys and values, respectively, and input into the reverse denoising process; and, The noisy video features and the image features are concatenated to obtain the second concatenation result, and the second concatenation result is input into the reverse denoising process.

[0026] In other embodiments, the camera angle of the first reference image is consistent with the camera angle of the driving video; the first target video uses the first object as the subject of the shot to replicate the camera angle and the action of the second object; the training module is configured to: Based on the first video features and the second video features, the motion encoder and the video generation network are updated in the first stage.

[0027] In other embodiments, the device further includes: The second acquisition module is configured to acquire a second target video and a second reference image; wherein the camera angle of the second reference image is inconsistent with the camera angle of the driving video; the second target video takes the first object as the subject of the shot and is used to replicate the camera angle of the second reference image and the action of the second object; The training module is further configured to perform a second-stage parameter update on the motion encoder and the video generation network based on the driving video, the second target video, and the second reference image after the first-stage parameter update.

[0028] In other embodiments, the model further includes a motion decoder; the training module is configured to: Based on the first video features and the second video features, a first loss is obtained; First human pose parameters are extracted from the first target video, and the one-dimensional motion features are mapped to second human pose parameters through the motion decoder; a second loss is obtained based on the first human pose parameters and the second human pose parameters. Based on the weight coefficients corresponding to the second loss, the weighted sum of the first loss and the second loss is obtained to obtain the total loss; wherein, the value of the weight coefficient gradually decays to zero after the start of training; Based on the total loss, update the parameters of the motion encoder and the video generation network.

[0029] According to a fourth aspect of the present disclosure, a video generation apparatus is provided, comprising: The third acquisition module is configured to acquire the driving video, the natural language text input by the user, and the reference image; The second generation module is configured to invoke a model for video generation to generate a target video based on the driving video, the natural language text, and the reference image; wherein the model for video generation is trained using the aforementioned model training device; the target video uses an object in the reference image as the subject and is used to replicate the orientation of the object in the reference image and the actions of the object in the driving video; the camera angle of the target video is consistent with the camera angle of the reference image, or the camera angle of the target video is consistent with the camera angle indicated by the natural language text.

[0030] According to a fifth aspect of the present disclosure, an electronic device is provided, the electronic device comprising: One or more processors; Memory used to store the executable program code of the processor; The processor is configured to execute the program code to implement the aforementioned model training method for video generation or the aforementioned video generation method.

[0031] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided, wherein program code in the computer-readable storage medium, when executed by a processor of an electronic device, enables the electronic device to perform the above-described model training method for video generation or the above-described video generation method.

[0032] According to a seventh aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor of an electronic device, implements the above-described model training method for video generation or the above-described video generation method.

[0033] In this embodiment, the model for video generation includes a motion encoder and a video generation network. During the training phase, this scheme uses a joint training mechanism to collaboratively optimize the motion encoder and the video generation network. Specifically, this scheme extracts one-dimensional motion features from the driving video using the motion encoder and uses these features as action control signals. This one-dimensional feature structure effectively filters two-dimensional spatial layout information, ensuring that the action control signals only represent the motion state of the object and are completely decoupled from the camera viewpoint of the driving video. Based on this decoupling design, this scheme achieves the separation of action behavior in the driving video from the original camera viewpoint. Therefore, when the trained model performs the video generation task during the inference phase, the camera viewpoint of the generated video is no longer limited to the camera viewpoint of the driving video. This makes it possible to freely control the camera viewpoint of the generated video, significantly improving the flexibility of video generation.

[0034] In addition to the aforementioned motion control signals, image features are extracted from the first reference image during the training phase and used as additional control signals. Given these control signals, this scheme performs denoising processing on the noisy video features through a video generation network. The noisy video features are obtained by adding noise to the first video features (extracted from the first target video), which uses the first object in the first reference image as the subject and is used to replicate the actions of the second object in the driving video. Thus, based on the first and second video features, this scheme can update the parameters of the motion encoder and the video generation network. This training method, through multi-signal collaborative constraints and end-to-end training with denoising, allows the model to fully learn the motion representation capabilities of the motion control signals and the subject feature restoration capabilities of the image features. Simultaneously, combined with the supervision of the first video features, it achieves accurate feature mapping and parameter optimization. This enables the trained model to accurately capture the core motion patterns of the actions in the driving video and stably restore the subject features of the object in the reference image. Furthermore, due to the decoupling between the action behavior and the camera perspective in the driving video, the model possesses stronger generalization ability and flexible perspective adjustment capabilities during the inference phase.

[0035] In summary, this solution can achieve high-quality video generation that closely matches user expectations.

[0036] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0037] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0038] Figure 1 This is a schematic diagram illustrating an implementation environment for video generation according to an exemplary embodiment.

[0039] Figure 2 This is a flowchart illustrating a model training method for video generation according to an exemplary embodiment.

[0040] Figure 3 This is a schematic diagram of the architecture of a model for video generation according to an exemplary embodiment.

[0041] Figure 4 This is a flowchart illustrating another model training method for video generation according to an exemplary embodiment.

[0042] Figure 5 This is a schematic diagram illustrating a single-view motion reconstruction and cross-view generation according to an exemplary embodiment.

[0043] Figure 6 This is a flowchart illustrating a video generation method according to an exemplary embodiment.

[0044] Figure 7 This is a schematic diagram illustrating a video generation process according to an exemplary embodiment.

[0045] Figure 8 This is a block diagram illustrating a model training apparatus for video generation according to an exemplary embodiment.

[0046] Figure 9 This is a block diagram illustrating a video generation apparatus according to an exemplary embodiment.

[0047] Figure 10 This is a block diagram illustrating a terminal according to an exemplary embodiment. Detailed Implementation

[0048] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0049] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0050] The information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this disclosure are all authorized by the user or fully authorized by the parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0051] The following section will introduce some terms or abbreviations used in the embodiments of this disclosure.

[0052] DiT (Diffusion Transformer) is a generative model that combines the Transformer architecture with a diffusion model. Its core is replacing the U-Net convolutional backbone in traditional diffusion models with a Transformer. DiT follows the forward noise addition and inverse denoising paradigm of DDPM (Denoising Diffusion Probabilistic Models) while incorporating the sequence modeling capabilities of the Transformer. In this embodiment, DiT is used as the basic backbone network for video generation.

[0053] SMPL (Skinned Multi-Person Linear Model): A parameterized 3D human body model that accurately represents the 3D geometric structure of the human body through shape and pose parameters. In this embodiment, the model is introduced only as an auxiliary geometric supervision module in the early stages of training, and the supervision constraint is gradually removed through an annealing strategy as the training progresses.

[0054] VAE (Variational Autoencoder): A generative neural network model that uses an encoder-decoder architecture to learn probabilistic representations of high-dimensional data. It is often used to compress high-dimensional image or video data into a low-dimensional, continuous latent space. In this embodiment, VAE is used to encode video frames included in the driving video.

[0055] Cross-Attention: An attention mechanism in Transformer networks used to fuse features from different modalities. In this embodiment, the cross-attention mechanism can be used to simultaneously inject multiple control signals (such as text control signals and action control signals) into DiT.

[0056] This disclosure proposes a video generation scheme based on implicit space motion representation, automatic alignment of character orientation, and free control of camera viewpoint (including shooting angle and camera movement). During the training phase, this scheme jointly trains a motion encoder and a video generation network (such as the DiT architecture, which is a pre-trained model and only undergoes parameter fine-tuning during training). Furthermore, the motion encoder compresses the driving video (two-dimensional frame sequence) into a view-agnostic one-dimensional motion token. That is, the one-dimensional motion token represents the character's motion behavior in three-dimensional space. Additionally, the motion encoder fits the motion distribution of temporal data related to human movement in the driving video to the implicit space and samples the one-dimensional motion token from the implicit space. Finally, this scheme injects multiple control signals into the video generation network through a cross-attention mechanism.

[0057] In summary, this solution can achieve the following aspects.

[0058] 1. Achieve free control over camera angles: During the inference phase, users can control the camera angles of the generated video using natural language, without being limited by the original camera angles in the driving video.

[0059] 2. Automatic Orientation Alignment: Without the need for additional manual calibration, this solution can automatically adjust the spatial orientation of the actions in the generated video based on the initial orientation of the person in the reference image, so that the generated video is naturally aligned with and connected to the reference image.

[0060] 3. High-fidelity motion transfer: Implicit representations are used to replace two-dimensional and three-dimensional explicit motion signals. This not only preserves richer details of hand and limb dynamics, but also avoids depth ambiguity (such as difficulty in distinguishing the front and back relationships of limbs), thereby achieving accurate, natural and highly expressive generation of controllable character motion videos.

[0061] In summary, this solution not only ensures the generation of high-fidelity videos, but also achieves complete decoupling between the actions in the video and the camera perspective, flexible control of the camera perspective through natural language, and automatic alignment of the actions in the generated video with the orientation of figures in the reference image.

[0062] It should be noted that the application scenarios of this solution include, but are not limited to: digital human dynamic performance in e-commerce and short video fields, virtual dress-up display, film-level pre-show and virtual shooting, special effects motion capture and drive control in game or film production, etc., and this disclosure does not limit these applications.

[0063] The video generation scheme provided in this disclosure will be described in detail below through the following embodiments.

[0064] Figure 1 This is a schematic diagram illustrating an implementation environment for video generation according to an exemplary embodiment. See also... Figure 1 The implementation environment includes: terminal 101 and server 102.

[0065] In this embodiment, terminal 101 is an electronic device used by a user. In this disclosure, a target application with video generation functionality is installed on terminal 101. Exemplarily, the target application is an instant messaging application, such as a short video application; this disclosure does not limit this application.

[0066] In some embodiments, the terminal 101 is a device such as a smartphone, tablet computer, or laptop computer. Figure 1 The example of terminal 101 being a smartphone is provided for illustration only. Furthermore, those skilled in the art will understand that the number of terminals described above can be more or less. For example, there may be only a few terminals, or dozens or hundreds of devices, or even more. This disclosure does not limit the number or type of terminals.

[0067] In other embodiments, server 102 provides background services for the target application. Server 102 runs a model for video generation. This model can be trained on server 102 or on another server; this disclosure does not limit its application to this.

[0068] Additionally, server 102 is connected to terminal 101 via a wireless network or a wired network. Furthermore, server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server; this disclosure does not limit this. Additionally, the servers involved in the embodiments of this disclosure may also include other functional servers to provide more comprehensive and diversified services.

[0069] Figure 2 This is a flowchart illustrating a model training method for video generation according to an exemplary embodiment. Figure 2 As shown, this method is applied to electronic devices, such as... Figure 1 The server 102 shown. The method includes the following steps.

[0070] In step 201, the electronic device acquires a driving video, a first target video, and a first reference image; wherein the first target video uses a first object in the first reference image as the subject of the shot and is used to replicate the action of a second object in the driving video.

[0071] Since this scheme may use multiple datasets for phased training, in order to distinguish between the target videos (also known as generated videos) and reference images included in different datasets, this scheme refers to the target videos included in the first dataset used for the first phase of training as the first target videos and the reference images as the first reference images. In other words, the first dataset includes multiple driving videos, and the target video corresponding to each driving video is referred to as the first target video in this paper. Similarly, the reference image corresponding to each driving video is referred to as the first reference image in this paper.

[0072] in, Figure 3 This is a schematic diagram illustrating the architecture of a model for video generation according to an exemplary embodiment. Figure 3 In the middle, V D It refers to driving video, V tgt It refers to the target video, I R This refers to the reference image. It should be noted that V... tgt These are sample videos acquired through professional collection or in compliance with regulations. The acquisition process can be as follows: capturing real human or digital human motion videos through high-definition shooting or motion capture rendering, and then performing standardized preprocessing such as quality screening, resolution normalization, frame rate alignment, and time series cropping to obtain reference images that serve as supervised ground truth.

[0073] In step 202, the electronic device extracts image features from the first reference image, extracts first video features from the first target video, and performs noise-adding processing on the first video features to obtain noisy video features.

[0074] like Figure 3 As shown, during the training phase, in addition to the reference image I R In addition, each driving video is paired with a natural language text. In other words, regardless of the dataset used in any training phase, the dataset includes multiple quadruples. Each quadruple represents a training sample. Each training sample includes the following four items: driving video, natural language text, reference image, and target video. In this embodiment, to enable users to control the camera angle or video style of the generated video via natural language during the inference phase, natural language text including camera angle instructions and / or video style instructions can be applied to the training process. In other words, the first natural language text is related to at least one of the camera angle or video style of the first target video.

[0075] In 203, the electronic device extracts one-dimensional motion features from the driving video through the motion encoder of the model; wherein, the one-dimensional motion features are used to characterize the motion pattern of the second object.

[0076] In this embodiment of the disclosure, the motion encoder is an implicit motion encoder. The implicit motion encoder fits the motion distribution of temporal data related to human motion in the driving video and maps it to an implicit space, obtaining one-dimensional motion features after sampling from the implicit space.

[0077] In 204, under given control conditions, the electronic device performs denoising processing on the noisy video features through the video generation network of the model to obtain the second video features; wherein, the control conditions include one-dimensional motion features extracted from the first target video and image features extracted from the first reference image.

[0078] Among them, such as Figure 3 As shown, the second video feature refers to the output of the video generation network.

[0079] In some embodiments, in addition to one-dimensional motion features and image features extracted from the first reference image, text features extracted from the first natural language text are also used as control conditions, which are not limited in this disclosure. That is, to achieve controllable generation of natural language during the inference stage, the user can adjust the camera angle or video style of the generated video through natural language commands. In other words, during the training stage, this scheme will use text features extracted from natural language text (including at least one of camera angle commands or video style commands) as additional conditional constraints, and integrate them together with one-dimensional motion features and image features extracted from the reference image into the training process, so that the model learns the accurate mapping from natural language commands to camera angle and video style features.

[0080] In 205, the electronic device updates the parameters of the motion encoder and video generation network of the model based on the first video features and the second video features.

[0081] In summary, during the training phase, this scheme employs a joint training mechanism to collaboratively optimize the motion encoder and the video generation network. Specifically, the motion encoder extracts one-dimensional motion features from the driving video and uses these as the action control signal. This one-dimensional feature structure effectively filters out two-dimensional spatial layout information, ensuring that the action control signal only represents the motion state of the object, completely decoupling it from the camera viewpoint of the driving video. Based on this decoupling design, this scheme separates the action behavior in the driving video from the original camera viewpoint. Therefore, when the trained model performs the video generation task during the inference phase, the camera viewpoint of the generated video is no longer limited to that of the driving video. This makes it possible to freely control the camera viewpoint of the generated video, significantly improving the flexibility of video generation.

[0082] In addition to the aforementioned motion control signals, image features are extracted from the first reference image during the training phase and used as additional control signals. Given these control signals, this scheme performs denoising processing on the noisy video features through a video generation network. The noisy video features are obtained by adding noise to the first video features (extracted from the first target video), which uses the first object in the first reference image as the subject and is used to replicate the actions of the second object in the driving video. Thus, based on the first and second video features, this scheme can update the parameters of the motion encoder and the video generation network. This training method, through multi-signal collaborative constraints and end-to-end training with denoising, allows the model to fully learn the motion representation capabilities of the motion control signals and the subject feature restoration capabilities of the image features. Simultaneously, combined with the supervision of the first video features, it achieves accurate feature mapping and parameter optimization. This enables the trained model to accurately capture the core motion patterns of the actions in the driving video and stably restore the subject features of the object in the reference image. Furthermore, due to the decoupling between the action behavior and the camera perspective in the driving video, the model possesses stronger generalization ability and flexible perspective adjustment capabilities during the inference phase.

[0083] The above Figure 2 The diagram shown is merely the basic process of this disclosure. The following section will further elaborate on the solution provided in this disclosure based on a specific implementation method. Figure 4 This is a flowchart illustrating another model training method for video generation according to an exemplary embodiment, such as... Figure 4 As shown, this video generation method is applied to electronic devices, such as... Figure 1 The server 102 shown. The method includes the following steps.

[0084] In 401, the electronic device acquires driving video, first target video, first reference image, and first natural language text.

[0085] To enable the model to truly understand motion in three-dimensional space, this approach employs a phased training strategy during the training phase, using pre-collected training data covering multiple camera angles.

[0086] Figure 5 This is a schematic diagram illustrating single-view motion reconstruction and cross-view generation according to an exemplary embodiment. For example... Figure 5 As shown, this scheme achieves progressive training from single-view motion reconstruction to cross-view generation. Steps 401-405 below describe the first stage of the training process, which is related to... Figure 5 This corresponds to the single-view motion reconstruction section. Steps 406 to 407 below describe the second phase of the training process, which is related to... Figure 5 The cross-perspective generation part corresponds to this.

[0087] For the first stage of training, the first target video uses the first object in the first reference image as the main subject, and is used to replicate the actions of the second object in the driving video, as well as the camera perspective of the first reference image or the first natural language text. It should be noted that since the purpose of this stage is to enable the model to learn and replicate the dynamic features of the actions in the driving video, the camera perspective of the first reference image and the camera perspective indicated by the first natural language text are consistent with the camera perspective of the driving video.

[0088] In step 402, the electronic device extracts image features from the first reference image, extracts text features from the first natural language text, extracts first video features from the first target video, and performs noise-adding processing on the first video features to obtain noisy video features.

[0089] In some embodiments, such as Figure 3 As shown, the model for video generation includes a motion encoder, a video generation network, a text encoder, and a VAE. The electronic device extracts text features from a first natural language text using the text encoder, extracts image features from a first reference image using the VAE encoder, and extracts first video features from a first target video using the VAE encoder; this disclosure does not limit the scope of the invention.

[0090] It should be noted that VAEs are usually pre-trained, and their parameters are not updated during the training phase.

[0091] In other embodiments, such as Figure 3 As shown, adding noise to the first video feature can be done by adding Gaussian random noise to the first video feature, and this disclosure does not limit this. Alternatively, taking DiT as an example, the first video feature can be subjected to multi-step noise addition through the forward noise addition process of DiT until the first video feature becomes pure noise, that is, the noisy video feature is a noise feature, and this disclosure does not limit this.

[0092] In 403, the electronic device extracts one-dimensional motion features from the driving video using the motion encoder of the model; wherein the one-dimensional motion features are used to characterize the motion pattern of the second object.

[0093] In some embodiments, the motion encoder of the model extracts one-dimensional motion features from the driving video, including the following steps 4031-4032.

[0094] 4031. Perform data augmentation processing on the video frames included in the driving video to obtain the processed video frames.

[0095] As an example, data augmentation is performed on video frames included in the driving video to obtain processed video frames, including at least one of the following: A. Perform random perspective transformation on the video frames included in the driving video to obtain the processed time-frequency frames.

[0096] Among them, performing random perspective transformation on the video frames contained in the driving video can be achieved by performing random homography matrix mapping on the pixel space of the video frames, and by randomly sampling the core parameters of the perspective transformation, controllable random perspective distortion can be achieved on the video frames contained in the driving video.

[0097] B. Perform color dithering on the video frames included in the driving video to obtain the processed time-frequency frames.

[0098] Color jitter processing includes adjusting the brightness, contrast, saturation, and hue of video frames.

[0099] It should be noted that data augmentation can disrupt the two-dimensional pixel layout of the driving video, forcing the motion encoder to learn essential spatial motion signals that do not depend on specific camera angles and character identities, thereby improving the model's generalization ability.

[0100] In addition, to extract pure motion information from the driving video and remove interference from background and camera perspective information, this solution employs a dual-scale motion encoder. See step 4032 below for details.

[0101] 4032. Extract local motion features and global motion features from the processed video frames.

[0102] Local motion features are used to characterize the motion pattern of a single part of the second object; global motion features are used to indicate the motion patterns of multiple parts of the second object. As an example, such as... Figure 3 As shown, the single part mentioned above refers to the hand area, and the multiple parts mentioned above refer to the limbs of a person. As another example, the single part mentioned above refers to the hand area or the head area, and the multiple parts mentioned above refer to the entire body area of ​​a person.

[0103] In some embodiments, such as Figure 3 As shown, the dual-scale motion encoder includes a first motion encoder (hand motion encoder). ) and second motion encoder (limb motion encoder) Correspondingly, extracting local motion features and global motion features from the processed video frames refers to: using the first motion encoder ( The processed video frames are mapped to the implicit space to obtain hand motion features. ; and, via the second motion encoder ( The processed video frames are mapped to the implicit space to obtain limb movement features. .

[0104] For example, both encoders employ a Transformer architecture, capable of compressing the input two-dimensional frame sequence into a compact one-dimensional motion token. This one-dimensional feature structure effectively filters two-dimensional spatial layout information, ensuring that the motion control signal only represents the motion state of the person, completely decoupled from the camera's perspective in the driving video. Furthermore, mapping the processed video frames to the implicit space means that the two motion encoders fit the motion distribution of the temporal data related to human motion in the driving video, learn the probability distribution patterns of the motion data, and sample from the fitted motion distribution to obtain motion features representing the single-dimensional motion pattern, i.e., one-dimensional motion features.

[0105] In addition, this scheme uses a dual-scale motion encoder to extract features in both local and global dimensions. For example, a limb motion encoder is used to extract limb motion features, and a hand motion encoder is used to extract gesture motion features, which can solve the problems of difficulty and ambiguity in estimating overall motion and hand motion.

[0106] In 404, under given control conditions, the electronic device performs denoising processing on the noisy video features through the video generation network of this model to obtain the second video features; wherein, the control conditions include extracted one-dimensional motion features, image features and text features.

[0107] In this embodiment of the disclosure, the control conditions, i.e., the injection of each control signal, include any one of the following: 1. Based on the cross-attention mechanism, each control condition is used as a key and value to input into the reverse denoising process of the video generation network.

[0108] This approach involves injecting all control signals into the reverse denoising process of the video generation network based on a cross-attention mechanism, such as into the layers of the DiT generator.

[0109] 2. The noisy video features and various control conditions are spliced ​​together to obtain the first splicing result, which is then input into the reverse denoising process of the video generation network.

[0110] This approach involves using feature-connected methods to inject all control signals, along with the noisy video features, into the reverse denoising process of the video generation network.

[0111] 3. Based on the cross-attention mechanism, the first part of the control conditions is used as the key and value and input into the reverse denoising process of the video generation network; and the noisy video features and the second part of the control conditions are concatenated to obtain the second concatenation result, which is then input into the reverse denoising process of the video generation network.

[0112] This approach is a hybrid of methods 1 and 2 described above. As an example, this scheme is based on a cross-attention mechanism, using one-dimensional motion features and text features as keys and values, inputting them into the inverse denoising process. In other words, the generator of the video generation network simultaneously receives two control signals: one from the text encoder's text features, and the other from the motion encoder's one-dimensional motion tokens. Furthermore, the noisy video features and image features are concatenated to obtain a second concatenated result, which is then input into the inverse denoising process.

[0113] It should be noted that, as mentioned earlier, the one-dimensional motion token is independent of the camera angle driving the video. It only expresses the motion behavior in three-dimensional space and will not interfere with the control of the camera angle by the text features, thus realizing the coexistence of motion replication and camera angle control.

[0114] In 405, the electronic device performs a first-stage parameter update on the motion encoder and video generation network of the model based on the first video features and the second video features.

[0115] Considering that simply training the motion encoder and video generation network together often results in model convergence difficulties, as the model may rely on the video generation network's own capabilities to generate videos, preventing the motion encoder from truly learning spatial motion information, this solution introduces an auxiliary geometric supervision and annealing strategy. Specifically, in the early stages of training, this solution uses a lightweight motion decoder (… Figure 3 In This approach maps one-dimensional motion tokens to human pose parameters, providing geometric prior guidance for implicit motion representation. Furthermore, as training progresses, the weight of the auxiliary geometric loss is gradually reduced until it decays to zero (annealing), allowing the model to overcome the accuracy limitations of SMPL and spontaneously learn more accurate and flexible implicit motion representations in a video data-driven manner.

[0116] In summary, based on the first and second video features, the motion encoder and video generation network of this model undergo a first-stage parameter update, including the following steps: 1. Obtain the first loss based on the first video features and the second video features.

[0117] The first loss is used to supervise the generation process of video features, calculating the difference between the predicted value and the true value. For example, this scheme calculates the first loss using the loss function shown in Formula 1 below.

[0118] Formula 1:

[0119] in, This refers to the video features output by the model, i.e., the second video features; This refers to the video features extracted from the first target video, i.e., the second video features; It refers to the L2 norm; This refers to the first loss.

[0120] 2. Extract the first human pose parameters from the first target video, and map the one-dimensional motion features to the second human pose parameters through a motion decoder; obtain the second loss based on the first human pose parameters and the second human pose parameters.

[0121] The second loss is used to constrain the motion decoder. The predicted human pose parameters are used to provide auxiliary supervision. That is, the second loss aims to constrain the motion representations obtained from the intermediate steps of the model. and It has a more accurate meaning of motion posture. For example, this scheme calculates the second loss using the loss function shown in Formula 2 below.

[0122] Formula 2:

[0123] in, Refers to motion decoder The predicted output, namely the second human pose parameters; It refers to the human pose parameters extracted from the first target video by the pre-trained pose estimator, i.e., the first human pose parameters. This refers to the second loss. For example, , Represents hand posture parameters. Represents limb posture parameters.

[0124] 3. Weighting coefficients based on the second loss Obtain the first loss Second loss The weighted sum yields the total loss. Based on total loss Update the parameters of the motion encoder and video generation network.

[0125] In this embodiment of the disclosure, A linear decay strategy is employed to guide the model to learn the correct human geometry in the early stages of training, while allowing the model to focus entirely on the quality of video generation in the later stages. Its mathematical expression is shown in Formula 3 below.

[0126] Formula 3:

[0127] in, This refers to the current number of training steps; This refers to the preset decay cutoff step number.

[0128] It should be noted that, From step 0 of training until the preset decay cutoff step number Up to this point, its value linearly decays from 1 to 0, and in subsequent training steps... It is always 0.

[0129] In addition, this solution calculates the total loss using the following formula 4. .

[0130] Formula 4:

[0131] In some embodiments, to address the problem of mismatch between the orientation of the person in the reference image and the direction of the driving video action that may occur during the inference stage, and to achieve the effect of automatically aligning the generated action with the orientation of the person in the reference image, this scheme uses the first frame of the first target video as the reference image during the training stage, and uses it as the spatiotemporal reference anchor point for the entire video sequence.

[0132] In other words, the first reference image is the first frame of the first target video. This setting enables a unified alignment of the reference construction logic between the training and inference phases. For example, if the character in the driving video is walking towards the front of the screen, while the character in the reference image is facing the left side of the screen, the model can automatically align the spatial motion signals extracted from the driving video to the character's orientation in the reference image, based on this reference construction mechanism. This ensures that the character in the generated video starts facing the left side of the screen, without requiring additional manual calibration.

[0133] It should be noted that after the motion encoder has learned basic motion encoding capabilities, this approach will use a different dataset to begin the second phase of training. During training, the goal is to require the model to generate a target video with a camera viewpoint B based on a driving video with camera viewpoint A. This forces the one-dimensional motion token to contain actual actions in three-dimensional space, rather than simply limb coordinate information on a two-dimensional plane.

[0134] In 406, the electronic device acquires a second target video, a second reference image, and second natural language text.

[0135] The second reference image has a different camera angle than the driving video, but it is consistent with the camera angle indicated by the second natural language text. The second target video takes the first object in the second reference image as the subject of the video and is used to replicate the camera angle of the second reference image and the action of the second object in the driving video.

[0136] To distinguish between target videos and reference images in different datasets, this scheme refers to the target videos in the second dataset used for the second-stage training as "second target videos" and the reference images as "second reference images." In other words, the second dataset includes multiple driving videos, and the target video corresponding to each driving video is referred to as the "second target video" in this paper. Similarly, the reference image corresponding to each driving video is referred to as the "second reference image" in this paper. Furthermore, the driving videos included in the second dataset and the first dataset can have identical content, differing only in camera angle.

[0137] In 407, after the parameter update in the first stage, the electronic device performs a second-stage parameter update on the motion encoder and the video generation network based on the driving video, the second target video, the second reference image, and the second natural language text.

[0138] This step can be referred to as steps 402-405 above, and will not be repeated here.

[0139] For example, the following is based on Figure 5 This paper introduces single-view motion reconstruction and cross-view generation.

[0140] exist Figure 6 middle, This refers to the driver video provided by the user. This refers to the final generated video. This refers to the reference image used for provision. Among them, Figure 6 Two reference images are shown. The upper reference image corresponds to single-view motion reconstruction. In this case, the user gives a natural language instruction: "Frontal view, camera stay still." As mentioned earlier, The included two-dimensional video frames are compressed into compact one-dimensional motion tokens, thus obtaining implicit representations. After encoding the reference image to obtain image features and encoding the natural language text to obtain text features, these three are injected as control conditions into the video generation network to guide video generation. Finally, the model outputs... Figure 6 The video shown in the upper right corner. In this image, the actions of the characters in the generated video are a complete replica of those in the driving video. Furthermore, the camera angle of the generated video is aligned not only with the user-provided instructions and the reference image above, but also with the camera angle in the driving video.

[0141] exist Figure 6In the middle, the reference image at the bottom corresponds to cross-viewpoint generation. In this case, the user provides a natural language instruction such as "right front view, camera moves in a right arc." After encoding the reference image to obtain image features and encoding the natural language text to obtain text features, these three are injected as control conditions into the video generation network to guide video generation. Finally, the model outputs... Figure 6 The video shown in the lower right corner. In this image, the actions of the characters in the generated video are a complete replica of those in the driving video. Furthermore, the camera angle of the generated video is aligned with the user-provided instructions and the reference image above, but the camera angle of the generated video differs from that of the driving video.

[0142] In summary, during the training phase, this scheme employs a joint training mechanism to collaboratively optimize the motion encoder and the video generation network. Specifically, the motion encoder extracts one-dimensional motion features from the driving video and uses these as the action control signal. This one-dimensional feature structure effectively filters out two-dimensional spatial layout information, ensuring that the action control signal only represents the motion state of the object, completely decoupling it from the camera viewpoint of the driving video. Based on this decoupling design, this scheme separates the action behavior in the driving video from the original camera viewpoint. Therefore, when the trained model performs the video generation task during the inference phase, the camera viewpoint of the generated video is no longer limited to that of the driving video. This makes it possible to freely control the camera viewpoint of the generated video, significantly improving the flexibility of video generation.

[0143] In addition to the aforementioned motion control signals, image features are extracted from the first reference image during the training phase and used as additional control signals. Given these control signals, this scheme performs denoising processing on the noisy video features through a video generation network. The noisy video features are obtained by adding noise to the first video features (extracted from the first target video), which uses the first object in the first reference image as the subject and is used to replicate the actions of the second object in the driving video. Thus, based on the first and second video features, this scheme can update the parameters of the motion encoder and the video generation network. This training method, through multi-signal collaborative constraints and end-to-end training with denoising, allows the model to fully learn the motion representation capabilities of the motion control signals and the subject feature restoration capabilities of the image features. Simultaneously, combined with the supervision of the first video features, it achieves accurate feature mapping and parameter optimization. This enables the trained model to accurately capture the core motion patterns of the actions in the driving video and stably restore the subject features of the object in the reference image. Furthermore, due to the decoupling between the action behavior and the camera perspective in the driving video, the model possesses stronger generalization ability and flexible perspective adjustment capabilities during the inference phase.

[0144] Figure 6 This is a flowchart illustrating a video generation method according to an exemplary embodiment. The method is applied to an electronic device, such as one powered by... Figure 1 The terminal 101 and server 102 shown jointly execute this method. For example... Figure 6 As shown, the method includes the following steps.

[0145] In 601, the electronic device acquires driving video, natural language text input by the user, and reference images.

[0146] In this embodiment of the disclosure, taking an electronic device as an example, the terminal can directly input natural language text through the input box on the display screen, or it can acquire the user's voice in natural language form through the microphone; this disclosure does not limit this. Taking the user input as voice as an example, the terminal will first convert the voice input into text form.

[0147] In 602, the electronic device invokes a model for video generation, generating a target video based on the driving video, natural language text, and reference images.

[0148] It should be noted that the model for video generation mentioned in this step is based on Figure 2 or Figure 4 The model is generated using the corresponding model training method. This model is responsible for performing the video generation task.

[0149] As an example, after acquiring natural language text, driving video, and reference images, the terminal sends these to the server. The server then calls a trained model to generate a target video based on the driving video, natural language text, and reference images. The server then returns the generated target video to the terminal, which finally presents it to the user. This disclosure does not limit the scope of the example. For instance, the video features output by the trained video generation network are decoded by a VAE to obtain the target video; this disclosure also does not limit the scope of the example.

[0150] Specifically, the target video uses objects in the reference image as the main subject, replicates the orientation of objects included in the reference image, and drives the movements of objects included in the video. Furthermore, the camera angle of the target video is consistent with the camera angle of the reference image, or the camera angle of the target video is consistent with the camera angle indicated by the natural language text.

[0151] For example, the following is based on Figure 7 The inference process of the model used for video generation is introduced.

[0152] exist Figure 7 middle, Refers to the reference image provided. This refers to user-provided video feed, along with user-defined natural language commands such as "frontal view, zoom out camera." As mentioned earlier, The included two-dimensional video frames are compressed into compact one-dimensional motion tokens, thus obtaining implicit representations. In the reference image... After encoding image features and encoding natural language text features, these three elements are injected as control conditions into the video generation network to guide video generation. Ultimately, the model outputs... Figure 7 The video shown at the bottom is the generated video. For example... Figure 7 As shown, the actions of the characters in the generated video are completely replicated from those in the driving video, but the characters' orientation is aligned with the orientation of the characters in the reference image. Additionally, the camera movement in the generated video is aligned with the user-provided natural language commands, achieving a zoom-out effect.

[0153] In summary, the video generation scheme provided in this disclosure, based on a model for video generation, can achieve the following aspects.

[0154] User-friendly camera perspective control: Users do not need to have computer graphics knowledge. They only need to input simple natural language (such as slowly zooming out or panning to the right) to change the camera perspective while replicating actions, which greatly reduces the threshold for video creation.

[0155] Automatic alignment without manual intervention: By using the reference image as the anchor point mechanism for the first frame of the sequence, the spatial alignment problem between the orientation of the person in the reference image and the driving action is solved, avoiding the drawbacks of incorrect person orientation or the need for manual calibration.

[0156] Rich and natural motion details: Implicit motion representations break free from the degree of freedom constraints of parametric models, and can more accurately reproduce hand grasping details, high-frequency motion features, and large-scale difficult movements that are difficult to capture by parametric models, such as splits and bending over.

[0157] It should be noted that this solution can also be extended to other application scenarios. For example, given a handheld camera video of a character with shaky movements, the video itself can be selected as the driving video, and the character in the first frame of the video can be used as the reference image. Given the text command "static camera movement," a video replicating the same action and featuring the same character can be generated, but with the camera stationary, achieving video stabilization. As another example, by inputting a single character image as the reference image, a driving video with fixed movements can be generated by copying this static reference frame. Combined with natural language commands describing camera movements such as "circling camera movement" as constraints, a new 3D perspective of the target character can be generated. The generated multi-view videos can serve as the basic data support for subsequent 3D character pose reconstruction tasks.

[0158] Figure 8 This is a block diagram illustrating a model training apparatus for video generation according to an exemplary embodiment. The model includes a motion encoder and a video generation network. (Refer to...) Figure 8 The device includes: The first acquisition module 801 is configured to acquire a driving video, a first target video, and a first reference image; wherein the first target video uses a first object in the first reference image as the subject of the shot and is used to replicate the action of a second object in the driving video. The first encoding module 802 is configured to extract image features from the first reference image and extract first video features from the first target video; The noise-adding module 803 is configured to perform noise-adding processing on the first video feature to obtain noisy video features; The second encoding module 804 is configured to extract one-dimensional motion features from the driving video through the motion encoder, the one-dimensional motion features being used to characterize the motion pattern of the second object; The first generation module 805 is configured to perform denoising processing on the noisy video features through the video generation network under given control conditions to obtain second video features; wherein, the control conditions include the one-dimensional motion features and the image features; Training module 806 is configured to update the parameters of the motion encoder and the video generation network based on the first video features and the second video features.

[0159] In this embodiment, the model for video generation includes a motion encoder and a video generation network. During the training phase, this scheme uses a joint training mechanism to collaboratively optimize the motion encoder and the video generation network. Specifically, this scheme extracts one-dimensional motion features from the driving video using the motion encoder and uses these features as action control signals. This one-dimensional feature structure effectively filters two-dimensional spatial layout information, ensuring that the action control signals only represent the motion state of the object and are completely decoupled from the camera viewpoint of the driving video. Based on this decoupling design, this scheme achieves the separation of action behavior in the driving video from the original camera viewpoint. Therefore, when the trained model performs the video generation task during the inference phase, the camera viewpoint of the generated video is no longer limited to the camera viewpoint of the driving video. This makes it possible to freely control the camera viewpoint of the generated video, significantly improving the flexibility of video generation.

[0160] In addition to the aforementioned motion control signals, image features are extracted from the first reference image during the training phase and used as additional control signals. Given these control signals, this scheme performs denoising processing on the noisy video features through a video generation network. The noisy video features are obtained by adding noise to the first video features (extracted from the first target video), which uses the first object in the first reference image as the subject and is used to replicate the actions of the second object in the driving video. Thus, based on the first and second video features, this scheme can update the parameters of the motion encoder and the video generation network. This training method, through multi-signal collaborative constraints and end-to-end training with denoising, allows the model to fully learn the motion representation capabilities of the motion control signals and the subject feature restoration capabilities of the image features. Simultaneously, combined with the supervision of the first video features, it achieves accurate feature mapping and parameter optimization. This enables the trained model to accurately capture the core motion patterns of the actions in the driving video and stably restore the subject features of the object in the reference image. Furthermore, due to the decoupling between the action behavior and the camera perspective in the driving video, the model possesses stronger generalization ability and flexible perspective adjustment capabilities during the inference phase.

[0161] In some embodiments, the second encoding module is configured to: Data enhancement processing is performed on the video frames included in the driving video to obtain the processed video frames; Extract local motion features and global motion features from the processed video frames; The local motion features are used to characterize the motion pattern of a single part of the second object; the global motion features are used to indicate the motion patterns of multiple parts of the second object.

[0162] In other embodiments, the motion encoder includes a first motion encoder and a second motion encoder; the second encoding module is configured to: The processed video frames are mapped to the implicit space using the first motion encoder to obtain the local motion features; The processed video frames are mapped to the implicit space using the second motion encoder to obtain the global motion features.

[0163] In other embodiments, the second encoding module is configured to perform at least one of the following: A random perspective transformation is performed on the video frames included in the driving video to obtain the processed time-frequency frame; Color dithering is performed on the video frames included in the driving video to obtain the processed time-frequency frame; wherein, color dithering includes adjusting brightness, contrast, saturation and hue.

[0164] In other embodiments, the control conditions further include: text features extracted from a first natural language text; wherein the first natural language text relates to at least one of the following: The camera angle of the first target video; The video style of the first target video.

[0165] In other embodiments, the apparatus further includes a condition injection module; the condition injection module is configured to perform any of the following: Based on the cross-attention mechanism, the control conditions are used as keys and values ​​and input into the reverse denoising process of the video generation network. The noisy video features and the control conditions are spliced ​​together to obtain a first splicing result, and the first splicing result is input into the reverse denoising process; Based on the cross-attention mechanism, the first part of the control conditions is used as a key and a value and input into the reverse denoising process; and the noisy video features and the second part of the control conditions are concatenated to obtain a second concatenation result, which is then input into the reverse denoising process.

[0166] In other embodiments, the conditional injection module is configured to: Based on the cross-attention mechanism, the one-dimensional motion features and the text features are used as keys and values, respectively, and input into the reverse denoising process; and, The noisy video features and the image features are concatenated to obtain the second concatenation result, and the second concatenation result is input into the reverse denoising process.

[0167] In other embodiments, the camera angle of the first reference image is consistent with the camera angle of the driving video; the first target video uses the first object as the subject of the shot to replicate the camera angle and the action of the second object; the training module is configured to: Based on the first video features and the second video features, the motion encoder and the video generation network are updated in the first stage.

[0168] In other embodiments, the device further includes: The second acquisition module is configured to acquire a second target video and a second reference image; wherein the camera angle of the second reference image is different from the camera angle of the driving video; the second target video takes the first object as the subject of the shot and is used to replicate the camera angle in the second reference image and the action of the second object; The training module is further configured to perform a second-stage parameter update on the motion encoder and the video generation network based on the driving video, the second target video, and the second reference image after the first-stage parameter update.

[0169] In other embodiments, the model further includes a motion decoder; the training module is configured to: Based on the first video features and the second video features, a first loss is obtained; First human pose parameters are extracted from the first target video, and the one-dimensional motion features are mapped to second human pose parameters through the motion decoder; a second loss is obtained based on the first human pose parameters and the second human pose parameters. Based on the weight coefficients corresponding to the second loss, the weighted sum of the first loss and the second loss is obtained to obtain the total loss; wherein, the value of the weight coefficient gradually decays to zero after the start of training; Based on the total loss, update the parameters of the motion encoder and the video generation network.

[0170] In other embodiments, the first reference image is the first frame of the first target video.

[0171] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.

[0172] Figure 9 This is a block diagram illustrating a video generation apparatus according to an exemplary embodiment. (Refer to...) Figure 9 The device includes: The third acquisition module 901 is configured to acquire the driving video, the natural language text input by the user, and the reference image; The second generation module 902 is configured to call a model for video generation and generate a target video based on the driving video, the natural language text, and the reference image; The model used for video generation is trained based on the aforementioned model training device; the target video uses the object in the reference image as the subject of the shot, and is used to replicate the orientation of the object included in the reference image and the action of the object included in the driving video; the camera angle of the target video is consistent with the camera angle of the reference image, or the camera angle of the target video is consistent with the camera angle indicated by the natural language text.

[0173] In summary, the video generation scheme provided in this disclosure, based on a model for video generation, can achieve the following aspects.

[0174] User-friendly camera perspective control: Users do not need to have computer graphics knowledge. They only need to input simple natural language (such as slowly zooming out or panning to the right) to change the camera perspective while replicating actions, which greatly reduces the threshold for video creation.

[0175] Automatic alignment without manual intervention: By using the reference image as the anchor point mechanism for the first frame of the sequence, the spatial alignment problem between the orientation of the person in the reference image and the driving action is solved, avoiding the drawbacks of incorrect person orientation or the need for manual calibration.

[0176] Rich and natural motion details: Implicit motion representations break free from the degree of freedom constraints of parametric models, and can more accurately reproduce hand grasping details, high-frequency motion features, and large-scale difficult movements that are difficult to capture by parametric models, such as splits and bending over.

[0177] It should be noted that this solution can also be extended to other application scenarios. For example, given a handheld camera video of a character with shaky movements, the video itself can be selected as the driving video, and the character in the first frame of the video can be used as the reference image. Given the text command "static camera movement," a video replicating the same action and featuring the same character can be generated, but with the camera stationary, achieving video stabilization. As another example, by inputting a single character image as the reference image, a driving video with fixed movements can be generated by copying this static reference frame. Combined with natural language commands describing camera movements such as "circling camera movement" as constraints, a new 3D perspective of the target character can be generated. The generated multi-view videos can serve as the basic data support for subsequent 3D character pose reconstruction tasks.

[0178] It should be noted that the video generation device provided in the above embodiments, when generating video, or the model training device for video generation provided in the above embodiments, when training a model, is only illustrating the division of the above functional units. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. Furthermore, the video generation device and the video generation method embodiments provided in the above embodiments belong to the same concept, and the model training device and the model training method embodiments provided in the above embodiments belong to the same concept. For details of their specific implementation, please refer to the method embodiments, which will not be repeated here.

[0179] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0180] In some embodiments, when the electronic device is provided as a server Figure 10This is a block diagram illustrating a server 1000 according to an exemplary embodiment. The server 1000 can vary considerably depending on its configuration or performance, and includes one or more Central Processing Units (CPUs) 1001 and one or more memories 1002. The memories 1002 store at least one line of program code, which is loaded and executed by the processor 1001 to implement the aforementioned model training method for video generation or the aforementioned video generation method. Of course, the server also has wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1000 also includes other components for implementing device functions, which will not be elaborated here.

[0181] In some embodiments, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, which are executed by a processor of an electronic device to implement the above-described model training method for video generation or the above-described video generation method. In other embodiments, the computer-readable storage medium is a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0182] In some embodiments, a computer program product is also provided, including a computer program that, when executed by a processor of an electronic device, implements the model training method for video generation described above or the video generation method described above.

[0183] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0184] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A model training method for video generation, characterized in that, The model includes a motion encoder and a video generation network; the method includes: Acquire a driving video, a first target video, and a first reference image; wherein the first target video uses a first object in the first reference image as the subject of the shot and is used to replicate the action of a second object in the driving video; Image features are extracted from the first reference image, and first video features are extracted from the first target video. Noise is added to the first video features to obtain noisy video features. The motion encoder extracts one-dimensional motion features from the driving video, and the one-dimensional motion features are used to characterize the motion pattern of the second object. Under given control conditions, the noisy video features are denoised by the video generation network to obtain second video features; wherein, the control conditions include the one-dimensional motion features and the image features; Based on the first video features and the second video features, the parameters of the motion encoder and the video generation network are updated.

2. The model training method according to claim 1, characterized in that, The extraction of one-dimensional motion features from the driving video using the motion encoder includes: Data enhancement processing is performed on the video frames included in the driving video to obtain the processed video frames; Extract local motion features and global motion features from the processed video frames; The local motion features are used to characterize the motion pattern of a single part of the second object; the global motion features are used to indicate the motion patterns of multiple parts of the second object.

3. The model training method according to claim 2, characterized in that, The motion encoder includes a first motion encoder and a second motion encoder; the extraction of local motion features and global motion features from the processed video frames includes: The processed video frames are mapped to the implicit space using the first motion encoder to obtain the local motion features; The processed video frames are mapped to the implicit space using the second motion encoder to obtain the global motion features.

4. The model training method according to claim 2, characterized in that, The data enhancement processing of the video frames included in the driving video to obtain the processed video frames includes at least one of the following: A random perspective transformation is performed on the video frames included in the driving video to obtain the processed time-frequency frame; Color dithering is performed on the video frames included in the driving video to obtain the processed time-frequency frame; wherein, color dithering includes adjusting brightness, contrast, saturation and hue.

5. The model training method according to claim 1, characterized in that, The control conditions further include: text features extracted from the first natural language text; wherein the first natural language text is related to at least one of the following: The camera angle of the first target video; The video style of the first target video.

6. The model training method according to claim 1, characterized in that, The first reference image is the first frame of the first target video.

7. The model training method according to claim 5, characterized in that, The method further includes any one of the following: Based on the cross-attention mechanism, the control conditions are used as keys and values ​​and input into the reverse denoising process of the video generation network. The noisy video features and the control conditions are spliced ​​together to obtain a first splicing result, and the first splicing result is input into the reverse denoising process; Based on the cross-attention mechanism, the first part of the control conditions is used as the key and value and input into the reverse denoising process; Furthermore, the noisy video features and the second part of the control conditions are spliced ​​together to obtain a second splicing result, and the second splicing result is input into the reverse denoising process.

8. The model training method according to claim 7, characterized in that, Based on the cross-attention mechanism, the first part of the control conditions is used as a key and a value, and then input into the reverse denoising process. And, the noisy video features and the second part of the control conditions are spliced ​​together to obtain a second splicing result, and the second splicing result is input into the reverse denoising process, including: Based on the cross-attention mechanism, the one-dimensional motion features and the text features are used as keys and values, and then input into the reverse denoising process. as well as, The noisy video features and the image features are concatenated to obtain the second concatenation result, and the second concatenation result is input into the reverse denoising process.

9. The model training method according to claim 1, characterized in that, The camera angle of the first reference image is consistent with the camera angle of the driving video; the first target video takes the first object as the subject of the shot and is used to replicate the camera angle and the action of the second object. The step of updating the parameters of the motion encoder and the video generation network based on the first video features and the second video features includes: Based on the first video features and the second video features, the motion encoder and the video generation network are updated in the first stage.

10. The model training method according to claim 9, characterized in that, Following the parameter update in the first phase, the method further includes: Acquire a second target video and a second reference image; wherein the camera angle of the second reference image is inconsistent with the camera angle of the driving video; the second target video takes the first object as the subject of the shot and is used to replicate the camera angle of the second reference image and the action of the second object; Based on the driving video, the second target video, and the second reference image, the motion encoder and the video generation network are updated in the second stage.

11. The model training method according to claim 1, characterized in that, The model further includes a motion decoder; updating the parameters of the motion encoder and the video generation network based on the first video features and the second video features includes: Based on the first video features and the second video features, a first loss is obtained; First human pose parameters are extracted from the first target video, and the one-dimensional motion features are mapped to second human pose parameters through the motion decoder; a second loss is obtained based on the first human pose parameters and the second human pose parameters. Based on the weight coefficients corresponding to the second loss, the weighted sum of the first loss and the second loss is obtained to obtain the total loss; wherein, the value of the weight coefficient gradually decays to zero after the start of training; Based on the total loss, update the parameters of the motion encoder and the video generation network.

12. A video generation method, characterized in that, The method includes: Acquire driving video, user-input natural language text, and reference images; A model for video generation is invoked to generate a target video based on the driving video, the natural language text, and the reference image; wherein the model for video generation is trained using the model training method described in claims 1 to 11; the target video uses an object in the reference image as the subject of the shot and is used to replicate the orientation of the object included in the reference image and the action of the object included in the driving video; the camera angle of the target video is consistent with the camera angle of the reference image, or the camera angle of the target video is consistent with the camera angle indicated by the natural language text.

13. A model training apparatus for video generation, characterized in that, The model includes a motion encoder and a video generation network; the device includes: The first acquisition module is configured to acquire a driving video, a first target video, and a first reference image; wherein the first target video uses a first object in the first reference image as the subject of the shot and is used to replicate the action of a second object in the driving video. The first encoding module is configured to extract image features from the first reference image and extract first video features from the first target video; The noise-adding module is configured to perform noise-adding processing on the first video features to obtain noisy video features; The second encoding module is configured to extract one-dimensional motion features from the driving video through the motion encoder, the one-dimensional motion features being used to characterize the motion pattern of the second object; The first generation module is configured to perform denoising processing on the noisy video features through the video generation network under given control conditions to obtain the second video features; wherein the control conditions include the one-dimensional motion features and the image features; The training module is configured to update the parameters of the motion encoder and the video generation network based on the first video features and the second video features.

14. A video generation apparatus, characterized in that, The device includes: The third acquisition module is configured to acquire the driving video, the natural language text input by the user, and the reference image; The second generation module is configured to invoke a model for video generation to generate a target video based on the driving video, the natural language text, and the reference image; wherein the model for video generation is trained using the model training apparatus shown in claims 1 to 11; the target video uses an object in the reference image as the subject of the shot and is used to replicate the orientation of the object included in the reference image and the action of the object included in the driving video; the camera angle of the target video is consistent with the camera angle of the reference image, or the camera angle of the target video is consistent with the camera angle indicated by the natural language text.

15. An electronic device, characterized in that, The electronic device includes: One or more processors; Memory used to store the executable program code of the processor; The processor is configured to execute the program code to implement the model training method for video generation as described in any one of claims 1 to 11; or the video generation method as described in claim 12.

16. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of an electronic device, the electronic device is enabled to perform the model training method for video generation as described in any one of claims 1 to 11; or, the video generation method as described in claim 12.

17. A computer program product, characterized in that, The method includes a computer program that, when executed by a processor of an electronic device, implements the model training method for video generation as described in any one of claims 1 to 11; or, the video generation method as described in claim 12.