Portrait video generation method and device, and electronic device
By acquiring audio and user instructions, and combining a pre-trained video generation model and a 3D face model, the problem of generating specific head movements in existing technologies has been solved, achieving efficient and accurate portrait video generation and meeting the precise control requirements of film and television.
Patent Information
- Application Number
- CN202511494591.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-10-20
AI Technical Summary
Existing portrait video generation technologies struggle to efficiently and accurately generate videos that perform specific head movements, especially in scenarios such as film and television production and e-commerce live streaming, failing to meet the needs of virtual characters to perform complex and specific head movements.
By acquiring audio information, reference images, and user instructions, and combining them with a pre-trained video generation model, a portrait video matching the head movements of the target person with the user instructions is generated using a progressive focus training strategy and a 3D face model.
It enables the efficient and accurate generation of portrait videos that perform specific head movements, preserving natural voice-related actions while meeting the precise control requirements of film and television, thus improving the realism and controllability of video generation.
Smart Images

Figure CN120956978B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of portrait video generation technology, and in particular to a portrait video generation method, apparatus and electronic device. Background Technology
[0002] In the field of portrait video generation, related technologies can use voice audio to drive the generation of animation from static portrait images. Among them, animation generated based on diffusion models has achieved significant advantages in lip-sync accuracy, facial expression realism, and natural head movement control, bringing new breakthroughs to portrait video generation.
[0003] However, in practical applications such as film and television production and e-commerce live streaming, simply achieving natural portrait video representation is far from sufficient. In these scenarios, virtual characters often have more complex and specific needs, such as needing to perform specific head movements according to a script, like turning towards a conversation partner, nodding, or shaking their head. Existing portrait video generation technologies based on diffusion models are insufficient to meet the needs of generating these specific head movements.
[0004] Therefore, how to efficiently and accurately generate videos of specific head movements has become a hot research topic that urgently needs to be addressed in this field. Summary of the Invention
[0005] This invention provides a portrait video generation method, apparatus, and electronic device, which enables the efficient and accurate generation of portrait videos featuring specific head movements.
[0006] This invention provides a portrait video generation method, comprising: acquiring audio information, a reference image, user instructions, and a noisy image, wherein the reference image is an image containing a target person; the noisy image is an image obtained by adding noise to the reference image; the user instructions are instructions for guiding the head movements of the target person in the generated target portrait video; invoking a pre-trained video generation model, wherein the video generation model includes a reference network and a denoising network, and the video generation model is trained using a progressive focus training strategy; inputting the reference image into the reference network to obtain target person features output by the reference network; obtaining 3D coefficients corresponding to the user instructions based on the user instructions and the audio information, wherein the 3D coefficients are used to characterize the head movement features and facial expression features of the target person; inputting the 3D coefficients, the target person features, and the noisy image into the denoising network to obtain the target portrait video output by the denoising network, wherein the head movements of the target person in the target portrait video match the user instructions.
[0007] According to a portrait video generation method provided by the present invention, the video generation model further includes a 3D face model; the step of obtaining 3D coefficients corresponding to the user instruction based on the user instruction and the audio information specifically includes: obtaining 3D deformation model coefficients based on the user instruction and the audio information, wherein the 3D deformation model coefficients are used to characterize the deformation features of the target person; inputting the 3D deformation model coefficients and the audio information into the 3D face model to obtain the 3D coefficients output by the 3D face model corresponding to the user instruction.
[0008] According to a portrait video generation method provided by the present invention, the video generation model further includes a three-dimensional deformable face model; the step of obtaining 3D deformation model coefficients based on the user instruction and the audio information specifically includes: inputting the user instruction and the audio information into the three-dimensional deformable face model to obtain the 3D deformation model coefficients output by the three-dimensional deformable face model.
[0009] According to a portrait video generation method provided by the present invention, the video generation model is trained through a progressive focus training strategy, which is implemented as follows: A training sample set is obtained, wherein the training sample set includes multiple training samples, and the training samples include audio sample information and user instruction sample information; the training samples are input into a three-dimensional deformable face model to obtain 3D deformation model coefficient samples output by the three-dimensional deformable face model, wherein the 3D deformation model coefficient samples include head motion feature samples and facial expression feature samples; based on the head motion feature samples and the facial expression feature samples, the video generation model is trained through a progressive focus training strategy to obtain a trained video generation model.
[0010] According to a portrait video generation method provided by the present invention, the step of training the video generation model based on the head motion feature samples and the facial expression feature samples using a progressive focusing training strategy to obtain a trained video generation model specifically includes: training the attention layer related to head motion in the video generation model based on the head motion feature samples to obtain a first-stage trained video generation model; freezing the attention layer related to head motion, and training the first-stage trained video generation model based on the facial expression feature samples to obtain a trained video generation model.
[0011] According to a portrait video generation method provided by the present invention, the user instruction includes any one or more of a first user instruction and a second user instruction, wherein the first user instruction is an instruction for generating the head movement of a target person by rotating the target person at a specified axis angle; and the second user instruction is a video or image for representing the head movement of the target person desired by the user.
[0012] The present invention also provides a portrait video generation apparatus, the apparatus comprising: an acquisition module for acquiring audio information, a reference image, a user instruction, and a noisy image, wherein the reference image is an image containing a target person; the noisy image is an image obtained by adding noise to the reference image; the user instruction is an instruction for guiding the head movements of the target person in the generated target portrait video; an invocation module for invoking a pre-trained video generation model, wherein the video generation model includes a reference network and a denoising network, and the video generation model is trained using a progressive focus training strategy; an input module for inputting the reference image into the reference network to obtain the target person features output by the reference network; a processing module for obtaining 3D coefficients corresponding to the user instruction based on the user instruction and the audio information, wherein the 3D coefficients are used to characterize the head movement features and facial expression features of the target person; and a generation module for inputting the 3D coefficients, the target person features, and the noisy image into the denoising network to obtain the target portrait video output by the denoising network, wherein the head movements of the target person in the target portrait video match the user instruction.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the portrait video generation method as described above.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the portrait video generation method as described above.
[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the portrait video generation method as described above.
[0016] This invention provides a portrait video generation method, apparatus, and electronic device that acquires audio information, a reference image, user instructions, and a noisy image. The reference image is an image containing a target person; the noisy image is an image obtained by adding noise to the reference image; the user instructions are commands used to guide the head movements of the target person in the generated portrait video; a pre-trained video generation model is invoked, comprising a reference network and a denoising network, trained using a progressive focus training strategy; the reference image is input into the reference network to obtain the target person features output by the reference network; based on the user instructions and audio information, 3D coefficients corresponding to the user instructions are obtained, where the 3D coefficients characterize the head movement and facial expression features of the target person; the 3D coefficients, the target person features, and the noisy image are input into the denoising network to obtain the target portrait video output by the denoising network, where the head movements of the target person in the target portrait video match the user instructions. By integrating the collaborative control of audio-driven and user-instructed commands, this invention achieves efficient and accurate generation of portrait videos that perform specific head movements, preserving natural voice-related movements while meeting the requirements for precise control at the film and television level. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the portrait video generation method provided by the present invention.
[0019] Figure 2 This is a flowchart illustrating how the present invention obtains 3D coefficients corresponding to the user instruction based on the user instruction and the audio information.
[0020] Figure 3 This is a schematic diagram of the process of training a video generation model using a progressive focus training strategy provided by the present invention.
[0021] Figure 4 This is a schematic diagram of the process provided by the present invention, which trains the video generation model based on the head motion feature samples and the facial expression feature samples through a progressive focusing training strategy to obtain a trained video generation model.
[0022] Figure 5 This is a schematic diagram of the portrait video generation device provided by the present invention.
[0023] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0025] The portrait video generation method provided by this invention integrates the collaborative control of audio driving and user commands. It generates basic motion trajectories based on audio, and allows real-time editing based on rotation parameters, video, or images through user instructions. This hybrid mode of "AI generation + human refinement" preserves natural voice-related movements (such as stress corresponding to eyebrow raising) while meeting the requirements of precise control at the film and television level.
[0026] Figure 1 This is one of the flowcharts illustrating the portrait video generation method provided by the present invention.
[0027] The following will combine Figure 1 The portrait video generation method provided by this invention will be described.
[0028] In an exemplary embodiment of the present invention, combined with Figure 1 As can be seen, the portrait video generation method may include steps 110 to 150, and each step will be described below.
[0029] In step 110, audio information, a reference image, user instructions, and a noisy image are acquired, wherein the reference image is an image containing the target person; the noisy image is an image obtained by adding noise to the reference image; and the user instructions are instructions used to guide the head movements of the target person in the generated target portrait video.
[0030] In one embodiment, audio information, a reference image, user instructions, and a noisy image can be acquired. The audio information can be a voice recording (such as a WAV or MP3 file), for example, its content could be "Welcome to our channel." The reference image can be a frontal half-body photograph (JPEG format) containing the target person (e.g., a company CEO). This image clearly shows the person's facial features, hairstyle, and clothing. The user instructions can be a text command, such as "nod, smile," or "turn head angle," which guides the head movements and expressions of the target person in the final generated video. The noisy image can be obtained by adding Gaussian noise to the reference image using a computer program (such as a Python library). The level (variance) of added noise can be determined according to a predefined scheduling algorithm (such as the noise scheduler in DDPM), the purpose of which is to provide an input that conforms to the model's expectations at the beginning of the denoising process.
[0031] In step 120, a pre-trained video generation model is invoked. The video generation model includes a reference network and a denoising network. The video generation model is trained using a progressive focus training strategy.
[0032] In another embodiment, a pre-trained video generation model can be loaded. This model can employ a deep learning architecture, comprising two core sub-modules: a reference network and a denoising network. The reference network can be an encoder (such as a CNN-based ResNet), whose function is to extract abstract feature representations from a reference image. The denoising network can be a diffusion model based on a U-Net structure, whose core function is to progressively remove noise from noisy images to generate clear video frames. By separating identity encoding (Reference Net, corresponding to the reference network mentioned above) and motion generation (Denoising Net, corresponding to the denoising network mentioned above) through a dual U-Net architecture, and combining this with FLAME coefficient space transformation, the common identity distortion problem in GAN methods is solved, and motion data can be transferred across identities.
[0033] In another embodiment, the video generation model can be trained using a progressively focused training strategy. Specifically, the training process involves: in the initial stage of training, coarse-grained training is performed using a high noise level and diverse training data (including various people and scenes) to allow the model to learn general prior knowledge about video generation. In the later stage of training, the noise level is gradually reduced, and more refined and specialized data (such as multi-angle and multi-expression data of specific people) is used to fine-tune the model, allowing it to gradually focus on learning how to generate high-quality, high-fidelity portrait videos of the target person.
[0034] In step 130, the reference image is input into the reference network to obtain the target person features output by the reference network.
[0035] In one embodiment, the acquired reference image can be input into a loaded reference network. The reference network performs multi-layer convolutional processing on the image, ultimately outputting a target person feature. This feature is a high-dimensional tensor that encodes the target person's identity information, including but not limited to the precise shape of facial features, skin color, hairstyle, beard, and other unique identifiers.
[0036] In step 140, based on user instructions and audio information, 3D coefficients corresponding to user instructions are obtained, wherein the 3D coefficients are used to characterize the head movement features and facial expression features of the target person.
[0037] In another embodiment, user instructions such as "nod" and "smile" along with the audio message "Welcome to our channel" can be input into a pre-trained driver signal generation module (which can be another neural network, such as a Transformer model). This module analyzes the prosody, rhythm, and content of the audio and combines it with the semantics of the text instructions to predict a series of time-continuous 3D coefficients. These coefficients can typically be understood as parameters of a 3D face model (such as a 3DMM - 3D Morphable Model), including, for example: head motion features: characterized by coefficients of rotation (Euler angles or quaternions), translation (X, Y, Z coordinates), etc., corresponding to the action of "nodding"; and facial expression features: characterized by blendshape coefficients that control the movement of muscles such as lips, eyebrows, and eyes, corresponding to the expression of "smiling," and synchronized with the audio speech lip movements.
[0038] In step 150, the 3D coefficients, target person features, and noisy image are input into the denoising network to obtain the target portrait video output by the denoising network, wherein the head movements of the target person in the target portrait video match the user's instructions.
[0039] In another embodiment, the generated temporal 3D coefficients, the target person's feature tensor, and the noisy image can be input together into a denoising network. Starting with the noisy image, the denoising network, guided by the target person's features, performs a multi-step iterative denoising process based on the head pose and expression defined by the 3D coefficients for each frame. Each iteration generates a less noisy intermediate image, ultimately outputting a series of continuous, clear video frames—the target portrait video. In this final video, the target person's (e.g., the company CEO) lip movements are perfectly synchronized with the audio "Welcome to our channel," and they accurately perform "nodding" and "smiling" head movements and expressions, perfectly matching the user's instructions. The generated video format can be MP4 or AVI.
[0040] In this embodiment, by introducing user instructions as control signals and converting them together with audio information into precise 3D coefficients (representing head movements and facial expressions), refined and directional control over the head movements and expressions of the person in the generated portrait video is achieved. This solves the problem of random and uncontrollable video movements generated by existing technologies, allowing users to directly guide the generated video content as needed.
[0041] This invention provides a portrait video generation method that acquires audio information, a reference image, user instructions, and a noisy image. The reference image is an image containing the target person; the noisy image is an image obtained by adding noise to the reference image; the user instructions are commands used to guide the head movements of the target person in the generated portrait video; a pre-trained video generation model is invoked, comprising a reference network and a denoising network, trained using a progressive focus training strategy; the reference image is input into the reference network to obtain the target person features output by the reference network; based on the user instructions and audio information, 3D coefficients corresponding to the user instructions are obtained, where the 3D coefficients characterize the head movement and facial expression features of the target person; the 3D coefficients, the target person features, and the noisy image are input into the denoising network to obtain the target portrait video output by the denoising network, where the head movements of the target person in the target portrait video match the user instructions. By integrating the collaborative control of audio-driven and user-instructed commands, this method achieves efficient and accurate generation of portrait videos that perform specific head movements, preserving natural voice-related movements while meeting the requirements for precise control at the film and television level.
[0042] Figure 2 This is a flowchart illustrating how the present invention obtains 3D coefficients corresponding to the user instruction based on the user instruction and the audio information.
[0043] The following will combine Figure 2 The process of obtaining 3D coefficients corresponding to the user instruction based on the user instruction and the audio information provided by the present invention will be described.
[0044] In an exemplary embodiment of the present invention, the video generation model may further include a 3D face model (e.g., a FLAME module); combined with Figure 2 As can be seen, obtaining the 3D coefficients corresponding to the user instruction based on the user instruction and the audio information may include steps 210 and 220, which will be described in detail below.
[0045] In step 210, based on user instructions and audio information, 3D deformation model coefficients are obtained, wherein the 3D deformation model coefficients are used to characterize the deformation features of the target person.
[0046] In step 220, the 3D deformation model coefficients and audio information are input into the 3D face model to obtain the 3D coefficients output by the 3D face model corresponding to the user's instructions.
[0047] In one embodiment, the processing can be based on user instructions (such as "show a surprised expression and slightly shake your head in rhythm with the speech") and audio information, handled by a lightweight analysis network or rule system. This analysis network parses the semantics of the user instructions and initially combines them with the rhythmic information of the audio to output a set of basic 3D deformation model coefficients. These coefficients primarily characterize the deformation features of the target person's face, such as the muscle movement parameters for controlling the raised eyebrows and open mouth required to control a surprised expression, but they are not yet fully integrated with precise head rotation or translation.
[0048] Furthermore, the initially obtained 3D deformation model coefficients and the original audio information are input into the 3D face model (e.g., the FLAME module). The 3D face model, acting as a more accurate and complex physics engine or neural network, performs the following operations:
[0049] By using audio information to drive the deformation of the lips, lip movement coefficients that are precisely matched with speech phonemes are generated, and the time sequence of facial expression coefficients is finely adjusted to synchronize with the audio rhythm.
[0050] Based on the implicit action requirements in the user's instructions (such as "slight head shaking"), and combined with the context of the current frame, the precise head rotation and translation coefficients (i.e., posture parameters) are calculated.
[0051] The finely adjusted facial expression coefficients, lip movement coefficients, and calculated head pose coefficients are fused and optimized to output a complete, accurate, and temporally coherent set of 3D coefficients. This set of coefficients includes both high-fidelity facial expression features and natural, fluid head movement features.
[0052] In one embodiment, a 3D head coefficient, incorporating head pose and facial expression coefficients, is introduced to control the diffusion model. The 3D head coefficient in this invention is constructed based on the parametric 3D head model FLAME. Specifically, global rotation is employed. As an additional diffusion condition for head control, this parameter Yes, it is a subvector representing the global head rotation in 6 degrees of freedom, generated about the axis. The three rotation angles constitute the facial expression coefficient. This invention uses... (Jaw movement) (Eye features) (Eyelid features) and (Expression parameters). Therefore, the 3D head condition used for the diffusion process in the model designed in this invention can be expressed as formula (1):
[0053] (1)
[0054] The overall conditions for the diffusion process can be summarized as follows: ,in It contains 3D header information that changes over time. Sequence, each Corresponding to a frame in the video, This represents appearance features, which can correspond to the target task features mentioned earlier. Its mathematical expression is formula (2):
[0055] (2)
[0056] For all the coefficients mentioned above, this invention follows the parameter settings of FLAME so that the designed model is compatible with other models in the FLAME ecosystem (such as EMOCA, EMOTE) and other upstream task models.
[0057] The goal of DenoisingNet is to generate videos that both conform to 3D head information and match appearance features.
[0058] Therefore, during the training phase, the FLAME fitting method is used to extract data from each frame of the training video. and the 3D head conditional sequence of the target frame. The target video is used as a supervisory signal and the reference image is used as input for training. The training objective function is given by formula (3):
[0059] (3)
[0060] in This represents the denoising network DenoisingNet. In Stable Diffusion (SD), Representative text prompt Image prompts Or a combination of both In the field of portrait animation, conditions This includes reference images, audio, context frames, key points, and other additional information; Indicates the diffusion time step; This represents the latent representation with noise, corresponding to the noisy image mentioned earlier; This refers to the predicted portrait video, also known as the target portrait video. This represents an operation that calculates the expectation. Portrait videos representing predictions It conforms to a standard normal distribution, with a variance of 1 and a mean of 0.
[0061] During the inference phase, given a portrait image input to ReferenceNet, it obtains... Simultaneously, a 3D head conditional sequence is generated through the audio-to-FLAME module. After being adjusted by the user signal, it is input into DenoisingNet, and the final output is a speaking head video precisely controlled by 3D head conditions.
[0062] In the aforementioned embodiments, the final 3D coefficients are generated through a 3D face model as an intermediate medium, allowing control signals (user instructions and audio) to be mapped into a structured parameter space that conforms to the principles of facial anatomy. This ensures that the generated head pose, expression, and lip movements are physically reasonable and natural, avoiding distortions, inconsistencies, or jitter that may occur with direct regression parameters, and greatly enhancing the realism of the generated video.
[0063] In yet another exemplary embodiment of the present invention, the preceding text continues... Figure 2 The above embodiment will be used as an example for illustration. The video generation model may also include a three-dimensional deformable face model; wherein, based on user instructions and audio information, the 3D deformation model coefficients (corresponding to step 220) can be obtained in the following way:
[0064] User instructions and audio information are input into a 3D deformable face model to obtain the 3D deformation model coefficients output by the 3D deformable face model.
[0065] In one embodiment, the acquired user instructions (such as "show a laughing expression") and audio information can be directly input into a 3D deformable face model. This 3D deformable face model can be an end-to-end trained neural network capable of simultaneously understanding the semantics of natural language instructions and the waveform features of audio, and directly mapping them to the face deformation parameter space. After internal processing, the model outputs a complete set of 3D deformation model coefficients over time. These coefficients directly and completely characterize the deformation features of the target person's face under the user's instruction and audio-driven conditions, including but not limited to facial expression parameters such as the corners of the mouth rising and eyes squinting corresponding to a laugh, as well as lip movement parameters synchronized with speech. In this embodiment, by directly inputting user instructions and audio information into a dedicated 3D deformable face model to obtain deformation coefficients, end-to-end mapping from high-level semantic instructions to low-level deformation parameters is achieved. This simplifies the processing flow, avoids the complexity and error accumulation that might have previously required multiple intermediate modules (such as separate expression analysis modules and lip-reading modules), and improves the overall efficiency and reliability of the system.
[0066] Among them, 3D deformable face models can accurately capture, parameterize, and reconstruct faces using statistical models. As a fundamental 3D deformable face model, FLAME has been widely used in numerous studies. This model employs linear blending skinning (LBS) based on standard vertices and blending shape techniques, including... Each vertex Joints (neck, jaw, and eyeballs) and Each facet. From a mathematical perspective, FLAME can describe the face modeling process using a function, as shown in formula (4):
[0067] (4)
[0068] This function parameterizes a 3D face into identity shape parameters. Facial expression parameters and attitude parameters (Include Rotation of individual joints (neck, jaw, and eyeballs) and global rotation. The rotation vector for each pose parameter is represented using axis-angle notation. Based on these input parameters, FLAME can output a list of vertices. and dough sheets A three-dimensional mesh structure.
[0069] Figure 3 This is a schematic diagram of the process of training a video generation model using a progressive focus training strategy provided by the present invention.
[0070] The following will combine Figure 3 The process of training a video generation model using a progressive focus training strategy provided by this invention will be described.
[0071] In an exemplary embodiment of the present invention, combined with Figure 3 As can be seen, training the video generation model through the progressive focus training strategy can include steps 310 to 330, which will be described in detail below.
[0072] In step 310, a training sample set is obtained, wherein the training sample set includes multiple training samples, and the training samples include audio sample information and user instruction sample information.
[0073] In one embodiment, a large-scale training sample set can be constructed by crawling from a database or the web. This set contains millions of training samples. Each training sample includes at least audio sample information and user instruction sample information. The audio sample information can be a real voice recording (such as "Hello, the weather is nice today"). The user instruction sample information can be a text instruction that matches the voice recording semantically and temporally (such as "Smile and speak"). These instructions can be manually annotated or automatically generated from video context or voice content.
[0074] In step 320, the training samples are input into the three-dimensional deformable face model to obtain the 3D deformation model coefficient samples output by the three-dimensional deformable face model. The 3D deformation model coefficient samples include head motion feature samples and facial expression feature samples.
[0075] In one embodiment, each acquired training sample (audio sample information + user instruction sample information) can be batch-input into a pre-trained, fixed 3D deformable face model. This 3D deformable face model can process the input and output corresponding 3D deformation model coefficient samples for each training sample. This set of coefficient samples is time-series data, precisely containing head motion feature samples and facial expression feature samples. The head motion feature samples can be posture parameters such as natural head rotation and nodding during speech. The facial expression feature samples can be smiling expression parameters and lip movement parameters precisely synchronized with the speech content.
[0076] At this point, the original (audio, text) training sample pairs have been transformed into more structured (audio, 3D coefficients) training sample pairs, providing high-quality driving signals for subsequent training.
[0077] In step 330, the video generation model is trained using a progressive focusing training strategy based on head motion feature samples and facial expression feature samples to obtain a trained video generation model.
[0078] In one embodiment, the video generation model (including a reference network and a denoising network) is trained using prepared high-quality training data. The training employs a progressively focused training strategy, which can be implemented as follows:
[0079] Phase 1 - Coarse-grained General Training: In the initial training phase, the model is trained using a relatively high level of noise (such as in the first few steps of noise scheduling in a diffusion model) and a wide variety of data (including facial data of various identities, skin colors, and ages). The main task of the model in this phase is to learn how to generate a structurally correct, reasonably animated, but potentially less detailed facial video from random or noisy images based on 3D coefficient-driven signals. This phase focuses on enabling the model to acquire general prior knowledge of video generation and motion synthesis, laying a stable foundation.
[0080] Phase Two - Progressive Fine-Tuning: After the model initially converges, the progressive fine-tuning phase begins. In this phase, the level of added noise is gradually reduced (using later steps in noise scheduling), and the training data is gradually focused on higher-quality, higher-resolution, or more style-specific subsets (e.g., cartoon, realistic). As noise decreases, the model is forced to learn to recover finer image details, such as skin texture, hair strands, and precise lighting effects. Through this progressive training from "coarse" to "fine," the model's generative capabilities are gradually sharpened, ultimately enabling the generation of high-definition, high-fidelity, and richly detailed target portrait videos.
[0081] Through the above multi-stage training process, a well-trained video generation model is finally obtained, which perfectly combines general generation capabilities with the ability to finely depict details.
[0082] In another embodiment, the Progressive Focusing Training (PFT) strategy is employed, the core idea of which is that the model typically learns global image features first and then focuses on details. This invention divides 3D head conditions into two groups: detailed facial expressions. and head movement conditions Specifically, it can include the following three stages: The first stage (head motion stage) involves image training, where all layers are trainable. (Conditions provided only.) The first stage forces the model to focus on head pose changes between the reference and target images; the second stage (expression stage) focuses on learning information including eye movements. Mandibular movement and global emoticons Facial expressions. The conditions for this stage are... ,in To avoid coupling between facial expressions and head movements, the attention layer related to head movements is frozen. The third stage (video training stage) involves video training following the corresponding settings.
[0083] In this embodiment, a progressively focused training strategy of "coarse first, then fine" is adopted. Common features are first learned on a large amount of noisy and diverse data to provide a stable and robust initialization for the model. This effectively avoids problems such as gradient explosion and mode collapse that are prone to occur when training directly on high-resolution, low-noise data, and greatly accelerates the convergence process of the model.
[0084] Figure 4 This is a schematic diagram of the process provided by the present invention, which trains the video generation model based on the head motion feature samples and the facial expression feature samples through a progressive focusing training strategy to obtain a trained video generation model.
[0085] The following will combine Figure 4 The process of training the video generation model based on the head motion feature samples and the facial expression feature samples provided by the present invention using a progressive focusing training strategy to obtain a trained video generation model is described below.
[0086] In another exemplary embodiment of the present invention, the video generation model is trained using a progressive focusing training strategy based on head motion feature samples and facial expression feature samples to obtain a trained video generation model. This may include steps 410 and 420, which will be described in detail below:
[0087] In step 410, the attention layer related to head movement in the video generation model is trained based on head movement feature samples to obtain the video generation model after the first stage of training.
[0088] In step 410, the attention layer related to head movement is frozen, and the video generation model trained in the first stage is trained based on facial expression feature samples to obtain the trained video generation model.
[0089] In one embodiment, training focused on head movements can begin first. This starts with an initialized or pre-trained video generation model. Certain layers in the model's denoising network are identified and defined as attention layers related to head movements. These attention layers typically refer to modules in the network responsible for modeling long-range dependencies and sensitive to global pose and structural information (such as the Cross-Attention layer in the Transformer architecture). At this stage, the model is trained using only the head movement feature samples (or as the primary supervision signal). Specifically, when calculating the loss, the primary constraint is that the head pose of the generated video must be consistent with the pose defined by the head movement feature samples. During this training stage, all parameters of the model (including the attention layers related to head movements and other layers) are updatable. Through training, the model learns how to generate videos with correct, natural head movements based on the head movement features. After this training stage is completed, the first-stage trained video generation model is obtained. At this point, the attention layers related to head movements in the model have learned how to understand and execute head movement commands effectively.
[0090] Next, training focused on facial expressions is restarted. First, the parameters of all attention layers related to head movements in the video generation model after the first stage of training are frozen. This means that the weights of these key layers will no longer be updated in subsequent backpropagation calculations, and their learned head movement generation capabilities are firmly preserved. Then, based on the facial expression feature samples (or using them as the primary supervision signal), the remaining unfrozen parameters of the video generation model after the first stage of training are trained. These updatable parameters typically include the network layers responsible for generating facial details, textures, and expressions. At this stage, since the head movement generation capability has been "locked in," the model can focus entirely on learning how to generate lifelike facial expressions and precise lip-sync based on fine facial expression features without worrying about disrupting the already learned head movement generation capability. After this stage of training is completed, the final trained video generation model is obtained. This model possesses both the stable and accurate head movement generation capability obtained in the first stage of training and the delicate and rich expression generation capability obtained in the second stage of training.
[0091] In this embodiment, by freezing the attention layer related to head movements, during the initial stage of training focused on facial expressions, all learning resources are forcibly directed to the subtask of "facial expression generation." This physical parameter freezing achieves hard decoupling of the two tasks during training, effectively avoiding inter-task interference that may occur when training all parameters simultaneously (e.g., optimizing expressions may unintentionally damage already learned head movement representations), ensuring that each subtask can be trained to its optimal state. This invention proposes a progressively focused training scheme to achieve decoupling of expression and movement, ensuring independent control of expression and head movement by progressively focusing on learning facial expression conditions and head posture.
[0092] In another exemplary embodiment of the present invention, the user instruction may include any one or more of a first user instruction and a second user instruction, wherein the first user instruction is an instruction for generating the head movement of the target person according to a specified axis rotation angle; and the second user instruction is a video or image for representing the head movement of the target person desired by the user.
[0093] In one embodiment, the first user instruction can be a specific text command that precisely characterizes the angle by which the target person's head needs to rotate around a specified axis. For example, the user could enter the command via a text box: "Rotate 30 degrees clockwise around the Y-axis, then slowly return to the starting position." Here, the "Y-axis" is the specified axis, and "30 degrees" is the angle of rotation. The second user instruction can be a short video (such as an MP4 file) or an image (such as a JPEG file) containing the desired head movement of the target person. For example, the user could upload a demonstration video in which a person is making a "nodding affirmative" gesture. Alternatively, the user could upload a side-view image of a person, hoping that the generated portrait video will have a head orientation similar to that image.
[0094] In another embodiment, the user can provide both instructions simultaneously. For example, the user uploads a demonstration video of "head shaking" (second user instruction) and simultaneously inputs the text instruction "but reduce the amplitude of the movement by half" (a variant of the first user instruction). The system will combine these two instructions to generate the final control signal.
[0095] In this embodiment, multiple input formats are supported, adapting to different user preferences and scenario needs. Whether you are a technical user who prefers precise control or an ordinary user who prefers intuitive demonstrations, you can find an interaction method that suits you. This enhances the understanding and adaptability to user intent, significantly improving its generalizability and practicality.
[0096] As described above, the portrait video generation method provided by this invention can generate natural head movements based on audio, and also supports replacing these movements with user-defined actions (such as turning the head, tilting, looking around, or even sudden shaking). It also supports six degrees of freedom (6DoF) head movement control, demonstrating a refined understanding of three-dimensional space. Regarding facial expression control, the method proposed in this invention can generate realistic expressions from three sources: audio input, user-specified expression coefficients, and coefficients extracted from images or videos. These sources can be used individually or in combination to create new expressions. The method proposed in this invention is based on a diffusion framework, achieving high controllability while maintaining high accuracy and naturalness in lip-sync. Furthermore, this invention is the first to use 3D head coefficients as a diffusion condition, which not only enhances controllability but also provides flexibility for integration with existing 3D head generation methods, bridging the gap between 3D model-based methods and end-to-end diffusion technology.
[0097] The portrait video generation apparatus provided by the present invention is described below. The portrait video generation apparatus described below and the portrait video generation method described above can be referred to in correspondence.
[0098] Figure 5 This is a schematic diagram of the portrait video generation device provided by the present invention.
[0099] The following will combine Figure 5 The structure of the portrait video generation apparatus provided by the present invention will be described.
[0100] In an exemplary embodiment of the present invention, combined with Figure 5 As can be seen, the portrait video generation device may include an acquisition module 510, a calling module 520, an input module 530, a processing module 540, and a generation module 550. Each module will be described in detail below.
[0101] The acquisition module 510 can be configured to acquire audio information, a reference image, user instructions, and a noisy image, wherein the reference image is an image containing a target person; the noisy image is an image obtained by adding noise to the reference image; and the user instructions are instructions for guiding the head movements of the target person in the generated target portrait video.
[0102] The calling module 520 can be configured to call a pre-trained video generation model, wherein the video generation model includes a reference network and a denoising network, and the video generation model is trained through a progressive focus training strategy;
[0103] Input module 530 can be configured to input the reference image into the reference network to obtain the target person features output by the reference network;
[0104] The processing module 540 can be configured to obtain 3D coefficients corresponding to the user instruction based on the user instruction and the audio information, wherein the 3D coefficients are used to characterize the head movement features and facial expression features of the target person.
[0105] The generation module 550 can be configured to input the 3D coefficients, the target person features, and the noisy image into the denoising network to obtain a target portrait video output by the denoising network, wherein the head movements of the target person in the target portrait video match the user's instructions.
[0106] In an exemplary embodiment of the present invention, the video generation model further includes a 3D face model; the processing module 540 can obtain 3D coefficients corresponding to the user instruction based on the user instruction and the audio information in the following manner:
[0107] Based on the user instructions and the audio information, 3D deformation model coefficients are obtained, wherein the 3D deformation model coefficients are used to characterize the deformation features of the target person;
[0108] The 3D deformation model coefficients and the audio information are input into the 3D face model to obtain the 3D coefficients output by the 3D face model corresponding to the user's instructions.
[0109] In an exemplary embodiment of the present invention, the video generation model further includes a three-dimensional deformable face model; the processing module 540 can obtain the 3D deformation model coefficients based on the user instruction and the audio information in the following manner:
[0110] The user instructions and the audio information are input into the three-dimensional deformable face model to obtain the 3D deformation model coefficients output by the three-dimensional deformable face model.
[0111] In an exemplary embodiment of the present invention, the calling module 520 may train the video generation model using a progressive focus training strategy in the following manner:
[0112] Obtain a training sample set, wherein the training sample set includes multiple training samples, and the training samples include audio sample information and user instruction sample information;
[0113] The training samples are input into a three-dimensional deformable face model to obtain 3D deformation model coefficient samples output by the three-dimensional deformable face model, wherein the 3D deformation model coefficient samples include head motion feature samples and facial expression feature samples.
[0114] Based on the head motion feature samples and the facial expression feature samples, the video generation model is trained using a progressive focusing training strategy to obtain a trained video generation model.
[0115] In an exemplary embodiment of the present invention, the calling module 520 may train the video generation model based on the head motion feature samples and the facial expression feature samples using a progressive focusing training strategy to obtain a trained video generation model:
[0116] Based on the head motion feature samples, the attention layer related to head motion in the video generation model is trained to obtain the video generation model after the first stage of training.
[0117] The attention layer related to head movement is frozen, and the video generation model trained in the first stage is trained based on the facial expression feature samples to obtain the trained video generation model.
[0118] In an exemplary embodiment of the present invention, the user instruction includes any one or more of a first user instruction and a second user instruction, wherein the first user instruction is an instruction for characterizing the generation of the head movement of the target person by rotating the target person at a specified axis angle; and the second user instruction is a video or image for characterizing the head movement of the target person desired by the user.
[0119] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communications bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other through the communications bus 640. The processor 610 can call logic instructions in the memory 630 to execute a portrait video generation method, which includes: acquiring audio information, a reference image, user instructions, and a noisy image, wherein the reference image is an image containing a target person; the noisy image is an image obtained by adding noise to the reference image; the user instructions are instructions for guiding the head movements of the target person in the generated target portrait video; calling a pre-trained video generation model, wherein the video generation model includes a reference network and a denoising network, and the video generation model is trained using a progressive focus training strategy; inputting the reference image into the reference network to obtain the target person features output by the reference network; based on the user instructions and the audio information, obtaining 3D coefficients corresponding to the user instructions, wherein the 3D coefficients are used to characterize the head movement features and facial expression features of the target person; inputting the 3D coefficients, the target person features, and the noisy image into the denoising network to obtain the target portrait video output by the denoising network, wherein the head movements of the target person in the target portrait video match the user instructions.
[0120] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0121] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the portrait video generation method provided by the above methods. The method includes: acquiring audio information, a reference image, user instructions, and a noisy image, wherein the reference image is an image containing a target person; the noisy image is an image obtained by adding noise to the reference image; the user instructions are instructions for guiding the head movements of the target person in the generated target portrait video; and invoking a pre-trained video generation model, wherein... The video generation model includes a reference network and a denoising network, and is trained using a progressive focus training strategy. The reference image is input into the reference network to obtain the target person features output by the reference network. Based on the user instruction and the audio information, 3D coefficients corresponding to the user instruction are obtained, wherein the 3D coefficients are used to characterize the head movement features and facial expression features of the target person. The 3D coefficients, the target person features, and the noisy image are input into the denoising network to obtain the target portrait video output by the denoising network, wherein the head movements of the target person in the target portrait video match the user instruction.
[0122] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the portrait video generation method provided by the methods described above. The method includes: acquiring audio information, a reference image, user instructions, and a noisy image, wherein the reference image is an image containing a target person; the noisy image is an image obtained by adding noise to the reference image; the user instructions are instructions for guiding the head movements of the target person in the generated target portrait video; and invoking a pre-trained video generation model, wherein the video generation model includes a reference network. The video generation model is trained using a progressive focus training strategy and a denoising network. The reference image is input into the reference network to obtain the target person features output by the reference network. Based on the user instruction and the audio information, 3D coefficients corresponding to the user instruction are obtained, wherein the 3D coefficients are used to characterize the head movement features and facial expression features of the target person. The 3D coefficients, the target person features, and the noisy image are input into the denoising network to obtain the target portrait video output by the denoising network, wherein the head movements of the target person in the target portrait video match the user instruction.
[0123] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0124] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating portrait videos, characterized in that, The method includes: The system acquires audio information, a reference image, user instructions, and a noisy image, wherein the reference image is an image containing the target person; the noisy image is an image obtained by adding noise to the reference image; and the user instructions are commands used to guide the head movements of the target person in the generated target portrait video. A pre-trained video generation model is invoked, wherein the video generation model includes a reference network and a denoising network, and the video generation model is trained using a progressive focus training strategy; The reference image is input into the reference network to obtain the target person features output by the reference network; Based on the user instruction and the audio information, 3D coefficients corresponding to the user instruction are obtained, wherein the 3D coefficients are used to characterize the head movement features and facial expression features of the target person; The 3D coefficients, the target person's features, and the noisy image are input into the denoising network to obtain the target portrait video output by the denoising network. The head movements of the target person in the target portrait video match the user's instructions. The video generation model is trained using a progressive focus training strategy, implemented in the following manner: Obtain a training sample set, wherein the training sample set includes multiple training samples, and the training samples include audio sample information and user instruction sample information; The training samples are input into a three-dimensional deformable face model to obtain 3D deformation model coefficient samples output by the three-dimensional deformable face model, wherein the 3D deformation model coefficient samples include head motion feature samples and facial expression feature samples. Based on the head motion feature samples and the facial expression feature samples, the video generation model is trained using a progressive focusing training strategy to obtain a trained video generation model.
2. The portrait video generation method according to claim 1, characterized in that, The video generation model further includes a 3D face model; the step of obtaining 3D coefficients corresponding to the user instruction based on the user instruction and the audio information specifically includes: Based on the user instructions and the audio information, 3D deformation model coefficients are obtained, wherein the 3D deformation model coefficients are used to characterize the deformation features of the target person; The 3D deformation model coefficients and the audio information are input into the 3D face model to obtain the 3D coefficients output by the 3D face model corresponding to the user's instructions.
3. The portrait video generation method according to claim 2, characterized in that, The video generation model also includes a 3D deformable face model; the process of obtaining 3D deformation model coefficients based on the user instructions and the audio information specifically includes: The user instructions and the audio information are input into the three-dimensional deformable face model to obtain the 3D deformation model coefficients output by the three-dimensional deformable face model.
4. The portrait video generation method according to claim 1, characterized in that, The process of training the video generation model based on the head motion feature samples and the facial expression feature samples using a progressive focusing training strategy to obtain a trained video generation model specifically includes: Based on the head motion feature samples, the attention layer related to head motion in the video generation model is trained to obtain the video generation model after the first stage of training. The attention layer related to head movement is frozen, and the video generation model trained in the first stage is trained based on the facial expression feature samples to obtain the trained video generation model.
5. The portrait video generation method according to claim 1, characterized in that, The user instruction includes any one or more of a first user instruction and a second user instruction, wherein the first user instruction is an instruction for generating the head movement of the target person by rotating the target person at a specified axis angle; and the second user instruction is a video or image for representing the head movement of the target person that the user expects.
6. A portrait video generation device, characterized in that, The apparatus is used to implement the portrait video generation method according to any one of claims 1 to 5, the apparatus comprising: The acquisition module is used to acquire audio information, a reference image, user instructions, and a noisy image, wherein the reference image is an image containing the target person; the noisy image is an image obtained by adding noise to the reference image; and the user instructions are instructions used to guide the head movements of the target person in the generated target portrait video. The calling module is used to call a pre-trained video generation model, wherein the video generation model includes a reference network and a denoising network, and the video generation model is trained through a progressive focus training strategy; The input module is used to input the reference image into the reference network to obtain the target person features output by the reference network; The processing module is used to obtain 3D coefficients corresponding to the user instruction based on the user instruction and the audio information, wherein the 3D coefficients are used to characterize the head movement features and facial expression features of the target person; The generation module is used to input the 3D coefficients, the target person features, and the noisy image into the denoising network to obtain the target portrait video output by the denoising network, wherein the head movements of the target person in the target portrait video match the user instructions.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the portrait video generation method as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the portrait video generation method as described in any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the portrait video generation method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Portrait video generation method and device, electronic equipment and storage medium
CN112750185A
Character animation generation method and device based on voice driving, equipment and medium
CN119338958A