Video generation method, apparatus, device, storage medium, and program product

By acquiring change control and driving information, and using a deep neural network model to generate realistic videos, the problem of unrealistic and unnatural virtual character videos in existing technologies is solved, and the naturalness and lighting adaptability of the videos are improved.

CN122457797APending Publication Date: 2026-07-24BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING CO WHEELS TECH CO LTD
Filing Date
2025-01-23
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

The virtual character videos generated by existing technologies are unrealistic, unnatural, and inflexible, with limited movement, simple expressions, and an inability to adapt to changes in different lighting environments.

Method used

By acquiring change control information and combining it with driving information to generate target video, and using deep neural network models such as the denoising U-NET model, combined with motion control information and illumination change control information, realistic video effects are generated.

Benefits of technology

It achieves realistic video effects, improves the naturalness and adaptability of videos, and enhances the ability of videos to perform under different lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122457797A_ABST
    Figure CN122457797A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video generation method, device, equipment, storage medium and program product. The method comprises: obtaining change control information used for controlling a change process of the reference image; and generating a target video based on driving information and the change control information. In the embodiments of the present disclosure, the target video is generated based on the driving information and the change control information, which can realize a realistic video effect and improve the naturalness of the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a video generation method, apparatus, device, storage medium, and program product. Background Technology

[0002] In existing technologies, the generated virtual character videos (such as digital human videos) have the following problems: they are not realistic, unnatural, and inflexible, such as limited movement, simple expressions, and inability to adapt to changes in different lighting environments. Summary of the Invention

[0003] This invention provides a video generation method, apparatus, device, storage medium, and program product, which can improve the video generation effect.

[0004] In a first aspect, embodiments of this disclosure provide a video generation method, comprising: acquiring change control information for controlling the change process of a reference image; and generating a target video based on driving information and the change control information.

[0005] Secondly, this disclosure also provides a video generation apparatus, including: a change control information acquisition module, used to acquire change control information for controlling the change process of a reference image; and a video generation module, used to generate a target video based on driving information and the change control information.

[0006] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method as described in embodiments of this disclosure.

[0007] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video generation method as described in embodiments of this disclosure.

[0008] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the video generation method as described in embodiments of this disclosure.

[0009] The technical solution of this disclosure involves acquiring change control information for controlling the change process of the reference image; and generating a target video based on the driving information and the change control information. This embodiment of the disclosure, by generating the target video based on the driving information and the change control information, can achieve realistic video effects and improve the naturalness of the video. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0011] Figure 1 This is a schematic diagram of a video generation method provided in an embodiment of the present invention;

[0012] Figure 2 This is a schematic diagram of another video generation method provided in an embodiment of the present invention;

[0013] Figure 3 This is a schematic diagram of another video generation method provided in an embodiment of the present invention;

[0014] Figure 4 This is a schematic diagram of another video generation method provided in an embodiment of the present invention;

[0015] Figure 5 This is a schematic diagram of the target video generation process provided in an embodiment of the present invention;

[0016] Figure 6 Schematic diagram of reference images provided for embodiments of the present invention;

[0017] Figure 7 This is a schematic diagram of each frame of the target video provided in an embodiment of the present invention;

[0018] Figure 8 This is a schematic diagram of each frame of the target video provided in an embodiment of the present invention;

[0019] Figure 9 This is a schematic diagram of a video generation device provided in an embodiment of the present invention;

[0020] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0021] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0022] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". It should be noted that the concepts of "first," "second," etc., mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should be noted that the modifications "a" and "a plurality" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless explicitly indicated otherwise in the context, they should be understood as "one or more". It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) shall comply with the requirements of applicable laws, regulations, and relevant provisions.

[0023] Figure 1 This is a schematic flowchart of a video generation method provided by an embodiment of the present invention. This embodiment is applicable to situations involving video generation. The method can be executed by a video generation device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server. Figure 1 As shown, the method includes:

[0024] S110. Obtain change control information for controlling the change process of the reference image.

[0025] The reference image guides the generation of the target video based on the background and portrait features in the reference image. That is, the background and portrait features of each frame in the target video are identical to those in the reference image. In this embodiment, the portrait features are not limited and can be digital humans, digital animals, etc. Portrait features can be understood as facial features and head features. In this embodiment, the source of the reference image is not limited; for example, it can be a user-inputted reference image or a pre-stored reference image (which can be understood as a default image). It should be noted that if the user does not input a reference image, a default image can be used as the reference image. The reference image can be any frame from any video, or it can be an image fused from an image containing portrait features and an image containing background features. The reference image can include facial features or full-body features; that is, the reference image can be a reference facial image or a reference full-body image.

[0026] In this context, change control information can be understood as information controlling the reference image's body movements (such as facial and limb movements), facial expressions, and lighting changes at different shooting points (equivalent to the camera being able to move, i.e., different camera movement conditions). The shooting point can be understood as the position chosen by the camera when photographing people or animals. Shooting positions can include various positions such as frontal shooting, side shooting, oblique shooting, low-angle shooting, and high-angle shooting at different shooting distances. Shooting distances can include close-up distances and far-away distances.

[0027] In this embodiment, there are no restrictions on the method of obtaining change control information. For example, multiple sets of motion parameters can be extracted from multiple frames of images in any video, and image rendering can be performed based on the multiple sets of motion parameters to obtain motion image sequences and illumination images. Feature extraction can be performed on the motion image sequences and illumination images to obtain change control information.

[0028] Optionally, the change control information includes: motion control information and / or illumination change control information.

[0029] Motion control information can be understood as body motion control information and facial expression motion information at different shooting points. Body motion control information can include facial motion control information and motion control information for other parts of the body (such as limbs and trunk). Facial expression motion information can be understood as facial expression motion information.

[0030] Among them, the illumination change control information can be understood as the illumination change control information under different shooting points.

[0031] In this embodiment, the change control information includes motion control information and / or illumination change control information, which can precisely control the changes in body movement, facial expressions, and illumination of the reference image, thereby generating more realistic video effects and improving the naturalness of the video.

[0032] S120: Generate the target video from the driving information and change control information.

[0033] The driving information is used to drive the generation of the target video. In this embodiment, the driving information is not limited and may include first driving information and second driving information. The first driving information is a reference image. The second driving information can be at least one of the following: text, image, text and image, driven speech, etc. The second driving information is used to guide the speech generation in the target video.

[0034] In this embodiment, the driving speech is used to drive the reference image to generate a corresponding target video according to the driving speech, and the audio in the target video is the driving speech. This embodiment does not limit the specific content of the driving speech and can be determined according to actual needs. This embodiment also does not limit the source of the driving speech; it can be extracted from any video containing audio or a recorded audio clip.

[0035] The target video is used to play audio using a reference image. In this embodiment, the audio can be obtained based on the second driving information.

[0036] In this embodiment, there are no restrictions on the method of generating the target video. For example, the target video can be generated by a video generation model, such as by inputting driving information and change control information into the video generation model to generate the target video; or the target video can be generated based on rules, such as by using a predefined video template and dynamically adjusting the content (such as voice and reference images) in the template by combining driving information and change control information to generate the target video.

[0037] In this embodiment, the video generation model is not limited and can be any deep neural network model that can generate the target video based on driving information and change control information.

[0038] In this embodiment, the video generation model enables the generated target video to play speech using a reference image while also controlling the reference image to change according to change control information. That is, the generated target video is a video that controls the body movement (such as facial movement), expression changes, lighting changes and speech of the reference image in sync.

[0039] Optionally, generating the target video based on driving information and change control information includes: extracting semantic information from the reference image through the feature extraction module of the video generation model; adjusting the semantic information using change control information through the image generation module of the video generation model; and generating the target video based on the adjusted semantic information and driving information.

[0040] In this embodiment, there are no restrictions on the method used to extract semantic information in the feature extraction module of the video generation model. For example, a ReferenceNet network can be used to extract semantic information. Here, semantic information can be understood as the visual appearance information of the portrait and its related background in the reference image.

[0041] In this embodiment, the feature extraction module of the video generation model can also be used to extract the driving audio information from the driving information.

[0042] In this embodiment, there are no restrictions on the method of extracting driving audio information by the feature extraction module of the video generation model. For example, a speech model (wav2vec) can be used to extract driving audio information from the driving information.

[0043] In this embodiment, there are no restrictions on the way the image generation module of the video generation model generates the target video. For example, a diffusion model can be used to fuse motion control information and / or illumination change control information with semantic information to adjust the semantic information, and then generate the target video based on the adjusted semantic information and driving information.

[0044] Optionally, the image generation module in the video generation model is a denoising U-NET model.

[0045] In this embodiment, the denoising U-NET model can be understood as a diffusion model using the U-Net architecture. The video generation model, including this embodiment, uses the denoising U-NET model to generate the target video. This method can integrate motion control information and / or illumination change control information, semantic information, and driving audio information, so that each frame in the generated target video is visually consistent with the reference image, while flexibly displaying various motion, expressions, and lighting effects.

[0046] Optionally, the training method for the video generation model includes: training the basic video generation model using a first training set to obtain a first video generation model; the first training set includes a first reference image set and a first video set; each first video in the first video set has the same portrait features, background features, and lighting features as the corresponding first reference image, but different three-dimensional motion features; adjusting the first video generation model using a second training set to obtain a second video generation model; the second training set includes a second reference image set and a second video set; each second video in the second video set has the same portrait features and background features as the corresponding second reference image, but different lighting features; adjusting the second video generation model using a third training set to obtain a third video generation model, which serves as the trained video generation model; wherein the third training set includes a third reference image set and a third video set, and each third video in the third video set includes the driving information.

[0047] In this embodiment, the basic video generation model can be understood as an untrained video generation model. The first training set includes a first reference image set and a first video set; each first video in the first video set has the same portrait features, background features, and lighting features as the corresponding first reference image, but different 3D motion features. 3D motion features can include head 3D motion features and facial expression 3D motion features, or 3D motion features of other parts of the body (such as limbs and torso motion features). In this embodiment, the range of motion corresponding to the 3D motion features (such as head 3D motion features) in the first video set is not limited; it can be a wide range of motion, such as a head rotation range of 60-80 degrees. Each frame in the first video set can be captured at different shooting points. By possessing a wide range of 3D motion features and a first video set captured at different shooting points, the basic video generation model is trained, resulting in a target video generated by the trained video generation model with a larger range of motion, richer 3D motion features, and a more natural video. The second training set includes a second reference image set and a second video set; each second video in the second video set has the same portrait features and background features as the corresponding second reference image, but different lighting features; the third training set includes a third reference image set and a third video set, and each third video in the third video set includes driving information, such as driving speech.

[0048] In this embodiment, the training of the video generation model can be divided into three stages: First, a first training set is input into the basic video generation model to generate a first training video set; a first loss value is obtained based on the first video set and the first training video set; the basic video generation model is iteratively trained based on the first loss value until a stopping condition is met (e.g., the first loss value is less than a set loss value, or the number of iterations is greater than a set number), thus obtaining the first video generation model. Second, a second training set is input into the first video generation model to generate a second training video set; a second loss value is obtained based on the second video set and the second training video set; the first video generation model is fine-tuned based on the second loss value until a stopping condition for fine-tuning is met (e.g., the second loss value is less than a set loss value, or the number of iterations is greater than a set number), thus obtaining the second video generation model. Third, a third training set is input into the second video generation model to generate a third training video set; a third loss value is obtained based on the third video set and the third training video set; the second video generation model is fine-tuned based on the third loss value until a stopping condition for fine-tuning is met (e.g., the third loss value is less than a set loss value, or the number of iterations is greater than a set number), thus obtaining the third video generation model.

[0049] In this embodiment, the basic video generation model is gradually trained and adjusted using a first training set, a second training set, and a third training set, enabling the video generation model to learn more diverse features, including portrait features, background features, lighting features, 3D motion features, and driving information, thereby effectively improving the performance of the video generation model.

[0050] The technical solution of this disclosure involves acquiring change control information for controlling the change process of a reference image; and generating a target video from the driving information and the change control information. This disclosure, by generating a target video based on the driving information and the change control information, can achieve realistic video effects and improve the naturalness of the video.

[0051] Figure 2 This is a schematic flowchart of another video generation method provided in an embodiment of the present invention. See also... Figure 2 The method provided in this embodiment of the invention specifically includes the following steps:

[0052] S201. Generate a sequence of motion images based on multiple pre-extracted sets of first motion parameters.

[0053] In this embodiment, any method can be used to generate a motion image sequence based on multiple sets of first motion parameters. For example, a single-view method can be used. Figure 3 The FLAME model within the DECA framework, a 3D reconstruction method, renders multiple sets of first motion parameters into a sequence of motion images. This sequence consists of multiple motion images, each containing rich 3D information, such as 3D facial motion features and 3D facial expression features. Specifically, each motion image is in the format of a 3D Morphable FaceModel (3DMM). The FLAME model can be understood as a type of 3DMM technology. Each set of first motion parameters corresponds to one motion image.

[0054] In this embodiment, the first motion parameter is not limited and can include any number of key parameters in 3D reconstruction.

[0055] Optionally, each set of first motion parameters includes a first variable parameter and a first fixed parameter; wherein, the first variable parameter includes at least one of the following: head shape parameter, posture motion parameter, facial expression motion parameter, and shooting parameter of the shooting point; the first fixed parameter includes the illumination parameter of the illumination ball; the first variable parameter is different for each set of first motion parameters.

[0056] In each group of first motion parameters, the first fixed parameter is the same.

[0057] In this embodiment, the posture motion parameters are not limited, and may include head motion parameters and motion parameters of other body parts. Motion parameters of other body parts may include limb motion parameters, trunk motion parameters, etc. For example, if the posture motion parameters only include head motion parameters, the corresponding motion image sequence may be a facial motion image sequence; if the posture motion parameters include motion parameters of the whole body, the corresponding motion image sequence may be a whole body motion image sequence.

[0058] Among these, head shape parameters can be understood as the head's shape characteristics, determined by factors such as genetics or age, and can be denoted as Shape Parameters. Lighting parameters can be understood as the lighting conditions in the environment and can be denoted as Lighting Parameters. Head motion parameters can be understood as the head's posture, including rotation and translation, and can be denoted as Pose Parameters. Facial expression parameters can include expressions such as smiling and frowning, and can be denoted as Expression Parameters. Shooting parameters for the shooting point can be understood as parameters related to the camera's posture, including the camera's position and angle.

[0059] In this embodiment, by utilizing 3DMM technology to generate a sequence of motion images based on multiple sets of different first motion parameters, pixel-level control of the motion images is achieved during the generation process (i.e., the movement and changes of each pixel are controlled by different sets of first motion parameters), thereby enabling a wider range of motion control (such as facial and expression movements), and thus giving the motion image sequence a wider range of three-dimensional motion features.

[0060] In this embodiment, a set of initial first motion parameters can be extracted from each frame of the first set video.

[0061] In this embodiment, each frame of the first set of images in the first set video corresponds to a different set of initial first motion parameters.

[0062] In this embodiment, a set of initial motion parameters can be extracted from each frame of the first set video using any open-source tool.

[0063] It should be noted that each frame in the target video corresponds one-to-one with each frame in the first set video. That is, the motion image corresponding to each frame in the target video is obtained based on a set of initial first motion parameters of the corresponding frame in the first set video.

[0064] In this embodiment, the process of generating a motion image sequence can be as follows: For each frame of an image in a first set of preset videos, the parameters in a corresponding set of extracted initial first motion parameters are adjusted to obtain a corresponding set of first motion parameters: the illumination parameters of the illumination sphere are set to a fixed value as the first fixed parameter, and the head shape parameters, posture motion parameters, facial expression motion parameters, and shooting parameters of the shooting point are dynamically adjusted as the first variable parameters. Specifically, the head shape parameters, posture motion parameters, facial expression motion parameters, and shooting parameters of the shooting point extracted from the first set of preset videos can be directly used as the first variable parameters, or the head shape parameters, posture motion parameters, facial expression motion parameters, and shooting parameters of the shooting point can be adjusted according to the actual situation, and the adjusted head shape parameters, posture motion parameters, facial expression motion parameters, and shooting parameters of the shooting point can be used as the first variable parameters. Using the FLAME model, the corresponding set of first motion parameters (first fixed parameters and first variable parameters) are rendered to obtain the motion image corresponding to each frame of the first set of preset videos, thereby obtaining a motion image sequence.

[0065] S202. Input the motion image sequence into the motion encoder and extract the three-dimensional motion features of each motion image as motion control information.

[0066] In this embodiment, the deep learning model used by the motion encoder is not limited, as long as it can extract the three-dimensional motion features of each motion image. Since the motion images are in 3DMM format, each motion image contains three-dimensional information, thus the three-dimensional motion features of each motion image can be extracted to improve the accuracy of motion control information.

[0067] Optionally, the motion encoder includes multiple residual units; each residual unit is connected in series.

[0068] In this embodiment, the internal structure of the residual unit is not limited and can include any one or more neural network models. For example, the residual unit may include a first residual unit and a second residual unit; the number of first residual units is a first predetermined number, such as 3, and the number of second residual units is a second predetermined number, such as 2. That is, the motion encoder includes a first predetermined number of first residual units and a second predetermined number of second residual units. Each first residual unit includes a residual block and a downsampling convolutional layer. Each first residual unit is connected in series. Each second residual unit includes a residual block, a self-attention layer, and a downsampling convolutional layer. Each second residual unit is connected in series. The first residual unit and the second residual unit are connected in series.

[0069] In this embodiment, by employing multiple serial residual units for three-dimensional motion feature extraction, the feature extraction capability of the motion encoder can be enhanced, thereby generating more accurate and natural motion control information.

[0070] In this embodiment, a motion image sequence is generated based on the pre-extracted first motion parameters; the motion image sequence is then input into a motion encoder to extract three-dimensional motion features as motion control information. This method can extract accurate three-dimensional motion features from the motion image sequence, thereby improving the accuracy of motion control information.

[0071] S203. Extract semantic information from the reference image using the feature extraction module of the video generation model.

[0072] S204. Extract features from semantic information to obtain deep semantic information.

[0073] For example, the image generation module of the video generation model takes the denoising U-NET model as an example. The denoising U-NET model includes multiple denoising units; each denoising unit is connected serially. The number of denoising units is a third predetermined number. Each denoising unit includes a first attention layer, a second attention layer, a third attention layer, and a fourth attention layer in sequence.

[0074] For example, the feature extraction module of the video generation model uses a reference network, which includes multiple reference units connected serially. The number of reference units is a predetermined number. Each reference unit includes a fifth attention layer and a sixth attention layer. The first and fifth attention layers are identical, both using a spatial attention mechanism. The second and sixth attention layers are also identical, both using a cross-attention mechanism. The third attention layer can be an audio attention mechanism. The fourth attention layer can be a temporal attention mechanism. The fourth attention layer enables temporal alignment, ensuring the coherence of each frame in the generated target video.

[0075] For example, the motion encoder uses a first predetermined number of first residual units and a second predetermined number of second residual units. The sum of the first and second predetermined numbers equals a third predetermined number, or equals a fourth predetermined number. The third and fourth predetermined numbers are equal.

[0076] For example, the operation process within each denoising unit is similar. Taking the current denoising unit as an example, the motion control information output by the residual block in the corresponding residual unit, the semantic information output by the corresponding reference unit, the input of the current denoising unit, and the driving audio information are fused to obtain the output of the current denoising unit, which is then used as the input of the next denoising unit. When the current denoising unit is the first denoising unit, the input of the current denoising unit is only the semantic information. When the current denoising unit is not the first denoising unit, the input of the current denoising unit is the output of the previous denoising unit and the semantic information. When the current denoising unit is the last denoising unit, the output of the current denoising unit, after being decoded by the decoder, can be used as the target video.

[0077] For example, when the current denoising unit is not the first denoising unit, the first attention layer extracts features from the semantic information output by the fifth attention layer and the input of the current denoising unit to obtain deep semantic information, which is then used as the output of the first attention layer.

[0078] For example, when the current denoising unit is the first denoising unit, the semantic information output by the fifth attention layer is used to extract features through the first attention layer to obtain deep semantic information, which is then used as the output of the first attention layer.

[0079] S205. Normalize the motion control information.

[0080] For example, continuing with the current denoising unit, the motion control information output by the residual block in the corresponding residual unit is normalized by setting an activation function, such as the tanh activation function, so that the dimension (i.e., matrix shape) of the normalized motion control information is the same as the dimension of the output of the first attention layer, so as to facilitate the processing of the input of the second attention layer.

[0081] S206. The deep semantic information and the normalized motion control information are fused together to obtain the adjusted semantic information.

[0082] For example, continuing with the current denoising unit, the normalized motion control information and the output (deep semantic information) of the first attention layer are multiplied element-wise to obtain the input of the second attention layer. The input of the second attention layer and the semantic information output of the sixth attention layer are then input into the second attention layer to obtain its output, which can be used to obtain the adjusted semantic information.

[0083] In this embodiment, feature extraction is performed on semantic information to obtain deep semantic information; motion control information is normalized; and the deep semantic information and the normalized motion control information are fused together to make the target video contain rich three-dimensional motion features and improve the naturalness of the target video.

[0084] S207. Normalize the facial expression motion parameters to obtain normalized facial expression information.

[0085] Among them, facial expression motion parameters are used to determine motion control information. These facial expression motion parameters can be the facial expression motion parameters from the first set of motion parameters.

[0086] In this embodiment, before normalizing the facial motion parameters, a first multilayer perceptron can be used to transform the dimensions of the facial motion parameters and the driving audio information, making the transformed facial motion parameters and the transformed driving audio information have the same dimension, thus facilitating the subsequent fusion of the facial motion parameters and the driving audio information. The driving audio information can be encoded driving audio information.

[0087] In this embodiment, a random embedding can be made between the transformed facial expression motion parameters and the transformed driving audio information.

[0088] For example, taking facial expression motion parameters as an example, the transformed facial expression motion parameters are input into the second multilayer perceptron to generate first target parameters (including scaling parameters and offset parameters). The first target parameters are then normalized using an adaptive layer normalization mechanism to obtain normalized facial expression information. The parameters in the first multilayer perceptron are randomized, while the parameters in the second multilayer perceptron are initially set to zero to improve the accuracy of the first target parameters.

[0089] For example, taking driving audio information as an example, the transformed driving audio information is input into the second multilayer perceptron to generate second target parameters (including scaling parameters and offset parameters). The second target parameters are then normalized using an adaptive layer normalization mechanism to obtain normalized audio information.

[0090] S208. Generate the target video based on normalized facial expression information, adjusted semantic information, and driving information.

[0091] For example, continuing with the current denoising unit, and using facial motion parameters as the embedding, the output of the second attention layer (i.e., the adjusted semantic information), the transformed driving audio information, and the normalized facial expression information are input into the third attention layer to obtain the output of the third attention layer; the output of the third attention layer is then input into the fourth attention layer to obtain the output of the current denoising unit.

[0092] For example, continuing with the current denoising unit and using driving audio information as the embedding, the output of the second attention layer (i.e., the adjusted semantic information), the transformed facial motion parameters, and the normalized audio information are input into the third attention layer to obtain its output. The output of the third attention layer is then input into the fourth attention layer to obtain the output of the current denoising unit. When the current denoising unit is the last denoising unit, its output, after decoding by the decoder, can be used as the target video. Specifically, if the reference image is a reference facial image and the motion image sequence is a facial motion image sequence, then each frame in the corresponding target video is a face-related image. If the reference image is a reference full-body image and the motion image sequence is a full-body motion image sequence, then each frame in the corresponding target video is a full-body related image.

[0093] In this embodiment, the facial motion parameters are normalized to obtain normalized facial information. Based on the normalized facial information, adjusted semantic information, and driving information, the target video is generated in a way that can increase the richness of facial expressions in the target video and improve the accuracy and naturalness of the synchronization between driving audio information and facial expressions, thereby improving the naturalness and realism of the target video.

[0094] In this embodiment, a motion image sequence is generated based on pre-extracted first motion parameters. The motion image sequence is then input into a motion encoder to extract the three-dimensional motion features of each motion image as motion control information. Semantic information is extracted from a reference image using a feature extraction module of a video generation model. Feature extraction is performed on the semantic information to obtain deep semantic information. The motion control information is normalized. The deep semantic information and the normalized motion control information are fused to obtain adjusted semantic information. Facial motion parameters are normalized to obtain normalized facial expression information. A target video is generated based on the normalized facial expression information, adjusted semantic information, and driving information. This embodiment, by extracting driving audio information from the driving information, extracting semantic information from the reference image, extracting deep semantic information, normalizing motion control information, fusing deep semantic information and normalized motion control information, normalizing facial motion parameters, and generating a target video based on the normalized facial expression information, adjusted semantic information, and driving audio information obtained from the above series of operations, makes the motion changes in the target video more reasonable, and the coordination with the driving audio information more natural and harmonious, thus improving the effect of the target video.

[0095] Figure 3 This is a schematic flowchart of another video generation method provided in an embodiment of the present invention. See also... Figure 3 The method provided in this embodiment of the invention specifically includes the following steps:

[0096] S301. Generate an illumination image based on the pre-extracted second motion parameters.

[0097] In this embodiment, any method can be used to generate the illumination image based on the second motion parameter. For example, a single-view method can be used. Figure 3 The FLAME model in the DECA framework, a 3D reconstruction method, generates illumination images. These illumination images can serve as the global illumination image in the target video, representing the illumination conditions for each frame. The global illumination image for each target video can be different.

[0098] In this embodiment, the second motion parameter is not limited and can include any number of key parameters in 3D reconstruction.

[0099] The second motion parameters include a second variable parameter and a second fixed parameter; wherein, the second fixed parameter includes at least one of the following: head shape parameter, posture motion parameter, facial expression motion parameter, and shooting parameter of the shooting point; the second variable parameter includes the illumination parameter of the illumination ball.

[0100] In this embodiment, the initial second motion parameters can be extracted from any frame of the second set video using any open-source tool. The second set video can be the same as or different from the first set video.

[0101] In this embodiment, the process of generating the illumination image can be as follows: For any frame of the second set video, the parameters in the initial second motion parameters are adjusted to obtain the corresponding second motion parameters: the head shape parameters, posture motion parameters, facial expression motion parameters, and shooting parameters of the shooting point are all set to their corresponding fixed values, and the illumination parameters of the illumination sphere are directly used as the second variable parameters. The illumination parameters of the illumination sphere can be parameters extracted from the second set video, or the illumination parameters can be adjusted according to the actual situation, and the adjusted illumination parameters are used as the second variable parameters. The second motion parameters are then rendered using the FLAME model to obtain the illumination image.

[0102] In this embodiment, during the generation of the target video, the lighting image can be used as the global image in the target video, that is, the lighting image of each frame in the target video is the same, so as to improve the naturalness of the target video.

[0103] In this embodiment, multiple sets of second motion parameters can also be obtained, and a sequence of illumination images can be generated based on these multiple sets of second motion parameters, so that the illumination image of each frame in the target video is different, thereby improving the flexibility of the target video. Specifically, for each set of second motion parameters, the second variable parameters are different, while the second fixed parameters are the same. Each set of second motion parameters corresponds to one illumination image.

[0104] S302. Input the illumination image into the illumination encoder and extract the illumination features of the illumination image as illumination change control information.

[0105] In this embodiment, there are no restrictions on the deep learning model used by the illumination encoder, as long as it can extract the illumination features of the illumination image.

[0106] Optionally, the illumination encoder includes multiple first convolutional layers and a self-attention layer and a second convolutional layer; each first convolutional layer is connected in series, and the multiple first convolutional layers, self-attention layers and second convolutional layers are connected in series sequentially; the first convolutional layers and the second convolutional layers are different.

[0107] In this embodiment, the self-attention layer in the illumination encoder is the same as the self-attention layer in the motion encoder. In this embodiment, the number of the first convolutional layers is not limited; for example, it can be eight. In this embodiment, the parameters in the first convolutional layer are randomly initialized, which helps with gradient flow and feature extraction during the learning process of the illumination encoder. The parameters in the second convolutional layer are initialized to all zeros, which simplifies the initialization process of the illumination encoder and reduces its complexity. Each convolutional layer is followed by an activation function, which helps introduce non-linearity and enhances the expressive power of the illumination encoder.

[0108] In this embodiment, by using multiple first convolutional layers and including self-attention layers and second convolutional layers to extract illumination features from illumination images, the feature extraction capability of the illumination encoder can be enhanced, thereby generating more accurate illumination change control information.

[0109] In this embodiment, an illumination image is generated based on the pre-extracted second motion parameters; the illumination image is then input into an illumination encoder to extract the illumination features of the illumination image, which are used as illumination change control information. This method can extract accurate illumination features from the illumination image and improve the accuracy of the illumination change control information.

[0110] S303. Extract semantic information from the reference image through the feature extraction module of the video generation model.

[0111] S304. The semantic information and the illumination change control information are fused to obtain illumination fusion information, which is then used as the adjusted semantic information.

[0112] For example, the image generation module of the video generation model takes the denoising U-NET model as an example. The denoising U-NET model includes multiple denoising units; each denoising unit is connected serially. The number of denoising units is a third predetermined number. Each denoising unit includes a first attention layer, a second attention layer, a third attention layer, and a fourth attention layer in sequence.

[0113] For example, the feature extraction module of the video generation model uses a reference network, which includes multiple reference units connected serially. The number of reference units is a predetermined number. Each reference unit includes a fifth attention layer and a sixth attention layer. The first and fifth attention layers are identical, both using a spatial attention mechanism. The second and sixth attention layers are also identical, both using a cross-attention mechanism. The third attention layer can be an audio attention mechanism. The fourth attention layer can be a temporal attention mechanism.

[0114] For example, the operation process within each denoising unit is similar. Taking the current denoising unit as an example, the semantic information output by the corresponding reference unit, the input of the current denoising unit, and the driving audio information are fused to obtain the output of the current denoising unit, which is then used as the input of the next denoising unit. When the current denoising unit is the first denoising unit, the input of the current denoising unit is the information obtained by adding the processed noise and the encoded illumination change control information, along with the semantic information, which can be called the processed illumination change control information. When the current denoising unit is not the first denoising unit, the input of the current denoising unit is the output of the previous denoising unit and the semantic information. When the current denoising unit is the last denoising unit, the output of the current denoising unit, after being decoded by the decoder, can be used as the target video.

[0115] For example, when the current denoising unit is not the first denoising unit, the semantic information output by the fifth attention layer and the input of the current denoising unit are fused by the first attention layer to form the adjusted semantic information, which is then used as the output of the first attention layer.

[0116] For example, when the current denoising unit is the first denoising unit, the semantic information output by the fifth attention layer and the processed illumination change control information are fused by the first attention layer to obtain illumination fusion information, which is used as the adjusted semantic information and can also be used as the output of the first attention layer.

[0117] In this embodiment, semantic information and illumination change control information are fused to obtain illumination fusion information, which is then used as adjusted semantic information. This makes the semantic information include illumination change control information, thereby improving the adaptability of the target video under different illumination conditions.

[0118] S305. Generate the target video based on the adjusted semantic information and driving information.

[0119] For example, continuing with the current denoising unit, the output of the first attention layer (adjusted semantic information) and the semantic information output of the sixth attention layer are input into the second attention layer to obtain the output of the second attention layer. The output of the second attention layer and the driving audio information are input into the third attention layer to obtain the output of the third attention layer; the output of the third attention layer is input into the fourth attention layer to obtain the output of the current denoising unit. When the current denoising unit is the last denoising unit, the output of the current denoising unit, after being decoded by the decoder, can be used as the target video.

[0120] In this embodiment, a lighting image is generated based on pre-extracted second motion parameters. The lighting image is then input into a lighting encoder to extract its lighting features, which are used as lighting change control information. Semantic information is extracted from a reference image using the feature extraction module of the video generation model. The semantic information and lighting change control information are fused to obtain lighting fusion information, which is used as adjusted semantic information. The target video is then generated based on the adjusted semantic information and driving information. This embodiment, by extracting semantic information from a reference image, fusing it with lighting change control information to obtain adjusted semantic information, and generating the target video based on the adjusted semantic information and driving information, enables the target video to adapt to different lighting conditions, thereby improving the naturalness and realism of the target video in different environments.

[0121] Figure 4 This is a schematic flowchart of another video generation method provided in an embodiment of the present invention. See also... Figure 4 The method provided in this embodiment of the invention specifically includes the following steps:

[0122] S401. Generate a sequence of motion images based on multiple pre-extracted sets of first motion parameters.

[0123] Optionally, each set of first motion parameters includes a first variable parameter and a first fixed parameter; wherein, the first variable parameter includes at least one of the following: head shape parameter, posture motion parameter, facial expression motion parameter, and shooting parameter of the shooting point; the first fixed parameter includes the illumination parameter of the illumination ball; the first variable parameter of each set of first motion parameters is different.

[0124] S402. Input the motion image sequence into the motion encoder and extract the three-dimensional motion features of each motion image as motion control information.

[0125] Optionally, the motion encoder includes multiple residual units; each residual unit is connected in series.

[0126] S403. Generate an illumination image based on the pre-extracted second motion parameters.

[0127] Optionally, the second motion parameter includes a second variable parameter and a second fixed parameter; wherein, the second fixed parameter includes at least one of the following: head shape parameter, posture motion parameter, facial expression motion parameter, and shooting parameter of the shooting point; the second variable parameter includes the illumination parameter of the illumination ball.

[0128] S404. Input the illumination image into the illumination encoder and extract the illumination features of the illumination image as illumination change control information.

[0129] Optionally, the illumination encoder includes multiple first convolutional layers and a self-attention layer and a second convolutional layer; each first convolutional layer is connected in series, and the multiple first convolutional layers, self-attention layers and second convolutional layers are connected in series sequentially; the first convolutional layers and the second convolutional layers are different.

[0130] S405. Extract semantic information from the reference image using the feature extraction module of the video generation model.

[0131] S406. The semantic information and the illumination change control information are fused to obtain illumination fusion information.

[0132] For example, the image generation module of the video generation model takes the denoising U-NET model as an example. The denoising U-NET model includes multiple denoising units; each denoising unit is connected serially. The number of denoising units is a third predetermined number. Each denoising unit includes a first attention layer, a second attention layer, a third attention layer, and a fourth attention layer in sequence.

[0133] For example, the feature extraction module of the video generation model uses a reference network, which includes multiple reference units connected serially. The number of reference units is a predetermined number. Each reference unit includes a fifth attention layer and a sixth attention layer. The first and fifth attention layers are identical, both using a spatial attention mechanism. The second and sixth attention layers are also identical, both using a cross-attention mechanism. The third attention layer can be an audio attention mechanism. The fourth attention layer can be a temporal attention mechanism.

[0134] For example, the motion encoder uses a first predetermined number of first residual units and a second predetermined number of second residual units. The sum of the first and second predetermined numbers equals a third predetermined number, or equals a fourth predetermined number. The third and fourth predetermined numbers are equal.

[0135] For example, the operation process within each denoising unit is similar. Taking the current denoising unit as an example, the motion control information output by the residual block in the corresponding residual unit, the semantic information output by the corresponding reference unit, the input of the current denoising unit, and the driving audio information are fused to obtain the output of the current denoising unit, which is then used as the input of the next denoising unit. When the current denoising unit is the first denoising unit, the input of the current denoising unit is the information obtained by adding the processed noise and the encoded illumination change control information, along with the semantic information, which can be called the processed illumination change control information. When the current denoising unit is not the first denoising unit, the input of the current denoising unit is the output of the previous denoising unit and the semantic information. When the current denoising unit is the last denoising unit, the output of the current denoising unit, after being decoded by the decoder, can be used as the target video.

[0136] For example, when the current denoising unit is not the first denoising unit, the semantic information output by the fifth attention layer and the input of the current denoising unit are fused by the first attention layer to form the adjusted semantic information, which is then used as the output of the first attention layer.

[0137] For example, when the current denoising unit is the first denoising unit, the semantic information output by the fifth attention layer and the processed illumination change control information are fused by the first attention layer to obtain illumination fusion information, which can be used as the output of the first attention layer.

[0138] S407. Normalize the motion control information.

[0139] For example, continuing with the current denoising unit, the motion control information output by the residual block in the corresponding residual unit is normalized by setting an activation function, such as the tanh activation function, so that the dimension (i.e., matrix shape) of the normalized motion control information is the same as the dimension of the output of the first attention layer, so as to facilitate the processing of the input of the second attention layer.

[0140] S408. The illumination fusion information and the normalized motion control information are fused together to obtain the adjusted semantic information.

[0141] For example, continuing with the current denoising unit, the normalized motion control information and the output (lighting fusion information) of the first attention layer are multiplied element-wise to obtain the input of the second attention layer. The input of the second attention layer and the semantic information of the output of the sixth attention layer are then input into the second attention layer to obtain the output of the second attention layer, which can be used to obtain the adjusted semantic information.

[0142] In this embodiment, semantic information and illumination change control information are fused to obtain illumination fusion information; the motion control information is normalized; and the illumination fusion information and the normalized motion control information are fused together so that the semantic information includes both illumination change control information and motion control information. This results in the target video not only containing rich three-dimensional motion features, but also improving the adaptability of the target video under different lighting conditions and enhancing the naturalness of the target video.

[0143] S409. Normalize the facial expression motion parameters to obtain normalized facial expression information.

[0144] Among them, facial expression motion parameters are used to determine motion control information. These facial expression motion parameters can be the facial expression motion parameters from the first set of motion parameters.

[0145] In this embodiment, before normalizing the facial motion parameters, a first multilayer perceptron can be used to transform the dimensions of the facial motion parameters and the driving audio information, making the transformed facial motion parameters and the transformed driving audio information have the same dimension, thus facilitating the subsequent fusion of the facial motion parameters and the driving audio information. The driving audio information can be encoded driving audio information.

[0146] In this embodiment, a random embedding can be made between the transformed facial expression motion parameters and the transformed driving audio information.

[0147] For example, taking facial expression motion parameters as an example, the transformed facial expression motion parameters are input into the second multilayer perceptron to generate first target parameters (including scaling parameters and offset parameters). The first target parameters are then normalized using an adaptive layer normalization mechanism to obtain normalized facial expression information. The parameters in the first multilayer perceptron are randomized, while the parameters in the second multilayer perceptron are initially set to zero to improve the accuracy of the first target parameters.

[0148] For example, taking driving audio information as an example, the transformed driving audio information is input into the second multilayer perceptron to generate second target parameters (including scaling parameters and offset parameters). The second target parameters are then normalized using an adaptive layer normalization mechanism to obtain normalized audio information.

[0149] S410: Generate the target video based on normalized facial expression information, adjusted semantic information, and driving information.

[0150] For example, continuing with the current denoising unit, and using facial motion parameters as the embedding, the output of the second attention layer (i.e., the adjusted semantic information), the transformed driving audio information, and the normalized facial expression information are input into the third attention layer to obtain the output of the third attention layer; the output of the third attention layer is then input into the fourth attention layer to obtain the output of the current denoising unit.

[0151] For example, continuing with the current denoising unit and using driving audio information as the embedding, the output of the second attention layer (i.e., the adjusted semantic information), the transformed facial motion parameters, and the normalized audio information are input into the third attention layer to obtain its output. The output of the third attention layer is then input into the fourth attention layer to obtain the output of the current denoising unit. When the current denoising unit is the last denoising unit, its output, after being decoded by the decoder, can be used as the target video.

[0152] In this embodiment, the facial motion parameters are normalized to obtain normalized facial information. Based on the normalized facial information, adjusted semantic information, and driving information, the target video is generated in a way that can increase the richness of facial expressions in the target video and improve the accuracy and naturalness of the synchronization between driving information and facial expressions, thereby improving the naturalness and authenticity of the target video.

[0153] In this embodiment, a sequence of motion images is generated based on multiple pre-extracted sets of first motion parameters; the sequence of motion images is input into a motion encoder to extract the three-dimensional motion features of each motion image as motion control information; an illumination image is generated based on pre-extracted second motion parameters; the illumination image is input into an illumination encoder to extract the illumination features of the illumination image as illumination change control information; semantic information is extracted from a reference image through the feature extraction module of the video generation model; the semantic information and illumination change control information are fused to obtain illumination fusion information; the motion control information is normalized; the illumination fusion information and the normalized motion control information are fused to obtain adjusted semantic information; facial expression motion parameters are normalized to obtain normalized facial expression information; and a target video is generated based on the normalized facial expression information, the adjusted semantic information, and the driving audio information. In this embodiment, semantic information is extracted from a reference image, and then fused with illumination change control information to obtain illumination fusion information. This fusion information is then fused with normalized motion control information to obtain adjusted semantic information. Finally, based on normalized facial expression information, adjusted semantic information, and driving audio information, a target video is generated. This method makes the generated target video not only more natural in terms of synchronization between body movement, facial expression movement, and driving information, but also more realistic in terms of adaptability to the lighting environment, thereby improving the overall quality of the video.

[0154] For example, Figure 5 This is a schematic diagram of the target video generation process provided in an embodiment of the present invention, such as... Figure 5 As shown, the image generation module of the video generation model takes the denoising U-NET model as an example. The denoising U-NET model includes multiple denoising units; each denoising unit is connected serially. The number of denoising units is a third predetermined number ( Figure 5 (Three are shown). Each denoising unit includes a first attention layer, a second attention layer, a third attention layer, and a fourth attention layer in sequence. The reference image is a reference facial image as an example.

[0155] The feature extraction module of the video generation model takes a reference network as an example. The reference network includes multiple reference units, and each reference unit is connected serially. The number of reference units is a fourth predetermined number ( Figure 5 (Three are shown). Each reference unit includes a fifth attention layer and a sixth attention layer. The first attention layer is identical to the fifth attention layer. The second and sixth attention layers are identical. Semantic information can be extracted from the reference facial image processed by the variational autoencoder using the reference network.

[0156] The driving speech in the driving information is encoded by an audio encoder, and the driving audio information can also be extracted from the encoded driving speech by the feature extraction module of the video generation model.

[0157] The number of residual units in a motion encoder ( Figure 5 The number of reference units is equal to the number of denoising units (shown in the diagram, which is 3).

[0158] The illumination image is input into the illumination encoder to extract the illumination features of the illumination image, which are then used as illumination change control information.

[0159] Extract a set of first motion parameters from each frame of the first set video, and generate a motion image sequence based on multiple sets of first motion parameters. Figure 5 In this example, the motion image sequence is a facial motion image sequence. Figure 5 In the first set video and the target video, each frame of the image is taken as a facial image.

[0160] When the current denoising unit is not the first denoising unit, the semantic information output by the fifth attention layer and the input of the current denoising unit are fused through the first attention layer to form the adjusted semantic information, which is then used as the output of the first attention layer. At this time, the input of the current denoising unit is the output of the previous denoising unit and the semantic information.

[0161] When the current denoising unit is the first denoising unit, the semantic information output by the fifth attention layer and the input of the current denoising unit are fused by the first attention layer to obtain the illumination fusion information, which can be used as the output of the first attention layer. At this time, the input of the current denoising unit is the information obtained by adding the processed noise and the encoded illumination change control information, as well as the semantic information.

[0162] Continuing with the current denoising unit as an example, the normalized motion control information and the output of the first attention layer are multiplied element-wise to obtain the input of the second attention layer. The input of the second attention layer and the semantic information from the output of the sixth attention layer are then input into the second attention layer to obtain its output, which can be used to obtain the adjusted semantic information. The output of the second attention layer (i.e., the adjusted semantic information) and the driving audio information are then input into the third attention layer to obtain its output; the output of the third attention layer is then input into the fourth attention layer to obtain the output of the current denoising unit. When the current denoising unit is the last denoising unit, its output, after being decoded by a variational autodecoder, can be used as the target video.

[0163] In this embodiment, semantic information is extracted from a reference facial image, and then fused with illumination change control information to obtain illumination fusion information. This fusion information is then fused with normalized motion control information to obtain adjusted semantic information. Finally, based on normalized facial expression information, adjusted semantic information, and driving audio information, a target video is generated. This method makes the generated target video not only more natural in terms of facial expressions, facial movements, and synchronized driving speech, but also more realistic in terms of adaptability to lighting environments, thereby improving the overall quality of the video.

[0164] For example, Figure 6 This is a schematic diagram of a reference image provided for an embodiment of the present invention. For example, Figure 7 This is a schematic diagram of each frame of the target video provided in an embodiment of the present invention. Figure 7 This refers to the effect when considering motion control information. Figure 8 This is a schematic diagram of each frame of the target video provided in an embodiment of the present invention. Figure 8 This is the effect achieved by combining motion control information and illumination change control information.

[0165] Figure 9 This is a schematic diagram of a video generation device provided in an embodiment of the present invention. The device includes: a change control information acquisition module 910 and a video generation module 920;

[0166] The change control information acquisition module 910 is used to acquire change control information for controlling the change process of the reference image;

[0167] The video generation module 920 is used to generate a target video based on the driving information and the change control information.

[0168] The technical solution of this disclosure involves a change control information acquisition module acquiring change control information for controlling the change process of the reference image; and a video generation module generating a target video based on driving information and the change control information. This embodiment of the disclosure, by generating a target video based on driving information and change control information, can achieve realistic video effects and improve the naturalness of the video.

[0169] Optionally, the change control information includes: motion control information and / or illumination change control information.

[0170] Optionally, the change control information acquisition module is specifically used to: generate a sequence of motion images based on multiple pre-extracted first motion parameters; input the sequence of motion images into a motion encoder, and extract the three-dimensional motion features of each motion image as the motion control information.

[0171] Optionally, each set of first motion parameters includes a first variable parameter and a first fixed parameter; wherein, the first variable parameter includes at least one of the following: head shape parameter, posture motion parameter, facial expression motion parameter, and shooting parameter of the shooting point; the first fixed parameter includes the illumination parameter of the illumination ball; the first variable parameter of each set of first motion parameters is different.

[0172] Optionally, the motion encoder includes multiple residual units; each of the residual units is connected in series.

[0173] Optionally, the change control information acquisition module is specifically used to: generate an illumination image based on the pre-extracted second motion parameters; input the illumination image into the illumination encoder to extract the illumination features of the illumination image as the illumination change control information.

[0174] Optionally, the second motion parameter includes a second variable parameter and a second fixed parameter; wherein, the second fixed parameter includes at least one of the following: head shape parameter, posture motion parameter, facial expression motion parameter, and shooting parameter of the shooting point; the second variable parameter includes the illumination parameter of the illumination ball.

[0175] The illumination encoder includes multiple first convolutional layers and a self-attention layer and a second convolutional layer; each first convolutional layer is connected in series, and the multiple first convolutional layers, the self-attention layer, and the second convolutional layer are connected in series sequentially; the first convolutional layer and the second convolutional layer are different.

[0176] Optionally, the video generation module is specifically used to: extract semantic information from a reference image through the feature extraction module of the video generation model; adjust the semantic information using the change control information through the image generation module of the video generation model; and generate a target video based on the adjusted semantic information and driving information.

[0177] Optionally, the video generation module is further configured to: extract features from the semantic information to obtain deep semantic information; normalize the motion control information; and fuse the deep semantic information and the normalized motion control information to obtain adjusted semantic information.

[0178] Optionally, the video generation module is further configured to: when the change control information is illumination change control information, fuse the semantic information and the illumination change control information to obtain illumination fusion information, and use it as the adjusted semantic information.

[0179] Optionally, the video generation module is further configured to: fuse the semantic information and the illumination change control information, whereby the change control information includes motion control information and illumination change control information, to obtain illumination fusion information; normalize the motion control information; and fuse the illumination fusion information and the normalized motion control information to obtain adjusted semantic information.

[0180] Optionally, the video generation module is further configured to: normalize the facial expression motion parameters to obtain normalized facial expression information; wherein the facial expression motion parameters are used to determine the motion control information; and generate a target video based on the normalized facial expression information, the adjusted semantic information, and the driving information.

[0181] Optionally, the above device further includes a training module, which is specifically used for: training the basic video generation model using a first training set to obtain a first video generation model; the first training set includes a first reference image set and a first video set; each first video in the first video set has the same portrait features, background features, and lighting features as the corresponding first reference image, but different three-dimensional motion features; adjusting the first video generation model using a second training set to obtain a second video generation model; the second training set includes a second reference image set and a second video set; each second video in the second video set has the same portrait features and background features as the corresponding second reference image, but different lighting features; adjusting the second video generation model using a third training set to obtain a third video generation model, which serves as the trained video generation model; wherein the third training set includes a third reference image set and a third video set, and each third video in the third video set includes the driving information.

[0182] The video generation apparatus provided in this disclosure can execute the video generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0183] Figure 10 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0184] like Figure 10 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0185] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0186] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the method of video generation.

[0187] In some embodiments, the method video generation may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method video generation described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform method video generation by any other suitable means (e.g., by means of firmware).

[0188] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0189] Computer programs used to implement the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0190] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0191] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0192] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0193] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0194] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the video generation method provided in any embodiment of this application.

[0195] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0196] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. A video generation method, characterized in that, include: Acquire change control information used to control the process of changing the reference image; The target video is generated based on the driving information and the change control information.

2. The method according to claim 1, characterized in that, The change control information includes: motion control information and / or illumination change control information.

3. The method according to claim 2, characterized in that, Acquiring motion control information for controlling the changes in the reference image includes: Based on multiple pre-extracted sets of first motion parameters, a sequence of motion images is generated; The motion image sequence is input into a motion encoder, and the three-dimensional motion features of each motion image are extracted as the motion control information.

4. The method according to claim 3, characterized in that, Each set of first motion parameters includes a first variable parameter and a first fixed parameter; wherein, the first variable parameter includes at least one of the following: head shape parameter, posture motion parameter, facial expression motion parameter, and shooting parameter of the shooting point; the first fixed parameter includes the illumination parameter of the illumination ball; the first variable parameter of each set of first motion parameters is different.

5. The method according to any one of claims 3-4, characterized in that, The motion encoder includes multiple residual units; each residual unit is connected in series.

6. The method according to any one of claims 2-5, characterized in that, Acquiring illumination change control information for controlling the change process of the reference image includes: Based on the pre-extracted second motion parameters, an illumination image is generated; The illumination image is input into the illumination encoder to extract the illumination features of the illumination image, which are used as the illumination change control information.

7. The method according to claim 6, characterized in that, in, The second motion parameters include a second variable parameter and a second fixed parameter; wherein, the second fixed parameter includes at least one of the following: head shape parameter, posture motion parameter, facial expression motion parameter, and shooting parameter of the shooting point; the second variable parameter includes the illumination parameter of the illumination ball.

8. The method according to any one of claims 6-7, characterized in that, in, The illumination encoder includes multiple first convolutional layers and a self-attention layer and a second convolutional layer; each first convolutional layer is connected in series, and the multiple first convolutional layers, the self-attention layer and the second convolutional layer are connected in series sequentially; the first convolutional layer and the second convolutional layer are different.

9. The method according to any one of claims 1-8, characterized in that, Generating the target video based on the driving information and the change control information includes: Semantic information is extracted from the reference image using the feature extraction module of the video generation model; The image generation module of the video generation model uses the change control information to adjust the semantic information; The target video is generated based on the adjusted semantic information and the driving information.

10. The method according to claim 9, characterized in that, When the change control information is motion control information, the semantic information is adjusted using the change control information, including: Feature extraction is performed on the semantic information to obtain deep semantic information; The motion control information is normalized. The deep semantic information and the normalized motion control information are fused together to obtain the adjusted semantic information.

11. The method according to claim 9, characterized in that, When the change control information is illumination change control information, the semantic information is adjusted using the change control information, including: The semantic information and the illumination change control information are fused to obtain illumination fusion information, which is then used as the adjusted semantic information.

12. The method according to claim 9, characterized in that, The change control information includes motion control information and illumination change control information. Therefore, the semantic information is adjusted using the change control information, including: The semantic information and the illumination change control information are fused to obtain illumination fusion information; The motion control information is normalized. The illumination fusion information and the normalized motion control information are fused together to obtain the adjusted semantic information.

13. The method according to claim 9, characterized in that, Generate a target video based on the adjusted semantic information and the driving information, including: The facial motion parameters are normalized to obtain normalized facial information; wherein, the facial motion parameters are used to determine motion control information. The target video is generated based on the normalized facial expression information, the adjusted semantic information, and the driving information.

14. The method according to any one of claims 9-13, characterized in that, The training methods for the video generation model include: The basic video generation model is trained using the first training set to obtain the first video generation model; the first training set includes a first reference image set and a first video set; each first video in the first video set has the same portrait features, background features, and lighting features as the corresponding first reference image, but different three-dimensional motion features; The first video generation model is adjusted using a second training set to obtain a second video generation model; the second training set includes a second reference image set and a second video set; each second video in the second video set has the same portrait features and background features as the corresponding second reference image, but different lighting features; The second video generation model is adjusted using a third training set to obtain a third video generation model, which serves as the trained video generation model. The third training set includes a third reference image set and a third video set, and each third video in the third video set includes the driving information.

15. A video generation apparatus, characterized in that, include: The change control information acquisition module is used to acquire change control information for controlling the change process of the reference image. The video generation module is used to generate a target video based on the driving information and the change control information.

16. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method as described in any one of claims 1-14.

17. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the video generation method as described in any one of claims 1-14.

18. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video generation method as described in any one of claims 1-14.