Method, apparatus, device, storage medium and program product for video generation
Patent Information
- Application Number
- CN202411355894.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2044-09-26
AI Technical Summary
[0009]应当理解,该部分中所描述的内容并非旨在限定本公开的实施例的关键特征或重要特征,也不用于限制本公开的范围。本公开的其他特征将通过以下的描述而变得容易理解。
Smart Images

Figure CN119364139B_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatuses, devices, storage media, and program products for video generation. Background Technology
[0002] With the continuous development of voice-driven face generation (TFG) technology, it has shown broad potential in application scenarios such as virtual character generation, video conferencing, and intelligent assistants.
[0003] Generating videos that are both high-quality and reflect personalized characteristics of the target object still faces many challenges, and effective solutions are urgently needed. Summary of the Invention
[0004] In a first aspect of this disclosure, a method for video generation is provided. The method may include: acquiring input information for generating the video, the input information including a source image containing a target object and driving conditions including at least target audio; determining a three-dimensional facial representation of the target object and source facial actions based on the source image; determining a sequence of driving facial actions synchronized with the target audio based at least on the target audio; and generating a video of the target object based on a difference sequence of actions between the source facial actions and the driving facial action sequence, and the three-dimensional facial representation of the target object, the video representing the target object speaking with the speech of the target audio.
[0005] In a second aspect of this disclosure, an apparatus for video generation is provided. The apparatus may include: an input information acquisition module configured to acquire input information for generating the video, the input information including a source image containing a target object and driving conditions including at least target audio; a target object information determination module configured to determine a three-dimensional facial representation of the target object and source facial actions based on the source image; a driving facial action sequence determination module configured to determine a driving facial action sequence synchronized with the target audio, at least based on the target audio; and a video generation module configured to generate a video of the target object based on a difference action sequence between the source facial actions and the driving facial action sequence, and the three-dimensional facial representation of the target object, the video representing the target object speaking with the speech of the target audio.
[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the electronic device to perform the method of the first aspect.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. A computer program is stored on the medium, which, when executed by a processor, implements the method of the first aspect.
[0008] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method of the first aspect.
[0009] It should be understood that the description in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;
[0012] Figure 2 A flowchart of a method for video generation according to some embodiments of the present disclosure is shown;
[0013] Figure 3 A schematic diagram of a method for video generation according to some embodiments of the present disclosure is shown;
[0014] Figure 4 An example diagram illustrating the fine-tuning process of a video synthesis model according to some embodiments of the present disclosure is shown;
[0015] Figure 5A A schematic diagram illustrating the audio-to-action model inference process according to some embodiments of the present disclosure is shown;
[0016] Figure 5B A schematic diagram illustrating the training of an audio-to-action model according to some embodiments of the present disclosure is shown;
[0017] Figure 5C A schematic diagram of training samples according to some embodiments of the present disclosure is shown;
[0018] Figure 6 A schematic structural block diagram of an apparatus for video generation according to some embodiments of the present disclosure is shown; and
[0019] Figure 7 A block diagram of an electronic device that can implement one or more embodiments of the present disclosure is shown. Detailed Implementation
[0020] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0021] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.
[0022] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.
[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0024] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, relevant users should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and authorization should be obtained from the relevant users. Among them, relevant users may include any type of rights holder, such as individuals, enterprises, and groups.
[0025] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly inform the user that the requested operation will require obtaining and using the user's information, thereby enabling the relevant user to choose whether to provide information to the software or hardware such as the electronic device, application, server, or storage medium that performs the operation of the technical solution disclosed herein based on the prompt message.
[0026] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide information to the electronic device.
[0027] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0028] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.
[0029] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, environment 100 may include electronic device 110.
[0030] In this example environment 100, electronic device 110 can acquire input information 102. Input information 102 includes at least target audio 114 and a source image 113 of a target object. For example, the target object can include humans, animals, cartoon characters, and virtual characters, etc. Electronic device 110 can generate a three-dimensional facial representation corresponding to the target object based on the source image 113 using a target model 115. Furthermore, electronic device 110 can also generate a facial motion sequence synchronized with the target audio 114 using the target model 115, based on the target audio 114. This facial motion sequence drives the movements of the three-dimensional facial representation corresponding to the target object. Since the movements of the three-dimensional facial representation are synchronized with the target audio, a video of the target object speaking using the target audio can be generated. Figure 1 The image only shows one target model 115 as an example; in reality, multiple different target models 115 may collaborate to complete the video generation.
[0031] Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry). Server-side equipment (not shown) can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc. Server-side equipment can, for example, provide background services for the applications of electronic device 110.
[0032] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0033] With the development of voice-driven facial generation technology, achieving personalization in video generation has become increasingly important. Personalization in video generation emphasizes the perceived similarity between the target person in the generated video and the real target person in terms of appearance and speaking style. Existing personalized video generation methods typically train an independent Neural Radiation Field (NeRF) model for each (class of) target person (or the type corresponding to the target task) to implicitly store the target person's feature information. However, this approach is inefficient and lacks generalization ability because each (class of) target person needs to be trained separately and training data is limited.
[0034] In embodiments of this disclosure, an improved scheme for video generation is proposed. In this scheme, an electronic device acquires input information for generating the video, including a source image containing a target object and driving conditions containing at least target audio. Based on the source image, a three-dimensional facial representation of the target object and source facial actions are determined. At least based on the target audio, a sequence of driving facial actions synchronized with the target audio is determined. Based on the difference sequence of actions between the source facial actions and the driving facial action sequence, and the three-dimensional facial representation of the target object, a video of the target object is generated, the video representing the target object speaking with the speech of the target audio.
[0035] Through the above process, a 3D facial representation of the target object is generated based on the source image, and a sequence of driving facial movements is generated synchronously with the target audio, enabling efficient generation of facial movements that match the audio. Furthermore, adjustments are made using the difference sequence between the source facial movements and the driving facial movements to ensure that the generated facial movements are more natural and coherent. This difference-driven generation method reduces reliance on large amounts of training data and makes the generation process more generalizable, adapting to different target objects and audio inputs, significantly improving the efficiency and adaptability of video generation.
[0036] Figure 2 An example flow diagram of a method 200 for video generation according to some embodiments of the present disclosure is shown. For ease of discussion, reference will be made to... Figure 1 The process 200 is described in the context of environment 100. In environment 100, video generation can be completed by electronic device 110, but some of the operations can be performed by requesting a server device (not shown) (such as determining the facial action sequence, video generation, or part of the model training process can be implemented at the server device).
[0037] In box 201, electronic device 110 acquires input information for generating video, including a source image containing the target object and driving conditions containing at least the target audio.
[0038] The electronic device 110 can obtain the input information for generating the video as a trigger condition for generating the video. Figure 3 A schematic diagram 300 of a method for video generation according to some embodiments of the present disclosure is shown. Input information may include at least a source image 301 containing a target object and target audio 305. Exemplarily, the target object may include humans, animals, cartoon characters, and virtual characters, etc. The source image 301 of the target object may be a still image or a video image obtained by extracting frames from a video containing the target object.
[0039] The source image 301 containing the target object can be used to determine a three-dimensional facial representation of the target object. The driving conditions containing the target audio 305 can be used to generate facial movements synchronized with the target object's speech content. Exemplarily, the driving conditions may also include other additional information (such as a reference video to reflect speaking style) to further optimize the generated video content.
[0040] In box 202, based on source image 301, electronic device 110 can determine a 3D facial representation 302 and source facial motion of the target object. The 3D facial representation 302 is a 3D geometric model generated by analyzing the source image 301 of the target object; it accurately describes the facial shape and structure of the target object to preserve its identity features. Source facial motion can be the initial facial pose or expression information of the target object in the source image, used to describe the state of the face. Source facial motion can be Projected Normalized Coordinate Code (PNCC), and the source facial motion can adopt PNCC... src express.
[0041] For example, the source image 301 containing the target object can be adopted as I src This indicates that 3D facial representation can be achieved using P... cano This indicates that P corresponds to this. cano =FaceRecon(I src FaceRecon can represent a model that realizes the reconstruction of an image into a 3D face, which is used to convert a received image into a 3D facial image.
[0042] In box 203, electronic device 110 determines a sequence of driven facial movements synchronized with the target audio, at least based on the target audio. Determining the sequence of driven facial movements can be accomplished using an audio-to-motion (ATO) model 310. The ATO model 310 can generate a sequence of driven facial movements synchronized with speech, at least based on the target audio. The sequence of driven facial movements can be multi-frame, and each frame can employ PNCC (Programmable Pixel Conversion)... tgt This sequence represents the driving sequence of facial movements generated under the guidance of target audio. This sequence reflects the dynamic changes in facial expressions and is synchronized with the target audio. The driving sequence of facial movements can be used to generate dynamic facial movements in videos.
[0043] To improve the accuracy of the facial motion sequence, the input information may also include a reference video 304. The reference video may contain a video of a reference subject speaking, reflecting the subject's speaking style. The reference subject can be the target subject or an object other than the target subject. Based on the motion reference provided by the reference video 304, the facial motion sequence can be made more realistic and closely resemble a real person.
[0044] In box 204, electronic device 110 generates a video of the target object based on the difference sequence between the source facial action and the driving facial action sequence, as well as the three-dimensional facial representation of the target object. The video represents the target object speaking with the speech of the target audio.
[0045] The source facial action can be combined with each driving facial action in the driving facial action sequence to form an action pair 307. The electronic device 110 can determine a difference action sequence based on the differences between the action pairs. The electronic device 110 uses a video synthesis model 320 to combine the difference action sequence with the three-dimensional facial representation 302 of the target object to generate a dynamic facial video 308 of the target object, which is synchronized with the audio to appear as if the target object is speaking.
[0046] For example, the generation of a video of the target object can be expressed by the formula:
[0047] I raw =VolumeRenderer(P cano +MotionAdapter[PNCC src PNCC tgt ],cam tgt (1)
[0048] MotionAdapter[PNCC src PNCC tgt [] can represent the difference between the source facial action and a given driving facial action in the driving facial action sequence. cam tgt It can represent the acquisition position of the driver camera that controls the head pose. VolumeRenderer can represent the processing procedure for volumetric rendering of the model. raw This can represent a low-resolution image obtained through volumetric rendering, which is then processed by a super-resolution model to generate the final high-resolution result: I pred =SuperResolution(I raw SuperResolution can represent the processing procedure of a super-resolution model.
[0049] Through the above process, combining 3D facial representation and driven facial motion sequences, the accuracy and personalization of video generation are effectively improved. The generated video not only maintains the personalized appearance of the target object but also ensures a high degree of synchronization between facial movements and the speech content of the target audio, ultimately presenting natural and smooth facial movements. By calculating differential motion sequences, the matching degree between facial expressions and movements is further optimized, making the generated results more realistic and efficient.
[0050] Depending on the information contained in the driving conditions, the process of determining the driven facial motion sequence can include two categories. The first category is based on the audio feature information of the target audio. The second category is based on the audio feature information of the target audio and the motion feature information of the reference video. The first category is described below: Electronic device 110 extracts the audio feature information of the target audio 305, which at least indicates the duration of the target audio 305 and the pronunciation characteristics of each sound unit. Based on the pronunciation characteristics of each sound unit, a driven facial motion sequence corresponding to the duration of the target audio 305 is determined.
[0051] Electronic device 110 extracts audio feature information from target audio 305. This feature information includes at least the duration of target audio 305 and the pronunciation features of each sound unit. The duration can represent the total duration of target audio 305, while the pronunciation features reflect the speech attributes of each sound unit in target audio 305, such as pitch, speech rate, and volume.
[0052] Based on the articulation features of each extracted sound unit, the electronic device 110 can use the audio-to-action model 310 to determine a sequence of driving facial movements corresponding to the duration of the target audio 305. By matching audio features with facial movements, it ensures that the generated facial movement sequence accurately reflects the speech changes in the target audio 305. Thus, facial movements synchronized with the target audio 305 can be generated to achieve dynamic matching between the target object's facial movements and the target audio 305.
[0053] The second type will now be introduced: combination Figure 3 As shown, in response to the driving condition, a reference video 304 is also included. The electronic device 110 extracts audio feature information of the target audio 305 and motion feature information of the reference video. The audio feature information at least indicates the duration of the target audio 305 and the pronunciation features of each sound unit, and the motion feature information at least indicates the facial motion features of the reference object. Based on the pronunciation features of each sound unit and the facial motion features, a driving facial motion sequence that meets the duration is generated.
[0054] The audio feature information of the target audio 305 is the same as in the previous example, and may include at least the duration of the target audio and the pronunciation features of each sound unit in the target audio 305. The reference video 304 may be a video containing the speech content of a reference subject, which can reflect the speech style and facial movement features of the reference subject. The reference subject may be the target subject itself, or it may be another object different from the target subject. The reference video 304 can provide reference information on speech style or facial expressions. The motion feature information of the reference video can reflect the facial movement features of the reference subject in the reference video, such as changes in facial expressions such as mouth opening and closing, and eye blinking.
[0055] Electronic device 110 generates a driven facial movement sequence that matches the duration of the target audio 305 by combining the pronunciation features of the target audio 305 with the facial movement features of the reference video 304. This driven facial movement sequence not only accurately reflects the speech changes in the target audio 305, but also reflects the facial movement changes of the reference object in the reference video 304, so that the final generated video can simultaneously possess audio synchronization and the facial expression style reflected in the reference video.
[0056] For the second type, the process of generating the driving facial action sequence may specifically include: iteratively executing the following steps until a preset condition is met: determining the first facial action change rate corresponding to each time step in the duration based on the pronunciation features and facial action features of each voice unit; determining the first facial action sequence based on the first facial action change rate and the initial facial action sequence; determining the second facial action change rate corresponding to each time step based on the first facial action sequence, the first facial action change rate, the pronunciation features and facial action features of each voice unit; determining the second facial action sequence based on the second facial action change rate and the first facial action sequence; and generating the driving facial action sequence based on the second facial action sequence determined after meeting the preset condition.
[0057] In generating the driving facial action sequence, each iteration processes a complete action sequence of the same length as the entire target audio 305. The purpose of iteration is to improve the accuracy of the sequence through multiple optimizations. For each time step within the duration, the electronic device 110 can use the audio-to-action model 310 to determine the first facial action change rate within the corresponding duration based on the pronunciation characteristics of each sound unit (such as pitch, speed, etc.) and the facial action features in the reference video. The first facial action change rate reflects the rate of change of facial actions in the driving facial action sequence within the time step, and combined with the characteristics of the audio, ensures that the facial action is consistent with the sound rhythm of the target audio 305.
[0058] A first facial action sequence is generated by applying a first facial action change rate to an initial facial action sequence. If it corresponds to the first time step, the initial facial action sequence can be a full-noise facial action sequence with the same length as the target audio. If it is not the first time step, the initial facial action sequence is the facial action sequence determined in the time step preceding the current time step.
[0059] Based on the generated first facial motion sequence, the first facial motion change rate, the pronunciation features of the voice unit, and the facial motion features in the reference video, a new second facial motion change rate can be iteratively determined. This process optimizes the facial motion change rate, ensuring that the facial motion more accurately matches the expression changes in the audio and reference video.
[0060] By combining the speed of change of the second facial movement with the sequence of the first facial movement, a more accurate second facial movement sequence can be generated. The second facial movement sequence has been optimized based on the first facial movement sequence, gradually improving the synchronization and precision between facial movements and audio.
[0061] Once the iterations meet the preset conditions, the optimized second facial motion sequence can be used to generate a driving facial motion sequence. This sequence is synchronized with the target audio at each time step, exhibiting natural and fluid facial movements, and is ultimately used to generate a video of the target object. The preset conditions may include the rate of motion change being less than a change threshold (indicating convergence), the generated facial motion sequence meeting the expected quality standards, the iterations reaching the set number of iterations, and so on.
[0062] Through the above process, the audio-to-action model 310 can obtain an action sequence that can be synchronized with the target audio 305, and can also realize the facial action features in the reference video 304. Each iteration optimizes the entire action sequence, gradually approximating a precise and natural facial action sequence from the initial coarse action sequence.
[0063] In some embodiments of this disclosure, the video of the target object is obtained using a trained video synthesis model 320, which includes a base model and an additional model. The base model is trained through pre-training, and the additional model is trained through fine-tuning, during which the parameters of the pre-trained base model remain unchanged.
[0064] The video synthesis model 320 can be composed of a base model and an additional model. The base model is trained via pre-training and primarily provides the basic features and structure required for video generation. The additional model, trained via fine-tuning, is responsible for handling the personalized features of the target object, such as facial expressions and dynamic changes. During the fine-tuning of the additional model, the parameters of the base model remain unchanged, allowing the additional model to converge faster and requiring fewer training samples and time. This is because the base model has already learned rich general features, while the additional model trained via fine-tuning only needs to focus on the detailed adjustments of personalized features, thus greatly improving training efficiency and effectiveness.
[0065] This approach, by fixing the basic model and fine-tuning the design of additional models, can improve efficiency and adapt to the personalized needs of different target objects.
[0066] The following example, using electronic device 110 as an example, details the fine-tuning process of the supplementary model. It should be noted that the fine-tuning process of the supplementary model can also be performed by other electronic devices or by a server-side device. Electronic device 110 generates a video prediction result for the object sample based on the difference action sequence samples between the source facial action samples and the driving facial action sequence samples of the object sample, and the 3D facial representation of the object sample. The driving facial action sequence samples are determined from the audio contained in the sample video of the object sample. The 3D facial representation of the object sample is initialized to a 3D facial representation determined from a reference frame of the sample video and is updated during the fine-tuning process of the supplementary model. The parameters of the supplementary model are fine-tuned based on the first difference between each video frame in the video prediction result and the corresponding video frames in the sample video.
[0067] Figure 4 A schematic diagram 400 showing a fine-tuning of an additional model according to some embodiments of the present disclosure is illustrated. (In conjunction with...) Figure 4 As shown, the basic model of the video synthesis model 320 may include a motion adaptation model 411-A, a volumetric rendering model 411-B, and a super-resolution model 411-C. Additional models of the video synthesis model 320 may include a low-rank adaptation (LoRA) model 421.
[0068] The sample video of the object sample can be a sample video including different facial styles. By parsing the audio contained in the sample video of the object sample using electronic device 110, the driving facial action sequence sample corresponding to the sample video of the object sample can be obtained. The source facial action sample of the object sample can be combined with each driving facial action in the driving facial action sequence sample to form a sample pair 422. By processing the sample pair 422 using the pre-trained motion adaptation model 411-A, the differential action sequence sample 403 can be obtained.
[0069] Figure 4 The 3D reconstruction model 412 is pre-trained. The 3D reconstruction model 412 processes reference frames 421 of a sample video containing object samples to obtain an initial 3D facial representation 404 of the object samples. The reference frame 421 of the sample video containing object samples can be the first frame of the sample video or any other frame. The initial 3D facial representation 404 can serve as a 3D facial representation of the object samples, and this 3D facial representation can serve as a learnable 3D facial representation 442.
[0070] The volumetric rendering model 411-B processes the 3D differential motion sequence sample 407 to obtain a low-resolution image. The 3D differential motion sequence sample 407 is obtained by fusing a learnable 3D facial representation 442 and the differential motion sequence sample 403. The low-resolution image is then processed by the super-resolution model 411-C to generate a high-resolution result. This high-resolution result corresponds to the video prediction result of the object sample.
[0071] During the processing of input data by the motion adaptation model 411-A, volumetric rendering model 411-B, and super-resolution model 411-C, the parameters of these basic models remain unchanged. On the other hand, the low-rank adaptation model 421 simultaneously participates in the processing of input data by the aforementioned basic models, adjusting the output results through its own parameters.
[0072] The video prediction result of the object sample obtained by adjusting the output result through the parameters of the low-rank adaptive model 421 can be expressed as:
[0073]
[0074] It can represent a learnable 3D facial representation 442, MotionAdapter([PNNC src ',PNCC tgt '] can represent the difference in facial motion samples between source facial motion samples and driving facial motion sequence samples based on object samples. cam drv ' can represent the acquisition position of the driving camera that controls head posture. SR can represent the processing procedure of the super-resolution model. I pred 'Can represent a video frame in the video prediction result. This can represent adjusting the learnable parameters in the low-rank adaptive model 421.
[0075] Based on the first difference between each video frame 431 of the video prediction result and the corresponding video frames 432 in the sample video, the parameters of the additional model are fine-tuned and the learnable 3D facial representation 442 is iteratively updated. That is, after each adjustment, the learnable 3D facial representation 442 and the learnable parameters of the low-rank adaptive model 421 can be adjusted based on the first difference. The training process continues after adjustment until the first difference between each video frame 431 of the video prediction result and the corresponding video frames 432 in the sample video converges.
[0076] The first difference between each video frame 431 in the video prediction result and the corresponding video frames 432 in the sample video can indicate at least one of the following: pixel differences between each video frame in the video prediction result and the corresponding video frames in the sample video, structural and texture differences between each video frame in the video prediction result and the corresponding video frames in the sample video, identity feature differences of the object samples contained in each video frame in the video prediction result and the corresponding video frames in the sample video, and authenticity differences between each video frame in the video prediction result and the corresponding video frames in the sample video.
[0077] Pixel differences can correspond to the pixel-level errors between each video frame 431 in the video prediction result and the corresponding video frames 432 in the sample video. For example, they can include absolute error loss (L1 loss), mean square error loss (L2 loss), and so on.
[0078] Structural and texture differences can correspond to the perceptual differences (LPIPS loss) between two images evaluated based on feature representations of video frames in the human visual system. By capturing the perceptual similarity of structural and texture details, the human visual system's understanding of image similarity is simulated.
[0079] The difference in identity features can correspond to the identity similarity (ID loss) of the objects appearing in each video frame 431 in the video prediction result and the corresponding video frames 432 in the sample video. That is, the training objective is not only to make facial movements similar, but also to ensure that the target facial features are consistent with identity features, such as facial contours, the shape and distance of the eyes and eyebrows, etc.
[0080] The realism difference corresponds to the realism assessment (GAN loss) between each video frame 431 in the video prediction result and the corresponding video frames 432 in the sample video. Specifically, the realism difference can be used to measure the difference in "realism" between each video frame 431 in the video prediction result and the corresponding video frames 432 in the sample video, that is, whether the discriminator can distinguish whether the generated video frame comes from the sample video.
[0081] For example, by assigning corresponding weights (λ) to different losses, the parameters of the additional model can be fine-tuned based on the first differences in each of the aforementioned dimensions, and the initial 3D facial representation of the object sample can be iteratively updated. The first difference can be expressed as:
[0082] L SD-Hybrid =λ1*L1+λ LPIPS *L LPIPS +λ ID *L ID +λ GAN *L GAN (3)
[0083] L SD-Hybrid L1 can represent the first difference, and L1 and λ1 can represent the pixel difference and its corresponding weight, respectively. LPIPS and λ LPIPS These can represent structural and textural differences, along with their corresponding weights. L ID and λ ID These can represent differences in identity characteristics and their corresponding weights, respectively. L GAN and λ GAN These can represent the differences in authenticity and their corresponding weights, respectively.
[0084] Therefore, the videos generated by the finely tuned model are not only accurate at the pixel level, but also of high quality in terms of structure, identity consistency and overall realism, which can better meet the multi-dimensional requirements of practical applications.
[0085] The above is based on Figure 4 The fine-tuning process of the video synthesis model 320 is described. During the inference phase, the 3D reconstruction model 412 processes the source image containing the target object to obtain a 3D facial representation. The motion adaptation model 411-A in the trained video synthesis model 320 processes the source facial motion and driving facial motion sequences output by the audio-to-motion model 310 to obtain a 3D differential motion sequence between the source and driving facial motion sequences. The volumetric rendering model 411-B processes the 3D differential motion sequence to obtain a low-resolution image. The 3D differential motion sequence is obtained by fusing the 3D facial representation and the driving facial motion sequence. The low-resolution image is processed by the super-resolution model 411-C to generate a high-resolution result. This high-resolution result corresponds to the video prediction result of the object sample. The low-rank adaptation model 421 simultaneously participates in the processing of input data by the motion adaptation model 411-A, the volumetric rendering model 411-B, and the super-resolution model 411-C.
[0086] The facial motion sequence in the aforementioned example is obtained using the audio-to-action model 310. This model is trained as follows: Occlusion processing is performed on a reference sample video containing object samples to obtain a processed reference sample video. The processed reference sample video is combined with a given sample audio to obtain the input sample. Based on each time step in the sample audio duration, the input sample is iteratively processed using the audio-to-action model to obtain the predicted facial motion sequence for each time step. This predicted result is generated based on the predicted results of the facial motion sequences at adjacent time steps. The parameters of the audio-to-action model are adjusted based on the second difference between the predicted facial motion sequence for each time step and the predicted facial motion sequence sample. The facial motion sequence sample is obtained based on the reference sample video.
[0087] Before introducing the training process of the audio-to-action model 310, we will first introduce the inference process of the audio-to-action model 310. Figure 5A A schematic diagram 500A illustrating the inference process of an audio-to-action model 310 according to some embodiments of the present disclosure is shown. Using the audio content contained in the reference video 304, the Mel spectrum 501 corresponding to the audio can be obtained. Using the audio contained in the reference video 304, a facial motion sequence 502 can be generated. The features of the reference video 304 can be represented as [B, T1, Ca1+Cm1]. B can represent the amount of data, T1 can represent the duration, Ca1 can represent the Mel spectrum 501, and Cm1 can represent the features corresponding to the facial motion sequence 502.
[0088] The features corresponding to the target audio 305 can be represented as [B, T2, Ca2]. B can represent the data volume, T2 can represent the duration, and Ca2 can represent the Mel spectrum 503 of the target audio 305. The features of the reference video 304 and the target audio 305 are combined and used as input data for the audio-to-action model 310. The audio-to-action model 310 determines the driving facial action sequence 511 based on the input data.
[0089] Figure 5BA schematic diagram 500B showing the specific structure of an audio-to-action model 310 according to some embodiments of the present disclosure is illustrated. The audio-to-action model 310 may include a forward inference submodule 310-A and a differential equation solving submodule 310-B. In the initial state x0 of the inference, the driving facial action sequence is blank or random noise. Subsequently, the forward inference submodule 310-A can begin to determine the initial change state x1 of the facial action based on the input data [B, T1+T2, Ca1+Ca2+Cm1]. The input data may be features (Ca1+Ca2+Cm1) corresponding to the reference video 304 and the target audio 305 in a duration of T(T1+T2), with a data volume of B. The differential equation solving submodule 310-B determines the rate of change of action dxt / dt at each time step based on the initial change of the facial action.
[0090] Through intermediate iterations (from state x2 to state x) tT-1 The differential equation solving submodule 310-B determines the rate of change of action dxt / dt for each iteration. This rate of change is combined with the change state output by the forward inference submodule 310-A to progressively generate a more coherent and natural sequence of facial movements. After multiple iterations, the final complete sequence of movements (x) can be generated. tT This result is highly synchronized with the input audio and video features and exhibits natural facial movements and expression changes.
[0091] The training process of the audio-to-action model 310 will be introduced next. Figure 5C A schematic diagram 500C of training samples according to some embodiments of the present disclosure is shown. The training samples include aligned sample audio 521 and sample video 522. Masking processing is performed on the sample video 522, whereby the masked portion can simulate missing data to generate a processed reference sample video. The training objective of the audio-to-action model 310 is to reconstruct the masked action sequence. Through training, the audio-to-action model 310 can fill in the missing action sequence based on the sample audio and the unmasked action portion, generating natural facial movement variations.
[0092] During training, the audio-to-action model 310 generates action sequences by calculating the rate of change of action (dxt / dt) at each time step. For example, in the initial state (x0), the input samples contain only sample audio and no sample video, and the model's output is a blank sequence of length T(T1+T2). In the first iteration (x1), the input samples include sample audio and the blank sequence from the initial state, and the audio-to-action model 310 can begin to generate preliminary changes in facial movements based on audio features (such as pitch, speech rate, stress, etc.). The audio-to-action model 310 determines the rate of change dxt / dt for the first time step, which is the predicted facial movement trend (such as mouth opening and closing or facial muscle movement) based on the audio signal.
[0093] The parameters of the audio-to-action model are adjusted based on the second difference between the predicted result of the driving facial action sequence and the driving facial action sequence sample at each time step. For example, the second difference can at least indicate the difference in the trajectory of the action change. Adjusting the parameters of the audio-to-action model 310 based on the second difference can be expressed as follows:
[0094]
[0095] L CFM This can be expressed as calculating the true velocity u at different time steps (t), under the true data distribution q(x) and the conditional probability distribution p(x|x1). t (x|x1) and model prediction speed v t The difference (i.e., error) between (x; θ) is calculated, and the expected value (mean) of that difference is taken. t (x; θ) can represent the prediction velocity (the result predicted at time step t based on the current state x and model parameters θ). t (x|x1) can represent the optimal path from the current state x to the target state x1 at time step t. t (x|x1) is the ideal rate of change of motion (true value) determined by driving facial motion sequence samples, which can be obtained based on reference sample videos.
[0096] For example, for data point x in the real data distribution q(x) t The true velocity at continuous time steps can be expressed as u. t = dxt / dt. Where the data points x0~p(x) follow a simple conditional probability distribution p(x) (such as a Gaussian distribution), starting at t=0, and as the time step approaches t=1, the velocity field pushes the distribution towards the true data q(x). This process can be represented by a differential equation:
[0097] dxt / dt=u t≈v t (x;θ)stx o ~p(x)t∈[0,1] (5)
[0098] The goal of the entire training process is to continuously optimize v t (x; θ), making it as close as possible to dxt / dt. Therefore, adjusting the parameters of the audio-to-action model 310 based on the second difference can be performed based on the following loss function:
[0099]
[0100] Through training, v can be obtained t (x t ;θ)≈dxt / dt, so the prediction result x1 can be obtained by expressing (5).
[0101] After multiple iterations of training, the input to the audio-to-action model contains rich facial motion information, and the rate of change dxt / dt gradually decreases because the generation of facial motions has become relatively stable. The final changes are more about fine-tuning details. This indicates the end of training.
[0102] In the aforementioned example, the second difference can indicate the difference in the trajectory of the change in motion (L). CFM Furthermore, the second difference can also refer to the audio-visual synchronization difference (Lsync) between the predicted facial action sequence and the sample facial action sequence. The Lsync difference assesses the degree of synchronization between the generated facial action sequence and the audio input; this difference is measured by a synchronization loss. The synchronization loss ensures that facial actions and speech rhythms remain synchronized in time, avoiding audio-visual asynchrony. For example, assigning corresponding weights (λ) to the action trajectory difference and the Lsync difference allows adjustment of the audio-to-action model parameters based on the second differences in each of the aforementioned dimensions. The second difference can be expressed as a loss function as follows:
[0103] L ICS-A2M =λ CFM *L CFM+ =λsync*Lsync (7)
[0104] L ICS-A2M This can represent the second difference. L CFM and λ CFM They can represent the differences in the trajectory of the action change and the corresponding weights, L CFM and λ CFM These can represent the differences in audio-visual synchronization and their corresponding weights.
[0105] Figure 6A schematic structural block diagram of an apparatus 600 for video generation according to some embodiments of the present disclosure is shown. The apparatus 600 may be implemented in or included in an electronic device 110, for example. The various modules / components in the apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.
[0106] like Figure 6 As shown, the device 600 may include an input information acquisition module 601, configured to acquire input information for generating a video, the input information including a source image containing a target object and driving conditions including at least target audio. A target object information determination module 602 is configured to determine a three-dimensional facial representation and source facial movements of the target object based on the source image. A driving facial movement sequence determination module 603 is configured to determine a driving facial movement sequence synchronized with the target audio, at least based on the target audio. A video generation module 604 is configured to generate a video of the target object based on a difference sequence of movements between the source facial movements and the driving facial movement sequence, and the three-dimensional facial representation of the target object, the video representing the target object speaking with the speech of the target audio.
[0107] In some embodiments of this disclosure, the facial motion sequence determination module 603 may be specifically configured to: extract audio feature information of the target audio, wherein the audio feature information at least indicates the duration of the target audio and the pronunciation features of each sound unit; and determine a facial motion sequence corresponding to the duration of the target audio based on the pronunciation features of each sound unit.
[0108] In some embodiments of this disclosure, the driving condition further includes a reference video. Based on this, the driving facial motion sequence determination module 603 can also be specifically configured to: extract audio feature information of the target audio and motion feature information of the reference video, wherein the audio feature information at least indicates the duration of the target audio and the pronunciation features of each sound unit, and the motion feature information at least indicates the facial motion features of the reference object. Based on the pronunciation features of each sound unit and the facial motion features, a driving facial motion sequence that meets the duration is generated.
[0109] In some embodiments of this disclosure, the facial motion sequence determination module 603 can be specifically configured to iteratively execute the following steps until a preset condition is met: Based on the pronunciation features and facial motion features of each sound unit, determine the first facial motion change rate corresponding to each time step in the duration. Based on the first facial motion change rate and the initial facial motion sequence, determine the first facial motion sequence. Based on the first facial motion sequence, the first facial motion change rate, the pronunciation features and facial motion features of each sound unit, determine the second facial motion change rate corresponding to each time step. Based on the second facial motion change rate and the first facial motion sequence, determine the second facial motion sequence. Based on the second facial motion sequence determined after meeting the preset condition, generate the driving facial motion sequence.
[0110] In some embodiments of this disclosure, the video of the target object is obtained using a trained video synthesis model, which includes a base model and an additional model. This also includes a model training module that trains the base model through pre-training and the additional model through fine-tuning. That is, the base model is trained through pre-training, and the additional model is trained through fine-tuning, with the parameters of the pre-trained base model remaining unchanged during the fine-tuning of the additional model.
[0111] In some embodiments of this disclosure, the model training module can be configured to: generate a video prediction result for the object sample based on the difference action sequence samples between the source facial action samples and the driving facial action sequence samples of the object sample, and the three-dimensional facial representation of the object sample; the driving facial action sequence samples are determined from the audio contained in the sample video of the object sample; the three-dimensional facial representation of the object sample is initialized to a three-dimensional facial representation determined from a reference frame of the sample video and is updated during the fine-tuning of the additional model; and the parameters of the additional model are fine-tuned based on the first difference between each video frame in the video prediction result and the corresponding video frames in the sample video.
[0112] In some embodiments of this disclosure, the first difference indicates at least one of the following: pixel differences between each video frame in the video prediction result and the corresponding video frames in the sample video; structural and texture differences between each video frame in the video prediction result and the corresponding video frames in the sample video; identity feature differences of object samples contained in each video frame in the video prediction result and the corresponding video frames in the sample video; and authenticity differences between each video frame in the video prediction result and the corresponding video frames in the sample video.
[0113] In some embodiments of this disclosure, the driving facial action sequence is obtained using an audio-to-action model. Based on this, the model training module can be configured to: perform occlusion processing on a reference sample video containing object samples to obtain a processed reference sample video; combine the processed reference sample video with a given sample audio to obtain an input sample; iteratively process the input sample using the audio-to-action model based on each time step in the sample audio duration to obtain a driving facial action sequence prediction result corresponding to each time step, wherein the driving facial action sequence prediction result corresponding to each time step is generated based on the driving facial action sequence prediction results corresponding to previously adjacent time steps; and adjust the parameters of the audio-to-action model based on a second difference between the driving facial action sequence prediction result corresponding to each time step and the driving facial action sequence sample, wherein the driving facial action sequence sample is obtained based on the reference sample video.
[0114] In some embodiments of this disclosure, the second difference indicates at least one of the following: a difference in the motion change trajectory between the driving facial motion sequence material and the driving facial motion sequence sample, and a difference in audio-visual synchronization between the driving facial motion sequence material and the driving facial motion sequence sample.
[0115] Figure 7 A block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 7 The electronic device 700 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 7 The illustrated electronic device 700 may include or be implemented as Figure 1 Electronic devices 110 or Figure 6 Device 600.
[0116] like Figure 7 As shown, electronic device 700 is in the form of a general-purpose electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processing unit 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 700.
[0117] Electronic device 700 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 700.
[0118] Electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 7 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 720 may include computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0119] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0120] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) via communication unit 740 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 700, or with any device that enables electronic device 700 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0121] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0122] According to an exemplary implementation of this disclosure, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform... Figure 2 The methods provided are among the various optional methods available in the code, so they will not be elaborated upon here.
[0123] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0124] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0125] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0127] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for video generation, comprising: Obtain input information for generating video, the input information including a source image containing a target object, and driving conditions containing at least target audio, the driving conditions also including a reference video; Based on the source image, determine the three-dimensional facial representation and source facial motion of the target object; Based at least on the target audio, determine a sequence of facial movements synchronized with the target audio, wherein determining the sequence of facial movements synchronized with the target audio includes: Extract audio feature information of the target audio and motion feature information of the reference video. The audio feature information at least indicates the duration of the target audio and the pronunciation features of each sound unit, and the motion feature information at least indicates the facial motion features of the reference object; and Based on the pronunciation features of each sound unit and the facial movement features, a driving facial movement sequence that satisfies the specified duration is generated; The generation of the facial movement sequence based on the pronunciation features of each sound unit and the facial movement features includes: Iteratively execute the following steps until the preset condition is met: Based on the pronunciation characteristics of each sound unit and the facial movement characteristics, the first facial movement change rate corresponding to each time step in the duration is determined; Based on the speed of change of the first facial movement and the initial facial movement sequence, the first facial movement sequence is determined; Based on the first facial movement sequence, the first facial movement change speed, the pronunciation characteristics of each sound unit, and the facial movement characteristics, the second facial movement change speed corresponding to each time step is determined; Based on the speed of change of the second facial movement and the first facial movement sequence, the second facial movement sequence is determined; and The driving facial action sequence is generated based on the second facial action sequence determined after satisfying the preset conditions; as well as Based on the difference sequence between the source facial movements and the driving facial movement sequence, and the three-dimensional facial representation of the target object, a video of the target object is generated, the video representing the target object speaking with the speech of the target audio.
2. The method according to claim 1, wherein the video of the target object is obtained using a trained video synthesis model, the video synthesis model comprising a base model and an additional model; The base model is trained through pre-training, and the additional model is trained through fine-tuning, wherein the parameters of the pre-trained base model remain unchanged during the fine-tuning of the additional model.
3. The method of claim 2, wherein the fine-tuning of the additional model comprises: Based on the difference action sequence samples between the source facial action samples and the driving facial action sequence samples of the object sample, and the 3D facial representation of the object sample, a video prediction result for the object sample is generated. The driving facial action sequence samples are determined from the audio contained in the sample video of the object sample. The 3D facial representation of the object sample is initialized to a 3D facial representation determined from a reference frame of the sample video and is updated during the fine-tuning of the additional model. as well as Based on the first difference between each video frame in the video prediction result and the corresponding video frames in the sample video, the parameters of the additional model are fine-tuned.
4. The method of claim 3, wherein the first difference indicates at least one of the following: The pixel differences between each video frame in the video prediction result and the corresponding video frames in the sample video. The structural and texture differences between each video frame in the video prediction result and the corresponding video frames in the sample video. The differences in identity features between each video frame in the video prediction result and the corresponding video frames in the sample video. The difference in authenticity between each video frame in the video prediction result and the corresponding video frames in the sample video.
5. The method of claim 1, wherein the driving facial motion sequence is obtained using an audio-to-action model, the audio-to-action model being trained in the following manner: The reference sample video containing the object sample is masked to obtain the processed reference sample video; The processed reference sample video and the given sample audio are combined to obtain the input sample; Based on each time step in the duration of the sample audio, the input sample is iteratively processed using an audio-to-action model to obtain the predicted result of the driving facial action sequence corresponding to each time step. The predicted result of the driving facial action sequence corresponding to each time step is generated based on the predicted result of the driving facial action sequence corresponding to the adjacent time step. Based on the second difference between the predicted result of the driven facial motion sequence and the driven facial motion sequence sample corresponding to each time step, the parameters of the audio-to-action model are adjusted, wherein the driven facial motion sequence sample is obtained based on the reference sample video.
6. The method of claim 5, wherein the second difference indicates at least one of the following: The difference between the predicted facial motion sequence and the motion change trajectory of the facial motion sequence sample is as follows. The audio-visual synchronization difference between the predicted results of the facial motion sequence and the samples of the facial motion sequence.
7. An apparatus for video generation, comprising: The input information acquisition module is configured to acquire input information for generating video, the input information including a source image containing a target object and a driving condition containing at least target audio, the driving condition further including a reference video; The target object information determination module is configured to determine the three-dimensional facial representation and source facial motion of the target object based on the source image; A facial motion sequence determination module is configured to determine a facial motion sequence synchronized with the target audio, based at least on the target audio, wherein determining the facial motion sequence synchronized with the target audio includes: Extract audio feature information of the target audio and motion feature information of the reference video. The audio feature information at least indicates the duration of the target audio and the pronunciation features of each sound unit, and the motion feature information at least indicates the facial motion features of the reference object; and Based on the pronunciation features of each sound unit and the facial movement features, a driving facial movement sequence that satisfies the specified duration is generated; The generation of the facial movement sequence based on the pronunciation features of each sound unit and the facial movement features includes: Iteratively execute the following steps until the preset condition is met: Based on the pronunciation characteristics of each sound unit and the facial movement characteristics, the first facial movement change rate corresponding to each time step in the duration is determined; Based on the speed of change of the first facial movement and the initial facial movement sequence, the first facial movement sequence is determined; Based on the first facial movement sequence, the first facial movement change speed, the pronunciation characteristics of each sound unit, and the facial movement characteristics, the second facial movement change speed corresponding to each time step is determined; Based on the speed of change of the second facial movement and the first facial movement sequence, the second facial movement sequence is determined; and The driving facial action sequence is generated based on the second facial action sequence determined after satisfying the preset conditions; as well as The video generation module is configured to generate a video of the target object based on the difference sequence between the source facial movements and the driving facial movement sequence, and the three-dimensional facial representation of the target object, the video representing the target object speaking with the speech of the target audio.
8. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 6.
10. A computer program product comprising computer-executable instructions that, when executed by a processor, implement the method of any one of claims 1 to 6.