Emotional speaking head video generation method and device, equipment and medium

By extracting audio emotion vectors from driving audio and decoupling them into facial expression features and continuous intensity values, separating identity features to eliminate the emotion bias of identity images, and using a pre-trained generative model and temporal extrapolation strategy to generate emotional speaking head videos, the problem of low accuracy of emotional expression in existing technologies is solved, and high accuracy and realistic expression synchronization are achieved.

CN121985198APending Publication Date: 2026-05-05PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-16
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing talking head video generation technology has low accuracy in terms of emotional expression. Traditional methods rely on the emotional state in the input image, lack dynamic changes, and have difficulty learning subtle and diverse emotional expressions, resulting in inconsistent emotional expression between the generated video and the audio or a single expression.

Method used

By extracting audio emotion vectors from driving audio and decoupling them into facial expression features and continuous intensity values, and separating identity features to eliminate the emotion bias of identity images, emotional speaking head videos are generated using a pre-trained generative model and a temporal extrapolation strategy.

Benefits of technology

It improves the accuracy and realism of generated emotional speaking head videos, synchronizes generated expressions with audio rhythm, and ensures the purity and accuracy of the emotional driving source.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985198A_ABST
    Figure CN121985198A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to the fields of financial science and technology and medical health, and discloses an emotional speaking head video generation method, device, equipment and medium, and the method comprises the steps: obtaining a driving audio and an identity image, and employing a pre-trained audio encoder to encode the driving audio to obtain an audio emotion vector; processing the identity image and the audio emotion vector based on an emotion face representation model to obtain a face identity feature, an expression feature and an emotion intensity value; generating a speaking head video clip through a pre-trained generation model according to the facial identity feature, the expression feature and the emotion intensity value; and generating an emotional speaking head video by adopting a time extrapolation strategy according to the speaking head video clip. The emotion expression accuracy and reality of emotion speaking head video generation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology and can be applied to the fields of fintech and healthcare. In particular, it relates to a method, apparatus, device, and medium for generating emotional speaking head videos. Background Technology

[0002] In recent years, audio-driven speech head generation technology has made significant progress in multimedia applications, and has been widely used in film production, metaverse, fintech (such as intelligent emotional interaction of virtual customer managers), and healthcare (such as remote mental health counseling and rehabilitation training). However, existing methods still have significant limitations in emotional expression: traditional methods can mostly only generate neutral lip movements, neglecting the importance of emotional interaction, resulting in generated facial expressions that heavily rely on the inherent emotional state of the input image and lack dynamic changes. On the other hand, although some emotion perception methods attempt to introduce emotional information, they still need to rely on external emotional reference videos or discrete emotion labels, thus causing input bias and emotion intensity saturation problems. This results in the emotional expression of the generated video being inconsistent with the emotional rhythm contained in the audio, or the expression being monotonous and lacking natural variation. In addition, due to the limitations of current datasets in terms of emotion category and intensity annotation, models struggle to learn nuanced and diverse emotional expressions, further restricting the naturalness and expressiveness of the generated videos. Summary of the Invention

[0003] This invention provides a method, apparatus, computer device, and medium for generating emotional speaking head videos, in order to solve the technical problem of low accuracy in emotional expression in existing speaking head video generation.

[0004] Firstly, a method for generating emotional speaking head videos is provided, including: Acquire driving audio and identity image, and use a pre-trained audio encoder to encode the driving audio to obtain an audio emotion vector; The identity image and the audio emotion vector are processed based on the emotion facial representation model to obtain facial identity features, expression features, and emotion intensity values; Based on the facial identity features, the expression features, and the emotion intensity value, a pre-trained generative model generates speaking head video clips. An emotional speaking video is generated based on the speaking video clip using a time extrapolation strategy.

[0005] Secondly, an emotional speaking head video generation device is provided, including: An encoding unit is used to acquire driving audio and identity image, and to encode the driving audio using a pre-trained audio encoder to obtain an audio emotion vector; The processing unit is used to process the identity image and the audio emotion vector based on the emotional facial representation model to obtain facial identity features, expression features and emotion intensity values; The first generation unit is used to generate a speaking head video segment based on the facial identity features, the expression features, and the emotion intensity value using a pre-trained generation model. The second generation unit is used to generate an emotional speaking video based on the speaking video segment using a time extrapolation strategy.

[0006] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described emotional speaking head video generation method.

[0007] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described emotional speaking head video generation method.

[0008] The aforementioned method, apparatus, computer device, and storage medium for generating emotional speaking head videos can acquire driving audio and an identity image, and encode the driving audio using a pre-trained audio encoder to obtain an audio emotion vector; process the identity image and the audio emotion vector based on an emotional facial representation model to obtain facial identity features, expression features, and emotion intensity values; generate speaking head video segments using a pre-trained generation model based on the facial identity features, expression features, and emotion intensity values; and generate emotional speaking head videos using a time extrapolation strategy based on the speaking head video segments. In this invention, the audio emotion vector is first extracted from the driving audio and decoupled into expression features and continuous intensity values; then, the identity features are separated to eliminate the emotional bias of the identity image; finally, an emotional speaking head video is generated using a time extrapolation strategy based on the speaking head video segments. This not only ensures the purity and accuracy of the emotional driving source but also synchronizes the generated expression with the audio rhythm, thus improving the accuracy and realism of the emotional expression in the generated emotional speaking head video. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a flowchart illustrating a method for generating emotional speaking head videos according to an embodiment of the present invention; Figure 2 yes Figure 1 A schematic diagram of a specific implementation of step S120; Figure 3 yes Figure 1 A schematic diagram of a specific implementation of step S130; Figure 4 yes Figure 1 A schematic diagram of a specific implementation of step S140; Figure 5 This is a flowchart illustrating a method for generating emotional speaking head videos in another embodiment of the present invention; Figure 6 This is a schematic block diagram of an emotional speaking head video generation device according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 8 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0012] The emotional speaking head-video generation method provided in this invention can be applied to either a client or a server. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. Currently, in the fintech and healthcare fields, the accuracy of emotional expression in existing speaking head-video generation is relatively low. To address this problem, this invention proposes an emotional speaking head-video generation method. This method first extracts audio emotion vectors from the driving audio and decouples them into facial expression features and continuous intensity values; then, it separates identity features to eliminate the emotional bias of the identity image; finally, it generates emotional speaking head-videos based on speaking head-video segments using a time extrapolation strategy. This not only ensures the purity and accuracy of the emotional driving source but also synchronizes the generated facial expressions with the audio rhythm, thus improving the accuracy and realism of emotional expression in the generated emotional speaking head-videos. The invention will be described in detail below through specific embodiments.

[0013] Please see Figure 1 As shown, Figure 1A flowchart of the emotional speaking head video generation method provided in the embodiment of the present invention includes the following steps: S110-S150.

[0014] S110. Obtain the driving audio and identity image, and use a pre-trained audio encoder to encode the driving audio to obtain an audio emotion vector.

[0015] Specifically, after acquiring the input driving audio and a single identity image, the driving audio and single identity image can be audio and images in the fintech field, such as emotional intelligent investment advisors. The driving audio is a recording of a senior investment advisor interpreting market trends: "Dear client, your technology sector fund has seen significant gains today, which is a very positive signal!"; the identity image is a standard professional photo of the investment advisor. The driving audio and single identity image can also be audio and images in the healthcare field, such as personalized mental health counseling assistants. The driving audio is a recording of a psychotherapist guiding meditation or conducting cognitive behavioral therapy: "Now, please gently close your eyes, feel the peace within, and let all the anxiety slowly dissipate..."; the identity image is a neutral-expression photo of the therapist or a doctor trusted by the user.

[0016] The processed driving audio is input into the audio encoder, allowing the feature extraction network in the audio encoder to extract semantic feature vectors. These semantic feature vectors are then mapped to a preset emotional feature space in the audio encoder via a projection layer or adapter network to output the audio emotional vector. It should be noted that in this embodiment, the feature extraction network is a deep Transformer network, which can capture prosody, pitch, and semantic information in the audio. It should also be noted that the processed driving audio is pre-processed and segmented. Pre-processing involves converting the original driving audio to mono, resampling its sampling rate to a standard sampling rate, automatically removing silent or background noise segments at the beginning and end of the audio using a speech activity detection algorithm, scaling the perceived volume of the entire audio segment to a uniform target level, and segmenting the continuous audio signal into a series of short-time overlapping frames (e.g., frame length 25ms, frame shift 10ms), applying a window function (e.g., Hamming window) to each frame to reduce spectral leakage.

[0017] S120. Based on the emotional facial representation model, the identity image and the audio emotion vector are processed to obtain facial identity features, expression features and emotion intensity values.

[0018] Specifically, the emotional facial representation model includes a neutral encoder and a regressive neural network. The neutral encoder is a neural network model that can eliminate emotional expression components from the input identity image, thereby obtaining facial identity features unrelated to emotion. Figure 2As shown, step S120 includes the following steps: S121-S122: S121, eliminating the emotional expression components in the identity image using the neutral encoder to obtain the facial identity features unrelated to emotion; S122, decoding the audio emotion vector using the regression neural network supervised by 2D continuous emotion labels to obtain the expression features and the emotion intensity value. More specifically, the step of eliminating the emotional expression components in the identity image using the neutral encoder to obtain the facial identity features unrelated to emotion includes: inputting the identity image into a pre-trained image encoder to obtain a hybrid feature vector that integrates identity information and expression information; inputting the hybrid feature vector into the neutral encoder for processing to separate the identity information and expression information in the hybrid feature vector, and removing the expression information to obtain the facial identity features. It should be noted that the single identity image is input into a pre-trained image encoder, which typically employs a deep convolutional neural network architecture to extract advanced semantic features from the input identity image through its multi-layer nonlinear transformations. In this process, the image encoder outputs a high-dimensional feature vector, which inevitably integrates two core types of facial information: static, innate identity information, including an individual's skeletal structure, facial features, skin texture, and unique facial features; and dynamic, rapidly changing facial expression information, encompassing emotional expression elements such as raised corners of the mouth, furrowed brows, and squinting eyes caused by muscle movements. This output is called the hybrid feature vector, a composite representation carrying all identity and expression signals. Subsequently, this hybrid feature vector is input into a neutral encoder for deep processing. This neutral encoder achieves "information decoupling" at the feature level. Internally, it learns to construct a decoupled latent space by introducing mechanisms such as domain adaptive adversarial training, feature projection, or selective inhibition. Within this space, identity and expression information are mapped to mutually orthogonal subspaces. The neutral encoder accurately identifies and separates expression-related components from the hybrid features and actively removes them through "feature subtraction" or masking operations. Finally, a clean, standardized facial identity feature vector is output, which retains unique identity information to the greatest extent possible while containing almost no trace of emotional expression. This lays a highly consistent identity foundation for subsequent audio-driven emotion generation, unaffected by input facial expressions. It should also be noted that the neutral encoder's optimization objective is to minimize the expression domain loss to ensure that only identity is preserved. Understandably, by implementing step S120, emotion-independent identity representation is achieved, eliminating the influence of input image emotion and reducing computational costs.

[0019] Furthermore, a regression neural network supervised by 2D continuous emotion labels is used to decode the audio emotion vector, obtaining fine-grained facial expression features and emotion intensity values. This regression neural network takes the temporal emotion vector extracted by the audio encoder as input and learns a nonlinear mapping from the audio feature space to the emotion semantic space through joint modeling using a multilayer perceptron and an attention mechanism. Specifically, facial expression features are constrained by a classification loss function (such as cross-entropy loss) to accurately correspond to discrete emotion categories such as happiness, sadness, and anger; emotion intensity values ​​are supervised by a regression loss function (such as mean squared error loss) to quantify the intensity of emotion in a continuous scalar form. This dual-branch design not only ensures the semantic interpretability of emotion attributes but also achieves a smooth transition and natural variation in emotion expression through dynamic adjustment of the intensity dimension, ultimately forming an emotion representation that combines category accuracy and intensity subtlety. It should be noted that, to adapt to a specific target identity, a small sample of video data for that target identity is collected; a low-rank adapter technique is used to fine-tune the regression neural network in the emotional facial representation model to capture and learn the detailed features of the target identity while reducing the number of training parameters.

[0020] S130. Based on the facial identity features, the expression features, and the emotion intensity value, a speaking head video segment is generated using a pre-trained generative model.

[0021] Specifically, such as Figure 3As shown, step S130 includes the following steps: S131-S133: S131, fusing the facial identity features, expression features, and emotion intensity values ​​to obtain a conditional feature vector; S132, inputting the conditional vector into the generation model so that the generation model generates facial image frames under the control of the conditional feature vector; S133, combining the facial image frames in sequence to generate the speaking head video segment. It should be noted that the facial identity features, expression features, and emotion intensity values ​​are fused using multimodal methods to construct the conditional feature vector driving video generation. Specifically, firstly, the three are aligned and dimensionality reduced through a fully connected layer or attention mechanism to ensure that the features are in the same semantic space. Subsequently, a feature concatenation or weighted fusion strategy is used for integration: the emotion intensity value serves as a dynamic weight, adjusting the salience of the expression features in real time to achieve continuous control of the emotion intensity; stable identity features serve as the main constraint for the generated content. The final fused conditional feature vector is a compact high-dimensional representation that integrates target identity, specified expression, and precise intensity information. The conditional feature vector is input into a pre-trained generative model (e.g., a diffusion model). During the iterative denoising and generation process or adversarial generation, the conditional feature vector is deeply involved in each layer of the model's computation through cross-attention or feature modulation mechanisms, precisely guiding the denoising direction or the generator's feature synthesis. This ensures that the final frame-by-frame generated facial images maintain a high degree of consistency with the original image in terms of identity, strictly match the audio emotion vector in terms of expression and intensity, and are completely synchronized with the driving audio in terms of lip movements. Finally, all facial image frames generated in time step sequence are sent to the post-processing module. Through optical flow estimation, frame interpolation, or a dedicated temporal consistency network, minor temporal jitter and flicker are smoothed, enhancing the coherence between adjacent frames. Finally, the processed sequence is synthesized into a high-fidelity, smooth-moving, naturally expressive, and precisely synchronized speaking head video clip, completing the end-to-end synthesis from static identity and dynamic audio to dynamic video. It should also be noted that in this embodiment, the neutral encoder, the regressive neural network, and the generative model are all trained based on the construction of a large-scale audio-action dataset during the training phase. The construction process of the large-scale audio-action dataset is as follows: the dataset consists of triplet samples composed of video, audio, and action annotations, which are usually derived from publicly available multimodal corpora. The video content covers diverse personal identities, shooting postures, and rich emotional expressions to ensure that the model has good generalization ability.In the preprocessing stage, the video is divided into segments of fixed length (e.g., each segment has N=34 frames), and two types of motion parameters are extracted from each frame: one is the body joint rotation parameters, represented in the form of a rotation matrix; the other is the facial motion parameters, including the pose of the jaw joint and the facial blending shape weights, which are used to accurately describe changes in facial expressions. These structured annotations provide reliable supervision information for the model to learn identity-independent emotional representations.

[0022] S140. Generate an emotional speaking video using a time extrapolation strategy based on the speaking video segment.

[0023] Specifically, after obtaining the spoken video segment, as follows: Figure 4 As shown, step S140 includes the following steps: S141-S144: S141, setting the initial segment length to a first preset frame number and setting the sliding generation length to a second preset frame number, wherein the second preset frame number is less than the first preset frame number; S142, generating a video segment of the first preset frame number based on the speaking head video segment; S143, using the frame of the first preset frame number minus the second preset frame number as context, iteratively generating a new video segment of the next second preset frame number; S144, smoothly splicing the video segment and the new video segment with the overlapping part of the first preset frame number minus the second preset frame number to obtain the emotional speaking head video. It should be noted that the initial segment length is set to N frames and the sliding generation length is N' frames, where N' < N; firstly, an N-frame video segment is generated; then, using the last N-N' frames as context, iteratively generating a new segment of the next N' frames; the new and old segments are smoothly spliced ​​with the overlapping part of the N-N' frames to obtain the emotional speaking head video, to ensure the continuity and consistency of the entire generated video in the time dimension. It should also be noted that the time extrapolation strategy supports the generation of videos of arbitrary length, breaking through the length limitations of traditional methods and making it suitable for long narrative scenarios. Please see Figure 5 , Figure 5 This is a flowchart illustrating a method for generating emotional speaking head videos according to another embodiment of the present invention, as shown below. Figure 5 As shown, in this embodiment, the method includes steps S110-S160. That is, in this embodiment, after step S140 in the above embodiment, the method further includes steps S150-S160.

[0024] S150. Based on the identity coordinates corresponding to the facial identity features, automatically generate a time mask; S160. Using the time mask, the lip region and / or eye region in the emotional speaking head video are regenerated or modified for a specified time period to obtain the target emotional speaking head video.

[0025] Specifically, based on the standardized identity coordinates corresponding to the facial identity features, an accurate temporal mask is automatically generated. This process first decodes the identity feature vector into geometric coordinates aligned with the image space using a predefined face mesh model and projection transformation, thereby accurately locating the pixel positions of key areas such as the lips and eyes. Subsequently, combined with audio-driven temporal activity detection (such as lip movement intervals based on phoneme activity or blink intervals based on emotional intensity), the starting and ending frames to be edited, along with their intensity weights, are dynamically determined, ultimately generating a binary or soft temporal mask that contains both spatial region information and temporal range information. This temporal mask is then used to finely regenerate or modify the initially generated emotional speaking head video for specified regions and time periods. Specifically, the original video frames, the temporal mask, and enhanced control conditions (such as corrected lip movements or subtle expression adjustments) are input into a pre-trained generative model, guiding the model to reconstruct content only for the spatiotemporal regions marked by the mask (such as the lip region within a certain time period), while keeping the rest completely unchanged. This selective generation mechanism ensures the visual temporal continuity and spatial consistency of the edited video, ultimately outputting a target emotional speaking head video with higher lip-sync or more accurate facial expression details, significantly improving the accuracy and controllability of the generated content.

[0026] To facilitate understanding of the emotional speaking head video generation method in this invention, the following example is provided: The user provides a driving audio clip (someone smiling and saying "Today is a good day") and an identity image (a neutral expression photo of the same person). A pre-trained audio encoder analyzes the rhythm and intonation of the audio, extracting its joyful emotional features and outputting an audio emotion vector. An emotional facial representation model processes the identity image and the audio emotion vector: a neutral encoder analyzes the identity image, removing any potential facial expression traces and outputting clean facial identity features (ensuring the correct facial identity in the resulting video); a regression network decodes the audio emotion vector, predicting the facial features corresponding to the "joy" category, as well as a high emotion intensity value (e.g., 0.8, representing intense joy); the facial identity features, the "joy" expression features, and the intensity value of 0.8 are fused and used as control conditions input to a pre-trained generative model (e.g., a diffusion model). Based on these conditions, the generative model generates a 4-second (128-frame) video clip starting from noise, where the person's lip movements are synchronized with the audio and exhibit a very obvious large smile; using a temporal extrapolation strategy, the last 64 frames of the above clip are used as context to generate the next 4-second clip, which is then smoothly spliced ​​together. This process is repeated to finally synthesize a 1-minute emotional speaking head video. The characters in the video maintain their correct identities throughout, and, in accordance with the natural changes in the emotional intensity of the audio, display a smooth and realistic expression of joy, transitioning from a smile to a hearty laugh.

[0027] The emotional speaking head video generation method of this invention first extracts the audio emotion vector from the driving audio and decouples it into facial expression features and continuous intensity values; then, it separates the identity features to eliminate the emotional bias of the identity image; finally, it generates the emotional speaking head video based on the speaking head video segment using a time extrapolation strategy. This not only ensures the purity and accuracy of the emotional driving source, but also synchronizes the generated facial expressions with the audio rhythm, thus improving the accuracy and realism of the emotional expression in the generated emotional speaking head video.

[0028] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0029] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.

[0030] In one embodiment, an emotional speaking head video generation device 200 is provided, which corresponds one-to-one with the emotional speaking head video generation method in the above embodiments. For example... Figure 6 As shown, the emotional speaking head video generation device includes an acquisition and encoding unit 201, a processing unit 202, a first generation unit 203, and a second generation unit 204. Detailed descriptions of each functional module are as follows: The acquisition encoding unit 201 is used to acquire the driving audio and the identity image, and to encode the driving audio using a pre-trained audio encoder to obtain an audio emotion vector. Processing unit 202 is used to process the identity image and the audio emotion vector based on the emotion facial representation model to obtain facial identity features, expression features and emotion intensity values; The first generation unit 203 is used to generate a speaking head video segment based on the facial identity features, the expression features, and the emotion intensity value through a pre-trained generation model. The second generation unit 204 is used to generate an emotional speaking video based on the speaking video segment using a time extrapolation strategy.

[0031] In one embodiment, the encoding unit 201 is specifically used for: The processed driving audio is input into the audio encoder so that the feature extraction network in the audio encoder can extract semantic feature vectors. The semantic feature vector is mapped to the preset emotional feature space of the audio encoder through a projection layer or adapter network to output the audio emotional vector.

[0032] In one embodiment, the processing unit 202 is specifically used for: The neutral encoder removes the emotional expression components from the identity image to obtain the facial identity features that are unrelated to emotion. The audio emotion vector is decoded using a regression neural network supervised by 2D continuous emotion labels to obtain the facial expression features and the emotion intensity value.

[0033] In one embodiment, the processing unit 202 is further configured to: The identity image is input into an image encoder to obtain a hybrid feature vector that integrates identity information and facial expression information; The hybrid feature vector is input into the neutral encoder for processing to separate the identity information and facial expression information in the hybrid feature vector, and the facial expression information is removed to obtain the facial identity feature.

[0034] In one embodiment, the first generating unit 203 is specifically used for: The facial identity features, the expression features, and the emotion intensity value are fused to obtain a conditional feature vector; The conditional vector is input into the generative model so that the generative model generates facial image frames under the control of the conditional feature vector. The facial image frames are combined sequentially to generate the speaking head video segment.

[0035] In one embodiment, the second generation unit 204 is specifically used for: The initial segment length is set to a first preset number of frames, and the sliding generation length is set to a second preset number of frames, wherein the second preset number of frames is less than the first preset number of frames; Generate a video segment with the first preset number of frames based on the spoken video segment; Using the last frame in the video segment (the first preset frame number minus the second preset frame number) as the context, a new video segment with the next second preset frame number is iteratively generated. The emotional speaking video is obtained by smoothly splicing the overlapping portion of the video clip and the new video clip by subtracting the second preset frame number from the first preset frame number.

[0036] In one embodiment, the emotional speaking head video generation device 200 further includes: The third generation unit is used to automatically generate a time mask based on the identity coordinates corresponding to the facial identity features; The fourth generation unit is used to regenerate or modify the lip region and / or eye region in the emotional speaking head video for a specified time period using the time mask to obtain the target emotional speaking head video.

[0037] The emotional speaking head video generation device of the present invention first extracts the audio emotion vector from the driving audio and decouples it into facial expression features and continuous intensity values; then separates the identity features to eliminate the emotional bias of the identity image; finally, it generates emotional speaking head videos based on speaking head video segments using a time extrapolation strategy. This not only ensures the purity and accuracy of the emotional driving source, but also synchronizes the generated facial expressions with the audio rhythm, thus improving the accuracy and realism of the emotional expression in the generated emotional speaking head videos.

[0038] Specific limitations regarding the emotional speaking head-video generation device can be found in the limitations of the emotional speaking head-video generation method described above, and will not be repeated here. Each unit in the aforementioned emotional speaking head-video generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These units can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0039] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side method for generating emotional speaking head videos.

[0040] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the client-side functions or steps of an emotion-based speech head video generation method.

[0041] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described emotional speaking head video generation method.

[0042] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described emotional speaking head video generation method.

[0043] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0044] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0045] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0046] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for generating emotional speaking head videos, characterized in that, include: Acquire driving audio and identity image, and use a pre-trained audio encoder to encode the driving audio to obtain an audio emotion vector; The identity image and the audio emotion vector are processed based on the emotion facial representation model to obtain facial identity features, expression features, and emotion intensity values; Based on the facial identity features, the expression features, and the emotion intensity value, a pre-trained generative model generates speaking head video clips. An emotional speaking video is generated based on the speaking video clip using a time extrapolation strategy.

2. The method for generating emotional speaking head videos as described in claim 1, characterized in that, The step of encoding the driving audio using a pre-trained audio encoder to obtain an audio emotion vector includes: The processed driving audio is input into the audio encoder so that the feature extraction network in the audio encoder can extract semantic feature vectors. The semantic feature vector is mapped to the preset emotional feature space of the audio encoder through a projection layer or adapter network to output the audio emotional vector.

3. The method for generating emotional speaking head videos as described in claim 1, characterized in that, The emotional facial representation model includes a neutral encoder and a regressive neural network. The step of processing the identity image and the audio emotion vector based on the emotional facial representation model to obtain facial identity features, expression features, and emotion intensity values ​​includes: The neutral encoder removes the emotional expression components from the identity image to obtain the facial identity features that are unrelated to emotion. The audio emotion vector is decoded using a regression neural network supervised by 2D continuous emotion labels to obtain the facial expression features and the emotion intensity value.

4. The method for generating emotional speaking head videos as described in claim 3, characterized in that, The step of removing emotional expression components from the identity image using the neutral encoder to obtain facial identity features unrelated to emotion includes: The identity image is input into an image encoder to obtain a hybrid feature vector that integrates identity information and facial expression information; The hybrid feature vector is input into the neutral encoder for processing to separate the identity information and facial expression information in the hybrid feature vector, and the facial expression information is removed to obtain the facial identity feature.

5. The method for generating emotional speaking head videos as described in claim 1, characterized in that, The step of generating a speaking head video segment using a pre-trained generative model based on the facial identity features, the expression features, and the emotion intensity value includes: The facial identity features, the expression features, and the emotion intensity value are fused to obtain a conditional feature vector; The conditional vector is input into the generative model so that the generative model generates facial image frames under the control of the conditional feature vector. The facial image frames are combined sequentially to generate the speaking head video segment.

6. The method for generating emotional speaking head videos as described in claim 1, characterized in that, The step of generating an emotional speaking video based on the speaking video segment using a time extrapolation strategy includes: The initial segment length is set to a first preset number of frames, and the sliding generation length is set to a second preset number of frames, wherein the second preset number of frames is less than the first preset number of frames; Generate a video segment with the first preset number of frames based on the spoken video segment; Using the last frame in the video segment (the first preset frame number minus the second preset frame number) as the context, a new video segment with the next second preset frame number is iteratively generated. The emotional speaking video is obtained by smoothly splicing the overlapping portion of the video clip and the new video clip by subtracting the second preset frame number from the first preset frame number.

7. The method for generating emotional speaking head videos as described in any one of claims 1-6, characterized in that, After the step of generating an emotional speaking video based on the speaking video segment using a time extrapolation strategy, the method further includes: Based on the identity coordinates corresponding to the facial identity features, a time mask is automatically generated; The target emotional speaking head video is obtained by regenerating or modifying the lip and / or eye regions in the emotional speaking head video for a specified time period using the time mask.

8. An emotional speaking head video generation device, characterized in that, include: An encoding unit is used to acquire driving audio and identity image, and to encode the driving audio using a pre-trained audio encoder to obtain an audio emotion vector; The processing unit is used to process the identity image and the audio emotion vector based on the emotional facial representation model to obtain facial identity features, expression features and emotion intensity values; The first generation unit is used to generate a speaking head video segment based on the facial identity features, the expression features, and the emotion intensity value using a pre-trained generation model. The second generation unit is used to generate an emotional speaking video based on the speaking video segment using a time extrapolation strategy.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the emotional speaking head video generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the emotional speaking head video generation method as described in any one of claims 1 to 7.