Method for realizing synchronization of character expression and lip shape in video through emotion in sound and cloned digital human system
By extracting multi-dimensional audio features and applying expressions and lip shape generation models, we generate expressions and lip shape movements that are synchronized with sound emotions, solving the problem of difficult to achieve highly realistic expression synchronization in the existing technology, and achieving high accuracy and natural cloned digital human performance.
Patent Information
- Application Number
- CN202510150775.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The prior art is difficult to fully capture and reproduce the subtle expression differences corresponding to emotional changes in sounds, and cannot achieve highly realistic expression synchronization effects.
By collecting audio information, multi-dimensional features of sound are extracted, and a pre-trained expression and lip generation model is used to generate expression parameters and lip parameters based on sound characteristics, and fusion processing is performed to generate continuous animation sequences. Finally, video information with expressions synchronized with lip and sound emotions is generated through rendering.
It realizes highly accurate expressions and lip shape synchronization, can deeply analyze sound characteristics and accurately map into cloned digital human facial expressions and lip movements, improving the natural fidelity and personalized expression capabilities of cloned digital humans.
Smart Images

Figure CN120163906A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of audio - video processing in artificial intelligence, and particularly relates to a method for synchronizing the facial expressions and lip shapes of characters in a video through emotions in sound and a cloned digital human system. Background Art
[0002] In the field of artificial intelligence, the facial processing of people in videos is mainly achieved in the following three ways:
[0003] First, the manual processing method of traditional video editing software. In common video editing software (such as Adobe Premiere Pro, Final Cut Pro, etc.), when processing real - person videos, if one wants to synchronize the facial expressions of characters with the emotions in the sound, editors mainly rely on manual operations. First, they will repeatedly listen to the audio part in the video and judge the expected emotional changes of the characters according to the characteristics of the sound, such as intonation, speech rate, volume, etc. For example, in a dialogue video, when hearing the sound pitch rise and the speech rate accelerate, the editor will find the corresponding frame of the character's picture on the video timeline, and then manually adjust the playback speed of the video, add key frames, and use special effect tools in the software (such as adjusting the deformation effect of facial features) to simulate the excited or agitated expressions that the character may show. For some complex scenarios, it may also be necessary for the editor to clip and splice video segments to ensure that the expression changes match the rhythm of the sound. For instance, at a plot turning point where the emotional change in the sound is significant, the editor may select a suitable facial expression picture from other segments, crop and adjust it, and then insert it into the current position to try to achieve the synchronization of expressions and sounds.
[0004] Second, the expression adjustment technology based on simple template matching. Some existing technologies attempt to use simple template matching methods to handle the expression synchronization problem in real - person videos. First, an expression template library is established, which contains some common expression templates (such as basic expressions like happy, sad, angry, etc.) and their corresponding sound feature templates (such as specific pitch ranges, speech rate intervals, etc.). When processing a video, the system analyzes the audio in the video, extracts the sound features (such as spectral features, prosodic features, etc.), and then matches them with the sound feature templates in the expression template library. Once a matching template is found, the corresponding expression template is applied to the character's picture in the video. For example, when it is detected that the sound has a high pitch and a fast speech rate, which matches the sound feature template of the happy expression, the system will add some preset happy expression special effects (such as the corners of the mouth turning up, the eyes widening, etc.) to the face of the character in the video, and adjust the parameters of the special effects to adapt to the character image in the video.
[0005] III. Preliminary Artificial Intelligence-Assisted Facial Expression Analysis and Adjustment Technology. Some existing research or technology products use artificial intelligence algorithms to analyze and adjust facial expressions and voices in real-person videos. For example, through deep learning algorithms, a model is trained to recognize the facial expressions and voice emotions of people in videos. The input of the model is a sequence of video frames (including facial images of people) and audio segments, and the output is the facial expression classification result (such as various emotion categories) and voice emotion features (such as emotion intensity, etc.). In practical applications, the system first preprocesses the input real-person video, extracts video frames and audio data, and then inputs them into the trained model. The model analyzes the facial expressions and voices according to the learned patterns. If it detects that the facial expression does not match the voice emotion, it tries to make adjustments. A common adjustment method is to change the facial expression by modifying the key feature points of the person's face in the video frame. For example, using technologies such as Generative Adversarial Networks (GANs), new coordinates of facial key feature points are generated according to the voice emotion features, and then these coordinates are applied to the face of the person in the original video frame, thereby changing the facial expression of the person to make it more consistent with the voice emotion.
[0006] The above three methods have the following problems respectively:
[0007] (I) Problems existing in the manual processing method of traditional video editing software
[0008] Highly dependent on manual experience and operation skills: This manual processing method requires editors to have rich video editing experience and sharp observation skills. The processing results of different editors may vary greatly, making it difficult to ensure consistency and accuracy.
[0009] Low efficiency and time-consuming: For long videos or frequent requirements for synchronizing facial expressions and voices, manual operations will consume a large amount of time and energy, greatly reducing the efficiency of video production.
[0010] Difficulty in ensuring accuracy: Due to human subjective judgment and errors in manual operations, it is difficult to achieve precise synchronization of facial expressions and voice emotions, and problems such as premature or late facial expression changes and mismatches with voice emotion intensity are likely to occur.
[0011] (II) Problems existing in the facial expression adjustment technology based on simple template matching
[0012] Single and unrealistic facial expression templates: The used facial expression templates are usually relatively simple and fixed, and cannot accurately reflect the real facial expression changes of real people under various complex emotions, resulting in the generated facial expressions looking rigid and unnatural.
[0013] Limited adaptability: It can only handle voice emotion situations that match the existing templates in the template library. For some special voice emotion changes not included in the template library, it cannot make effective facial expression adjustments, and the flexibility is poor.
[0014] Difficulty in achieving personalization: There may be individual differences when different people express the same emotion. This method cannot generate personalized facial expressions according to the characteristics of different people and cannot meet diverse needs.
[0015] (III) Problems existing in the preliminary artificial intelligence-assisted facial expression analysis and adjustment technology
[0016] High difficulty in model training: A large amount of labeled data (including real-person video data with accurate facial expression and voice emotion annotations) is required to train the model. The data collection and annotation work is costly and time-consuming. Moreover, the training effect of the model is easily affected by the quality and quantity of the data. If the data is insufficient or biased, the accuracy and generalization ability of the model will be limited.
[0017] Insufficient fineness of facial expression synchronization: Although some correlation analysis and adjustment of facial expressions and voice emotions can be carried out, there are still deficiencies in the fineness of facial expression changes. For example, for the subtle changes in emotion in the voice, it may not be accurately reflected in the facial expression, resulting in inaccurate and unnatural synchronization of facial expressions and voice emotions. Especially in the processing of the real-person version of facial expressions of cloned digital humans, due to the richness and complexity of real-person facial expressions, existing technologies are difficult to fully capture and reproduce the subtle facial expression differences corresponding to the emotional changes in the voice, and cannot achieve a highly realistic facial expression synchronization effect. Summary of the Invention
[0018] The technical problem to be solved by the present invention is to provide a method for synchronizing the facial expressions and lip shapes of a person in a video through the emotion in the voice and a cloned digital human system, which solves the problem that it is difficult to fully capture and reproduce the subtle facial expression differences corresponding to the emotional changes in the voice in the prior art and cannot achieve a highly realistic facial expression synchronization effect.
[0019] The present invention adopts the following technical solutions to solve the above technical problems:
[0020] A method for synchronizing the facial expressions and lip shapes of a person in a video through the emotion in the voice, collecting audio information and extracting multi-dimensional features of the voice; applying a pre-trained facial expression and lip shape generation model to generate corresponding facial expression parameters and lip shape parameters according to the multi-dimensional features of the voice; performing fusion processing on the facial expression parameters and lip shape parameters, and generating a continuous animation sequence according to the fused parameters; rendering the continuous animation sequence to generate video information with synchronized facial expressions, lip shapes and voice emotions.
[0021] After collecting the audio information, first, perform noise reduction processing on the audio information, and then extract multi-dimensional features from the noise-reduced audio information.
[0022] The multi-dimensional features include pitch features, timbre features, speech rate features, prosody features, and emotional features in the voice.
[0023] The expression and lip shape generation model includes an expression generation sub-module, a lip shape generation sub-module, and a model training module. Among them, the expression generation sub-module is constructed based on a generative adversarial network, and its internal generator part takes the voice emotion features as input and the facial expression parameters as output; the lip shape generation sub-module is constructed using a variational autoencoder, takes the speech features in the voice as input, and the lip movement parameters as output; the model training module uses a supervised learning algorithm, takes the voice features in the preprocessed video data as input, and the expression parameters and lip shape parameters as output.
[0024] The model training module repeatedly trains the models in the expression generation sub-module and the lip shape generation sub-module. During the training process, by continuously adjusting the weight parameters of the models, the models generate appropriate expressions and lip movements according to the input voice features. There is a connection relationship for parameter update between the model training module and the corresponding model structures in the expression generation sub-module and the lip shape generation sub-module, and finally the trained parameters are updated to the corresponding models in a timely manner.
[0025] The cloned digital human system includes a cloned digital human three-dimensional model and its control system. The control system applies the method to obtain video information with synchronized expressions, lip shapes, and voice emotions, and sends it to the cloned digital human three-dimensional model for synchronous output.
[0026] The control system includes an audio processing module, an expression and lip shape generation module, a model training module, and a fusion and rendering module; among them,
[0027] The audio processing module includes an audio acquisition unit, a noise reduction unit, and a feature extraction unit connected in sequence. The audio acquisition unit receives an externally input audio signal, and after noise reduction by the noise reduction unit, it is transmitted to the feature extraction unit for multi-dimensional feature extraction of the sound.
[0028] The expression and lip shape generation module is constructed based on a deep learning generation model, and includes an expression generation sub-module and a lip shape generation sub-module. The expression generation sub-module generates the facial expression parameters of the cloned digital human, and the lip shape generation sub-module generates the corresponding lip movement parameters.
[0029] The model training module is used to train the deep learning generation model, and includes a data collection unit, a data preprocessing unit, and a model training unit. The data collection unit collects real-person video data, and annotates the expressions, lip shapes, and voices in the video. The data preprocessing unit cleans and normalizes the annotated data. The model training unit uses a supervised learning algorithm, takes the voice features in the preprocessed video data as input, and the expression and lip shape parameters as output, and repeatedly trains the deep learning generation model in the expression and lip shape generation module until the optimal parameters of the model are obtained.
[0030] The fusion and rendering module includes a parameter fusion unit, an animation generation unit and a rendering unit which are connected in sequence. The parameter fusion unit fuses the facial expression parameters and the lip shape movement parameters so that the expression and lip shape movement are coordinated with the facial movement of the cloned digital human. The animation generation unit generates a continuous animation sequence for the face of the cloned digital human according to the fused parameters, renders the cloned digital human model with expression and lip shape animation, and generates a video output with expression, lip shape and sound emotion synchronization.
[0031] The feature extraction unit uses a deep learning algorithm to analyze the noise-reduced audio signal and extracts multiple features of the sound including pitch, timbre, speaking speed, rhythm, volume and emotion.
[0032] The expression generation submodule generates facial expression parameters of the cloned digital human, including muscle movements of the eyes, eyebrows, and cheeks, based on the sound and emotional features extracted by the audio processing module. The lip shape generation submodule generates corresponding lip shape movement parameters, including the degree of lip opening and closing, the movement position of the corners of the mouth, and the coordination relationship between the lips and teeth, based on the speech content, speech speed, and rhythmic features in the audio.
[0033] The video data collected by the data collection unit includes character expressions, lip movements and corresponding sound signals, and the annotated content includes expression types, corresponding relationships between lip movements and voice, and emotional features in the sound.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] 1. Achieve highly accurate expression and lip shape synchronization. On the one hand, the accurate mapping of sound features to expression and lip shape changes can deeply analyze various features in the sound, including but not limited to pitch, timbre, speaking speed, rhythm, volume, etc., and accurately map these sound features to subtle changes in the facial expressions of the cloned digital person (such as the frequency of eye blinking, the degree of eyebrow fluctuation, cheek muscle movement, etc.) and precise lip shape movements (including the opening and closing amplitude of the lips, the upward or downward angle of the corners of the mouth, the coordination of the lips and teeth, etc.). For example, when the pitch in the voice rises and the speaking speed increases, the model can accurately drive the face of the cloned digital person to show an excited expression, with eyes wide open, eyebrows raised, lips opening and closing quickly, and the corners of the mouth raised, achieving a high degree of synchronization between expression and lip shape changes and the emotions of the sound, avoiding delays, dislocations or unnatural movements.
[0036] In view of the pronunciation characteristics of different languages, especially in scenarios involving interactions among multiple languages, ensure that the lip movements of the cloned digital human can accurately correspond to the phonetic pronunciations of different languages. Whether it is the plosives and fricatives in English or the tone changes in Chinese, through precise analysis of the voice, the lip movements and expressions of the cloned digital human can naturally match the language expression, just as smooth and natural as when a real person is speaking.
[0037] On the other hand, adapt to diverse voice input scenarios. Whether it is professionally recorded clear audio or audio in actual scenarios with certain environmental noise or slightly degraded sound quality (such as on-site interviews, remote meetings, etc.), the method of the present invention can effectively extract voice features and accurately generate corresponding expressions and lip movements. For example, for audio recorded in a noisy outdoor environment, the system can, through advanced signal processing technologies and artificial intelligence algorithms, filter out noise interference, accurately identify the voice emotions and speech content of the speaker, and thus generate appropriate expression and lip synchronization actions for the cloned digital human, ensuring its performance stability and naturalness in various complex voice environments.
[0038] 2. Improve the natural realism of the cloned digital human. On the one hand, capture subtle expression changes and enhance emotional expressiveness. Pay attention to the simulation of subtle expression changes in humans, enabling the cloned digital human to show emotional levels and nuances similar to those of real people. In addition to common basic expressions (such as happy, sad, angry, surprised, etc.), it can also capture expression changes corresponding to subtle emotions such as hesitation, embarrassment, shyness, etc., and present them through precise facial muscle movement control and lip shape adjustment. For example, when there is a hint of hesitation in the voice, the eyes of the cloned digital human may slightly narrow, the eyebrows may slightly wrinkle, and the lips may pause and slightly purse. These subtle expression changes will greatly enhance the emotional expressiveness of the cloned digital human, making it closer to the emotional expression mode of real people.
[0039] Simulate the dynamic change rules of expressions during real human communication, such as the natural and smooth transition of expressions, and avoid abrupt expression conversions. When the emotion in the voice gradually changes from calm to excited, the expression of the cloned digital human can follow the natural emotional change pattern of humans and gradually transition from the normal state to the excited state. The movement of facial muscles and the change of lip shape present a continuous and smooth process, rather than a rigid switch, making it difficult for the audience to perceive that it is a digital human and as if they are communicating with a real person.
[0040] On the other hand, personalized expression generation fits different character styles. According to different cloned digital human character settings or the real-life prototypes being imitated, expressions and lip movements with personalized characteristics are generated. For example, for a cloned digital human set to be cheerful and lively, under the same voice emotion, its expressions and lip movements may be more exaggerated and vivid than those of a cloned digital human set to be calm and introverted. By learning and analyzing the characteristic styles of different characters, the cloned digital human has unique personality in terms of expression and lip performance, better meeting the needs of various application scenarios. Whether it is role-playing in entertainment shows or image building in enterprise customer service scenarios, it can interact with users in the most natural and character-appropriate way.
[0041] 3. Expanding the application scope and value of cloned digital humans is mainly manifested in the following aspects:
[0042] 1) Optimizing the application experience in the entertainment industry
[0043] In the entertainment fields such as film and television production, virtual live streaming, and games, cloned digital humans can be more realistically integrated into the plot or interactive scenarios. In film and television production, cloned digital human actors can perform with more natural expressions and lip synchronization, achieving seamless docking with real actors and bringing more shocking visual and emotional experiences to the audience. In virtual live streaming, the cloned digital human image of the anchor can real-time and accurately display rich and diverse expressions and lip movements according to the voice emotion and speech content, enhancing the interactivity and immersion with the audience, attracting more audience attention and participation in the interaction, and improving the quality and influence of the live stream. In games, if non-player characters (NPCs) are played by cloned digital humans, their realistic expressions and lip synchronization will make the game world more vivid and interesting, enabling players to better understand the emotions and intentions of NPCs, thus enhancing the playability and immersion of the game.
[0044] 2) Improving the effects of education and training
[0045] In the fields of education and training, cloned digital humans can serve as virtual teachers or training instructors, and better convey knowledge and emotions through natural expression and lip synchronization. For example, in language teaching, cloned digital human teachers can vividly display the mouth shape changes and corresponding expressions during pronunciation according to the emotional color and intonation of the teaching content, helping students more intuitively understand the pronunciation points and the emotional connotations behind the language, and improving learning efficiency. In vocational skills training, cloned digital humans can simulate the conversations and expressions of people in various work scenarios, enabling trainees to more vividly experience the communication situations in actual work, and enhancing the practicality and effectiveness of the training.
[0046] 3) Enhancing the quality of commercial services and communication
[0047] In the business field, such as customer service, product display, etc., cloned digital humans can communicate with customers in a more humanized way. The customer service cloned digital human can adjust its expression and tone in a timely manner according to the customer's voice emotion. When the customer expresses dissatisfaction, the cloned digital human can respond to the customer with appropriate expressions (such as concerned eyes, slightly frowning) and a sincere tone, making the customer feel understood and valued, and improving customer satisfaction. In product display, the cloned digital human can introduce product features and advantages through vivid expressions synchronized with lip movements, attracting the attention of consumers, better conveying product information, promoting sales conversion, and creating greater commercial value for enterprises. Description of the Drawings
[0048] Figure 1 It is a flowchart of the application method of the digital cloning human system of the present invention.
[0049] Figure 2 It is a flowchart of the specific method for image cloning of the digital cloning human system of the present invention.
[0050] Figure 3 It is a flowchart of the specific method for voice cloning of the digital cloning human system of the present invention.
[0051] Figure 4 It is a flowchart of the specific method for semantic emotion judgment and expression generation of the digital cloning human system of the present invention. Detailed Embodiments
[0052] The structure and working process of the present invention will be further described below with reference to the accompanying drawings.
[0053] The present invention aims to construct an audio-visual processing method and system based on artificial intelligence for realizing the precise synchronization of the expressions, lip movements and voice emotions of cloned digital humans in videos, making the cloned digital humans behave more naturally and realistically, just like real people. The system mainly consists of an audio processing module, an expression and lip movement generation module, a model training module, and a fusion and rendering module. Each module works together to complete the entire process from voice input to the synchronous output of the expressions and lip movements of the cloned digital human.
[0054] The specific solution includes a method for realizing the synchronization of the expressions and lip movements of the characters in the video through the emotions in the voice, and a cloned digital human system. Among them, the method for realizing the synchronization of the expressions and lip movements of the characters in the video through the emotions in the voice collects audio information and extracts multi-dimensional features of the voice; applies a pre-trained expression and lip movement generation model to generate corresponding expression parameters and lip movement parameters according to the multi-dimensional features of the voice; performs fusion processing on the expression parameters and lip movement parameters, and generates a continuous animation sequence according to the fused parameters; renders the continuous animation sequence to generate video information with synchronized expressions, lip movements and voice emotions.
[0055] The cloned digital human system includes a cloned digital human three-dimensional model and its control system. The control system applies the method to obtain video information with synchronized expressions, lip shapes, and voice emotions, and sends it to the cloned digital human three-dimensional model for synchronous output.
[0056] Specific Embodiment 1
[0057] A method for synchronizing the expressions and lip shapes of a person in a video through emotions in the voice includes the following steps:
[0058] Step 1: Collect audio information and extract multi-dimensional features of the voice; collect the target audio signal through an audio collection unit (which can be real-time recording or importing existing audio files). The collected audio signal enters a noise reduction unit for noise reduction processing, and then a feature extraction unit extracts the multi-dimensional features of the voice. The order of this step cannot be randomly interchanged because only after collecting the audio can subsequent noise reduction and feature extraction operations be carried out. If the order is reversed, the work cannot be carried out normally.
[0059] Step 2: Apply a pre-trained expression and lip shape generation model to generate corresponding expression parameters and lip shape parameters based on the multi-dimensional features of the voice;
[0060] The expression and lip shape generation model includes an expression generation sub-module, a lip shape generation sub-module, and a model training module.
[0061] Among them, the expression generation sub-module is constructed based on the Generative Adversarial Network (GAN). The internal generator part takes the voice emotion features as input and the facial expression parameters as output; the lip shape generation sub-module is constructed using the Variational Autoencoder (VAE), taking the speech features in the voice as input and the lip movement parameters as output. The work of these two sub-modules is carried out in parallel, jointly providing basic parameters for subsequent fusion. This step needs to be carried out after the first step of audio feature extraction and depends on the voice features extracted in the first step as input.
[0062] The model training module adopts a supervised learning algorithm, taking the voice features in the preprocessed video data as input and the expression parameters and lip shape parameters as output. The data collection unit in the model training module collects and organizes real-person video data in advance. After the data preprocessing unit processes it, the model training unit uses this data to train the GAN and VAE models in the expression and lip shape generation model. During the training process, by continuously adjusting the weight parameters of the model, the model can generate appropriate expressions and lip shape actions according to the input voice features. There is a connection relationship for parameter update between the model training module and the corresponding model structures in the expression generation sub-module and the lip shape generation sub-module. Finally, the trained parameters are updated to the corresponding models in a timely manner. This process can be carried out when the system is initially built or the model is updated. The trained model has the ability to generate expression and lip shape action parameters according to voice features. This step can be carried out in parallel with the first step in terms of time sequence. However, before actually using the expression and lip shape generation module to generate parameters, the model needs to complete the training work.
[0063] Step 3: Perform fusion processing on the expression parameters and lip shape parameters, and generate a continuous animation sequence according to the fused parameters; The parameter fusion unit in the fusion and rendering module receives the expression parameters and lip shape parameters transmitted from the expression and lip shape generation module, performs fusion processing, and then transmits the fused parameters to the animation generation unit. The animation generation unit generates a continuous animation sequence based on this, enabling the expression and lip shape actions to smoothly transition and change. This step needs to be carried out after the generation of the expression and lip shape parameters in the third step, and the order cannot be reversed because only with parameters can the fusion and animation generation operations be carried out.
[0064] Step 4: Render the continuous animation sequence to generate video information with synchronized expression, lip shape, and voice emotion.
[0065] The animation sequence data generated by the animation generation unit is transmitted to the rendering unit. The rendering unit uses ray tracing technology to render the animation sequence, and finally generates video information with synchronized expression, lip shape, and voice emotion.
[0066] Specific Embodiment 2
[0067] The cloned digital human system includes a cloned digital human three-dimensional model and its control system. The control system uses the above method to obtain video information with synchronized expression, lip shape, and voice emotion, and sends it to the cloned digital human three-dimensional model for synchronous output.
[0068] The control system includes an audio processing module, an expression and lip shape generation module, a model training module, and a fusion and rendering module; among them,
[0069] The audio processing module includes an audio acquisition unit, a noise reduction unit, and a feature extraction unit that are connected in sequence. The audio acquisition unit receives an externally input audio signal. After being denoised by the noise reduction unit, it is transmitted to the feature extraction unit for multi-dimensional feature extraction of the sound. The audio acquisition unit uses a professional audio capture card (which has the ability to capture high-fidelity sound signals and is compatible with a variety of audio input sources) to receive externally input audio signals. For example, it can be connected to a microphone to record real-time human voices or access stored audio files. The captured audio signal is transmitted to the noise reduction unit in real time. The noise reduction unit is designed based on the Adaptive Filtering Algorithm (a kind of algorithm that can dynamically adjust filtering parameters according to the characteristics of the input signal and noise, thereby effectively removing noise). Its function is implemented through software programming. It is connected to the audio acquisition unit through a data line, receives the original captured audio signal, and performs noise reduction processing on it. The processed audio signal is then transmitted to the feature extraction unit. The feature extraction unit is built using a Convolutional Neural Network (CNN, a deep learning network structure that is good at automatically extracting features from data). Its input port is connected to the output port of the noise reduction unit, receives the denoised audio signal, and performs in-depth feature extraction operations.
[0070] The functions and operation processes of this audio module are as follows:
[0071] First, after the audio acquisition unit is started, it obtains the audio signal in real time to ensure the accurate input of sound data. For example, when we want to capture the audio of a service conversation for a virtual customer service cloning digital human, the speaking voice of the customer service staff is accurately captured through a microphone. Then, the noise reduction unit filters out the possible environmental noises (such as the noisy voices in the office background, the humming sound generated by the operation of computer equipment, etc.) in the captured audio, improves the purity of the audio signal, and ensures that subsequent feature extraction is not interfered by noise. Finally, the feature extraction unit extracts multi-dimensional sound features from the denoised audio, including pitch features (obtained by analyzing the frequency changes of the audio signal. For example, a high pitch corresponds to higher frequency components), timbre features (distinguishing the voice characteristics of different people based on factors such as the harmonic structure of the sound), speech rate features (judging by calculating the number of speech syllables per unit time), prosody features (such as reflected by the cadence and stress distribution of speech), and emotional features in the sound (judging whether it is happy, sad or other emotions and their intensities by virtue of the classification and recognition ability of the deep learning model for sound emotions). These extracted features will be used as important bases for subsequent expression and lip shape generation and transmitted to the expression and lip shape generation module.
[0072] The expression and lip shape generation module is built based on the generative model of deep learning, including the expression generation submodule and the lip shape generation submodule. The expression generation submodule generates the facial expression parameters of the cloned digital human, and the lip shape generation submodule generates the corresponding lip shape movement parameters.
[0073] The expression generation submodule is built on the Generative Adversarial Network (GAN, which consists of a generator and a discriminator. It continuously optimizes the generation effect through adversarial training and can generate realistic data). Its internal generator part receives the sound emotion features as input and outputs the facial expression parameters of the cloned digital person. These parameters specifically cover the opening and closing angle of the eyes, the curvature of the eyebrows, the contraction and relaxation of the cheek muscles, etc., which are used to control the expression changes of the cloned digital person. The lip shape generation submodule is built using the Variational Autoencoder (VAE, a deep learning model that can learn the potential distribution of data and generate data). It inputs the speech features in the sound (such as phonemes, syllable information, speaking speed, rhythm, etc.) and outputs the lip shape movement parameters, such as the opening and closing amplitude of the lips, the moving coordinates of the mouth corners, the coordination state of the lips and teeth, etc., so as to accurately control the lip shape movement of the cloned digital person. The output ports of the two submodules are respectively connected to the parameter fusion unit in the fusion and rendering module to transmit the generated expression parameters and lip shape parameters.
[0074] The function and operation process of this module are as follows:
[0075] After receiving the sound features from the audio processing module, the generator of the expression generation submodule generates expression parameters based on the emotional features of the sound. For example, if the sound contains a happy emotion with a high intensity, the generator will output the corresponding expression parameters to make the cloned digital person's eyes open wide and crescent-shaped, eyebrows raised, cheek muscles slightly raised, showing an obvious happy expression. At the same time, the lip shape generation submodule generates accurate lip shape movement parameters based on the voice features to ensure that when the cloned digital person speaks the corresponding voice content, the lip movement and pronunciation are perfectly matched. For example, when pronouncing the phoneme "O", the lips will naturally open in a circle. The parameters generated by these two submodules together provide basic data support for the subsequent expression and lip shape synchronization presentation of the cloned digital person.
[0076] The model training module is used to train the generative model of deep learning, including a data collection unit, a data preprocessing unit, and a model training unit. The data collection unit collects real-person video data and annotates the expressions, lip shapes, and voices in the videos. The data preprocessing unit cleans and normalizes the annotated data. The model training unit uses a supervised learning algorithm, taking the voice features in the preprocessed video data as input and the expression and lip shape parameters as output, and repeatedly trains the generative model of deep learning in the expression and lip shape generation module until the optimal parameters of the model are obtained;
[0077] The data collection unit is responsible for collecting real-person video data from multiple channels. These data sources can be public film and television material libraries (such as [specific film and television material library name], which contains a large number of video resources of different scenes and different characters), professionally recorded expression and voice datasets (such as [specific dataset name, organized and released by relevant scientific research institutions or enterprises, with detailed and standardized annotations]), etc. The collected video data contains rich human expressions, lip movements, and corresponding voice signals, and the expressions, lip shapes, and voices in each video are accurately annotated. The annotation content is detailed to the expression type (such as specifically smiling, laughing, frowning, etc.), the correspondence between lip movements and voices (the change in lip shape corresponding to each syllable), the emotional characteristics in the voice (specific emotions and intensity values), etc. The data collection unit transmits the collected data to the data preprocessing unit. The data preprocessing unit uses technical means such as data normalization (unifying data in different ranges to a specific interval for convenient model training) and data cleaning (removing duplicate, incorrect, or incomplete data records) to process the collected data, ensuring the quality and consistency of the data. The processed data is then passed to the model training unit. The model training unit uses a supervised learning algorithm (such as the Back Propagation Algorithm, which adjusts the weight parameters of the neural network based on the reverse propagation of errors to optimize the model performance), taking the voice features in the preprocessed video data as input and the expression and lip shape parameters as output, and repeatedly trains the model in the expression and lip shape generation module (i.e., the GAN and VAE models mentioned above). During the training process, by continuously adjusting the weight parameters of the model, the model can accurately generate appropriate expressions and lip movements according to the input voice features. There is a connection relationship for parameter update between the model training unit and the corresponding model structure in the expression and lip shape generation module, ensuring that the trained parameters can be updated to the corresponding model in a timely manner.
[0078] The functions and operation processes of this module are as follows:
[0079] First, the data collection unit widely collects real-person video data to accumulate sufficient materials for model training. For example, it collects videos of people speaking in different age groups, genders, and language environments, covering various rich facial expressions and lip movement situations. Next, the data preprocessing unit normalizes the collected data, removing data that may interfere with the model training effect, making the data input into the model more accurate and reliable. Finally, the model training unit uses the supervised learning algorithm to make the model learn the mapping rules between voice features and facial expression and lip parameter according to the set input-output relationship. After multiple iterative trainings (such as setting the number of training rounds to [specific number of rounds, e.g., 100 rounds]), the prediction ability of the model is continuously optimized until the model reaches a satisfactory accuracy rate (such as an accuracy rate above [specific accuracy rate value, e.g., 95%]) on the validation data set (a part of the data divided from the collected data for validating the model effect), completing the model training process, and updating the trained model parameters to the corresponding model structure of the facial expression and lip generation module, enabling it to generate facial expressions and lip movements according to the actual input voice features.
[0080] The fusion and rendering module includes a parameter fusion unit, an animation generation unit, and a rendering unit connected in sequence. The parameter fusion unit fuses the facial expression parameters and lip movement parameters, making the facial expressions and lip movements coordinated with the facial movements of the cloned digital human. The animation generation unit generates a continuous animation sequence for the face of the cloned digital human according to the fused parameters, and renders the cloned digital human model with facial expressions and lip animations to generate a video output with synchronized facial expressions, lip movements, and voice emotions.
[0081] The parameter fusion unit receives the expression parameters and lip shape parameters from the expression and lip shape generation module, and fuses the two through a specific fusion algorithm (such as a weighted fusion algorithm based on time synchronization and action coordination, which assigns corresponding weights to the parameters of the two according to the corresponding relationship between the expression and lip shape actions in the time dimension and the overall visual coordination) to ensure that the expression and lip shape actions are coordinated on the face of the cloned digital human, avoiding the phenomenon of disconnection or incoordination between the expression and lip shape actions. The fused parameters are transmitted to the animation generation unit. The animation generation unit, based on the fused parameters, uses animation generation technologies such as keyframe interpolation (generating intermediate transition frames between the starting and ending keyframes to make the action changes more smooth and natural) to generate a continuous animation sequence for the face of the cloned digital human, enabling the expression and lip shape actions to transition and change smoothly. The generated animation sequence data is then passed to the rendering unit. The rendering unit adopts ray tracing technology (Ray Tracing Technology, a rendering technology that generates high-quality images by simulating physical processes such as light propagation, reflection, and refraction), combined with the three-dimensional model of the cloned digital human (previously created by a professional 3D modeling software, and the model structure includes details such as facial bones and skin materials, such as created using [specific 3D modeling software name and version]), to render the cloned digital human model with expression and lip shape animations, and finally generates a high-quality and visually realistic video output, presenting the image of the cloned digital human with synchronized expression, lip shape, and voice emotion.
[0082] The functions and operation processes of this module are as follows:
[0083] After receiving the expression parameters and lip shape parameters, the parameter fusion unit performs a detailed fusion operation to ensure that while the cloned digital human makes an expression, the lip shape action matches it perfectly, just like the natural synchronization of expression and mouth shape when a real person speaks. For example, when expressing a surprised expression and saying the syllable "ah" at the same time, the fused parameters can make the eyes of the cloned digital human widen and the mouth open in a suitable round shape, and the two are highly coordinated in terms of time and action amplitude. Then, the animation generation unit converts the fused parameters into a continuous animation sequence, giving vividness and smoothness to the facial actions of the cloned digital human, enabling the expression and lip shape actions to change naturally. Finally, the rendering unit uses advanced ray tracing technology to render the cloned digital human model with animations, making it more realistic in appearance, and the expression and lip shape of the cloned digital human in the output video can be precisely synchronized with the voice emotion, presenting a natural and real visual effect.
[0084] One of the core inventions of this solution is the ability to simultaneously integrate multiple features in the sound (such as emotional features, voice features, etc.) and accurately map them to the expressions and lip movements of the cloned digital person. Through the comprehensive analysis of multimodal features such as pitch, timbre, speech speed, rhythm, volume, and emotional intensity by the deep learning model, the expressions and lip movements are highly synchronized with the sound in terms of time and amplitude. For example, in a voice clip with strong emotions, the model can not only make the face of the cloned digital person show the corresponding expression based on the emotional characteristics, but also ensure that the lip movements are perfectly matched with the voice content based on the voice characteristics, and the changing rhythms of the two are completely synchronized, avoiding the problem of the expression and lip movements being out of sync or uncoordinated with the sound in the prior art.
[0085] For the deep learning models in the expression and lip shape generation module, in addition to the generative adversarial network (GAN) and variational autoencoder (VAE) mentioned above, you can also consider using recurrent neural network (RNN) and its variants (such as long short-term memory network LSTM, gated recurrent unit GRU, etc.). The RNN series of models has advantages in processing time series data (such as audio signals and video frame sequences), and can better capture the dynamic relationship between sound and expression and lip shape movements. For example, LSTM can remember the changing trend of sound features in time series through its unique memory unit, so as to more accurately generate corresponding expressions and lip shape movements.
[0086] Another important invention is the ability to generate personalized expressions and lip movements based on different cloned digital human character settings or real-life prototypes. By learning from a large amount of real-life video data with different styles and different character features, the model can capture the personalized differences between different characters when expressing the same emotion or voice. For example, for a cloned digital human with a cheerful personality, when expressing a happy emotion, his expression and lip movements may be more exaggerated and vivid than those of a cloned digital human with a calm personality. His eyes will open wider, his mouth corners will rise higher, and his lip movements will be more dynamic. This personalized generation technology makes the cloned digital human more realistic and unique.
[0087] In terms of feature extraction of the audio processing module, in addition to deep learning algorithms, traditional signal processing techniques combined with machine learning algorithms can also be used. For example, traditional methods such as Mel Frequency Cepstral Coefficients (MFCC) are used to extract the basic features of the audio, and then machine learning algorithms such as support vector machines (SVM) are used to classify these features and extract emotional features. This method may have certain advantages when the amount of data is relatively small, and the computational cost is relatively low. However, its disadvantage is that its ability to express complex sound emotions and diverse speech features may not be as good as deep learning algorithms.
[0088] like Figures 1 to 4As shown, the specific application principle and working process of the cloned digital human system are as follows:
[0089] Step 1: Image cloning. Specifically,
[0090] 1.1 Image acquisition: Use a mobile phone, camera, or other professional equipment to capture clear image videos of the person to be cloned with different expressions and postures, including the face.
[0091] 1.2 Feature extraction and model creation: Use image processing algorithms to extract the facial features of the person to be cloned, such as the shape, proportion, and skin color of the facial features, and create a basic model of the digital human based on the extracted features.
[0092] 1.3 Detail drawing: On the basis of the basic outline, draw more delicate details such as textures, hair, and makeup to enhance the realism.
[0093] Step 2: Voice cloning. Specifically,
[0094] 2.1 Audio acquisition: Record a large number of clear voice samples of the person to be cloned, covering various tones, speech rates, and emotional expressions.
[0095] 2.2 Voice feature analysis: Analyze the features such as spectrum, pitch, and duration of the collected speech.
[0096] 2.3 Model training: Use deep learning technology to train the voice cloning model so that it can learn the voice characteristics of the person to be cloned.
[0097] 2.4 Speech synthesis: When inputting text, the voice cloning model generates a voice similar to that of the person to be cloned according to the learned features.
[0098] Step 3: Semantic emotion judgment and expression generation. Specifically,
[0099] 3.1 Natural language processing: Perform semantic parsing on the input text to understand its meaning and emotional tendency.
[0100] 3.2 Emotion classification: Use the trained emotion classification model to judge the type of emotion expressed in the text.
[0101] 3.3 Expression library generation: Generate a rich 2D expression library through algorithms, including expressions of various common emotions.
[0102] 3.4 Expression matching: Select the corresponding 2D expression from the expression library according to the judged emotion.
[0103] Step 4: Video generation. Specifically,
[0104] 4.1 Synthesis and Rendering: Synthesize the image of the digital human, the matching expressions, the generated voice, and the designed actions, and perform rendering optimization to generate the final digital human video.
[0105] For example, when applying this method and the digital human cloning system to generate a New Year greeting video, first, according to the requirements, obtain the materials needed for the New Year greeting video. Based on these materials, generate a cloned digital human, synthesize various digital human images, expressions, voices, and designed actions that conform to the New Year greeting elements, and perform rendering optimization to generate a smooth video file that conforms to the theme.
[0106] Those skilled in the art should understand that those skilled in the art can implement variations in combination with the prior art and the above embodiments. Such variations do not affect the essence of this solution and will not be elaborated here.
[0107] It should be understood that this solution is not limited to the above specific implementation manners. The devices and structures not described in detail should be understood to be implemented in a common manner in this field; any person skilled in the art, without departing from the scope of this solution's technical solution, can make many possible changes and modifications to this solution's technical solution by using the methods and technical content disclosed above, or modify it into equivalent embodiments with equivalent changes, which does not affect the essence of this solution. Therefore, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of this solution without departing from the content of this solution's technical solution still fall within the scope of protection of this solution's technical solution.
Claims
1. A method for realizing synchronization of facial expressions and lip shapes of characters in a video by using emotions in sound, characterized in that: Collect audio information and extract multi-dimensional features of sound; A pre-trained expression and lip shape generation model is used to generate corresponding expression parameters and lip shape parameters according to the multi-dimensional features of the sound; the expression parameters and lip shape parameters are fused, and a continuous animation sequence is generated based on the fused parameters; the continuous animation sequence is rendered to generate video information with expression, lip shape and sound emotion synchronization.
2. The method for realizing lip synchronization of a person's facial expression in a video through emotions in sound according to claim 1, characterized in that: After collecting the audio information, first, the audio information is subjected to noise reduction processing, and then multi-dimensional feature extraction is performed on the noise-reduced audio information.
3. The method for realizing lip synchronization of a person's facial expression in a video by using emotions in sound according to claim 2, characterized in that: The multi-dimensional features include pitch features, timbre features, speech speed features, rhythm features and emotional features in the sound.
4. The method for realizing lip synchronization of a person's facial expression in a video by using emotions in sound according to claim 1, characterized in that: The expression and lip shape generation model includes an expression generation submodule, a lip shape generation submodule and a model training module, wherein the expression generation submodule is constructed based on a generative adversarial network, and its internal generator part takes sound emotion features as input and facial expression parameters as output; the lip shape generation submodule is constructed using a variational autoencoder, taking the speech features in the sound as input and the lip shape movement parameters as output; the model training module adopts a supervised learning algorithm, taking the sound features in the preprocessed video data as input, and the expression parameters and lip shape parameters as output.
5. The method for realizing lip synchronization of a person's facial expression in a video by using emotions in sound according to claim 4, characterized in that: The model training module repeatedly trains the models in the expression generation submodule and the lip shape generation submodule. During the training process, the weight parameters of the model are continuously adjusted so that the model can generate appropriate expressions and lip movements based on the input sound features. There is a connection relationship for parameter updating between the model training module and the corresponding model structures in the expression generation submodule and the lip shape generation submodule, and finally the trained parameters are updated to the corresponding model in a timely manner.
6. A digital human cloning system, characterized in that: The invention comprises a cloned digital human three-dimensional model and a control system thereof. The control system applies the method described in any one of claims 1 to 5 to obtain video information with synchronized expression, lip shape and voice emotion, and sends it to the cloned digital human three-dimensional model for synchronized output.
7. The digital human cloning system according to claim 6, characterized in that: The control system includes an audio processing module, an expression and lip shape generation module, a model training module, and a fusion and rendering module; wherein, The audio processing module includes an audio acquisition unit, a noise reduction unit, and a feature extraction unit connected in sequence. The audio acquisition unit receives an external input audio signal, and after noise reduction by the noise reduction unit, transmits the signal to the feature extraction unit for multi-dimensional feature extraction of the sound. The expression and lip shape generation module is built based on the generative model of deep learning, including the expression generation submodule and the lip shape generation submodule. The expression generation submodule generates the facial expression parameters of the cloned digital human, and the lip shape generation submodule generates the corresponding lip shape movement parameters. The model training module is used to train the deep learning generation model, including a data collection unit, a data preprocessing unit and a model training unit. The data collection unit collects real-person video data and annotates the expressions, lip shapes and sounds in the video. The data preprocessing unit cleans and normalizes the annotated data. The model training unit adopts a supervised learning algorithm, takes the sound features in the preprocessed video data as input, and the expression and lip shape parameters as output, and repeatedly trains the deep learning generation model in the expression and lip shape generation module until the optimal parameters of the model are obtained; The fusion and rendering module includes a parameter fusion unit, an animation generation unit and a rendering unit which are connected in sequence. The parameter fusion unit fuses the facial expression parameters and the lip shape movement parameters so that the expression and lip shape movement are coordinated with the facial movement of the cloned digital human. The animation generation unit generates a continuous animation sequence for the face of the cloned digital human according to the fused parameters, renders the cloned digital human model with expression and lip shape animation, and generates a video output with expression, lip shape and sound emotion synchronization.
8. The digital human cloning system according to claim 7, characterized in that: The feature extraction unit uses a deep learning algorithm to analyze the noise-reduced audio signal and extracts multiple features of the sound including pitch, timbre, speaking speed, rhythm, volume and emotion.
9. The digital human cloning system according to claim 8, characterized in that: The expression generation submodule generates facial expression parameters of the cloned digital human, including muscle movements of the eyes, eyebrows, and cheeks, based on the sound and emotional features extracted by the audio processing module. The lip shape generation submodule generates corresponding lip shape movement parameters, including the degree of lip opening and closing, the movement position of the corners of the mouth, and the coordination relationship between the lips and teeth, based on the speech content, speech speed, and rhythmic features in the audio.
10. The digital human cloning system according to claim 9, characterized in that: The video data collected by the data collection unit includes character expressions, lip movements and corresponding sound signals, and the annotated content includes expression types, corresponding relationships between lip movements and voice, and emotional features in the sound.
Citation Information
Patent Citations
Three-dimensional virtual image lip shape generation method and device and electronic equipment
CN113256821A
Method, system and device for driving image through voice and storage medium
CN116597857A
Method for generating digital human voice and facial animation through text
CN116863038A
Emotion-controlled three-dimensional virtual image expression animation generation method
CN117765137A
Method and device for improving fluency of facial expressions and actions of digital human and storage medium
CN119091012A
Cited By
Method and device for generating digital human video
CN121078285A
Digital human generation method based on sound driving
CN121304866A
A digital human generation method based on sound driving
CN121304866B
ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion
CN121354597A
Digital population type synchronization method and device based on phonon driving, equipment and medium
CN121545542A