A method for voice-driven digital mouth shapes
By extracting the time-frequency characteristics of the voice signal, generating personalized voiceprint characteristics, combining nonlinear dynamics and timing modeling, the problems of inaccurate synchronization and unnatural facial expressions in digital verbal generation are solved, and efficient and personalized digital verbal video generation is achieved, suitable for medical popularization, virtual assistants and entertainment fields.
Patent Information
- Application Number
- CN202510331501.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-20
AI Technical Summary
The prior art is difficult to achieve precise synchronization of speech and facial movements in digital verb generation, and personalized adjustments are difficult, facial expressions lack natural and realistic sense, and fail to effectively process the timing characteristics of the voice signal, resulting in unnatural video content.
By receiving speech signals, extracting time-frequency features, generating personalized voiceprint features, using nonlinear dynamics modeling and timing modeling methods, combining chaotic dynamics models and Hamiltonian systems, ensuring the synchronization of digital verbal and speech signals and biomechanical rationality, using a bidirectional long and short-term memory network for time synchronization, and generating personalized digital verbal videos.
The precise synchronization of voice signals and digital verbs is achieved. The generated video content can truly reflect the doctor's pronunciation style and emotional expression, improve the generation efficiency and nature, and enhance the immersion and interactivity of the video.
Smart Images

Figure CN119864040B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital humans, and in particular to a method for driving digital human figures with voice. Background Art
[0002] Traditional methods for generating digital population models usually rely on manual recording and manual design of lip shapes and facial expressions. This process is complex and time-consuming, and the generated digital population models often cannot accurately reflect the dynamic characteristics of speech signals. These traditional methods make it difficult to achieve precise synchronization between speech and facial movements. Especially in personalized scenarios, the generation of lip shapes and facial expressions often lacks naturalness and realism. Existing automated generation methods mostly rely on general models. Although these models can handle basic lip shape generation tasks, they are difficult to personalize according to the speech characteristics of different users, resulting in the generated digital population models being unable to perfectly adapt to the pronunciation habits and speaking speed of different users, showing unnatural or uncoordinated visual effects.
[0003] In addition, most existing technologies do not fully consider the complex dynamic relationship between lip shape and facial expression. The correspondence between speech signals and facial expressions is nonlinear, and the movement of facial muscles has a high degree of biological regularity. In traditional methods, the lip shape generation process often ignores subtle changes in speech signals and fails to consider the emotional fluctuations of speech and its impact on facial movements, resulting in the lack of emotional expression and dynamic changes in the generated digital mouth shape.
[0004] Furthermore, existing synchronization methods generally rely on static algorithm models and are unable to process the temporal features in speech signals in real time and efficiently, resulting in a less than smooth dynamic synchronization process between speech signals and facial expressions. Even in some efficient deep learning models, although they can process some temporal features of speech signals, they are still insufficient in capturing long-term dependencies and emotional features, making it difficult for the generated video content to fully present the emotional fluctuations and detailed changes in speech signals.
[0005] Therefore, the existing technology has significant deficiencies in terms of personalized lip shape generation, natural synchronization of speech and facial expressions, and dynamic changes of facial expressions. Summary of the invention
[0006] In view of the shortcomings of the prior art, the present invention provides a method for driving a digital population type with voice, which solves the problems in the prior art of inaccurate synchronization between the generation of a digital population type and a voice signal, difficulty in personalized adjustment, and unnatural facial expressions.
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: A method for voice-driven digital population type, comprising the following steps:
[0008] Receive a voice signal, preprocess the voice signal, and extract the time-frequency features of the voice signal;
[0009] Generate personalized voiceprint features based on the time-frequency features of the voice signal, where the voiceprint features are related to the personalized voice features of the doctor;
[0010] Use nonlinear dynamics modeling to generate a digital mouth shape synchronized with the voice signal, where the nonlinear dynamics modeling describes the nonlinear relationship between the voice signal and the mouth shape based on a chaotic dynamics model;
[0011] Synchronize the voice signal and the digital mouth shape through a time series modeling method to ensure that the mouth shape is aligned with the voice content;
[0012] Generate a final digital mouth shape video according to the generated synchronization result and output the video.
[0013] Preferably, the time-frequency features of the voice signal are extracted by short-time Fourier transform, and the short-time Fourier transform is based on frame segmentation of the voice signal to calculate the spectral information of each frame to obtain a joint representation of time and frequency.
[0014] Preferably, the generation of the personalized voiceprint features is jointly achieved by a convolutional neural network and a long short-term memory network. The convolutional neural network is used to extract local features of the voice signal, and the long short-term memory network is used to capture the temporal relationship in the voice signal to obtain the personalized voice features of the doctor.
[0015] Preferably, the step of using nonlinear dynamics modeling to generate a digital mouth shape synchronized with the voice signal includes:
[0016] Based on the time-frequency features of the voice signal, construct a nonlinear dynamics model, where the change of the voice signal is associated with the facial movements of the digital mouth shape through the state equation of a dynamic system;
[0017] The state equation is as follows:
[0018]
[0019] Where, is the state vector of the system, representing the state of the digital mouth shape; is the input vector of the voice signal, containing the time-frequency features extracted from the voice signal; is the parameter of the model, is the function describing the change of the system state;
[0020] Model the nonlinear relationship between the voice signal and the digital mouth shape through the state equation to generate a digital mouth shape synchronized with the voice signal.
[0021] Preferably, the step of generating a digital mouth shape synchronized with the speech signal using non-linear dynamics modeling further includes:
[0022] Applying the Lyapunov exponent to the modeling system for stability analysis, where the Lyapunov exponent is calculated by the following formula:
[0023]
[0024] where, represents the perturbation term of the system at time and represents the perturbation term of the system at the initial time, and
[0025] is the Lyapunov exponent;
[0026] Preferably, the step of generating a digital mouth shape synchronized with the speech signal using non-linear dynamics modeling further includes:
[0027] Constraining the digital mouth shape generation process using a Hamiltonian system, where the equation of motion of the Hamiltonian system is:
[0028]
[0029] where, is the Hamiltonian; is the displacement coordinate of the digital mouth shape face, representing the physical positions of the various parts of the face in the digital mouth shape; is the momentum, representing the motion state of the facial muscles; is the mass of the facial muscles, representing the motion inertia of the facial muscles; is the potential energy function, representing the potential energy of the facial muscles; and respectively represent the derivatives of the Hamiltonian with respect to position and momentum;
[0030] By constraining with the Hamiltonian system, it is ensured that the generation of the digital mouth shape follows physical laws, and the motion of the facial muscles matches the changes in the speech signal.
[0031] Preferably, the step of time-synchronizing the speech signal and the digital mouth shape by a time series modeling method includes:
[0032] Using a time series modeling method to time-align the speech signal and the digital mouth shape, and learning the time series relationship between the speech signal and the mouth shape through a bidirectional long short-term memory network;
[0033] The time-frequency features and digital mouth shape features of the speech signal are input and processed through a bidirectional long short-term memory network to obtain a feature output containing forward and backward temporal information, and the time delay of mouth shape generation is adjusted to ensure precise synchronization between speech and mouth shape.
[0034] Preferably, the step of generating the final digital mouth shape video according to the generated synchronization result and outputting the video includes:
[0035] Combining the digital mouth shape synchronized by the temporal modeling method with the speech signal;
[0036] According to the synchronization result, adjust the facial expression of the digital mouth shape to match the emotional information of the speech signal;
[0037] Based on each frame image of the generated digital mouth shape, the image includes facial movements and mouth shape dynamics;
[0038] Combine the generated image sequence into a video file in chronological order, where each frame image represents the facial expression and mouth shape of the digital mouth shape at a specific moment;
[0039] Output the video file so that it can be viewed and used by the user.
[0040] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method as described above is implemented.
[0041] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method as described above is implemented.
[0042] The present invention provides a method for voice-driven digital mouth shape. It has the following beneficial effects:
[0043] The present invention realizes the precise synchronization of the speech signal and the digital mouth shape through nonlinear dynamics modeling and temporal modeling methods. The pronunciation of each syllable can be accurately mapped to the movement of facial muscles, ensuring that the generated video is highly consistent with the speech content in terms of mouth shape changes.
[0044] By extracting the personalized voiceprint features of the doctor, the generated digital mouth shape is not only synchronized with the speech but also can reflect the unique voice characteristics of the doctor, such as pitch, speech rate, and timbre. This makes the generated video able to more realistically restore the doctor's pronunciation style and emotional expression.
[0045] The present invention introduces a Hamiltonian system to physically constrain the generation of facial expressions, ensuring that the movement of facial muscles conforms to biomechanical laws. Facial expressions not only synchronize with the emotional information of speech signals but also present a more natural and realistic visual effect.
[0046] Through the technology of the present invention, the generation process of digital lip-sync videos is optimized, the synchronization between speech signals and lip movements is greatly improved, and at the same time, the video output is fast, avoiding the complex manual production process in traditional methods and significantly improving the efficiency.
[0047] The digital lip-sync video generation technology of the present invention can be widely applied to multiple fields such as medical popularization, virtual assistants, entertainment, and education. It can not only provide personalized virtual character performances but also enhance the audience's immersion and interactivity, with broad market potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a schematic flowchart of the method of the present invention;
[0049] Figure 2 is a schematic structural diagram of the device of the present invention;
[0050] Figure 3 is a schematic structural diagram of the computer device of the present invention.
[0051] Among them, 100 is a speech signal input module; 200 is a speech signal processing module; 300 is a voiceprint feature extraction module; 400 is a lip movement generation module; 500 is a synchronization module; 600 is an output module; 40 is a computer device; 41 is a processor; 42 is a memory; 43 is a storage medium. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] Please refer to the attached Figure 1 , the present invention provides a method for voice-driven digital lip-sync, aiming to generate a digital lip-sync video synchronized with the doctor's speech signal according to the doctor's speech signal. This method combines multiple technical means to ensure that the lip movements of the generated video are highly consistent with the speech content and can reflect the personalized speech characteristics of the doctor. The core of this embodiment includes feature extraction of speech signals, generation of personalized voiceprints, non-linear dynamics modeling, timing modeling synchronization, and finally the generation and output of digital lip-sync videos.
[0054] As shown in Figure 1 the method for driving a digital human mouth shape by voice may include the following steps:
[0055] S1. Receive a voice signal, preprocess the voice signal, and extract the time-frequency features of the voice signal;
[0056] S2. Generate personalized voiceprint features based on the time-frequency features of the voice signal;
[0057] S3. Use nonlinear dynamics modeling to generate a digital human mouth shape synchronized with the voice signal;
[0058] S4. Synchronize the voice signal and the digital human mouth shape in time through a time series modeling method;
[0059] S5. Generate a final digital human mouth shape video according to the generated synchronization result and output the video.
[0060] The method steps of the present invention and their technical implementation will be described in detail below.
[0061] For step S1, in this embodiment, by receiving and processing the voice signal, necessary input is provided for subsequent personalized voiceprint feature extraction and mouth shape generation. The specific operation process is as follows:
[0062] In this step, first, a voice signal of a doctor is received through a suitable voice input device (such as a microphone, a voice acquisition card, etc.). The received voice signal is usually a time-domain signal and contains the pronunciation information of the doctor. The quality of this voice signal directly affects the subsequent processing results. Therefore, the signal needs to be preprocessed first.
[0063] The key to preprocessing the voice signal lies in removing noise and eliminating unnecessary signal interference. In order to improve the clarity of the signal, a band-pass filter or an adaptive filtering method can be used for noise removal. For example, a Wiener filter is used to smooth the voice signal to reduce the influence of background noise and other irrelevant frequencies. In addition, the amplitude of the voice signal can be adjusted through a normalization process so that the dynamic range of the signal is suitable for the subsequent feature extraction process.
[0064] After the preprocessing of the voice signal is completed, time-frequency feature extraction is performed. Since the voice signal is essentially non-stationary, that is, its spectrum changes with time, it is necessary to process it through time-frequency analysis techniques. In the present invention, the short-time Fourier transform (STFT) and the wavelet transform are the main methods for extracting the time-frequency features of the voice signal.
[0065] Specifically, the short-time Fourier transform (STFT) is used to frame the voice signal and calculate the spectrum information of each frame to obtain a joint representation of time and frequency. The specific operation can be realized through the following formula:
[0066]
[0067] Among them, is the input voice signal, is the window function, is the time-frequency representation after short-time Fourier transform. By frame division and Fourier transform, the spectral information of each frame can be obtained, thereby capturing the changes of the signal in time and frequency.
[0068] In addition, as another time-frequency analysis method, wavelet transform can capture the local features of the signal more precisely. Through multi-scale analysis, wavelet transform can not only capture low-frequency information, but also effectively extract high-frequency components. In this embodiment, wavelet transform is used for time-frequency feature extraction, and the specific formula is as follows:
[0069]
[0070] Among them, is the wavelet basis function, and are the scale and translation parameters respectively, is the wavelet transform coefficient. Through this method, multi-scale analysis of the voice signal can be obtained, making the time-frequency feature extraction more refined.
[0071] Through STFT and wavelet transform, the extracted time-frequency features provide a basis for the generation of personalized voiceprint features in subsequent steps. These time-frequency features capture the spectral information of the voice signal and can effectively reflect personalized features such as pitch and speech rate in the voice.
[0072] The time-frequency features extracted in this step have the following advantages: First, they can describe the changes of the voice signal in time and frequency, providing rich voice information; Second, the time-frequency features can help accurately match the pronunciation habits of doctors in the generation of medical popular science videos, providing accurate input for the generation of personalized digital lip shapes.
[0073] In summary, step S1 receives the voice signal, performs preprocessing, and uses the time-frequency feature extraction method to obtain the time-frequency features of the voice, laying a foundation for the subsequent generation of personalized voiceprint features and digital lip shapes.
[0074] For step S2, in this embodiment, step S2 generates voiceprint features related to the personalized features of the doctor's voice based on the time-frequency features of the voice signal extracted in step S1. By extracting the voiceprint features, the personalized pronunciation features of each doctor can be effectively captured to ensure that the generated digital lip shape can highly match the voice features of the doctor.
[0075] In this embodiment, the time-frequency features of the voice signal are the spectral information extracted from the doctor's voice signal, and these time-frequency features can reflect personalized features such as pitch, timbre, and speech rate in the voice. For these time-frequency features, the convolutional neural network (CNN) and long short-term memory network (LSTM) in the deep learning method are used to extract the voiceprint features.
[0076] First, the convolutional neural network (CNN) is used to extract local features from the time-frequency map. The time-frequency features of the voice signal are usually spectrograms obtained through short-time Fourier transform or wavelet transform. The convolutional neural network can process these spectrograms through its convolutional layers to capture local feature patterns. The advantage of the convolutional neural network is that through convolutional operations, it can effectively extract local information in the spectrum, such as syllables, pitch, and tone changes.
[0077] Secondly, the long short-term memory network (LSTM) is used to capture the temporal information of the voice signal, especially dynamic changes such as pitch and intonation. As a recurrent neural network with long-term memory capabilities, LSTM can handle long-term dependencies in the voice signal. Therefore, the LSTM network plays a key role in voiceprint feature extraction, and can associate the local features extracted from the CNN with the time series to generate more accurate personalized voiceprint features.
[0078] The specific operation process is as follows:
[0079] Input of time-frequency features: The time-frequency features of the voice signal obtained in step S1 are used as the input of the CNN. These features can reflect the changes of the voice signal at different time points and frequencies, thus helping the network capture the local features of the voice.
[0080] Processing by convolutional layers: The CNN performs convolutional processing on the time-frequency map through multiple convolutional layers to extract local features, such as syllables and timbre. These local features are crucial for describing the personalized features of the doctor's voice. The convolutional process slides the convolutional kernel to process the voice features within each time window, thereby capturing local changes.
[0081] Extraction of temporal information: After extracting the local features, the LSTM network is used to further process these features to capture the temporal dependencies in the voice signal. LSTM can handle the dynamic changes with a long time span in the voice signal and generate feature vectors related to the personalization of the doctor's voice.
[0082] The generated personalized voiceprint features represent the uniqueness of the doctor's voice, such as pitch, speech rate, and timbre, and provide a basis for the subsequent generation of digital mouth shapes. These voiceprint features can not only ensure the high synchronization of the digital mouth shape with the doctor's voice, but also reflect the personalized characteristics in the doctor's voice, making the generated digital mouth shape more personalized.
[0083] In this embodiment, by combining CNN and LSTM, the deep mining of the time-frequency features of speech signals is successfully realized, and voiceprint features closely related to the personalized speech features of doctors are generated. This process not only improves the accuracy of digital mouth shape generation but also enhances the personalization of the generation results, making the generated video content more realistic and conforming to the speech features of doctors.
[0084] For step S3, in this embodiment, a digital mouth shape synchronized with the speech signal is generated through nonlinear dynamics modeling. In this step, a chaotic dynamics model is used to capture the nonlinear relationship between the speech signal and the digital mouth shape and ensure that the generated mouth shape is highly synchronized with the speech signal.
[0085] During the generation of the digital mouth shape, the relationship between the speech signal and the mouth shape is complex and nonlinear. The change of the speech signal will directly affect the movement of facial muscles, and this kind of influence is nonlinear and time-varying. To accurately describe this complex relationship, this embodiment adopts a chaotic dynamics model, which can effectively capture the nonlinear characteristics existing in the system.
[0086] First, based on the time-frequency features of the speech signal (for example, the spectrogram obtained through short-time Fourier transform or wavelet transform), a state equation of a dynamic system is established. This equation describes how the input of the speech signal affects the facial muscle dynamics of the digital mouth shape.
[0087] The state equation of this dynamic system is expressed as:
[0088]
[0089] Where:
[0090] represents the state vector of the digital mouth shape, describing the movement of facial muscles, such as the dynamic changes of the lips, tongue, and chin;
[0091] is the input signal, containing the time-frequency features of the speech signal. Specifically, these features include time-varying features such as the frequency, amplitude, and syllables of the audio;
[0092] is the parameter set of the system, controlling the response characteristics of facial muscles, such as the tension of muscles and the movement characteristics of facial bones, etc.;
[0093] is a nonlinear function describing the change of the system state, which links the speech signal with the dynamic response of facial muscles.
[0094] This model is based on the chaotic dynamics method, through the nonlinear function Precisely simulating the non-linear mapping relationship between speech and lip shapes ensures that the lip shapes at each time point are consistent with the instantaneous characteristics of the speech signal.
[0095] To ensure that the generated digital lip shapes are physically stable and reasonable, this embodiment introduces the Lyapunov exponent to analyze the stability of the system. The Lyapunov exponent is an important indicator for measuring the sensitivity of a dynamic system to initial conditions. By calculating the Lyapunov exponent, the stability of the model during lip shape generation can be evaluated, avoiding unnatural lip shape transformations in the system.
[0096] In this embodiment, the calculation formula of the Lyapunov exponent is as follows:
[0097]
[0098] Where:
[0099] represents the perturbation term of the system at time , reflecting the change in the state of facial muscles;
[0100] is the perturbation term at the initial time of the system;
[0101] is the Lyapunov exponent, representing the stability of the system. If , the system is unstable; if , the system tends to be stable.
[0102] In this embodiment, by calculating the Lyapunov exponent, the stability of the model can be evaluated, and the system parameters (i.e., ) can be adjusted according to the results to ensure that the generated lip shape changes smoothly and naturally, avoiding unstable facial movements.
[0103] To ensure that the movement of facial muscles conforms to biomechanical laws, this embodiment further introduces the Hamiltonian system to impose physical constraints on the lip shape generation process. The Hamiltonian system can simulate the movement of physical systems and is particularly suitable for describing the behavior of particles in a force field. In this embodiment, the Hamiltonian system is used to describe the dynamic response of facial muscles, thus ensuring that the generated digital lip shapes conform to physical laws.
[0104] Specifically, the equations of the Hamiltonian system are as follows:
[0105]
[0106] Where:
[0107] is the Hamiltonian, representing the total energy of the system, including kinetic energy and potential energy;
[0108] are the displacement coordinates of the digital population-type facial muscles, representing the physical positions of various parts of the face (such as the mouth, eyes, chin, etc.);
[0109] is the momentum, representing the motion state of the facial muscles;
[0110] is the mass of the facial muscles, describing the inertia of the facial muscles;
[0111] is the potential energy function, representing the potential energy of the facial muscles, which is related to the geometric shape of the face and the tension of the muscles;
[0112] and respectively represent the partial derivatives of the Hamiltonian with respect to position and momentum, describing the dynamic changes of the facial muscles.
[0113] Through the Hamiltonian system, the motion of the facial muscles is not only driven by the speech signal but also follows physical laws, thus ensuring that the generated mouth shapes are biologically reasonable.
[0114] Through the above non-linear dynamics modeling, Lyapunov exponent stability analysis, and Hamiltonian system constraints, the speech signal is combined with the facial dynamic changes of the digital population type. The generated digital population type is not only highly synchronized with the speech signal but also meets the requirements of biomechanics in terms of motion, avoiding unnatural mouth shapes and facial expressions.
[0115] Through this model, each syllable, speech rate change, etc. of the speech signal can be accurately converted into facial movements, ensuring that the generation of the digital population type is both accurate and natural. In particular, the introduction of the Lyapunov exponent and Hamiltonian system constraints greatly enhances the stability and physical rationality of the generation process, ensuring that the generated mouth shapes and facial expressions not only conform to the rhythm and emotional changes of the speech content but also conform to biomechanical characteristics, making the finally generated video more realistic and credible.
[0116] For step S4, in this embodiment, the speech signal and the digital population type are ensured to be accurately synchronized in time through a timing modeling method. During the generation process of the digital population type, the changes in the mouth shapes must highly match the rhythm, syllables, and their dynamic characteristics of the speech signal to ensure that the finally generated video content is natural and conforms to the actual pronunciation rules.
[0117] To this end, the present invention uses a Bidirectional Long Short-Term Memory Network (Bi-LSTM) for time series modeling. Bi-LSTM is a deep learning model with powerful time series learning capabilities, capable of capturing the long-term dependencies between speech signals and lip movements, ensuring precise temporal alignment between lip movements and speech signals.
[0118] In this embodiment, first, the time-frequency features of the speech signals extracted from steps S1 and S2 are used as the input sequence of the Bi-LSTM model. These time-frequency features can comprehensively reflect the changes of speech signals in time and frequency, providing key data support for time series modeling.
[0119] The Bi-LSTM model can capture the temporal information in speech signals more comprehensively through the learning mechanisms in both the forward and backward propagation directions. By forward and backward propagation of the input speech signal features, Bi-LSTM can process information in two time directions respectively, thereby capturing the past and future features of speech simultaneously. This feature enables Bi-LSTM to model lip movements more accurately and synchronize lip movements with speech signals precisely.
[0120] In the specific implementation process, the state update formula of the Bi-LSTM network is:
[0121]
[0122] Where:
[0123] is the output at the current moment, representing the generated lip movement state;
[0124] and are the states at the previous and next moments respectively, representing the forward and backward information of the speech signal;
[0125] is the feature of the speech signal input at the current moment;
[0126] is the weight matrix, responsible for mapping the input information to the output space;
[0127] is the bias term, used to adjust the offset of the model output.
[0128] Through this formula, the Bi-LSTM network can capture the dynamic changes at each moment in the time series, thereby generating lip movements synchronized with the speech signal.
[0129] In addition, during the temporal synchronization process, the Bi-LSTM model can automatically adjust the duration, intensity, and transitions of the mouth shapes to highly match the syllables, pitch, pauses, and other features of the speech signal. By continuously learning the rhythm, speech rate, and pitch changes in the speech signal, the model dynamically adjusts the mouth shape changes for each frame to ensure that the generated digital human mouth shapes have a natural and accurate synchronization effect.
[0130] Through temporal modeling, Bi-LSTM can not only handle short-term speech variations but also capture the dynamic features with long time spans in the speech. In this way, the generation process of the digital human mouth shapes is not only synchronized with the current syllable but also can adapt to the long-term speech feature changes, thus ensuring the coherence and fluency of the mouth shape changes.
[0131] Through step S4, the synchronization between the digital human mouth shape changes and the speech signal is significantly enhanced. The generation of each frame of the mouth shape is dynamically adjusted based on the temporal features of the speech signal, making the finally generated digital human mouth shapes highly natural and consistent, and capable of accurately reflecting the rhythm, emotion, and pronunciation features of the speech.
[0132] In this embodiment, by introducing the Bi-LSTM network for temporal modeling, the high synchronization between the mouth shape and the speech signal is ensured. The bidirectional information transfer mechanism of the Bi-LSTM network enables the mouth shape generation to not only consider the speech features at the current moment but also make full use of the speech information at the previous and subsequent moments for prediction, thereby improving the accuracy of temporal synchronization. This method effectively eliminates the unnatural feeling caused by the asynchronous changes between the speech signal and the mouth shape in the traditional method, while improving the fluency and realism of the digital human mouth shape generation.
[0133] Through this step, the synchronization process between the speech signal and the mouth shape is precisely aligned. The generated digital human mouth shapes are consistent with the doctor's speech signal, being not only visually real and vivid but also capable of reflecting the detailed changes in the speech. This method provides an efficient and natural digital human mouth shape generation technology for applications such as medical popular science videos and virtual doctors, enhancing the user experience and the conveying effect of the video content.
[0134] For step S5, in this embodiment, according to the synchronization result generated by the previous temporal modeling method, the digital human mouth shape video is finally generated and output for users to use. The technologies involved in this step mainly include converting the facial movements and mouth shapes after synchronizing the digital human mouth shapes with the speech signal into each frame of image and combining them into a complete video file, ensuring that the video content is natural, fluent, and highly consistent with the mouth shape, rhythm, and emotion changes of the speech signal.
[0135] In this embodiment, first, according to the synchronization result generated by the timing modeling method in step S4, the facial expression of the digital mouth shape is adjusted to match the emotional information of the voice signal. In this way, the digital mouth shape is not only synchronized with the voice signal in terms of mouth shape, but also synchronized with the emotional changes of the voice in terms of expression.
[0136] The generation process of facial movements for each frame is as follows:
[0137] Facial movement and mouth shape generation: According to the synchronization result, the movements of facial muscles are generated by the model. These movements of facial muscles include dynamic changes in parts such as the lips, chin, and eyes. Each frame of the image reflects the mouth shape and facial expression of the digital mouth shape at a certain moment.
[0138] Mathematical modeling of the generation process: When generating each frame of the image, facial models in computer graphics (such as 3D facial modeling) and muscle-driven technologies (such as facial animation-driven models) can be used to model the movements of facial parts. Based on the non-linear dynamics model provided in step S3, the movement states of facial muscles are calculated and driven by parametric equations (such as facial muscle models).
[0139] The generation of each frame of the digital mouth shape image takes the following factors into consideration:
[0140] Dynamic response of facial muscles: The actions and tensions of each facial muscle will be dynamically adjusted according to the syllables and speech rate in the voice signal. The specific facial movement model can be described by the movement equations of facial muscles and adjusted accordingly according to the generated timing synchronization result.
[0141] Matching of facial expressions: In addition to mouth shape synchronization, facial expressions (such as smiling, frowning, etc.) will also be matched according to the emotional information of the voice signal. Through emotional analysis algorithms, the emotional features in the voice signal (such as happiness, sadness, surprise, etc.) are identified, and then these emotional features are mapped to the facial expressions of the digital mouth shape.
[0142] Each generated frame of the image is arranged in chronological order to form a series of continuous image frames. The specific process includes:
[0143] Image synthesis: The generated digital mouth shape facial image for each frame is synthesized with the background scene to generate an image containing the complete facial expression and background. The synthesized image maintains the synchronization and emotional consistency of the voice signal.
[0144] Video synthesis: The continuous image frames are combined into a final video file through video synthesis technologies (such as encoding technologies). The frame rate, resolution, and format of the video can be adjusted according to specific requirements. Common video formats include MP4, AVI, MOV, etc.
[0145] The generated digital lip-synced video not only has a high degree of synchronization but also can truly reflect the syllables, emotions, and speech rate changes in the speech, making the video content natural and having strong interactivity and immersion.
[0146] Finally, the generated video file will be output in a standard video coding format, enabling users to view, share, and use the generated video content. The output format of the video can include but is not limited to MP4 format, AVI format, etc., ensuring wide compatibility and application scenarios.
[0147] Through step S5 in this embodiment, the generated digital lip-synced video can be applied in multiple fields such as medical science popularization, virtual assistants, and entertainment. The digital lip-sync in the video is highly synchronized with the doctor's voice signal, and the lip movements and facial expressions not only conform to the rhythm and emotional changes of the speech content but also can reflect the doctor's personalized voice characteristics. The realization of this technology makes the video content more vivid and real, and improves the acceptance and understanding of the video content by users.
[0148] In addition, during the generation process of the digital lip-synced video, the physical constraints and emotional mapping of facial muscles are fully considered, so that the generated video is not only precisely synchronized in visual effects but also can convey the emotional information in the speech, thereby enhancing the interactivity and expressiveness of the video.
[0149] Through this step, the present invention can achieve an efficient conversion from voice signals to digital lip-synced videos, providing a new and personalized digital video generation method, which has important application value and broad market prospects.
[0150] Generally speaking, the present invention realizes a high degree of synchronization between the digital lip-sync and the voice signal by receiving the voice signal, extracting time-frequency features, combining personalized voiceprint features to generate non-linear dynamics modeling. This method uses a chaotic dynamics model and time series modeling technology to synchronize the voice signal and the digital lip-sync in time through a bidirectional long short-term memory network (Bi-LSTM), and combines the Hamiltonian system to physically constrain the generation of facial expressions, ensuring that the generated digital lip-sync completely matches the speech rhythm, emotion, and speech rate. Finally, by generating and synthesizing the video file frame by frame, the generation of a digital lip-synced video that is highly consistent with the speech content, natural, vivid, and has personalized features is realized, which is widely applicable to fields such as medical science popularization and virtual assistants, and has important application value and market prospects.
[0151] The device for voice-driven digital lip-sync described below can be correspondingly referred to the method for voice-driven digital lip-sync described above.
[0152] Please refer to the appendix Figure 2 , the present invention also provides a device for voice-driven digital lip-sync, including:
[0153] A voice signal input module 100 for receiving voice signals;
[0154] A voice signal processing module 200 for preprocessing the voice signals and extracting the time-frequency features of the voice signals;
[0155] A voiceprint feature extraction module 300 for generating personalized voiceprint features related to the personalized features of the doctor's voice;
[0156] A mouth shape generation module 400 for generating a digital human mouth shape synchronized with the voice signals based on the time-frequency features and personalized voiceprint features using nonlinear dynamics modeling;
[0157] A synchronization module 500 for synchronizing the time of the voice signals and the digital human mouth shape through a timing modeling method;
[0158] An output module 600 for generating and outputting a final digital human mouth shape video.
[0159] The device of this embodiment can be used to execute the above method embodiment, and its principle and technical effects are similar, which will not be elaborated here.
[0160] Please refer to the attached Figure 3 , the present invention also provides a computer device 40, including: a processor 41 and a memory 42, where the memory 42 stores a computer program executable by the processor, and when the computer program is executed by the processor, it executes the above method.
[0161] The present invention also provides a storage medium 43, on which a computer program is stored, and when the computer program is run by the processor 41, it executes the above method.
[0162] Among them, the storage medium 43 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (Static Random Access Memory, abbreviated as SRAM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), an erasable programmable read-only memory (Erasable Programmable Read-Only Memory, abbreviated as EPROM), a programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), a read-only memory (Read-OnlyMemory, abbreviated as ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disc.
[0163] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for voice-driven digital mouth shapes, characterized in that, It includes the following steps: Receive a voice signal, preprocess the voice signal, and extract the time-frequency features of the voice signal; Generate personalized voiceprint features based on the time-frequency features of the voice signal; Use nonlinear dynamics modeling to generate a digital mouth shape synchronized with the voice signal, and the nonlinear dynamics modeling describes the nonlinear relationship between the voice signal and the mouth shape based on a chaotic dynamics model; The step of using nonlinear dynamics modeling to generate a digital mouth shape synchronized with the voice signal includes: Based on the time-frequency features of the voice signal, construct a nonlinear dynamics model, where the change of the voice signal is associated with the facial movements of the digital mouth shape through the state equation of the dynamic system; Model the nonlinear relationship between the voice signal and the digital mouth shape through the state equation to generate a digital mouth shape synchronized with the voice signal; The step of using nonlinear dynamics modeling to generate a digital mouth shape synchronized with the voice signal further includes: Apply the Lyapunov exponent to the modeling system for stability analysis, and the Lyapunov exponent is calculated by the following formula: where δ(t) represents the perturbation term of the system at time t, δ(0) represents the perturbation term of the system at the initial time, and λ is the Lyapunov exponent; By calculating the Lyapunov exponent, ensure the stability of the digital mouth shape generation process and avoid the generation of unnatural mouth shapes; The step of using nonlinear dynamics modeling to generate a digital mouth shape synchronized with the voice signal further includes: Use the Hamiltonian system to constrain the digital mouth shape generation process, and the motion equation of the Hamiltonian system is: where \(H\) is the Hamiltonian; \(q\) is the displacement coordinate of the digital humanoid face, representing the physical positions of the various parts of the face in the digital humanoid; \(p\) is the momentum, representing the motion state of the facial muscles; \(m\) is the mass of the facial muscles, representing the motion inertia of the facial muscles; \(V(q)\) is the potential energy function, representing the potential energy of the facial muscles; and respectively represent the derivatives of the Hamiltonian with respect to position and momentum; Through the constraint of the Hamiltonian system, ensure that the digital mouth shape generation follows physical laws, and the movement of facial muscles matches the change of the voice signal; Synchronize the voice signal and the digital mouth shape in time through a time series modeling method to ensure that the mouth shape is aligned with the voice content; Generate a final digital mouth shape video according to the generated synchronization result and output the video.
2. The method for driving the digital mouth shape by voice according to claim 1, wherein The time-frequency features of the voice signal are extracted by short-time Fourier transform, and the short-time Fourier transform is based on the frame-by-frame processing of the voice signal to calculate the spectral information of each frame to obtain a joint representation of time and frequency.
3. The method for voice-driven digital mouth shapes according to claim 1, wherein The generation of the personalized voiceprint features is jointly realized by a convolutional neural network and a long short-term memory network. The convolutional neural network is used to extract the local features of the voice signal, and the long short-term memory network is used to capture the time series relationship in the voice signal to obtain the personalized voice features of the doctor.
4. The method for voice-driven digital mouth shape according to claim 1, wherein The state equation is as follows: where x(t) is the state vector of the system, representing the state of the digital mouth shape; u(t) is the input vector of the voice signal, containing the time-frequency features extracted from the voice signal; θ is the parameter of the model, and f is the function describing the change of the system state.
5. The method for driving a digital mouth shape by voice according to claim 1, characterized in that, The step of synchronizing the voice signal and the digital mouth shape in time through a time series modeling method includes: Use a timing modeling method to perform time alignment between the speech signal and the digital mouth shape, and learn the timing relationship between the speech signal and the mouth shape through a bidirectional long short-term memory network; Input and process the time-frequency features of the speech signal and the digital mouth shape features through a bidirectional long short-term memory network to obtain a feature output containing forward and backward timing information, and adjust the time delay of mouth shape generation to ensure precise synchronization between the speech and the mouth shape.
6. The method for driving a digital mouth shape by voice according to claim 1, wherein The steps of generating the final digital mouth shape video according to the generated synchronization result and outputting the video include: Combine the digital mouth shape synchronized by the timing modeling method with the speech signal; According to the synchronization result, adjust the facial expression of the digital mouth shape to match the emotional information of the speech signal; Based on each frame image of the generated digital mouth shape, the image includes facial movements and mouth shape dynamics; Combine the generated image sequence into a video file in chronological order, where each frame image represents the facial expression and mouth shape of the digital mouth shape at a specific moment; Output the video file so that it can be viewed and used by the user.
7. A device for voice-driven digital mouth shapes, for performing the method of voice-driven digital mouth shapes according to any one of claims 1-6, characterized in that, Including: A speech signal input module for receiving a speech signal; A speech signal processing module for preprocessing the speech signal and extracting the time-frequency features of the speech signal; A voiceprint feature extraction module for generating personalized voiceprint features related to the personalized features of the doctor's speech; A mouth shape generation module that generates a digital mouth shape synchronized with the speech signal using nonlinear dynamics modeling based on the time-frequency features and personalized voiceprint features; A synchronization module for time-synchronizing the speech signal and the digital mouth shape through a timing modeling method; An output module for generating and outputting the final digital mouth shape video.
8. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the method of speech-driven digital mouth shape according to any one of claims 1-6.
Citation Information
Patent Citations
Method for fusing brain electricity and muscle electricity signal chaos characteristics for hand motion identification
CN101732110A
Method for extracting characteristic parameters in speech recognition
CN102646415A