Digital human speech lip synchronization control system based on dynamic fusion of multimodal features

Through the dynamic fusion technology of multimodal feature, using historical speech data and audio and video feature recognition, personalized voice files are constructed and lip-like correction is performed, which solves the personalization and fluency problems in digital lip-like synchronization, and achieves more natural and accurate lip-like generation.

CN120319262BActive Publication Date: 2025-08-19HANGZHOU XINGMAI YUNSHANG TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510806848.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-08-19
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The existing digital human voice lip sync technology lacks personalized modeling, resulting in lip typing generation that does not conform to user habits, and the quality of lip animations decreases when pronunciation is unclear or speech speed changes, affecting the authenticity and credibility of the interaction.

Method used

Through dynamic fusion of multimodal features, a personalized voice archive and speech mapping model is constructed using historical speech data, combined with audio and video features for identification and correction, and personalized lip control parameters are generated, including the recognition and correction of content abnormalities, matching abnormalities and smooth abnormalities.

Benefits of technology

It improves the accuracy and consistency of lip shape generation, enhances the personalized expressiveness of digital lip shapes, ensures smooth matching of lip shapes and voice content under different pronunciation conditions, and improves the naturalness and credibility of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120319262B_ABST
    Figure CN120319262B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of speech recognition control technology, specifically to a digital human speech and lip-sync control system based on dynamic fusion of multimodal features. The system comprises obtaining historical speech data of a user; performing feature recognition on the historical speech data to obtain historical speech features and historical lip-sync features, and constructing a personalized speech profile; and constructing a speech mapping model based on the historical speech features and historical lip-sync features; receiving data input by the user to obtain first data; performing feature recognition based on the first data to obtain first data features; generating second data based on the speech mapping model and the first data features; and constructing a mapping correction model to correct the second data and output a third control parameter. The present invention achieves synchronized control of the digital human speech and lip-sync via the third control parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition control, and in particular to a digital human speech lip synchronization control system based on dynamic fusion of multimodal features. Background Art

[0002] Voice-driven lip synchronization is a hot research topic in the field of digital human technology. While mainstream lip synchronization technology has made significant progress using methods such as deep learning, enabling the generation of corresponding lip animations from input voice signals or text, it still faces numerous challenges, limiting the naturalness and personalization of digital human interactions.

[0003] First, the lip-syncing styles generated by many existing systems are relatively general and single, lacking the ability to fine-tune the modeling and adaptation of the user's unique speaking habits, lip-syncing characteristics, and the mapping relationship between pronunciation and lip-syncing. This makes it difficult for digital human images to reflect individual differences and personal charm in lip-syncing expression.

[0004] Secondly, when the user inputs unclear pronunciation, rapid changes in speech speed, and unnatural speech rhythm, the quality of the lip animation generated by the existing system will often drop significantly, resulting in a mismatch between the lip shape and the actual speech content, lip movements becoming stuck or jumping, and even the visual information conveyed by the lip shape deviating from the semantics of the speech content, seriously affecting the authenticity and credibility of digital human interaction.

[0005] Therefore, a digital human speech lip synchronization control system based on dynamic fusion of multimodal features is proposed. Summary of the Invention

[0006] The object of the present invention is to provide a digital human voice and lip synchronization control system based on dynamic fusion of multimodal features, which uses the user's historical speech data; performs feature recognition on the historical speech data to obtain historical voice features and historical lip features, and constructs a personalized voice file; and constructs a voice mapping model based on the historical voice features and historical lip features; receives data input by the user to obtain first data; performs feature recognition based on the first data to obtain first data features; generates second data based on the voice mapping model and the first data features; constructs a mapping correction model to correct the second data and outputs third control parameters; and realizes synchronous control of the digital human voice and lip through the third control parameters.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] The digital human voice lip synchronization control system based on dynamic fusion of multimodal features includes:

[0009] A data input module, configured to receive data input by a user and obtain first data;

[0010] The archive construction module obtains the user's historical speech data; performs feature recognition on the historical speech data to obtain historical voice features and historical lip shape features, and constructs a personalized voice archive; and constructs a voice mapping model based on the historical voice features and historical lip shape features;

[0011] a correction and adjustment module that performs feature recognition based on the first data to obtain first data features; generates second data based on the speech mapping model and the first data features; constructs a mapping correction model to correct the second data, including distinguishing standard pattern segments from non-standard pattern segments of the second data, adapting the standard pattern segments, and correcting the non-standard pattern segments, and outputting a third control parameter;

[0012] The data output module generates fused output data based on the second data and the third control parameter.

[0013] The historical speech features include the user's acoustic features, prosodic patterns, speech rate features, and pronunciation statistics of phonemes;

[0014] The historical lip shape features include the three-dimensional motion trajectory of key lip points, the opening and closing amplitude of the mouth shape, the personalized distribution of speed, and the user's co-articulation visual pattern in continuous speech;

[0015] The speech mapping model is obtained based on deep learning network training and is used to predict lip feature sequences.

[0016] The first data includes video data; performing feature recognition on the first data to obtain first data features, the first data features including voice features and video features;

[0017] The first data feature is input into a speech mapping model for recognition, wherein the speech mapping model generates a lip shape feature sequence based on the speech feature in the first data feature.

[0018] Then, the video feature in the first data feature is recognized to obtain a first lip shape feature;

[0019] The second data is generated by fitting the lip shape feature sequence and the first lip shape feature and then combining the speech feature.

[0020] The mapping correction model is trained based on historical speech data and historical fusion data, and includes content anomaly recognition, matching anomaly recognition, and fluency anomaly recognition;

[0021] The content anomaly identification is based on the speech text content, speech features and video features, and the content anomaly data is obtained by analyzing the context of the speech text content and the anomaly of the speech features and video features;

[0022] The matching anomaly identification is obtained based on the corresponding matching degree of the voice features and the lip shape features in the historical speech data and the historical fusion data; it is used to analyze the synchronization matching degree of the voice features and the lip shape features, and identify the data that does not meet the synchronization matching requirements as matching anomaly data;

[0023] The fluency anomaly recognition is based on the fluency of the speech and lip shape in the historical speech data and the historical fusion data; it is used to analyze the fluency of the speech and lip shape in the generated fusion data and identify the data that does not meet the fluency requirements as fluency anomaly data;

[0024] The second data is divided based on the content abnormal data, the matching abnormal data and the fluency abnormal data to obtain standard pattern segments and non-standard pattern segments.

[0025] extracting lip-shape dynamic features and synchronized speech features from the second data;

[0026] Analyzing content anomaly data, matching anomaly data, and fluency anomaly data between the lip dynamic features and the synchronized speech features using a mapping correction model to distinguish between standard pattern segments and non-standard pattern segments;

[0027] For standard mode segments, adaptive control parameters are generated based on personalized voice files to enhance personalized style;

[0028] For non-standard pattern segments, generating correction control parameters for correcting the non-standard pattern segments based on the personalized voice profile;

[0029] A third control parameter is obtained based on the adapted control parameter and the corrected control parameter.

[0030] The data output module generates fused output data when the second data and the third control parameter are combined;

[0031] In combination with the expression linkage rules in the personalized voice file, the facial expression area other than the lip shape in the second data is adjusted, and the audio and video synchronization accuracy in the second data is adjusted.

[0032] The profile building module obtains the user's explicit rating of the fused output data generated by the data output module, and obtains implicit feedback by analyzing the user's behavior pattern of repeatedly correcting and adjusting the lip sync segment;

[0033] Based on explicit scoring and implicit feedback, the personalized voice archive and the voice mapping model are continuously learned and iteratively optimized.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] 1. By using historical speech features and historical lip shape features, the present invention enables the archive construction module and the speech mapping model to learn based on more comprehensive and detailed data, thereby more accurately capturing the user's pronunciation and lip shape habits, and improving the richness and accuracy of the model input.

[0036] 2. This technical solution simultaneously utilizes audio features and lip-shape features directly extracted from the video. Based on the quality of speech features and video features and the user's demand tendencies, it fits the lip-shape information from these two sources, making the generated data closer to the actual situation. In particular, when the audio information is insufficient or ambiguous, the video information can provide a powerful supplement.

[0037] 3. Through the mapping correction model, the lip shape is identified in multiple dimensions such as content anomalies, matching anomalies, and fluency anomalies, which can more comprehensively and accurately locate various problems that may occur in the mouth shape generation process; the lip shape is divided into standard mode and non-standard mode, so that subsequent correction adjustments can adopt differentiated strategies; the non-standard mode segments with real problems are corrected in particular, while the standard mode segments that conform to user habits may be style-maintained or enhanced; the lip shape correction is no longer a blind adjustment, but is based on intelligent judgment and classification based on the understanding of the anomaly type, which improves the efficiency and effect of the correction and ensures the quality and consistency of the output lip shape. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 Schematic diagram of the structure of the digital human voice lip synchronization control system based on dynamic fusion of multimodal features of the present invention;

[0039] Figure 2 Schematic diagram of the structure of the mapping correction model of the present invention;

[0040] Figure 3 This is a schematic diagram of the fusion output data acquisition process of the present invention. DETAILED DESCRIPTION

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0042] Example 1:

[0043] The present invention proposes a digital human voice lip synchronization control system based on dynamic fusion of multimodal features, the structure of which is as follows: Figure 1 Shown, including:

[0044] A data input module, configured to receive data input by a user and obtain first data;

[0045] The archive construction module obtains the user's historical speech data; performs feature recognition on the historical speech data to obtain historical voice features and historical lip shape features, and constructs a personalized voice archive; and constructs a voice mapping model based on the historical voice features and historical lip shape features;

[0046] a correction and adjustment module that performs feature recognition based on the first data to obtain first data features; generates second data based on the speech mapping model and the first data features; constructs a mapping correction model to correct the second data, including distinguishing standard pattern segments from non-standard pattern segments of the second data, adapting the standard pattern segments, and correcting the non-standard pattern segments, and outputting a third control parameter;

[0047] The data output module generates fused output data based on the second data and the third control parameter.

[0048] The historical speech features include the user's acoustic features, prosodic patterns, speech rate features, and pronunciation statistics of phonemes;

[0049] The historical lip shape features include the three-dimensional motion trajectory of key lip points, the opening and closing amplitude of the mouth shape, the personalized distribution of speed, and the user's co-articulation visual pattern in continuous speech;

[0050] The speech mapping model is obtained based on deep learning network training and is used to predict lip feature sequences.

[0051] The acoustic features include Mel-frequency cepstral coefficients, pitch, energy and zero-crossing rate;

[0052] Mel-frequency cepstral coefficients are usually extracted with 13-39 dimensions (including first-order and second-order differences), which can effectively characterize the short-term power spectrum envelope of the speech signal and have a high degree of phoneme discrimination; pitch, which describes the pitch of the sound, is related to rhythm and emotion, and indirectly affects the mouth shape. For example, the rise in pitch at the end of a question sentence may be accompanied by a specific mouth shape or eyebrow movement; energy, the amplitude of the speech signal, is related to the strength of pronunciation and affects the degree of opening and closing of the mouth shape; zero-crossing rate, the number of times the signal crosses the zero point per unit time, plays a certain role in distinguishing between clear and voiced sounds.

[0053] The prosodic pattern includes tone length, pauses, and intonation contours;

[0054] Sound length refers to the duration of phonemes, syllables, and words. Different sound lengths will affect the time the mouth shape is maintained; pauses refer to the silent period between sentences, corresponding to the closed or natural state of the mouth shape; intonation contour refers to the pattern of intonation changes in a sentence, such as rising tone, falling tone, and flat tone, which carries semantics and emotions.

[0055] The speech rate features include the number of syllables per second and the number of phonemes per second, which directly affect the speed and clarity of mouth movements.

[0056] The pronunciation statistical characteristics of the phonemes include the mean and variance of the pronunciation duration of the personalized phonemes, and the average pronunciation pattern of a specific phoneme sequence.

[0057] Acoustic features are the foundation, prosody and speaking rate features reflect speaking style, and phoneme statistics are more inclined to personalized pronunciation details; these features are the cornerstones for training speech mapping models and building personalized voice profiles, and together constitute a comprehensive description of the user's pronunciation method.

[0058] The three-dimensional motion trajectory of the lip key points usually selects several key points on the lip contour, such as the 8+8 points of the inner and outer lips defined by MPEG-4FDP, or the driving vertices implicit in the lip-related key points in ARKit's 52 BlendShapes, and records the three-dimensional coordinate sequence of these points in each frame, directly describing the dynamic shape changes of the lips.

[0059] The opening and closing range of the mouth includes: the maximum distance between the upper and lower lips, and the distance between the left and right corners of the mouth.

[0060] The personalized distribution of speed includes: the displacement speed of key points and the rate of change of the opening and closing amplitude.

[0061] The statistical distribution (mean, variance, range) of these parameters in different phonemes and different contexts is calculated to form personalized parameters that reflect the habits and dynamic characteristics of the user's mouth opening and closing when speaking.

[0062] Visual patterns of co-articulation in continuous speech: Co-articulation refers to the fact that the pronunciation of a phoneme is influenced by the preceding and following phonemes, resulting in changes in the standard mouth shape. By analyzing mouth shape data from large amounts of continuous speech, the user's stable, habitual mouth shape transition patterns and specific shapes for specific phoneme combinations are identified. This is key to natural mouth shape. Simply splicing the mouth shapes of isolated phonemes can appear mechanical. Capturing co-articulation patterns can greatly enhance fluency and realism.

[0063] The introduction of the collaborative pronunciation visual model helps to solve the unnatural feeling caused by the splicing of single phoneme mouth shapes, making the transition of digital people's mouth shapes when speaking continuously smoother and more in line with the pronunciation rules of real humans.

[0064] The speech mapping model uses a sequence-to-sequence architecture based on deep learning. This embodiment builds a speech mapping model based on the Transformer. The encoder receives a speech feature sequence and captures speech context through a multi-head self-attention mechanism and a feedforward network. The decoder then autoregressively generates the current lip shape features based on the encoder output and the predicted lip shape features from the previous moment. The model parameters are optimized using a backpropagation algorithm, enabling the model to accurately predict lip shape features based on the input speech features. The model uses mean squared error as the loss function for optimization.

[0065] By using historical speech features and historical lip shape features, the present invention enables the archive construction module and the speech mapping model to learn based on more comprehensive and detailed data, thereby more accurately capturing the user's pronunciation and lip shape habits, and improving the richness and accuracy of the model input.

[0066] The first data includes video data; performing feature recognition on the first data to obtain first data features, the first data features including voice features and video features;

[0067] The first data feature is input into a speech mapping model for recognition, wherein the speech mapping model generates a lip shape feature sequence based on the speech feature in the first data feature.

[0068] Then, the video feature in the first data feature is recognized to obtain a first lip shape feature;

[0069] The second data is generated by fitting the lip shape feature sequence and the first lip shape feature and then combining the speech feature.

[0070] In the process of fitting and generating the second data, based on the user's mode selection, the lip shape feature sequence and the video lip shape feature are fitted through a neural network.

[0071] The user's mode selections include: audio priority mode, video priority mode, balanced mode, and dynamic adaptive mode.

[0072] Audio-first mode: This mode is used when the video quality is low (e.g., poor lighting, occlusion, low resolution) or when the user trusts the naturalness of audio-driven lip syncing.

[0073] Video priority mode: When the video quality is high, or the user wants the lip movements to be as faithful to the actual lip movements in the video as possible.

[0074] Balanced Mode: Balances the coherence of audio drivers and the realism of video drivers.

[0075] Dynamic Adaptive Mode: In dynamic adaptive mode, the system continuously monitors characteristics such as the audio signal's signal-to-noise ratio and energy distribution, as well as indicators such as the stability of face detection in the video, and the smoothness and confidence of key point tracking. Based on these real-time evaluation results, a decision model or preset rules are used to dynamically determine whether the current moment should prioritize audio-driven lip shape or video-extracted lip shape, or to determine the optimal ratio for the fusion of the two. For example, when the audio signal is detected to be clear and expressive, but the video lighting is poor or the face is obscured, the system can automatically increase the weight of the audio-driven lip shape features; otherwise, the weight of the video lip shape features is increased. This decision model can be trained using historical data and aims to maximize the naturalness and accuracy of the output lip shape.

[0076] This technical solution simultaneously utilizes audio features and lip shape features directly extracted from the video. Based on the quality of voice features and video features and the user's demand tendencies, it fits the lip shape information from these two sources, making the generated data closer to the actual situation. In particular, when the audio information is insufficient or ambiguous, the video information can provide a powerful supplement.

[0077] The mapping correction model is trained based on historical speech data and historical fusion data, and its structure is constructed based on a deep learning classifier (such as MLP, CNN). It is used to classify each lip segment (or frame) in the second data as a standard mode segment or a non-standard mode segment, and further identify the types of content abnormalities, matching abnormalities, and fluency abnormalities.

[0078] Its structure is as follows Figure 2 As shown, it includes content anomaly recognition, matching anomaly recognition, and fluency anomaly recognition;

[0079] The content anomaly identification is based on the speech text content context analysis of historical speech data and historical fusion data, and the anomaly analysis of speech features and video features to identify content anomaly data;

[0080] The matching anomaly identification is obtained based on the corresponding matching degree of voice features and lip features in the historical speech data and the historical fusion data, and is used to analyze the synchronization matching degree of the voice features and lip features, and identify the data that does not meet the synchronization matching requirements as matching anomaly data;

[0081] The fluency anomaly recognition is based on the fluency of the speech and lip shape in the historical speech data and the historical fusion data, and is used to analyze the fluency of the speech and lip shape in the generated fusion data, and identify the data that does not meet the fluency requirements as fluency anomaly data;

[0082] The second data is divided based on the content abnormal data, the matching abnormal data and the fluency abnormal data to obtain standard pattern segments and non-standard pattern segments.

[0083] Historical fusion data refers to the use of historical user speech data to generate corresponding lip animation data (such as a lip feature sequence) through an initial version or pre-trained "speech mapping model." The generated lip animation data will be evaluated and annotated, including content anomaly labels, matching anomaly labels, and fluency anomaly labels.

[0084] Content Anomaly Tags: These tags mark segments where the lip shapes clearly don't match the speech content, and where there are discrepancies between the historical speech data and the historical fusion data. For example, the pronunciation is / a / but the lip shape is closed. Match Anomaly Tags: These tags mark segments where the lip shapes are out of sync with the speech. For example, the lip movements are ahead of or behind the corresponding pronunciation. Fluency Anomaly Tags: These tags mark segments where the lip sequences themselves are unnatural, jittery, or jumpy.

[0085] By mapping the correction model to perform multi-dimensional identification of lip shape anomalies such as content anomalies, matching anomalies, and fluency anomalies, it is possible to more comprehensively and accurately locate various problems that may arise in the lip shape generation process; the lip shapes are divided into standard modes and non-standard modes, so that subsequent correction adjustments can adopt differentiated strategies; non-standard mode segments that really have problems are corrected, while standard mode segments that conform to user habits may be style-maintained or enhanced; lip shape correction is no longer a blind adjustment, but is based on intelligent judgment and classification based on the understanding of the anomaly type, which improves the efficiency and effectiveness of correction and ensures the quality and consistency of the output lip shape.

[0086] extracting lip-shape dynamic features and synchronized speech features from the second data;

[0087] Analyzing content anomaly data, matching anomaly data, and fluency anomaly data between the lip dynamic features and the synchronized speech features using a mapping correction model to distinguish between standard pattern segments and non-standard pattern segments;

[0088] For standard pattern segments, adaptive control parameters are generated based on personalized voice files to enhance the personalized style of speech lip shape;

[0089] For non-standard pattern segments, generating correction control parameters for correcting the speech lip shape of the non-standard pattern segments based on the personalized voice file;

[0090] A third control parameter is obtained based on the adapted control parameter and the corrected control parameter.

[0091] When the mapping correction model identifies the current lip shape as belonging to a standard pattern segment, the system generates adaptive control parameters to enhance the personalized style. The personalized style is derived from the statistical analysis results of the above-mentioned historical voice features and historical lip shape features stored in the personalized voice archive.

[0092] Generation and application of adaptive control parameters:

[0093] Mouth-Shaping Amplitude Adjustment Based on Speech Intensity: The profile records the average range of the user's mouth opening and closing at different speech energies and intensities. If the current speech energy is high and the mouth shape is standard, the adaptation parameters may slightly increase the mouth opening and closing to better match the user's excited or stressed speaking style.

[0094] Adjusting the smoothness and clarity of lip movements based on speaking speed: The profile records the rate of change and clarity of the user's lip movements at different speaking speeds (for example, lip movements may be simplified at fast speeds and fuller at slow speeds). If the current speaking speed is slow, the adaptive parameters can make the lip movements smoother and fuller; if the speaking speed is fast, the lip movement transitions of some non-critical phonemes may be appropriately simplified to maintain overall fluency and conform to the user's habits at fast speaking speeds.

[0095] Personalized exaggerated / restricted mouth shape for specific phonemes: The profile analyzes whether the user has special pronunciation habits (more exaggerated or more restrained than the average person) for certain specific phonemes (such as the degree of lip rounding of the vowel / o / or the width of the mouth opening of / i / ).

[0096] The adaptation parameter can be a scaling factor applied to the standard mouth shape features of the corresponding phoneme to match the user's personalized preferences. For example, if the user habitually pouts when pronouncing / u / , the adaptation parameter will increase the Blendshape value of mouthFunnel or mouthPucker.

[0097] Personalized enhancement of co-articulation patterns: User-specific co-articulation visual patterns are stored in the profile. When a corresponding phoneme sequence is recognized, the adaptation parameters guide the lip animation to favor the user's accustomed co-articulation pattern rather than a generic or average pattern.

[0098] The correction process of the correction control parameters includes:

[0099] When the mapping correction model identifies a lip-sync segment as non-standard, it identifies the cause of the mismatch; for example, if the lip animation lags behind the speech features by 50 milliseconds, it generates a correction control parameter: {Type: 'Timing Adjustment', Value: 'Advance by 50ms', Target: 'Lip-Sync Dynamic Feature Sequence'}. The data output module applies this parameter by resampling the lip-sync dynamic feature sequence in the second data or adjusting its keyframe timestamps to bring the lip animation 50 milliseconds ahead on the timeline, thereby aligning it with the speech features.

[0100] If slight 'smoothness anomalies' are detected (e.g., high-frequency jitter in a specific Blendshape weight), the correction control parameters may also include {Type: 'Smoothing Filter', Target: 'Blendshape_X Weight Curve', Algorithm: 'Gaussian Smoothing', Window Size: '3 Frames'}, and the data output module will apply Gaussian smoothing to the specified Blendshape weight sequence.

[0101] This solution generates adaptive control parameters for standard pattern segments to enhance personalized style, and generates correction control parameters for non-standard pattern segments to modify them. This differentiated processing strategy not only effectively corrects lip errors and improves the accuracy of lip animation, but also preserves and strengthens the user's unique pronunciation and lip style. By enhancing personalized style, the digital human's lip shape is more in line with the user's own characteristics and more expressive.

[0102] The data output module generates fused output data when the second data and the third control parameter are combined;

[0103] Combined with the expression linkage rules in the personalized voice file, the facial expression area in the second data except the lip shape is adjusted, and the audio and video synchronization accuracy in the second data is adjusted. The acquisition process of the fusion output data is as follows: Figure 3 shown.

[0104] Expression linkage rules are obtained in the following ways:

[0105] Statistical learning: Analyze historical user audio and video data to identify patterns in the co-occurrence of speech features (such as specific pitch patterns, emotional overtones, and keywords) and non-lip-forming facial expressions (such as eyebrows, eyes, and forehead). For example, this can be achieved through methods such as association rule mining.

[0106] Template-based + personalized fine-tuning: Provides a set of universal expression linkage templates (such as raising eyebrows when surprised and frowning when questioning), allowing users to make personalized fine-tuning with a small amount of annotation or feedback.

[0107] When the expression linkage rules are triggered, the system will generate corresponding control parameters based on the rules and act on the Blendshapes or bone control points in these areas to achieve more comprehensive facial expressions that are synchronized with the lip movements.

[0108] Optimization of audio and video synchronization accuracy includes:

[0109] Timestamp alignment: Speech feature extraction and lip shape feature generation are strictly based on unified timestamps to ensure frame-level correspondence.

[0110] Model prediction delay compensation: Deep learning model inference has a certain delay. For real-time applications, this delay needs to be considered or compensated in the design of speech mapping models and mapping correction models.

[0111] Matching anomaly correction: When the mapping correction model identifies a matching anomaly (lip-syncing ahead or behind), the third control parameter contains instructions for fine-tuning the timing of the lip-syncing sequence (for example, by interpolating or deleting certain transition frames, or adjusting the timing nodes of the lip-syncing animation curve), and dynamically adjusts the facial expression based on the expression linkage rules.

[0112] Applying the third control parameter to modify lip-sync animation and optimizing audio-visual synchronization ensures that the results of the correction and adjustment module are effectively reflected in the final output, directly improving the viewing experience of lip-sync animation. By introducing expression linkage rules and adjusting other facial expression areas besides the lip (such as eyebrows, eyes, cheeks, etc.), the digital human's facial expressions are no longer limited to the mouth, but instead present a more natural, coordinated, and emotional overall dynamic, greatly enhancing the digital human's vividness and expressiveness.

[0113] The profile building module obtains the user's explicit rating of the fused output data generated by the data output module, and obtains implicit feedback by analyzing the user's behavior pattern of repeatedly correcting and adjusting specific lip segments;

[0114] Based on explicit scoring and implicit feedback, the personalized voice archive and the voice mapping model are continuously learned and iteratively optimized.

[0115] The implicit feedback acquisition process includes identifying the user's behavioral characteristics:

[0116] Repeated generation: The user inputs the same text or voice and repeatedly clicks the "Generate Lip Shape" button.

[0117] Manual editing frequency / amplitude: If the system provides lip-sync editing tools (such as adjusting keyframes and modifying Blendshape curves), record which segments and parameters the user frequently or significantly modifies.

[0118] Playback control behavior: repeatedly slow down, pause, and view a specific clip frame by frame.

[0119] Version backtracking / selection: If the system generates multiple lip sync candidates for the same input (possibly due to different random seeds or fine-tuning parameters), record which version the user ultimately selected and which versions were discarded.

[0120] Abandoned operation: After the user inputs content, generates lip shape, but quickly deletes the segment or re-enters the content.

[0121] Count the phoneme combinations, words, speaking speeds, or emotional expressions that generate the most implicit negative user feedback (e.g., repetitive speech and extensive editing). These areas are considered "difficult samples" where the current model performs poorly. These difficult samples and their contextual information are given higher weights and added to the "historical fusion data" for targeted training, or used to guide the "mapping correction model" to learn more refined correction strategies.

[0122] Personalized profile update: If a user repeatedly adjusts a certain mouth shape to a specific pattern, this may reflect their deeper personalized preferences and can be used to update the co-articulation pattern or mouth shape parameter distribution in the "Personalized Voice Profile".

[0123] By introducing explicit user ratings and implicit feedback mechanisms, the system is equipped with the ability to continuously learn and evolve. Models and personalized profiles are no longer static; instead, they can be iteratively optimized based on user usage and feedback. This allows the system to increasingly meet users' personalized needs, thereby improving user satisfaction.

[0124] Example 2:

[0125] The digital human voice lip synchronization control system based on dynamic fusion of multimodal features includes:

[0126] A data input module, configured to receive data input by a user and obtain first data;

[0127] The archive construction module obtains the user's historical speech data; performs feature recognition on the historical speech data to obtain historical voice features and historical lip shape features, and constructs a personalized voice archive; and constructs a voice mapping model based on the historical voice features and historical lip shape features;

[0128] a correction and adjustment module that performs feature recognition based on the first data to obtain first data features; generates second data based on the speech mapping model and the first data features; constructs a mapping correction model to correct the second data, including distinguishing standard pattern segments from non-standard pattern segments of the second data, adapting the standard pattern segments, and correcting the non-standard pattern segments, and outputting a third control parameter;

[0129] The data output module generates fused output data based on the second data and the third control parameter.

[0130] The historical speech features include the user's acoustic features, prosodic patterns, speech rate features, and pronunciation statistics of phonemes;

[0131] The historical lip shape features include the three-dimensional motion trajectory of key lip points, the opening and closing amplitude of the mouth shape, the personalized distribution of speed, and the user's co-articulation visual pattern in continuous speech;

[0132] The speech mapping model is obtained based on deep learning network training and is used to predict lip feature sequences.

[0133] By using historical speech features and historical lip shape features, the present invention enables the archive construction module and the speech mapping model to learn based on more comprehensive and detailed data, thereby more accurately capturing the user's pronunciation and lip shape habits, and improving the richness and accuracy of the model input.

[0134] The first data includes video data; performing feature recognition on the first data to obtain first data features, the first data features including voice features and video features;

[0135] Inputting the first data feature into a speech mapping model for recognition, the speech mapping model generating a lip shape feature sequence based on the speech feature in the first data feature;

[0136] Then, the video feature in the first data feature is recognized to obtain a first lip shape feature;

[0137] The second data is generated by fitting the lip shape feature sequence and the first lip shape feature and then combining the speech feature.

[0138] This technical solution simultaneously utilizes audio features and lip shape features directly extracted from the video. Based on the quality of voice features and video features and the user's demand tendencies, it fits the lip shape information from these two sources, making the generated data closer to the actual situation. In particular, when the audio information is insufficient or ambiguous, the video information can provide a powerful supplement.

[0139] The mapping correction model is trained based on historical speech data and historical fusion data, and includes content anomaly recognition, matching anomaly recognition, and fluency anomaly recognition;

[0140] The content anomaly identification is based on the speech text content, speech features and video features, and the content anomaly data is obtained by analyzing the context of the speech text content and the anomaly of the speech features and video features;

[0141] The matching anomaly identification is obtained based on the corresponding matching degree of the voice features and the lip shape features in the historical speech data and the historical fusion data; it is used to analyze the synchronization matching degree of the voice features and the lip shape features, and identify the data that does not meet the synchronization matching requirements as matching anomaly data;

[0142] The fluency anomaly recognition is based on the fluency of the speech and lip shape in the historical speech data and the historical fusion data; it is used to analyze the fluency of the speech and lip shape in the generated fusion data and identify the data that does not meet the fluency requirements as fluency anomaly data;

[0143] The second data is divided based on the content abnormal data, the matching abnormal data and the fluency abnormal data to obtain standard pattern segments and non-standard pattern segments.

[0144] The present invention uses a mapping correction model to perform multi-dimensional recognition of content anomalies, matching anomalies, and fluency anomalies on lip shapes, and can locate various problems that may occur in the process of lip shape generation in a more comprehensive and accurate manner; the lip shapes are divided into standard modes and non-standard modes, so that subsequent correction adjustments can adopt differentiated strategies; non-standard mode segments with real problems are corrected in a focused manner, and standard mode segments that conform to user habits may be style-maintained or enhanced; lip shape correction is no longer a blind adjustment, but is based on intelligent judgment and classification based on the understanding of the anomaly type, thereby improving the efficiency and effect of correction and ensuring the quality and consistency of the output lip shapes.

[0145] extracting lip-shape dynamic features and synchronized speech features from the second data;

[0146] Analyzing content anomaly data, matching anomaly data, and fluency anomaly data between the lip dynamic features and the synchronized speech features using a mapping correction model to distinguish between standard pattern segments and non-standard pattern segments;

[0147] For standard mode segments, adaptive control parameters are generated based on personalized voice files to enhance personalized style;

[0148] For non-standard pattern segments, generating correction control parameters for correcting the non-standard pattern segments based on the personalized voice profile;

[0149] A third control parameter is obtained based on the adapted control parameter and the corrected control parameter.

[0150] The present invention generates adaptive control parameters for standard pattern segments to enhance personalized style, and generates correction control parameters for non-standard pattern segments to modify them. This differentiated processing strategy not only effectively corrects lip errors and improves the accuracy of lip animation, but also preserves and strengthens the user's unique pronunciation and lip style. By enhancing personalized style, the digital human's lip shape better matches the user's characteristics and is more expressive.

[0151] The data output module generates fused output data when the second data and the third control parameter are combined;

[0152] In combination with the expression linkage rules in the personalized voice file, the facial expression area other than the lip shape in the second data is adjusted, and the audio and video synchronization accuracy in the second data is adjusted.

[0153] This invention applies the third control parameter to modify lip-sync animation and optimizes audio-visual synchronization accuracy, ensuring that the results of the correction and adjustment module are effectively reflected in the final output, directly improving the viewing experience of lip-sync animation. By introducing expression linkage rules and adjusting other facial expression areas besides the lip (such as eyebrows, eyes, cheeks, etc.), the digital human's facial expressions are no longer limited to the mouth, but instead present a more natural, coordinated, and emotional overall dynamic, greatly enhancing the digital human's vividness and expressiveness.

[0154] The profile building module obtains the user's explicit rating of the fused output data generated by the data output module, and obtains implicit feedback by analyzing the user's behavior pattern of repeatedly correcting and adjusting the lip sync segment;

[0155] Based on explicit scoring and implicit feedback, the personalized voice archive and the voice mapping model are continuously learned and iteratively optimized.

[0156] By introducing explicit user ratings and implicit feedback mechanisms, this invention enables the system to continuously learn and evolve. Models and personalized profiles are no longer static; instead, they can be iteratively optimized based on user usage and feedback. This allows the system to increasingly better meet users' personalized needs, thereby improving user satisfaction.

[0157] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A digital human voice lip synchronization control system based on dynamic fusion of multimodal features, characterized by: include: A data input module, configured to receive data input by a user and obtain first data; the first data includes video data; The archive construction module obtains the user's historical speech data; performs feature recognition on the historical speech data to obtain historical voice features and historical lip shape features, and constructs a personalized voice archive; and constructs a voice mapping model based on the historical voice features and historical lip shape features; a correction and adjustment module, performing feature recognition based on the first data to obtain first data features; generating second data based on the speech mapping model and the first data features; constructing a mapping correction model to correct the second data, including distinguishing between standard pattern segments and non-standard pattern segments of the second data, adapting the standard pattern segments, and correcting the non-standard pattern segments, and outputting a third control parameter; The first data features include voice features and video features; The second data acquisition process includes: inputting the first data features into a speech mapping model for recognition, wherein the speech mapping model generates a lip shape feature sequence based on the speech features in the first data features; then recognizing the video features in the first data features to obtain a first lip shape feature; fitting the lip shape feature sequence and the first lip shape feature, and then combining the lip shape feature with the speech features to generate the second data; The process of obtaining the third control parameter includes: extracting lip-shape dynamic features and synchronized speech features from the second data; Using the mapping correction model, we analyze the content anomaly data, matching anomaly data, and fluency anomaly data between the lip dynamic features and the synchronized speech features to distinguish between standard and non-standard pattern segments. For standard mode segments, adaptive control parameters are generated based on personalized voice files to enhance personalized style; For non-standard pattern segments, generating correction control parameters for the lip shape of the non-standard pattern segments based on the personalized voice archive; obtaining a third control parameter based on the adapted control parameter and the corrected control parameter; The data output module generates fused output data based on the second data and the third control parameter.

2. The digital human voice lip-sync control system based on dynamic fusion of multimodal features according to claim 1, characterized in that: The historical speech features include the user's acoustic features, prosodic patterns, speech rate features, and pronunciation statistics of phonemes; The historical lip shape features include the three-dimensional motion trajectory of key lip points, the opening and closing amplitude of the mouth shape, the personalized distribution of speed, and the user's co-articulation visual pattern in continuous speech; The speech mapping model is obtained based on deep learning network training and is used to predict lip feature sequences.

3. The digital human voice lip-sync control system based on dynamic fusion of multimodal features according to claim 1, characterized in that: The mapping correction model is trained based on historical speech data and historical fusion data, and includes content anomaly recognition, matching anomaly recognition, and fluency anomaly recognition; The content anomaly identification is based on the speech text content context analysis of historical speech data and historical fusion data, and the anomaly analysis of speech features and video features to identify content anomaly data; The matching anomaly recognition is based on the learning and training of the corresponding matching degree of voice features and lip features in the historical speech data and the historical fusion data; it is used to analyze the synchronization matching degree of voice features and lip features, and identify the data that does not meet the synchronization matching requirements as matching anomaly data; The fluency anomaly recognition is based on the fluency of speech and lip shape in historical speech data and historical fusion data. It is used to analyze the fluency of speech and lip shape in the generated fusion data, and identify data that does not meet the fluency requirements as fluency abnormal data; The second data is divided based on the content abnormal data, the matching abnormal data and the fluency abnormal data to obtain standard pattern segments and non-standard pattern segments.

4. The digital human voice lip-sync control system based on dynamic fusion of multimodal features according to claim 1, characterized in that: The data output module generates fused output data when the second data and the third control parameter are combined; In combination with the expression linkage rules in the personalized voice file, the facial expression area excluding the lip shape in the second data is adjusted, and the audio and video synchronization accuracy in the second data is adjusted.

5. The digital human voice lip-sync control system based on dynamic fusion of multimodal features according to claim 1, characterized in that: The profile building module obtains the user's explicit rating of the fused output data generated by the data output module, and obtains implicit feedback by analyzing the user's behavior pattern of repeatedly correcting and adjusting the lip sync segment; Based on explicit scoring and implicit feedback, the personalized voice archive and the voice mapping model are continuously learned and iteratively optimized.

Citation Information

Patent Citations

  • Dynamic tweet intelligent generation method and system based on visual feature recognition

    CN120277280A