Mouth Shape Generation for Digital Humans Using Speech Speed Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current digital human applications face challenges in accurately generating mouth shapes in face images when dealing with variations in speech speed, leading to issues like 'word missing' and 'liaison' due to the mouth shape not being adequately synchronized with audio data, especially at higher speech speeds.
Innovation Solution
A method and apparatus that utilize audio features, including speech speed and semantic features, to process preset face images and generate accurate mouth shapes by determining the speech speed feature and semantic feature from audio data, and using these features to adjust the mouth shape in the face image, with separate processing for low and high speech speeds to maintain accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If audio-driven mouth shape technology is used to generate face images, then the mouth shape can be synchronized with speech, but accuracy deteriorates when speech speed varies due to word missing and liaison errors
Solution Approach 1:
The patent segments the mouth shape generation process into two distinct processing paths based on speech speed: a first processing path for normal speech speeds and a second processing path for fast speech speeds. This segmentation allows each path to be optimized for its specific speed range, preventing word missing and liaison errors that occur when a single path handles all speeds. The system divides the problem space to maintain precision across varying conditions.
Solution Approach 2:
The patent implements dynamic processing by adjusting the mouth shape generation approach based on the detected speech speed. The system dynamically selects between different processing paths and adjusts processing parameters according to the actual speech speed, enabling adaptive optimization that maintains accuracy whether speech is slow or fast. This dynamic adaptation resolves the contradiction between synchronization reliability and generation precision.
2Device complexity
If a single processing path is used for all speech speeds, then the system complexity is low, but mouth shape accuracy deteriorates at high speech speeds
Solution Approach 1:
The patent segments the processing system into multiple paths with different complexity levels. The first processing path handles normal speech speeds with standard processing, while the second processing path handles fast speech speeds with enhanced processing. This segmentation allows the system to use higher complexity only when necessary, maintaining overall simplicity while achieving high precision when needed.
Solution Approach 2:
The patent changes processing parameters based on speech speed detection. When fast speech is detected, the system activates the second processing path with adjusted parameters optimized for high-speed speech. This parameter adaptation allows the system to maintain low complexity for normal operations while achieving high precision when speech speed requires it, resolving the contradiction between simplicity and accuracy.
3Productivity
If speech speed variations are not considered in processing, then the processing efficiency is high, but mouth shape accuracy deteriorates due to word missing and liaison
Solution Approach 1:
The patent performs preliminary detection of speech speed before proceeding with mouth shape generation. By detecting and categorizing speech speed in advance, the system can select the appropriate processing path beforehand, avoiding the need for complex real-time adjustments during processing. This preliminary action maintains efficiency while ensuring accuracy by preparing the correct processing approach in advance.
Solution Approach 2:
The patent implements dynamic processing that adapts to speech speed variations without sacrificing efficiency. The system dynamically selects processing paths and adjusts parameters based on detected speech speed, enabling efficient processing for normal speeds and enhanced processing only when fast speech is detected. This dynamic approach maintains high overall efficiency while ensuring accuracy when needed.
Data Source
AI summary
The present disclosure provides a mouth shape-based method for generating a face image, a method for training a model, and a device, which relates to the field of artificial intelligence, in particular to the field of cloud computing and digital human. The specific implementation solution is as follows: acquiring audio data to be recognized and a preset face image; determining an audio feature of the audio data to be recognized; where the audio feature includes a speech speed feature and a semantic feature; and performing, according to the speech speed feature and the semantic feature, processing on the preset face image, to generate a face image having a mouth shape.


