A streaming speech inference driven digital human lip shape animation synchronization method
By jointly enhancing speech emotion features in the time and frequency domains and introducing an uncertainty modeling mechanism, the problems of insufficient emotional expression and unstable mapping in speech-driven lip animation generation are solved, achieving more realistic emotional expression and higher cross-scene adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING XILIANLIAN TECHNOLOGY CO LTD
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies for voice-driven lip animation generation suffer from insufficient emotional expression modeling and unstable and undiversified speech-lip mapping, resulting in lip animations that are monotonous and unrealistic in emotional expression and lack cross-speaker generalization ability.
By jointly enhancing speech emotion features in the time and frequency domains and introducing an uncertainty modeling mechanism based on information entropy, a mapping relationship between speech and lip movement latent space is constructed. A dynamic compensation mechanism is also introduced to adaptively correct deviations in the alignment process, thereby improving the emotional expression ability and diversity, as well as the robustness of the model.
It significantly improves the emotional expressiveness and diversity of lip-sync animation, and enhances the model's stability and cross-scene generalization ability in different contexts and with different speakers.
Smart Images

Figure CN122415804A_ABST