A streaming speech inference driven digital human lip shape animation synchronization method

By jointly enhancing speech emotion features in the time and frequency domains and introducing an uncertainty modeling mechanism, the problems of insufficient emotional expression and unstable mapping in speech-driven lip animation generation are solved, achieving more realistic emotional expression and higher cross-scene adaptability.

CN122415804APending Publication Date: 2026-07-17BEIJING XILIANLIAN TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING XILIANLIAN TECHNOLOGY CO LTD
Filing Date
2026-04-30
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies for voice-driven lip animation generation suffer from insufficient emotional expression modeling and unstable and undiversified speech-lip mapping, resulting in lip animations that are monotonous and unrealistic in emotional expression and lack cross-speaker generalization ability.

Method used

By jointly enhancing speech emotion features in the time and frequency domains and introducing an uncertainty modeling mechanism based on information entropy, a mapping relationship between speech and lip movement latent space is constructed. A dynamic compensation mechanism is also introduced to adaptively correct deviations in the alignment process, thereby improving the emotional expression ability and diversity, as well as the robustness of the model.

Benefits of technology

It significantly improves the emotional expressiveness and diversity of lip-sync animation, and enhances the model's stability and cross-scene generalization ability in different contexts and with different speakers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415804A_ABST
    Figure CN122415804A_ABST
Patent Text Reader

Abstract

本发明涉及人工智能技术领域,具体是指一种流式语音推理驱动的数字人唇形动画同步方法,包括语音特征编码、语音情感时频增强、不确定性引导特征对齐、强一致性融合建模和动画输出,通过对语音情感特征进行时域与频域联合增强,并引入基于信息熵的不确定性建模机制,对情感特征进行自适应强化,使模型能够捕捉语音中不同情绪状态下的细粒度变化,从而显著提升唇形动画的情感表达能力与多样性;本发明构建语音到唇形运动隐空间的映射关系,并引入动态补偿机制对对齐过程中的偏差进行自适应修正,使唇形隐空间表示能够在不同语境、不同说话人及复杂语音条件下保持稳定性,从而显著提升模型的鲁棒性与跨场景泛化能力。
Need to check novelty before this filing date? Find Prior Art