Audio-driven talking-head video generation method based on 3D gaussian splatting and proxy attention mechanism

By combining a multi-resolution three-plane hash grid with an efficient spatial-audio attention module, the problem of low generation efficiency in existing technologies is solved, and high-fidelity dynamic facial expressions and natural and smooth video effects are generated efficiently.

CN122415807APending Publication Date: 2026-07-17XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XINJIANG UNIVERSITY
Filing Date
2025-01-16
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing audio-driven speech head video generation methods struggle to maintain facial region cohesion during dynamic changes in facial expressions. Furthermore, Gaussian function operations involve a large number of parameters and complex high-dimensional spaces, resulting in low generation efficiency and making it difficult to preserve facial details while ensuring efficient generation.

Method used

A static 3D Gaussian representation is generated using a multi-resolution three-plane hash grid. An efficient spatial-audio attention module is used to fuse audio and spatial features. Dynamic facial animations are generated through a Gaussian warp decoder. Finally, a generative adversarial network based on face priors is used for video enhancement.

Benefits of technology

It improves the accuracy and rendering speed of dynamic facial expressions in generated talking head videos, ensuring the fidelity of facial details and the natural smoothness of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415807A_ABST
    Figure CN122415807A_ABST
Patent Text Reader

Abstract

本发明涉及计算机视觉和语音合成领域,特别是一种基于高效空间‑音频注意力机制的端到端高保真说话头生成方法。该方法解决了现有音频驱动的动态面部生成技术中,面部表情同步和渲染速度较慢的问题。通过使用空间结构编码器提取视频帧中的空间特征,结合高效的空间‑音频注意力模块将音频特征与空间特征融合,生成与音频信号同步的动态面部动画。进一步通过高斯变形解码器对每帧图像进行面部变形,提升了生成的面部细节和视频质量。该方法能够高效生成高质量的说话头视频,广泛应用于数字人、虚拟化身、电影制作和虚拟会议等领域。
Need to check novelty before this filing date? Find Prior Art