Audio-driven talking-head video generation method based on 3D gaussian splatting and proxy attention mechanism
By combining a multi-resolution three-plane hash grid with an efficient spatial-audio attention module, the problem of low generation efficiency in existing technologies is solved, and high-fidelity dynamic facial expressions and natural and smooth video effects are generated efficiently.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XINJIANG UNIVERSITY
- Filing Date
- 2025-01-16
- Publication Date
- 2026-07-17
AI Technical Summary
Existing audio-driven speech head video generation methods struggle to maintain facial region cohesion during dynamic changes in facial expressions. Furthermore, Gaussian function operations involve a large number of parameters and complex high-dimensional spaces, resulting in low generation efficiency and making it difficult to preserve facial details while ensuring efficient generation.
A static 3D Gaussian representation is generated using a multi-resolution three-plane hash grid. An efficient spatial-audio attention module is used to fuse audio and spatial features. Dynamic facial animations are generated through a Gaussian warp decoder. Finally, a generative adversarial network based on face priors is used for video enhancement.
It improves the accuracy and rendering speed of dynamic facial expressions in generated talking head videos, ensuring the fidelity of facial details and the natural smoothness of the video.
Smart Images

Figure CN122415807A_ABST