A single-sample audio-driven speaker face generation method based on three-dimensional Gaussian splash
CN122415883APending Publication Date: 2026-07-17HEFEI UNIV OF TECH +1
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2026-04-28
- Publication Date
- 2026-07-17
Smart Images

Figure CN122415883A_ABST
Abstract
The application relates to the technical field of face generation, and discloses a single-sample audio-driven speaker face generation method based on three-dimensional Gaussian splashing, which comprises the following steps: generating a complete three-dimensional Gaussian head model according to an input single face image through a visible area reconstruction branch and an occlusion area completion branch; extracting audio features of an input audio signal, inputting the audio features into a coarse-grained motion field and a fine-grained motion field respectively, predicting deformation parameters of the three-dimensional Gaussian head model, and fusing the deformation parameters predicted by the coarse-grained motion field and the fine-grained motion field to update the three-dimensional Gaussian head model; rendering the updated three-dimensional Gaussian head model into a two-dimensional image through differentiable Gaussian rasterization, and refining the two-dimensional image through a neural renderer to generate a speaker face video frame. The application can significantly improve the authenticity and stability of the generated video and realize high-efficiency rendering.
Need to check novelty before this filing date? Find Prior Art