A single-sample audio-driven speaker face generation method based on three-dimensional Gaussian splash

CN122415883APending Publication Date: 2026-07-17HEFEI UNIV OF TECH +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI UNIV OF TECH
Filing Date
2026-04-28
Publication Date
2026-07-17

Smart Images

  • Figure CN122415883A_ABST
    Figure CN122415883A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of face generation, and discloses a single-sample audio-driven speaker face generation method based on three-dimensional Gaussian splashing, which comprises the following steps: generating a complete three-dimensional Gaussian head model according to an input single face image through a visible area reconstruction branch and an occlusion area completion branch; extracting audio features of an input audio signal, inputting the audio features into a coarse-grained motion field and a fine-grained motion field respectively, predicting deformation parameters of the three-dimensional Gaussian head model, and fusing the deformation parameters predicted by the coarse-grained motion field and the fine-grained motion field to update the three-dimensional Gaussian head model; rendering the updated three-dimensional Gaussian head model into a two-dimensional image through differentiable Gaussian rasterization, and refining the two-dimensional image through a neural renderer to generate a speaker face video frame. The application can significantly improve the authenticity and stability of the generated video and realize high-efficiency rendering.
Need to check novelty before this filing date? Find Prior Art