This invention relates to the fields of
computer vision and
speech processing technology, and particularly to an interactive 3D head generation method and apparatus based on interleaved multimodal context. It primarily addresses the problems of decoupling speech and 3D head motion, lack of
context awareness, inconsistent generation results, and insufficient multimodal
information fusion, proposing the following technical solution: Step 1, extracting speech
signal features; Step 2, fusing head motion features with speech features; Step 3, dividing speech features and 3D head motion features into fixed-length segments; Step 4, inputting multimodal context segments to a multilayer
transformer encoder; Step 5, generating a 3D head dynamic sequence; Step 6, adjusting the generation result through a guidance mechanism. This invention unifies the speech and head motion feature spaces to achieve precise
adaptation and fusion, and combines multimodal context modeling and
diffusion-guided generation to improve the
synergy between head dynamics and speech,
context awareness, and the consistency and diversity of generation.