Speech head generation method and apparatus, device, medium, product

By extracting speech and facial motion features, injecting social perception features, and generating a Gaussian displacement-corrected set of three-dimensional Gaussian primitives, the problem of speaker realism and computational efficiency in multi-turn dialogue scenarios is solved. This achieves speaker generation with high visual realism, temporal consistency, and social behavior consistency, thereby enhancing the immersion and fluency of VR social communication.

CN122115652APending Publication Date: 2026-05-29INST OF SOFTWARE - CHINESE ACAD OF SCI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF SOFTWARE - CHINESE ACAD OF SCI
Filing Date
2026-01-13
Publication Date
2026-05-29

Smart Images

  • Figure CN122115652A_ABST
    Figure CN122115652A_ABST
Patent Text Reader

Abstract

The application provides a speaking head generation method, device, equipment, medium and product. The speaking head generation method comprises the following steps: a step of extracting speech features and facial motion features of a speaker and a listener from a multi-round dialogue scene comprising the speaker and the listener, processing the speech features and the facial motion features by using a motion coding network to obtain interaction representation; a step of injecting social perception features into the interaction representation to obtain Gaussian displacement correction and joint representation; a step of generating facial network animation according to the joint representation; a step of correcting the facial network animation according to the Gaussian displacement correction to obtain a dynamic three-dimensional Gaussian primitive set; and a final generation step of bidirectionally generating a speaking head according to the dynamic three-dimensional Gaussian primitive set. Thus, the problems of the contradiction between realism and interactivity, the bottleneck of computational efficiency and practicality, and the lack of social context factors in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and virtual reality technology, specifically to a method, apparatus, computer device, and storage medium for generating speaker heads with high visual realism, temporal consistency, and social behavioral consistency for multi-turn dialogue scenarios. Background Technology

[0002] In recent years, advancements in deep learning and computer graphics have significantly propelled the development of digital human "talking head" generation technology. This technology is increasingly valuable in the field of virtual reality (VR), particularly in immersive social scenarios involving two interacting subjects, where multi-turn dialogue has become a core application. In such multi-turn exchanges, the roles of "speaker" and "listener" dynamically alternate as the dialogue progresses, placing higher demands on the temporal consistency and role-switching capabilities of the generation system. Existing talking head generation methods can be broadly categorized into two representation paradigms: those based on 3D meshes and those based on 2D images.

[0003] In 3D mesh-driven methods, the mainstream approach typically predicts mesh vertices or blendshapes based on a transformer or diffusion model, then drives a 3D head model (such as the FLAME model) to generate speaking head animations. Representative works include FaceFormer, EmoTalk, DiffusionTalker, and Learning to Listen. However, most of these methods only support one-way generation in either a "pure speaker mode" or a "pure listener mode." DualTalk has made preliminary explorations into modeling two-person multi-turn dialogues, but it still struggles to generate realistic textures, accurate colors, and natural dynamic behaviors (such as micro-expressions and subtle movements). This can easily cause "discomfort" in VR applications because the generated results differ perceptibly from the details of real human behavior, leading to physiological discomfort in the immersive experience.

[0004] In contrast, generation methods based on 2D images generate highly realistic talking heads by directly driving RGB faces, effectively overcoming the aforementioned limitations. These methods mainly include two technical approaches: Large-scale diffusion model approach: such as pre-trained models based on Stable Diffusion or Diffusion Transformer (DIT). These models have strong generalization ability, but have a large number of parameters and extremely high computational cost, making them unsuitable for real-time deployment in VR scenarios.

[0005] Highly efficient rendering approaches based on NeRF or 3D Gaussian Splashing (3DGS), such as ER-NeRF, SyncTalk, and GaussianTalker, offer high rendering efficiency, high accuracy, and low computational resource consumption, making them more suitable for VR scenarios. However, these methods are primarily geared towards "speaker-centric" scenarios and have not yet established a dynamic modeling mechanism for dual-agent scenarios in multi-turn dialogues.

[0006] Furthermore, in multi-round social interactions, the social relationship between the two parties significantly influences facial expressions, head posture, and other nonverbal behaviors—a crucial factor generally overlooked in existing technologies. In real-world scenarios, social relationships significantly modulate a character's expressions and reactions: for example, conversations between close friends are typically accompanied by relaxed expressions, frequent smiles, and richer body language; while professional conversations between superiors and subordinates are more restrained, with limited eye contact and relatively formal expressions. These relationship characteristics accumulate over multiple rounds of interaction, directly affecting the naturalness of the character's transitions between speaker and listener, thus determining the realism and fluency of the overall interactive experience. Ignoring this social relationship factor may generate visually plausible speaking heads, but their behavior often contradicts the context, resulting in significant social inconsistency in VR social communication and limiting its practical application.

[0007] In summary, the existing technology has the following main drawbacks: 1. The contradiction between realism and interactivity: 3D mesh-based methods can achieve dynamic multi-turn interactions between speakers and listeners, but the generated results lack realism in texture and color, easily triggering the "uncanny valley" effect in VR environments, causing user discomfort. Meanwhile, speech head generation methods based on 3DGS and other photorealistic techniques are currently limited to "single-person speaking" scenarios and fail to model interactions in multi-turn dialogues.

[0008] 2. Bottlenecks in computational efficiency and practicality: Although 2D methods based on large-scale pre-trained models can generate realistic appearances, their high computational cost and slow inference speed seriously hinder their practical deployment in VR environments.

[0009] 3. Lack of Social Context Factors: In multi-turn dialogue-based social interactions, the social relationships between the speakers (such as closeness or distance, superior-subordinate relationships) have a decisive influence on shaping nonverbal behaviors such as facial expressions and head movements. Existing technologies do not consider this crucial factor, resulting in generated talking heads that may appear visually realistic but exhibit inconsistent social behavior and a lack of natural interactive dynamics, thus limiting their effectiveness in VR social communication.

[0010] Therefore, to solve the above problems, a new speaking head generation technology with high visual realism, temporal consistency and social behavior consistency is needed for multi-turn dialogue scenarios. Summary of the Invention

[0011] The problem to be solved by the present invention This invention was made in consideration of the above problems, and its purpose is to solve the contradiction between realism and interactivity, the bottleneck of computational efficiency and practicality, and the lack of social context factors in the speech head generation technology.

[0012] Methods for solving problems The first embodiment of the present invention relates to a method for generating a speaking head, comprising the following steps: The extraction step involves extracting the speech features and facial motion features of the speaker and listener from a multi-turn dialogue scenario that includes both the speaker and the listener. Then, the action coding network is used to process these features to obtain the interaction representation. The injection step injects socially perceived features into the interactive representation, resulting in Gaussian displacement correction and joint representation. The initial generation step involves generating a facial network animation based on joint representations. The correction step involves correcting the facial network animation based on Gaussian displacement correction to obtain a dynamic 3D Gaussian element set. The final generation step involves bidirectionally generating the speech head based on a dynamic three-dimensional Gaussian meta-set.

[0013] Preferably, the social relationship in the social perception feature is either blood-related or non-blood-related.

[0014] Preferably, the social relations in the social perception features are either equal or unequal.

[0015] Preferably, in the initial generation step, the joint representation is converted into target mesh motion parameters, and then parameter training is performed, in which the mesh motion sequence is introduced as a supervision signal.

[0016] Preferably, an anchor Gaussian-neural Gaussian hierarchical structure is adopted, and social perception features are input.

[0017] Preferably, in the final generation step, joint training is performed in three stages: in the first stage, pre-training of speaker-listener motion generation is completed; in the second stage, pre-training of speaker-listener avatar rendering module is completed; and in the third stage, the speech-mesh-image triplet dataset is used as a unified source of supervision.

[0018] A second embodiment of the present invention relates to a speech head generation device, comprising: The extraction module extracts the speech features and facial motion features of the speaker and listener from a multi-turn dialogue scenario that includes the speaker and listener, and then processes them using an action coding network to obtain the interaction representation. The injection module injects socially perceived features into the interactive representation, resulting in Gaussian displacement correction and joint representation. The initial generation module generates facial network animations based on joint representations; The correction module corrects the facial network animation based on Gaussian displacement correction to obtain a dynamic 3D Gaussian element set; The final generation module generates a speech head bidirectionally based on a dynamic three-dimensional Gaussian meta-set.

[0019] The third embodiment of the present invention relates to a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the speech head generation method according to the first embodiment of the present invention.

[0020] The fourth embodiment of the present invention relates to a computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of the speech head generation method according to the first embodiment of the present invention.

[0021] The fifth aspect of the present invention relates to a computer program product, including a computer program that, when executed by a processor, implements the steps of the speech head generation method of the first aspect.

[0022] The effects of the invention Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Achieve highly realistic two-person interactive speech generation; 2. Low computational overhead, suitable for real-time deployment in VR scenarios; 3. It can generate facial expressions and movements that are appropriate for specific social situations. Attached Figure Description

[0023] Figure 1 This is a flowchart of a speech header generation method according to the first embodiment of the present invention.

[0024] Figure 2 This is a structural diagram of a computer device according to a third embodiment of the present invention. Detailed Implementation

[0025] The following is a detailed description of the speech head generation method involved in this invention.

[0026] Figure 1 This is a flowchart of a speech header generation method according to the first embodiment of the present invention. Figure 1As shown, the specific process of this speaker head generation method is as follows: First, an extraction step (step S100) is performed, extracting the speech features and facial motion features of the speaker and listener from a multi-turn dialogue scenario containing the speaker and listener, and then processing them using an action coding network to obtain an interaction representation. Next, an injection step (step S101) is performed, injecting socially perceptual features into the interaction representation to obtain Gaussian shift correction and a joint representation. Then, a preliminary generation step (step S102) is performed, generating a facial network animation based on the joint representation. Next, a correction step (step S103) is performed, correcting the facial network animation based on Gaussian shift correction to obtain a dynamic three-dimensional Gaussian primitive set. Finally, a final generation step (step S104) is performed, bidirectionally generating the speaker head based on the dynamic three-dimensional Gaussian primitive set.

[0027] First, let's explain step S100.

[0028] Step S100 extracts the speech features and facial motion features of the speaker and listener from the multi-turn dialogue scenario simultaneously, providing a unified temporal representation for subsequent social perception modeling and 3D reconstruction.

[0029] First, the input two-person dialogue video and corresponding audio are preprocessed. The original audio signal is resampled, denoised, and energy normalized. The multi-turn dialogue is then sliced ​​according to the audio content, ensuring that each segment includes an audio clip from one speaker and a synchronized video frame sequence for both parties. To ensure timeline consistency across multiple turns, this step uniformly aligns the audio and video streams, resamples all segments to a fixed frame rate, and interpolates missing frames as needed.

[0030] Secondly, to achieve parameterized face motion tracking, face detection and 3D face regression are performed on each frame of the video. Facial mesh motion encoding, including expression parameters, jaw rotation parameters, and head pose parameters, is extracted based on a 3D deformable face model (e.g., the FLAME model). By sorting the parameters of consecutive frames in the same video by time, two independent temporal mesh motion trajectories for the speaker and listener can be obtained, denoted as: Speaker movement sequence: ; Listener movement sequence: .

[0031] Here, each This represents the facial motion parameters of the digital human representing the speaker in frame i. These facial motion parameters include expression parameters, jaw rotation parameters, and head pose parameters. The facial motion parameters of the listener in frame i are represented by expression parameters, jaw rotation parameters, and head posture parameters.

[0032] Subsequently, a deep speech coding network is used to extract features from the speech signals of the speaker and the listener, mapping continuous audio segments into high-dimensional speech feature sequences, denoted as follows: Speaker's speech characteristics: ; Listener's voice characteristics: .

[0033] Here, each The audio features representing the speaker in the i-th frame are extracted by the audio encoder after processing the raw audio input. The audio features representing the listener in frame i are extracted by the audio encoder after processing the raw audio input.

[0034] Finally, the speaker's grid motion sequence was processed using an action coding network. Embedding is performed to obtain the speaker's temporal action representation. This representation (interactive representation) is related to speech features. In subsequent steps, we will jointly construct the speaker-listener interaction semantics, providing a foundation for injecting socially perceptual features.

[0035] The following describes step S101.

[0036] Step S101 explicitly transforms the social relationships between people into a learnable vector form and injects it into the speaker-listener interaction representation, thereby modulating the listener's facial response behavior and movement style.

[0037] First, the relationship between the two participants in the dialogue is structurally labeled. Social relationships are decomposed into multiple orthogonal dimensions, such as "blood ties / non-blood ties" and "equality / inequality," with each dimension corresponding to multiple discrete categories. For each category, a learnable embedding vector is constructed, and the embedding vectors of each dimension are concatenated or linearly combined to obtain the overall social relationship embedding. .

[0038] in, These represent relationship labels in the dimensions of blood ties and power, respectively. This indicates a vector concatenation operation.

[0039] Secondly, to enable social relation semantics to influence the motion decoding process, this invention constructs a social perception network to process embedded vectors. A nonlinear mapping is performed to obtain the social feature vector for motion modulation. and social feature vectors used for 3D Gaussian shift modulation : .

[0040] in, and There are two sets of multilayer perceptrons, where t represents time, i.e., the frame index number, which corresponds to the frame index number in step S100.

[0041] Subsequently, a speaker-listener interaction code was constructed, based on the speaker's speech features. Speaker's action characteristics and listener's voice characteristics Using multi-head attention-based temporal interaction mechanisms as input, the speaker's speech-action information is mapped into a joint representation that has predictive significance for the listener. In attention calculation, social features are introduced. As a query or modulation vector, attention weights are biased to produce differentiated response patterns under different relationship types. For example, in intimate relationship scenarios, the model tends to amplify facial expressions and nodding frequency, while in hierarchical relationship scenarios, the model tends to produce relatively restrained facial reactions.

[0042] Through the above steps, a comprehensive latent variable sequence that simultaneously encodes speech semantics, speaker action, and social relation semantics is obtained. This provides high-level guidance for the subsequent generation of facial mesh animations.

[0043] The following describes step S102.

[0044] Step S102 is based on the joint characterization obtained in step S101. Generate listener facial mesh animations that match the speaker's speech content and social context.

[0045] First, a motion decoding network based on temporal Trasformer is constructed, with the network input being a sequence of synthesized representations, i.e., joint representations. The output is the target mesh motion parameters of the listener at each time step: .

[0046] Each of them This includes factors such as facial expression coefficients and jaw rotation. The network employs a multi-layered attention structure to model temporal dependencies, enabling the model to understand speech rhythm, semantic stress, and dialogue turn boundaries, thereby generating appropriate responses such as nodding, shaking the head, and smiling at the right moments.

[0047] Next, parameter training is performed, during which real listener mesh motion sequences are introduced. As a supervisory signal, the mean squared error loss (MSE) is used to constrain the consistency between the predicted sequence and the true sequence: .

[0048] After the above decoding process, a listener grid animation trajectory that conforms to the rhythm of speech and reflects the differences in social relationships is obtained, providing parameter input for the motion drive of three-dimensional Gaussian elements.

[0049] The following describes step S103.

[0050] In step S103, under the framework of three-dimensional Gaussian scattering, the face mesh animation generated in step S102 is mapped into a set of renderable three-dimensional Gaussian primitives, and the positions of these primitives are corrected by social perception features, so that facial geometric deformation and social behavioral style can be uniformly expressed in three-dimensional space.

[0051] First, construct a set of Gaussian anchor points in the normalized space based on the neutral face mesh. For each face triangle, initialize a Gaussian anchor point with normalized coordinates as follows: The initial rotation matrix is ​​the identity matrix, and the scale parameter is a preset constant. As the mesh animation progresses, the vertex positions of the corresponding triangular facets change, and their affine transformation in the current frame can be calculated. ,in For rotation matrix, As a scale factor, It is the translation vector of the centroid of the triangular facet.

[0052] Without considering social factors, the position of the anchor Gauss in the dynamic space can be represented as: .

[0053] Secondly, to inject social relation semantics into the 3D geometric details, a Gaussian offset network is constructed, using the social features obtained in step S101. As input, output Gaussian displacement offset : .

[0054] in, This is a multilayer perceptron. Ultimately, the actual position of each Gaussian cell in the current frame is updated as follows: .

[0055] Among them, offset Continuous variation between adjacent frames allows for the continuous presentation of subtle facial expressions and inertial facial movements relevant to social relationships; for example, in intimate relationships, a listener might briefly nod after the speaker finishes speaking. For neural Gaussians generated through Gaussian splitting and cloning, this invention maintains that they share the same offset as the corresponding anchor Gaussians. To maintain the geometric consistency of local areas.

[0056] Finally, the updated Gaussian set The input is a 3D Gaussian scattering rasterization module, which generates intermediate features such as color maps and depth maps for each frame, providing a rendering foundation for the final high-fidelity speaking digital human generation.

[0057] The following describes step S104.

[0058] Step S104, based on the completion of three-dimensional Gaussian pose correction, generates a speaking digital human image sequence with high visual fidelity and social behavior consistency, i.e. bidirectional speaking head generation, which can be used in real-time scenarios such as virtual anchors, virtual customer service, and VR / AR interaction.

[0059] Step S104, based on the anchor point-neural Gaussian set for pose correction completed in step S103, generates a speaking digital human image sequence with high visual realism, temporal consistency and social behavior consistency under the constraints of three-stage joint training and multimodal triplet data supervision.

[0060] (a) High-consistency rendering generation based on anchor-neural Gaussian structure: After completing the 3D Gaussian element displacement correction in step S103, this step uses a hierarchical structure of "anchor Gaussian – neural Gaussian" as the rendering input unit. Each anchor Gaussian maintains a one-to-one binding relationship with its corresponding face mesh triangle, and its position, rotation, and scale are jointly determined by the mesh animation parameters generated in step three and the social relationship modulation offset obtained in step S103. The neural Gaussians obtained by cloning or splitting the anchor Gaussians share the pose and social offset parameters of the same anchor Gaussian during spatial transformation, thus ensuring the topological and motion consistency of the local Gaussian set during the adaptive adjustment of Gaussian density.

[0061] During frame-by-frame rendering, the 3D Gaussian scattering renderer projects the updated Gaussian set onto a 2D pixel plane given the camera's intrinsic and extrinsic parameters. It then performs weighted accumulation based on the color, opacity, and spatial distribution of each Gaussian element to generate the corresponding 2D head image for each frame. Due to the consistent constraints of the anchor-neural Gaussian structure in both spatial and temporal dimensions, the rendering results effectively avoid Gaussian drift, texture stretching, and facial expression misalignment during multi-turn dialogues, ensuring continuity and stability in facial expression transitions, head movements, and micro-expression details between the speaker and listener.

[0062] (II) Image-level refinement and optimization combining a three-stage joint training strategy: During the model training phase, the final end-to-end stage of the three-stage training strategy is jointly optimized. Specifically, after pre-training the speaker-listener motion generation and avatar rendering modules in the first two stages, social perception is introduced in this stage, and end-to-end joint training is performed on the motion generation, 3D Gaussian pose correction, and 2D rendering processes.

[0063] In this stage, the rendered digital human image is aligned with real video frames, and multiple optimization objectives, including pixel reconstruction loss, perceptual similarity loss, and adversarial loss, are constructed to constrain the generated image to approximate real samples in terms of brightness distribution, texture detail, and structural consistency. Simultaneously, a regularization constraint for Gaussian shift is introduced to limit the spatial offset caused by social relationship modulation, preventing excessive deformation from affecting visual stability. Through these joint optimization strategies, the digital human image generated in step five maintains high realism while also considering temporal smoothness and the rationality of social behavior in multi-turn dialogue scenarios.

[0064] (III) Supervised generation based on the multimodal triplet dataset (RSATalker): To support the stable generation and generalization of high-fidelity digital humans in step S104, the multimodal triplet dataset (RSATalker) is used as a unified source of supervision during training. This dataset contains time-aligned speech signals, 3D face mesh motion parameters, and corresponding real video frames, along with explicit social relationship labels.

[0065] In the training implementation of step S104, the generated 2D digital human images are aligned one-to-one with the corresponding real video frames in the dataset, ensuring that the model simultaneously meets the requirements of speech-expression consistency, mesh-Gaussian geometry consistency, and appearance consistency between the rendered results and real images within the same training framework. Through this joint supervision method based on the "speech-mesh-image" triplet, the model can stably generate speaking digital human image sequences that conform to real social contexts under different identities, social relationships, and dialogue rounds.

[0066] (iv) Output of generated results: After the above rendering and joint optimization processing, this step finally outputs a video sequence of a speaking digital human that is strictly synchronized with the input speech content, can reflect the differences in social relationships in multi-turn dialogues, and has high visual realism and real-time synthesis capabilities, thus achieving a high-fidelity visualization of the speaker-listener interaction process.

[0067] Example The following specific embodiment illustrates the above-mentioned speech header generation method.

[0068] The following is an example of the implementation of steps S100 to S103 in a two-person, two-round dialogue.

[0069] The scenario involves two short rounds of dialogue between speaker A and speaker B: Round 1: a. Say one sentence, b. Listen; Round 2: b says a sentence, a listens.

[0070] (1) First round: a. Say a sentence, b. Listen.

[0071] The processing procedure for step S100 is as follows: enter: a's speaking facial motion parameter sequence ; The audio of a; The audio of b.

[0072] deal with: The audio of 'a' is processed using an audio encoder to obtain the audio features of 'a'. ; The audio of b is processed using an audio encoder to obtain the audio features of b. .

[0073] Using an action coding network to analyze the facial motion parameter sequence of a during speech Processing is performed to obtain the temporal action representation of a. .

[0074] Output: audio characteristics of a ; audio features of b ; temporal action representation of a .

[0075] The processing procedure for step S101 is as follows: enter: audio characteristics of a ; audio features of b ; temporal action representation of a ; Social relationship labels between a and b.

[0076] deal with: Map social relationship tags to social relationship embeddings q; Embed social relationships into q inputs to action social networks ,get ; The social relationship embedding q and the time frame number t are input into a Gaussian displacement network. ,get ; Will , , and By fusing cross-attention mechanisms, a comprehensive latent variable sequence Z is obtained.

[0077] Output: Gaussian displacement correction ; The sequence of hidden variables is Z.

[0078] The processing procedure for step S102 is as follows: enter: The sequence of hidden variables is Z.

[0079] deal with: The facial motion parameters of the interlocutor b are obtained by processing Z using an action decoder. The loss is calculated between the predicted facial motion parameters of the interlocutor b and the actual values, and the network is optimized.

[0080] Output: Facial movement parameters of interlocutor b.

[0081] The processing procedure in step S103 is as follows: enter: Facial movement parameters of interlocutor B; A pre-constructed set of 3D Gaussian anchor points for the human face; Gaussian displacement correction .

[0082] deal with: The facial motion parameters of the interlocutor b are used to drive the face mesh model; Update the position and pose of the anchor Gaussian based on the moving face mesh model; Combined with Gaussian displacement correction Spatial displacement correction is performed on the Gaussian elements.

[0083] Output: A dynamic set of three-dimensional Gaussian meta-elements that match facial movements and social relationships in the current round.

[0084] The processing procedure for step S104 is as follows: enter: The dynamic three-dimensional Gaussian element set obtained in step S103; Camera parameters.

[0085] deal with: 3DGS technology is used to render three-dimensional Gaussian primitives frame by frame to generate two-dimensional human face images.

[0086] Output: A digital human video was synthesized, in which the speaker b acts as the listener.

[0087] (2) In the second round of dialogue, b says a sentence and a listens.

[0088] Then swap the positions of a and b, and repeat the same steps as above.

[0089] Therefore, the speaking head generation method according to the first embodiment of the present invention adopts a hybrid representation of "FLAME mesh + 3DGS": firstly, the three-dimensional facial motion of the FLAME mesh is generated by voice-driven generation; then, the three-dimensional Gaussian primitives are bound one by one to the mesh triangular faces; and the Gaussian is mapped from the normal coordinate system to the deformable coordinate system through coordinate system transformation, realizing precise linkage of head rendering with mesh deformation. Utilizing the fast rendering characteristics of Gaussian splashing, frame-level efficient rendering is achieved while maintaining head posture, facial details, and texture accuracy, making it suitable for deployment in immersive scenes with high real-time requirements, such as VR. In this way, it combines the structural controllability of traditional mesh methods with the efficient and realistic rendering capabilities of 3DGS. Compared with schemes using only meshes or NeRF, it can obtain higher image quality and more natural facial dynamics at a lower computational cost, significantly reducing the "uncanny valley" effect and enhancing the immersion and credibility of virtual humans.

[0090] Furthermore, a speaker-listener motion generation mechanism was designed. Using the voice and facial movements of both participants as input, multimodal features of speaker A are extracted through a voice encoder and a motion encoder. These features are then combined with the voice features of listener B. Cross-attention and a Transformer decoder are used to generate target facial motion parameters for listener B, which drive the FLAME head model to complete role switching and facial expression / head movement generation in multi-turn dialogues. This module can uniformly model non-verbal behaviors (nodding, smiling, blinking, etc.) in both "speaking" and "listening" states, maintaining temporal smoothness and coherence. By explicitly modeling the dynamic coupling relationship between the speaker and listener, this avoids the problem of poor facial expressions and stiff movements caused by relying solely on filler words in the listening rounds, which is common in traditional "single-speaker" methods. This enables more natural role switching and interactive feedback in multi-turn dialogue scenarios, significantly improving the fluency and interactivity of the generated avatar in long conversations.

[0091] Furthermore, a social perception mechanism is employed to encode interpersonal relationships from orthogonal dimensions such as "blood ties / non-blood ties" and "equality / inequality," constructing a discrete social relationship embedding. A structured interpersonal relationship query vector is then generated through a learnable query mechanism. This query is transformed into motion social features via a motion social network to modulate the generation of the listener's facial movements. Simultaneously, it is input into a Gaussian offset network along with the time step, outputting spatial offset corrections for different anchor points, achieving geometric-level social relationship perception. In this way, the design transforms abstract interpersonal relationships into learnable conditional signals, influencing both macroscopic motion patterns such as facial expressions and head movements, and finely adjusting the local geometry of the Gaussian distribution. This enables the generation of facial behaviors consistent in role, reasonable in emotion, and in sync with the interaction rhythm in different scenarios such as "superior-subordinate," "parent-child," "lovers," and "siblings." Compared to traditional methods that do not model social relationships, this significantly improves the performance of multi-turn dialogue avatars in terms of "social consistency" and "role credibility."

[0092] Furthermore, a hierarchical structure of "anchor Gaussian – neural Gaussian" is introduced in the 3DGS part: in the initial stage, anchor Gaussian is assigned to each FLAME triangle, and clone / split Gaussian generated by adaptive density control is used as neural Gaussian, and the consistency of topology and motion is maintained by sharing anchor point offsets. At the same time, a "three-stage training method" is adopted: the first stage pre-trains speaker-listener motion generation on large-scale multi-turn dialogue mesh data; the second stage pre-trains the avatar rendering module to learn stable Gaussian-mesh alignment relationship; the third stage introduces a social awareness module for end-to-end joint optimization, supplemented by multiple constraints such as image reconstruction loss, position constraint loss and Gaussian offset regularization. In this way, the anchor-neural Gaussian structure ensures the controllability and alignment accuracy of the Gaussian distribution during the density adaptation process, avoiding facial misalignment and texture stretching caused by Gaussian drift. The three-stage training effectively alleviates the non-convergence and artifact problems that are prone to occur in direct end-to-end training, enabling the system to obtain better image quality and temporal consistency while ensuring training stability. Therefore, it outperforms existing representative methods in objective indicators such as L1, PSNR, SSIM, and LPIPS, as well as in user subjective evaluations, verifying the comprehensive advantages of this design in terms of realism, fluency, and accuracy of social relationships.

[0093] Furthermore, based on existing multi-turn dialogue data, a "voice-grid-video" triplet is established and supplemented with refined social relationship annotations, forming the RSATalker dataset specifically for multi-turn social dialogue avatar generation, providing a unified multimodal supervision signal for the model. In this way, by simultaneously modeling voice-driven processes, 3D grid motion, and 2D avatar rendering within the same dataset, this invention significantly improves the model's comprehensive generalization ability in cross-modal consistency, identity preservation, and social relationship expression, providing a standardized data foundation and evaluation benchmark for subsequent comparative evaluation and industrialization of related algorithms.

[0094] The second embodiment of the present invention provides a speaker head generation device comprising: an extraction module, which extracts the speech features and facial motion features of the speaker and listener from a multi-turn dialogue scenario involving a speaker and a listener, and then processes them using an action coding network to obtain an interaction representation; an injection module, which injects socially perceptual features into the interaction representation to obtain a Gaussian shift correction and a joint representation; a preliminary generation module, which generates a facial network animation based on the joint representation; a correction module, which corrects the facial network animation based on the Gaussian shift correction to obtain a dynamic three-dimensional Gaussian primitive set; and a final generation module, which bidirectionally generates a speaker head based on the dynamic three-dimensional Gaussian primitive set. The speaker head generation device corresponds to the speaker head generation method of the first embodiment; therefore, various modifications of the first embodiment are also applicable to the second embodiment, and will not be described further here.

[0095] As described above, the speech head generation apparatus according to the second embodiment of the present invention solves the problems of static and rigid routing rules in the speech head generation process of the prior art, which cannot adapt to the dynamic changes of business, lack understanding of business semantics, and have poor scalability and maintainability.

[0096] The third embodiment of the present invention provides a computer device, the internal structure of which can be shown in the figure below. Figure 2 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to connect to external devices for data interaction. When the computer program is executed by the processor, it implements the speech header generation method according to the first embodiment of the present invention.

[0097] Those skilled in the art will understand that Figure 2The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0098] The fourth embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein a processor executes the computer program to implement the speech head generation method involved in the first embodiment of the present invention.

[0099] The fifth embodiment of the present invention provides a computer program product, including a computer program, wherein when a processor executes the computer program, it implements the speech head generation method involved in the first embodiment of the present invention.

[0100] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0101] Industrial application The speaking head generation method, apparatus, computer equipment, storage medium, and computer program products of the present invention solve the problems of the contradiction between realism and interactivity, the bottleneck of computational efficiency and practicality, and the lack of social context factors in the prior art.

[0102] Although the invention has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the invention.

Claims

1. A method for generating a speaker's head, the method being designed for multi-turn dialogue scenarios, characterized in that, Includes the following steps: The extraction step involves extracting the speech features and facial motion features of the speaker and listener from a multi-turn dialogue scenario that includes both the speaker and the listener. Then, the action coding network is used to process these features to obtain the interaction representation. The injection step injects socially perceived features into the interactive representation, resulting in Gaussian displacement correction and joint representation. The initial generation step involves generating a facial network animation based on joint representations. The correction step involves correcting the facial network animation based on Gaussian displacement correction to obtain a dynamic 3D Gaussian element set. The final generation step involves bidirectionally generating the speech head based on a dynamic three-dimensional Gaussian meta-set.

2. The method for generating a speaking head according to claim 1, characterized in that, The social relationship in the characteristics of social perception is either blood-related or non-blood-related.

3. The method for generating a speaking head according to claim 1, characterized in that, Social relations in the characteristics of social perception are either equal or unequal.

4. The method for generating a speaking head according to claim 1, characterized in that, In the initial generation step, the joint representation is converted into target mesh motion parameters, and then parameter training is performed. During parameter training, the mesh motion sequence is introduced as a supervision signal.

5. The method for generating a speaking head according to claim 1, characterized in that, In the correction step, an anchor Gaussian-neural Gaussian hierarchical structure is adopted, and social perception features are input.

6. The method for generating a speaking head according to claim 1, characterized in that, In the final generation step, joint training is performed in three stages. In the first stage, pre-training of speaker-listener motion generation is completed. In the second stage, pre-training of speaker-listener avatar rendering module is completed. In the third stage, the speech-mesh-image triplet dataset is used as a unified source of supervision.

7. A speaker generation device, the device being designed for multi-turn dialogue scenarios, characterized in that, include: The extraction module extracts the speech features and facial motion features of the speaker and listener from a multi-turn dialogue scenario that includes the speaker and listener, and then processes them using an action coding network to obtain the interaction representation. The injection module injects socially perceived features into the interactive representation, resulting in Gaussian displacement correction and joint representation. The initial generation module generates facial network animations based on joint representations; The correction module corrects the facial network animation based on Gaussian displacement correction to obtain a dynamic 3D Gaussian element set; The final generation module generates a speech head bidirectionally based on a dynamic three-dimensional Gaussian meta-set.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the speech head generation method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the speech head generation method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the speech head generation method according to any one of claims 1 to 6.