Interactive multi-round dialogue digital human modeling system and method

The self-supervised method extracts and aligns the multimodal features of the two-person, and generates natural 3D speaker animations, solving the problem of insufficient role conversion and non-verbal feedback in the existing technology, and realizing natural interaction and high synchronization in multiple rounds of dialogue.

CN120279176APending Publication Date: 2025-07-08RENMIN UNIVERSITY OF CHINA
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510364366.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing 3D speaking generation technology cannot support the alternating transitions of speaking and listening characters, multiple rounds of continuous dialogue, and realistic nonverbal feedback, resulting in a lack of coherence and interactivity in dialogue.

Method used

The self-supervised method is adopted to extract the double-player multimodal features in interactive multi-round dialogue scenarios, perform timing alignment and enhancement, and a codec with a fusion attention mechanism is used to generate the listener's voice text and expression parameters, and combine the 3D facial animation frame sequence to generate natural dialogue animation.

Benefits of technology

It realizes seamless transformation of speaker and listener roles, supports natural rotation in multiple rounds of interaction, generates rich nonverbal feedback and highly synchronized voice and lip styling, improving the fluency of conversation, interactive realism and expressiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279176A_ABST
    Figure CN120279176A_ABST
Patent Text Reader

Abstract

The invention relates to an interactive multi-round dialogue digital human modeling method, which comprises the following steps of: extracting double-person multi-modal characteristics in an interactive multi-round dialogue scene, including voice characteristics and expression characteristics of a speaker in a current dialogue round and voice characteristics of a listener in a previous dialogue round in the current dialogue round; performing time sequence alignment and reinforcement on the extracted double-person multi-modal features based on a time dimension to obtain a combined feature sequence; according to the joint feature sequence, generating a voice text of the listener in the current dialogue round and synchronous expression parameters based on a codec fused with an attention mechanism; and generating a corresponding 3D facial animation frame sequence according to the expression parameters of the listener in the current dialogue round.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision technology and artificial intelligence interaction, and particularly to an interactive multi-round dialogue digital human modeling system and method. Background Art

[0002] With the development of virtual digital humans and conversational AI, the market urgently needs 3D speaker models that can simulate face-to-face communication with real people to enhance the interaction experience.

[0003] Ideally, dialogue participants should be able to smoothly switch between the speaker and listener roles and make the communication more natural and credible through non-verbal feedback such as facial expressions and nodding.

[0004] However, the existing 3D speaker generation technologies have obvious deficiencies: most methods only model a single role, either only generating the lip movements and expressions of the speaker (ignoring the reactions during listening), or only generating simple responses of the listener. For example, some speaker models can synchronize lip movements according to speech, but lack the portrayal of key signals such as nodding and frowning during listening, resulting in a lack of coherence and interactivity in the dialogue. On the contrary, some listener models can only generate brief nodding or expression responses based on the other party's speech and cannot continuously integrate into the multi-round dialogue scenario. This mode of role fragmentation ignores the dynamics of role rotation and mutual influence in real conversations, and is prone to unnatural interactions and awkward role transitions.

[0005] In summary, the current technology cannot meet the requirements of natural two-way dialogue, cannot support the alternating conversion of the speaker and listener roles, multi-round continuous dialogue, and realistic non-verbal feedback at the same time, and there is an urgent need for an improved solution to solve the above problems. Summary of the Invention

[0006] In view of the above problems, the purpose of the present invention is to provide a self-supervised 3D face modeling method, which aims to directly reconstruct a high-precision 3D face model from a video through a self-supervised method without relying on a pre-trained model.

[0007] To achieve the above purpose, the present invention adopts the following technical solutions:

[0008] In a first aspect, the present application provides an interactive multi-round dialogue digital human modeling method, including:

[0009] Extracting dual-person multimodal features in an interactive multi-round dialogue scenario, including the speech features and expression features of the speaker in the current dialogue turn, and the speech features of the listener in the previous dialogue turn in the current dialogue turn;

[0010] Performing temporal alignment and enhancement on the extracted dual-person multimodal features in the time dimension to obtain a joint feature sequence;

[0011] Based on the joint feature sequence, a codec based on a fusion attention mechanism generates the speech text of the listener in the current conversation turn and the synchronized expression parameters.

[0012] According to the expression parameters of the listener in the current conversation turn, a corresponding 3D facial animation frame sequence is generated.

[0013] In one embodiment, in the extraction of dual-person multi-modal features in an interactive multi-turn conversation scenario, the speech signal of the speaker in the current conversation turn and the speech signal of the listener in the previous conversation turn are collected by a microphone, and a pre-trained audio encoder is used to extract the high-dimensional speech features in the speech signal.

[0014] In one embodiment, in the extraction of dual-person multi-modal features in an interactive multi-turn conversation scenario, the 3D face expression parameters of the speaker in the current conversation turn are extracted as expression features.

[0015] In a second aspect, an interactive multi-turn conversation digital human modeling system is provided, including:

[0016] A dual-person multi-modal feature extraction module, configured to extract dual-person multi-modal features in an interactive multi-turn conversation scenario, including the speech features and expression features of the speaker in the current conversation turn, and the speech features of the listener in the previous conversation turn in the current conversation turn

[0017] A cross-modal temporal enhancement module, configured to perform temporal alignment and enhancement on the extracted dual-person multi-modal features based on the time dimension to obtain a joint feature sequence;

[0018] A dual-speaker interaction module, configured to generate the speech text of the listener in the current conversation turn and the synchronized expression parameters based on the joint feature sequence by a codec based on a fusion attention mechanism;

[0019] An expression synthesis module, configured to generate a corresponding 3D facial animation frame sequence according to the expression parameters of the listener in the current conversation turn.

[0020] Due to the above technical solutions adopted by the present invention, it has the following advantages:

[0021] The present invention realizes the seamless conversion of the speaker and listener roles, supports the natural turn-taking of both parties in multi-round interactions without generating a rigid transition. Secondly, the present invention can generate appropriate non-verbal feedback (such as nodding, expressions) during the listening stage, enriching the interactivity and realism of the conversation; during the speaking stage, it ensures a high degree of synchronization between the voice and lip movements, and the expressions are consistent with the semantic emotions. By considering the behaviors of both roles simultaneously, the conversation generated by the present invention is more coherent and fluent, and the interaction rhythm is closer to that of real people, significantly improving the naturalness and expressiveness of the virtual 3D speaker conversation. Experimental results show that compared with the existing models that only support single roles, this solution has obvious improvements in aspects such as the fluency of the conversation, the realism of the interaction, and the richness of expressions, and can bring a more immersive communication experience to users. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic diagram of the framework structure of the DualTalk dual-speaker interaction 3D speaker conversation generation method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention fall within the scope of protection of the present invention.

[0024] In view of the problems of the prior art, an interactive multi-round dialogue digital human modeling method is provided in an embodiment of the present invention, including:

[0025] S1, extracting dual-person multi-modal features in an interactive multi-round dialogue scenario, including the voice features and expression features of the speaker in the current dialogue turn, and the voice features of the listener in the previous dialogue turn in the current dialogue turn;

[0026] S2, performing temporal alignment and enhancement on the extracted dual-person multi-modal features based on the time dimension to obtain a joint feature sequence;

[0027] S3, based on the joint feature sequence, generating the voice text of the listener in the current dialogue turn and the synchronized expression parameters based on an encoder-decoder with a fusion attention mechanism;

[0028] S4, generating a corresponding 3D facial animation frame sequence according to the expression parameters of the listener in the current dialogue turn.

[0029] The following is further described in detail based on the attached drawings of the present application Figure 1 , and further detailed description is carried out.

[0030] The present invention proposes a dual-speaker interactive 3D speaker dialogue generation framework, namely the "DualTalk" unified model, which is used to depict the dynamic behaviors of both parties in a conversation simultaneously. Different from the traditional method of separately modeling speakers and listeners, this solution regards the dialogue participants as roles that can both speak and listen, and integrates the behaviors of both roles in a single model. Its overall architecture is as shown in Figure 1 and includes the following four core modules, and each module works together to generate natural and coherent two-person dialogue animations:

[0031] Dual-speaker joint encoding module: Obtain and encode the multimodal features of both parties, and fuse the speech and visual signals of the two speakers in a unified feature space. Specifically, this module separately uses a pre-trained audio encoder to extract high-dimensional speech features for each speaker's speech, and at the same time encodes the facial movements of the main speaker (such as 3D face expression parameters, that is, blendshape coefficients). By projecting the speech features of both parties and the facial movement features of one party into a shared representation space, the joint representation of multimodal information of both parties in the conversation is realized, providing a feature basis that integrates the information of speakers and listeners for the subsequent modules.

[0032] Cross-modal temporal enhancement module: Under the unified feature space, perform temporal alignment and enhancement on the fused multimodal feature sequence to ensure the coherence of dialogue features in the time dimension. This module introduces a cross-modal attention mechanism to enable the speech features to guide the adjustment of the corresponding frame's facial movement features, thereby closely aligning the visual expressions with the speech rhythm. Subsequently, a bidirectional long short-term memory network (Bi-LSTM) is used to model the context of the feature sequence, capture the information of past and future frames, and strengthen the temporal correlation of the feature sequence. After being processed by this module, the speech and facial action features correspond one by one and are synchronized in time, outputting a temporal feature representation that integrates speech content and visual dynamics, laying a foundation for the reproduction of the real dialogue rhythm.

[0033] More specifically, the cross-modal temporal enhancement module processes based on the following mathematical method: Based on the coherence in the time dimension, use the high-dimensional speech features to guide the adjustment process of the 3D face expression parameters of the corresponding frames, and the specific description is as follows:

[0034] Represent the high-dimensional speech features in the current dialogue turn as The face expression parameter of the corresponding frame is represented as Then the adjustment process of the expression parameter can be expressed as:

[0035] M′t = M t + F(A t ; θ exp )

[0036] where F represents the mapping function from speech features to expression adjustment amounts, and this function is determined by the deep neural network parameters θ exp ;

[0037] Subsequently, the adjusted multi-modal feature sequence {X t}, where X t = [A t ; M't], is input into a bidirectional long short-term memory network (Bi-LSTM) for context modeling:

[0038] H t = Bi-LSTM(X t , H t-1 , H t+1 ; θ lstm )

[0039] where H t is the hidden state output by the bidirectional LSTM, representing the context-aware feature sequence after modeling in the time dimension.

[0040] Two-speaker interaction module: Models the interaction relationship between the two parties' roles and generates context-aware dialogue features. This module adopts an encoder-decoder architecture based on Transformer: First, the joint feature sequence output by the previous module is processed by the Transformer encoder to capture the long-distance dependencies and complex interaction patterns across frames between the speaker and the listener. To further highlight the association between different modalities, this module introduces a "modal alignment attention" mechanism (Modal Alignment Attention), which draws on the bias attention strategy to dynamically adjust the attention weights according to the dialogue context. This mechanism uses an alignment mask to reallocate the attention focus of speech and facial features in the Transformer, enabling the generated response to be synchronized with the other party's expression and speech rhythm, thereby achieving context alignment of the actions of the speaker and the listener at the feature level. Then, the Transformer decoder decodes the feature sequence optimized by alignment to obtain a representation containing rich interaction information, which depicts both the expected expression changes of the current speaker and the influence of the non-verbal feedback of the other party (the current listener). Through this interaction module, the present invention can simulate the dynamic relationship of back-and-forth interaction between the two parties in a real dialogue, enabling the generated behavior to respond to both the current context and the actions of the other party.

[0041] More specifically, the operation process of the two-speaker interaction module is as follows:

[0042] Obtain a context-aware feature representation through the Transformer encoder:

[0043] Z = TransformerEncoder(H t ; θenc )

[0044] Then, the Modal Alignment Attention (MAA) mechanism is introduced to dynamically adjust the attention weights:

[0045]

[0046] Among them, W Q , W K , W V are learnable projection matrices, and d k represents the dimensions of the query and the key;

[0047] Next, the Transformer decoder decodes the aligned and optimized feature sequence Z′:

[0048] D = TransformerDecoder(Z′; θ dec )

[0049] Among them, D contains rich interaction information and can simultaneously depict the expected expression changes of the current speaker and the context interaction relationship.

[0050] Furthermore, it also includes fine-tuning the features output by the Transformer decoder using an adaptive expression modulation mechanism:

[0051] Specifically, the adaptive expression modulation is defined as:

[0052] D′ = D + α · σ(DW m + b m )

[0053] Among them, D is the output of the Transformer decoder, is a learnable parameter, σ is a non-linear activation function (such as ReLU), and α is a modulation coefficient used to dynamically control the expression amplitude and emotional intensity;

[0054] The modulated feature sequence D′ then passes through a fully connected mapping layer to obtain the actual face parameter sequence

[0055]

[0056] Among them, is the weight matrix and bias term for the final expression parameter mapping.

[0057] Facial expression synthesis module: Generates the final 3D speaker facial animation frame sequence based on the features output by the interaction module. This module predicts the facial expression parameters (such as blendshape coefficients) of the listener (the party whose turn it is to generate the head currently), thereby driving the 3D head model to produce corresponding actions. To make the generated expressions more vivid and realistic, this module includes an adaptive expression modulation mechanism to fine-tune the features output by the Transformer decoder. Specifically, first extract the emotional components in the features through a learnable non-linear mapping, multiply them by the adjustment coefficient, and then add them to the original features to dynamically adjust the expression amplitude and emotional intensity. The modulated features are then mapped to the actual face parameter space via a fully connected layer to obtain the final facial motion sequence. This module ensures that the mouth shape is highly synchronized with the speech during speech, and the facial expressions and feedback actions during listening are consistent with the dialogue context, and can generate delicate and emotional facial animations.

[0058] For Figure 1 the provided framework, its implementation method process is as follows:

[0059] Input data acquisition: Collect and prepare the audio and head movement data of both parties in the conversation. Assume that in the current conversation turn, speaker A is the speaker and speaker B is the listener, then the inputs include: the speech signal A of speaker A A (such as a voice audio) and the corresponding head movement parameters M A (such as the sequence of video face blendshape coefficients for each frame), and the speech signal A of speaker B in this conversation turn B (the short speech that the listener may make, such as "um" or a short answer). These data can be collected through a microphone and a camera and obtained after necessary preprocessing (such as audio separation, face tracking). The target output of the model is the head movement sequence M of speaker B during listening and when it is their turn to speak B , so that it is synchronized with the conversation context. Formally, it can be expressed as: where f represents the dual-speaker conversation generation mapping relationship implemented by the model of the present invention.

[0060] Feature extraction and joint encoding: Send the above inputs into the "dual-speaker joint encoding module" for multi-modal feature extraction and fusion. Specifically, respectively send the audio A of speaker A A , A B into their respective speech encoders (such as a speech feature extraction model based on a deep neural network) to extract the high-dimensional speech feature representations H A and H B . At the same time, send the head movement parameter sequence M of speaker A AInput the facial feature encoder, map the blendshape coefficients of each frame into a low-dimensional vector representation, and obtain the facial motion feature sequence M'. A . Then, project the speech feature H A and H B into the same feature space as the facial features through linear transformation and fuse them with M' A to form a unified multi-modal feature representation Z. This fused feature encodes the speech content and expression dynamics of speaker A, as well as the speech content of speaker B, providing a feature input that integrates the information of both parties for the subsequent steps.

[0061] Cross-modal temporal feature enhancement: Feed the joint feature Z into the "cross-modal temporal enhancement module" to perform temporal synchronization and correlation modeling on speech and facial features. First, the cross-modal attention mechanism is adopted inside the module, using the speech feature of speaker A as the query, and its facial motion features as the key and value to calculate the attention weighted sum. Through this speech-guided visual attention calculation, the model can dynamically adjust the facial feature representation at the corresponding moment according to the speech content, making the mouth shape expression correspond to the speech rhythm. Subsequently, input the fused feature sequence into a bidirectional LSTM network to model the dependency relationship of features from the time dimension (context). The bidirectional LSTM utilizes the information of past frames on the one hand and combines the trend of future frames on the other hand, thereby outputting an enhanced feature sequence T containing the context of each moment. After being processed by this module, the features of the speech modality and the visual modality are highly aligned on the time axis and contain the temporal correlation information of the speech and expression changes during the conversation, ensuring that the subsequent generated action sequence conforms to the rhythm of natural conversation.

[0062] Dual-speaker Interaction Modeling: The enhanced multi-modal feature sequence T is input into the "Dual-speaker Interaction Module" to model the interaction influence relationship between speaker A and listener B. First, the Transformer encoder is used to encode the sequence features to capture the long-distance dependencies and global interaction patterns in the conversation. The features output by the encoder have initially integrated the behavioral dynamics of both parties. Next, the Modal Alignment Attention mechanism is introduced to adjust the encoded features. Through a pre-designed alignment mask, this mechanism emphasizes the information of relevant modalities, enabling the model to reasonably allocate the attention focus according to the current context during decoding. For example, when speaker A pauses in speech, more attention is paid to speaker A's facial expression or listener B's reaction. After the modal alignment process, a more coordinated interaction feature sequence T' that contains both speech and visual context can be obtained, where the time series of the actions of both parties is further aligned and the response relationship is more coherent. Then, the aligned feature sequence is input into the Transformer decoder for step-by-step decoding to generate an implicit representation D that contains the dialogue context information. This representation synthesizes the expression intention of the main speaker A and the feedback influence of the secondary participant B, which is equivalent to an "interaction feature snapshot" of the dialogue in the current round. Through the processing of the interaction module, the model can understand and simulate the action linkage between the "speaker-listener" parties, laying the foundation for generating natural responses.

[0063] Facial Expression Synthesis and Head Animation Generation: The representation D output by the interaction module is input into the "Facial Expression Synthesis Module" to generate the head movement sequence of listener B. First, the adaptive facial expression modulation mechanism is used to process D to enhance the emotional expression component therein. Specifically, a learnable mapping function Mod extracts the facial expression offset in the representation and multiplies it by the factor α and then superimposes it back onto D to obtain the modulated representation D'. The modulation coefficient α is used to control the amplitude of facial expression enhancement, thereby giving an appropriate emotional intensity while ensuring basic lip synchronization. Subsequently, D' is mapped through a fully connected layer to an output vector with the same dimension as the blendshape parameters to generate the facial expression parameter sequence of listener B frame by frame, which is the final head movement prediction. These parameters can directly drive a 3D face model (such as the FLAME model) to render the facial animation of speaker B in this round of conversation, including lip movement, eyebrow and eye movement, and head pose, etc. At this time, the generation process of this round of conversation is completed, and the obtained is synchronized with the speech and facial expressions of speaker A, and the dialogue simulation enters the next round.

[0064] Multi-round dialogue interaction: In practical applications, the dialogue proceeds in multiple continuous rounds. After completing the above steps, the original listener B may become the speaker in the next round, while the speaker A becomes the listener. At this time, the roles of both parties are exchanged, and steps 1 to 5 are repeated according to the new round of inputs (the speech and head movements of speaker B, and the speech of speaker A) to generate the head response movements of speaker A when listening to B's speech. Since the present invention uses a unified model to process both roles, it can naturally transition to the situation of role exchange without additional model switching, thus supporting continuous dialogue generation with multiple rounds and back-and-forth alternation of dialogue participants. By iterating the above process in turn, the present invention can simulate the cumulative effect of the interaction between both parties in a long dialogue and generate a smooth and realistic 3D speaker animation throughout the multi-round dialogue. This method ensures that no matter which role speaks, the other role can give coherent and appropriate facial feedback, realizing natural communication like real people coming and going.

[0065] In one embodiment, the processing process of the above identity conversion and recognition is specifically as follows:

[0066] Define the dialogue round as The speaker identity in the current round is S(r), and the listener identity is L(r); if the speaker and listener in the r-th round are denoted as A and B respectively, then the identity conversion in the next round is achieved through the following rules:

[0067] S(r + 1) = L(r), L(r + 1) = S(r)

[0068] The dialogue round count can be achieved through real-time audio segment detection, that is:

[0069] Define the voice activity detection (VAD) signal of the speaker in the current round as V S(r) (t). When the interruption time τ of the continuous voice signal exceeds the threshold γ, that is:

[0070]

[0071] Then it is considered that the dialogue round conversion has occurred, and the round count is incremented by 1;

[0072] Through the above method, each dialogue round is automatically marked, and the rotation of the roles of both parties in the dialogue is recognized in real time, realizing continuous modeling of multi-round dialogue.

[0073] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0074] The integrated units implemented in the form of software function units can be stored in a computer-readable storage medium. The above software function units stored in a storage medium include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0075] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. An interactive multi-round dialogue digital human modeling method, characterized in that Including: S1, extracting the dual-modal features of two persons in an interactive multi-turn dialogue scenario, including the speech features and facial expression features of the speaker in the current dialogue turn, and the speech features of the listener in the previous dialogue turn in the current dialogue turn; S2, performing temporal alignment and enhancement on the extracted dual-modal features of two persons based on the time dimension to obtain a joint feature sequence; S3, based on the joint feature sequence, using an encoder-decoder based on a fusion attention mechanism to generate the speech text of the listener in the current dialogue turn and the synchronized facial expression parameters; S4, generating a corresponding 3D facial animation frame sequence according to the facial expression parameters of the listener in the current dialogue turn.

2. The interactive multi-round dialogue digital human modeling method according to claim 1, wherein In the extraction of the dual-modal features of two persons in the interactive multi-turn dialogue scenario in S1, the speech signals of the speaker in the current dialogue turn and the speech signals of the listener in the previous dialogue turn are collected through a microphone, and the high-dimensional speech features in the speech signals are extracted using a pre-trained audio encoder.

3. The interactive multi-round dialogue digital human modeling method according to claim 2, wherein In the extraction of the dual-modal features of two persons in the interactive multi-turn dialogue scenario, the 3D face expression parameters of the speaker in the current dialogue turn are extracted as the facial expression features.

4. The interactive multi-round dialogue digital human modeling method according to claim 3, wherein In S2, it specifically includes: Based on the coherence in the time dimension, using the high-dimensional speech features to guide the adjustment process of the 3D face expression parameters of the corresponding frames, specifically described as: Represent the high-dimensional voice features in the current conversation turn as Represent the facial expression parameters of the corresponding frame as Then the adjustment process of the expression parameters can be expressed as: M't = M t + F(A t ; θ exp ) Among them, F represents a mapping function from speech features to expression adjustment amounts, and this function is determined by the deep neural network parameters θ exp determined; Subsequently, the adjusted multi-modal feature sequence {X t}, where X t = [A t ; W′t], is input into a bidirectional long short-term memory network (Bi-LSTM) for context modeling: H t = Bi-LSTM(X t , H t-1 , H t+1 ; θ lstm ) Among which H t is the hidden state output by the bidirectional LSTM, representing the context-aware feature sequence after modeling in the time dimension.

5. The interactive multi-round dialogue digital human modeling method according to claim 4, wherein In S3, the joint feature sequence H output by the Bi-LSTM is processed by the Transformer encoder t , which is specifically described as follows: First, obtaining a context-aware feature representation through a Transformer encoder: Z = TransformerEncoder(H t ; θ enc ) Then, introducing a Modal Alignment Attention (MAA) mechanism to dynamically adjust the attention weights: Among them, W Q , W K , W V is a learnable projection matrix, and d k represents the dimension of the query and the key; Next, the Transformer decoder decodes the feature sequence Z' optimized by alignment: D = TransformerDecoder(Z′; θ dec ) where D contains rich interaction information and can simultaneously depict the expected facial expression changes of the current speaker and the context interaction relationship.

6. The interactive multi-round dialogue digital human modeling method according to claim 5, wherein In S3, it also includes fine-tuning the features output by the Transformer decoder using an adaptive facial expression modulation mechanism: Specifically, defining the adaptive facial expression modulation as: D′ = D + α·σ(DW m + b m ) where D is the output of the Transformer decoder, are learnable parameters, σ is a non-linear activation function (such as ReLU), and α is a regulation coefficient used to dynamically control the expression amplitude and emotion intensity; The modulated feature sequence D′ is then passed through a fully connected mapping layer to obtain the actual face parameter sequence Among them, is the weight matrix and bias term for the final expression parameter mapping.

7. The self-supervised 3D face modeling method according to claim 1, characterized in that The method also includes: recording of the dialogue turns and identification of the role identity transitions, specifically implemented as follows: Define the dialogue turn as In the current turn, the identity of the speaker is S(r) and the identity of the listener is L(r); if we denote the speaker and listener in the r-th turn as A and B respectively, then the identity conversion in the next turn is achieved through the following rules: S(r + 1) = L(r), L(r + 1) = S(r) The dialogue turn counting can be achieved through real-time audio segment detection, that is: Define the voice activity detection (VAD) signal of the current speaker as V S(r) (t). When the interruption time τ of the continuous voice signal exceeds the threshold γ, i.e.: Then it is considered that a dialogue turn transition occurs, and the turn count is incremented by 1; Through the above method, each dialogue turn is automatically marked, and the rotation of the roles of the two parties in the dialogue is identified in real time, realizing the continuous modeling of multi-turn dialogues.

8. An interactive multi-round dialogue digital human modeling system, characterized in that, Including: A dual-modal feature extraction module for two persons, used to extract the dual-modal features of two persons in an interactive multi-turn dialogue scenario, including the speech features and facial expression features of the speaker in the current dialogue turn, and the speech features of the listener in the previous dialogue turn in the current dialogue turn A cross-modal temporal enhancement module for temporally aligning and enhancing the extracted dual-modal features of two persons based on the time dimension to obtain a joint feature sequence; A two-speaker interaction module for generating the speech text of the listener in the current dialogue turn and the synchronized facial expression parameters based on the joint feature sequence using an encoder-decoder based on a fusion attention mechanism; An expression synthesis module, which is used to generate a corresponding 3D facial animation frame sequence according to the expression parameters of the listener in the current conversation turn.

Citation Information

Cited By

  • Communication session interaction method and system based on AI digital person

    CN120848739A

  • Communication session interaction method and system based on AI digital person

    CN120848739B

  • Portrait dialogue video generation method, multi-person dialogue video generation method, product, equipment and storage medium

    CN121000955A

  • Portrait dialogue video generation method, multi-person dialogue video generation method, product, device and storage medium

    CN121000955B