Methods, apparatus, and computer equipment for constraining multimodal output of digital humans

By using multimodal emotion fusion and closed-loop feedback mechanisms, the problem of insufficient user emotion perception in digital human systems is solved, achieving coordination and consistency of multimodal output of digital humans and improving the naturalness and continuity of interaction.

CN121580336BActive Publication Date: 2026-04-03VASTAI TECH (SHANGHAI) INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing digital human systems lack the ability to perceive and continuously respond to users' emotions in real time, resulting in rigid interaction processes, inconsistent multimodal outputs, and reduced realism and immersion in the interaction.

Method used

By acquiring multiple input data (speech, image, text) to generate a fused emotion vector, evaluating and adjusting the multiple modal emotion components of the digital human to achieve consistent output, using a deep learning model for emotion feature extraction and fusion, and establishing a perception-decision closed-loop feedback mechanism.

Benefits of technology

It achieves coordinated consistency in the multimodal output of digital humans, improves the naturalness and continuity of interaction, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121580336B_ABST
    Figure CN121580336B_ABST
Patent Text Reader

Abstract

This disclosure provides a method, apparatus, and computer device for constraining the multimodal output of a digital human, comprising: acquiring a first input dataset for recognizing user emotions at a first moment based on a first time interval; generating a fused first user emotion vector based on the first input dataset; generating a first digital human emotion vector based on the first user emotion vector, and determining digital human emotion components corresponding to multiple modalities of the digital human; evaluating the consistency of the multiple digital human emotion components; adjusting at least one of the multiple digital human emotion components to meet the consistency requirement in response to multiple digital human emotion components not meeting the consistency requirement; and adjusting the output parameters of the corresponding modal of the digital human based on the multiple digital human emotion components that meet the consistency requirement, so as to achieve constrained multimodal output of the digital human. This ensures the consistency of the multimodal output of the digital human, thereby improving the naturalness of interaction and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of digital human interaction, and in particular to a method, apparatus, computer device, computer-readable storage medium, and computer program product for constraining the multimodal output of a digital human. Background Technology

[0002] With the rapid development of artificial intelligence technology, digital humans have evolved from early pre-recorded animated characters into intelligent interactive subjects capable of real-time voice dialogue and facial expression interaction, and are thus widely used in real-time interactive scenarios such as virtual assistants, digital customer service, educational tutoring, virtual anchors, and medical care.

[0003] In digital human interaction scenarios, multimodal outputs such as facial expressions, speech rhythm, and gestures are characterized by high visualization, multiple degrees of freedom, and strong perceptuality. However, existing digital human systems mostly rely on fixed scripts to drive speech and facial animation, lacking the ability to perceive and continuously respond to user emotions in real time, resulting in stiff interaction processes and a lack of natural emotional expression. Even with real-time adjustment, inconsistencies in emotional expression between different modalities are prone to occur. For example, a facial expression showing pleasure but a low voice tone can easily create a sense of "unnaturalness" or "mismatch" on the user's end, significantly reducing the realism and immersion of the interaction. Existing digital human interaction mode adjustment methods often fail to consider the diversity of digital human interaction scenarios and the special characteristics of digital humans as multimodal, high-dimensional, continuously controllable objects.

[0004] Therefore, a method is needed that can adjust the multimodal output of digital humans in real time, continuously, and naturally to improve the naturalness of interaction and user experience. Summary of the Invention

[0005] This disclosure provides a method, apparatus, computer device, computer-readable storage medium, and computer program product for constraining the multimodal output of a digital human.

[0006] According to one aspect of this disclosure, a method for constraining the multimodal output of a digital human is provided. The method includes: acquiring a first input dataset for recognizing user emotions at a first moment based on a first time interval, the first input dataset including at least two of first speech input data, first image input data, and first text input data; generating a fused first user emotion vector based on the first input dataset; generating a first digital human emotion vector based on the first user emotion vector; determining a plurality of digital human emotion components based on the first digital human emotion vector, wherein the plurality of digital human emotion components include facial expression modal emotion components, action modal emotion components, and speech modal emotion components; evaluating the consistency of the plurality of digital human emotion components, wherein, in response to the plurality of digital human emotion components not meeting the consistency requirement, at least one of the plurality of digital human emotion components is adjusted to meet the consistency requirement; and adjusting the output parameters of the corresponding modality of the digital human based on the plurality of digital human emotion components that meet the consistency requirement, so as to achieve constrained multimodal output of the digital human.

[0007] By introducing multimodal consistency constraints, the emotional expression of digital humans in facial, voice, and motion output is coordinated and consistent, thereby improving the naturalness and continuity of digital human multimodal interaction.

[0008] In some embodiments, assessing the consistency of multiple digital human emotion components includes: calculating a consistency metric between every two digital human emotion components; and assessing the consistency of the multiple digital human emotion components based on a comparison of the consistency metric and a consistency threshold.

[0009] In some embodiments, adjusting at least one of the multiple digital human emotion components to meet the consistency requirement in response to multiple digital human emotion components not meeting the consistency requirement includes: determining a first adjustment priority based on the interaction scenario of the digital human; determining a second adjustment priority based on a consistency metric; and determining at least one digital human emotion component that needs to be adjusted based on the first adjustment priority and the second adjustment priority.

[0010] In some embodiments, generating a fused first user emotion vector based on a first input dataset includes: extracting a first emotion feature set based on the first input dataset, the first emotion feature set including at least two of first speech emotion features, first image emotion features, and first text emotion features; generating a first emotion vector set based on the first emotion feature set, the first emotion vector set including at least two of first speech emotion vectors, first image emotion vectors, and first text emotion vectors; and fusing at least two emotion vectors in the first emotion vector set based on a deep learning model to generate the first user emotion vector.

[0011] In some embodiments, the method further includes: acquiring a second input dataset for identifying user emotions at a second time point, wherein the second time point is later than the first time point by a first time interval, and the second input dataset includes at least two of second voice input data, second image input data, and second text input data; generating a fused second user emotion vector based on the second input dataset; updating the second user emotion vector based on the first user emotion vector via a weighted update algorithm; determining a second digital human emotion vector based on the updated second user emotion vector and the first digital human emotion vector; and adjusting the output parameters of the corresponding modality of the digital human based on the second digital human emotion vector.

[0012] In some embodiments, generating a fused second user emotion vector based on a second input dataset includes: extracting a second emotion feature set based on the second input dataset, the second emotion feature set including at least two of second speech emotion features, second image emotion features, and second text emotion features; generating a second emotion vector set based on the second emotion feature set, the second emotion vector set including at least two of second speech emotion vectors, second image emotion vectors, and second text emotion vectors; and fusing at least two emotion vectors in the second emotion vector set based on a deep learning model to generate a second user emotion vector.

[0013] In some embodiments, determining the second digital human emotion vector based on the updated second user emotion vector and the first digital human emotion vector includes: determining an initial second digital human emotion vector based on the updated second user emotion vector; and adjusting the initial second digital human emotion vector based on a comparison between the initial second digital human emotion vector and the first digital human emotion vector and a preset digital human emotion change threshold.

[0014] In some embodiments, the method further includes: acquiring a third input dataset for identifying user emotions at a third time point, wherein the third time point is later than the first time point by at least a first time interval, and the third input dataset includes at least two of third voice input data, third image input data, and third text input data; extracting a third emotion feature set based on the third input dataset, the third emotion feature set including at least two of third voice emotion features, third image emotion features, and third text emotion features; generating a third emotion vector set based on the third emotion feature set, the third emotion vector set including at least two of third voice emotion vectors, third image emotion vectors, and third text emotion vectors; fusing at least two emotion vectors in the third emotion vector set based on a deep learning model to generate a third user emotion vector; and updating the first time interval to a second time interval based on a comparison between the difference between the third user emotion vector and the first user emotion vector and a preset difference in user emotion vectors.

[0015] In some embodiments, the deep learning model includes a multilayer perceptron model, a convolutional neural network model, or an attention model.

[0016] In some embodiments, the first image emotion features are extracted by: identifying multiple user facial muscle units in the first image input data based on key point detection; determining the intensity value of each user facial muscle unit among the multiple user facial muscle units; and extracting the first image emotion features based on the intensity value of each user facial muscle unit.

[0017] According to another aspect of this disclosure, an apparatus for constraining multimodal output of a digital human is provided. The apparatus includes: a user emotion data acquisition unit configured to acquire, at a first moment, a first input dataset for identifying user emotions based on a first time interval, the first input dataset including at least two of first voice input data, first image input data, and first text input data; a user emotion vector generation unit configured to generate a fused first user emotion vector based on the first input dataset; a digital human emotion vector generation unit configured to generate a first digital human emotion vector based on the first user emotion vector; and a digital human emotion component determination unit configured to determine the emotion components based on the first user emotion vector. A first digital human emotion vector is defined, identifying multiple digital human emotion components, wherein the multiple digital human emotion components include facial expression modal emotion components, action modal emotion components, and speech modal emotion components; a digital human emotion component adjustment unit is configured to evaluate the consistency of the multiple digital human emotion components, wherein, in response to the multiple digital human emotion components not meeting the consistency requirements, at least one of the multiple digital human emotion components is adjusted to meet the consistency requirements; a digital human multimodal output adjustment unit is configured to adjust the output parameters of the corresponding modality of the digital human based on the multiple digital human emotion components that meet the consistency requirements, so as to achieve constrained multimodal output of the digital human.

[0018] According to another aspect of this disclosure, a computer device is provided, comprising: at least one processor; and a memory storing a computer program thereon, wherein, when executed by the at least one processor, the computer program causes the at least one processor to perform the method described above for constraining the multimodal output of a digital human.

[0019] According to another aspect of this disclosure, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, causes the processor to perform the method described above for constraining the multimodal output of a digital human.

[0020] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, causes the processor to perform the method described above for constraining the multimodal output of a digital human.

[0021] According to one or more embodiments of this disclosure, during the interaction between a digital human and a user, tone of voice, facial expressions, and behavioral actions can be adjusted based on the user's emotions and feedback. Furthermore, during the adjustment process, multimodal features such as voice, facial expressions, and text are integrated to generate continuously adjustable digital human output parameters. Considering the multimodal output characteristics of digital human interaction, the digital human interaction adjustment method of this disclosure sets consistency constraints for the output of each modality of the digital human, thereby achieving natural digital human interaction adjustment that is beneficial to user experience.

[0022] These and other aspects of this disclosure will be apparent from the embodiments described below, and will be elucidated with reference to the embodiments described below. Attached Figure Description

[0023] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of this disclosure. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0024] Figure 1 This is a flowchart illustrating a portion of the process of a method for constraining the multimodal output of a digital human according to an exemplary embodiment.

[0025] Figure 2 This is a flowchart illustrating the consistency evaluation process of digital human emotion components according to an exemplary embodiment.

[0026] Figure 3 This is a flowchart illustrating the process of adjusting the emotional components of a digital human according to an exemplary embodiment.

[0027] Figure 4 This is a flowchart illustrating another part of the process of a method for constraining the multimodal output of a digital human according to an exemplary embodiment.

[0028] Figure 5 This is a flowchart illustrating another part of the process of a method for constraining the multimodal output of a digital human according to an exemplary embodiment.

[0029] Figure 6 This is a schematic block diagram illustrating an apparatus for constraining the multimodal output of a digital human according to an exemplary embodiment.

[0030] Figure 7This is a block diagram illustrating an exemplary computer device that can be applied to an exemplary embodiment. Detailed Implementation

[0031] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0032] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0033] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. As used herein, the term "multiple" means two or more, and the term "based on" should be interpreted as "at least partially based on". Furthermore, the terms "and / or" and "at least one of..." cover any one of the listed items and all possible combinations thereof.

[0034] In related technologies, interactive digital human systems still rely on fixed scripts or rule engines as their core driving mechanism. The system triggers corresponding speech synthesis and facial animation sequences based on pre-defined text flows. While this "playback-style" interaction can complete basic question-and-answer sessions, it severely lacks the ability to perceive and understand the user's emotional state in real time. This results in monotonous and stiff digital human responses, failing to simulate the crucial emotional continuity and resonance in human dialogue, severely limiting the naturalness of the interaction and the user's immersive experience. At the emotion perception level, existing technologies largely rely on single-modal information, such as obtaining single-dimensional emotional signals solely through voice emotion analysis (e.g., judging emotions from tone and speed) or analyzing user facial expressions (e.g., smiling, frowning) via cameras; or offline processing, such as recording first, then analyzing, and then driving the digital human to react, making it difficult to achieve real-time, integrated multimodal emotion capture. At the emotion expression and regulation level, existing technologies lack continuous, adaptive dynamic adjustment mechanisms. Specifically, the emotional output of traditional digital humans is usually based on a simple "trigger-response" rule base or a limited state machine. For example, when the keyword "happy" is detected, the digital human plays a fixed smiling animation. This output is discrete and programmed, unable to adapt smoothly and coherently to subtle changes in the user's emotions (such as the transition from pleasure to surprise) or fluctuations in intensity. The digital human's tone, speech rate, and facial expressions thus appear mechanical and lack dynamic variation. At the interaction loop level, existing technologies lack a feedback loop that takes into account the user's real-time reactions to the digital human's feedback and adjusts the digital human's behavior accordingly, resulting in a rigid interaction process and failing to achieve truly empathetic and dynamic dialogue.

[0035] To address this, embodiments of this disclosure provide a method for constraining the multimodal output of a digital human. This method enables real-time, multimodal, continuous, and closed-loop feedback-enabled adaptive adjustment of the digital human's emotions, ensuring consistent emotional expression across facial expressions, speech, and actions, thereby enhancing the naturalness and continuity of multimodal interaction. The method simultaneously analyzes the user's multimodal emotional signals, including speech, facial expressions, and actions, and drives the digital human's tone and expressions to make subtle, dynamic, and continuous adjustments. It also monitors user feedback in real time to form a closed-loop optimization, ultimately achieving highly human-like, natural, and fluent emotional interaction. This overcomes the limitations of existing technologies and significantly improves the digital human interaction experience.

[0036] Figure 1 This is a flowchart illustrating a portion of the process of a method 100 for constraining multimodal output of a digital human according to an exemplary embodiment. Figure 1 As shown, the method for constraining the multimodal output of a digital human includes steps S110, S120, S130, S140, S150 and S160.

[0037] In step S110, based on a first time interval, a first input dataset for recognizing user emotions is acquired at a first moment. The first input dataset includes at least two of the following: first voice input data, first image input data, and first text input data.

[0038] In the example, the first time interval can be a preset time interval, such as 10s, 5s, 1s, or less. At each first time interval, a first input dataset is collected at the first moment, which is then used to identify the user's real-time emotions.

[0039] In the example, different types of input data in the first input dataset can appear in different application scenarios. For example, the first voice input data may appear in application scenarios such as intelligent customer service, virtual anchors, and online education; the first image input data may appear in application scenarios such as emotional companionship, intelligent customer service, and haptic interaction; and the first text input data may appear in application scenarios such as virtual anchors and document processing.

[0040] In the example, different types of input data in the first input dataset can be acquired in different ways. For example, the first voice input data can be captured by the terminal's microphone, the first image input data can be captured by the terminal's camera or a separate camera device, and the first text input data can be acquired by the corresponding text box control.

[0041] In the example, by acquiring at least two types of input data, multimodal emotional features can be extracted and used in the subsequent process of identifying user emotions.

[0042] In step S120, a fused first user sentiment vector is generated based on the first input dataset. Step S120 further includes:

[0043] A first emotion feature set is extracted based on a first input dataset. The first emotion feature set includes at least two of the following: first speech emotion features, first image emotion features, and first text emotion features.

[0044] A first emotion vector set is generated based on a first emotion feature set, the first emotion vector set including at least two of a first speech emotion vector, a first image emotion vector, and a first text emotion vector; and

[0045] At least two emotion vectors from the first emotion vector set are fused based on a deep learning model to generate a first user emotion vector.

[0046] In the example, a first emotion feature set is extracted based on the acquired first input dataset. Different types of input data can correspond to different feature extraction methods. The first speech emotion feature can be obtained by extracting acoustic features from the first speech input data. Acoustic features can include speech intensity, speech rate, pitch, Mel-frequency cepstral coefficients (MFCC), etc.

[0047] In the example, the first text sentiment features can be extracted through text preprocessing and text encoding. Text preprocessing can include standardizing character format (such as converting full-width characters to half-width characters and normalizing English capitalization), word segmentation (i.e., dividing the first text input data into basic units of the vocabulary), and formatting. Text encoding uses Transformer models such as BERT / RoBERTa to transform the preprocessed first text input data into first text sentiment features.

[0048] In the example, the emotional features of the first image can be obtained by extracting the user's facial expressions and movements from the first image input data. The process of obtaining the emotional features of the first image further includes:

[0049] Based on key point detection, multiple user facial muscle units are identified in the first image input data;

[0050] Determine the intensity value of each user's facial muscle unit in multiple user facial muscle units;

[0051] Based on the intensity value of each user's facial muscle unit, the emotion features of the first image are extracted.

[0052] In the example, the first image input data can be a video frame from a camera video stream. The location of the user's face is identified from the video frame, and the location of feature points on the user's face is determined using a facial keypoint model (e.g., a 68-point model or a 98-point model). This allows the determination of the deformation features of the user's face to identify multiple facial motion units (AUs), which are the smallest constituent units of facial expressions in a Face Activity Coding System (FACS).

[0053] In the example, the intensity value of each user's facial muscle unit can be determined based on the distance and angle of multiple feature points corresponding to each user's facial muscle unit.

[0054] In the example, based on the set of intensity values ​​for each user's facial muscle units, a first image emotion feature with a multi-dimensional representation can be obtained.

[0055] In the example, based on the categories of emotion features in the obtained first emotion feature set, corresponding first speech emotion vector, first image emotion vector, and first text emotion vector can be generated respectively, thereby obtaining a first emotion vector set containing at least two modal emotion vectors.

[0056] In the example, the initial speech emotion features can be input into deep learning models such as GRU, LSTM, and Transformer to generate an initial speech emotion vector. This can identify emotions such as excitement / calm, harshness / gentleness, tension / relaxation, etc.

[0057] In the example, the first image emotion vector is generated by inputting the intensity values ​​of multiple facial muscle units into models such as MLP, CNN, and Transformer.

[0058] In the example, the first text sentiment feature can be further processed by Transformer models such as BERT / RoBERTa to output the first text sentiment vector according to the set sentiment vector dimension.

[0059] In the example, an emotion embedding (or emotion vector) is a way of representing an emotional state using multidimensional continuous numerical values. For example, an emotion vector can be described based on an emotion valence-arousal-dominance (VAD) model space, a pleasure-activation-advantage (PAD) model space, or other combinations of meta-emotions.

[0060] In the example, a user's sentiment vector might contain the following dimensions:

[0061] Joy; Surprise; Anxiety; Calm.

[0062] At any given moment, the vector [0.6, 0.2, 0.5, 0.1] represents a user primarily experiencing a mixed state of pleasure and anxiety, with a slight sense of surprise and relaxation. The digital human can adjust its facial expressions, body language, and vocal characteristics based on this vector to achieve multimodal, continuous, and natural emotional expression.

[0063] The user's emotional state at the first moment t is represented as the user's emotional state vector E_u(t):

[0064] E_u(t) = [e_1(t), e_2(t), …, e_N(t)]

[0065] In this context, each dimension element e_1(t), e_2(t), …, e_N(t) can be a continuous value, used to characterize the intensity of the user's state across different emotional dimensions. The user vector is updated along preset time intervals (e.g., the first time interval) during continuous interaction.

[0066] In the example, the resulting first speech emotion vector, first image emotion vector, or first text emotion vector is unimodal, derived from a specific type of input data. Therefore, the user's emotional state described by a single emotion vector is one-sided and easily affected by interference. Specifically, the text in the first text input data may conceal the user's true emotions, the speech in the first speech input data is susceptible to noise interference, and the user's facial expressions in the first image input data are brief and difficult to capture accurately. These factors all contribute to the inability of the corresponding first speech emotion vector, first image emotion vector, or first text emotion vector to fully reflect the user's emotional state. The emotion vector fusion of this embodiment, through complementarity and cross-validation, can construct a more complete, accurate, and robust emotion representation—it can both compensate for the information gaps of a single modality and resolve ambiguities and contradictions between modalities, thereby achieving dynamic understanding and tracking of the user's complex and continuous emotions, providing a reliable basis for digital humans to achieve natural and empathetic feedback interactions.

[0067] In the example, the deep learning model can be a multilayer perceptron model, a convolutional neural network model, or an attention model, which are commonly used in multimodal fusion scenarios.

[0068] Continue to refer to Figure 1 In step S130, a first digital human emotion vector is generated based on the first user emotion vector. In this example, the first digital human emotion vector can be generated by combining multiple pieces of information. For instance, the interaction scenarios of digital humans are highly diverse. Digital humans acting as emotional companions, professional digital humans, etc., have different role styles and settings, and their responses to user emotional states will also differ. Therefore, a first digital human emotion vector E_target(t) can be generated based on the first user emotion vector E_u(t) and the digital human role style. The first digital human emotion vector is used to bridge user emotions and the multimodal action output of the digital human, realizing the coordinated mapping of emotion-driven actions of the digital human.

[0069] In step S140, based on the first digital human emotion vector, multiple digital human emotion components are determined, including an expression modal emotion component c_{face}(t), a gesture modal emotion component c_{gesture}(t), and a speech modal emotion component c_{prosody}(t). In the example, in the digital human's interaction scenario, multimodal outputs such as facial expressions, speech prosody, and gestures are characterized by high visualization, multiple degrees of freedom, and strong perceptuality. The actions of each modality will be based on the emotion components of different modalities. The output is not a complete emotion vector, but rather a low-dimensional or scalarized expression obtained from the same digital human emotion vector for each modality, used to characterize the dominant emotion in that modality.

[0070] In step S150, the consistency of multiple digital human emotion components is evaluated, wherein, in response to multiple digital human emotion components not meeting the consistency requirements, at least one of the multiple digital human emotion components is adjusted to meet the consistency requirements.

[0071] Figure 2 This is a flowchart illustrating the consistency evaluation process of digital human emotion components according to an exemplary embodiment. Further reference... Figure 2 The consistency assessment process of the digital human emotion components in step S150 further includes S210 and S220.

[0072] In step S210, a consistency metric is calculated between every two digital human emotion components among the multiple digital human emotion components.

[0073] In the example, the digital human emotion component is a scalar, and the consistency metric is the absolute value of the difference between any two digital human emotion components.

[0074] In step S220, the consistency of multiple digital human emotion components is evaluated based on a comparison of the consistency metric and the consistency threshold. Specifically, the consistency metric needs to simultaneously satisfy:

[0075] |c_{face}(t) c_{prosody}(t)|<ε_1;

[0076] |c_{face}(t) c_{gesture}(t)|<ε_2;

[0077] |c_{prosody}(t) c_{gesture}(t)|<ε_3.

[0078] Among them, ε_1, ε_2, and ε_3 are the consistency thresholds preset according to the actual application needs.

[0079] Figure 3 This is a flowchart illustrating the process of adjusting the emotion components of a digital human according to an exemplary embodiment, such as... Figure 3 As shown, the digital human emotion component adjustment process in step S150 further includes steps S310, S320 and S330.

[0080] In step S310, the first adjustment priority is determined based on the interaction scenario of the digital human.

[0081] In step S320, a second adjustment priority is determined based on the consistency metric value.

[0082] In the example, as mentioned above, the character style and persona of the digital human vary greatly across different interaction scenarios. The facial expressions, vocal rhythms, and gestures output by different digital humans have different weights in terms of their importance to the interactive expression. Therefore, the priority of adjusting the emotional components of each modality of the digital human can be determined based on this weight. For example, in an intelligent customer service scenario, users may pay more attention to the digital human's voice expression; therefore, the voice modality emotional component will have the highest priority. Furthermore, the priority order of the facial expression modality emotional component and the action modality emotional component can be determined based on the interaction scenario.

[0083] In the example, during the calculation of the consistency metric, it can be further determined that the emotional component of a certain modality has a larger deviation than the other two emotional components. Therefore, a second adjustment priority can be determined accordingly to prioritize the adjustment of the emotional components that cause inconsistency in the digital human's emotional expression.

[0084] In step S330, at least one digital human emotion component that needs to be adjusted is determined according to the first adjustment priority and the second adjustment priority.

[0085] In the example, the first adjustment priority and the second adjustment priority can be combined (e.g., weighted) to ultimately determine at least one digital human emotion component that needs adjustment.

[0086] Continue to refer to Figure 1 In step S160, based on multiple digital human emotion components that meet consistency requirements, the output parameters of the corresponding modal of the digital human are adjusted to achieve constrained multimodal output of the digital human.

[0087] In the example, the digital human's multimodal output at the first time t corresponds to the digital human's motion state vector S_{motion}(t):

[0088] S_{motion}(t) = [S_{face}(t), S_{gesture}(t), S_{prosody}(t)]

[0089] Where S_{face}(t) represents facial expression parameters; S_{gesture}(t) represents limb movement parameters; and S_{prosody}(t) represents speech emotion parameters.

[0090] Figure 4 This is a flowchart illustrating another part of the process of a method 100 for constraining the multimodal output of a digital human according to an exemplary embodiment. Figure 2 As shown, the method 100 for constraining the multimodal output of a digital human further includes steps S410, S420, S430, S440 and S450.

[0091] In step S410, a second input dataset for identifying user emotions is acquired at a second time point, wherein the second time point is later than the first time point by a first time interval, and the second input dataset includes at least two of the following: second voice input data, second image input data, and second text input data.

[0092] In step S420, a fused second user sentiment vector is generated based on the second input dataset.

[0093] In steps S410-S420, the method for obtaining the second input dataset and the method for generating the second user emotion can be corresponding to steps S110-S120, so they will not be described in detail here.

[0094] It should be noted that the first time interval can be a preset fixed sampling period. The first moment and the second moment are any two adjacent moments in the digital human interaction process based on this fixed sampling period.

[0095] In step S430, the second user emotion vector is updated based on the first user emotion vector using a weighted update algorithm.

[0096] In the example, the second user emotion vector initially generated at the second time t+1 is equivalent to the user feedback vector E_u^{f}(t) after the digital human has completed emotion state adjustment, and the user emotion state vector is updated as follows:

[0097] E_u(t+1) = (1 β)·E_u(t) + β·E_u^{f}(t)

[0098] Wherein, β is the state update coefficient, which is used to adjust the response speed to new feedback emotions in order to suppress the instantaneous fluctuations of user emotions and improve system stability.

[0099] By combining the first user's emotion vector and the second user's emotion vector, a perception-decision closed loop is constructed. Based on this, the method of this disclosure can not only react based on the user's current emotional state, but also evaluate the user's feedback on the emotional state regulated by the digital human, and use this evaluation to update the second digital human's emotion vector, thereby improving the naturalness of the interaction.

[0100] In step S440, the second digital human emotion vector is determined based on the updated second user emotion vector and the first digital human emotion vector.

[0101] In the example, the second digital human's emotion vector at time t+1 is determined according to the following relationship:

[0102] E_target(t+1) = G(E_target(t), E_u(t+1)).

[0103] G() is a mapping function used to coordinate and update the digital human's emotion vector at the previous moment and the user's current emotion vector.

[0104] In the example, step S440 further includes:

[0105] Based on the updated second user emotion vector, determine the initial second digital human emotion vector;

[0106] The initial second digital human emotion vector is adjusted based on a comparison between the initial second digital human emotion vector and the first digital human emotion vector and a preset digital human emotion change threshold.

[0107] In the example, the digital human's facial expressions and body movements are high-dimensional, continuously controlled objects, and humans are extremely sensitive to their rate of change. If the parameters of emotion-driven actions change drastically within a short period of time, they are easily perceived as unnatural or even abrupt movements. This problem is significantly different from traditional interactive systems that only contain voice or text output, and is a unique technical challenge in the regulation of emotion-driven actions in digital humans. Therefore, in order to make the changes in the digital human's interactive output smooth, it is necessary to determine the modulus ||E_target(t+1) of the initial second digital human emotion vector and the first digital human emotion vector. Whether E_target(t)|| does not exceed a preset threshold δ for digital human emotion changes. Here, δ is used to limit the rate of change of the digital human's emotional actions over time. To maximize the protection scope, δ can be a scalar threshold, used to limit the overall amplitude of the digital human's emotional actions. In other embodiments, δ can also be a vector threshold, used to limit the rate of change of different emotional action dimensions separately. Through the above constraints, the digital human's facial expressions, body movements, and vocal emotions maintain a smooth evolution over continuous time, thereby effectively avoiding the problem of action jumps caused by short-term fluctuations in the user's emotions.

[0108] If the magnitude of the initial second digital human emotion vector and the first digital human emotion vector does not exceed the preset digital human emotion change threshold, then the initial second digital human emotion vector is used as the second digital human emotion vector to drive digital human interaction. If the magnitude of the initial second digital human emotion vector and the first digital human emotion vector exceeds the preset digital human emotion change threshold, then the value of at least one element in the initial second digital human emotion vector is adjusted until its magnitude and that of the first digital human emotion vector do not exceed the preset digital human emotion change threshold.

[0109] In step S450, the output parameters of the corresponding modality of the digital human are adjusted based on the second digital human emotion vector.

[0110] In step S450, the constrained digital human multimodal output method at the second time moment can correspond to steps S130-S160, so it will not be described in detail here.

[0111] Figure 5 This is a flowchart illustrating another part of a method for constraining the multimodal output of a digital human according to an exemplary embodiment. Figure 2 As shown, the method 100 for constraining the multimodal output of a digital human further includes steps S510, S520, S530, S540 and S550.

[0112] In step S510, a third input dataset for identifying user emotions is acquired at a third time point, wherein the third time point is at least one first time interval later than the first time point, and the third input dataset includes at least two of the following: third voice input data, third image input data, and third text input data.

[0113] In step S520, a third emotion feature set is extracted based on the third input dataset. The third emotion feature set includes at least two of the following: third speech emotion features, third image emotion features, and third text emotion features.

[0114] In step S530, a third emotion vector set is generated based on the third emotion feature set. The third emotion vector set includes at least two of the following: third speech emotion vector, third image emotion vector, and third text emotion vector.

[0115] In step S540, at least two emotion vectors in the third emotion vector set are fused based on a deep learning model to generate a third user emotion vector.

[0116] In steps S510-S540, the methods for obtaining the third input dataset, extracting the third emotion feature set, and generating the third emotion vector set can correspond to those in step S120, so they will not be described in detail here.

[0117] In step S550, based on the comparison between the difference between the third user emotion vector and the first user emotion vector and the preset difference between user emotion vectors, the first time interval is updated to the second time interval.

[0118] It should be noted that the third time point is separated from the first time point by one or more first time intervals. Therefore, the third input dataset, the third emotion feature set, and the third emotion vector set can be directly obtained from the second input dataset, the second emotion feature set, and the second emotion vector set at the second time point without needing to be obtained separately.

[0119] In the example, different users exhibit varying degrees of emotional fluctuation. For users with minor emotional fluctuations, the frequency of the digital human's multimodal output adjustment can be reduced, while for users with significant emotional fluctuations, the frequency can be appropriately increased. Based on this, at least one preset user emotion vector difference can be established. This preset difference can represent user emotional states such as no emotional fluctuation, minor emotional fluctuation, and / or significant emotional fluctuation. By comparing the difference between the third user's emotion vector and the first user's emotion vector with the preset difference, the user's emotional fluctuation at that moment can be determined, thereby adaptively adjusting the frequency of the digital human's multimodal output adjustment, thus contributing to personalized digital human interaction.

[0120] To better illustrate the method 100 for constraining the multimodal output of a digital human according to the embodiments of this disclosure, and especially to illustrate the application background of multimodal emotion fusion and the method for constraining the multimodal output of a digital human, the following will describe in detail the constraints on the multimodal output of a digital human and the adjustment of the interaction mode in two specific application scenarios.

[0121] Example 1: Intelligent customer service scenario.

[0122] 1. User interaction methods.

[0123] Users interact with the digital human customer service representative in real time via camera and microphone. Users may speak and display facial expressions simultaneously. Additionally, users can input messages via text.

[0124] 2. User emotion recognition.

[0125] • Voice emotion recognition: Extract acoustic features (energy, pitch, speech rate, etc.) and use LSTM or GRU deep learning models to generate voice emotion vectors E_voice.

[0126] • Video facial expression recognition: Generates image emotion vectors E_face through facial landmark detection (68-point model) and muscle action unit (AU) analysis.

[0127] • Text sentiment recognition: Generate text sentiment vectors E_text from user text input using language models such as BERT.

[0128] • Multimodal emotion fusion: Using a cross-modal attention mechanism, the emotion vectors of speech, images, and text are fused to generate a fused user emotion vector E_u(t) to represent the user's real-time emotional state.

[0129] 3. Digital human target emotion mapping.

[0130] Based on the user's emotion vector E_u(t), the system determines the digital human's target emotion state vector E_target(t), which drives the digital human's facial expressions, vocal emotions, and gestures. The target emotion mapping can be achieved through a trained mapping function, mapping the user's complex emotional states to the digital human's responsive emotional states.

[0131] 4. Generation of emotion-driven action state parameters.

[0132] Based on the digital human's target emotion state vector E_target(t), which drives the digital human's facial expressions, vocal emotions, and gestures, the digital human's target emotion state vector is mapped to multimodal action state parameters:

[0133] • Voice parameters: speech rate, pitch, and volume can be continuously adjusted;

[0134] • Facial expression parameters: When using BlendShape for digital human facial expression modeling, the continuous change of BlendShape weights reflects expressions such as smiling and frowning;

[0135] • Micro-expression movements: Continuous adjustment of blinking frequency and breathing amplitude.

[0136] • The digital human's output parameters change smoothly, ensuring a natural interaction.

[0137] 5. Multimodal output constraints.

[0138] • Emotional consistency constraint: Ensure that the dominant emotional components of digital human facial expressions, speech rhythm and gestures are consistent to avoid modality mismatch.

[0139] • Smooth Change Constraint: Limits the range of change in the digital human's motion parameters between adjacent time points to avoid abrupt changes in expressions or movements, ensuring continuity and naturalness.

[0140] 6. Closed-loop adaptive adjustment.

[0141] • Capture user feedback emotional information (changes in voice tone, facial expression, etc.) to the digital human's output, forming a feedback emotional vector E_u^f(t).

[0142] • Using a simple weighted update algorithm, the feedback emotion vector is combined with the original emotion vector to dynamically adjust the digital human's output parameters, enabling the digital human to continuously adapt to changes in the user's emotions:

[0143] • E_u(t+1) = (1 β)·E_u(t) + β·E_u^f(t).

[0144] • The interaction process involves cyclically recognizing user emotions, adjusting the digital human's emotional state, and providing user emotional feedback, thereby achieving a continuous and natural closed-loop adaptive interaction.

[0145] Example 2: Virtual anchor scenario.

[0146] 1. User interaction methods.

[0147] Users can interact with virtual anchor digits in real time via voice and bullet comments.

[0148] 2. User emotion recognition.

[0149] • Speech emotion recognition: Extract acoustic embeddings and generate speech emotion vectors using a Transformer model.

[0150] • Text sentiment recognition: The sentiment vector of the bullet screen text is extracted using an NLP model (such as the sentiment analysis BERT model).

[0151] • Multimodal emotion fusion: Using an attention fusion method, speech and text emotion vectors are combined to generate a fused user emotion vector to represent the user's real-time emotional state.

[0152] 3. Digital human output adjustment.

[0153] The digital human's output parameters are continuously adjusted based on the fused user emotion vectors. For example:

[0154] • When users are excited: the digital human's voice tone rises, the speech rate increases, facial expressions become exaggerated, and gestures become more pronounced;

[0155] • When the user is anxious: the digital human lowers its tone of voice, slows down its speech, softens its facial expressions, and uses gentle gestures.

[0156] 4. Closed-loop adaptive adjustment (emotional closed-loop control module)

[0157] • Continuously capture user emotional feedback (voice, facial expressions, changes in the emotional state of bullet comments) to form a feedback emotional vector.

[0158] • By using weighted update or reinforcement learning strategies, the feedback emotion vector is fused with the original emotion vector to dynamically adjust the digital human's output parameters.

[0159] • Circularly update the user's emotion vector to achieve continuous and adaptive closed-loop digital human multimodal output adjustment.

[0160] The method for constraining the multimodal output of a digital human according to the embodiments of this disclosure has the following beneficial effects:

[0161] 1. Model user emotions as a continuously updated system state variable to form a clear closed-loop control path;

[0162] 2. Introduce a target emotion state vector to ensure a stable mapping between user emotions and digital human actions;

[0163] 3. To address the high sensitivity of digital human motion output, constraints on consistency of emotion and motion and smooth changes are introduced;

[0164] 4. It is significantly different from existing technical solutions that rely solely on speech or text feature fusion or dynamic weight adjustment;

[0165] 5. Significantly improves the naturalness, stability, and interactive experience of digital human emotional expression.

[0166] Embodiments of this disclosure also provide an apparatus for constraining the multimodal output of a digital human.

[0167] Figure 6 This is a schematic block diagram illustrating an apparatus for constraining the multimodal output of a digital human according to an exemplary embodiment.

[0168] like Figure 6 As shown, in some embodiments, the apparatus 600 for constraining the multimodal output of a digital human includes a user emotion data acquisition unit 610, a user emotion vector generation unit 620, a digital human emotion vector generation unit 630, a digital human emotion component determination unit 640, a digital human emotion component adjustment unit 650, and a digital human multimodal output adjustment unit 660.

[0169] The user emotion data acquisition unit 610 is configured to acquire a first input dataset for recognizing user emotions at a first moment based on a first time interval. The first input dataset includes at least two of first voice input data, first image input data, and first text input data.

[0170] The user emotion vector generation unit 620 is configured to generate a fused first user emotion vector based on the first input dataset;

[0171] The digital human emotion vector generation unit 630 is configured to generate a first digital human emotion vector based on a first user emotion vector.

[0172] The digital human emotion component determination unit 640 is configured to determine multiple digital human emotion components based on a first digital human emotion vector, wherein the multiple digital human emotion components include facial expression modal emotion components, action modal emotion components and voice modal emotion components;

[0173] The digital human emotion component adjustment unit 650 is configured to evaluate the consistency of multiple digital human emotion components, wherein, in response to multiple digital human emotion components not meeting the consistency requirements, at least one of the multiple digital human emotion components is adjusted to meet the consistency requirements.

[0174] The digital human multimodal output adjustment unit 660 is configured to adjust the output parameters of the corresponding modality of the digital human based on multiple digital human emotion components that meet consistency requirements, so as to achieve constrained multimodal output of the digital human.

[0175] While specific functions have been discussed above with reference to specific units, it should be noted that the functions of the units discussed herein may be divided into multiple units, and / or at least some functions of multiple units may be combined into a single unit. The specific unit performing an action discussed herein includes the specific unit itself performing the action, or alternatively, the specific unit calling or otherwise accessing another component or unit that performs the action (or performs the action in conjunction with the specific unit). Therefore, a specific unit performing an action may include the specific unit performing the action itself and / or another unit that performs the action, called or otherwise accessed by the specific unit.

[0176] It should also be understood that the above regarding Figure 6 The described units can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these units can be implemented as computer program code / instructions configured to execute in one or more processors and stored in a computer-readable storage medium. Alternatively, these units can be implemented as hardware logic / circuit. For example, these units can be implemented together in a System on Chip (SoC). An SoC may include an integrated circuit chip (which includes one or more components in a processor (e.g., a Central Processing Unit (CPU), microcontroller, microprocessor, digital signal processor (DSP), etc.), memory, one or more communication interfaces, and / or other circuitry) and may optionally execute received program code and / or include embedded firmware to perform functions.

[0177] According to one aspect of this disclosure, a computer device is also provided, including a memory, a processor, and a computer program stored in the memory. The processor is configured to execute the computer program to implement the steps of any of the method embodiments described above.

[0178] According to one aspect of this disclosure, a non-transitory computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the method embodiments described above.

[0179] According to one aspect of this disclosure, a computer program product is also provided, which includes a computer program that, when executed by a processor, implements the steps of any of the method embodiments described above.

[0180] In the following text, combined with Figure 7 Illustrative examples describing such computer devices, non-transitory computer-readable storage media, and computer program products.

[0181] Figure 7 An example configuration of a computer device 700 that can be used to implement the methods described herein is shown.

[0182] Computer device 700 can be a variety of different types of devices. Examples of computer device 700 include, but are not limited to: desktop computers, server computers, laptop or netbook computers, mobile devices (e.g., tablet computers, cellular or other wireless phones (e.g., smartphones), notebook computers, mobile stations), wearable devices (e.g., glasses, watches), entertainment devices (e.g., entertainment appliances, set-top boxes communicatively coupled to a display device, game consoles), televisions or other display devices, automotive computers, and so on.

[0183] Computer device 700 may include at least one processor 702, memory 704, multiple communication interfaces 706, display device 708, other input / output (I / O) devices 710, and one or more mass storage devices 712 capable of communicating with each other, such as via system bus 714 or other suitable connections.

[0184] Processor 702 may be a single processing unit or multiple processing units, and all processing units may include single or multiple computing units or multiple cores. Processor 702 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operating instructions. Among other capabilities, processor 702 may be configured to acquire and execute computer-readable instructions stored in memory 704, mass storage device 712, or other computer-readable media, such as program code of operating system 716, program code of application program 718, program code of other program 720, etc.

[0185] Memory 704 and mass storage device 712 are examples of computer-readable storage media for storing instructions that are executed by processor 702 to perform the various functions described above. For example, memory 704 can generally include both volatile and non-volatile memory (e.g., RAM, ROM, etc.). Furthermore, mass storage device 712 can generally include hard disk drives, solid-state drives, removable media, including external and removable drives, memory cards, flash memory, floppy disks, optical disks (e.g., CDs, DVDs), storage arrays, network-attached storage, storage area networks, etc. Both memory 704 and mass storage device 712 can be collectively referred to herein as memory or computer-readable storage media, and can be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program code, which can be executed by processor 702 as a specific machine configured to perform the operations and functions described in the examples herein.

[0186] Multiple programs may be stored on mass storage device 712. These programs include operating system 716, one or more application programs 718, other programs 720, and program data 722, and they may be loaded into memory 704 for execution. Examples of such application programs or program modules may include computer program logic (e.g., computer program code or instructions) of the methods described herein and / or other embodiments described herein.

[0187] Although the operating system 716, one or more applications 718, other programs 720, and program data 722 are in Figure 7 The data is illustrated as being stored in memory 704 of computer device 700, but operating system 716, one or more application programs 718, other programs 720, and program data 722, or portions thereof, may also be implemented using any form of computer-readable medium accessible by computer device 700. As used herein, "computer-readable medium" includes at least two types of computer-readable media: computer-readable storage media and communication media.

[0188] Computer-readable storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, DVD, or other optical storage devices, magnetic cassettes, magnetic tapes, disk storage devices or other magnetic storage devices, or any other non-transmission medium that can be used to store information for access by computer devices. In contrast, communication media can embody computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms. Computer-readable storage media as defined herein do not include communication media.

[0189] One or more communication interfaces 706 are used for exchanging data with other devices, such as via a network, direct connection, etc. Such communication interfaces can be one or more of the following: any type of network interface (e.g., a network interface card (NIC)), wired or wireless (such as IEEE 802.11 Wireless LAN (WLAN)) wireless interface, Wi-MAX interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth™ interface, Near Field Communication (NFC) interface, etc. Communication interface 706 can facilitate communication across a variety of network and protocol types, including wired networks (e.g., LAN, cable, etc.) and wireless networks (e.g., WLAN, cellular, satellite, etc.), the Internet, etc. Communication interface 706 can also provide communication with external storage devices (not shown), such as storage arrays, network-attached storage, storage area networks, etc.

[0190] In some examples, a display device 708, such as a monitor, may be included for displaying information and images to the user. Other I / O devices 710 may be devices that receive various inputs from the user and provide various outputs to the user, and may include touch input devices, gesture input devices, cameras, keyboards, remote controls, mice, printers, audio input / output devices, and so on.

[0191] The technologies described herein can be supported by these various configurations of computer device 700, and are not limited to specific examples of the technologies described herein. For example, the functionality can also be implemented wholly or partially on a “cloud” using a distributed system. A cloud includes and / or represents a platform for resources. The platform abstracts the underlying functionality of the cloud’s hardware (e.g., servers) and software resources. Resources may include applications and / or data that can be used when performing computational processing on a server remote from computer device 700. Resources may also include services provided via the Internet and / or via subscriber networks such as cellular or Wi-Fi networks. The platform can abstract resources and functionality to connect computer device 700 to other computer devices. Therefore, the implementation of the functionality described herein can be distributed throughout the cloud. For example, the functionality can be implemented partly on computer device 700 and partly through a platform that abstracts the functionality of the cloud.

Claims

1. A method for constraining the multimodal output of a digital human, characterized in that, The method includes: Based on a first time interval, a first input dataset for recognizing user emotions is acquired at a first moment, the first input dataset including at least two of first voice input data, first image input data and first text input data; Based on the first input dataset, a fused first user sentiment vector is generated; Based on the first user's emotion vector, a first digital human emotion vector is generated; Based on the first digital human emotion vector, multiple digital human emotion components are determined, wherein the multiple digital human emotion components include facial expression modal emotion components, action modal emotion components, and voice modal emotion components; The consistency of the plurality of digital human emotion components is evaluated, wherein, in response to the plurality of digital human emotion components not meeting the consistency requirements, at least one of the plurality of digital human emotion components is adjusted to meet the consistency requirements. Based on multiple digital human emotion components that meet consistency requirements, the output parameters of the corresponding modal of the digital human are adjusted to achieve constrained multimodal output of the digital human.

2. The method according to claim 1, characterized in that, Assessing the consistency of the multiple digital human emotion components includes: Calculate the consistency metric between every two digital human emotion components among the plurality of digital human emotion components; The consistency of the multiple digital human emotion components is evaluated based on a comparison of the consistency metric and the consistency threshold.

3. The method according to claim 2, characterized in that, In response to the plurality of digital human emotion components not meeting the consistency requirement, adjusting at least one of the plurality of digital human emotion components to meet the consistency requirement includes: Based on the interaction scenario of the digital human, determine the first adjustment priority; A second adjustment priority is determined based on the consistency metric value; Based on the first adjustment priority and the second adjustment priority, determine the at least one digital human emotion component that needs to be adjusted.

4. The method according to any one of claims 1-3, characterized in that, Based on the first input dataset, the fused first user sentiment vector is generated by including: A first emotion feature set is extracted based on the first input dataset. The first emotion feature set includes at least two of the following: first speech emotion features, first image emotion features, and first text emotion features. A first emotion vector set is generated based on the first emotion feature set, and the first emotion vector set includes at least two of the following: a first speech emotion vector, a first image emotion vector, and a first text emotion vector. At least two emotion vectors from the first emotion vector set are fused using a deep learning model to generate the first user emotion vector.

5. The method according to any one of claims 1-3, characterized in that, The method further includes: A second input dataset for identifying user emotions is acquired at a second time point, wherein the second time point is later than the first time point by a first time interval, and the second input dataset includes at least two of second voice input data, second image input data, and second text input data; Based on the second input dataset, a fused second user sentiment vector is generated; The second user sentiment vector is updated based on the first user sentiment vector using a weighted update algorithm. The second digital human emotion vector is determined based on the updated second user emotion vector and the first digital human emotion vector; Based on the second digital human's emotion vector, the output parameters of the corresponding modality of the digital human are adjusted.

6. The method according to claim 5, characterized in that, Based on the second input dataset, the fused second user sentiment vector is generated as follows: A second emotion feature set is extracted based on the second input dataset. The second emotion feature set includes at least two of the following: second speech emotion features, second image emotion features, and second text emotion features. A second emotion vector set is generated based on the second emotion feature set, and the second emotion vector set includes at least two of the following: second speech emotion vector, second image emotion vector, and second text emotion vector. At least two emotion vectors in the second emotion vector set are fused based on a deep learning model to generate the second user emotion vector.

7. The method according to claim 5, characterized in that, Based on the updated second user emotion vector and the first digital human emotion vector, the second digital human emotion vector is determined as follows: Based on the updated second user emotion vector, determine the initial second digital human emotion vector; Based on the comparison between the initial second digital human emotion vector and the first digital human emotion vector and a preset digital human emotion change threshold, the initial second digital human emotion vector is adjusted.

8. The method according to claim 1, characterized in that, The method further includes: A third input dataset for identifying user emotions is acquired at a third time point, wherein the third time point is later than the first time point by at least one first time interval, and the third input dataset includes at least two of third voice input data, third image input data, and third text input data; A third emotion feature set is extracted based on the third input dataset, and the third emotion feature set includes at least two of the third speech emotion features, third image emotion features, and third text emotion features. A third emotion vector set is generated based on the third emotion feature set, and the third emotion vector set includes at least two of the third speech emotion vector, third image emotion vector, and third text emotion vector. At least two emotion vectors in the third emotion vector set are fused based on a deep learning model to generate a third user emotion vector. Based on the comparison between the difference between the third user emotion vector and the first user emotion vector and the preset difference between user emotion vectors, the first time interval is updated to the second time interval.

9. The method according to claim 4, characterized in that, The deep learning model includes a multilayer perceptron model, a convolutional neural network model, or an attention model.

10. The method according to claim 4, characterized in that, The emotion features of the first image are extracted using the following method: Based on key point detection, multiple user facial muscle units are identified in the first image input data; Determine the intensity value of each user facial muscle unit among the plurality of user facial muscle units; Based on the intensity value of each user's facial muscle unit, the emotion features of the first image are extracted.

11. A device for constraining the multimodal output of a digital human, characterized in that, The device includes: The user emotion data acquisition unit is configured to acquire a first input dataset for identifying user emotions at a first moment based on a first time interval, wherein the first input dataset includes at least two of first voice input data, first image input data and first text input data. The user emotion vector generation unit is configured to generate a fused first user emotion vector based on the first input dataset; The digital human emotion vector generation unit is configured to generate a first digital human emotion vector based on the first user's emotion vector; The digital human emotion component determination unit is configured to determine multiple digital human emotion components based on the first digital human emotion vector, wherein the multiple digital human emotion components include facial expression modal emotion components, action modal emotion components, and voice modal emotion components. A digital human emotion component adjustment unit is configured to evaluate the consistency of the plurality of digital human emotion components, wherein, in response to the plurality of digital human emotion components not meeting the consistency requirement, at least one of the plurality of digital human emotion components is adjusted to meet the consistency requirement. The digital human multimodal output adjustment unit is configured to adjust the output parameters of the corresponding modality of the digital human based on multiple digital human emotion components that meet consistency requirements, so as to achieve constrained multimodal output of the digital human.

12. A computer device, characterized in that, The computer device includes: At least one processor; A memory having a computer program stored thereon, wherein, when executed by the at least one processor, the computer program causes the at least one processor to perform the method of any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the method according to any one of claims 1-10.

14. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, causes the processor to perform the method of any one of claims 1-10.

Citation Information

Patent Citations

  • Multi-modal interaction method and system of digital human intelligent agent

    CN120653118A

  • Digital human construction method and device based on heterogeneous emotion semantic graph and long sequence emotion modeling

    CN121350948A