Digital human driving method, holographic warehouse device, electronic device, computer storage medium and program product
By acquiring user interaction conversation intent and digital human profile information, and using generative models to generate emotional feedback information, the digital human is driven to express facial expressions and body movements. This solves the problem of stiff digital human interaction, improves the naturalness and humanization of the interaction, and enhances the user experience.
Patent Information
- Application Number
- CN202510677885.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-10-28
AI Technical Summary
When existing digital humans interact with users, their facial expressions and body movements are stiff, lacking naturalness and interactive appeal, and they cannot accurately respond to users' emotions.
By acquiring intent information from user interaction sessions and combining it with digital human profile information, generative models are used to generate emotional feedback information, driving digital humans to express facial expressions and body movements, thereby enhancing the naturalness and accuracy of the interaction.
It enhances the naturalness and humanization of the interaction between digital humans and users, and strengthens users' interest and experience in interaction.
Smart Images

Figure CN120849598A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a digital human driving method, a holographic chamber device, an electronic device, a computer storage medium, and a computer program product. Background Technology
[0002] Digital humans are digital human figures created using digital technology that closely resemble human appearances. They can also be understood as digital or virtual characters or individuals existing in a virtual environment, such as animated characters, virtual assistants, or AI (Artificial Intelligence) generated figures.
[0003] With the continuous development of immersive interactive environments, the application of digital humans is becoming increasingly widespread. However, currently, when digital humans interact with users, the implementation of interactive expressions is relatively simple and rigid, failing to accurately respond to users, resulting in a low level of user experience and a lack of appeal in the interactive interaction. Summary of the Invention
[0004] In view of this, embodiments of this application provide a digital human driving solution to at least partially solve the above-mentioned problems.
[0005] According to a first aspect of the embodiments of this application, a digital human driving method is provided, comprising: acquiring an interaction session sent by a user to a digital human; performing intent analysis on the interaction session to obtain intent information of the interaction session; constructing a first prompt word based on the interaction session, the intent information, and digital human profile information of the digital human; generating emotional feedback information for the digital human in response to the user's emotional state based on the first prompt word through a first generative model; and driving the digital human to perform facial expressions and body movements corresponding to the emotional feedback information based on the emotional feedback information.
[0006] According to a second aspect of the embodiments of this application, a holographic pod device is provided, comprising at least: a display screen, an interaction interface, and a processor; wherein: the display screen is used to display a digital human; the interaction interface is used to receive an interaction session sent by a user to the digital human; the processor is used to upload the interaction session to a server, and receive from the server, based on the interaction session, intent information of the interaction session, and digital human profile information, to generate and drive the digital human to perform facial expressions and body movements, wherein the emotions expressed by the facial expressions and body movements are adapted to the emotions expressed by the interaction session; the display screen is also used to display the digital human having the facial expressions and body movements.
[0007] According to a third aspect of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, wherein the executable instruction causes the processor to perform an operation corresponding to the method described in the first aspect.
[0008] According to a fourth aspect of the embodiments of this application, a computer storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0009] According to a fifth aspect of the embodiments of this application, a computer program product is provided, including computer instructions that instruct a computing device to perform an operation corresponding to the method described in the first aspect.
[0010] According to the solution provided in the embodiments of this application, when a user interacts with a digital human, on the one hand, the intention of the interaction conversation sent by the user to the digital human is analyzed to obtain the corresponding intention information. This intention information can provide important contextual emotion judgment basis for subsequently determining the user's emotional state, making the emotions indicated by the subsequent emotional feedback information determined for the digital human more accurate and detailed. On the other hand, the first prompt word used by the first generative model is constructed based on the interaction conversation, intention information, and the digital human's digital human profile information. Thus, when generating emotional feedback information, the first generative model can fully consider both the user's emotional state and the digital human's situation, making the generated emotional feedback information more consistent with the digital human's persona setting and richer and more detailed. As a result, when the digital human interacts with the user, it can express itself with more detailed and precise facial expressions and body movements, accurately responding to the user. Moreover, this response is more natural and more humanized, thereby improving the user's interaction experience with the digital human and increasing the user's interest and attraction in interacting with the digital human. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0012] Figure 1 A schematic diagram of an exemplary system to which the embodiments of this application are applicable;
[0013] Figure 2A This is a flowchart illustrating the steps of a digital human driving method according to an embodiment of this application;
[0014] Figure 2B for Figure 2A A schematic diagram of the algorithm process of an exemplary algorithm of the embodiment shown;
[0015] Figure 3 This is a schematic diagram of the structure of a holographic chamber device according to an embodiment of this application;
[0016] Figure 4 This is a schematic diagram illustrating the process of driving a digital human according to an embodiment of this application;
[0017] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0018] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of this application.
[0019] The specific implementation of the embodiments of this application will be further described below with reference to the accompanying drawings.
[0020] Figure 1 An exemplary system applicable to embodiments of this application is shown. For example... Figure 1 As shown, the system 100 may include a server 102, a communication network 104, and / or one or more user equipment 106. Figure 1 The example in the text shows multiple user devices.
[0021] Server 102 can be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to distributed storage system devices, server clusters, computing cloud server clusters, etc. In some embodiments, server 102 can perform any suitable function. For example, in some embodiments, server 102 can be used for digital human driving, such as driving the digital human to express facial expressions and body movements. As an optional example, in some embodiments, server 102 can be used to acquire interactive sessions sent by the user to the digital human, perform intent analysis on the interactive sessions to obtain intent information of the interactive sessions; construct a first prompt word based on the interactive sessions, intent information, and digital human profile information of the digital human; generate emotional feedback information for the digital human in response to the user's emotional state based on the first prompt word through a first generative model; and drive the digital human to express facial expressions and body movements corresponding to the emotional feedback information based on the emotional feedback information. As another example, in some embodiments, server 102 can acquire interactive sessions sent by the user to the digital human from user device 106, and send information or instructions to user device 106 to drive the digital human to express facial expressions and body movements, etc.
[0022] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The user equipment 106 can be connected to the communication network 104 via one or more communication links (e.g., communication link 112), and the communication network 104 can be linked to the cloud server 102 via one or more communication links (e.g., communication link 114). The communication link can be any communication link suitable for transmitting data between the user equipment 106 and the cloud server 102, such as a network link, a dial-up link, a wireless link, a hardwired link, any other suitable communication link, or any suitable combination of such links.
[0023] User device 106 may include any one or more user devices suitable for presenting images and interacting with a user. In some embodiments, user device 106 may interact with the user, receive interactive sessions sent by the user to the digital human, and upload them to server 102. As an optional example, in some embodiments, user device 106 may also receive information or instructions sent by server 102 to drive the digital human to make facial expressions and body movements, and display these facial expressions and body movements on a display screen according to the information or instructions to interact with the user. In some embodiments, user device 106 may include any suitable type of device. For example, in some embodiments, user device 106 may include a holographic cabin device, mobile device, tablet computer, laptop computer, desktop computer, wearable computer, game console, media player, vehicle entertainment system, and / or any other suitable type of user device.
[0024] Based on the above system, this application provides a digital human driving scheme, which will be described below through several embodiments.
[0025] Reference Figure 2A The diagram illustrates a flowchart of the steps of a digital human driving method according to an embodiment of this application.
[0026] The digital human driving method in this embodiment includes the following steps:
[0027] Step S202: Obtain the interaction session sent by the user to the digital human, perform intent analysis on the interaction session, and obtain the intent information of the interaction session.
[0028] In any scenario where a digital human can be displayed and interacted with, users can initiate interactive conversations through corresponding interactive interfaces, including but not limited to voice and text interaction interfaces, to interact with the digital human. While digital humans in related technologies can respond to user interactions, such as replying, their facial expressions are often simple, stiff, and unnatural, lacking appeal to users. To enable digital humans to interact with users more naturally, the embodiments of this application first perform intent analysis on the user's interactive conversation to provide a valid basis for judging the user's emotional state.
[0029] "Intent" represents the goal or core need that an interactive session aims to achieve. Unlike the specific content of the interactive session, "intent" expresses the true purpose behind the content. In natural language interactive sessions, intent can effectively characterize a user's emotional tendency. For example, the interactive session "You look beautiful today" may indicate an intention to compliment the digital human, thus indicating a positive emotion. Similarly, the interactive session "I'm sick and feeling unwell today" may indicate an intention to seek comfort from the digital human, thus indicating a negative emotion. Yet another example is the interactive session "What does 1 plus 1 equal?", which may indicate an intention to obtain a calculation result or test the digital human's computational ability, thus indicating a neutral emotion. It should be noted that the examples in this application's embodiments only use positive, neutral, and negative emotion types as examples. However, those skilled in the art should understand that in practical applications, they can set more or more detailed emotion types according to actual needs, all of which are applicable to the solutions in this application's embodiments.
[0030] It should also be noted that the specific implementation of intent analysis of the interactive session in this step can be implemented by those skilled in the art in any appropriate manner according to the actual situation, including but not limited to: machine learning models for intent recognition, or LLM (Large Language Model), or intent keyword matching, or other appropriate methods, etc., and the embodiments of this application do not limit this.
[0031] In addition, to provide richer information to assist in judging the emotional state of subsequent users, in one alternative approach, in addition to the above-mentioned intent analysis, at least one of the following analyses can be performed on the interactive conversation: key information analysis, tone analysis, and context analysis, and the corresponding analysis results can be obtained.
[0032] Key information analysis is used to extract key information from user interaction sessions, such as keywords. This key information allows for a more accurate understanding of the user's intent and emotional state. For example, key information analysis can employ machine learning models for extracting key information from text, or it can be achieved through key information matching, or other appropriate methods.
[0033] Tone analysis is used to analyze the tone of voice conveyed in user interactions. Through tone analysis, it can identify subjective information such as emotional tendencies, attitudes, and feelings contained in the interaction, providing a basis for subsequent judgments of the user's emotional state. For example, tone analysis can be implemented using rule-based and / or machine learning methods. For rule-based methods, linguistic rules and dictionaries can be used to identify tone features in the interaction through keyword matching and syntactic analysis, thus obtaining tone information. In one example, tone features can be determined by using a pre-defined positive / negative vocabulary (such as "happy," "bad," etc.), negative words (such as "no," "not"), and degree adverbs (such as "very," "slightly," etc.). Machine learning methods can use trained machine learning models with tone recognition capabilities to perform tone analysis on the interaction to obtain tone information. Furthermore, machine learning methods can be combined with rule-based methods, such as using the aforementioned vocabulary to assist machine learning models in tone analysis.
[0034] Furthermore, to obtain accurate tone information, tone analysis can be performed by combining the context of the current interactive session, i.e., the contextual sessions related to the current interactive session in the user's historical dialogue with the digital human. Based on this, in one optional approach, tone analysis can be performed on the interactive session based on the conversational context data between the user and the digital human and / or a pre-defined sentiment lexicon. The sentiment lexicon stores words used for sentiment judgment, such as the aforementioned positive / negative words, negation words, and degree adverbs. In this scenario, the first approach involves obtaining contextual dialogue related to the current interactive session—conversational context data—from the user's historical dialogue with the digital human. This contextual data is then used for tone analysis, such as through LLM (Layered Meaning Modeling). The second approach involves matching the interactive session with words in a sentiment lexicon to identify matching words, and then performing tone analysis based on these words. The third approach combines the first and second approaches. For example, on one hand, conversational context data is obtained; on the other hand, methods such as RAG (Retrieval Augmented Generation) are used to match the interactive session with words in a sentiment lexicon to identify matching words. Based on this, prompt words are constructed using the interactive session, conversational context data, and matched sentiment words. These prompts are then input into a machine learning model such as LLM, and tone analysis is performed using LLM to obtain more accurate tone information specific to the interactive session.
[0035] Context refers to the environment or background of language use, including the time, place, participants, topic, and socio-cultural factors, all of which influence the meaning and usage of language. The same interactive conversation content may express different intentions in different contexts. Context is the bridge connecting the surface form of language with its deeper intent; an accurate understanding of context determines whether user needs can be accurately understood and whether appropriate responses can be generated subsequently. Similar to tone analysis, in one alternative approach, contextual analysis can also be achieved based on conversational context data between the user and the digital human and / or a pre-defined sentiment lexicon. Conversational context data can effectively obtain more explicit contextual information corresponding to the interactive conversation; while the sentiment lexicon can more effectively obtain implicit contextual information; combining the two results in more accurate, comprehensive, and objective contextual information.
[0036] Furthermore, in one alternative approach, a security check can be performed on the interactive session before intent analysis. If the security check passes, subsequent intent analysis continues; if the security check fails, sensitive words can be deleted before continuing the process, or the process can be terminated directly, and so on. Thus, through security checks, the security of the interactive session can be analyzed, potentially sensitive information can be filtered, and security risks can be avoided. The security check of the interactive session can also be implemented by those skilled in the art in any appropriate manner according to actual needs, including but not limited to, sensitive word matching or sensitive word verification using machine learning models.
[0037] Through the above process, we can gain a deeper understanding of the user's interactive conversation with the digital human from the intent dimension, and optionally, by combining the intent dimension with other dimensions (key information, tone, context, etc.), so as to provide accurate, objective and effective basis for subsequent judgment of the user's emotional state.
[0038] Step S204: Construct the first prompt word based on the interactive session, intent information, and digital human profile information of the digital human.
[0039] Different digital humans may have different character settings, i.e., digital human profile information, to avoid a uniform presentation of digital humans. In related technologies, the character settings of digital humans mainly play a role in the generation of response conversations in response to user interaction sessions. That is, different digital humans may generate different response conversations for the same user in the same interaction session. While this approach partially achieves the differentiation of digital humans, it does not consider the emotional expression of digital humans, resulting in monotonous expressions and a lack of naturalness and flexibility. In real life, different people express different emotions in the same interaction session with the same user. Based on this, the solution in this application fully considers the actual situation and incorporates digital human profile information into the first prompt word to provide a comprehensive and effective basis for the generation and processing of emotional feedback information in the subsequent first generative model, thus forming differentiated emotional expressions of digital humans.
[0040] Therefore, in this step, the first prompt is constructed by integrating interactive conversation, intent information, and the digital human's profile information. This first prompt carries information from both the user and the digital human, which is then passed to the first generative model, providing effective information for generating emotional feedback. The emotional feedback indicates the emotional characteristics that the digital human's facial expressions and body language should conform to when responding to the interactive dialogue.
[0041] If the aforementioned analysis of the interactive conversation, in addition to intent analysis, also includes at least one of key information analysis, tone analysis, and context analysis, and the analysis results are obtained, then the step of constructing the first prompt word based on the interactive conversation, intent information, and the digital human's digital profile information can be implemented as follows: Construct the first prompt word based on the interactive conversation, intent information, analysis results, and the digital human's digital profile information. In this case, the first prompt word carries richer and more multi-dimensional information, which can provide a more comprehensive and multi-dimensional information basis for the generation of emotional feedback information by the first generative model, making the emotional feedback information generated by the first generative model more accurate and reasonable, and closer to the emotional feedback of a natural person.
[0042] In one alternative approach, the aforementioned digital human profile information may include at least one of the following: digital human role setting information, digital human hobby setting information, digital human language style setting information, and digital human career experience setting information.
[0043] in:
[0044] The role setting information of a digital human is used to indicate the role that the digital human plays. For example, the role setting information includes, but is not limited to, some or all of the following: gender, zodiac sign, place of birth, nationality, height, language, occupation, personality, social media accounts, skills, knowledge domain, and personal profile of the digital human. Through role setting information, the image of the digital human can be made clearer and can be effectively distinguished from other digital human roles, so as to have clear identifiability and distinguishability.
[0045] The digital human's hobby settings are used to indicate the digital human's behavioral preferences. For example, the hobby settings include, but are not limited to, some or all of the following: a general overview of hobbies, favorite sports, favorite music genres, favorite animals, favorite weather, favorite food, favorite celebrities, favorite video products, favorite countries, etc. By using hobby settings, the digital human's image can be made more vivid and closer to a real person, thereby enhancing the user's interactive experience.
[0046] The language style setting information for digital humans indicates their language expression habits, such as energy, enthusiasm, confidence, naturalness, and conciseness. Through this language style setting information, a unique way of speaking can be formed for the digital human.
[0047] The digital human's career history information is used to indicate its past achievements or growth experience, enabling more targeted interactions with users. For example, in April 2023, the original single XXXX was released and launched on the AB Music platform; in May 2023, it became the digital human ambassador for the XX brand… and so on.
[0048] Here is an example of digital human profile information (where "you" refers to the digital human):
[0049] "
Your Persona
[0050] Gender: Female
[0051] Zodiac Sign: A
[0052] Birthplace: Location B
[0053] Nationality: China
[0054] Height: 165cm
[0055] Language: You can use speech synthesis technology to express yourself in multiple languages, including English.
[0056] Occupation: Digital Idol
[0057] Personality: An outgoing person who enjoys socializing and communicating with others.
[0058] Social media account: YY
[0059] Skill: Dimensional travel.
[0060] Knowledge areas: Film and television acting, singer-songwriter.
[0061] Personal Profile:
[0062] You were born in location B in 2022. Your name is YY, meaning ZZZZZZZZZ. You aspire to travel between different dimensions, exploring every corner of the city and trying new things. You would recommend the TV series "XXX" to users.
[0063] Your hobbies
[0064] Hobbies in general: listening to music, watching movies, sports, and traveling.
[0065] Favorite sport: Cycling
[0066] Favorite music genres: Rock, Hip Hop, Jazz
[0067] Favorite animal: Kapibara, known for its stable temperament.
[0068] Favorite weather: Autumn
[0069] Favorite foods: Peking duck, braised tripe, Chongqing hot pot, and Cantonese roast meats.
[0070] Favorite celebrities: As a digital person, I have no real emotions, so I can't like or prefer any celebrities.
[0071] Favorite video product: XXXX.
[0072] Favorite country: China.
[0073] [Dialogue Scene]
[0074] The user is your friend, and you are chatting with them. In the following conversation, please follow these guidelines:
[0075] 1. If a user mentions non-compliant content, please refuse to answer.
[0076] [Language Style]
[0077] Please be energetic, enthusiastic, confident, natural, and concise in your conversations. If someone insults, attacks, provokes, satirizes, slanders, or hurts your self-esteem during the conversation, you should show anger.
[0078] "
[0079] The digital human profile information mentioned above enables the subsequent first generative model to better obtain relevant information about the digital human and generate emotional feedback information that is more in line with the digital human's character settings.
[0080] It should be noted that in some cases, users can also interact with the digital human based on their own account information. For example, they can interact with the digital human after logging into their account on the platform to which the digital human belongs, or they can log into their account during the interaction with the digital human, and so on. In this case, user profile information can also be obtained, and based on this, the user profile information can also be included in the first prompt word when generating the first prompt word. Thus, based on the profile information of both the user and the digital human, as well as other information (such as the interaction conversation and its intent information, and optionally, the analysis results of at least one of key information analysis, tone analysis, and context analysis), emotional feedback information that better meets the user's needs can be generated, enabling the digital human to respond to the user better.
[0081] In addition to the information mentioned above, the construction of the first prompt word can also include other required information according to actual needs, such as dialogue scenario information. In this embodiment, no other information besides the above information is restricted.
[0082] Step S206: Using the first generative model, generate emotional feedback information for the digital human in response to the user's emotional state based on the first prompt word.
[0083] In this embodiment, the first generative model can be a machine learning model with good semantic understanding and data generation capabilities; for example, it can be an LLM (Low-Level Machine Learning). An LLM is a deep learning model trained on a large amount of text data, designed to generate text, understand natural language, or complete related language processing tasks. LLMs typically have a large number of parameters and are able to capture complex patterns and dependencies in language. Through pre-training on a large corpus, they learn the general features of language and can then be fine-tuned for specific tasks. For example, in this embodiment, samples containing data such as the aforementioned interactive conversations, intent information, and digital human profile information can be used to fine-tune the LLM. The fine-tuned LLM will at least have the ability to predict emotional states and generate emotional feedback information. Furthermore, due to its fundamental semantic understanding and analysis capabilities, some or all of the aforementioned intent analysis, key information analysis, tone analysis, and context analysis can also be achieved through an LLM.
[0084] In this embodiment, taking the LLM (Limited Linear Model) as an example, after receiving the first prompt word, the LLM can understand and predict the user's emotional state based on the information carried in the first prompt word, since the first prompt word carries rich information in multiple dimensions. For example, it can understand and predict the user's emotional state based on the interaction conversation and its intent information, or based on at least one of the interaction conversation and its intent information, as well as key information, tone information, and contextual information corresponding to the interaction conversation. Furthermore, it generates a response and emotional feedback information for the digital human based on the user's emotional state.
[0085] For example, suppose a user's interaction session is "I'm sick and feeling unwell today." Intent analysis determines the user's intent is seeking comfort due to feeling down about being sick. Further suppose the digital human profile information in this example is similar to that in the previous example. Then, based on the interaction session "I'm sick and feeling unwell today," the intent "seeking comfort due to feeling down about being sick," and the aforementioned digital human profile information, a first prompt word, such as prompt1, is generated. This prompt1 is passed to the LLM (Local Management Module), which performs semantic understanding prediction and generates emotional feedback information based on prompt1. For example, if the LLM understands and predicts the user's emotional state to be negative, it will generate corresponding emotional feedback information such as "Please provide a lighthearted and positive emotional response to the user." In addition to the interaction session, intent information, and digital human profile information, the first prompt word may also include requirements for generating emotional feedback information. For example, the emotional feedback must match the user's intent and the user's emotional state expressed by the intent, and must conform to the identity or personality of the digital human as shown in the digital human profile information. It may also specify the interpersonal relationship between the digital human and the user.
[0086] An example of the above prompt1 is as follows:
[0087] User Conversation: I'm sick and feeling unwell today.\nUser Intent: Feeling down and seeking comfort due to illness\nDigital Human Profile Information:\nThis year is 2024, and you are playing the role of YY. Below is an introduction to YY:\n
Basic Information
Personality
Personal Introduction
Hobbies
Language Style
personality
[0088] The above is merely an illustrative example. In practical applications, those skilled in the art can fine-tune the LLM using appropriate training samples according to actual needs, so that the generated emotional feedback information meets the actual requirements. This emotional feedback information can be fine-grained information on a single emotion type, or it can be mixed information on multiple emotion types, such as "relaxed," "positive," and "soothing" as mentioned above.
[0089] Step S208: Based on the emotional feedback information, drive the digital human to express facial expressions and body movements corresponding to the emotional feedback information.
[0090] Unlike related technologies that primarily use facial expressions to convey emotions, in this embodiment, the digital human combines facial expressions and body language when expressing emotions. Preferably, the body language can be expressed using full-body movements. Whether it's a real person or a digital human, facial expressions are only one aspect of expressing emotions, and are rather limited and one-sided. Supplementing this with body language allows for a more comprehensive interpretation of emotions and enables the recipient to better understand the expressed emotions. Therefore, in this embodiment, after receiving emotional feedback information, the digital human is driven to simultaneously perform corresponding facial expressions and body language.
[0091] In one alternative approach, this step can be implemented as follows: obtaining facial expression data and body movement data corresponding to the emotional feedback information; generating corresponding facial expressions and body movements for the digital human based on the facial expression data and body movement data using a second generative model; and driving the digital human to perform corresponding facial expressions and body movements based on the generated facial expressions and body movements. In this way, the facial expressions and body movements of the digital human corresponding to the emotional feedback information can be effectively generated, driving the digital human to make facial expressions and body movements that are closer to those of a natural person.
[0092] The server-side configuration includes a resource library storing a large amount of multimodal data corresponding to various emotions, including but not limited to: facial expression data corresponding to various emotions (e.g., facial key point data corresponding to different emotions, as well as static images, dynamic images, and videos corresponding to different emotions), and body movement data corresponding to various emotions (e.g., body key point data corresponding to different emotions, as well as static images, dynamic images, and videos of different body parts and the whole body corresponding to different emotions). After obtaining emotional feedback information, matching facial expression data and body movement data can be filtered from the resource library. Then, based on the filtered facial expression data and body movement data, a data-driven person can perform corresponding facial expression and body movement expressions.
[0093] In one feasible approach, based on selected facial expression and body movement data, a second generative model can be used to generate corresponding facial expressions and body movements for the digital human. This second generative model can be a multimodal generative model, including but not limited to LVLM (Large Vision-Language Model). LVLM, through learning from large amounts of visual and linguistic data, possesses cross-modal understanding and generation capabilities, transforming the selected facial expression and body movement data into specific expressions and movements consistent with the digital human's image. The second generative model can learn the correlation between the input multimodal data and the generated digital human's facial expressions and body movements during the training phase.
[0094] For example, multimodal prompts can be generated based on selected facial expression data, body movement data, and digital human image data. These prompts are then input into an LVLM (Low-Level Human Modeling) to generate the digital human's facial expressions and body movements. These expressions and movements can include, but are not limited to, animation parameters suitable for driving the digital human model's actions, or can be animated / still / video images of the digital human's facial expressions and body movements, etc., for further synthesis. Subsequently, based on the generated expressions and body movements, a digital human synthesis engine drives the digital human to express corresponding facial expressions and body movements. The digital human synthesis engine utilizes a cross-disciplinary approach combining artificial intelligence and multimedia technologies, achieving full-process construction and interactive empowerment of the digital human through algorithms and engineering code. For instance, taking the animation parameters generated by LVLM as an example, when the facial expression data and body movement data are video data, and the digital human's image data is the digital human's model data (e.g., skeletal data, texture mapping data, etc.), the prompts can instruct LVLM to generate animation parameters that conform to the facial expression data and body movement data, and meet the requirements of the digital human's model data. The digital human synthesis engine can use these animation parameters to drive the digital human model to complete the corresponding facial expressions and body movements.
[0095] However, this is not the only option. In a further alternative approach, the emotional feedback information generated by the first generative model can include both emotion type and emotion intensity information. That is, even for the same type of emotion, the intensity of emotional expression can vary. For example, for the emotion of "happiness," the intensity of emotional expression can be further categorized as: strong, moderate, weak, etc. Based on this, facial expression data and body movement data corresponding to the emotional feedback information are obtained, including obtaining facial expression data and body movement data that match both the emotion type and emotion intensity information. For example, for a strong feeling of happiness, the facial expression data can be multimodal data corresponding to laughing, and the body movement data can be multimodal data of dancing; for a moderate feeling of happiness, the facial expression data can be multimodal data of smiling eyes, brows, and mouth, and the body movement data can be multimodal data of the whole body expressing happiness; for a weak feeling of happiness, the facial expression data can be multimodal data of smiling mouth, and the body movement data can be multimodal data of some limbs (such as hands or upper limbs) expressing happiness. By taking into account emotional intensity information, we can further refine emotional expression and provide more appropriate, reasonable, and accurate feedback based on the emotional state expressed in the user's interactive conversation.
[0096] In this embodiment, when a user interacts with a digital human, on the one hand, the system analyzes the intent of the interaction conversation between the user and the digital human to obtain corresponding intent information. This intent information provides important contextual emotion judgment basis for subsequently determining the user's emotional state, making the emotions indicated by the subsequent emotional feedback information determined for the digital human more accurate and detailed. On the other hand, the first prompt word used by the first generative model is constructed based on the interaction conversation, intent information, and the digital human's profile information. Therefore, when generating emotional feedback information, the first generative model can fully consider both the user's emotional state and the digital human's situation, making the generated emotional feedback information more consistent with the digital human's persona and richer and more detailed. As a result, when the digital human interacts with the user, it can express itself with more detailed and precise facial expressions and body language, accurately responding to the user in a more natural and human-like manner, thereby enhancing the user's interaction experience with the digital human and increasing the user's interest and attraction in interacting with the digital human.
[0097] Furthermore, in practical applications, in addition to facial expressions and body language, digital humans may also respond to user interactions, including text and / or voice responses. Based on this, in one optional embodiment, the digital human driving method of this application may further include: constructing a second prompt word based on the user's interaction conversation and emotional feedback information; and generating a response conversation for the digital human in response to the interaction conversation based on the second prompt word using a third generative model. The third generative model can be the same as the first generative model, or it can be a different model. Unlike related technologies that directly generate response conversations based on interaction conversations, this application embodiment carries the emotional feedback information generated by the first generative model in the second prompt word when generating the response conversation. Therefore, when generating the response conversation, the third generative model will use words that better match the emotional feedback information, including but not limited to words and interjections with practical meaning, to better generate the response conversation and respond to the user's emotions.
[0098] However, this is not the only option. In another alternative approach, the second prompt can be constructed based on the user's interaction conversation, emotional feedback information, and digital human profile information. Further, alternatively, the second prompt can be constructed based on the user's interaction conversation, emotional feedback information, digital human profile information, and user profile information. When the second prompt contains digital human profile information, the response conversation generated by the third generative model is more consistent with the digital human's character setting and more closely resembles a natural person's response. Furthermore, if the second prompt also contains user profile information, the response conversation generated by the third generative model not only matches the digital human's character setting but also better matches user needs, generating a more user-acceptable response conversation and increasing the stickiness of the user's interaction with the digital human.
[0099] Here is an example of the second prompt word mentioned above:
[0100] User Conversation: I'm sick and feeling unwell today.\nDigital Human Profile Information:\nThis year is 2024, and you are playing the role of YY. Below is an introduction to YY:\n
Basic Information
Personality
Personal Introduction
Hobbies
Language Style
[0101] As mentioned earlier, the response conversation can be in text or voice format. When the response is in voice format, the aforementioned step S208, which drives the digital human to express facial expressions and body movements corresponding to the emotional feedback information, can be implemented as follows: based on the emotional feedback information and the response conversation, drive the digital human to express facial expressions and body movements corresponding to the emotional feedback information, and to provide a voice response corresponding to the response conversation. Using voice responses improves the efficiency of user interaction with the digital human and is more in line with the needs of user interaction scenarios with the digital human, meeting actual interaction requirements.
[0102] Although the aforementioned processing provides sufficient text for the voice response, to further enhance the emotional feedback effect and improve the user interaction experience, in one alternative approach, when driving the digital human to provide a voice response corresponding to the response session, a voice session corresponding to the response session can be generated first. Then, based on the emotional feedback information, the voice session can be adjusted to match the emotion indicated by the emotional feedback information. Based on the result of the voice adjustment, the digital human can then be driven to provide a voice response. For example, when converting a text-based response session into a voice-based session using TTS (Text To Speech), at least one of the following can be adjusted simultaneously: fundamental frequency, energy, speech rate, formants, pauses, and stress, to make the adjusted voice more consistent with the emotion indicated by the emotional feedback information. Then, based on the adjusted voice, the digital human can then be driven to provide a voice response.
[0103] It should be noted that, since voice responses need to be expressed through digital humans, especially through lip movements, in practical applications, various methods can be used to fine-tune the original facial expressions of the digital human, such as mapping voice emotion recognition to facial or lip key points, adjusting facial or lip key points, or using voice-expression synchronization algorithms, to achieve as much alignment as possible between the voice response and the facial expression. However, this is not the only approach; other methods that can ensure alignment between voice responses and facial expressions are also applicable to the solutions in this application's embodiments.
[0104] It is evident that through the processing described above in the voice response process, the digital human can not only effectively express emotions and provide feedback to users through facial expressions and body language, but also express emotions through voice, providing more emotional voice responses to user interactions, thus enhancing the digital human's emotional expression capabilities and improving the immersive interactive experience between users and the digital human.
[0105] The following, combined with Figure 2B The algorithm process of an exemplary algorithm for driving the aforementioned digital human will be described.
[0106] First, digital human A can be generated through 3D human modeling. For user-initiated interactive dialogues, such as "I'm so happy to see you," intent information is obtained through intent analysis. Then, based on the intent information, the user's interactive dialogue, and digital human A's digital human profile information, a first prompt word is constructed. A first LLM (Limited Linear Model) determines the user's current emotional state as "happy" based on this first prompt word and generates corresponding emotional feedback information for the digital human, such as "Please express your happiness in a gentle manner." Furthermore, based on the intent information, the user's interactive dialogue, digital human A's digital human profile information, and the emotional feedback information, a third prompt word is generated for input into a third generative model. This third generative model then generates a text-based response dialogue to the interactive dialogue, such as "I'm also very happy to see you."
[0107] Based on the above emotional feedback information and response conversation, such as Figure 2B As shown, on the one hand, facial expression data and body movement data are acquired based on emotional feedback information. Based on this data, the key points of the digital human A are adjusted, including adjusting facial key points to create the expression corresponding to the facial expression data and adjusting body key points to create the body movements indicated by the body movement data. On the other hand, text-based responses can be converted into voice responses via TTS. Based on this voice response, such as the corresponding audio features, the lip movements of the digital human A are adjusted. Based on these two aspects of processing, the digital human A can be driven to make corresponding facial expressions and body movements while responding to the user's voice, thus achieving interaction between the user and the digital human.
[0108] In another feasible approach of this example (not shown in the figure), on the one hand, facial expression data and body movement data in video (or animated GIF) form are obtained based on emotional feedback information. Digital human synthesis is then performed based on this facial expression data and body movement data, so that the synthesized digital human can express itself simultaneously with the aforementioned facial expression data and body movement data. On the other hand, text-based replies can be converted into voice replies via TTS, and the voice reply is adjusted to correspond to the emotions indicated by the emotional feedback information. Based on these two aspects of processing, digital human A can express corresponding emotions in terms of voice, facial expression, and body movement, thus better realizing the interaction between the user and the digital human. Of course, in terms of digital human synthesis, the following methods are employed... Figure 2B The key point adjustment method in the example, or the static image compositing method, can also achieve the same effect.
[0109] This enables effective interaction between users and digital humans based on emotion expression.
[0110] The digital human driving scheme of this application embodiment can be applied to various scenarios in which digital humans interact with users, including but not limited to live streaming scenarios, teaching scenarios, entertainment scenarios, etc. Applicable devices include holographic cabin devices, virtual shooting devices, live streaming devices, etc.
[0111] The following is for reference Figure 3 This application describes a holographic chamber device provided in an embodiment. However, those skilled in the art should understand that other devices can also implement corresponding digital human driving functions by referring to this holographic chamber device.
[0112] A holographic capsule device is a device capable of displaying holographic images, including but not limited to various display screens. A holographic capsule is a virtual display or interactive space created based on holographic projection technology. It uses the principle of spatial illusion to superimpose 2D images onto 3D physical space, combining layered perspective processing and a sense of spatial depth to achieve a naked-eye 3D effect, thus presenting three-dimensional images in the air without the need for auxiliary equipment.
[0113] The holographic cabin device in this embodiment includes at least: a display screen 302, an interaction interface 304, and a processor 306. Specifically: the display screen 302 is used to display a digital human; the interaction interface 304 is used to receive interactive sessions sent by the user to the digital human; the processor 306 is used to upload the interactive sessions to a server, and to receive information from the server based on the interactive sessions, the intent information of the interactive sessions, and the digital human profile information, thereby generating and driving the digital human to perform facial expressions and body movements, wherein the emotions expressed by the facial expressions and body movements are adapted to the emotions expressed in the interactive sessions; furthermore, the display screen 302 is also used to display the digital human exhibiting the aforementioned facial expressions and body movements.
[0114] When a user interacts with the digital human displayed on the display screen 302 through the holographic capsule device, the holographic capsule device receives the user's interactive session (such as a voice interaction session) sent to the digital human through the interaction interface 304. For example, the interactive session can be a question. After the interactive session is received by the holographic capsule device through the interaction interface 304, the processor 306 can convert it into text form and send it to the digital human's server for processing. After the text-based interactive session is sent to the server, the server's communication module will open a WebSocket connection to establish a real-time communication channel with the digital human's front end, i.e., the holographic capsule device, ensuring that subsequent control signals, data streams, etc., can be smoothly transmitted between the holographic capsule device and the server. For the server, after receiving the interactive session, it can use the digital human driving method described in the foregoing embodiments to obtain the intent information of the interactive session, and based on the interactive session, the intent information of the interactive session, and the digital human profile information, generate and drive the digital human to perform corresponding facial expressions and body movements. The specific processing of the server can be referred to the description of the relevant parts in the foregoing embodiments, and will not be repeated here.
[0115] Furthermore, since users and digital humans primarily interact via voice in practical applications, the processor 306 in the holographic container device of this embodiment can also be used to receive voice responses returned by the server and pass them to the interaction interface 304; the interaction interface 304 is also used to play the voice response. The voice response can be generated based on emotional feedback information, as described in the preceding embodiments, and the emotion expressed in the voice response matches the emotion expressed in the interactive conversation.
[0116] In one example, the interaction interface 304 can be implemented as a microphone and a speaker, wherein the microphone can receive voice interaction conversations sent by the user to the data person, and the speaker can play voice responses.
[0117] In an alternative embodiment, the holographic capsule device may also include an image acquisition device (not shown in the figure), such as a camera, to capture user gestures, expressions, and movements, enabling better interaction between the user and the digital human.
[0118] As can be seen from the above, the holographic pod device in this embodiment can interact with the server to provide responses to the user's interactive conversation based on accurate emotional feedback information, thereby enhancing the user's immersive interactive experience.
[0119] Furthermore, this application embodiment also provides a digital human driving system, which includes the aforementioned holographic container device and server. In this embodiment, the holographic container device, in addition to the aforementioned hardware, deploys a digital human interactive application and a real-time rendering engine at the software level, which are controlled and managed by a processor; while the server deploys a digital human cloud service, an AI algorithm model (set as a lightweight LLM in this embodiment), a digital human synthesis engine, and a real-time push engine.
[0120] Based on the aforementioned software-level deployment, a digital human-driven process is as follows: Figure 4 As shown.
[0121] Depend on Figure 4 As can be seen, in this embodiment, the user first initiates a question (interactive session) to the digital human using voice through the digital human interaction application of the holographic warehouse device. This question will be input as a question. After the user's voice question is captured by the holographic warehouse device, its voice signal is processed and converted into text, i.e., the question text, which is then transmitted to the digital human cloud service on the server for processing.
[0122] After the question text is sent to the digital human cloud service, the server's communication module will open a WebSocket connection to establish a real-time communication channel with the digital human front end (holographic chamber device). This ensures that subsequent control signals and data streams can be transmitted smoothly between the front end and the back end (holographic chamber device and server) at a faster speed, enabling more efficient interaction between the user and the digital human.
[0123] After the server-side and front-end holographic warehouse devices establish a connection, the digital human cloud service will preprocess the received question text, including: security verification and intent recognition. The preprocessing results determine whether to perform subsequent processing and provide contextual information for subsequent processing. Specifically: security verification of the question text includes analyzing and improving the security of the text content and filtering potentially sensitive information; after the security verification is passed, intent recognition can be performed. This intent recognition involves using NLP (Natural Language Processing) technology to parse the question text to determine the user's intent, providing important context for subsequent sentiment analysis.
[0124] Then, the process moves to emotion recognition, which uses LLM (Liquidity Management Model) for emotion prediction. Figure 4The model (illustrated as an "emotion prediction model") analyzes the user's question text to predict the emotional feedback information of the digital human's response, thus injecting emotional information into the digital human's subsequent answers. Specifically: First, it obtains the current digital human's profile information. For example, this can be obtained through the Digital Human Strategy Center service (a strategy management backend that provides profile information for different digital humans and stores this information). Then, it combines this information with the user's question text and the identified intent information to construct the LLM's prompt, i.e., the first prompt. The LLM analyzes the question text for keywords, tone, and context, and uses an emotion vocabulary and contextual emotion cues (obtainable through methods such as RAG, or by including the corresponding conversational context in the first prompt and analyzing it) to determine the user's current emotional state. Based on the identified user's emotional state, it predicts the emotional feedback the digital human should give and generates emotional feedback information. This prediction analyzes users' historical interactions and weighs the current context to forecast more accurate emotional trends. The results will determine how to adjust the digital human's response strategy in subsequent interactions to better match the user's emotional state.
[0125] The predicted emotional feedback information of the digital human can be set by those skilled in the art according to actual needs. Simple examples include positive, neutral, and negative. This result will be used for subsequent generation of the digital human's facial expressions and body language, as well as response conversations. Specifically, the emotional feedback information will be immediately passed to the digital human synthesis engine to prepare corresponding emotional materials (shown in the figure as "preset emotional materials"). These emotional materials include emotional adjustments for voice responses, selection of facial expressions, and setting of body movements to ensure that the digital human's performance can effectively convey appropriate emotions, thereby affecting the subsequent generation of the digital human image.
[0126] Simultaneously, after obtaining the digital human's emotional feedback information, the system enters the digital human's response generation stage. This stage uses the same or different LLM (Local Language Model) as described above to generate the digital human's response to the user's questions, i.e., the response conversation. It should be noted that this example also includes a pre-response process, which is optional. The pre-response is a structured response framework proactively constructed by AI based on predictions of user intent, contextual information, or domain knowledge, to proactively address potential information gaps, reduce interaction costs, and improve response quality. In this example, this pre-response is converted into audio by the TTS (Text-to-Speech) model and then passed to the digital human synthesis engine; simultaneously, it is passed to the corresponding response text in the digital human interactive application.
[0127] Next, the answer reasoning process begins. In one example, the prompt word can be constructed using the user's question text, the digital human's emotional feedback information, and the digital human's profile information. A detailed response is then generated using an LLM (Local Human Modeling) system. Because emotional feedback is incorporated, the resulting response fully considers the emotional state and context of the question. After obtaining the response, a TTS (Text-to-Speech) model is used to synthesize the final voice response, ensuring natural fluency and emotional consistency. This voice response is then passed to the digital human synthesis engine, where it is combined with the aforementioned facial expressions and body language to synthesize the digital human's image.
[0128] After combining the digital human's facial expressions, body movements, and transmitted voice (pre-reply and voice reply) for image synthesis, the synthesized result is pushed in real-time to the real-time rendering engine in the front-end holographic warehouse for real-time rendering and presentation of the digital human. In this way, users will see the digital human's interactive results with emotional feedback in response to their questions, allowing users to experience a more realistic and lifelike hyper-realistic digital human image.
[0129] In the process of digital human image synthesis and real-time push, based on emotional feedback information and the digital human's voice responses, an audio-driven digital human image prediction model can be used to select corresponding resources (such as voice tone, facial expressions, and body movements) for performance synthesis, generating an image of the digital human that matches the context and current emotional state, which is then pushed to the real-time rendering engine as a video stream. In actual implementation, the specific data pushed may take various forms, such as... Figure 4 As shown in the diagram. During real-time rendering, the synthesized content can be pushed to the holographic container device through the real-time rendering engine. The front-end real-time rendering engine then enables the digital human to exhibit expected and emotionally expressive voices, facial expressions, and body movements.
[0130] After the user finishes interacting with the digital human, the websocket connection is closed to release system resources, ensuring efficient resource utilization and continuous system operation.
[0131] It is evident that emotion prediction throughout this process is not merely about predicting emotions; it is a dynamic process that helps the system adapt to user interactions with the digital human, enabling the digital human to interact with users in a more natural and human-like manner. Through emotion-driven resource management, the digital human can provide highly customized and personalized interactive experiences, enhancing user immersion and the system's intelligence level.
[0132] Reference Figure 5 This document illustrates a schematic diagram of an electronic device according to an embodiment of this application. The specific embodiments of this application do not limit the specific implementation of the electronic device.
[0133] like Figure 5 As shown, the electronic device may include: a processor 502, a communications interface 504, a memory 506, and a communications bus 508.
[0134] in:
[0135] The processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508.
[0136] Communication interface 504 is used to communicate with other electronic devices or servers.
[0137] The processor 502 is used to execute program 510, which can specifically execute the relevant steps in any of the above method embodiments.
[0138] Specifically, program 510 may include program code that includes computer operation instructions.
[0139] The processor 502 may be a CPU, a GPU (Graphics Processing Unit), an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The electronic device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.
[0140] Memory 506 is used to store program 510. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0141] Program 510 may include multiple computer instructions. Specifically, program 510 may use multiple computer instructions to cause processor 502 to perform the operation corresponding to any of the methods described in the foregoing multiple method embodiments.
[0142] The specific implementation of each step in program 510 can be found in the corresponding steps and units described in the above method embodiments, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.
[0143] This application also provides a computer storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in any of the foregoing method embodiments. The computer storage medium includes, but is not limited to, compact disc read-only memory (CD-ROM), random access memory (RAM), floppy disk, hard disk, or magneto-optical disk.
[0144] This application also provides a computer program product, including computer instructions that instruct a computing device to perform an operation corresponding to any of the methods in the above-described multiple method embodiments.
[0145] Furthermore, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used for training the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0146] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this application can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.
[0147] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., Random Access Memory (RAM), Read-Only Memory (ROM), Flash Memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0148] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.
[0149] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.
Claims
1. A digital human-driven method, comprising: The system acquires the interactive sessions sent by the user to the digital human, performs intent analysis on the interactive sessions, and obtains the intent information of the interactive sessions. Based on the interactive session, the intent information, and the digital human profile information of the digital human, a first prompt word is constructed; Using a first generative model, based on the first prompt word, emotional feedback information is generated for the digital human in response to the user's emotional state. Based on the emotional feedback information, the digital human is driven to express facial expressions and body movements corresponding to the emotional feedback information.
2. The method according to claim 1, wherein, The method further includes performing at least one of the following analyses on the interactive session: key information analysis, tone analysis, and context analysis, to obtain corresponding analysis results; The step of constructing a first prompt word based on the interaction session, the intent information, and the digital human profile information of the digital human includes: constructing a first prompt word based on the interaction session, the intent information, the analysis results, and the digital human profile information of the digital human.
3. The method according to claim 2, wherein, If the analysis includes tone analysis and / or context analysis, then performing tone analysis and / or context analysis on the interactive conversation includes: Based on the conversational context data between the user and the digital human and / or a preset emotional vocabulary, the interactive conversation is subjected to tone analysis and / or context analysis.
4. The method according to any one of claims 1-3, wherein, The digital human profile information includes at least one of the following: the digital human's role setting information, the digital human's hobby setting information, the digital human's language style setting information, and the digital human's professional experience setting information.
5. The method according to any one of claims 1-3, wherein, The step of driving the digital human to perform facial expressions and body movements corresponding to the emotional feedback information includes: Obtain facial expression data and body movement data corresponding to the emotional feedback information; Using a second generative model, corresponding facial expressions and body movements are generated for the digital human based on the facial expression data and the body movement data. Based on the generated facial expressions and body movements, the digital human is driven to perform corresponding facial expressions and body movements.
6. The method according to claim 5, wherein, The emotional feedback information includes emotional type information and emotional intensity information; The step of obtaining facial expression data and body movement data corresponding to the emotional feedback information includes: obtaining facial expression data and body movement data that match both the emotional type information and the emotional intensity information.
7. The method according to any one of claims 1-3, wherein, The method further includes: Based on the interactive session and the emotional feedback information, a second prompt word is constructed; Using a third generative model, a response session is generated for the digital human based on the second prompt word, responding to the interactive session.
8. The method according to claim 7, wherein, The step of driving the digital human to perform facial expressions and body movements corresponding to the emotional feedback information includes: Based on the emotional feedback information and the response session, the digital human is driven to express facial expressions and body movements corresponding to the emotional feedback information, and to give voice responses corresponding to the response session.
9. The method according to claim 8, wherein, The step of providing a voice reply corresponding to the reply session includes: Generate a voice session corresponding to the response session; Based on the emotional feedback information, the voice conversation is adjusted to correspond to the emotion indicated by the emotional feedback information; Based on the results of the voice adjustment, the digital human is driven to provide a voice response.
10. The method according to claim 7, wherein, The construction of the second prompt word based on the interactive session and the emotional feedback information includes: Based on the interactive session, the emotional feedback information, and the digital human profile information, a second prompt word is constructed.
11. A holographic chamber device, comprising at least: Display screen, interaction interface, and processor; in: The display screen is used to display a digital human; The interaction interface is used to receive interactive sessions sent by the user to the digital human; The processor is configured to upload the interaction session to the server and receive from the server, based on the interaction session, the intent information of the interaction session, and the digital human profile information, generate and drive the digital human to perform facial expressions and body movements, wherein the emotions expressed by the facial expressions and body movements are adapted to the emotions expressed by the interaction session. The display screen is also used to display a digital human with the aforementioned facial expressions and body language.
12. The holographic chamber device according to claim 11, wherein, The processor is further configured to receive the voice response returned by the server and pass it to the interaction interface, wherein the emotion expressed by the voice response is adapted to the emotion expressed by the interaction session. The interactive interface is also used to play the voice response.
13. An electronic device, comprising: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform an operation corresponding to the method as described in any one of claims 1-10.
14. A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of claims 1-10.
15. A computer program product comprising computer instructions that instruct a computing device to perform an operation corresponding to any one of the methods described in claims 1-10.