A method, system, device and storage medium for virtual digital human interactive Q&A

By collecting user facial images and dialogue content, combining natural language processing and emotion analysis technology, the virtual digital human interaction question and answer system can more accurately identify the user's emotional state and adjust the performance of virtual digital humans, solving the problem of insufficient emotional recognition and feedback in the existing system, and achieving a more natural and personalized interactive experience.

CN119476313BActive Publication Date: 2025-05-27BEIJING ZHUOYUE WEILAI INT MEDICINE TECH DEV CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510038320.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-27
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

The existing virtual reality interactive question and answer system has shortcomings in user emotional recognition and feedback, and cannot provide a personalized and humanized interactive experience.

Method used

By collecting user's facial images and dialogue content, combining natural language processing and emotion analysis technology, the system can more accurately identify the user's emotional state, and adjust the facial expressions, body movements and voice tone of virtual digital people according to the emotional state to generate answers that are more in line with user's emotions.

Benefits of technology

It realizes more accurate user emotional recognition and more natural virtual digital expression, which improves the naturalness and fluency of interaction and enhances the authenticity and immersion of user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119476313B_ABST
    Figure CN119476313B_ABST
Patent Text Reader

Abstract

A method, system, device and storage medium for virtual digital human interactive Q&A, which relates to the field of virtual reality. In this method, the dialogue content of the user is subjected to a first process, and the facial image of the user is subjected to a second process to obtain a first emotional state; the text after the first process is subjected to a third process to obtain the second emotional state of the user, and the target emotional state is determined by combining the first emotional state and the second emotional state; the text after the first process is subjected to a fourth process based on a deep learning model to generate a first answer, and the expression mode of the first answer is adjusted according to the target emotional state and the dialogue context to generate a second answer; a virtual digital human is constructed by using virtual reality technology, and the facial expression, body movement and intonation of the virtual digital human are determined according to the emotional tendency in the second answer to display the second answer. Implementing the technical solution provided by this application can more accurately identify the emotional state of the user, and then generate an answer that better conforms to the user's emotion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of virtual reality, and particularly relates to a method, system, device, and storage medium for virtual digital human interactive question and answer. Background Art

[0002] With the development of technology, the human-computer interaction technology has advanced by leaps and bounds. Traditional question-and-answer systems mainly communicate with users through text or voice, and this monotonous communication method can no longer meet the users' needs for rich interactive experiences. In recent years, virtual reality technology has emerged, and interacting through the construction of virtual digital humans has become a new research direction.

[0003] Currently, there are already some virtual reality-based interactive question-and-answer systems on the market, but these systems have made certain innovations in visual presentation, and there are still deficiencies in user emotion recognition and feedback, and they cannot provide personalized and user-friendly interactive experiences.

[0004] Therefore, how to incorporate more accurate user emotion recognition into virtual reality interactive question-and-answer systems, and how to make virtual digital humans express answers more naturally, have become problems that need further research and improvement currently. Summary of the Invention

[0005] The present application provides a method, system, device, and storage medium for virtual digital human interactive question and answer, which can more accurately identify the emotional state of users, and then generate answers that are more in line with the emotions of users, and can make virtual digital humans express answers more naturally.

[0006] In a first aspect of the present application, a method for virtual digital human interactive question and answer is provided, which is applied to an interactive question-and-answer platform. The method includes:

[0007] Collect the facial images and conversation content of users, perform a first processing on the conversation content through natural language processing technology, and perform a second processing on the facial images using a preset emotion analysis model to obtain a first emotional state. The first processing includes word segmentation and semantic role labeling, and the second processing includes facial expression feature extraction and emotion recognition;

[0008] Use a preset emotion analysis model to perform a third processing on the text after the first processing to obtain the second emotional state of the user, and determine the target emotional state by combining the first emotional state and the second emotional state;

[0009] Perform a fourth processing on the text after the first processing based on a deep learning model to generate a first answer, and adjust the expression of the first answer according to the target emotional state and conversation context to generate a second answer;

[0010] Construct a virtual digital human using virtual reality technology, and determine the facial expressions, body movements, and speech intonations of the virtual digital human according to the emotional tendency in the second answer to display the second answer.

[0011] Among them, the second processing of the facial image using a preset emotion analysis model to obtain the first emotional state includes:

[0012] Extract feature points from the facial image, match the feature points with preset expression categories, convert the matched expression categories into corresponding first emotional states, and generate corresponding first emotional values according to the first emotional states. The first emotional states include positive emotions, negative emotions, and neutral emotions.

[0013] The third processing of the text after the first processing using a preset emotion analysis model to obtain the user's second emotional state includes:

[0014] Extract preset features related to emotion analysis from the text after the first processing. The preset features include emotion words, negation words, and degree words.

[0015] Input the preset features into a preset emotion analysis model to obtain an output result.

[0016] Map the emotional tendency of the text to a preset emotion category according to the output result, determine the second emotional state according to the mapped emotion category, and generate a corresponding second emotional value according to the second emotional state. The second emotional states include positive emotions, negative emotions, and neutral emotions.

[0017] By adopting the above technical solutions, collecting the user's facial images and conversation content, and combining natural language processing and sentiment analysis technologies, the system can more accurately understand the user's intentions and emotional states. This multi-dimensional information input makes the interaction process more similar to natural communication between humans, enhancing the realism and immersion of the user experience. By combining the first emotional state (based on facial images) and the second emotional state (based on text content) to determine the target emotional state, the system can more comprehensively grasp the user's emotional changes and adjust the expression of the answer accordingly. This personalized feedback mechanism helps to establish a deeper emotional connection and enhance the user's trust and satisfaction with the system. The answers generated based on the deep learning model already have high accuracy and pertinence, but further adjusting the expression according to the target emotional state and conversation context makes the answers more suitable for the current situation and the user's needs. This ability to dynamically adjust enables the system to handle more complex interaction scenarios and provide more effective help and suggestions. The virtual digital human constructed using virtual reality technology not only has vivid facial expressions and body movements but also can adjust the intonation according to the emotional tendency of the answer. This multi-modal presentation makes the virtual digital human more lifelike and enhances the emotional communication and interaction experience with the user. During the interaction with the virtual digital human, users can feel more real and natural emotional feedback. By accurately extracting feature points from facial images (such as the position and shape changes of key parts like eyes, eyebrows, and mouth), the system can capture subtle facial expression changes. These feature points, as the key basis for emotion recognition, help to improve the accuracy and reliability of emotion recognition. The rapid matching process between the feature points and the preset expression categories enables the system to capture and recognize the user's emotional state in real time. This instant feedback mechanism helps the system quickly adjust the interaction strategy to better fit the user's current emotional state. Converting the matched expression category into the corresponding first emotional state (positive emotion, negative emotion, and neutral emotion) not only realizes the basic classification of emotions but also provides more refined emotional feedback to users. This detailed emotion classification helps the system better understand the user's emotional changes and thus provide more considerate and personalized services. Generating the corresponding first emotional value according to the first emotional state enables the emotional state to be quantitatively represented. This quantitative processing not only facilitates subsequent data analysis and processing but also provides an important reference basis for the system to adjust the answer expression, select appropriate interaction strategies, etc. Accurately identifying the user's emotional state and adjusting the interaction method according to the emotional state helps to promote the naturalness and fluency of human-computer interaction. Users can feel the sensitivity and attention of the system to their emotional changes, thus enhancing their trust and satisfaction with the system. By adopting the above technical solutions, extracting preset features closely related to sentiment analysis (such as sentiment words, negation words, and degree words), the system can focus on the key parts expressing emotions in the text, reduce the interference of irrelevant information, and thus improve the accuracy of sentiment analysis.The selection of preset features not only considers the words directly expressing emotions (emotional words), but also covers negative words and degree words that affect the intensity of emotions. This comprehensive feature extraction method enables the system to more accurately grasp the overall emotional tendency of the text and avoid the one-sidedness that may be brought by a single feature. Mapping the emotional tendency of the text to preset emotional categories (positive emotion, negative emotion, and neutral emotion) and generating corresponding second emotional values according to the mapping results enables the emotional state to be quantitatively represented. This quantitative evaluation method not only facilitates subsequent data processing and analysis, but also provides strong support for the system to adjust interaction strategies and optimize the user experience. The entire process, from feature extraction to emotional analysis, then to the mapping of emotional categories and the generation of values, is automatically completed based on a preset emotional analysis model. This highly automated processing method greatly improves the real-time performance of emotional analysis, enabling the system to quickly respond to the emotional changes of users. Accurately identifying the emotional tendency in the text and adjusting the interaction method according to the emotional state helps to improve the intelligent level of human-computer interaction. The system can more sensitively capture the emotional needs of users, provide more considerate and personalized services, thereby enhancing user satisfaction and loyalty.

[0018] Optionally, determining the target emotional state by combining the first emotional state and the second emotional state includes:

[0019] When the first emotional state and the second emotional state are consistent, determining the first emotional state as the target emotional state;

[0020] When the first emotional state and the second emotional state are inconsistent, performing weighted summation on the first emotional value and the second emotional value to obtain a third emotional value, determining the third emotional state corresponding to the third emotional value through a preset numerical mapping table, and determining the third emotional state as the target emotional state.

[0021] By adopting the above technical solution, when the first emotional state (based on facial expressions) and the second emotional state (based on text content) are consistent, the emotional state is directly used as the target emotional state, which simplifies the decision-making process and ensures the accuracy of emotion recognition. When the two are inconsistent, the contributions of the two emotional states are comprehensively considered through weighted summation, enabling the system to more comprehensively evaluate the user's emotional tendency, thereby enhancing the robustness of emotion recognition. In the case of inconsistency, a third emotional value is obtained through weighted summation, and the third emotional state is determined as the target emotional state by means of a preset numerical mapping table. This method not only takes into account the reliability differences of different emotional sources (for example, facial expressions and text content may have different weights when expressing emotions), but also achieves precise quantification of emotional states through the numerical mapping table, further enhancing the accuracy of emotion recognition. Determining the target emotional state by comprehensively considering facial expressions and text content enables the system to more accurately grasp the user's true emotional needs. This emotion recognition method based on multi-source information helps the system to provide more considerate and personalized services, thereby optimizing the user experience. For example, in a chatbot or customer service system, adjusting the tone and wording of the response according to the user's emotional state can improve user satisfaction and trust. By combining multiple emotion recognition technologies (such as facial expression recognition and text emotion analysis) and adopting strategies such as weighted summation to determine the target emotional state, the system can more intelligently handle complex emotional interaction scenarios. This intelligent processing method not only improves the response speed and accuracy of the system, but also enables the system to better adapt to the emotional expression methods and habits of different users.

[0022] Optionally, the determining the target emotional state by combining the first emotional state and the second emotional state includes:

[0023] Using time series analysis method to analyze the changing trend of the user's emotional state over time to predict the future emotional state;

[0024] Adjusting the third emotional state according to the future emotional state, and determining the adjusted third emotional state as the target emotional state.

[0025] By adopting the above technical solution, through the time series analysis method, the system can capture the changing trend of the user's emotional state over time and predict the future emotional state. This forward-looking ability enables the system to prepare and adjust the interaction strategy in advance to better cope with the possible emotional changes of the user, thereby providing more considerate and timely services. The user's emotional state is dynamically changing, while traditional emotion recognition methods can often only process static emotional data. By introducing time series analysis, the system can track the changes in the user's emotional state in real time and dynamically adjust the target emotional state according to the changing trend. This dynamic adaptability enables the system to more flexibly cope with complex emotional interaction scenarios and improve the accuracy and effectiveness of emotion recognition. Adjusting the target emotional state based on the prediction of the future emotional state helps the system provide a more user-expected interaction experience. For example, when it is predicted that the user may have a negative emotion, the system can adjust the tone and wording of the response in advance to relieve the user's negative emotion; when it is predicted that the user may have a positive emotion, the system can further enhance the fun and affinity of the interaction to deepen the user's positive experience. The introduction of the time series analysis method and emotion prediction marks a further improvement in the intelligent level of the system in emotion recognition. By learning and understanding the changing rules of the user's emotional state, the system can gradually build an emotional model of the user and provide more intelligent and personalized services based on this. This improvement in the intelligent level not only improves the response speed and accuracy of the system but also enables the system to better meet the emotional needs of the user. The embodiments of the present application demonstrate the latest application results of time series analysis and emotion prediction in the field of emotion computing. By continuously exploring and optimizing emotion recognition technologies and strategies, the development of the emotion computing field can be promoted, providing strong emotion analysis ability support for more intelligent applications. This technological innovation not only helps to improve the user experience and satisfaction but also brings more commercial value and social benefits to enterprises and society.

[0026] Optionally, the adjusting the expression of the first answer according to the target emotional state and the conversation context to generate a second answer includes:

[0027] Traverse the preset rule library, match the first target rule according to the current target emotional state and the conversation context, where the conversation context includes the user's previous question, the system's answer, and background information, and the first target rule includes adding preset positive words and sentences when the target emotional state is a positive emotion and there are no preset negative words in the conversation context, or adding preset comforting words and sentences when the target emotional state is a negative emotion and the conversation context contains preset negative words;

[0028] Adjust the expression of the first answer according to the first target rule;

[0029] When multiple first target rules are matched, select the finally executed second target rule according to the priorities of the multiple first target rules or a preset conflict resolution mechanism.

[0030] By adopting the above technical solution, according to the user's current emotional state and conversation context, dynamically adjust the expression mode of the answer, making the answer closer to the user's emotional needs. For example, adding preset positive words and sentences in a positive emotional state can enhance the user's sense of pleasure; adding preset comforting words and sentences in a negative emotional state can relieve the user's negative emotions. This personalized answering method helps to improve user satisfaction and loyalty. Through the analysis of the conversation context, the system can more accurately understand the user's intention and emotional state, so as to generate a more natural and fluent answer. This context-based understanding ability makes the conversation process more coherent and reduces the occurrence of misunderstandings and ambiguities. The rules in the preset rule library are summarized based on a large amount of data and experience, and have high accuracy and pertinence. By matching the target rules and adjusting the answer expression, the system can generate an answer that better meets the user's expectations and needs. When multiple first target rules are matched, the system can select the finally executed second target rule according to the priorities of the rules or a preset conflict resolution mechanism. This flexible conflict resolution mechanism ensures that the system can make reasonable decisions when facing complex situations and avoids conflicts and contradictions between rules.

[0031] Optionally, determining the facial expressions, body movements, and speech intonations of the virtual digital human according to the emotional tendency in the second answer to display the second answer includes:

[0032] When the emotional tendency in the second answer is a positive emotion, set the facial expression of the virtual digital human to a smile, a wink, or a raised eyebrow, set the body movement of the virtual digital human to a wave, a jump, or applause, and set the speech intonation of the virtual digital human to a speech rate greater than the first speech rate threshold, a volume greater than the first volume threshold, and a rising pitch;

[0033] When the emotional tendency in the second answer is a negative emotion, set the facial expression of the virtual digital human to a frown, a lowered head, or a pout, set the body movement of the virtual digital human to a lowered head, crossed arms, or a slow shake of the head, and set the speech intonation of the virtual digital human to a speech rate less than the second speech rate threshold, a volume less than the second volume threshold, and a falling pitch, where the second speech rate threshold is less than the first speech rate threshold and the second volume threshold is less than the first volume threshold.

[0034] By adopting the above technical solutions, closely associating the facial expressions, body movements, and speech intonations of the virtual digital human with the emotional tendencies in the second answer can ensure the accuracy of emotional expression. Whether it is a positive emotion or a negative emotion, it can be accurately conveyed through corresponding non-verbal signals, enabling users to more intuitively feel the emotional color behind the answer. The emotional expression method makes the interaction of the virtual digital human more vivid and interesting, capable of attracting users' attention and enhancing their sense of participation. When the virtual digital human presents a positive answer with positive gestures such as smiling and waving, it can enhance users' sense of pleasure and satisfaction; while when presenting a negative answer with negative gestures such as frowning and lowering the head, it can arouse users' resonance and attention, thereby deepening their understanding and memory of the answer. Through refined facial expression, body movement, and speech intonation design, the virtual digital human can present more realistic emotional responses, making users feel as if they are in a real conversation scenario. This sense of immersion and substitution helps to enhance users' trust and dependence on the virtual digital human, thereby promoting more in-depth interaction and cooperation. The embodiments of this application are not limited to the expression of only two emotional tendencies, positive and negative, but can also be extended and customized according to specific requirements. For example, more combinations of facial expressions, body movements, and speech intonations can be set to express different emotional levels and nuances, so as to meet more diverse and complex emotional expression needs.

[0035] In the second aspect of this application, a virtual digital human interactive Q&A system is provided, including a collection module, an emotion module, an answer module, and a virtual module, where:

[0036] The collection module is configured to collect the user's facial images and conversation content, perform a first process on the conversation content through natural language processing technology, and perform a second process on the facial images using a preset emotion analysis model to obtain a first emotional state. The first process includes word segmentation and semantic role annotation, and the second process includes facial expression feature extraction and emotion recognition;

[0037] The emotion module is configured to perform a third process on the text after the first process using a preset emotion analysis model to obtain the user's second emotional state, and determine the target emotional state by combining the first emotional state and the second emotional state;

[0038] The answer module is configured to perform a fourth process on the text after the first process based on a deep learning model to generate a first answer, and adjust the expression of the first answer according to the target emotional state and the conversation context to generate a second answer;

[0039] The virtual module is configured to construct a virtual digital human using virtual reality technology, and determine the facial expressions, body movements, and speech intonations of the virtual digital human according to the emotional tendency in the second answer to display the second answer.

[0040] Among them, the second processing of the facial image by using a preset sentiment analysis model to obtain a first emotional state includes:

[0041] Extracting feature points from the facial image, matching the feature points with preset expression categories, converting the matched expression categories into corresponding first emotional states, and generating corresponding first emotional values according to the first emotional states. The first emotional states include positive emotions, negative emotions, and neutral emotions.

[0042] The third processing of the text after the first processing by using a preset sentiment analysis model to obtain the user's second emotional state includes:

[0043] Extracting preset features related to sentiment analysis from the text after the first processing. The preset features include sentiment words, negation words, and degree words;

[0044] Inputting the preset features into a preset sentiment analysis model to obtain an output result;

[0045] Mapping the sentiment tendency of the text to a preset sentiment category according to the output result, determining the second emotional state according to the mapped sentiment category, and generating a corresponding second emotional value according to the second emotional state. The second emotional states include positive emotions, negative emotions, and neutral emotions.

[0046] In a third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. Both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes the method described in any one of the above.

[0047] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method described in any one of the above is executed.

[0048] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0049] 1. By collecting the user's facial image and conversation content and performing sentiment analysis respectively, the user's emotional state can be understood more accurately. This sentiment recognition method based on multi-modal information enables the system to integrate more naturally into the user's emotional communication, improving the naturalness and fluency of the interaction;

[0050] 2. By combining the sentiment analysis of facial images and the sentiment analysis of text content, the system can capture the user's sentiment tendency from multiple dimensions, thereby more accurately determining the target sentiment state. This multi-modal sentiment recognition technology significantly improves the accuracy and reliability of sentiment recognition;

[0051] 3. The first answer generated based on the deep learning model is adjusted according to the target sentiment state and the conversation context, making the expression of the answer more in line with the user's emotional needs and context. This personalized answer generation method can enhance the user's satisfaction and trust;

[0052] 4. The virtual digital human constructed using virtual reality technology can automatically adjust facial expressions, body movements, and speech intonations according to the sentiment tendency in the second answer. This emotional expression method enables the virtual digital human to convey information more vividly and vividly, enhancing the user's sense of immersion and substitution. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a schematic flowchart of the method for virtual digital human interactive question and answer disclosed in the embodiments of the present application;

[0054] Figure 2 is a schematic block diagram of the system for virtual digital human interactive question and answer disclosed in the embodiments of the present application;

[0055] Figure 3 is a schematic structural diagram of an electronic device disclosed in the embodiments of the present application.

[0056] Description of the reference numerals: 201, acquisition module; 202, sentiment module; 203, answer module; 204, virtual module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0058] In the description of the embodiments of the present application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "for example" or "for instance" is intended to present related concepts in a specific manner.

[0059] In the description of the embodiments of the present application, the term "plural" means two or more. For example, plural systems mean two or more systems, and plural screen terminals mean two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0060] This embodiment discloses a method for virtual digital human interactive question answering, which is applied to an interactive question answering platform. Figure 1 It is a schematic flowchart of the method for virtual digital human interactive question answering disclosed in the embodiments of the present application. As Figure 1 shown, the method includes the following steps:

[0061] S110. Collect the facial image and conversation content of the user, perform a first process on the conversation content through natural language processing technology, and perform a second process on the facial image by using a preset sentiment analysis model to obtain a first emotional state. The first process includes word segmentation and semantic role labeling, and the second process includes facial expression feature extraction and emotion recognition.

[0062] Use a camera or other image capture device to obtain the facial image of the user in real time or pre-recorded. These images contain rich non-verbal information, such as expressions, eye contact, etc., which are crucial for sentiment analysis. Capture the user's speech through an audio input device such as a microphone and convert it into text form. This usually involves speech recognition technology to ensure the accurate acquisition of the conversation content.

[0063] Natural language processing of the conversation content (i.e., the first process): (1) Split the conversation text into individual words or phrases, which is a basic step in natural language processing and helps with subsequent more complex analysis. For example, in English, the sentence "I am happy" will be word-segmented into "I", "am", "happy"; (2) Semantic role labeling (SRL): On the basis of word segmentation, further analyze the semantic relationships between the various components in the sentence. Semantic role labeling can identify the arguments of the verb, such as the agent (the initiator of the action), the patient (the object of the action), etc., so as to more accurately understand the deep meaning of the sentence. For example, in the sentence "I saw the cat with binoculars", SRL can identify that "I" is the agent of "saw", "the cat" is the patient, and "with binoculars" is the instrumental adverbial.

[0064] Emotional analysis of facial images (i.e., the second processing): (1) Facial expression feature extraction: Using computer vision technology to extract key features from facial images, such as the shape and position of eyes, eyebrows, and mouth. These features are crucial for identifying different emotional states (such as happiness, sadness, anger, etc.); (2) Emotional recognition: Comparing the extracted facial expression features with the patterns in the preset emotional analysis model to determine the user's current emotional state. The emotional analysis model is usually trained based on a large number of labeled datasets and can identify multiple emotional categories and their corresponding expression features.

[0065] Combining the natural language processing results of the conversation content with the emotional analysis results of the facial images can provide users with a more comprehensive and in-depth understanding. For example, in a customer service system, this can help customer service staff better understand the emotions and needs of users, thereby providing more considerate and effective services. In a human-computer interaction scenario, this multi-modal emotional analysis technology can also be used to enhance the emotional intelligence of robots, making their interaction with users more natural and smooth.

[0066] The second processing of the facial image using the preset emotional analysis model to obtain the first emotional state includes:

[0067] Extracting feature points from the facial image, matching the feature points with the preset expression categories, converting the matched expression categories into the corresponding first emotional state, and generating the corresponding first emotional value according to the first emotional state. The first emotional state includes positive emotions, negative emotions, and neutral emotions.

[0068] Using facial recognition technology, key feature points are extracted from the captured facial images, such as the contours, positions, and shapes of eyes, eyebrows, and mouths. The extracted feature points are encoded into numerical vectors, which can describe the relative positions and shape changes of various parts of the face and are the basis for recognizing different expressions. The interactive Q&A platform incorporates an emotion analysis model trained with a large amount of labeled data. This model associates different facial feature vectors with preset expression categories (such as happy, sad, angry, surprised, disgusted, fearful, and neutral). These expression categories are further grouped into three emotional states: positive emotions (such as happy), negative emotions (such as sad, angry, disgusted, fearful), and neutral emotions. The encoded facial feature vectors are input into the emotion analysis model, and the model finds the most matching expression category through comparison and calculation. Then, the matched expression category is converted into the corresponding first emotional state (positive, negative, or neutral). To further quantify the emotional state, the interactive Q&A platform can assign an emotional numerical range to each emotional state. For example, positive emotions can be assigned a numerical value range from 1.0 to 0.5 (1.0 represents very positive, and 0.5 represents slightly positive but close to neutral), negative emotions from -0.5 to -1.0 (-0.5 represents slightly negative but close to neutral, and -1.0 represents very negative), and neutral emotions are 0.0. According to the matched emotional state, the system generates the corresponding emotional numerical value. For example, if the user's facial image is recognized as "happy", the emotional state is positive, and the platform assigns it an emotional numerical value of 0.8.

[0069] By extracting fine feature points from facial images (such as the position and shape changes of the corners of eyes and mouths) and using a preset emotion analysis model trained with a large amount of data for matching, the accuracy of emotion recognition can be significantly improved. This method is more reliable than simply relying on the overall facial contour or a single feature point. Converting the recognized expression category into the corresponding emotional state (positive, negative, neutral) and further generating an emotional numerical value enables the emotional state to be quantitatively represented. This quantitative representation is not only convenient for computer processing but also for human understanding and comparison of the emotional changes of different users or the same user at different time points. In a human-computer interaction scenario, accurately recognizing the user's emotional state and adjusting the interaction method accordingly (such as adjusting the intonation, changing the reply content, etc.) can enhance the user experience and make the user feel more considerate and personalized services. For example, in an online education platform, adjusting the teaching strategy according to the student's emotional state can stimulate the student's learning interest and enthusiasm. In some decision-making scenarios that require emotion analysis (such as market research, customer feedback analysis, etc.), it can help decision-makers more accurately understand the emotional tendencies of the target group, thereby formulating more reasonable and effective decision-making plans.

[0070] S120. Use a preset sentiment analysis model to perform a third processing on the first processed text to obtain the second emotional state of the user, and determine the target emotional state by combining the first emotional state and the second emotional state;

[0071] Identify emotion-related words from the text. These words may include adjectives, adverbs, emotion verbs, etc., which play a key role in expressing emotions. Based on a preset emotion dictionary or sentiment analysis model, judge the polarity (positive, negative or neutral) of each emotion word. The emotion dictionary usually contains a large number of words and their corresponding emotion polarity annotations. In addition to polarity, it is also necessary to evaluate the intensity of emotion words, that is, the degree of intensity when they express emotions. This can be achieved through the weight or score of emotion words. Integrate the polarity and intensity of all emotion words in the text, and use an aggregation algorithm (such as weighted average) to judge the overall emotional state of the entire text, that is, the second emotional state. Ensure that the first emotional state (based on the facial image) and the second emotional state (based on the text content) are consistent in representation, for example, both are divided into three categories: positive, negative and neutral. According to the requirements of the application scenario, assign different weights to the first emotional state and the second emotional state. For example, in some scenarios, facial expressions may more directly reflect the user's immediate emotions, so the first emotional state should be given a higher weight; while in other scenarios, the text content may contain more context information, making the second emotional state more important. According to the assigned weights, perform weighted average or other forms of comprehensive calculation on the first emotional state and the second emotional state to determine the target emotional state. The target emotional state is a more comprehensive emotional representation that integrates information from both text and image modalities.

[0072] The use of a preset sentiment analysis model to perform a third processing on the first processed text to obtain the second emotional state of the user includes:

[0073] Extract preset features related to sentiment analysis from the text that has undergone the first processing. The preset features include emotion words, negation words and degree words;

[0074] Input the preset features into a preset sentiment analysis model to obtain an output result;

[0075] Map the emotional tendency of the text to a preset emotion category according to the output result, determine the second emotional state according to the mapped emotion category, and generate a corresponding second emotional value according to the second emotional state. The second emotional state includes positive emotion, negative emotion and neutral emotion.

[0076] Emotional words are the words in the text that directly express emotions, such as "happy", "sad", etc. These words usually have a clear emotional polarity (positive or negative), and are important bases for judging the emotional tendency of the text. Negative words such as "not", "no", etc., can change the emotional polarity of emotional words. For example, "not happy" and "happy" are opposite in emotional tendency, and here "not" is a negative word. Degree words such as "very", "slightly", etc., are used to modify the intensity of emotional words and affect the degree of emotional expression in the text. For example, "very happy" is stronger in emotional intensity than "happy". When extracting these preset features, predefined emotion lexicons, negative word lexicons, and degree word lexicons are usually used to assist in recognition. In addition, machine learning or deep learning methods can also be used to automatically identify these features through training models. The extracted preset features (emotional words, negative words, degree words and their related information) are used as inputs and fed into a pre-trained sentiment analysis model. This model may be a rule-based system or a model based on machine learning or deep learning. A series of rules are defined to match the features in the text, and the emotional tendency of the text is inferred based on these rules. This method is simple and intuitive, but the design and maintenance of the rules are relatively complex. A large amount of labeled data is used to train the model so that it can automatically learn the relationship between the features in the text and the emotional tendency. This method has good generalization ability when dealing with complex texts and unknown words.

[0077] After the model processes the input features, it will output one or more numerical values or vectors representing the emotional tendency of the text. These output results need to be mapped according to the preset emotional categories to determine the second emotional state of the text. Map the results output by the model to the preset emotional categories, such as positive emotion, negative emotion, and neutral emotion. This is usually achieved by comparing the output results with the emotional category thresholds or performing multi-classification. According to the mapping results, determine the second emotional state of the text. For example, if the results output by the model indicate that the text has a positive emotional tendency, the second emotional state is positive emotion. To further quantify the emotional state of the text, an emotional numerical range can be assigned to each emotional category. For example, positive emotion can be assigned a numerical value range from 1.0 to 0.5 (1.0 represents very positive, 0.5 represents slightly positive but close to neutral), negative emotion is from -0.5 to -1.0 (-0.5 represents slightly negative but close to neutral, -1.0 represents very negative), and neutral emotion is 0.0. Generate the corresponding emotional numerical value according to the second emotional state.

[0078] By meticulously extracting the predefined features related to sentiment analysis (such as sentiment words, negation words, and degree words) from the text, the sentiment information in the text can be captured more comprehensively. These features play a crucial role in sentiment expression, and their accurate identification helps improve the accuracy of sentiment analysis. The predefined sentiment analysis model is usually trained with a large amount of labeled data and can handle various complex text situations. Inputting the predefined features into such a model can make full use of the learning results of the model and improve the robustness of sentiment analysis in different text types and contexts. Mapping the sentiment tendency of the text to the predefined sentiment categories (positive sentiment, negative sentiment, and neutral sentiment) and generating corresponding sentiment values enables the results of sentiment analysis to be quantitatively represented. This quantitative representation is not only convenient for computer processing but also for human understanding and comparison of the sentiment changes of different texts or the same text at different time points.

[0079] Optionally, determining the target sentiment state by combining the first sentiment state and the second sentiment state includes:

[0080] When the first sentiment state and the second sentiment state are consistent, determining the first sentiment state as the target sentiment state;

[0081] When the first sentiment state and the second sentiment state are inconsistent, performing a weighted sum of the first sentiment value and the second sentiment value to obtain a third sentiment value, determining the third sentiment state corresponding to the third sentiment value through a predefined value mapping table, and determining the third sentiment state as the target sentiment state.

[0082] When the first sentiment state and the second sentiment state are consistent, it indicates that the sentiment tendency analyzed from both the facial image and the text content is the same. In this case, the first sentiment state (or the second sentiment state, since they are the same) can be directly determined as the target sentiment state. This is because the information of two different modalities (image and text) points to the same sentiment direction, increasing the credibility of sentiment judgment.

[0083] When the first emotional state and the second emotional state are inconsistent, the situation becomes complex. This may be caused by various reasons. For example, there are contradictions in the user's emotional expression (such as saying one thing but meaning another), or the facial image and the text content capture the user's emotional states at different times or in different situations. To handle this situation, the process adopts the following strategy: First, the first emotional value and the second emotional value are weighted and summed to obtain a comprehensive emotional value (the third emotional value), and the weight allocation can be adjusted according to the actual situation, such as setting according to the importance of the facial image and the text in a specific application scenario; then, through a preset numerical mapping table, the third emotional state corresponding to the third emotional value is determined. This mapping table may be a range division table that maps different emotional value ranges to specific emotional categories (such as positive, negative, neutral, or more detailed categories). By looking up this mapping table, the third emotional value can be converted into the corresponding emotional state; finally, the third emotional state is determined as the target emotional state. This state is the result of integrating two modalities of information, namely the facial image and the text, and more comprehensively and accurately reflects the user's emotional tendency.

[0084] When the first emotional state and the second emotional state are consistent, one of them is directly adopted as the target emotional state. This strategy is based on the consistency principle, which simplifies the decision-making process while maintaining a high-accuracy emotional judgment. When they are inconsistent, the target emotional state is determined through weighted summation and a preset numerical mapping table. This method comprehensively considers multiple emotional information sources, reduces the bias that may be brought by a single information source, and thus improves the accuracy of emotional analysis. Multimodal emotional analysis (combining facial images and text) can inherently improve the robustness of the system because different information sources may have different advantages and reliabilities in different situations. By integrating these information sources, the system can better handle complex and changing emotional expression situations. The use of weighted summation and a preset numerical mapping table further enhances the robustness of the system, enabling the system to still make reasonable judgments in the face of inconsistent emotional information. In scenarios where user interaction is based on emotional analysis (such as intelligent customer service, online education, etc.), accurately judging the user's emotional state is crucial for enhancing the user experience. By integrating multiple emotional information sources to determine the target emotional state, the system can more accurately understand the user's emotional needs and thus provide more considerate and personalized services. This technical process is not only applicable to the emotional analysis of facial images and text but can also be extended to the combination of other multimodal emotional information sources (such as voice, physiological signals, etc.). This enables this technology to be applied to a wider range of scenarios, such as emotional computing, mental health monitoring, and human-computer interaction.

[0085] Optionally, the determining the target emotional state by combining the first emotional state and the second emotional state includes:

[0086] Use time series analysis methods to analyze the changing trend of the user's emotional state over time to predict the future emotional state;

[0087] Adjust the third emotional state according to the predicted future emotional state, and determine the adjusted third emotional state as the target emotional state.

[0088] Use time series analysis methods (such as ARIMA models, LSTM networks, etc.) to analyze the changing trend of the user's emotional state over time in the past period. This includes collecting the emotional values of the user at each time point in the past period (which can be the first emotional value, the second emotional value, or the third emotional value, depending on the actual situation and requirements), and constructing time series data. Through the time series model, the emotional state of the user at a future time point can be predicted, that is, the future emotional state. According to the predicted future emotional state, it may be found that there are differences between it and the current third emotional state obtained through multimodal analysis. At this time, the third emotional state can be fine-tuned according to the predicted future emotional state to better reflect the future emotional changes of the user. The adjusted third emotional state is the finally determined target emotional state. It not only considers the current multimodal emotional information but also incorporates the prediction of the changing trend of the emotional state over time, so it is more comprehensive and accurate.

[0089] The ways to adjust the third emotional state according to the future emotional state can be as follows: (1) Combine the change trend of the future emotional state with the current third emotional state to form a new emotional state that integrates the current state and the future trend. Specifically, first analyze the change trend of the future emotional state (such as rising, falling, stable, etc.), and then make fine-tuning to the third emotional state according to the intensity and direction of the trend. For example, if the prediction shows that the future emotional state will rise significantly, a certain positive emotional component can be added to the current third emotional state. (2) Set thresholds for emotional states. When the predicted future emotional state exceeds or is lower than these thresholds, trigger the adjustment of the third emotional state. Specifically, according to the requirements of the application scenario, set different emotional thresholds (such as positive emotional threshold, negative emotional threshold, etc.). When the predicted future emotional state exceeds or is lower than these thresholds, automatically adjust the third emotional state to reflect this change. For example, if the predicted future emotional state is lower than the negative emotional threshold, it may be necessary to adjust the third emotional state to a more negative state. (3) Consider the current scenario and context information where the user is located and make an adaptive adjustment to the third emotional state. Specifically, analyze the current scenario and context where the user is located through various data sources such as the user's behavior data, location information, time information, etc. Then, make a targeted adjustment to the third emotional state based on this information and the predicted future emotional state. For example, if the user is currently in a work scenario and the predicted future emotional state will become negative, it may be necessary to adjust the third emotional state to reduce its impact on work.

[0090] Through time series analysis methods, the system can capture the dynamic trend of the user's emotional state changing over time, thus more accurately predicting the future emotional state. This prediction is not only based on the current emotional state but also takes into account the changing patterns of historical emotional data, so it has higher accuracy. Traditional emotional analysis often only focuses on the emotional state at a certain moment or within a certain period of time, while time series analysis methods can span the time axis and regard the user's emotional state as a continuously changing process. This continuity makes emotional analysis more comprehensive and in-depth, helping to discover the laws and patterns of emotional changes. The prediction of the future emotional state provides a strong basis for adjusting the third emotional state. By making fine-tuning to the third emotional state according to the prediction results, the system can ensure that the target emotional state is closer to the user's future real emotional state. This optimization not only improves the accuracy of emotional analysis but also enhances the adaptability and flexibility of the system. Accurately predicting and adjusting the target emotional state helps to provide more personalized and considerate services for users. For example, on a social media platform, the system can recommend appropriate content or activities according to the user's future emotional state to relieve negative emotions or enhance positive experiences. This service based on emotional prediction can significantly improve user satisfaction and loyalty.

[0091] S130. Perform a fourth processing on the first processed text based on a deep learning model to generate a first answer, and adjust the expression of the first answer according to the target emotional state and the conversation context to generate a second answer;

[0092] A deep learning model is a neural network model that has been trained to understand and process natural language text. Such models typically include, but are not limited to, recurrent neural networks (RNNs), long short-term memory networks (LSTMs), Transformer, etc. They can capture semantic information, context relationships, etc. in the text. Use the deep learning model to perform further processing and analysis on the first processed text to generate a preliminary answer (i.e., the first answer). This process may include multiple steps such as text understanding, intent recognition, information extraction, etc., depending on the training objectives and task requirements of the model. Based on the fourth processing result of the deep learning model, the system can generate a preliminary answer. This answer may be a simple reply, an information summary, or a complex sentence, depending on the complexity of the question and the training effect of the model. The conversation context refers to the historical record of the current conversation, including the user's previous questions, the system's responses, and their interaction relationships. The conversation context is crucial for understanding the user's intent, predicting the user's behavior, and generating appropriate answers. After obtaining the first answer, the system will adjust the expression of the answer according to the target emotional state and the conversation context. Such adjustments may include changing the tone, adding or subtracting emotional words, adjusting the sentence structure, etc., with the aim of making the answer more in line with the user's emotional needs and conversation situation. After the above adjustments, the system finally generates a more accurate, considerate, and user-expected answer (i.e., the second answer). This answer not only answers the user's question but also takes into account the user's emotional state and the conversation context, thus enhancing the user experience and satisfaction.

[0093] Optionally, the adjusting the expression of the first answer according to the target emotional state and the conversation context to generate a second answer includes:

[0094] Traverse the preset rule library, match the first target rule according to the current target emotional state and the conversation context, where the conversation context includes the user's previous question, the system's answer, and background information, and the first target rule includes adding preset positive words or sentences when the target emotional state is a positive emotion and there are no preset negative words in the conversation context, or adding preset comforting words or sentences when the target emotional state is a negative emotion and the conversation context contains preset negative words;

[0095] Adjust the expression of the first answer according to the first target rule;

[0096] When multiple first target rules are matched, select the finally executed second target rule according to the priorities of the multiple first target rules or a preset conflict resolution mechanism.

[0097] The preset rule library is a database or collection that stores various adjustment rules, which are used to guide how to adjust the expression of the answer according to the emotional state and conversation context. Each rule in the rule library defines specific conditions (such as emotional state, keyword vocabulary in the conversation context, etc.) and corresponding operations (such as adding specific words and sentences, changing the tone, etc.). The system will first traverse all the rules in the preset rule library and check one by one whether they are applicable to the current target emotional state and conversation context. Through traversal, the system can find all the rules that may be applicable and prepare for subsequent matching and adjustment.

[0098] The system will match the most suitable rule, that is, the first target rule, according to the current target emotional state and conversation context (including the user's previous questions, the system's answers, and background information). When the target emotional state is a positive emotion and there are no preset negative words in the conversation context, the system may choose to add preset positive words and sentences; while when the target emotional state is a negative emotion and the conversation context contains preset negative words, the system may choose to add preset comforting words and sentences. When the first target rule is found, the system will adjust the expression of the first answer according to the requirements of the rule. This may include adding, deleting, or replacing specific words and sentences, as well as changing the strength of the tone. The adjusted answer (i.e., the second answer) will be more in line with the user's emotional needs and conversation situation. In some cases, the system may match multiple first target rules, and there may be conflicts or overlaps between these rules. To handle this situation, the system usually selects the finally executed second target rule according to the priorities of the rules or a preset conflict resolution mechanism. The priorities can be determined based on the importance of the rules, the scope of application, or other preset criteria; while the conflict resolution mechanism may include strategies such as rule merging, rule replacement, or rule ignoring.

[0099] By considering the target emotional state and the conversation context, the system can generate more personalized answers. These answers not only answer the user's questions but also incorporate the user's emotional needs and conversation situation, thus enhancing user satisfaction and a sense of belonging. When the system detects the user's positive emotion, by adding preset positive words and phrases, it can further stimulate the user's positive mood and enhance the emotional resonance with the user. Similarly, when detecting the user's negative emotion, adding preset comforting words and phrases helps relieve the user's negative mood and reflects the system's care and support. By traversing the preset rule library and matching the most applicable rules, the system can ensure that the generated answers are both accurate and relevant. This rule-based adjustment method avoids the randomness and uncertainty of answers and improves the reliability and stability of the system. This application can handle situations containing multiple emotional states and complex conversation contexts. Through the preset conflict resolution mechanism, when multiple rules are matched, the system can select the most appropriate rule to execute, thus ensuring that the generated answers meet the user's expectations and avoid internal contradictions. By continuously updating and optimizing the preset rule library, the system can continuously learn and adapt to new conversation situations and emotional states. This intelligence and adaptability enable the system to better meet the user's needs and enhance the user experience. Since the rules in the rule library are predefined, the system can quickly traverse and match the most applicable rules, thus quickly generating the second answer. This efficient processing method improves the system's response speed and efficiency, enabling the user to obtain a satisfactory answer faster.

[0100] S140. Construct a virtual digital human using virtual reality technology, and determine the facial expression, body movement, and intonation of the virtual digital human according to the emotional tendency in the second answer to display the second answer.

[0101] Based on VR (Virtual Reality) technology, highly realistic virtual digital humans can be constructed. These digital humans not only have a realistic appearance but also can perform complex interactions and action performances. Through technologies such as 3D modeling, texture mapping, and skeletal animation, virtual digital humans with various forms and smooth movements can be created. The second answer is generated according to the user's question and the conversation context, which contains a specific emotional tendency, such as positive, negative, or neutral. This emotional tendency is a direct reflection of the user's emotional state and is also an important basis for adjusting the performance of virtual digital humans. By using natural language processing (NLP) and sentiment analysis technologies, the emotional tendency of the text in the second answer can be identified. By analyzing the lexical, syntactic, and semantic information in the text, the emotional type expressed by the answer can be judged. According to the emotional tendency in the second answer, the facial expressions of virtual digital humans can be dynamically adjusted. For example, when the answer expresses a positive emotion, the digital human can be made to show a smile or a happy expression. In addition to facial expressions, the body movements of virtual digital humans can also be adjusted according to the emotional tendency. Appropriate body movements can enhance the expression effect of emotions and make the performance of digital humans more vivid and realistic. For example, when expressing a positive emotion, the digital human can be made to wave or jump. Voice is one of the important ways to express emotions. By adjusting the intonation of virtual digital humans, the expression effect of emotions can be further enhanced. For example, when expressing a positive emotion, the voice of the digital human can be made brighter and more cheerful. Users can not only see the text content of the second answer but also feel the emotional tendency expressed by the answer through the facial expressions, body movements, and intonation of virtual digital humans.

[0102] Optionally, determining the facial expressions, body movements, and intonation of the virtual digital human according to the emotional tendency in the second answer to display the second answer includes:

[0103] When the emotional tendency in the second answer is a positive emotion, set the facial expression of the virtual digital human to a smile, a wink, or a raised eyebrow, set the body movement of the virtual digital human to wave, jump, or clap, and set the intonation of the virtual digital human to a speech rate greater than the first speech rate threshold, a volume greater than the first volume threshold, and a rising pitch;

[0104] When the emotional tendency in the second answer is a negative emotion, set the facial expression of the virtual digital human to a frown, a lowered head, or a pout, set the body movement of the virtual digital human to a lowered head, crossed arms, or a slow shake of the head, and set the intonation of the virtual digital human to a speech rate less than the second speech rate threshold, a volume less than the second volume threshold, and a falling pitch, where the second speech rate threshold is less than the first speech rate threshold and the second volume threshold is less than the first volume threshold.

[0105] When the emotional tendency in the second answer is positive, the facial expressions of the virtual digital human will be set to actions expressing positive emotions such as smiling, blinking, or raising eyebrows. These expressions can intuitively convey positive emotions such as joy, happiness, or satisfaction, enhancing the user's positive feelings. At the same time, the body movements of the virtual digital human will also be adjusted accordingly, such as waving, jumping, or clapping. These actions further strengthen the expression of positive emotions, enabling users to feel the excitement and vitality of the digital human. In terms of voice, the speech rate of the virtual digital human will be greater than the first speech rate threshold, indicating that its speech is fluent and full of vitality; the volume will be greater than the first volume threshold to ensure clear communication of information; and the pitch will rise, presenting a relaxed and pleasant atmosphere.

[0106] On the contrary, when the emotional tendency in the second answer is negative, the facial expressions of the virtual digital human will turn into actions expressing negative emotions such as frowning, lowering the head, or pouting. These expressions accurately convey negative emotions such as sadness, frustration, or dissatisfaction, enabling users to feel the emotional resonance of the digital human. In terms of body movements, the virtual digital human will show actions such as lowering the head, crossing the arms, or slowly shaking the head. These actions reflect the heaviness and uneasiness in its heart, further deepening the expression of negative emotions. In terms of voice, the speech rate of the virtual digital human will be less than the second speech rate threshold (this threshold is less than the first speech rate threshold when the emotion is positive), indicating that its speech is slow and heavy; the volume will be less than the second volume threshold (this threshold is less than the first volume threshold when the emotion is positive) to reflect its low mood; and the pitch will drop, creating a sad or depressing atmosphere.

[0107] By precisely matching the emotional tendency with the corresponding facial expressions, body movements, and intonation of speech, the embodiments of the present application can ensure that the virtual digital human accurately conveys the emotional information in the second answer. Whether it is a positive emotion or a negative emotion, users can intuitively feel the emotional color contained in the answer through the performance of the virtual digital human, thereby improving the accuracy and effectiveness of information transmission. As a medium for interacting with users, the vivid and realistic performance of the virtual digital human can greatly enhance the user experience. When the virtual digital human can make corresponding responses according to the emotional tendency, users will feel more natural and cordial, as if communicating with a real individual. This immersive experience can enhance the user's sense of participation and satisfaction, and improve the user's trust and dependence on the system. By dynamically adjusting the facial expressions, body movements, and intonation of speech of the virtual digital human, the embodiments of the present application make the interaction process more vivid and interesting. Users no longer simply receive text or voice information, but can feel the ups and downs of emotions by observing the performance of the virtual digital human, thereby increasing the fun and attractiveness of the interaction. The embodiments of the present application are not limited to the expression of only two emotional tendencies, positive and negative, but can also be extended to more types of emotional expressions according to actual needs. By presetting or learning more emotional rules and expression methods, the virtual digital human can handle more complex and changeable emotional communication scenarios and provide more comprehensive and personalized services.

[0108] This embodiment also discloses a system for interactive question and answer of a virtual digital human, Figure 2 which is a schematic diagram of the modules of the system for interactive question and answer of a virtual digital human disclosed in the embodiments of the present application. As Figure 2 shown, the system includes a collection module 201, an emotion module 202, an answer module 203, and a virtual module 204, where:

[0109] The collection module 201 is configured to collect the facial images and conversation content of the user, perform a first processing on the conversation content through natural language processing technology, and perform a second processing on the facial images using a preset emotion analysis model to obtain a first emotional state. The first processing includes word segmentation and semantic role annotation, and the second processing includes facial expression feature extraction and emotion recognition;

[0110] The emotion module 202 is configured to perform a third processing on the text after the first processing using a preset emotion analysis model to obtain the second emotional state of the user, and determine the target emotional state by combining the first emotional state and the second emotional state;

[0111] The answer module 203 is configured to perform a fourth processing on the text after the first processing based on a deep learning model to generate a first answer, and adjust the expression mode of the first answer according to the target emotional state and the conversation context to generate a second answer;

[0112] The virtual module 204 is configured to construct a virtual digital human using virtual reality technology, and determine the facial expressions, body movements, and intonations of the virtual digital human according to the emotional tendency in the second answer to display the second answer.

[0113] Among them, the second processing of the facial image using the preset emotion analysis model to obtain the first emotional state includes:

[0114] Extract feature points from the facial image, match the feature points with preset expression categories, convert the matched expression categories into corresponding first emotional states, and generate corresponding first emotional values according to the first emotional states. The first emotional states include positive emotions, negative emotions, and neutral emotions.

[0115] The third processing of the text after the first processing using the preset emotion analysis model to obtain the user's second emotional state includes:

[0116] Extract preset features related to emotion analysis from the text after the first processing. The preset features include emotion words, negative words, and degree words.

[0117] Input the preset features into the preset emotion analysis model to obtain an output result.

[0118] Map the emotional tendency of the text to the preset emotion categories according to the output result, determine the second emotional state according to the mapped emotion categories, and generate corresponding second emotional values according to the second emotional states. The second emotional states include positive emotions, negative emotions, and neutral emotions.

[0119] Optionally, the emotion module 202 is configured to:

[0120] When the first emotional state and the second emotional state are consistent, determine the first emotional state as the target emotional state.

[0121] When the first emotional state and the second emotional state are inconsistent, perform a weighted sum of the first emotional value and the second emotional value to obtain a third emotional value, determine the third emotional state corresponding to the third emotional value through a preset numerical mapping table, and determine the third emotional state as the target emotional state.

[0122] Optionally, the emotion module 202 is configured to:

[0123] Use time series analysis methods to analyze the changing trend of the user's emotional state over time to predict the future emotional state.

[0124] Adjust the third emotional state according to the future emotional state, and determine the adjusted third emotional state as the target emotional state.

[0125] Optionally, the answer module 203 is configured to:

[0126] Traverse a preset rule library, match a first target rule according to the current target emotional state and the conversation context, where the conversation context includes the user's previous questions, the system's answers, and background information, and the first target rule includes that when the target emotional state is a positive emotion and there is no preset negative vocabulary in the conversation context, add preset positive words and sentences, or when the target emotional state is a negative emotion and the conversation context contains preset negative vocabulary, add preset comforting words and sentences;

[0127] Adjust the expression mode of the first answer according to the first target rule;

[0128] When multiple first target rules are matched, select the finally executed second target rule according to the priorities of the multiple first target rules or a preset conflict resolution mechanism.

[0129] Optionally, the virtual module 204 is configured to:

[0130] When the emotional tendency in the second answer is a positive emotion, set the facial expression of the virtual digital human to smiling, winking or raising eyebrows, set the body movements of the virtual digital human to waving, jumping or applauding, and set the speech intonation of the virtual digital human to a speech rate greater than the first speech rate threshold, a volume greater than the first volume threshold, and a rising pitch;

[0131] When the emotional tendency in the second answer is a negative emotion, set the facial expression of the virtual digital human to frowning, lowering the head or pouting, set the body movements of the virtual digital human to lowering the head, crossing the arms or slowly shaking the head, and set the speech intonation of the virtual digital human to a speech rate less than the second speech rate threshold, a volume less than the second volume threshold, and a falling pitch, where the second speech rate threshold is less than the first speech rate threshold, and the second volume threshold is less than the first volume threshold.

[0132] It should be noted that: when the device provided in the above embodiment realizes its functions, only the above-mentioned division of each functional module is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be repeated here.

[0133] This embodiment also discloses an electronic device, referring toFigure 3 , the electronic device may include: at least one processor 301, at least one communication bus 302, a user interface 303, a network interface 304, and at least one memory 305.

[0134] Among them, the communication bus 302 is used to implement connection communication between these components.

[0135] Among them, the user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.

[0136] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0137] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server using various interfaces and lines, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, as well as by calling data stored in the memory 305, it performs various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.

[0138] Among them, the memory 305 may include a Random Access Memory (RAM), or may include a Read-Only Memory. Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store data involved in the above-mentioned method embodiments. Optionally, the memory 305 may also be at least one storage device located far from the aforementioned processor 301. As Figure 3 shown, in the memory 305 as a computer storage medium, an operating system, a network communication module, a user interface module, and an application program of the method for virtual digital human interactive Q&A may be included.

[0139] In Figure 3 the electronic device shown, the user interface 303 is mainly used to provide an input interface for the user and obtain the data input by the user; and the processor 301 can be used to call the application program of the method for virtual digital human interactive Q&A stored in the memory 305. When executed by one or more processors 301, the electronic device is caused to execute the method of one or more of the above-mentioned embodiments.

[0140] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0141] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0142] In several embodiments provided in this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0143] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0144] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0145] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. And the aforementioned memory 305 includes: various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0146] The above are only exemplary embodiments of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, all equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will easily think of other implementation schemes of the present disclosure after considering the disclosure of the specification. This application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A method for interactive question and answer of a virtual digital human, characterized in that: Applied to an interactive question-answering platform, the method comprises: Collecting a user's facial image and conversation content, performing a first process on the conversation content using a natural language processing technology, and performing a second process on the facial image using a preset sentiment analysis model to obtain a first sentiment state, wherein the first process includes word segmentation and semantic role labeling, and the second process includes facial expression feature extraction and sentiment recognition; Performing a third processing on the first processed text using a preset sentiment analysis model to obtain a second sentiment state of the user, and determining a target sentiment state by combining the first sentiment state and the second sentiment state; Performing a fourth processing on the first processed text based on the deep learning model to generate a first answer, and adjusting the expression of the first answer according to the target emotional state and the conversation context to generate a second answer; A virtual digital person is constructed using virtual reality technology, and facial expressions, body movements, and voice intonation of the virtual digital person are determined according to the emotional tendency in the second answer to display the second answer. The step of performing a second processing on the facial image using a preset emotion analysis model to obtain a first emotion state includes: Extracting feature points from the facial image, matching the feature points with preset expression categories, converting the matched expression categories into corresponding first emotional states, and generating corresponding first emotional values ​​according to the first emotional states, wherein the first emotional states include positive emotions, negative emotions, and neutral emotions, The performing a third processing on the first processed text using a preset sentiment analysis model to obtain a second sentiment state of the user comprises: Extracting preset features related to sentiment analysis from the text after the first processing, wherein the preset features include sentiment words, negative words, and degree words; Inputting the preset features into a preset sentiment analysis model to obtain an output result; According to the output result, the emotional tendency of the text is mapped to a preset emotional category, a second emotional state is determined according to the mapped emotional category, and a corresponding second emotional value is generated according to the second emotional state, wherein the second emotional state includes positive emotion, negative emotion and neutral emotion, The step of adjusting the expression of the first answer according to the target emotional state and the conversation context to generate a second answer includes: Traversing the preset rule library, matching the first target rule according to the current target emotional state and the conversation context, wherein the conversation context includes the user's previous questions, the system's answers, and background information, and the first target rule includes adding preset positive words when the target emotional state is positive and the conversation context does not contain preset negative words, or adding preset comforting words when the target emotional state is negative and the conversation context contains preset negative words; adjusting the expression of the first answer according to the first target rule; When multiple first target rules are matched, the second target rule to be finally executed is selected according to the priorities of the multiple first target rules or a preset conflict resolution mechanism.

2. The method for interactive question and answer of a virtual digital human according to claim 1, characterized in that: The determining a target emotional state by combining the first emotional state and the second emotional state comprises: When the first emotional state and the second emotional state are consistent, determining the first emotional state as a target emotional state; When the first emotional state and the second emotional state are inconsistent, the first emotional value and the second emotional value are weightedly summed to obtain a third emotional value, and the third emotional state corresponding to the third emotional value is determined through a preset value mapping table, and the third emotional state is determined as the target emotional state.

3. The method for interactive question and answer of a virtual digital human according to claim 2, characterized in that: The determining a target emotional state by combining the first emotional state and the second emotional state comprises: Use time series analysis methods to analyze the changing trend of users' emotional states over time to predict future emotional states; The third emotional state is adjusted according to the future emotional state, and the adjusted third emotional state is determined as the target emotional state.

4. The method for interactive question and answer of a virtual digital human according to claim 1, characterized in that: Determining the facial expression, body movement and voice tone of the virtual digital person according to the emotional tendency in the second answer to display the second answer includes: When the emotional tendency in the second answer is a positive emotion, the facial expression of the virtual digital person is set to smile, blink or raise eyebrows, the body movement of the virtual digital person is set to wave, jump or applaud, and the voice tone of the virtual digital person is set to a speed greater than a first speed threshold, a volume greater than a first volume threshold and an upward tone; When the emotional tendency in the second answer is a negative emotion, the facial expression of the virtual digital person is set to frown, lower the head or pout, the body movement of the virtual digital person is set to lower the head, cross the arms or shake the head slowly, and the voice tone of the virtual digital person is set to a speaking speed less than a second speaking speed threshold, a volume less than a second volume threshold, and a descending tone, the second speaking speed threshold is less than the first speaking speed threshold, and the second volume threshold is less than the first volume threshold.

5. A virtual digital human interactive question-answering system, characterized in that: It includes acquisition module, emotion module, answer module and virtual module, among which: A collection module configured to collect a user's facial image and conversation content, perform a first process on the conversation content using a natural language processing technology, and perform a second process on the facial image using a preset sentiment analysis model to obtain a first sentiment state, wherein the first process includes word segmentation and semantic role labeling, and the second process includes facial expression feature extraction and sentiment recognition; an emotion module, configured to perform a third processing on the first processed text using a preset emotion analysis model to obtain a second emotion state of the user, and determine a target emotion state by combining the first emotion state and the second emotion state; an answer module, configured to perform a fourth processing on the first processed text based on a deep learning model to generate a first answer, and adjust an expression of the first answer according to the target emotional state and the conversation context to generate a second answer; a virtual module configured to construct a virtual digital person using virtual reality technology, and determine the facial expression, body movement and voice intonation of the virtual digital person according to the emotional tendency in the second answer to display the second answer, The step of performing a second processing on the facial image using a preset emotion analysis model to obtain a first emotion state includes: Extracting feature points from the facial image, matching the feature points with preset expression categories, converting the matched expression categories into corresponding first emotional states, and generating corresponding first emotional values ​​according to the first emotional states, wherein the first emotional states include positive emotions, negative emotions, and neutral emotions, The performing a third processing on the first processed text using a preset sentiment analysis model to obtain a second sentiment state of the user comprises: Extracting preset features related to sentiment analysis from the text after the first processing, wherein the preset features include sentiment words, negative words, and degree words; Inputting the preset features into a preset sentiment analysis model to obtain an output result; According to the output result, the emotional tendency of the text is mapped to a preset emotional category, a second emotional state is determined according to the mapped emotional category, and a corresponding second emotional value is generated according to the second emotional state, wherein the second emotional state includes positive emotion, negative emotion and neutral emotion, The step of adjusting the expression of the first answer according to the target emotional state and the conversation context to generate a second answer includes: Traversing the preset rule library, matching the first target rule according to the current target emotional state and the conversation context, wherein the conversation context includes the user's previous questions, the system's answers, and background information, and the first target rule includes adding preset positive words when the target emotional state is positive and the conversation context does not contain preset negative words, or adding preset comforting words when the target emotional state is negative and the conversation context contains preset negative words; adjusting the expression of the first answer according to the first target rule; When multiple first target rules are matched, the second target rule to be finally executed is selected according to the priorities of the multiple first target rules or a preset conflict resolution mechanism.

6. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 4 is performed.

Citation Information

Patent Citations

  • Emotion icon recommending method based on timing sequence analysis user conversation emotion tendency

    CN107729320A

  • Intelligent question and answer method and device based on emotion recognition, electronic equipment, and medium

    CN114036280A

  • Emotion recognition method and device, electronic equipment and storage medium

    CN116361739A

  • Human-computer interaction method and device, computer equipment and storage medium

    CN117993395A

  • Virtual digital human system and method for learning disorder group

    CN118519524A