Interaction processing method and system based on digital human
Through the digital human interaction system integrating multimodal information processing module, the existing digital humans' lack of emotional expression and understanding of complex emotional needs is solved, and a more natural and realistic interactive experience is achieved.
Patent Information
- Application Number
- CN202510134908.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-06
AI Technical Summary
When existing digital people interact, their emotional expression lacks the emotional depth and complexity of real humans, and it is difficult to accurately understand and respond to complex emotional needs, resulting in insufficient naturalness and sense of reality of the interactive experience.
Design an interactive processing system based on digital people, integrating speech recognition module, natural language processing module, emotion recognition module, dialogue management module, speech synthesis module, visual perception module and behavioral decision-making and feedback module. Through the combination of multimodal information and deep learning algorithms, users' voice slight and heavy data and facial expression data are obtained and analyzed in real time, and the digital people's response method is adjusted to match the user's emotional state.
It improves the emotional reasoning ability of digital people, enables them to provide responses that are more in line with human emotional logic, enhances emotional recognition ability, improves understanding and response to users' emotional needs, and creates a more warm interactive experience.
Smart Images

Figure CN120103969A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital humans, and in particular relates to an interactive processing method and system based on digital humans. Background Art
[0002] With the advancement of AI technology, especially the breakthroughs in natural language processing (NLP), computer vision and speech recognition technology, digital humans have been widely used as virtual assistants or service personnel in multiple fields. They can simulate human appearance, behavior, emotional expression and other characteristics, and provide a more natural and humanized interactive experience. Digital humans can provide 24 / 7 services without being restricted by time and space, and provide real-time interaction in education, medical care, entertainment and other industries to improve user experience. Through more intuitive and friendly interaction methods, they can enhance user participation and satisfaction. In summary, the interactive processing methods and systems based on digital humans have important background significance, which not only promotes the development of artificial intelligence technology, but also changes the way humans interact with machines, bringing new opportunities and challenges to multiple industries.
[0003] Although existing digital humans can simulate certain emotional responses when interacting, their emotional expressions are often based on preset models and lack the emotional depth and complexity of real humans. Digital humans may not be able to accurately understand and respond to some complex emotional needs, resulting in a lack of naturalness and realism in the interactive experience. In addition, digital humans may still misunderstand or fail to accurately reason when understanding complex contexts, metaphors, or ambiguous language. When faced with complex problems, digital humans may not be able to provide the same level of in-depth analysis and reasoning as humans. Summary of the invention
[0004] The technical problem to be solved by the present invention is to overcome the shortcomings of the above-mentioned prior art and provide an interactive processing method and system based on digital human.
[0005] The technical solution adopted to solve the above technical problems is: an interactive processing system based on digital human, characterized in that: the interactive processing system includes a speech recognition module, a natural language processing module, an emotion recognition module, a dialogue management module, a speech synthesis module, a visual perception module and a behavior decision and feedback module;
[0006] The emotion recognition module is responsible for automatically generating soothing sentences based on the position of the voice interaction user, and obtaining the voice severity data and different facial expression data of the voice interaction user of the digital human in real time;
[0007] The speech recognition module recognizes different speech inputs and converts them into processable text information through speech recognition technology;
[0008] The natural language processing module is used to understand and parse the text input by the user to extract valid information and understand the user's intention;
[0009] The dialog management module is used to process user input, generate responses and maintain the context of the dialog;
[0010] The speech synthesis module is used to convert text into speech output, so that the digital human is presented to the user in a natural and fluent manner;
[0011] The visual perception module captures the user's visual signals through a camera to identify the user's emotion and action non-verbal information;
[0012] The behavior decision and feedback module makes corresponding behavior decisions according to the user's input, emotional state and target task.
[0013] An interactive method based on a digital human interactive processing system comprises the following specific steps:
[0014] S1. Performing contextual reinforcement learning based on the device for voice interaction through the voice recognition module and the visual perception module, introducing a deep learning algorithm and establishing an emotional memory library, acquiring the voice weight data and different facial expression data for the digital human in real time, and determining the changes in the voice weight and facial expression of the user according to the voice weight data and different facial expression data;
[0015] S2, based on the voice weight and facial recognition results of the voice interaction user, determining the historical interaction of the emotion data of the voice interaction user, and determining the degree of change of the emotion data of the voice interaction user according to the historical interaction, and determining the interactive voice data of the digital person and the facial voice coordination through the analysis results of the degree of change of the emotion data and the voice weight data of the voice interaction user;
[0016] S3, the dialogue management module determines and stores emotion recognition results of different interaction times based on the voice weight data and different facial expression data of the user in the historical interaction times, and adjusts the interactive emotions of the digital human according to the changes in the emotion results of different interaction times in different time periods;
[0017] S4. Determine the adjustment result of the facial and voice emotions of the interactive digital person through the change of the emotion results in different interaction times in different time periods and the voice weight and facial change results in the current interaction times.
[0018] Furthermore, the emotion recognition module includes a knowledge graph and an inference engine, and the facial expression data is determined based on the knowledge graph and the inference engine in combination with analysis and inference results of facial image analysis of voice interaction users in different time periods.
[0019] Further, determining that there is no abnormality in the user's emotions specifically includes:
[0020] A1. Analyze the user's speech, text and voice to identify the user's emotional state, and use a natural language processing module to analyze the user's tone, vocabulary and emotional color to determine the abnormal fluctuation of the emotional value and determine the abnormal emotional value of the voice;
[0021] A2. Determine the expression emotion outliers by determining the changes of different facial data in the historical data and the results of the changes of different facial data, and based on the emotion recognition results of the changes of different facial data in the historical data;
[0022] A3. Determine the user's emotion abnormality value based on the determination of the voice emotion abnormality value and the determination of the expression emotion abnormality value, and determine the abnormal emotional state of the user according to the determination of the user's emotion abnormality value.
[0023] Further, determining the abnormal situation of the user's emotion according to the abnormal value of the user's emotion specifically includes: when the abnormal value of the user's emotion does not meet the requirement, determining that the user's emotion is abnormal.
[0024] Further, determining that there is no abnormality in the user's emotions specifically includes:
[0025] Analyze the user's speech content through the natural language processing module to determine his emotional tendency, compare the user's emotional fluctuations in different dialogue stages and situations, and if the emotional fluctuations are too large, determine that the user's emotional expression is abnormal;
[0026] The emotional features in the user's voice, such as speaking speed, pitch and intonation fluctuations, are analyzed by the behavioral decision and feedback module to determine the user's emotional state. At the same time, the user's facial reaction is analyzed by the visual perception module to determine that the user exhibits an emotional state consistent with the voice and language.
[0027] Furthermore, the number of interaction sentences that determine whether the emotion recognition results meet the specified requirements includes: identifying sentences with inconsistent emotions, extreme emotions, and sentences that cannot be correctly parsed through dialogue framework analysis, and recording the number of all sentences that "do not meet the requirements" during the interaction process.
[0028] The beneficial effects of the present invention are as follows:
[0029] (1) The present invention improves emotional reasoning capabilities through emotional data from historical interactions, enabling the digital human to adjust its current response method based on the user's past emotional state, produce a response that is more in line with human emotional logic, enhance the digital human's emotional recognition ability, and combine multimodal information such as voice, facial expressions, context, and body language to more comprehensively understand the user's emotional state. It also combines knowledge graphs with reasoning engines to enable the digital human to analyze and reason about complex problems from multiple perspectives.
[0030] (2) The present invention improves the digital human's ability to understand the emotional needs of each user by collecting and analyzing the user's personalized data. The digital human can make more personalized emotional responses based on the user's interests, emotional state, and historical interaction data, creating a warmer interactive experience. Through continuous learning and optimization, the digital human can continuously improve its emotional understanding and reasoning capabilities. The digital human can analyze user feedback after each interaction, self-adjust and optimize its own emotional response and language understanding, so that it gradually tends to be more in line with human emotions and cognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a framework diagram of an interactive processing system based on digital human;
[0032] Figure 2 The invention is a flow chart of an interactive processing method based on a digital human. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0034] like Figure 1-Figure 2 As shown, this embodiment is an interactive processing system based on a digital human, and the interactive processing system includes a speech recognition module, a natural language processing module, an emotion recognition module, a dialogue management module, a speech synthesis module, a visual perception module and a behavior decision and feedback module.
[0035] The emotion recognition module is responsible for automatically generating soothing sentences based on the position of the voice interaction user, and obtaining the voice weight data and different facial expression data of the voice interaction user of the digital human in real time. The emotion recognition module includes a knowledge graph and an inference engine. The facial expression data is determined based on the knowledge graph and the inference engine combined with the analysis and inference results of the facial image analysis of the voice interaction user in different time periods. The abnormal emotion of the user is determined based on the abnormal value of the user's emotion, specifically including: when the abnormal value of the user's emotion does not meet the requirements, it is determined that the user's emotion is abnormal, and it is determined that the user's emotion is not abnormal, specifically including:
[0036] The natural language processing module is used to analyze the user's speech content, determine their emotional tendencies, and compare the user's emotional fluctuations in different dialogue stages and situations. If the emotional fluctuations are too large, it is determined that the user's emotional expression is abnormal.
[0037] The behavioral decision and feedback module analyzes the emotional characteristics of the user's voice, such as speaking speed, pitch, and intonation, to determine the user's emotional state. At the same time, the visual perception module analyzes the user's facial reactions to determine whether the user exhibits an emotional state consistent with the voice and language.
[0038] The speech recognition module uses speech recognition technology to identify different voice inputs and convert them into processable text information.
[0039] The natural language processing module is used to understand and parse the text entered by the user to extract valid information and understand the user's intentions, and to determine whether the user's emotions are abnormal. Specifically, it includes:
[0040] A1. Analyze the user's speech, text and voice to identify the user's emotional state, and use the natural language processing module to analyze the user's tone, vocabulary and emotional color to determine the abnormal fluctuation of the emotional value and determine the abnormal value of the voice emotion;
[0041] A2. Determine the expression emotion outliers by determining the changes of different facial data in the historical data and the results of the changes of different facial data, and based on the emotion recognition results of the changes of different facial data in the historical data;
[0042] A3. Determine the user's emotion abnormal value based on the determination of the speech emotion abnormal value and the determination of the expression emotion abnormal value, and determine the abnormal emotional state of the user based on the determination of the user's emotion abnormal value, and judge the number of interaction sentences whose emotion recognition results meet the specified requirements, specifically including: through dialogue framework analysis, identify sentences with inconsistent emotions, extreme emotions and sentences that cannot be correctly parsed, and record the number of all "unsatisfactory" sentences during the interaction process.
[0043] The dialogue management module is used to process user input, generate responses and maintain the context of the dialogue. The speech synthesis module is used to convert text into speech output, enabling the digital human to be presented to the user in a natural and fluent manner. The visual perception module captures the user's visual signals through the camera to identify their emotional and action non-verbal information. The behavioral decision and feedback module makes corresponding behavioral decisions based on the user's input, emotional state and target tasks.
[0044] An interactive method based on a digital human interactive processing system comprises the following specific steps:
[0045] S1. Perform contextual reinforcement learning based on the device for voice interaction through the voice recognition module and the visual perception module, introduce deep learning algorithms and establish an emotional memory library, obtain the voice weight data and different facial expression data for the digital human in real time, and determine the changes in the user's voice weight and facial expression based on the voice weight data and different facial expression data;
[0046] S2. Determine the historical interaction of the emotion data of the voice interaction user based on the voice weight and facial recognition results of the voice interaction user, and determine the degree of change of the emotion data of the voice interaction user according to the historical interaction, and determine the interactive voice data of the digital human and the facial voice coordination through the analysis results of the degree of change of the emotion data and the voice weight data of the voice interaction user;
[0047] S3, the dialogue management module determines and stores the emotion recognition results of different interaction times based on the voice weight data and different facial expression data of the user in the historical interaction times, and adjusts the interactive emotions of the digital human according to the changes in the emotion results of different interaction times in different time periods;
[0048] S4. Determine the adjustment result of the facial and voice emotions of the interactive digital person through the changes in the emotional results in different interaction times in different time periods and the voice severity and facial changes in the current interaction times.
[0049] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.
Claims
1. An interactive processing system based on digital human, characterized by: The interactive processing system includes a speech recognition module, a natural language processing module, an emotion recognition module, a dialogue management module, a speech synthesis module, a visual perception module, and a behavior decision and feedback module; The emotion recognition module is responsible for automatically generating soothing sentences based on the position of the voice interaction user, and obtaining the voice severity data and different facial expression data of the voice interaction user of the digital human in real time; The speech recognition module recognizes different speech inputs and converts them into processable text information through speech recognition technology; The natural language processing module is used to understand and parse the text input by the user to extract valid information and understand the user's intention; The dialog management module is used to process user input, generate responses and maintain the context of the dialog; The speech synthesis module is used to convert text into speech output, so that the digital human is presented to the user in a natural and fluent manner; The visual perception module captures the user's visual signals through a camera to identify the user's emotion and action non-verbal information; The behavior decision and feedback module makes corresponding behavior decisions according to the user's input, emotional state and target task.
2. The interactive method based on the interactive processing system of digital human according to claim 1, characterized in that: The specific steps include: S1. Performing contextual reinforcement learning based on the device for voice interaction through the voice recognition module and the visual perception module, introducing a deep learning algorithm and establishing an emotional memory library, acquiring the voice weight data and different facial expression data for the digital human in real time, and determining the changes in the voice weight and facial expression of the user according to the voice weight data and different facial expression data; S2, based on the voice weight and facial recognition results of the voice interaction user, determining the historical interaction of the emotion data of the voice interaction user, and determining the degree of change of the emotion data of the voice interaction user according to the historical interaction, and determining the interactive voice data of the digital person and the facial voice coordination through the analysis results of the degree of change of the emotion data and the voice weight data of the voice interaction user; S3, the dialogue management module determines and stores emotion recognition results of different interaction times based on the voice weight data and different facial expression data of the user in the historical interaction times, and adjusts the interactive emotions of the digital human according to the changes in the emotion results of different interaction times in different time periods; S4. Determine the adjustment result of the facial and voice emotions of the interactive digital person through the change of the emotion results in different interaction times in different time periods and the voice weight and facial change results in the current interaction times.
3. The interactive method based on the interactive processing system of digital human according to claim 2, characterized in that: The emotion recognition module includes a knowledge graph and an inference engine, and the facial expression data is determined based on the knowledge graph and the inference engine in combination with analysis and inference results of facial image analysis of voice interaction users in different time periods.
4. The interactive method based on the interactive processing system of digital human according to claim 3, characterized in that: Determining that the user's emotions are not abnormal, specifically includes: A1. Analyze the user's speech, text and voice to identify the user's emotional state, and use a natural language processing module to analyze the user's tone, vocabulary and emotional color to determine the abnormal fluctuation of the emotional value and determine the abnormal emotional value of the voice; A2. Determine the expression emotion outliers by determining the changes of different facial data in the historical data and the results of the changes of different facial data, and based on the emotion recognition results of the changes of different facial data in the historical data; A3. Determine the user's emotion abnormality value based on the determination of the voice emotion abnormality value and the determination of the expression emotion abnormality value, and determine the abnormal emotional state of the user according to the determination of the user's emotion abnormality value.
5. The interactive method based on the interactive processing system of digital human according to claim 4, characterized in that: Determining the abnormal emotion situation of the user according to the abnormal emotion value of the user specifically includes: when the abnormal emotion value of the user does not meet the requirement, determining that the emotion of the user is abnormal.
6. The interactive method based on the interactive processing system of digital human according to claim 5, characterized in that: Determining that the user's emotions are not abnormal, specifically includes: Analyze the user's speech content through the natural language processing module to determine his emotional tendency, compare the user's emotional fluctuations in different dialogue stages and situations, and if the emotional fluctuations are too large, determine that the user's emotional expression is abnormal; The emotional features in the user's voice, such as speaking speed, pitch and intonation fluctuations, are analyzed by the behavioral decision and feedback module to determine the user's emotional state. At the same time, the user's facial reaction is analyzed by the visual perception module to determine that the user exhibits an emotional state consistent with the voice and language.
7. The interactive method based on the interactive processing system of digital human according to claim 6, characterized in that: The indicator for judging the number of interactive sentences whose emotion recognition results meet the specified requirements includes: identifying sentences with inconsistent emotions, extreme emotions and sentences that cannot be correctly parsed through dialogue framework analysis, and recording the number of all sentences that "do not meet the requirements" during the interaction process.