AI digital human interaction system
By designing the collaborative work of multi-modal information acquisition, emotion recognition, dialogue management, action and expression generation, reinforcement learning and user feedback processing modules, the problem of insufficient natural and vivid interaction experience in the existing AI digital human system is solved, and a more natural and vivid interaction experience and real-time optimization of the system is achieved.
Patent Information
- Application Number
- CN202510137299.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-27
AI Technical Summary
The existing AI digital human system lacks the fusion processing of multimodal information in interaction, which leads to the lack of natural and vivid interaction experience, and ignores comprehensiveness and accuracy when processing user feedback, resulting in unclear direction of system improvement.
An AI digital human interaction system is designed, including a multimodal information acquisition module, an emotion recognition module, a dialogue management module, an action and expression generation module, a reinforcement learning module and a user feedback processing module. Through the coordinated work of these modules, the fusion processing and real-time feedback analysis of multimodal information is realized.
Through multimodal information acquisition and emotion recognition, the coherence and personalization of the conversation are enhanced, the learning module optimizes the interaction strategy, the user feedback processing module ensures real-time improvement, and the overall system interaction effect is more natural and vivid.
Smart Images

Figure CN120045069A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to an AI digital human interaction system. Background Art
[0002] With the rapid development of artificial intelligence technology, AI digital human interaction systems have gradually become a research hotspot. In the prior art, AI digital human systems usually interact based on single text or voice information, lacking the fusion processing of multi-modal information, resulting in an unnatural and vivid interaction experience. However, in dealing with users' feedback opinions, existing AI digital humans often ignore the comprehensiveness and accuracy of feedback data, leading to an unclear direction for system improvement and the inability to improve and optimize in a timely manner, resulting in a deteriorated interaction experience. Summary of the Invention
[0003] Therefore, the present invention provides an AI digital human interaction system to solve the problem that the interaction experience in the prior art needs to be further improved.
[0004] To achieve the above object, the present invention provides the following technical solutions:
[0005] An AI digital human interaction system, comprising:
[0006] A multi-modal information acquisition module, responsible for acquiring various types of information from different sources and transmitting the collected raw data to subsequent modules for processing;
[0007] An emotion recognition module, which preprocesses and extracts features from the raw data transmitted by the multi-modal information acquisition module, then analyzes the information based on the extracted features, recognizes the user's emotional state, and transmits the recognized emotional state information to the dialogue management module;
[0008] A dialogue management module, which receives the emotional state information transmitted by the emotion recognition module and other information input by the user feedback processing module; tracks the dialogue state, selects an appropriate response method according to the dialogue strategy, and generates corresponding system behaviors; transmits the system behavior instructions to the action and expression generation module;
[0009] An action and expression generation module: generates corresponding actions and expressions according to the system behavior instructions transmitted by the dialogue management module;
[0010] A reinforcement learning module, which receives the state transition information generated by the dialogue management module during the dialogue process; updates the behavior strategy of the dialogue management module according to the state transition information to improve the system performance and user satisfaction; transmits the updated behavior strategy to the dialogue management module;
[0011] The user feedback processing module collects and analyzes user feedback data; analyzes and processes the feedback data, extracts valuable information, and transmits the analysis results to the dialogue management module and the reinforcement learning module for guiding the improvement and optimization of the system.
[0012] Furthermore, other information input by the user includes text and speech.
[0013] Furthermore, the state transition information includes the current dialogue state, the actions performed, and the rewards obtained.
[0014] Furthermore, the multi-modal information acquisition module realizes its functions through voice data acquisition, facial expression data acquisition, and action data acquisition, as follows:
[0015] Voice data acquisition: Use a high-sensitivity microphone array to collect user voice signals in real time; then preprocess the collected voice signals; extract voice features from the preliminarily preprocessed voice signals.
[0016] Facial expression data acquisition: Use a high-definition camera to capture the user's facial image; then apply a face detection algorithm to locate the facial area and perform alignment processing.
[0017] Action data acquisition: Obtain the user's body posture and gesture data through a depth sensor or a pose estimation algorithm; then preprocess the action data; extract action features from the preprocessed action data.
[0018] Furthermore, the emotion recognition module improves the accuracy and robustness of emotion recognition by constructing a multi-modal fusion emotion recognition model; introduces physiological signal data and user feedback to realize real-time evaluation of emotion stability; and adopts a fuzzy logic algorithm to dynamically adjust the emotion recognition strategy according to real-time data, enhancing the adaptive ability of the system.
[0019] Furthermore, the dialogue management module improves the accuracy of intent recognition and the naturalness of dialogue generation through a natural language processing NLP model; adopts context management technology to achieve the coherence and personalization of the dialogue; and introduces a user portrait and preference learning mechanism to adjust the dialogue strategy according to user characteristics, improving user satisfaction.
[0020] Furthermore, the reinforcement learning module optimizes the interaction strategy by adopting a reinforcement learning algorithm, improving the decision-making ability and performance effect of the digital human; introduces a user feedback mechanism to adjust and optimize the strategy in real time through online learning or offline learning.
[0021] Furthermore, the processing flow of the reinforcement learning module is as follows:
[0022] a. Set the initial state and environmental state of the AI digital human.
[0023] b. The AI digital human perceives the user input and the current environmental state;
[0024] c. Based on the current state and strategy, the AI digital human selects a response action;
[0025] d. The AI digital human executes the selected response action;
[0026] e. The AI digital human receives the user's feedback and calculates the reward based on the feedback;
[0027] f. The AI digital human updates its strategy according to the reward to better select actions in the future;
[0028] g. Repeat the loop: The process returns to step a to continue the next round of interaction and learning process.
[0029] Furthermore: The action and expression generation module is integrated with a language generation module, and the dialogue management module can transfer the system behavior instructions to the language generation module.
[0030] The present invention has the following advantages: The present invention realizes comprehensive information collection through the multi-modal information acquisition module. The dialogue management module can enhance coherence and personalization. The reinforcement learning module optimizes the interaction strategy. The user feedback processing module ensures real-time feedback to the dialogue management module and the reinforcement learning module to achieve real-time improvement, making the overall system interaction effect more natural and vivid.
[0031] Other features and advantages of the present invention will be described in the subsequent specification, and some of them will be obvious from the specification or understood by implementing the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] To more intuitively illustrate the prior art and the present application, exemplary drawings are given below. It should be understood that the specific shapes and structures shown in the drawings generally should not be regarded as limiting conditions when implementing the present application. For example, those skilled in the art are capable of making routine adjustments or further optimizations to the addition / deletion / attribution division of certain units (components), specific shapes, positional relationships, connection methods, dimensional proportional relationships, etc. based on the technical concept disclosed in the present application and the exemplary drawings.
[0033] Figure 1 It is a block diagram of an AI digital human interaction system provided for an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. It should be understood that these embodiments are only for further explaining the present invention and cannot be construed as limiting the protection scope of the present invention. Technical engineers in this field can make some non-essential improvements and adjustments to the present invention according to the content of the above invention; based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0035] Please refer to Figure 1 , an AI digital human interaction system, including
[0036] Multimodal information acquisition module: responsible for obtaining various types of information from different sources, including text, images, audio, and video, etc.; providing a rich data basis for subsequent processing and analysis;
[0037] Emotion recognition module, used to analyze the user's emotional state, including happiness, sadness, anger, etc.; providing a more user-friendly interaction experience for the intelligent system;
[0038] Dialogue management module, used to control the dialogue process, including tracking the dialogue state, formulating dialogue strategies, and transferring the dialogue, etc.; ensuring that the intelligent system can interact with the user naturally and smoothly;
[0039] Action and expression generation module, generating corresponding actions and expressions according to the user's input and the system's requirements; providing a more vivid and vivid interaction method for the intelligent system; in addition, the action and expression generation module can also integrate a language generation module;
[0040] Reinforcement learning module, optimizing the behavior strategy of the intelligent system through continuous trial-and-error learning; improving the adaptability and decision-making ability of the intelligent system in complex environments;
[0041] User feedback processing module, collecting and analyzing user feedback data, evaluating the performance of the intelligent system; providing data support for the improvement and optimization of the system.
[0042] The multimodal information acquisition module mainly realizes its functions through voice data acquisition, facial expression data acquisition, and action data acquisition.
[0043] Voice data acquisition:
[0044] Real-time collect the user's voice signal using a high-sensitivity microphone array; then preprocess the collected voice signal, including noise suppression, gain control, sampling rate normalization, etc.; extract voice features from the preliminarily preprocessed voice signal, such as MFCC (Mel Frequency Cepstral Coefficients), pitch, energy, etc.
[0045] Facial expression data acquisition:
[0046] Capture the user's facial image using a high-definition camera; then apply a face detection algorithm to locate the facial area and perform alignment processing; use CNN (Convolutional Neural Network) to extract facial expression features, such as changes in the shape of the eyes and mouth.
[0047] Action data acquisition:
[0048] Obtain the user's body posture and gesture data through a depth sensor (such as Kinect) or a pose estimation algorithm; then preprocess the action data, such as denoising and smoothing; extract action features from the preprocessed action data, such as joint angles and movement trajectories.
[0049] In this embodiment, a microphone array, a high-definition camera, and a depth sensor can be integrated to achieve synchronous acquisition of multi-modal information; signal processing and image processing algorithms are used to improve the data quality and the accuracy of feature extraction; deep learning models (such as CNN) are used for feature extraction to enhance the robustness and generalization ability of the system.
[0050] The emotion recognition module improves the accuracy and robustness of emotion recognition by constructing a multi-modal fusion emotion recognition model; introduces physiological signal data and user feedback to achieve real-time evaluation of emotional stability; and uses a fuzzy logic algorithm to dynamically adjust the emotion recognition strategy according to real-time data to enhance the adaptive ability of the system.
[0051] Construction of comprehensive emotion feature vector:
[0052] Fuse voice features, facial expression features, and action features into a comprehensive emotion feature vector; use a multi-modal emotion recognition model (such as an LSTM-Attention network) to classify the feature vector to obtain a preliminary emotional state.
[0053] Emotional stability evaluation:
[0054] Collect emotion recognition feedback data of the user during the interaction process; combine physiological signal data (such as heart rate variability and skin conductivity) for emotional stability analysis; use a decision tree or a random forest algorithm to evaluate the stability of the emotional change trend.
[0055] Improvement of emotion recognition accuracy:
[0056] Analyze the feature weights at low levels (such as basic emotion features) and high levels (such as situation-related emotions); calculate the comprehensive emotion recognition accuracy coefficient through weighted average; use a fuzzy logic system to dynamically adjust the parameters of the emotion recognition model to improve recognition accuracy.
[0057] The dialogue management module improves the accuracy of intent recognition and the naturalness of dialogue generation through an NLP model; adopts context management techniques such as dialogue state tracking (DST) to achieve dialogue coherence and personalization; and introduces a user profile and preference learning mechanism to adjust the dialogue strategy according to user characteristics and improve user satisfaction.
[0058] Specifically:
[0059] Intent recognition and dialogue generation:
[0060] Use an NLP (Natural Language Processing) model (such as BERT, GPT) to perform intent recognition on user input; generate corresponding dialogue response content according to the recognized intent.
[0061] Dialogue context management:
[0062] Maintain the context information of the dialogue, including the user's historical input, the system's historical responses, etc.; use the context information to generate more coherent and personalized dialogue responses.
[0063] Personalized dialogue adjustment:
[0064] Dynamically adjust the dialogue style and content based on the user's emotional state and preferences; introduce an external knowledge base and a common sense reasoning engine to enrich the dialogue content and depth.
[0065] The action and expression generation module adopts action generation algorithms (such as motion graphs, action synthesis networks) and expression generation technologies (such as facial muscle models, expression animation libraries); and introduces synchronization mechanisms such as timeline alignment and event-driven synchronization to ensure the coordination of actions and expressions; continuously optimize the generation effect of actions and expressions through user feedback and reinforcement learning algorithms to improve the naturalness and realism of the interaction.
[0066] The function implementation of the action and expression generation module is mainly achieved through the following three steps:
[0067] Action generation: Generate corresponding action sequences according to the instructions output by the dialogue management module and the user's emotional state; use skeletal animation technology or motion capture technology to convert the action sequences into executable animation effects;
[0068] Expression generation: Generate corresponding expression animations based on the emotional state output by the emotion recognition module; use facial capture technology or 3D modeling technology to achieve delicate expression of expressions;
[0069] Synchronization of Actions and Expressions: Ensure the synchronization and coordination of actions and expressions to enhance the realism and naturalness of interactions.
[0070] The reinforcement learning module optimizes the interaction strategy by adopting reinforcement learning algorithms, improving the decision-making ability and performance of the digital human; introduce a user feedback mechanism to adjust and optimize the strategy in real time through online learning or offline learning.
[0071] The realization of the reinforcement module function is mainly based on the following aspects:
[0072] Policy Definition: Define the interaction strategies of the digital human in different scenarios, including dialogue strategies, action strategies, expression strategies, etc.;
[0073] Environment Simulation: Build a virtual interaction environment to simulate interaction scenarios in the real world; train the interaction strategy of the digital human in the simulation environment to improve its adaptability in complex scenarios;
[0074] Policy Optimization: Use reinforcement learning algorithms (such as Q-learning, DQN, PPO) to optimize the interaction strategy of the digital human; dynamically adjust the policy parameters according to user feedback and real-time evaluation results to achieve continuous optimization of the strategy.
[0075] The reinforcement learning module interacts through an agent with the environment (in this embodiment, the AI digital human interaction scenario), executes actions and receives feedback (rewards or punishments) from the environment to achieve learning.
[0076] The agent is in the environment, can change the state of the environment by executing actions, and obtain rewards or punishments from the environment. The goal of the agent is to learn a strategy, that is, to select the optimal action in different states to maximize the long-term cumulative reward.
[0077] State: A complete description of the environment at a certain moment, and the agent selects actions based on the current state.
[0078] Action: The set of behaviors that the agent can take, and each action changes the state of the environment.
[0079] The environment gives rewards or punishments according to the actions executed by the agent; rewards are the core driving force of reinforcement learning, and the goal of the agent is to maximize the cumulative reward.
[0080] Policy is the mapping from state to action of the agent, which determines which action the agent selects in a given state. The goal of reinforcement learning is to find an optimal policy.
[0081] The value function is used to evaluate the value of a state or a state-action pair.
[0082] Common value functions include the state value function V(s) and the action value function Q(s,a).
[0083] The state value function V(s): represents the expected cumulative reward obtained by following the current policy in state s.
[0084] The action value function Q(s,a): represents the expected cumulative reward obtained by following the current policy after taking action a in state s.
[0085] In this embodiment, Q-learning is taken as an example, and its update formula is as follows:
[0086] Q(s,a)←Q(s,a)+α[r+γmax_a'Q(s',a')-Q(s,a)]
[0087] Where: α is the learning rate, which controls the update step size; r is the immediate reward; γ is the discount factor, which determines the current value of future rewards; s' is the new state after executing action a; a' is the action that may be taken in the new state s'.
[0088] To further understand the working process of the reinforcement learning module, in this embodiment, the agent is defined as the AI digital human itself. The AI digital human can perceive user inputs (such as text, voice, images, etc.), execute corresponding response actions (such as generating response text, playing audio, adjusting expressions and actions, etc.), and learn how to provide a more natural, smooth and personalized interaction experience based on user feedback (such as satisfaction, interaction duration, etc.). The environment is defined as the interaction scenario where the AI digital human is located, including text chat rooms, voice call interfaces, virtual reality scenarios, etc. The state of the environment can include user input information, the current state of the digital human (such as expressions, actions, etc.), historical interaction records, etc. The agent (AI digital human) selects actions according to the environmental state and receives feedback from the environment.
[0089] The current interaction state of the AI digital human includes the information input by the user (such as text, voice, expressions, etc.), the internal state of the digital human (such as emotions, memories, etc.), and the environmental context (such as time, location, topic, etc.).
[0090] Actions: Response actions that the AI digital human can execute in the current state, such as generating text responses, playing voice responses, adjusting expressions and actions, etc. The selection of actions should be based on the principle of maximizing cumulative rewards.
[0091] The reward function is used to evaluate whether the response actions of the AI digital human are good; the reward can be determined according to user feedback, such as the user's satisfaction score, interaction duration, whether to continue the interaction, etc. For example, if the user gives a high score or has a long interaction, the reward is positive; if the user expresses dissatisfaction or interrupts the interaction, the reward is negative.
[0092] Its example process is as follows:
[0093] a. Initialization: Set the initial state of the AI digital human and the environmental state;
[0094] b. Perception state: The AI digital human perceives the user input and the current environmental state;
[0095] c. Select action: According to the current state and strategy, the AI digital human selects a response action;
[0096] d. Execute action: The AI digital human executes the selected response action, such as generating a text response or adjusting the expression, etc.;
[0097] e. Receive feedback: The AI digital human receives the user's feedback (such as satisfaction score, interaction duration, etc.), and calculates the reward based on the feedback;
[0098] f. Update strategy: The AI digital human updates its strategy according to the reward to better select actions in the future;
[0099] g. Repeat loop: The process returns to the perception state step to continue the next round of interaction and learning process.
[0100] The user feedback processing module ensures the comprehensiveness and accuracy of feedback data through various feedback collection methods; uses analysis algorithms to deeply mine and analyze the feedback data, extracts valuable information; and establishes a feedback application mechanism to convert the feedback results into actual improvement measures through a closed-loop control system to continuously optimize the interaction effect of the digital human.
[0101] The implementation of the functions of the user feedback processing module is mainly achieved through the following three aspects:
[0102] Feedback collection: Collect real-time feedback data of users through various methods such as voice, text, expression or scoring; preprocess and extract features from the feedback data to provide a basis for subsequent analysis.
[0103] Feedback analysis: Analyze the feedback data using NLP or machine learning algorithms (such as sentiment analysis, topic model); identify key information such as user satisfaction, emotional changes, improvement suggestions, etc.
[0104] Feedback application: Apply the feedback analysis results to the optimization of the interaction strategy and personalized adjustment of the digital human; regularly push the improved interaction experience to users to improve user satisfaction and loyalty.
[0105] The interaction system of the present invention can mainly be divided into four stages: the information acquisition stage, the information analysis and processing stage, the information output stage, and the system optimization and learning stage.
[0106] 1. Information acquisition stage
[0107] The multimodal information acquisition module is responsible for acquiring various types of information from different sources, such as text, images, audio, and video, etc., and passing the collected raw data to subsequent modules for processing.
[0108] 2. Information Analysis and Processing Phase
[0109] Emotion Recognition Module: Analyze the information passed by the multimodal information acquisition module to identify the user's emotional state.
[0110] Receive the raw data passed by the multimodal information acquisition module, preprocess and extract features from the raw data, then input it into the emotion classification model for emotion recognition, and pass the recognized emotional state information to the dialogue management module.
[0111] Dialogue Management Module: Control the dialogue flow and formulate dialogue strategies based on the user's emotional state and other information.
[0112] Receive the emotional state information passed by the emotion recognition module, as well as other information input by the user feedback processing module (such as text, speech, etc.). Track the dialogue state, select an appropriate response method according to the dialogue strategy, and generate corresponding system behaviors; pass the system behavior instructions to the action and expression generation module or the language generation module (if the system needs to generate a natural language response).
[0113] 3. Information Output Phase
[0114] Action and Expression Generation Module: Generate corresponding actions and expressions according to the system behavior instructions passed by the dialogue management module.
[0115] Receive the system behavior instructions passed by the dialogue management module, generate corresponding action and expression data according to the instructions (such as motion data, facial expression parameters, etc.); pass the generated action and expression data to the execution mechanism (such as a robot, virtual character, etc.) for display.
[0116] 4. System Optimization and Learning Phase
[0117] Reinforcement Learning Module: Optimize the behavior strategy of the dialogue management module through continuous trial-and-error learning.
[0118] Receive the state transition information generated by the dialogue management module during the dialogue (such as the current dialogue state, executed actions, obtained rewards, etc.); update the behavior strategy of the dialogue management module according to the state transition information to improve the system performance and user satisfaction; pass the updated behavior strategy to the dialogue management module.
[0119] User Feedback Processing Module: Collect and analyze the user's feedback data to provide data support for system improvement and optimization.
[0120] Receive feedback data provided by users through various channels (such as ratings, comments, usage duration, etc.); analyze and process the feedback data to extract valuable information, such as user satisfaction, system performance bottlenecks, etc. Transmit the analysis results to the dialogue management module and the reinforcement learning module for guiding the improvement and optimization of the system.
[0121] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An AI digital human interaction system, characterized in that: include: The multimodal information acquisition module is responsible for acquiring various types of information from different sources and passing the collected raw data to subsequent modules for processing; The emotion recognition module preprocesses and extracts features from the raw data transmitted by the multimodal information acquisition module, then identifies the user's emotional state based on the extracted feature analysis information, and transmits the identified emotional state information to the dialogue management module; The dialogue management module receives the emotional state information transmitted by the emotion recognition module and other information input by the user feedback processing module; Track the state of the conversation, select appropriate responses based on the conversation strategy, and generate corresponding system behaviors; Pass the system behavior instructions to the action and expression generation module; Action and expression generation module: generates corresponding actions and expressions according to the system behavior instructions transmitted by the dialogue management module; A reinforcement learning module receives state transition information generated by the dialogue management module during the dialogue process; Update the behavior strategy of the dialogue management module according to the state transition information to improve the system performance and user satisfaction; The updated behavior strategy is passed to the dialogue management module; User feedback processing module, collects and analyzes user feedback data; Analyze and process the feedback data, extract valuable information, and pass the analysis results to the dialogue management module and reinforcement learning module to guide the improvement and optimization of the system.
2. The AI digital human interaction system according to claim 1, characterized in that: Other information input by the user includes text and voice.
3. The AI digital human interaction system according to claim 1, characterized in that: State transfer information includes the current dialogue state, the actions performed, and the rewards obtained.
4. The AI digital human interaction system according to claim 1, characterized in that: The multimodal information acquisition module realizes its functions by acquiring voice data, facial expression data, and action data, as follows: Voice data acquisition: Use a high-sensitivity microphone array to collect user voice signals in real time; then pre-process the collected voice signals; extract voice features from the pre-processed voice signals; Facial expression data acquisition: Use a high-definition camera to capture the user's facial image; then use the face detection algorithm to locate the facial area and perform alignment processing; Motion data acquisition: obtain user body posture and gesture data through depth sensors or posture estimation algorithms; then preprocess the motion data; extract motion features from the preprocessed motion data.
5. The AI digital human interaction system according to claim 1, characterized in that: The emotion recognition module improves the accuracy and robustness of emotion recognition by constructing a multimodal fusion emotion recognition model; introduces physiological signal data and user feedback to achieve real-time evaluation of emotional stability; and uses fuzzy logic algorithms to dynamically adjust emotion recognition strategies based on real-time data to enhance the system's adaptive capabilities.
6. The AI digital human interaction system according to claim 1, characterized in that: The dialogue management module uses the natural language processing (NLP) model to improve the accuracy of intent recognition and the naturalness of dialogue generation. It adopts context management technology to achieve dialogue consistency and personalization. It also introduces user profiling and preference learning mechanisms to adjust dialogue strategies based on user characteristics and improve user satisfaction.
7. The AI digital human interaction system according to claim 1, characterized in that: The reinforcement learning module optimizes the interaction strategy by adopting reinforcement learning algorithms to improve the decision-making ability and performance of digital humans; it introduces a user feedback mechanism to adjust and optimize the strategy in real time through online or offline learning.
8. The AI digital human interaction system according to claim 7, characterized in that: The processing flow of the reinforcement learning module is as follows: a. Set the initial state and environment state of the AI digital human; b. AI digital humans perceive user input and current environment status; c. Based on the current status and strategy, the AI digital human selects a response action; d. The AI digital human performs the selected response action; e. The AI digital human receives user feedback and calculates rewards based on the feedback; f. The AI updates its strategy based on the reward to better choose actions in the future; g. Repeat the loop: The process returns to step a and continues the next round of interaction and learning.
9. An AI digital human interaction system according to any one of claims 1 to 8, characterized in that: The action and expression generation module is integrated with the language generation module, and the dialogue management module can pass the system behavior instructions to the language generation module.
Citation Information
Cited By
Adaptive interaction method based on multi-modal emotion calculation and robot
CN120631175A
Digital human dynamic interaction method and system based on deep learning
CN120688535A
A digital human dynamic interaction method and system based on deep learning
CN120688535B
Digital human interaction method and system based on large model
CN121118962A
Live broadcast method for AI digital human multi-modal cloning and real-time interaction
CN121711501A