AI virtual real person interaction system and method based on emotion recognition
By introducing technologies such as multimodal emotion recognition and generative adversarial networks into the virtual character system, the problems of low emotional recognition accuracy and stiff response in the existing system are solved, high-precision emotion recognition and personalized interaction are achieved, and the system's computing efficiency and user experience are improved.
Patent Information
- Application Number
- CN202510384904.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
AI Technical Summary
There are problems in the existing virtual character system with low emotional recognition accuracy, stiff response, low computing efficiency, and lack of personalized interaction.
The AI virtual real-person interaction system based on emotion recognition is adopted, including interaction recognition module, emotion generation module, sound generation module, action generation module, quantum computing optimization module and AI training module. Through multi-modal emotion recognition, generative adversarial network, neural network speech synthesis, variational autoencoder, particle system simulation and quantum computing optimization technologies, high-precision emotion recognition and personalized interaction are achieved.
It improves the accuracy and nature of emotional recognition, enhances the personalized interaction ability of virtual characters, improves the system's computing efficiency and response speed, and provides an interactive experience that is closer to humans.
Smart Images

Figure CN120215715A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to an AI virtual human interaction system and method based on emotion recognition. Background Art
[0002] With the rapid development of artificial intelligence (AI) and virtual reality technologies, virtual characters have gradually been applied to various fields such as entertainment, education, and customer service. Virtual characters in the prior art, especially virtual assistants that interact with users, usually rely on fixed rules and programs to respond. Although these virtual characters can perform basic tasks, they have obvious deficiencies in emotional expression and the naturalness of user interaction, and cannot provide a more human-like interaction experience. The deficiencies of the prior art are mainly reflected in the following aspects.
[0003] Existing virtual characters mainly rely on speech or text analysis for emotion recognition. Although this method can capture some emotional signals, the accuracy is usually low. In particular, emotional expressions are often complex and diverse, and different modalities of data such as speech, facial expressions, and text often need to be comprehensively analyzed to accurately judge the user's emotional state. The single emotion analysis method in the prior art lacks a comprehensive understanding of emotional expression, resulting in low emotion recognition accuracy and inability to accurately recognize the complex emotional state of users.
[0004] Many virtual characters rely on preset reaction patterns or simple rule bases to perform interactions. Although this method can handle basic tasks, the reactions of virtual characters often lack natural fluency. Especially when facing the emotional changes of users, they often appear rigid and single. For example, when a user expresses pleasure or anger, the reaction of the virtual character may be fixed, lacking personalization and dynamic adjustment that conform to the user's emotional state. This lack of flexibility and authenticity in emotional expression cannot meet the user's needs for emotional interaction with virtual characters, resulting in a poor interaction experience.
[0005] Most existing virtual character systems interact based on static models. Even after multiple interactions, the reactions of virtual characters usually do not adjust according to the changes of users. This static reaction method makes it difficult for virtual characters to provide personalized experiences when facing different users and situations. Especially when facing users with long-term interactions, virtual characters fail to continuously optimize their performance according to the behavior and emotional changes of users, resulting in poor adaptability of the system and inability to provide continuous and personalized interaction experiences for users. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the present invention provides an AI virtual human interaction system and method based on emotion recognition, which solves the problems of low emotion recognition accuracy, rigid reaction, low calculation efficiency, and lack of personalized interaction in existing virtual character systems.
[0007] To achieve the above object, the present invention is realized through the following technical solutions: an AI virtual human interaction system based on emotion recognition, including; An interaction recognition module, configured to recognize the emotional state input by the user, obtain interaction data, and output the emotion classification and its intensity; An emotion generation module, connected to the emotion recognition module, and generating the emotional expression of the virtual character according to the emotion classification and intensity; A voice generation module, connected to the emotion generation module, and generating the voice performance of the virtual character according to the emotion data output by the emotion generation module; A motion generation module, connected to the emotion generation module, and generating the motion of the virtual character according to the emotional expression of the virtual character; A quantum computing optimization module, connected to the emotion generation module and the motion generation module, and optimizing the computing efficiency in the emotion generation and motion generation processes through quantum computing; An AI training module, connected to the interaction recognition module, and performing training based on the interaction data.
[0008] Preferably, the interaction recognition module includes: A voice recognition unit, configured to convert the voice input of the user into text data, extract voice features and perform emotion analysis; A facial expression analysis unit, capturing the facial expression of the user through image recognition technology and analyzing its emotional state; A text analysis unit, configured to analyze the text input of the user and recognize the emotional content therein.
[0009] Preferably, the interaction data includes the user's voice, facial expression, text input, and the feedback data of the virtual character.
[0010] Preferably, the emotion generation module uses a generative adversarial network for emotion generation. The generative adversarial network includes a generator and a discriminator. The generator generates the emotional expression of the virtual character according to the input emotion data, and the discriminator is used to evaluate whether the generated emotional expression conforms to the real emotional expression.
[0011] Preferably, the voice generation module adopts neural network speech synthesis technology to generate the voice output of the virtual character, and the voice performance is adjusted by tone, speech rate, and tone parameters to adapt to the emotional expression of the virtual character.
[0012] Preferably, the motion generation module includes: A variational autoencoder, configured to map the emotional expression to the latent space and generate the motion of the virtual character according to the latent space; A particle system simulation unit, generating the natural motion of the virtual character based on the latent space, including facial expressions and limb movements.
[0013] Preferably, the quantum computing optimization module adopts a quantum optimization algorithm to accelerate the calculations in the emotion generation and action generation modules.
[0014] Preferably, the AI training module adopts deep learning and reinforcement learning technologies for model training and optimization, and self-optimizes through continuous interaction data of users to dynamically adjust the voice, actions, and emotional expressions of virtual characters.
[0015] Preferably, the deep learning is used for the training of emotion recognition, voice generation, and action generation, and the reinforcement learning dynamically adjusts the reactions and behaviors of virtual characters according to the interaction data of users.
[0016] An AI virtual human interaction method based on emotion recognition includes the following steps: Obtain the emotional data of users through the interaction recognition module, including voice input, facial expression data, and text input. This data reflects the current emotional state of users and provides a basis for subsequent emotion generation and behavioral responses. Analyze the obtained emotional data of users through the emotion recognition module to identify the emotional state of users and determine the emotion classification and intensity. This step is the core of the entire interaction process, and accurately identifying the emotional state is the premise for generating real virtual character reactions. Through the emotion generation module, generate the emotional expressions of virtual characters according to the identified emotion classification and intensity, and use generative adversarial network technology to ensure that the generated emotional expressions are realistic and conform to the logic of human emotions. According to the emotional expressions of virtual characters, the action generation module generates corresponding actions of virtual characters, including facial expression changes and body movements, to ensure that the performance of virtual characters is consistent with their emotional states. According to the emotional expressions of virtual characters, the voice generation module generates voice outputs that conform to the emotional states, including the adjustment of pitch, speech rate, and tone characteristics, to ensure that the voice feedback of virtual characters is consistent with their emotional states. Through the quantum computing optimization module, optimize the computational efficiency required in the emotion generation and action generation processes. Quantum computing technology can accelerate the calculations in the emotion generation and action generation processes, thereby improving the response speed and real-time interaction ability of the system. Through the AI training module, continuously optimize the performance of the emotion recognition, emotion generation, action generation, and voice generation modules based on the interaction data between users and virtual characters. This module adopts deep learning and reinforcement learning technologies to enable virtual characters to continuously adjust their reactions according to the interactions of different users to provide a more personalized interaction experience.
[0017] The present invention provides an AI virtual human interaction system and method based on emotion recognition. It has the following beneficial effects: 1. Through multi-modal emotion recognition technology, the present invention combines voice, facial expressions, and text data, achieving a higher-precision emotion recognition effect. Compared with the existing emotion analysis methods that solely rely on voice or text, this multi-dimensional fusion method significantly improves the accuracy of emotion classification and effectively solves the problem of incomplete emotion recognition from a single data source.
[0018] 2. Through Generative Adversarial Network (GAN) technology, the present invention ensures that the emotional expressions of virtual characters are natural and realistic. Different from virtual characters in the prior art that rely on fixed response patterns, the generative adversarial network enables virtual characters to dynamically adjust according to the user's emotional state, avoiding the shortcoming of rigid performance and making each interaction more personalized and flexible.
[0019] 3. By adopting quantum computing optimization technology, the present invention significantly improves the computing efficiency during the process of emotion generation and action generation. Compared with traditional computing methods, quantum computing can accelerate complex calculations through parallel processing, especially greatly improving the processing speed and response time on large-scale data sets, effectively solving the latency problem in real-time interaction in the prior art.
[0020] 4. Through the AI training module, the present invention continuously optimizes the responses of virtual characters through deep learning and reinforcement learning. Different from virtual characters with static responses in the prior art, the AI training module can dynamically adjust the behaviors and emotional expressions of virtual characters according to user interaction data, providing a continuously personalized interaction experience and solving the problem of the lack of long-term adaptability in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is the system framework diagram of the present invention; Figure 2 is the schematic diagram of the interaction recognition module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] Please refer to the attached Figure 1 - attached Figure 2 , the embodiment of the present invention provides an AI virtual human interaction system based on emotion recognition, including; An interaction recognition module, which is used to recognize the emotional state input by the user, obtain interaction data, and output the emotion classification and its intensity; Specifically, in this embodiment, the interaction recognition module is responsible for identifying the user's emotional state and generating interaction data, providing effective input for subsequent modules such as emotion generation, speech generation, and action generation. The interaction recognition module integrates multiple technical units and combines the analysis of voice, facial expressions, and text input to accurately capture the user's emotions. Through this module, the system can perceive and respond to the user's emotional fluctuations in real time during the interaction, providing a basis for the virtual character to generate rich and natural emotional responses.
[0024] In this embodiment, the interaction recognition module includes a speech recognition unit, a facial expression analysis unit, and a text analysis unit. These units work in parallel to comprehensively analyze the multi-modal input data from the user, so as to more accurately understand and identify the user's emotional state. These emotional data will then be passed to the emotion generation module to help generate the virtual character's performance that matches the user's emotional state.
[0025] In this embodiment, the speech recognition unit is used to receive and convert the user's speech input. The user's speech is not only converted into text data, but this unit also extracts the emotional features in the speech. These emotional features include, but are not limited to, intonation, speech rate, pauses, etc.
[0026] Specifically, when processing the speech signal, first, the features in the audio signal, such as pitch, volume, and timbre, are extracted through speech signal processing algorithms. Then, an emotion analysis model is used to analyze these features to determine whether the speech contains specific emotional signals. For example, in a state of anger, the pitch of the speech is usually higher, the speech rate is faster, and the speech rhythm is rapid; while in a state of sadness, the intonation is lower, the speech rate is slower, and there are more pauses. Through these features, the speech recognition unit can infer the user's emotional state, providing accurate data support for subsequent emotion expression generation.
[0027] The facial expression analysis unit captures the user's facial expressions and analyzes the emotional information therein through image recognition technology. This unit usually uses a deep convolutional neural network (CNN) model to process the input facial images. Specifically, first, the user's facial images are obtained through a camera or other image acquisition devices, and then the facial feature points (such as key parts like eyes, mouth, eyebrows, etc.) are extracted using face detection algorithms.
[0028] Based on the relative movements of these key parts and the changes in facial expressions, the facial expression analysis unit can judge the user's emotional state. For example, raised eyebrows and upturned corners of the mouth may indicate happiness or pleasure, while lowered eyebrows and downturned corners of the mouth may indicate sadness or anger. Combining the changes in facial expressions, the system can further infer the user's emotional state, such as joy, anger, surprise, etc.
[0029] The text analysis unit is responsible for sentiment analysis of the text input by the user. This process usually uses natural language processing (NLP) technology to convert the user's text into sentiment data that can be understood by the machine. The core of text analysis is to analyze the emotional color of the sentence input by the user to determine its emotional tendency.
[0030] In this embodiment, the text analysis unit processes the text through the sentiment dictionary and sentiment classification model. The sentiment dictionary lists a large number of words with emotional colors (such as "happy", "angry", "sad", etc.), and each word corresponds to a specific emotion category. On this basis, the system can also perform semantic analysis based on the context to further determine the specific classification of the user's emotions. Using sentiment analysis techniques such as sentiment dictionary matching, sentiment polarity analysis and deep learning models (such as long short-term memory network LSTM), this module can more accurately extract sentiment information from the text input by the user.
[0031] The various sub-units of the interaction recognition module work together to improve the accuracy of emotion recognition through the fusion of multimodal data. The speech recognition unit, facial expression analysis unit, and text analysis unit obtain emotion information from different input sources (speech, image, text) and summarize this information into a comprehensive emotion classification and intensity data.
[0032] For example, when a user expresses anger through voice and facial expressions, the voice recognition unit will detect a fast speaking speed and a high tone, while the facial expression analysis unit will identify facial expressions of frowning eyebrows and drooping mouth corners. The text analysis unit may recognize that there are angry words in the user's language. Combining this information, the system can accurately determine that the user's emotional state is "angry" and give the corresponding emotional intensity.
[0033] The interactive recognition module can not only accurately analyze the user's voice, facial expressions and text input, but also judge the user's emotional state based on this information. This module uses advanced speech processing, image recognition and natural language processing technologies to effectively improve the accuracy of emotion recognition and the responsiveness of the system.
[0034] An emotion generation module, which is connected to the emotion recognition module and generates the emotion expression of the virtual character according to the emotion classification and intensity; Specifically, the main task of the emotion generation module in this embodiment is to generate the emotional expression of the virtual character based on the emotion classification and intensity data provided by the interaction recognition module. This module uses the generative adversarial network (GAN) technology to ensure that the emotional expression of the virtual character is natural and real, and matches the emotional state of the user. The emotion generation module is a key module for the emotional interaction between the virtual character and the user, and its accuracy directly affects the interactive experience of the entire system.
[0035] In the embodiment, the emotion generation module includes two main components: a generator and a discriminator. The generator generates the emotional expressions of the virtual character based on the emotional data input by the user (including emotion classification and emotion intensity), while the discriminator is responsible for evaluating whether the generated emotional expressions conform to the real emotional expressions. The generator and the discriminator gradually improve the authenticity and diversity of the virtual character's emotional expressions through adversarial training.
[0036] In this embodiment, the generator is the core component of the emotion generation module. The generator receives the emotion classification and intensity data output by the interaction recognition module as input and generates the emotional expressions of the virtual character based on this data. Specifically, the generator converts the emotion classification and intensity into specific forms of emotional expressions, such as facial expressions, intonation, and speech rate.
[0037] The working principle of the generator can be expressed as; ; Where: represents the input emotion classification data; represents the emotion intensity; represents the generator is the output emotional expression of the virtual character.
[0038] Emotion classification data and emotion intensity are processed through a deep neural network (such as a multi-layer perceptron MLP or a convolutional neural network CNN), and finally the emotional expressions of the virtual character are output.
[0039] The generator first extracts features from the input data through a multi-layer neural network, converting the emotional data into a representation in the latent space. Then, these latent representations are used to generate corresponding emotional expressions. For example, in a happy emotional state, the generator may output a smiling facial expression and a fast, cheerful intonation; while in a sad emotional state, the generator may generate a drooping facial expression and a slower speech rate.
[0040] The role of the discriminator is to evaluate whether the emotional expressions output by the generator are real. The discriminator helps the generator adjust its output by comparing with the real emotional expressions, so that the emotional expressions of the virtual character are more in line with the real emotional expressions of humans. The output of the discriminator is a probability value indicating whether the generated emotional expressions are consistent with the real emotional expressions.
[0041] Specifically, the output of the discriminator can be expressed as: ; Wherein: represents the generated emotional expression; is the weight of the classifier; is the bias; is an activation function, often used to map the input value between 0 and 1; is the output of the classifier.
[0042] The goal of the discriminator is to maximize the probability that its output is 1, that is, to judge that the generated emotional expression is "real". Conversely, if the output is 0, it means that the generated emotional expression does not conform to the real emotional expression.
[0043] Through adversarial training, the generator and the discriminator play against each other. The generator gradually improves the quality of the emotional expressions it generates, making it difficult for the discriminator to distinguish between the generated emotional expressions and the real emotional expressions. This training method ensures that the generated emotional expressions have higher realism and diversity.
[0044] The generative adversarial network (GAN) is an unsupervised learning method. When applying this technology in the emotional generation module, the generator and the discriminator are jointly trained. The generator continuously generates emotional expressions, and the discriminator continuously evaluates the authenticity of these expressions. During the training process, the generator tries to generate more and more realistic emotional expressions, while the discriminator tries to improve its own ability by accurately distinguishing between the generated expressions and the real expressions.
[0045] In the early stage of training, the emotional expressions output by the generator may be quite different from the real emotional expressions. However, as the training progresses, the generator will gradually adjust its generation strategy based on the feedback from the discriminator. Eventually, the generator can generate emotional expressions with high realism and diversity.
[0046] In this embodiment, the emotional generation module can generate natural, real and diverse virtual character emotional expressions through the generative adversarial network (GAN) technology. The generator generates the emotional expressions of the virtual character according to the input emotional classification and intensity data, while the discriminator continuously optimizes the output of the generator by comparing with the real emotional expressions. The design of this module ensures that the virtual character can make real and personalized responses according to the user's emotional state, providing a more natural interaction experience for the user.
[0047] The voice generation module, which is connected to the emotional generation module, generates the voice expression of the virtual character according to the emotional data output by the emotional generation module; Specifically, in this embodiment, the voice generation module is responsible for generating a voice performance that conforms to the emotional state according to the virtual character emotional data output by the emotion generation module. The voice generation module adjusts the characteristics of the voice, such as pitch, speaking speed, and tone, so that the voice of the virtual character can naturally and realistically reflect its emotional state. The accuracy of this module directly affects the realism of the virtual character's emotional expression and the user's interaction experience.
[0048] In this embodiment, the voice generation module uses neural network speech synthesis technology and combines the emotional data (such as emotion type and intensity) output by the emotion generation module to adjust the various parameters of the voice. Specifically, the voice generation module includes a speech synthesis network, an emotion adjustment unit, and a voice output unit. Through the coordinated work of these units, the generated voice will be highly consistent with the emotional state of the virtual character.
[0049] In this embodiment, the speech synthesis network is the core of the voice generation module. It receives the emotional data output by the emotion generation module and generates a voice that conforms to the emotional state based on this data. This process first converts the emotional data into a set of characteristic parameters, including information such as pitch, speaking speed, and tone.
[0050] Specifically, the speech synthesis network uses advanced deep learning technologies such as recurrent neural network (RNN), long short-term memory network (LSTM), or generative adversarial network (GAN) to process the emotional data and generate a voice waveform that matches the emotion. The voice generation process can be expressed as: ; Where: is the output voice signal; is the input emotional feature data; are the parameters of the speech synthesis network; is the synthesis function.
[0051] This network adjusts the voice parameters such as pitch, speaking speed, and tone according to the user's emotional state to ensure that the generated voice expression conforms to the user's emotional input.
[0052] The speech synthesis network not only generates natural and fluent speech but also can adjust the various parameters of the voice according to the characteristics of different emotions. For example, in a pleasant emotion, the generated voice may have a higher pitch, a faster speaking speed, and a bright tone; while in a sad emotion, the pitch of the voice is lower, the speaking speed is slower, and the tone is more dull.
[0053] The voice output unit is responsible for converting the adjusted voice features into actual voice signals and outputting them to the user. This unit usually uses neural network speech synthesis technologies (such as WaveNet, Tacotron, etc.) to convert the emotion-adjusted voice features into continuous audio waveforms.
[0054] In this embodiment, the voice output unit maps the voice features to the time domain through a deep generative network (such as WaveNet) to generate real and natural voice waveforms. This process can be expressed as: ; Where: is the finally output voice waveform; is the emotion-adjusted voice feature data is the voice generation function.
[0055] Through this generation process, the system can accurately generate voices in different emotional states, ensuring that the voice performance of the virtual character matches the emotional state.
[0056] The voice output unit can also be adjusted according to the user's feedback to ensure that the generated voice conforms to the user's expectations and interaction experience as much as possible. For example, the system can adjust parameters such as the speed, intonation, or tone of the voice output in real time according to the input of the speech recognition module, thereby improving the naturalness and fluency of the interaction.
[0057] The voice generation module can generate natural, real, and diverse voice performances according to the emotion data output by the emotion generation module. Through the collaborative work of the speech synthesis network, the emotion adjustment unit, and the voice output unit, the system can adjust features such as the pitch, speed, and tone of the voice according to the user's emotional state, ensuring that the voice performance of the virtual character is highly consistent with its emotional state. The close cooperation of this module with the interaction recognition module and the emotion generation module enables the virtual character to show more natural and flexible voice responses during the interaction process, improving the overall realism and user experience of the interaction.
[0058] The action generation module, which is connected to the emotion generation module, generates the actions of the virtual character according to the emotional expression of the virtual character; Specifically, the main task of the action generation module in this embodiment is to generate limb actions that match the emotional state of the virtual character. This module generates action performances with high naturalness and consistency by analyzing the emotion data from the emotion generation module. The accuracy and reaction speed of the action generation module directly determine the realism and interactivity of the virtual character during the interaction process.
[0059] In this embodiment, the action generation module generates the limb actions of the virtual character through an action generation network (such as a convolutional neural network CNN or a recurrent neural network RNN). The action generation network receives the emotion data from the emotion generation module and generates actions consistent with the emotional state based on this data. Specifically, the emotion classification and intensity data output by the emotion generation module are used as inputs, and the network generates the limb actions of the virtual character according to these data. Every detail of the actions (such as the swing of the arms, the change of facial expressions, the adjustment of body postures, etc.) will be affected by the emotion data.
[0060] The action generation process can be represented by the following formula: ; Where: represents the action performance of the virtual character at time ; is the emotion classification data; is the emotion intensity data; are the parameters of the action generation network is the action generation function.
[0061] Through the deep neural network, the system generates a matching action output according to the input emotion data.
[0062] The action generation network can capture the temporal features of the actions through methods such as long short-term memory network (LSTM) or self-attention mechanism, ensuring that the generated actions not only conform to the emotional expression but also have good temporal consistency and natural transitions. In this way, the system can generate more delicate and smooth limb actions, enhancing the interaction experience between the virtual character and the user.
[0063] In this embodiment, the emotion-driven action adjustment mechanism ensures that the limb actions of the virtual character are highly consistent with its emotional state. The emotion generation module affects the output of the action generation network through emotion classification and emotion intensity data. For example, when the user shows happy or excited emotions, the actions of the virtual character may be fast, jumping, or waving; while when the user shows sad or worried emotions, the actions may be slow, head-down, shoulders-drooping, etc.
[0064] The emotion-driven adjustment mechanism synchronizes the actions of the virtual character with the emotional state by adjusting the action feature parameters generated by the network. For example, a happy emotional state may make the actions of the virtual character more tense and energetic, resulting in fast limb reactions; while in a tense or angry emotion, the actions of the virtual character may be more oppressive, and the body language tends to be tense or rigid.
[0065] The mathematical model of this mechanism can be expressed as: ; Where: represents the adjustment amount of behavioral characteristics; is a function representing the relationship between emotional characteristics and behaviors, usually determined by analyzing emotional data; is the emotional state at time ; is the behavioral characteristic at time ;
[0066] By analyzing the emotional data, the system can automatically adjust the generated action characteristics, so as to accurately reflect the user's emotional state in the body language of the virtual character.
[0067] To further improve the interaction experience of the virtual character, the action generation module further includes a real-time adjustment and feedback mechanism. In some embodiments, the system can make fine adjustments to the actions of the virtual character according to the user's real-time feedback. Specifically, the system obtains the user's real-time emotional changes through the interaction recognition module and adjusts the action performance of the virtual character through the feedback mechanism.
[0068] For example, if the user shows emotions of surprise or confusion during the interaction, the system can immediately adjust the actions of the virtual character to make them more in line with the user's emotional state. The action generation network will adjust the generated action characteristics according to the new emotional data to ensure that the body actions of the virtual character can respond to the user's emotional changes in real time.
[0069] This feedback mechanism can be expressed by the following formula: ; Where: represents the action characteristics adjusted according to the user feedback; is the feedback function; and are the emotional classification and intensity of the user; is the action characteristic generated last time.
[0070] Through the feedback mechanism, the system can dynamically adjust the actions of the virtual character so that its action performance is synchronized with the user's emotional changes.
[0071] The action generation module can generate natural, realistic, and diverse body movements based on the emotion data output by the emotion generation module. The action generation network generates actions highly consistent with the emotional state of the virtual character according to the input emotion data, ensuring that the body language of the virtual character can accurately reflect the emotional state. In addition, the system also uses methods such as an emotion-driven adjustment mechanism and a real-time feedback mechanism to ensure the smoothness and naturalness of the actions, improving the user's immersion and interaction experience.
[0072] A quantum computing optimization module, which is connected to the emotion generation module and the action generation module, and optimizes the computational efficiency in the emotion generation and action generation processes through quantum computing; Specifically, the main task of the quantum computing optimization module in this embodiment is to optimize various computational processes in the system through the powerful computing power of quantum computing, especially in tasks that require large-scale parallel computing, search optimization, and complex decision-making analysis. By applying quantum computing algorithms, this module improves the computational efficiency of the system. Especially when dealing with big data processing, machine learning model optimization, and multi-objective optimization, it can significantly improve performance and response speed.
[0073] The quantum computing optimization module accelerates the computational tasks of the system based on quantum algorithms. It mainly breaks through the bottleneck of classical computers in processing large-scale data through the characteristic of qubits processing information in multiple states in parallel. Quantum computing can achieve efficient optimization through characteristics such as quantum superposition and quantum entanglement, showing great advantages especially in optimization problems.
[0074] Generally, the quantum computing optimization module optimizes the models in the system through quantum optimization algorithms (such as the Quantum Approximate Optimization Algorithm QAOA, the Variational Quantum Eigensolver VQE, etc.). For example, in the process of emotion recognition, the quantum computing optimization module can accelerate the training and inference speed of the model. Especially on large-scale datasets, quantum computing can significantly improve computational efficiency and optimization accuracy. In this embodiment, the quantum computing optimization module is mainly applied to optimize the training process of models, especially the training of deep learning models and machine learning models. Traditional optimization algorithms (such as gradient descent) may encounter problems such as slow convergence speed and large consumption of computing resources when dealing with large-scale datasets, while the quantum computing optimization module can significantly improve the training efficiency through quantum algorithms.
[0075] Specifically, the quantum computing optimization module performs parallel computing on multiple candidate solutions through quantum superposition states, thus accelerating the search process. For example, the Quantum Approximate Optimization Algorithm (QAOA) can be used to optimize the loss function of the model to help find the global optimal solution. Its basic process can be represented by the following formula: ; Where: is the loss function; is the quantum state; is the Hamiltonian are the optimization parameters.
[0076] By adjusting the quantum algorithm can find the optimal parameters that minimize the loss function, thereby improving the optimization speed and accuracy of the model.
[0077] The quantum variational algorithm (VQE) can also be used to optimize the weight update process of neural networks. By adjusting the weights of each layer of the neural network through quantum computing, the training process of the neural network can be accelerated, especially with higher efficiency when dealing with large-scale data.
[0078] The quantum computing optimization module can significantly improve the response speed of the system. Specifically, during the interaction between the virtual character and the user, the system needs to quickly process the user's emotional input and generate corresponding speech, actions, etc. The quantum computing optimization module accelerates the reasoning process through quantum algorithms, especially when dealing with large-scale parallel computing, it can greatly enhance the response ability of the system.
[0079] For example, the quantum computing optimization module can accelerate the reasoning process through quantum graph algorithms. Graph algorithms often encounter large computational problems when dealing with complex relational networks and decision trees, while quantum graph algorithms can simultaneously calculate on multiple paths through the superposition effect of qubits, thereby accelerating the overall reasoning process. This process can be expressed as: ; where: is the reasoning result; is the probability amplitude of each path; is at time under the path corresponding quantum state.
[0080] Through quantum computing, the reasoning results can be calculated in parallel on multiple paths, significantly improving the reasoning speed.
[0081] The quantum computing optimization module can also achieve the function of multi-objective optimization. In some embodiments, the system may need to balance multiple objectives. For example, the emotional expression of the virtual character needs to be weighed between accuracy and naturalness. In this case, the quantum computing optimization module can find the optimal balance point between multiple objectives through quantum multi-objective optimization algorithms.
[0082] By introducing the quantum computing optimization module, the system can significantly improve the processing efficiency and response speed. Especially when dealing with complex model optimization, real-time inference, and multi-objective optimization, quantum computing can greatly enhance the performance. Through technologies such as the Quantum Approximate Optimization Algorithm (QAOA), the Variational Quantum Eigensolver (VQE), and the quantum multi-objective optimization algorithm, the quantum computing optimization module can play an important role in tasks such as the emotional analysis, action generation, and speech synthesis of virtual characters, improving the system's computing efficiency and interaction experience.
[0083] The AI training module, which is connected to the interaction recognition module, conducts training based on interaction data; Specifically, the core task of the AI training module in this embodiment is to train each module of the virtual character, especially the deep learning models of related modules such as emotion generation, action generation, and speech generation. The AI training module improves the system's performance and accuracy in tasks such as emotion recognition, emotion generation, action, and speech generation by continuously optimizing and adjusting the model parameters. Through an efficient training process, the AI training module enables the virtual character to better understand and respond to the user's emotional input, thus providing a more natural, smooth, and realistic interaction experience.
[0084] In this embodiment, the AI training module closely cooperates with the aforementioned quantum computing optimization module, emotion generation module, action generation module, and speech generation module, etc., to optimize the overall performance of the system through an efficient training mechanism. By adopting advanced machine learning algorithms and quantum computing optimization algorithms, the AI training module conducts effective training on large-scale datasets to ensure that the virtual character can generate accurate actions and speech in different emotional states.
[0085] In this embodiment, the AI training module completes the training task by constructing multiple neural network models. The training process of the model relies on a large number of datasets, including multi-dimensional data such as the user's emotional input, interaction behavior, actions, and speech. By training these models, the AI training module enables the virtual character to generate corresponding body actions, speech, and emotional feedback according to different emotional states.
[0086] Generally, the AI training module uses deep learning frameworks (such as TensorFlow, PyTorch, etc.) to train the model. The training process includes steps such as data preprocessing, model initialization, loss function design, gradient calculation, and optimizer selection, with the goal of optimizing the weights and parameters of the model by minimizing the loss function. During the training process, the AI training module needs to continuously iterate and update the model parameters to improve the system's performance in tasks such as emotion analysis, action generation, and speech synthesis.
[0087] Specifically, the AI training module describes the training process through the following mathematical model; ; Where: is the total loss function; is the th sample's loss function; is the prediction result of the model; is the actual label; are the parameters of the model; is the total number of samples.
[0088] By minimizing the loss function, the system can optimize the training model to obtain the most suitable parameters.
[0089] In this embodiment, data preprocessing plays an important role in the AI training module. To ensure the quality and diversity of training data, the system needs to perform operations such as cleaning, standardizing, and normalizing the original data. Sentiment data, speech data, motion data, etc. may have different scales and dimensions. Therefore, the data preprocessing process can ensure the efficiency and accuracy during model training.
[0090] For example, when training a sentiment generation model, sentiment data may include different sentiment types (such as joy, sadness, anger, etc.) and their intensities. To enable this data to be effectively input into the neural network, the AI training module standardizes the sentiment data to meet the model input requirements. In addition, motion and speech data also need to undergo similar preprocessing to ensure the compatibility and consistency of different data sources.
[0091] In the AI training module, the choice of optimization algorithm directly affects the efficiency and convergence speed of model training. Generally, the AI training module uses classical gradient descent methods (such as Stochastic Gradient Descent SGD, Adam optimizer, etc.) to optimize model parameters. By continuously updating the model parameters, the loss function gradually decreases, thereby improving the accuracy of the system in tasks such as sentiment generation, motion generation, and speech generation.
[0092] Specifically, the optimization algorithm used by the AI training module can be expressed as: ; Where: is the th iteration's model parameters; is the learning rate; is the loss function with respect to the model parameters gradient.
[0093] When designing the loss function, the AI training module will select different loss functions according to the nature of the task. For example, during the training process of the emotion generation module, the cross-entropy loss function (Cross-EntropyLoss) can be used to measure the difference between the emotion classification predicted by the model and the true label; in the action generation module, the mean squared error (MSE) may be used to measure the difference between the generated action and the true action.
[0094] Quantum computing optimization algorithms can be used in the AI training module to accelerate the model training process. Through the parallelism of quantum computing, the AI training module can complete the training of large-scale datasets in a shorter time. Especially when dealing with complex models, quantum computing can significantly improve the computing efficiency and optimization accuracy.
[0095] For example, the quantum approximate optimization algorithm (QAOA) can accelerate the optimization process of the emotion generation model. Through quantum computing, QAOA can optimize multiple candidate solutions simultaneously, thereby improving the training speed. When optimizing the training of large-scale neural networks, quantum computing can break through the performance bottleneck of traditional computers and achieve faster convergence.
[0096] This process can be represented by the following formula: ; where; is the expected value of the quantum system, representing the quantum state under the Hamiltonian expected value; is the optimization parameter; is the Hamiltonian.
[0097] Quantum algorithms can accelerate the calculation process of the loss function through parallel computing, thereby improving the training efficiency.
[0098] The AI training module completes the training of tasks such as virtual character emotion generation, action generation, and speech synthesis through deep learning algorithms and quantum computing optimization algorithms. Through data preprocessing, optimization algorithms, and the design of loss functions, the AI training module can improve the training efficiency and accuracy of the model. The introduction of quantum computing further accelerates the training process and breaks through the bottleneck of traditional computers on large-scale datasets.
[0099] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. AI virtual real person interaction system based on emotion recognition, characterized by: include; The interaction recognition module is used to recognize the emotional state of the user input, obtain the interaction data, and output the emotion classification and its intensity; An emotion generation module, which is connected to the emotion recognition module and generates the emotion expression of the virtual character according to the emotion classification and intensity; A sound generation module, which is connected to the emotion generation module and generates a voice performance of the virtual character according to the emotion data output by the emotion generation module; An action generation module, which is connected to the emotion generation module and generates actions of the virtual character according to the emotional expression of the virtual character; A quantum computing optimization module, which is connected to the emotion generation module and the action generation module, and optimizes the computing efficiency in the emotion generation and action generation processes through quantum computing; The AI training module is connected to the interaction recognition module and is trained based on the interaction data.
2. The AI virtual real person interaction system based on emotion recognition according to claim 1 is characterized in that: The interaction identification module comprises: Speech recognition unit, used to convert user's speech input into text data, extract speech features and perform sentiment analysis; Facial expression analysis unit, which captures the user's facial expression and analyzes their emotional state through image recognition technology; The text analysis unit is used to analyze the user's text input and identify the emotional content therein.
3. The AI virtual real person interaction system based on emotion recognition according to claim 1 is characterized in that: The interaction data includes the user's voice, facial expressions, text input and feedback data of the virtual character.
4. The AI virtual real person interaction system based on emotion recognition according to claim 1 is characterized in that: The emotion generation module uses a generative adversarial network to generate emotions. The generative adversarial network includes a generator and a discriminator. The generator generates the emotional expression of the virtual character according to the input emotional data, and the discriminator is used to evaluate whether the generated emotional expression conforms to the real emotional expression.
5. The AI virtual real person interaction system based on emotion recognition according to claim 1 is characterized in that: The sound generation module uses neural network speech synthesis technology to generate the voice output of the virtual character. The voice performance adjusts the pitch, speaking speed, and tone parameters to adapt to the emotional expression of the virtual character.
6. The AI virtual real person interaction system based on emotion recognition according to claim 1 is characterized in that: The action generation module comprises: A variational autoencoder for mapping emotional expressions into a latent space and generating actions for the avatar based on the latent space; The particle system simulation unit generates natural movements of virtual characters based on the latent space, including facial expressions and body movements.
7. The AI virtual real person interaction system based on emotion recognition according to claim 1 is characterized in that: The quantum computing optimization module adopts a quantum optimization algorithm to accelerate the calculations in the emotion generation and action generation modules.
8. The AI virtual real person interaction system based on emotion recognition according to claim 1 is characterized in that: The AI training module uses deep learning and reinforcement learning technologies to perform model training and optimization, performs self-optimization through continuous user interaction data, and dynamically adjusts the virtual character's voice, movement, and emotional expression.
9. The AI virtual real person interaction system based on emotion recognition according to claim 8 is characterized in that: The deep learning is used for training of emotion recognition, sound generation and action generation, and the reinforcement learning dynamically adjusts the reaction and behavior of the virtual character according to the user's interaction data.
10. An AI virtual real person interaction method based on emotion recognition, according to the AI virtual real person interaction system based on emotion recognition according to any one of claims 1 to 9, characterized in that: The following steps are involved: The user's emotional data is obtained through the interactive recognition module, including voice input, facial expression data and text input. This data reflects the user's current emotional state and provides a basis for subsequent emotional generation and behavioral response; The emotion recognition module analyzes the acquired user emotion data, identifies the user's emotion state, and determines the emotion classification and emotion intensity. This step is the core of the entire interaction process. Accurately identifying the emotion state is the prerequisite for generating realistic virtual character reactions. Through the emotion generation module, the emotional expression of the virtual character is generated according to the identified emotion classification and intensity, and the generative adversarial network technology is used to ensure that the generated emotional expression is realistic and consistent with the logic of human emotions; According to the emotional expression of the virtual character, the action generation module generates corresponding virtual character actions, including facial expression changes and body movements, to ensure that the performance of the virtual character is consistent with its emotional state; According to the emotional expression of the virtual character, the sound generation module generates speech output that matches the emotional state, including the adjustment of pitch, speech speed, and tone characteristics to ensure that the speech feedback of the virtual character is consistent with its emotional state; Through the quantum computing optimization module, the computing efficiency required in the process of emotion generation and action generation is optimized. Quantum computing technology can accelerate the calculation in the process of emotion generation and action generation, thereby improving the system's response speed and real-time interaction capabilities; Through the AI training module, the performance of the emotion recognition, emotion generation, action generation and voice generation modules is continuously optimized based on the interaction data between users and virtual characters. This module uses deep learning and reinforcement learning techniques to enable the virtual characters to continuously adjust their reactions according to the interactions of different users to provide a more personalized interactive experience.