Virtual digital human real-time interaction system based on AI copies and use method

Through speech recognition and natural language processing combined with multilingual databases and contextual reasoning, the problem of digital people misjudging user intentions in distance education is solved, and the accurate voice and expression output of virtual digital people in multilingual environments is realized, which improves the naturalness of teaching and emotional expression.

CN120544554AInactive Publication Date: 2025-08-26SHENZHEN TONGZHU CLOUD TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510475477.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the field of distance education, when digital people play virtual teachers, homonyms or similar pronunciations/spellings in Chinese and English lead to the system misjudging the user's intention or context, and the intonation and emotional expression during pronunciation are incorrect, reducing the naturalness and credibility of teaching.

Method used

By obtaining the language signal data collected by the microphone, using speech recognition, natural language processing and multilingual pronunciation semantic comparison database, combining cosine similarity algorithm to judge semantic ambiguity, and generating accurate speech and expression output through context reasoning and expression and intonation correction mechanisms to realize the natural interaction between virtual digital people and users.

Benefits of technology

It improves the accuracy and natural fluency of virtual digital people in multilingual interaction, enhances emotional expression and sense of reality, is suitable for diversified application scenarios such as education, medical care and companionship, and improves adaptability in an international environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544554A_ABST
    Figure CN120544554A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of virtual digital humans, in particular to a virtual digital human real-time interaction system based on an AI (artificial intelligence) duplicate and a use method, comprising the following steps: step 1, acquiring original language signal data # imgabs0 # and text input data # imgabs1 # acquired by a microphone, inputting the acquired # imgabs2 # into a voice recognition module, and inputting the acquired # imgabs2 # into a voice recognition module; the voice is converted into a text T through an automatic voice recognition algorithm # imgabs3 # arranged in the voice recognition module, the calculation formula for converting the voice into the text T is # imgabs4 #, and text input data # imgabs5 # and the text T are combined to generate comprehensive text information T0; according to the method, AI technologies such as speech recognition, natural language understanding, semantic embedding, context reasoning and the like are fused, so that the virtual digital human can accurately recognize the intention of the user and naturally generate the response, and particularly has stronger semantic judgment and adjustment capability when processing multiple rounds of dialogues and ambiguous words; therefore, the accuracy and the natural fluency of virtual interaction are greatly improved, and the user experience is better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of virtual digital human technology, and specifically to a real-time interactive system and method for using a virtual digital human based on an AI avatar. Background Art

[0002] The AI-powered virtual human real-time interactive system is a system built by integrating multiple cutting-edge technologies including artificial intelligence, big data, computer graphics, natural language processing (NLP), speech synthesis, motion capture, and real-time rendering. Its core goal is to create a digital "self" or virtual avatar that can highly realistically reproduce human appearance, voice, expressions, and behavior, enabling real-time, natural interaction with real humans. At the same time, the AI-based real-time interaction system uses deep learning and large-scale model technology to drive the virtual digital human, which is not only highly simulated in vision and sound, but can also understand user commands, engage in voice conversations, and even generate corresponding body language and facial expressions in real time in videos. In applications in the field of distance education, since digital humans often need to act as virtual teachers, in multilingual interactive systems, homographs in Chinese and English or words with similar pronunciations or spellings can lead to: the system misjudging user intent or context, incorrect intonation and emotional expression during speech synthesis, and mismatches between facial expressions and semantics, reducing the naturalness and credibility of teaching. Therefore, to address the above issues, a real-time interactive system and usage method for virtual digital humans based on AI avatars are proposed. Summary of the Invention

[0003] The purpose of the present invention is to provide a real-time interactive system and method for using AI-powered virtual digital humans to address the problem in distance education applications where, because digital humans often need to act as virtual teachers, Chinese and English homophones or words with similar pronunciations or spellings in multilingual interactive systems can lead to: the system misjudging user intent or context, incorrect intonation and emotional expression during speech synthesis, and a mismatch between facial expressions and semantics, reducing the naturalness and credibility of teaching.

[0004] To achieve the above object, the present invention provides the following technical solutions: The AI-based virtual digital human real-time interaction system and usage method include the following steps: Step 1: Get the original language signal data collected by the microphone and text input data , will obtain Input into the speech recognition module, through the automatic speech recognition algorithm set inside the speech recognition module , convert the speech into text T, the calculation formula for converting the speech into text T is: ; Entering text into data Combined with text T to generate comprehensive text information T0; Step 2: Use the pre-trained natural language processing model to encode the comprehensive text T0 and generate a semantic vector representation: ; in, Represents the NLP encoding function, the output is the semantic embedding representation of the text, and d is the vector dimension; And use the attention mechanism to construct the context vector C and keyword set K; Step 3: Use the cosine similarity formula to match each candidate word w with the term d in the database to determine whether there is semantic ambiguity, and form the lexicon that meets the conditions into an ambiguous set A: Calculate the similarity between each candidate word w in the text T0 and the corresponding terms in the multilingual pronunciation and semantic comparison database, and let each term in the database be ,in, It is a collection of data including language entries, corresponding pronunciation information, semantic tags, etc. For any word w and database term d, calculate the cosine similarity: ; in, represents the embedding vector of the corresponding word, is the vector norm; Setting the ambiguity threshold , define the ambiguous word set as: ; If the set A is not empty, it is considered that there is semantic ambiguity; Step 4: When ambiguity is detected, the system calls the multilingual pronunciation and semantic comparison database and contextual reasoning module for further confirmation and generates corresponding intonation and expression correction parameters.

[0005] As a further optimized content of the present invention, the following steps are also included: Step 5: Input the optimized speech synthesis parameters into the speech synthesis module to generate speech output with correct semantics and emotions; Step 6: Synchronize the voice output with the expression rendering module to drive the virtual digital human to render the corresponding lip shape, expression and posture to achieve real-time response to user input.

[0006] As a further optimization of the present invention, in step 2, the keyword set is the vocabulary set T0 after text segmentation, which is recorded as , each word With embedding vector , the calculation formula of the attention score of the attention mechanism is: ; in, is the query vector representing the current conversation context, Expressive words Weight in building context representation.

[0007] As a further optimization of the present invention, the mathematical expression of the component context vector C in step 2 is: ; At the same time, select words with higher weights to form a keyword total: ; in, The threshold for keyword extraction.

[0008] As a further optimization of the present invention, in step 4, the context-assisted semantic confirmation process is as follows: Use the context vector C to match the semantic information of the term in the database: ; For each candidate word , select the best matching term : ; in, is the coefficient that balances the word vector similarity and context similarity.

[0009] As a further optimization of the present invention, in step 4, the process of generating intonation correction parameters is as follows: Define the intonation correction parameter p as: ; in, is the intonation correction coefficient, It reflects the semantic deviation between the current candidate word and the target word, and then adjusts the output intonation.

[0010] As a further optimization of the present invention, in step 4, the process of generating expression correction parameters is as follows: Define emoticons The correction parameters are: ; in, is the expression control coefficient, which is used to determine the expression changes of the virtual digital human according to the context.

[0011] As a further optimized content of the present invention, it includes: a speech recognition module for receiving user speech input and converting it into text information; A natural language processing module, used for performing semantic analysis on the text information; A word meaning ambiguity detection module is used to detect whether there are homographs or easily confused words with similar pronunciations or spellings in the text information; A multilingual pronunciation and semantics database, which provides the corresponding information of spelling, pronunciation and meaning of multilingual words required for word meaning ambiguity detection; Contextual reasoning module, used to infer target semantics based on the current conversation content and user input; A speech synthesis module, used to synthesize speech output with emotional attributes based on the inference results; The expression rendering module is used to drive the virtual digital human to generate corresponding facial expressions, lip shapes and movements according to the emotional parameters of the voice output, so as to achieve synchronous output of voice and movement.

[0012] As a further optimized content of the present invention, wherein: the context reasoning module further includes: Context state tracking unit, used to record historical multi-round dialogue information and generate dialogue state; Large language model calling interface, used to call pre-trained large-scale language models to perform multiple rounds of logical reasoning on the current text semantics; The decision control unit is used to combine the user's historical input behavior with the current context information to determine the user's true intention and select the voice tone and emotional expression strategy that matches the intention.

[0013] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention integrates AI technologies such as speech recognition, natural language understanding, semantic embedding, and contextual reasoning to enable virtual digital humans to accurately identify user intent and naturally generate responses. In particular, it possesses stronger semantic judgment and adjustment capabilities when handling multi-round conversations and ambiguous words, thereby significantly improving the accuracy and natural fluency of virtual interactions and providing a better user experience. 2. In this invention, the system integrates a tone and expression correction mechanism, which can generate corresponding emotional tone and facial expressions based on context and keywords, allowing the virtual digital human to not only "speak" but also "perform", with its tone and movements becoming more humanized. This mechanism enhances emotional expression and realism in interaction, effectively breaking the rigid limitations of traditional human-computer interaction and is suitable for diverse application scenarios such as education, medical care, and companionship. 3. In the present invention, by constructing a multilingual pronunciation semantic comparison database and combining it with the cosine similarity algorithm to judge word meaning ambiguity, the system can identify and handle cross-language and cross-cultural lexical ambiguity problems, significantly improving its adaptability in an international environment. At the same time, combined with contextual semantic matching and large language model reasoning, it can further enhance intelligent understanding capabilities in multiple contexts, providing a technical foundation for global promotion. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is a flow chart of the method for using the AI ​​avatar virtual digital human real-time interaction system of the present invention; Figure 2 This is a system block diagram of the real-time interactive system of AI avatar virtual digital human based on the present invention. DETAILED DESCRIPTION

[0015] See also Figure 1-2 , the present invention provides a technical solution: The AI-based virtual digital human real-time interaction system and usage method include the following steps: Step 1: Get the original language signal data collected by the microphone and text input data , will obtain Input into the speech recognition module, through the automatic speech recognition algorithm set inside the speech recognition module ,Convert speech into text T. The calculation formula for converting speech into text T is: ; Entering text into data Combined with text T to generate comprehensive text information T0; Step 2: Use the pre-trained natural language processing model to encode the comprehensive text T0 and generate a semantic vector representation: ; in, Represents the NLP encoding function, the output is the semantic embedding representation of the text, and d is the vector dimension; And use the attention mechanism to construct the context vector C and keyword set K; Step 3: Use the cosine similarity formula to match each candidate word w with the term d in the database to determine whether there is semantic ambiguity, and form the lexicon that meets the conditions into an ambiguous set A: Calculate the similarity between each candidate word w in the text T0 and the corresponding terms in the multilingual pronunciation and semantic comparison database, and let each term in the database be ,in, It is a collection of data including language entries, corresponding pronunciation information, semantic tags, etc. For any word w and database term d, calculate the cosine similarity: ; in, represents the embedding vector of the corresponding word, is the vector norm; Setting the ambiguity threshold , define the ambiguous word set as: ; If the set A is not empty, it is considered that there is semantic ambiguity; Step 4: When ambiguity is detected, the system calls the multilingual pronunciation and semantic comparison database and contextual reasoning module for further confirmation and generates corresponding intonation and expression correction parameters. By combining speech recognition technology with text input, it effectively enhances the data acquisition accuracy of human-computer interaction and realizes the fusion processing of multimodal input, providing a more accurate input basis for subsequent semantic understanding and improving the system's ability to recognize user intentions.

[0016] As a technical solution for further implementation of this plan, the following steps are also included: Step 5: Input the optimized speech synthesis parameters into the speech synthesis module to generate speech output with correct semantics and emotions; Step 6: Synchronize the speech output with the expression rendering module to drive the virtual digital human to render the corresponding lip shape, expression, and posture, achieving real-time response to user input. Introducing the synchronous processing of speech synthesis and expression rendering to achieve natural interaction between the virtual digital human and the user, giving the virtual character a "humanized" expression, making the interactive experience more realistic and emotional. As a technical solution for further implementation of this solution, in step 2, the keyword set is the vocabulary set T0 after text segmentation, denoted as , each word With embedding vector , the calculation formula of the attention score of the attention mechanism is: ; in, is the query vector representing the current conversation context, Expressive words The weights used in constructing context representations and filtering keywords through the attention mechanism help the system focus on the most critical information in the text, improve the accuracy of semantic analysis, and lay a good foundation for subsequent contextual reasoning; As a technical solution for further implementing this solution, the mathematical expression of the component context vector C in step 2 is: ; At the same time, select words with higher weights to form a keyword total: ; in, The construction of the context vector C and the sum of keywords as the threshold for keyword extraction enables the system to accurately model and extract the context, thereby improving the system's ability to understand the user's continuous speech and its logical coherence; As a technical solution for further implementation of this solution, in step 4, the context-assisted semantic confirmation process is as follows: Use the context vector C to match the semantic information of the term in the database: ; For each candidate word , select the best matching term : ; in, To balance the coefficients of word vector similarity and context similarity, a context-assisted semantic confirmation mechanism can be used to introduce contextual reasoning in ambiguous situations, improve the system's intelligent inference capabilities, effectively avoid semantic misjudgments or misunderstandings, and enhance system robustness. As a technical solution for further implementing this solution, in step 4, the process for generating intonation correction parameters is as follows: Define the intonation correction parameter p as: ; in, is the intonation correction coefficient, Reflect the semantic deviation between the current candidate word and the target word, and then adjust the output intonation. Introducing an intonation correction mechanism makes the speech output more consistent with the semantic emotional expression, thus allowing the virtual digital human to have more realistic speech emotional changes and achieve the unity of speech naturalness and emotional accuracy. As a technical solution for further implementing this solution, in step 4, the process of generating expression correction parameters is as follows: Define emoticons The correction parameters are: ; in, The expression control coefficient is used to determine the expression changes of the virtual digital human according to the context. By setting the expression correction parameters, the expression changes of the virtual human can be adjusted according to the context and semantics, significantly enhancing its emotional expression ability and affinity, and improving the credibility of the interaction; As a further technical solution for implementing this solution, it includes: a speech recognition module for receiving user speech input and converting it into text information; Natural language processing module, used for semantic analysis of text information; The word meaning ambiguity detection module is used to detect whether there are homographs or easily confused words with similar pronunciations or spellings in text information; A multilingual pronunciation and semantics database, which provides the corresponding information of spelling, pronunciation and meaning of multilingual words required for word meaning ambiguity detection; Contextual reasoning module, used to infer target semantics based on the current conversation content and user input; A speech synthesis module, used to synthesize speech output with emotional attributes based on the inference results; The expression rendering module is used to drive the virtual digital human to generate corresponding facial expressions, lip shapes, and movements based on the emotional parameters of the speech output, achieving the synchronous output of speech and movement. The system has a clear structure and clear module division of labor. It supports multi-dimensional collaborative processing of speech, semantics, emotion, and expression, effectively improving the overall interactive capabilities of the virtual digital human and providing support for multi-language and multi-scenario applications. As a technical solution for further implementing this solution, the contextual reasoning module further includes: Context state tracking unit, used to record historical multi-round dialogue information and generate dialogue state; Large language model calling interface, used to call pre-trained large-scale language models to perform multiple rounds of logical reasoning on the current text semantics; The decision control unit is used to combine the user's historical input behavior with the current context information to determine the user's true intention and select the voice tone and emotional expression strategy that matches the intention. The contextual reasoning module enhances the memory ability and intention recognition of multiple rounds of dialogue. Combined with the large model reasoning capability and decision unit strategy, it makes the system more adaptable and "understanding" capable, achieving a more natural human-computer dialogue process.

[0017] This article uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only used to help understand the method of the present invention and its core ideas. The above is only a preferred implementation method of the present invention. It should be pointed out that due to the limitations of textual expression, there are objectively infinite specific structures. For ordinary technicians in this technical field, without departing from the principles of the present invention, they can make several improvements, modifications or changes, and can also combine the above technical features in an appropriate manner; these improvements, modifications, changes or combinations, or the direct application of the inventive concept and technical solution to other occasions without improvement, should be regarded as the scope of protection of the present invention.

Claims

1. The method for using the AI ​​avatar virtual digital human real-time interactive system is characterized by: The following steps are involved: Step 1: Get the original language signal data collected by the microphone and text input data , will obtain Input into the speech recognition module, through the automatic speech recognition algorithm set inside the speech recognition module , convert the speech into text T, the calculation formula for converting the speech into text T is: ; Entering text into data Combined with text T to generate comprehensive text information T0; Step 2: Use the pre-trained natural language processing model to encode the comprehensive text T0 and generate a semantic vector representation: ; in, Represents the NLP encoding function, the output is the semantic embedding representation of the text, and d is the vector dimension; And use the attention mechanism to construct the context vector C and keyword set K; Step 3: Use the cosine similarity formula to match each candidate word w with the term d in the database to determine whether there is semantic ambiguity, and form the lexicon that meets the conditions into an ambiguous set A: Calculate the similarity between each candidate word w in the text T0 and the corresponding terms in the multilingual pronunciation and semantic comparison database, and let each term in the database be ,in, It is a collection of data including each language word, corresponding pronunciation information, semantic tags, etc. For any word w and database term d, calculate the cosine similarity: ; in, represents the embedding vector of the corresponding word, is the vector norm; Setting the ambiguity threshold , define the ambiguous word set as: ; If the set A is not empty, it is considered that there is semantic ambiguity; Step 4: When ambiguity is detected, the system calls the multilingual pronunciation and semantic comparison database and contextual reasoning module for further confirmation and generates corresponding intonation and expression correction parameters.

2. The method for using the AI ​​avatar virtual digital human real-time interaction system according to claim 1, characterized in that: The following steps are also included: Step 5: Input the optimized speech synthesis parameters into the speech synthesis module to generate speech output with correct semantics and emotions; Step 6: Synchronize the voice output with the expression rendering module to drive the virtual digital human to render the corresponding lip shape, expression and posture to achieve real-time response to user input.

3. The method for using the AI ​​avatar virtual digital human real-time interaction system according to claim 1, characterized in that: In step 2, the keyword set is the vocabulary set T0 after text segmentation, denoted as , each word With embedding vector , the calculation formula of the attention score of the attention mechanism is: ; in, is the query vector representing the current conversation context, Expressive words Weight in building context representation.

4. The method for using the AI ​​avatar virtual digital human real-time interaction system according to claim 1, characterized in that: The mathematical expression of the component context vector C in step 2 is: ; At the same time, select words with higher weights to form a keyword total: ; in, The threshold for keyword extraction.

5. The method for using the AI ​​avatar virtual digital human real-time interaction system according to claim 1 is characterized by: In step 4, the context-assisted semantic confirmation process is as follows: Use the context vector C to match the semantic information of the term in the database: ; For each candidate word , select the best matching term : ; in, is the coefficient that balances the word vector similarity and context similarity.

6. The method for using the AI ​​avatar virtual digital human real-time interaction system according to claim 1, characterized in that: In step 4, the process of generating intonation correction parameters is as follows: Define the intonation correction parameter p as: ; in, is the intonation correction coefficient, It reflects the semantic deviation between the current candidate word and the target word, and then adjusts the output intonation.

7. The method for using the AI ​​avatar virtual digital human real-time interaction system according to claim 1, characterized in that: In step 4, the process of generating expression correction parameters is as follows: Define emoticons The correction parameters are: ; in, is the expression control coefficient, which is used to determine the expression changes of the virtual digital human according to the context.

8. The AI ​​avatar-based virtual digital human real-time interaction system according to any one of claims 1 to 7, characterized in that: include: A speech recognition module, which is used to receive user speech input and convert it into text information; A natural language processing module, used for performing semantic analysis on the text information; A word meaning ambiguity detection module is used to detect whether there are homographs or easily confused words with similar pronunciations or spellings in the text information; A multilingual pronunciation and semantics database, which provides the corresponding information of spelling, pronunciation and meaning of multilingual words required for word meaning ambiguity detection; Contextual reasoning module, used to infer target semantics based on the current conversation content and user input; A speech synthesis module, used to synthesize speech output with emotional attributes based on the inference results; The expression rendering module is used to drive the virtual digital human to generate corresponding facial expressions, lip shapes and movements according to the emotional parameters of the voice output, so as to achieve synchronous output of voice and movement.

9. The AI ​​avatar-based virtual digital human real-time interaction system according to claim 8 is characterized by: The contextual reasoning module further comprises: Context state tracking unit, used to record historical multi-round dialogue information and generate dialogue state; Large language model calling interface, used to call pre-trained large-scale language models to perform multiple rounds of logical reasoning on the current text semantics; The decision control unit is used to combine the user's historical input behavior with the current context information to determine the user's true intention and select the voice tone and emotional expression strategy that matches the intention.

Citation Information

Cited By

  • Real-time voice stream dialogue interaction method and system based on large language model

    CN121191519A