Digital human system based on interactive artificial intelligence

By integrating user interface, natural language processing, speech recognition synthesis and emotion analysis understanding modules in the digital human system, the problem of digital human difficulty in identifying user emotions is solved and the user interaction experience is improved.

CN120012817AInactive Publication Date: 2025-05-16SHAANXI SHIHE QIFU TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510097645.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for existing digital people to accurately identify the user's underlying emotional state, resulting in a decrease in the experience of the user's interaction with digital people.

Method used

A digital human system based on interactive artificial intelligence is designed to collect text and voice input information through user interface modules, combine natural language processing, speech recognition synthesis and emotion analysis understanding modules to conduct comprehensive analysis to identify the user's potential emotional state.

Benefits of technology

By accurately identifying the user's underlying emotional state, the user's experience during interaction with digital people is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012817A_ABST
    Figure CN120012817A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a digital human system based on interactive artificial intelligence. According to the system, character information input by a user through a textbox and information input by voice through a microphone are collected through a user interface module, and a natural language processing module analyzes the captured text information and identifies intention, key information, entities and contexts of the user; the speech recognition synthesis module converts speech input information into texts through speech recognition, and the sentiment analysis and understanding module performs keyword analysis according to the texts input by the user and speech data, extracts sentiment features of the user, performs sentiment classification, calculates sentiment recognition accuracy and sentiment intensity scores, and performs sentiment analysis and understanding according to the sentiment recognition accuracy and sentiment intensity scores. The response generation module generates a corresponding text reply according to the user sentiment classification and the sentiment intensity score made by the sentiment analysis and understanding module, and the response output module transmits the text reply made by the system to the user through the user interface module to complete interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an interactive artificial intelligence digital human system. Background Art

[0002] Digital humans refer to virtual characters based on artificial intelligence technology, computer graphics and natural language processing. They usually have human appearance and behavior and can interact with users. They can simulate human emotions, language and behavior and are often used in various application scenarios. Digital humans can be used as teaching assistants to help students learn through interaction and provide personalized learning experiences. They can simulate different scenarios to enhance learning effects. In video games and film and television productions, digital humans can participate in the plot as characters to enhance the user's immersion and experience. As an important direction for future technological development, digital humans have shown great potential and application value in various fields.

[0003] Although modern digital humans can synthesize language and simulate emotions, they lack an understanding of complex emotions and have difficulty accurately identifying users' potential emotional states, which reduces the user experience during interaction with digital humans. Summary of the invention

[0004] 1. Technical issues to be resolved

[0005] In view of the shortcomings of the prior art, the present invention provides an interactive artificial intelligence digital human system, which has a user interface module that collects text information input by the user through a text box and information input by voice through a microphone, and calculates the input delay Sryc and the voice duration Ycsj to evaluate the complexity of the input information. The natural language processing module analyzes the captured text information, identifies the user's intention, key information, entities and context, calculates the intention recognition accuracy Sbzq, the success rate of entity resolution Sjcg and the context retention rate Sxbc, and the speech recognition and synthesis module converts the voice input information into text through speech recognition, calculates the word error rate Sccw of speech recognition and the naturalness of speech synthesis The score Hcpf is used to evaluate the fluency and naturalness of the output speech. The sentiment analysis and understanding module performs keyword analysis based on the text and voice data input by the user, extracts the user's emotional features and makes sentiment classification, and calculates the sentiment recognition accuracy Qgsb and the sentiment intensity score Qqpf. The response generation module generates the corresponding text reply based on the user's sentiment classification and sentiment intensity score made by the sentiment analysis and understanding module. The response output module transmits the text reply made by the system to the user through the user interface module to complete the interaction. By comprehensively analyzing the voice and text information input by the user, the user's potential emotional state is accurately identified, which improves the user's experience in the interaction process with the digital human and solves the above problems.

[0006] (II) Technical solution

[0007] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: an interactive artificial intelligence digital human system, comprising a user interface module, a natural language processing module, a speech recognition and synthesis module, a sentiment analysis and understanding module, a response generation module and a response output module;

[0008] The user interface module collects text information input by the user through the text box and voice input information through the microphone, and calculates the input delay Sryc and the voice duration Ycsj to evaluate the complexity of the input information. The collected text and voice information are transmitted to the natural language processing module;

[0009] The natural language processing module analyzes the captured text information, identifies the user's intention, key information, entities and context, and calculates the intention recognition accuracy Sbzq, the entity resolution success rate Sjcg and the context retention rate Sxbc;

[0010] The speech recognition and synthesis module converts the speech input information into text through speech recognition, calculates the word error rate Sccw of speech recognition and the naturalness score Hcpf of speech synthesis, which are used to evaluate the fluency and naturalness of the output speech;

[0011] The sentiment analysis and understanding module performs keyword analysis based on the text and voice data input by the user, extracts the user's sentiment characteristics and makes sentiment classification, and calculates the sentiment recognition accuracy Qgsb and the sentiment intensity score Qqpf;

[0012] The response generation module generates a corresponding text reply based on the user emotion classification and emotion intensity score made by the emotion analysis and understanding module;

[0013] The response output module transmits the text reply made by the system to the user through the user interface module to complete the interaction.

[0014] Preferably, the formula for calculating the input delay Sryc by the user interface module is as follows:

[0015] Sryc=Kssr-Srsj+Clsj

[0016] In the formula, Sryc represents input delay, Kssr represents the time when the user starts inputting, Srsj represents the time when the system receives the input, and Clsj represents the time required for the system to process the input. The above values ​​are obtained through the system timestamp.

[0017] Preferably, the formula for calculating the speech duration Ycsj by the user interface module is as follows:

[0018] Ycsj=Yh js-Yh ks

[0019] In the formula, Ycsj represents the duration of speech, Yh js represents the time when the user stops speaking, and Yhks represents the time when the user starts speaking. The above values ​​are obtained through the system timestamp.

[0020] Preferably, the formula for calculating the intention recognition accuracy Sbzq by the natural language processing module is as follows:

[0021]

[0022] In the formula, Sbzq represents the accuracy of intent recognition, Zqsb represents the number of correctly recognized intents, and Zyts represents the total number of intents, including correct and incorrect intents.

[0023] Preferably, the formula for calculating the success rate Sjcg of entity resolution by the natural language processing module is as follows:

[0024]

[0025] In the formula, Sjcg represents the success rate of entity resolution, Cgjx represents the number of successfully resolved entities, and Wbzl represents the total number of entities in the input text.

[0026] Preferably, the natural language processing module calculates the context retention rate Sxb c The formula is as follows:

[0027]

[0028] In the formula, Sxbc represents the context retention rate, Cgbc represents the amount of context information successfully retained, and Zsxw represents the total amount of context information in the conversation.

[0029] Preferably, the speech recognition synthesis module calculates the word error rate Scc of speech recognition w The formula is as follows:

[0030]

[0031] In the formula, Sccw represents the word error rate of speech recognition, Cwsb represents the number of words that are incorrectly recognized, Ldsb represents words that exist in the original text but are not recognized, Ccwr represents words that exist in the recognition result but are not in the original text, and Yszc represents the total number of words in the original audio.

[0032] Preferably, the formula for calculating the naturalness score Hcpf of speech synthesis by the speech recognition and synthesis module is as follows:

[0033]

[0034] In the formula, Hcpf represents the naturalness score of speech synthesis, N represents the total number of users who rated it, and Rk i It represents the rating given by the i-th user, and i represents the counting subscript.

[0035] Preferably, the formula for calculating the emotion recognition accuracy Qgsb by the emotion analysis and understanding module is as follows:

[0036]

[0037] In the formula, Qgsb represents the emotion recognition accuracy, Qdsz represents the number of correctly recognized emotions, and Zqsl represents the total number of emotions.

[0038] Preferably, the formula for calculating the sentiment intensity score Qqpf by the sentiment analysis and understanding module is as follows:

[0039]

[0040] In the formula, Qqpf represents the sentiment intensity score, M represents the total number of samples used to calculate the intensity score, and Qm i Represents the sentiment intensity score of the i-th sample.

[0041] Compared with the prior art, the present invention provides an interactive artificial intelligence digital human system, which has the following beneficial effects:

[0042] The present invention collects text information input by the user through a text box and information input by voice through a microphone through a user interface module, and calculates the input delay Sryc and the voice duration Ycsj to evaluate the complexity of the input information. The natural language processing module analyzes the captured text information, identifies the user's intention, key information, entity and context, calculates the intention recognition accuracy Sbzq, the success rate of entity resolution Sjcg and the context retention rate Sxbc. The speech recognition and synthesis module converts the voice input information into text through speech recognition, calculates the word error rate Sccw of speech recognition and the naturalness score Hcpf of speech synthesis to evaluate the complexity of the input information. The fluency and naturalness of the output speech are estimated. The sentiment analysis and understanding module performs keyword analysis based on the text and speech data input by the user, extracts the user's emotional features and makes sentiment classification, and calculates the sentiment recognition accuracy Qgsb and the sentiment intensity score Qqpf. The response generation module generates the corresponding text reply based on the user's sentiment classification and sentiment intensity score made by the sentiment analysis and understanding module. The response output module transmits the text reply made by the system to the user through the user interface module to complete the interaction. By comprehensively analyzing the voice and text information input by the user, the user's potential emotional state is accurately identified, thereby improving the user's experience in the interaction process with the digital human. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic diagram of the system flow of the present invention. DETAILED DESCRIPTION

[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0045] In view of the problem that the current digital human system is difficult to accurately identify the user's potential emotional state, which reduces the user's experience in the process of interaction with the digital human, a digital human system based on interactive artificial intelligence is proposed. Figure 1 , the system includes a user interface module, a natural language processing module, a speech recognition and synthesis module, a sentiment analysis and understanding module, a response generation module and a response output module;

[0046] The user interface module collects text information input by the user through the text box and voice input information through the microphone, and calculates the input delay Sryc and the voice duration Ycsj to evaluate the complexity of the input information, where:

[0047] The input delay calculation formula is as follows:

[0048] Sryc=Kssr-Srsj+Clsj

[0049] Lower input latency can significantly improve the user's interactive experience. When users input information, the fast response time makes users feel that the system is responsive and can communicate more naturally. In the formula, Sryc represents the input latency, Kssr represents the time when the user starts inputting, Srsj represents the time when the system receives the input, and Clsj represents the time required for the system to process the input. The above values ​​are obtained through the system timestamp. By analyzing the input latency, developers can identify performance bottlenecks in the system, such as insufficient processing power, network latency, or data transmission problems, and thus perform targeted optimization.

[0050] The calculation formula of speech duration is as follows:

[0051] Ycsj=Yh js-Yh ks

[0052] Analysis of speech duration can help the system better identify the user's intention. Longer speeches may indicate that the user is expressing complex ideas, while shorter speeches may indicate simple requests or questions. The system can adjust its processing logic according to the speech duration. In the formula, Ycsj represents the speech duration, Yhjs represents the time when the user stops speaking, and Yhks represents the time when the user starts speaking. The above values ​​are obtained through the system timestamp. Speech duration can be used as part of sentiment analysis. Longer duration may be related to the user's emotional investment. The system can use this information to better understand the user's emotional state and make more appropriate responses.

[0053] The natural language processing module analyzes the captured text information, identifies the user's intention, key information, entities and context, and calculates the intention recognition accuracy Sbzq, the entity resolution success rate Sjcg and the context retention rate Sxbc, where:

[0054] The calculation formula for intent recognition accuracy is as follows:

[0055]

[0056] High intent recognition accuracy ensures that the system provides highly relevant responses, allowing users to feel the system's intelligence and understanding during interaction, thereby improving satisfaction. In the formula, Sbzq represents the accuracy of intent recognition, Zqsb represents the number of correctly recognized intents, and Zyts represents the total number of intents, including correct and incorrect intents. By analyzing the accuracy of intent recognition, the system can understand the preferences of different users and provide a more personalized interaction experience.

[0057] The calculation formula for the success rate of entity resolution is as follows:

[0058]

[0059] Accurate entity resolution helps extract important information from user input and ensures that the system can make more relevant responses based on the user's specific needs. In the formula, Sjcg represents the success rate of entity resolution, Cgjx represents the number of successfully resolved entities, and Wbzl represents the total number of entities in the input text. By resolving multiple entities in user input, the system can obtain more comprehensive information, enabling it to have better contextual understanding when analyzing data.

[0060] The calculation formula of context retention rate is as follows:

[0061]

[0062] A high context retention rate can ensure that the user's intentions and information are accurately understood in multiple rounds of interactions, avoid confusion, and improve the naturalness and fluency of the conversation. In the formula, Sxbc represents the context retention rate, Cgbc represents the amount of context information successfully retained, and Zsxw represents the total amount of context information in the conversation. Many users' queries are complex and multi-step in nature. A high context retention rate can effectively solve these complex interactions and provide high-quality responses.

[0063] The speech recognition and synthesis module converts the speech input information into text through speech recognition, calculates the word error rate Sccw of speech recognition and the naturalness score Hcpf of speech synthesis, which are used to evaluate the fluency and naturalness of the output speech, where:

[0064] The formula for calculating the word error rate of speech recognition is as follows:

[0065]

[0066] The word error rate provides a quantitative standard to evaluate the accuracy of speech recognition and help developers understand the performance of the system. In the formula, Sccw represents the word error rate of speech recognition, Cwsb represents the number of words that are incorrectly recognized, Ldsb represents words that exist in the original text but are not recognized, Ccwr represents words that exist in the recognition result but are not in the original text, and Yszc represents the total number of words in the original audio. By monitoring the word error rate in different environments and language scenarios, the system can be continuously adjusted to meet the needs of multiple languages ​​and dialects, improving its versatility.

[0067] The naturalness score of speech synthesis is calculated as follows:

[0068]

[0069] Highly natural speech synthesis makes users sound more like they are communicating with humans rather than machines, making users feel more comfortable and intimate during interaction. In the formula, Hcpf represents the naturalness score of speech synthesis, N represents the total number of users who have scored, and Rk i represents the rating given by the i-th user, where i represents the count subscript. An accurate naturalness rating can help ensure that the synthesized audio can be easily understood in different situations, especially in noisy environments, where clear speech output is particularly important.

[0070] The sentiment analysis and understanding module performs keyword analysis based on the text and voice data input by the user, extracts the user's sentiment characteristics and makes sentiment classification, and calculates the sentiment recognition accuracy Qgsb and the sentiment intensity score Qqpf, where:

[0071] The calculation formula for emotion recognition accuracy is as follows:

[0072]

[0073] Accurate emotion recognition can help the system understand the user's emotional state in the conversation, thereby providing a more psychological response. In the formula, Qgsb represents the emotion recognition accuracy, Qdsz represents the number of correctly recognized emotions, and Zqsl represents the total number of emotions. Through emotion recognition, the digital human can adjust its response according to the user's emotional state, making the interaction more humane and interactive.

[0074] The calculation formula for the sentiment intensity score is as follows:

[0075]

[0076] The emotion intensity score can help the system understand the depth and intensity of the user's emotions more carefully so as to make a more appropriate response. In the formula, Qqpf represents the emotion intensity score, M represents the total number of samples used to calculate the intensity score, and Qm i represents the emotion intensity score of the i-th sample. By analyzing the emotion intensity, the system can select a more appropriate intervention strategy, such as providing more support or comfort when the user shows strong negative emotions;

[0077] The response generation module is closely integrated with the sentiment analysis and understanding module to achieve accurate classification and sentiment intensity scoring of user emotions. This process uses natural language processing technology and sentiment analysis algorithms to parse the text input by the user, extract potential sentiment features and sentiment polarity, such as positive, negative or neutral, and then use the sentiment intensity scoring system to quantify the depth of the user's emotions, such as distinguishing between happy and very happy, and assigning different intensity scores. Once the emotional state and intensity are accurately identified, the response generation module uses the rule engine and deep learning generation technology to generate the most appropriate text response. The system will dynamically adjust the tone, wording and content of the response according to the classification and intensity of the user's emotions to ensure that the response can not only reflect the understanding of the user's emotions, but also remain natural and fluent when providing information or suggestions;

[0078] The response output module is responsible for effectively delivering the text replies generated by the system to the user through the user interface module, completing the last step of intelligent interaction. Specifically, the response output module first converts the generated text into a format suitable for the user interface. This process may include encoding, tokenization and formatting of the text to ensure that the information can be displayed correctly on different platforms, such as mobile applications, web pages or smart devices. In the user interface module, the front-end development uses responsive design technology to ensure that the interface remains friendly and consistent regardless of the device the user uses. Then, the module uses the application programming interface to interact with the back-end for data, and ensures that information can be transmitted instantly without affecting other user operations through asynchronous requests AJAX. At the same time, the emotional feedback mechanism allows the system to dynamically adjust subsequent interactive content based on the user's real-time response, further enhancing the smoothness of the interaction and user experience. In addition, to improve accessibility and barrier-free experience, the response output module may also convert text into voice output, and use speech synthesis technology to present the reply to the user in a natural and fluent voice form, so that it can receive information more easily.

[0079] By comprehensively analyzing the voice and text information input by the user, the user's potential emotional state can be accurately identified, thereby improving the user's experience in the interaction process with the digital human.

[0080] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An interactive artificial intelligence digital human system, characterized by: It includes a user interface module, a natural language processing module, a speech recognition and synthesis module, a sentiment analysis and understanding module, a response generation module, and a response output module; The user interface module collects text information input by the user through the text box and voice input information through the microphone, and calculates the input delay Sryc and the voice duration Ycsj to evaluate the complexity of the input information, and transmits the collected text and voice information to the natural language processing module; The natural language processing module analyzes the captured text information, identifies the user's intention, key information, entities and context, and calculates the intention recognition accuracy Sbzq, the entity resolution success rate Sjcg and the context retention rate Sxbc; The speech recognition and synthesis module converts the speech input information into text through speech recognition, calculates the word error rate Sccw of speech recognition and the naturalness score Hcpf of speech synthesis, which are used to evaluate the fluency and naturalness of the output speech; The sentiment analysis and understanding module performs keyword analysis based on the text and voice data input by the user, extracts the user's sentiment characteristics and makes sentiment classification, and calculates the sentiment recognition accuracy Qgsb and the sentiment intensity score Qqpf; The response generation module generates a corresponding text reply based on the user emotion classification and emotion intensity score made by the emotion analysis and understanding module; The response output module transmits the text reply made by the system to the user through the user interface module to complete the interaction.

2. The interactive artificial intelligence digital human system according to claim 1, characterized in that: The formula for calculating the input delay Sryc by the user interface module is as follows: Sryc=Kssr-Srsj+Clsj In the formula, Sryc represents input delay, Kssr represents the time when the user starts inputting, Srsj represents the time when the system receives the input, and Clsj represents the time required for the system to process the input. The above values ​​are obtained through the system timestamp.

3. The interactive artificial intelligence digital human system according to claim 2, characterized in that: The formula for calculating the speech duration Ycsj by the user interface module is as follows: Ycsj=Yh js-Yh ks In the formula, Ycsj represents the speech duration, Y h js represents the time when the user stops speaking, and Y h ks represents the time when the user starts speaking. The above values ​​are obtained through the system timestamp.

4. The interactive artificial intelligence digital human system according to claim 3, characterized in that: The formula for calculating the intent recognition accuracy Sbzq by the natural language processing module is as follows: In the formula, Sbzq represents the accuracy of intent recognition, Zqsb represents the number of correctly recognized intents, and Zyts represents the total number of intents, including correct and incorrect intents.

5. The interactive artificial intelligence digital human system according to claim 4, characterized in that: The formula for calculating the success rate Sjcg of entity resolution by the natural language processing module is as follows: In the formula, Sjcg represents the success rate of entity resolution, Cgjx represents the number of successfully resolved entities, and Wbzl represents the total number of entities in the input text.

6. The interactive artificial intelligence digital human system according to claim 5, characterized in that: The formula for calculating the context retention rate Sxbc by the natural language processing module is as follows: In the formula, Sxbc represents the context retention rate, Cgbc represents the amount of context information successfully retained, and Zsxw represents the total amount of context information in the conversation.

7. The interactive artificial intelligence digital human system according to claim 6, characterized in that: The formula for calculating the word error rate Sccw of speech recognition by the speech recognition synthesis module is as follows: In the formula, Sccw represents the word error rate of speech recognition, Cwsb represents the number of words that are incorrectly recognized, Ldsb represents words that exist in the original text but are not recognized, Ccwr represents words that exist in the recognition result but are not in the original text, and Yszc represents the total number of words in the original audio.

8. The interactive artificial intelligence digital human system according to claim 7, characterized in that: The formula for calculating the naturalness score Hcpf of speech synthesis by the speech recognition and synthesis module is as follows: In the formula, Hcpf represents the naturalness score of speech synthesis, N represents the total number of users who rated it, and Rk i It represents the rating given by the i-th user, and i represents the counting subscript.

9. The interactive artificial intelligence digital human system according to claim 8, characterized in that: The formula for calculating the emotion recognition accuracy Qgsb by the emotion analysis and understanding module is as follows: In the formula, Qgsb represents the emotion recognition accuracy, Qdsz represents the number of correctly recognized emotions, and Zqsl represents the total number of emotions.

10. The interactive artificial intelligence digital human system according to claim 9, characterized in that: The formula for calculating the sentiment intensity score Qqpf by the sentiment analysis and understanding module is as follows: In the formula, Qqpf represents the sentiment intensity score, M represents the total number of samples used to calculate the intensity score, and Qm i Represents the sentiment intensity score of the i-th sample.

Citation Information

Patent Citations

  • Test method and system, processing equipment, electronic equipment and storage medium

    CN117116297A

  • Virtual digital human interaction method and system

    CN118426593A

  • Method and system for generating digital human

    CN118607343A

  • Interaction system applied to multi-scene digital human

    CN119052521A

  • Intelligent glasses streaming voice dialogue interaction system and method based on large language model

    CN119091878A

Cited By

  • Plateau intelligent medical care cabin and man-machine interaction method thereof

    CN120472902A

  • Interactive automatic explanation method for converting traditional video into artificial intelligence digital human

    CN121126039A