Voice interaction system and method based on AI intelligence

Through the voice interaction system based on AI intelligence, simulated dialogue is entered and excitement index is obtained in real time, the problem of single and low intelligence of traditional human-computer interaction systems is solved, and a more intelligent and user-friendly interactive experience is achieved.

CN120496512APending Publication Date: 2025-08-15JIANGXI YUANSHENG LANGHE PHARM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510180158.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional human-computer interaction tools lack diversity and entertainment, low intelligence and poor user experience.

Method used

Design a voice interaction system based on AI intelligence, enter simulated dialogue, obtain user excitement index in real time through the excitement module, record historical data for priority sorting, adjust the interaction behavior according to the excitement, and prompt the user whether to end the conversation and enter the next stage.

Benefits of technology

It improves the intelligence of human-computer interaction, brings users an excellent interactive experience, and adjusts simulated dialogue based on the excitement index to meet user interest needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496512A_ABST
    Figure CN120496512A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice interaction, in particular to a voice interaction system and method based on AI intelligence. The invention discloses a voice interaction system and method based on AI intelligence. A plurality of simulation dialogues capable of motivating the emotion of a user are input; the input simulation dialogue is selected by the user, and dialogue with the user is carried out based on the simulation dialogue selected by the user; calculating an average value of the excitement indexes of the user within a preset time threshold; recording an average value corresponding to each simulation dialogue and taking the average value as historical data; judging whether the average value is higher than a preset first excitement threshold value or not; and when the excitement index of the user reaches a preset second excitement threshold value, prompting the user whether to end the simulated dialogue and enter the next stage or not. The system has a voice interaction function, simulative dialogues more interested by the user are accurately found out according to historical data, the intelligent degree is high, and excellent emotional experience is brought to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice interaction technology, and in particular to a voice interaction system and method based on AI intelligence. Background Art

[0002] The interaction mode of traditional human-computer interaction tools is very simple, lacks diversity and entertainment, and has a poor user experience. In the existing field, tools with certain interactive operation functions have also been developed.

[0003] The interactive functions of human-computer interaction tools are single, they do not have voice interaction functions, and their intelligence level is low. Therefore, it is necessary to design a voice interaction system and method based on AI intelligence. Summary of the Invention

[0004] The purpose of the present invention is to solve at least one of the technical problems existing in the prior art and to provide a voice interaction system and method based on AI intelligence.

[0005] To achieve the above objectives, the present invention adopts the following technical solutions: A voice interaction system based on AI intelligence, comprising:

[0006] A recruitment module configured to record a number of simulated dialogues that can arouse user emotions;

[0007] a dialogue module configured to provide the recorded simulated dialogue for the user to select, and to conduct a dialogue with the user based on the simulated dialogue selected by the user;

[0008] An excitement module is configured to obtain the user's excitement index in real time during a conversation, and when a preset time threshold is reached, calculate the average value of the user's excitement index within the preset time threshold; record the average value corresponding to each simulated conversation and use it as historical data, and prioritize all simulated conversations based on the historical data, with simulated conversations with higher average values being ranked higher in priority;

[0009] a judgment module configured to judge whether the average value is higher than a preset first excitement threshold, and if so, continue the simulated conversation; if not, automatically change the simulated conversation according to the priority of the historical data;

[0010] The prompt module is configured to prompt the user whether to end the simulated conversation and enter the next stage when the user's excitement index reaches a preset second excitement threshold.

[0011] Furthermore, the recruitment module is specifically configured to name the simulated dialogues according to actual features of the simulated dialogues, and classify the simulated dialogues into several topics according to the names of the simulated dialogues.

[0012] Furthermore, the dialogue module is specifically configured as follows:

[0013] Displaying several topics to the user, and when the user selects a topic, displaying several simulated conversations in the selected topic to the user;

[0014] The simulated dialogue consists of system dialogue sentences and user dialogue sentences. The system dialogue sentences are played to the user and the user dialogue sentences are received as feedback.

[0015] After playing the system dialogue sentences, the user dialogue sentences that the user currently needs feedback on are displayed to the user.

[0016] Furthermore, the excitement module is specifically configured as follows:

[0017] The user's physiological sign status is obtained and the user's excitement index is calculated based on the analysis of the user's physiological sign status. The physiological sign status includes one or more of body temperature, heart rate, conversation speed, conversation sound decibel, exhalation frequency and intensity.

[0018] The present invention also provides a voice interaction method based on AI intelligence, comprising the following steps:

[0019] Step 1: Record several simulated conversations that can arouse the user's emotions;

[0020] Step 2: The recorded simulated dialogue is provided to the user for selection, and a dialogue is conducted with the user based on the simulated dialogue selected by the user;

[0021] Step 3: The user's excitement index is obtained in real time during the conversation. When a preset time threshold is reached, the average value of the user's excitement index within the preset time threshold is calculated; the average value corresponding to each simulated conversation is recorded and used as historical data. Based on the historical data, all simulated conversations are prioritized, and the simulated conversations with higher average values are ranked higher in priority.

[0022] Step 4: Determine whether the average value is higher than a preset first excitement threshold. If so, continue the simulated conversation. If not, automatically change the simulated conversation according to the priority of the historical data.

[0023] Step 5: When the user's excitement index reaches a preset second excitement threshold, the user is prompted whether to end the simulated conversation and enter the next stage.

[0024] Furthermore, in step 1, the simulated dialogues are named according to actual characteristics of the simulated dialogues, and the simulated dialogues are classified and grouped into several topics according to the names of the simulated dialogues.

[0025] Furthermore, in said step 2, it specifically includes:

[0026] Step 2.1, displaying several topics to the user, and after the user selects a topic, displaying several simulated conversations in the selected topic to the user;

[0027] Step 2.1: The simulated dialogue consists of system dialogue sentences and user dialogue sentences. The system dialogue sentences are played to the user and the user dialogue sentences are received as feedback.

[0028] Step 2.3: After playing the system dialogue sentence, the user dialogue sentence that the user currently needs to feedback is displayed to the user.

[0029] Furthermore, in step 3, the user's physiological sign status is obtained and the user's excitement index is calculated based on the analysis of the user's physiological sign status. The physiological sign status includes one or more of body temperature, heart rate, conversation speed, conversation sound decibels, exhalation frequency and intensity.

[0030] From the above description of the present invention, it can be seen that compared with the prior art, the AI-based voice interaction system and method of the present invention records a number of simulated dialogues that can mobilize the user's emotions; the recorded simulated dialogues are provided for the user to choose, and the user is engaged in a dialogue based on the simulated dialogue selected by the user; the user's excitement index is obtained in real time during the dialogue, and when the preset time threshold is reached, the average value of the user's excitement index within the preset time threshold is calculated; the average value corresponding to each simulated dialogue is recorded and used as historical data, and all simulated dialogues are prioritized based on the historical data, and the higher the average value, the higher the priority of the simulated dialogue; it is determined whether the average value is higher than the preset first excitement threshold. If the average value is higher than the preset first excitement threshold, the simulated dialogue is continued; if the average value is not higher than the preset first excitement threshold, the simulated dialogue is automatically replaced according to the priority ranking of the historical data; when the user's excitement index reaches the preset second excitement threshold, the user is prompted whether to end the simulated dialogue and enter the next stage. The present invention has a voice interaction function and adjusts the interactive behavior according to the user's excitement index. The average value of the excitement index of each simulated conversation is used as historical data. Based on the historical data, the simulated conversation that the user is more interested in is accurately found. It has a high degree of intelligence and brings an excellent interactive experience to the user. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a structural block diagram of an AI-based voice interaction system in a preferred embodiment of the present invention;

[0032] Figure 2 This is a flowchart of a voice interaction system based on AI intelligence in a preferred embodiment of the present invention;

[0033] Explanation of the numbers in the figure: 1. Recruitment module; 2. Dialogue module; 3. Excitement module; 4. Judgment module; 5. Prompt module. DETAILED DESCRIPTION

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.

[0035] In the description of the present invention, it should be noted that the terms "upper," "lower," "inner," "outer," "top / bottom," and the like, indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0036] Reference Figure 1 As shown, a preferred embodiment of the present invention is a voice interaction system based on AI intelligence, comprising:

[0037] The recruitment module 1 is configured to record a number of simulated dialogues that can arouse the user's emotions;

[0038] The dialogue module 2 is configured to provide the recorded simulated dialogue for the user to select, and to conduct a dialogue with the user based on the simulated dialogue selected by the user; specifically, different users are interested in different characters and topics, so allowing users to select simulated characters and simulated dialogues by themselves can bring a better experience to users.

[0039] Excitement module 3 is configured to obtain the user's excitement index in real time during a conversation. When a preset time threshold is reached, the module calculates the average of the user's excitement index within the preset time threshold. The average corresponding to each simulated conversation is recorded as historical data. Based on this historical data, all simulated conversations are prioritized, with simulated conversations with higher averages being prioritized higher. Specifically, the present invention, as an interactive system, obtains the user's excitement index to determine the user's excitement state and, based on this state, implements more intelligent interactive behaviors to enhance the user experience. The user's excitement state is graded using the excitement index. The following is an example of an excitement index: assuming the lowest excitement index is 0 and the highest excitement index is 10. The following is an example of calculating the average excitement index: For a preset time threshold of 5 minutes, if the first minute is 0, the second and third minutes are 2, and the fourth and fifth minutes are 5, then the average excitement index is 2*2+5*2 / 5=2.8. More specifically, a curve of the excitement index over time can be plotted, and the average excitement index can be derived from the curve. The average value corresponding to each simulated conversation is used as historical data, and all simulated conversations are prioritized based on the historical data. The higher the average value, the more interested the user is in the simulated conversation, and therefore the higher the priority is. If the average values calculated for a simulated conversation are different multiple times, the new average value is averaged again to obtain the basis for priority sorting.

[0040] The judgment module 4 is configured to judge whether the average value is higher than the preset first excitement threshold. If the average value is higher than the preset first excitement threshold, the simulated dialogue continues; if the average value is not higher than the preset first excitement threshold, the simulated dialogue is automatically replaced according to the priority of historical data; specifically, the time threshold starts from the moment the dialogue starts. Within the range of this time threshold, as long as the average value of the user's excitement index is higher than the preset first excitement threshold, it means that the user is more interested in the current simulated dialogue, so the simulated dialogue continues; if the average value of the user's excitement index is not higher than the preset first excitement threshold, it means that the user is not interested in the current simulated dialogue, so the user is notified to change the simulated dialogue, and voice interaction is performed with the user according to the user's own situation to give the user a better interactive experience.

[0041] Prompt Module 5 is configured to prompt the user whether to end the simulated conversation and proceed to the next stage when the user's excitement index reaches a preset second excitement threshold. When the user's excitement index reaches the preset second excitement threshold, the prompt module 5 prompts the user whether to end the simulated conversation and proceed to the next stage, thereby helping the user control the overall rhythm and improving the user experience.

[0042] As a preferred embodiment of the present invention, it may also have the following additional technical features:

[0043] In this embodiment, the recruitment module 1 is specifically configured to name the simulated conversations based on their actual characteristics, and to categorize the simulated conversations into several topics based on the names of the simulated conversations. Naming the simulated conversations facilitates user search. The actual characteristics of the simulated conversations include character identity characteristics, location characteristics, and situational characteristics. Since there are many simulated conversations, categorizing them into several topics facilitates search. The simulated conversations can be categorized based on both their names and actual characteristics.

[0044] In this embodiment, the dialogue module 2 is specifically configured as follows:

[0045] Displaying several topics to the user, and when the user selects a topic, displaying several simulated conversations in the selected topic to the user;

[0046] The simulated dialogue consists of system dialogue sentences and user dialogue sentences. The system dialogue sentences are played to the user and the user dialogue sentences are received as feedback.

[0047] After playing the system dialogue sentences, the user dialogue sentences that the user currently needs feedback on are displayed to the user.

[0048] In this embodiment, the excitement module 3 is specifically configured as follows:

[0049] Acquire the user's physiological status and analyze and calculate the user's excitement index based on the user's physiological status. The physiological status includes one or more of body temperature, heart rate, conversation speed, conversation volume decibels, exhalation frequency and intensity. Body temperature, heart rate, conversation speed, conversation volume decibels, exhalation frequency and intensity can reflect the user's excitement to a certain extent. Therefore, by acquiring physiological status such as body temperature, heart rate, conversation speed, conversation volume decibels, exhalation frequency and intensity through an acquisition unit, the user's excitement can be analyzed and calculated more accurately. The acquisition unit includes at least one sensor, including but not limited to a temperature sensor, a heart rate sensor, a sound sensor, and an EEG sensor.

[0050] Reference Figure 2 As shown, the present invention also provides a voice interaction method based on AI intelligence, comprising the following steps:

[0051] Step 1: Record several simulated conversations that can arouse the user's emotions;

[0052] Step 2: The recorded simulated dialogue is provided for the user to select, and a dialogue is conducted with the user based on the simulated dialogue selected by the user. Specifically, different users are interested in different characters and topics, so allowing users to select simulated characters and simulated dialogues by themselves can bring a better experience to users.

[0053] Step 3: During the conversation, the user's excitement index is obtained in real time. When a preset time threshold is reached, the average value of the user's excitement index within the preset time threshold is calculated. The average value corresponding to each simulated conversation is recorded and used as historical data. Based on the historical data, all simulated conversations are prioritized. The higher the average value, the higher the priority of the simulated conversation. Specifically, the present invention, as an interactive system, obtains the user's excitement index to determine the user's excitement state, and performs more intelligent interactive behaviors based on the user's excitement state to improve the user experience. The average value corresponding to each simulated conversation is used as historical data, and all simulated conversations are prioritized based on the historical data. The higher the average value, the more interested the user is in the simulated conversation, and therefore the higher the priority. If the average values calculated for a certain simulated conversation are different multiple times, the multiple average values are averaged again to obtain a new average value as the basis for priority ranking.

[0054] Step 4, determine whether the average value is higher than the preset first excitement threshold. If the average value is higher than the preset first excitement threshold, continue the simulated conversation; if the average value is not higher than the preset first excitement threshold, automatically change the simulated conversation according to the priority of historical data; specifically, the time threshold starts from the moment the conversation starts. Within the range of this time threshold, as long as the average value of the user's excitement index is higher than the preset first excitement threshold, it means that the user is more interested in the current simulated conversation, so continue the simulated conversation; if the average value of the user's excitement index is not higher than the preset first excitement threshold, it means that the user is not interested in the current simulated conversation, so notify the user to change the simulated conversation, and conduct voice interaction with the user according to the user's own situation to give the user a better interactive experience.

[0055] Step 5: When the user's excitement index reaches a preset second excitement threshold, the user is prompted to decide whether to end the simulated conversation and proceed to the next stage. This helps the user control the overall rhythm and improve the user experience.

[0056] In this embodiment, in step 1, the simulated dialogues are named according to actual features of the simulated dialogues, and the simulated dialogues are classified and grouped into several topics according to the names of the simulated dialogues.

[0057] In this embodiment, step 2 specifically includes:

[0058] Step 2.1, displaying several topics to the user, and after the user selects a topic, displaying several simulated conversations in the selected topic to the user;

[0059] Step 2.1: The simulated dialogue consists of system dialogue sentences and user dialogue sentences. The system dialogue sentences are played to the user and the user dialogue sentences are received as feedback.

[0060] Step 2.3: After playing the system dialogue sentence, the user dialogue sentence that the user currently needs to feedback is displayed to the user.

[0061] In this embodiment, in step 3, the user's physiological sign status is obtained and the user's excitement index is calculated based on the analysis of the user's physiological sign status. The physiological sign status includes one or more of body temperature, heart rate, conversation speed, conversation sound decibels, exhalation frequency and intensity.

[0062] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. That is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A voice interaction system based on AI intelligence, characterized by: include: A recruitment module (1) is configured to record a number of simulated dialogues that can arouse user emotions; A dialogue module (2) is configured to provide the recorded simulated dialogue for the user to select, and to conduct a dialogue with the user based on the simulated dialogue selected by the user; The excitement module (3) is configured to obtain the user's excitement index in real time during the conversation, and when a preset time threshold is reached, calculate the average value of the user's excitement index within the preset time threshold; record the average value corresponding to each simulated conversation and use it as historical data, and prioritize all simulated conversations based on the historical data, where the higher the average value, the higher the priority of the simulated conversation; A judgment module (4) is configured to judge whether the average value is higher than a preset first excitement threshold, and if the average value is higher than the preset first excitement threshold, continue the simulated dialogue; if the average value is not higher than the preset first excitement threshold, automatically change the simulated dialogue according to the priority of the historical data; The prompt module (5) is configured to prompt the user whether to end the simulated conversation and enter the next stage when the user's excitement index reaches a preset second excitement threshold.

2. The AI-based voice interaction system according to claim 1, characterized in that: The recruitment module (1) is specifically configured to name the simulated dialogues according to actual features of the simulated dialogues, and classify the simulated dialogues into several topics according to the names of the simulated dialogues.

3. The AI-based voice interaction system according to claim 1, characterized in that: The dialogue module (2) is specifically configured as follows: Displaying several topics to the user, and when the user selects a topic, displaying several simulated conversations in the selected topic to the user; The simulated dialogue consists of system dialogue sentences and user dialogue sentences. The system dialogue sentences are played to the user and the user dialogue sentences are received as feedback. After playing the system dialogue sentences, the user dialogue sentences that the user currently needs feedback on are displayed to the user.

4. The AI-based voice interaction system according to claim 1, characterized in that: The excitement module (3) is specifically configured as follows: The user's physiological sign status is obtained and the user's excitement index is calculated based on the analysis of the user's physiological sign status. The physiological sign status includes one or more of body temperature, heart rate, conversation speed, conversation sound decibel, exhalation frequency and intensity.

5. A voice interaction method based on AI intelligence, characterized in that: The following steps are involved: Step 1: Record several simulated conversations that can arouse the user's emotions; Step 2: The recorded simulated dialogue is provided to the user for selection, and a dialogue is conducted with the user based on the simulated dialogue selected by the user; Step 3: The user's excitement index is obtained in real time during the conversation. When a preset time threshold is reached, the average value of the user's excitement index within the preset time threshold is calculated; the average value corresponding to each simulated conversation is recorded and used as historical data. Based on the historical data, all simulated conversations are prioritized, and the simulated conversations with higher average values are ranked higher in priority. Step 4: Determine whether the average value is higher than a preset first excitement threshold. If so, continue the simulated conversation. If not, automatically change the simulated conversation according to the priority of the historical data. Step 5: When the user's excitement index reaches a preset second excitement threshold, the user is prompted whether to end the simulated conversation and enter the next stage.

6. The AI-based voice interaction method according to claim 5, characterized in that: In step 1, the simulated dialogues are named according to actual characteristics of the simulated dialogues, and several simulated dialogues are classified and grouped into several topics according to the names of the simulated dialogues.

7. The AI-based voice interaction method according to claim 5, characterized in that: In the step 2, it specifically includes: Step 2.1, displaying several topics to the user, and after the user selects a topic, displaying several simulated conversations in the selected topic to the user; Step 2.1: The simulated dialogue consists of system dialogue sentences and user dialogue sentences. The system dialogue sentences are played to the user and the user dialogue sentences are received as feedback. Step 2.3: After playing the system dialogue sentence, the user dialogue sentence that the user currently needs to feedback is displayed to the user.

8. The AI-based voice interaction method according to claim 5, characterized in that: In step 3, the user's physiological sign status is obtained and the user's excitement index is calculated based on the analysis of the user's physiological sign status. The physiological sign status includes one or more of body temperature, heart rate, conversation speed, conversation sound decibels, exhalation frequency and intensity.