Intention understanding driven large model AI multi-role cooperative interaction method and device

Through the large-model AI multi-role collaborative interaction method driven by intent understanding, and utilizing the collaborative work of the host, slave and server, the problem that the existing AI system cannot recognize user intentions is solved, flexible interaction between multiple roles is achieved, and the interaction accuracy and user experience are improved.

CN120704636APending Publication Date: 2025-09-26SHANGHAI LINGYUN TECHNOLOGY DEVELOPMENT CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510841530.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing AI dialogue systems are unable to effectively identify user intentions, resulting in an inability to accurately determine which character the user wants to interact with in one-to-many interaction scenarios, reducing the user experience.

Method used

Through the large-model AI multi-role collaborative interaction method driven by intent understanding, the host, slave and server work together to select the designated role and generate the corresponding speech according to the user instructions and real-time monitoring distance, thus realizing flexible interaction between multiple roles.

Benefits of technology

It improves the accuracy and flexibility of interaction, conforms to human conversation habits, enhances user experience, can understand needs in real time based on user speech and generate role speech, and adapts to interaction scenarios between multiple roles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704636A_ABST
    Figure CN120704636A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to an intention understanding driven large model AI multi-role collaborative interaction method and device. The method comprises the following steps: a host understands user intention according to a user instruction, selects one or more specified roles from dialogue roles and awakens the roles; uploading the ID of the slave corresponding to each specified role to a server; receiving the greeting of the specified role returned by the server, and sending the greeting to the corresponding slave for playing; user speech is collected in real time and uploaded to a server, so that the server understands user intentions according to the user speech, and speech of a specified role is generated through a large model AI; and forwarding the speech of the specified role to the corresponding slave for playing. Wherein the host is worn on the body of a user, the number of the slave machines is one or more, each slave machine is arranged on a physical entity, and each physical entity is used for playing a dialogue role. According to the invention, the interaction accuracy and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a large-model AI multi-role collaborative interaction method and device driven by intent understanding. Background Art

[0002] Currently, most AI (Artificial Intelligence) dialogue systems (such as smart speakers and chatbots) only support single-character interactions and are unable to simulate multi-character collaboration scenarios. For example, in applications such as education and training, entertainment games, or virtual social networking, users may need to interact with multiple AI characters simultaneously. However, due to the lack of effective recognition of user intent, traditional systems cannot accurately determine which character the user wishes to interact with or the specific content of the user's instructions in one-to-many interactive dialogue scenarios, thus reducing the user experience. Summary of the Invention

[0003] In order to solve the above problems in the prior art, the present invention proposes a large-model AI multi-role collaborative interaction method and device driven by intent understanding, which improves the accuracy of interaction and user experience.

[0004] In one aspect of the present invention, a large-model AI multi-role collaborative interaction method driven by intent understanding is proposed. The method is applicable to a host in a multi-role interaction system, wherein the multi-role interaction system includes: the host, a slave, and a server;

[0005] The host is worn on the user, and there are one or more slaves, each of which is set on a physical entity, and each of the physical entities is used to play a dialogue role;

[0006] The method comprises:

[0007] Understanding the user's intention according to the user's instruction, selecting one or more designated characters from the dialogue characters and waking them up;

[0008] Uploading the ID of the slave machine corresponding to each designated role to the server;

[0009] receiving the greeting of the designated character returned by the server, and sending the greeting to the corresponding slave device for playing;

[0010] Collect user speeches in real time and upload them to the server, so that the server can understand the user's intention based on the user's speech and generate the speech of the designated role through the large model AI;

[0011] Receive the speech of the designated role sent by the server, and send it to the corresponding slave device for playing;

[0012] in,

[0013] Each of the dialogue characters is pre-set with character attributes, which include: character name, character timbre, character task and corresponding slave ID.

[0014] Preferably, the user instruction includes: a voice instruction, a gesture instruction, a key instruction or an approach instruction;

[0015] The method further comprises:

[0016] Real-time monitoring of the distance between the user and each conversational character using Bluetooth, ultra-wideband, or visual recognition positioning technology;

[0017] If the distance between the user and one or more dialogue characters is less than a preset first distance threshold, the approach instruction is triggered.

[0018] Preferably, the step of "understanding the user's intention according to the user instruction, selecting one or more designated roles from the dialogue roles and waking them up" includes:

[0019] When the user instruction is the voice instruction, the voice instruction is uploaded to the server, and the designated role is determined and awakened according to the information returned by the server after using the large model AI reasoning;

[0020] When the user instruction is a gesture instruction or a key instruction, the designated role is selected and awakened according to a preset matching rule between the gesture or key and the role;

[0021] When the user instruction is the approach instruction, the dialogue character whose distance to the user is less than the preset first distance threshold is selected as the designated character and awakened.

[0022] Preferably, the method further comprises:

[0023] If the user instruction is not obtained, then

[0024] Search for one or more target tasks corresponding to the current time according to the preset time-task list;

[0025] Randomly select one of the target tasks as a task to be executed;

[0026] Obtaining the role pre-bound to the task to be executed from the preset time-task list and determining it as the designated role;

[0027] If the distance between the designated character and the user is less than a preset second distance threshold, waking up the designated character;

[0028] Uploading the task to be executed and the ID of the slave corresponding to the designated role to the server, so that the server generates a speech for the designated role using a large model AI algorithm according to the task to be executed and the role attributes of the designated role;

[0029] Receive the speech of the designated role from the server and send it to the corresponding slave device for playback;

[0030] Collect user speeches in real time and upload them to the server, so that the server can understand the user's intention based on the user's speech and generate a new round of speeches for the designated role through the large model AI;

[0031] A new round of speeches of the designated role is received from the server, and sent to the corresponding slave device for playing.

[0032] A second aspect of the present invention provides another large-model AI multi-role collaborative interaction method driven by intent understanding, the method being applicable to a server in a multi-role interaction system, the multi-role interaction system comprising: a host, a slave, and the server;

[0033] The host is worn on the user, and there are one or more slaves, each of which is set on a physical entity, and each of the physical entities is used to play a dialogue role;

[0034] The method comprises:

[0035] Receive the ID of the slave corresponding to each specified role sent by the host;

[0036] generating a greeting message of the designated role and sending the message to the host, so that the host forwards the greeting message of the designated role to the corresponding slave for playback;

[0037] Receiving user speeches uploaded by the host;

[0038] Understanding the user's intention based on the user's speech, and generating the speech of the designated role through the large model AI and sending it to the host, so that the host forwards the speech of the designated role to the corresponding slave for playback;

[0039] in,

[0040] The designated role is one or more, and is selected and awakened by the host from the dialogue roles according to a user instruction;

[0041] Each of the dialogue characters is pre-set with character attributes, which include: character name, character timbre, character task and corresponding slave ID.

[0042] Preferably, the user command includes: a voice command, a gesture command, a key command or an approach command;

[0043] The approach instruction is triggered when the host detects through Bluetooth, ultra-wideband or visual recognition positioning technology that the distance between the user and one or more of the dialogue characters is less than a preset first distance threshold;

[0044] The method further comprises:

[0045] Receive the voice command sent by the host, use the large model AI to infer the user's intention, and then determine the designated role, and send the slave ID corresponding to the designated role to the host, so that the host performs a wake-up operation.

[0046] Preferably, the step of "generating a greeting message for the designated character and sending it to the host" includes:

[0047] Obtaining a preset greeting for each designated role and sending it to the host; or,

[0048] According to the role attributes of each of the designated roles, the large model AI is used to generate a greeting for each of the designated roles and send it to the host.

[0049] Preferably, the step of "understanding the user's intention based on the user's speech, and generating the speech of the designated role through the large model AI and sending it to the host" includes:

[0050] Generate the speech of the designated role using a large model AI algorithm based on the user speech;

[0051] Sending the speeches of the designated roles to the host in sequence;

[0052] Receive a new round of user speeches sent by the host, and generate a new round of speeches for the designated role using a large model AI algorithm based on the new round of user speeches;

[0053] A new round of speeches by the designated role is sent to the host in sequence.

[0054] Preferably, before “generating the speech of the designated role using a large model AI algorithm according to the user speech”, the method further includes:

[0055] Determine the type of target task using a large model AI based on the user's speech;

[0056] If the target task is a learning or game, the server determines the user's completion status based on the learning or game progress and the user's speech, and adjusts the difficulty level of the target task based on the completion status;

[0057] The types of the target tasks include: learning, gaming, chatting or audio playback.

[0058] Preferably, the method further comprises:

[0059] Receive the task to be executed and the ID of the slave corresponding to the specified role sent by the host;

[0060] According to the task to be executed and the role attributes of the designated role, a large model AI algorithm is used to generate a speech of the designated role and send it to the host, so that the host forwards the speech of the designated role to the corresponding slave for playback;

[0061] Receive user speeches from the host, understand user intentions based on the user speeches, and generate a new round of speeches for the designated role through the large model AI and send it to the host, so that the host forwards the new round of speeches for the designated role to the corresponding slave for playback;

[0062] The tasks to be executed and the designated roles are obtained by the host from a preset time-task list according to the current time.

[0063] Preferably, the method further comprises:

[0064] The completion progress of the task to be executed is recorded, and words of guidance and encouragement to the user are generated according to the completion progress, and are sent to the host as a speech of a certain designated role, so that the host forwards the words of guidance and encouragement to the corresponding slave for playback.

[0065] According to a third aspect of the present invention, a computer-readable storage device is provided, storing a computer program that can be loaded by a processor and execute the above method.

[0066] The present invention has the following beneficial effects:

[0067] The intention understanding-driven large-model AI multi-role collaborative interaction method proposed in the present invention wakes up the specified role through understanding the user's instructions, rather than fixed role interaction, and can flexibly adapt to the interaction scenario between multiple roles; sets the role name, role timbre, role task and corresponding slave ID and other attributes for each role, realizes the rapid deployment of role templates, and facilitates the expansion of new roles; based on intention understanding rather than fixed wake-up words (such as "Role A answers"), users can directly express their needs through natural language (such as "My stomach is uncomfortable"), which is in line with human conversation habits. In addition, the user's needs are understood in real time based on the user's speech, and then the speech of each slave (role) is generated by the large model AI without relying on a fixed script. Through the above means, the present invention improves the accuracy and flexibility of the interaction, thereby improving the user experience.

[0068] When the user is detected approaching the character, the approach command is triggered, allowing the character to actively greet the user, which is closer to actual social occasions.

[0069] The character can initiate a conversation based on the preset time-task list, reminding or guiding the user to complete a task at a specified time (such as going out to attend a class reunion at 7 o'clock in the evening, learning an ancient Chinese article every day, etc.). BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 This is a schematic diagram of the scenario of the intention understanding-driven large-scale AI multi-role collaborative interaction in the present invention;

[0071] Figure 2 This is a schematic diagram of the main steps of Example 1 of the large-model AI multi-role collaborative interaction method driven by intention understanding in the present invention;

[0072] Figure 3 This is a schematic diagram of the main steps of Example 2 of the large-model AI multi-role collaborative interaction method driven by intent understanding in the present invention;

[0073] Figure 4 This is a schematic diagram of the main steps of Example 3 of the large-model AI multi-role collaborative interaction method driven by intent understanding in the present invention;

[0074] Figure 5 This is a schematic diagram of the main steps of Example 4 of the large-model AI multi-role collaborative interaction method driven by intention understanding in the present invention. DETAILED DESCRIPTION

[0075] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0076] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0077] It should be noted that, in the description of the present invention, the terms "first" and "second" are merely for the convenience of description, and do not indicate or imply the relative importance of the devices, elements or parameters, and therefore should not be understood as limiting the present invention. In addition, the term "and / or" in the present invention is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this document, unless otherwise specified, generally indicates that the associated objects are in an "or" relationship.

[0078] The method of this embodiment is based on a multi-role interaction system, which includes a host, slaves, and a server. The host is worn by the user; there are one or more slaves, each of which is installed on a physical entity (such as a toy or refrigerator), and each physical entity is used to play a role in the conversation; the server can be a local, remote, or cloud server.

[0079] Figure 1 This is a schematic diagram of the scenario of the intention understanding driven large model AI multi-role collaborative interaction in the present invention. Figure 1 As shown, the black circular host is worn on the user's chest with a lanyard, and the three black square slaves are respectively attached to three toys. The server is set up in the cloud, and the arrows in the figure indicate the data flow.

[0080] Each dialogue character has pre-set attributes, including name, voice, role, and corresponding slave ID. For example, character A's attributes include: name Akang, voice of an adult baritone, role of answering and playing the role of a doctor, and corresponding slave ID 01; character B's attributes include: name Aka, voice of a 7-year-old boy, role of playing a playmate, singing, and telling stories, and corresponding slave ID 02; character C's attributes include: name Lili, voice of an adult mezzo-soprano, role of playing the role of a teacher, and corresponding slave ID 03.

[0081] The host is mainly used to determine which characters the user wants to talk to, thereby waking up the corresponding slaves, uploading the slave ID to the server, and then forwarding the character speech generated by the server to the corresponding slaves; the server is mainly used to use the large model AI to generate the speech of the specified character based on the content uploaded by the host, and send it to the host; the slave is mainly used to respond to the host's wake-up operation and play the audio data sent by the host.

[0082] Figure 2 This is a schematic diagram of the main steps of the first embodiment of the intention understanding driven large model AI multi-role collaborative interaction method of the present invention. The method of this embodiment is applicable to the host in the multi-role interaction system, such as Figure 2 As shown, the method of this embodiment includes steps A10-A50:

[0083] Step A10: The host understands the user's intention according to the user's instruction, selects one or more designated roles from the dialogue roles, and wakes them up.

[0084] The user instructions include: voice instructions, gesture instructions, key instructions or approach instructions (indicating that the user approaches a certain character).

[0085] Specifically, step A10 may include:

[0086] Step A11: When the user instruction is a voice instruction, upload the voice instruction to the server, determine the designated role and wake it up based on the information returned by the server after using the large model AI reasoning, and go to step A20.

[0087] Specifically, after receiving the voice command, the server first converts the voice command into text information, and then uses the large model AI to infer the user's intention based on the text information, and then determines which designated characters the user wants to talk to, and sends the slave ID corresponding to the designated character to the host. The host can then wake up the designated character (actually wake up the slave corresponding to the designated character).

[0088] For example, if a user issues a voice command like "Where is Aka?", the server can use large-scale AI reasoning to determine that the user's designated role is "Aka." Another example is if the voice command is "I feel sick in my stomach," the server can use large-scale AI reasoning to determine that "Akang," playing the role of a doctor, is the designated role.

[0089] Step A12: When the user instruction is a gesture instruction or a key instruction, select a designated role according to a preset matching rule and wake it up, and then go to step A20.

[0090] Combined with video surveillance and image recognition, user gesture commands can be captured. In this case, the preset matching rules could be: whichever character is pointed at is selected as the designated character, or a specific gesture could be used to indicate a conversation with a specific character (for example, raising one finger upwards indicates a conversation with character A, while raising two fingers upwards indicates a conversation with character B).

[0091] The user can also issue key commands through the buttons set on the host to select the "designated character" with whom he wants to talk. In this case, the preset matching rules can be: pre-bind a certain button or button combination to a dialogue character, and when this button or button combination is pressed, the bound dialogue character is selected as the designated character.

[0092] Step A13: When the user instruction is a get-close instruction, a dialogue character whose distance to the user is less than a preset first distance threshold is selected as a designated character and awakened, and then the process goes to step A20.

[0093] In this step, the approach command is triggered when the host detects that the user is approaching one or more of the dialogue characters. If the approach command is present at the same time as a voice command, gesture command, or key command, the program automatically ignores the approach command.

[0094] Step A20: Upload the ID of the slave machine corresponding to each designated role to the server.

[0095] Among them, the server can be a local, remote or cloud server, used to perform tasks such as large model reasoning.

[0096] Step A30: Receive the greeting message of the designated character returned by the server and send it to the corresponding slave device for playback.

[0097] Specifically, the server may obtain a pre-saved preset greeting for each designated character, such as: "Hello! I'm Aka, nice to meet you!".

[0098] The server can also use the large-scale AI model to generate a greeting for each designated character based on their attributes. For example, "I'm Aka, I'm good at storytelling and singing. What would you like to hear?"

[0099] Step A40: Collect user speeches in real time and upload them to the server, so that the server can understand the user's intention based on the user's speech and generate the speech of the specified role through the large model AI.

[0100] Step A50: Receive the speech of the designated role sent by the server, and send it to the corresponding slave machine for playback.

[0101] Depending on the actual application scenario, steps A40-A50 may be repeatedly executed to implement multiple rounds of dialogue between the user and the specified character.

[0102] In an optional embodiment, before step A10, steps A5-A6 of triggering an approach instruction may be further included:

[0103] Step A5: Use Bluetooth, ultra-wideband, or visual recognition positioning technology to monitor the distance between the user and each dialogue character in real time.

[0104] Step A6: If the distance between the user and one or more dialogue characters is less than a preset first distance threshold (such as 1 meter), a move closer instruction is triggered.

[0105] Figure 3 This is a schematic diagram of the main steps of the second embodiment of the intention understanding driven large model AI multi-role collaborative interaction method of the present invention. This embodiment is executed when the host does not obtain the user's instructions, and the user instructions include: voice instructions, gesture instructions, key instructions or proximity instructions. Figure 3 As shown, this embodiment may include steps B10-B80:

[0106] Step B10: Search for one or more target tasks corresponding to the current time according to a preset time-task list.

[0107] Step B20: Randomly select a target task as the task to be executed.

[0108] Step B30: Obtain the role pre-bound to the task to be executed from the preset time-task list and determine it as the designated role.

[0109] Step B40: If the distance between the designated character and the user is less than a preset second distance threshold (eg, 3 meters), the designated character is awakened.

[0110] Step B50: Upload the ID of the slave machine corresponding to the task to be executed and the designated role to the server, so that the server generates a speech for the designated role using a large model AI algorithm based on the task to be executed and the role attributes of the designated role.

[0111] For example, when the task to be performed is "learn English words" and character C is bound, the AI ​​can generate character C's speech based on the previous learning progress: "Today we are going to review words about animal names. How do you say 'panda' in English?"

[0112] Step B60: Receive the speech of the designated character from the server and send it to the corresponding slave machine for playback.

[0113] Step B70: Collect user speeches in real time and upload them to the server, so that the server can understand the user's intention based on the user's speech and generate a new round of speeches for the specified role through the large model AI.

[0114] Step B80: Receive a new round of speeches by the designated role from the server, and send it to the corresponding slave device for playback.

[0115] Depending on the actual application scenario, steps B70-B80 may be repeatedly executed to achieve multiple rounds of dialogue between the user and the specified character.

[0116] For example, the preset time-task list sets two target tasks for 7 o'clock in the evening: playing songs and learning classical Chinese, and pre-bound dialogue character A for playing songs and pre-bound dialogue character B for learning classical Chinese. Then at 7 o'clock in the evening, if the host does not receive the user's instructions, it will query the preset time-task list and find the above two target tasks. If the randomly selected task to be executed is playing songs, as long as the distance between character A and the user is less than the preset second distance threshold, the host will wake up character A, and the server will generate character A's speech "Hello, it's time to listen to songs! Your favorite singer recently released a new song, do you want to listen to it?" The user then speaks "I still want to listen to the song from yesterday", so the server calls out the song segment played yesterday and sends it to the host, which then forwards it to the slave corresponding to character A for playback.

[0117] For example, parents set a tooth-brushing task for their children at 8 o'clock every night. According to steps B10-B80 above, the "doctor" doll will automatically remind them at 8 o'clock in the evening: "Let's play the 'Defeat Tooth Bacteria' game. You can get star rewards if you brush for 3 minutes!" At this time, the host can collect the sound of brushing teeth to know how many minutes the child has brushed. The server then generates a character speech and forwards it to the "doctor" character for playback: "You brushed for 4 minutes today! You have received 10 stars this month. Please keep it up!"

[0118] Figure 4 This is a schematic diagram of the main steps of the third embodiment of the intention understanding driven large model AI multi-role collaborative interaction method of the present invention. The method of this embodiment is applicable to the server in the multi-role interaction system, such as Figure 4 As shown, the method of this embodiment includes steps C10-C40:

[0119] Step C10: The server receives the ID of the slave corresponding to each designated role sent by the host.

[0120] Among them, there are one or more designated roles, which are selected and awakened by the host from the dialogue roles according to user instructions; each dialogue role has pre-set role attributes, which include: role name, role timbre, role task and corresponding slave ID.

[0121] User commands include: voice commands, gesture commands, key commands or approach commands; the approach command is triggered when the host detects through Bluetooth, ultra-wideband or visual recognition positioning technology that the distance between the user and one or more dialogue characters is less than a preset first distance threshold.

[0122] Step C20: Generate a greeting message for the designated character and send it to the host, so that the host forwards the greeting message for the designated character to the corresponding slave for playback.

[0123] Specifically, the method for generating a greeting for a designated role may include: obtaining a preset greeting for each designated role; or, generating a greeting for each designated role using a large model AI based on the role attributes of each designated role.

[0124] Step C30: Receive the user speech uploaded by the host.

[0125] In step C20 above, the designated character proactively greets the user (e.g., "Would you like to listen to a new song today?"). The user may then accept the target task proposed by the designated character (e.g., "Okay"), or may propose other target tasks (e.g., "Let's listen to a story today"). Therefore, the following step C40 needs to understand the user's intention based on the user's speech and analyze what the target task to be performed next is.

[0126] Step C40: understand the user's intention based on the user's speech, and generate the speech of the specified role through the large model AI and send it to the host, so that the host forwards the speech of the specified role to the corresponding slave for playback.

[0127] Specifically, step C40 may include steps C41-C46:

[0128] Step C41: Use the large model AI to determine the type of target task based on the user's speech.

[0129] Step C42: If the target task is learning or gaming, the user's completion status is determined based on the learning or gaming progress and user comments, and the difficulty level of the target task is adjusted based on the completion status.

[0130] Step C43: Generate the speech of the designated role using the large model AI algorithm based on the user's speech.

[0131] Step C44: Send the speech of the designated role to the host in sequence, so that the host forwards it to the corresponding slave for playback.

[0132] Step C45: Receive a new round of user speeches from the host, and generate a new round of speeches for the designated role using the large model AI algorithm based on the new round of user speeches.

[0133] Step C46: Send the new round of speeches of the designated role to the master in sequence, so that the master forwards them to the corresponding slaves for playback.

[0134] In actual application scenarios, steps C45-C46 may be repeatedly executed to achieve multiple rounds of dialogue between the user and the specified character.

[0135] In an optional embodiment, step C5 may be further included before step C10:

[0136] Step C5: Receive the voice command from the host, use the large model AI to infer the user's intention, and then determine the designated role, and send the slave ID corresponding to the designated role to the host, so that the host performs the wake-up operation.

[0137] Specifically, the voice command can be converted into text information first, and then the user intention can be inferred based on the text information using the large model AI to determine which characters the user wants to talk to, thereby determining the "designated role" and sending the slave ID corresponding to the "designated role" to the host.

[0138] Figure 5 This is a schematic diagram of the main steps of the fourth embodiment of the intention understanding driven large model AI multi-role collaborative interaction method of the present invention. The method of this embodiment is applicable to the server in the multi-role interaction system, such as Figure 5 As shown, the method of this embodiment includes steps D10-D40:

[0139] Step D10: The server receives the task to be executed and the ID of the slave corresponding to the specified role sent by the host.

[0140] Step D20: Based on the task to be executed and the role attributes of the designated role, the large model AI algorithm is used to generate the speech of the designated role and send it to the host, so that the host forwards the speech of the designated role to the corresponding slave for playback.

[0141] Step D30: Receive user speeches from the host, understand user intentions based on the user speeches, and generate a new round of speeches for the specified role through the large model AI and send it to the host, so that the host forwards the new round of speeches for the specified role to the corresponding slave for playback.

[0142] The tasks to be executed and the designated roles are obtained by the host from a preset time-task list according to the current time.

[0143] Step D40 records the completion progress of the task to be executed, and generates words to guide and encourage the user based on the completion progress, and sends them to the host as the speech of a designated role, so that the host forwards the words of guidance and encouragement to the corresponding slave for playback.

[0144] For example, if the target task is to learn the ancient Chinese text "Preface to the Expedition", when the user completes the first two paragraphs, the designated role C who plays the "teacher" can provide guidance: "You have learned two paragraphs. Approaching the sages can enlighten our wisdom. Next, let's listen to what talents Zhuge Liang recommended to Liu Chan!"

[0145] Depending on the actual application scenario, steps D30-D40 may be repeated multiple times to achieve multiple rounds of dialogue between the user and the specified character.

[0146] Although the various steps in the above embodiment are described in the above-mentioned order, those skilled in the art will understand that in order to achieve the effect of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reverse order. These simple changes are within the scope of protection of the present invention.

[0147] Based on the above method embodiment, the present invention further provides an embodiment of a computer-readable storage device. The storage device of this embodiment stores a computer program that can be loaded by a processor and execute the method described above.

[0148] The computer-readable storage device may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes.

[0149] Those skilled in the art should be able to appreciate that the method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0150] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is clearly not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent modifications or substitutions to the relevant technical features, and the technical solutions after such modifications or substitutions will fall within the scope of protection of the present invention.

Claims

1. A large-model AI multi-role collaborative interaction method driven by intent understanding, characterized by: The method is applicable to a host in a multi-role interaction system, wherein the multi-role interaction system comprises: the host, a slave and a server; The host is worn on the user, and there are one or more slaves, each of which is set on a physical entity, and each of the physical entities is used to play a dialogue role; The method comprises: Understanding the user's intention according to the user's instruction, selecting one or more designated characters from the dialogue characters and waking them up; Uploading the ID of the slave machine corresponding to each designated role to the server; receiving the greeting of the designated character returned by the server, and sending the greeting to the corresponding slave device for playing; Collect user speeches in real time and upload them to the server, so that the server can understand the user's intention based on the user's speech and generate the speech of the designated role through the large model AI; Receive the speech of the designated role sent by the server, and send it to the corresponding slave device for playing; in, Each of the dialogue characters is pre-set with character attributes, which include: character name, character timbre, character task and corresponding slave ID.

2. The intention understanding driven large model AI multi-role collaborative interaction method according to claim 1 is characterized in that: The user command includes: a voice command, a gesture command, a key command or an approach command; The method further comprises: Real-time monitoring of the distance between the user and each conversation character through Bluetooth, ultra-wideband or visual recognition positioning technology; If the distance between the user and one or more dialogue characters is less than a preset first distance threshold, the approach instruction is triggered.

3. The intention understanding driven large model AI multi-role collaborative interaction method according to claim 2 is characterized in that: The step of "understanding the user's intention according to the user's instruction, selecting one or more designated characters from the dialogue characters and waking them up" includes: When the user instruction is the voice instruction, the voice instruction is uploaded to the server, and the designated role is determined and awakened according to the information returned by the server after using the large model AI reasoning; When the user command is a gesture command or a key command, the designated role is selected and awakened according to a preset matching rule; When the user instruction is the approach instruction, the dialogue character whose distance to the user is less than the preset first distance threshold is selected as the designated character and awakened.

4. The intention understanding driven large model AI multi-role collaborative interaction method according to claim 1 is characterized in that: The method further comprises: If the user instruction is not obtained, then Search for one or more target tasks corresponding to the current time according to the preset time-task list; Randomly select one of the target tasks as a task to be executed; Obtaining the role pre-bound to the task to be executed from the preset time-task list and determining it as the designated role; If the distance between the designated character and the user is less than a preset second distance threshold, waking up the designated character; Uploading the task to be executed and the ID of the slave corresponding to the designated role to the server, so that the server generates a speech for the designated role using a large model AI algorithm according to the task to be executed and the role attributes of the designated role; Receive the speech of the designated role from the server and send it to the corresponding slave device for playback; Collect user speeches in real time and upload them to the server, so that the server can understand the user's intention based on the user's speech and generate a new round of speeches for the designated role through the large model AI; A new round of speeches of the designated role is received from the server, and sent to the corresponding slave device for playing.

5. A large-model AI multi-role collaborative interaction method driven by intention understanding, characterized by: The method is applicable to a server in a multi-role interactive system, wherein the multi-role interactive system comprises: a host, a slave and the server; The host is worn on the user, and there are one or more slaves, each of which is set on a physical entity, and each of the physical entities is used to play a dialogue role; The method comprises: Receive the ID of the slave corresponding to each specified role sent by the host; generating a greeting message of the designated role and sending the message to the host, so that the host forwards the greeting message of the designated role to the corresponding slave for playback; Receiving user speeches uploaded by the host; Understanding the user's intention based on the user's speech, and generating the speech of the designated role through the large model AI and sending it to the host, so that the host forwards the speech of the designated role to the corresponding slave for playback; in, The designated role is one or more, and is selected and awakened by the host from the dialogue roles according to a user instruction; Each of the dialogue characters is pre-set with character attributes, which include: character name, character timbre, character task and corresponding slave ID.

6. The intention understanding driven large model AI multi-role collaborative interaction method according to claim 5 is characterized in that: The user command includes: a voice command, a gesture command, a key command or an approach command; The approach instruction is triggered when the host detects through Bluetooth, ultra-wideband or visual recognition positioning technology that the distance between the user and one or more of the dialogue characters is less than a preset first distance threshold; The method further comprises: Receive the voice command sent by the host, use the large model AI to infer the user's intention, and then determine the designated role, and send the slave ID corresponding to the designated role to the host, so that the host performs a wake-up operation.

7. The intention understanding driven large model AI multi-role collaborative interaction method according to claim 5 is characterized in that: The step of "generating a greeting message for the designated character and sending it to the host" includes: Obtaining a preset greeting for each designated role and sending it to the host; or, According to the role attributes of each of the designated roles, the large model AI is used to generate a greeting for each of the designated roles and send it to the host.

8. The intention understanding driven large model AI multi-role collaborative interaction method according to claim 5 is characterized in that: The step of "understanding the user's intention based on the user's speech, and generating the speech of the designated role through the large model AI and sending it to the host" includes: Generate the speech of the designated role using a large model AI algorithm based on the user speech; Sending the speeches of the designated roles to the host in sequence; Receive a new round of user speeches sent by the host, and generate a new round of speeches for the designated role using a large model AI algorithm based on the new round of user speeches; A new round of speeches by the designated role is sent to the host in sequence.

9. The intention understanding driven large model AI multi-role collaborative interaction method according to claim 8 is characterized in that: Before "generating the speech of the designated role using a large model AI algorithm based on the user speech", the method further includes: Determine the type of target task using a large model AI based on the user's speech; If the target task is a learning or game, the user's completion status is determined based on the learning or game progress and the user's speech, and the difficulty level of the target task is adjusted based on the completion status; The types of the target tasks include: learning, gaming, chatting or audio playback.

10. The intention understanding driven large model AI multi-role collaborative interaction method according to claim 5 is characterized in that: The method further comprises: Receive the task to be executed and the ID of the slave corresponding to the specified role sent by the host; According to the task to be executed and the role attributes of the designated role, a large model AI algorithm is used to generate a speech of the designated role and send it to the host, so that the host forwards the speech of the designated role to the corresponding slave for playback; Receive user speeches from the host, understand user intentions based on the user speeches, and generate a new round of speeches for the designated role through the large model AI and send it to the host, so that the host forwards the new round of speeches for the designated role to the corresponding slave for playback; The tasks to be executed and the designated roles are obtained by the host from a preset time-task list according to the current time.

11. The intention understanding driven large model AI multi-role collaborative interaction method according to claim 10 is characterized in that: The method further comprises: The completion progress of the task to be executed is recorded, and words of guidance and encouragement to the user are generated according to the completion progress, and are sent to the host as a speech of a certain designated role, so that the host forwards the words of guidance and encouragement to the corresponding slave for playback.

12. A computer-readable storage device, characterized in that: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Multi-role interactive method based on voice recognition

    CN104932862A

  • Multi-role intelligent chatting method and system

    CN105975622A

  • Multi-role intelligent sound box partner system

    CN111696516A

  • Voice interaction device control method, server and voice interaction devices

    CN113470634A

  • Same-robot conversation control method and device, computer equipment and storage medium

    CN115757748A