System and device based on interactive learning of multiple doll agents

The multi-doll intelligent agent interactive learning system enables intelligent collaboration and division of labor among multiple roles, dynamically adjusts tone parameters, solves the problem of insufficient interactivity in existing devices, and improves the interactive quality and learning depth of children's educational devices.

CN120913461APending Publication Date: 2025-11-07BEIJING LINGJI TIANCI TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511034978.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing children's educational devices lack multi-role intelligent interaction mechanisms, cannot create a group discussion and exchange atmosphere, have insufficient interactivity, lack guided dialogue capabilities, cannot generate personalized learning topics based on user interests, and have a high threshold for children's participation.

Method used

The system adopts a multi-doll intelligent agent interactive learning approach, which realizes NFC sensing and voice interaction through physical dolls, a central base and service terminal. It generates personalized chat topics based on user profiles, sets role attributes, uses a large language model to recognize user intent, and dynamically adjusts voice parameters to achieve multi-doll intelligent agent collaboration and role division.

Benefits of technology

It improves the quality and depth of interaction, lowers the barrier to participation, enhances immersion and fun, cultivates critical thinking through multi-role dialogue, and ensures logical coherence and educational consistency in dialogue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913461A_ABST
    Figure CN120913461A_ABST
Patent Text Reader

Abstract

The invention discloses a system and device based on interactive learning of multiple doll intelligent agents, and belongs to the technical field of large model intelligent agents, and the system comprises an entity doll which is provided with an NFC label and is used for being placed on a central base and establishing NFC connection with the central base, and the central base carries out NFC induction and recognition on the entity doll and integrates radio and loudspeaker functions. The method comprises the following steps: acquiring and playing an audio signal, performing voice or touch interaction with a user, and establishing communication connection with a service terminal, wherein the service terminal selects a chat theme in combination with a portrait of the user, establishes a virtual chat room, performs role attribute setting on an identified entity doll, identifies the intention of the user, and performs chat interaction with the user; the interactive discussion is output to the central base for playing in combination with the chat theme, so that a plurality of intelligent agents can be supported to deeply discuss around a specific theme, the logic continuity and education consistency of multi-role interactive contents are ensured, real group discussion experience is simulated, and an interactive atmosphere similar to group discussion is formed, thereby improving the interactive quality and learning depth.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of large model intelligent agents, and in particular relates to a system and device based on multi-doll intelligent agent interactive learning. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, education intelligent hardware products are undergoing a profound transformation. Core AI technologies such as speech recognition, natural language processing, and speech synthesis have been widely applied in the field of children's education, enabling intelligent devices to engage in increasingly complex language interactions with users. A large number of children's education products and AI interactive dialogue hardware that focus on audio playback have emerged in the market. These products typically provide story playback, nursery rhyme singing, and basic voice question and answer functions for children through built-in speakers and trigger mechanisms combined with physical dolls, greatly enriching children's early learning experience. Although existing products have made some progress, their interactive modes still have significant limitations.

[0003] Such devices generally support content triggering and playback for only one doll at a time, such as a single doll triggering a story or simple voice interaction with one question and one answer. This design requires a certain level of initiative expression ability from children and is not suitable for all children, especially introverted or slow language development children.

[0004] Current systems lack guided dialogue capabilities and often can only passively wait for user questions, resulting in insufficient interactivity and limited learning depth. Most systems require users to ask explicit questions for interaction, which is too high a threshold for children users who have not yet developed systematic expression abilities. There is a lack of "active interaction," "guided questioning," and "thematic companionship" in the functional design.

[0005] Moreover, such products generally lack intelligent interaction mechanisms between multiple characters, and cannot form a small group discussion atmosphere. In real-time dialogue among multiple characters, there is no special dialogue scheduling mechanism, which can easily lead to interference phenomena such as talking over each other, repeated speaking, and meaningless arguments, affecting user experience and making it difficult to support complex role division such as "moderator," "opponent," and "supplementer." At the same time, they lack the ability to automatically generate topics based on user interests, profiles, or knowledge levels, and cannot achieve truly personalized learning and interactive experiences. SUMMARY

[0006] To solve the above problems and technical defects, the application adopts the following technical solution: a system for multi-doll intelligent agent interactive learning, comprising: a physical doll, a central base, and a service terminal.

[0007] The physical doll is provided with an NFC tag and is placed on the central base to establish an NFC connection with the central base.

[0008] The central base is used for NFC sensing and identification of the entity doll, integrates the functions of radio and speaker, realizes the collection and playing of audio signals, carries out voice or touch interaction with the user, establishes communication connection with the service terminal, and realizes real-time communication.

[0009] The service terminal is used for selecting a chat topic in combination with the portrait of the user, establishing a virtual chat room, setting the role attribute of the identified entity doll, recognizing the intention of the user according to the collected voice information, and outputting interactive discussion to the central base for playing in combination with the chat topic.

[0010] Preferably, the role attribute of the entity doll includes static attribute and dynamic attribute.

[0011] The static attribute includes name, avatar, tone, language code, world view, role biography, and knowledge base.

[0012] The dynamic attribute includes the positioning and responsibility of the entity doll in the virtual chat room, and the changes of emotion, tone, and speed.

[0013] The dynamic attribute is set according to the learning topic and the virtual chat room, and the dynamic attribute of the same entity doll is different in different chat rooms.

[0014] Preferably, the establishment of the virtual chat room includes:

[0015] According to the historical learning portrait, interest preference, knowledge mastery level of the user, and the education target preset by the parents, a suitable personalized learning topic is dynamically generated or filtered through the built-in intelligent recommendation algorithm.

[0016] A unique prompt word is preset for each entity doll selected by the user, each prompt word corresponds to a system prompt word in the large language model, and each entity doll is given a distinctive tone expression.

[0017] The role definition of the entity doll is set by professional content teachers and copyright parties, a role allocation mechanism is established, and the entity doll is dynamically positioned in combination with the role definition and the personalized learning topic.

[0018] According to the promotion of the current learning topic, the emotional tendency of the dialogue content, and the role positioning of the entity doll, the tone parameters of the entity doll are dynamically adjusted in real time.

[0019] Further, the virtual chat room is provided with a doll agent selection mechanism, which is used for accurately matching the doll agent most suitable for responding to the user's request or promoting the dialogue process in an environment where multiple entity dolls coexist.

[0020] Automatic promotion when the user is eavesdropping and intelligent response when the user actively interrupts, when the user chooses to interrupt the current dialogue between dolls, the agent selection mechanism will additionally consider the factor based on the intention recognition result;

[0021] The doll agent selection mechanism has multiple agent matching algorithms built in, which automatically select the most suitable doll agent to handle the current task or respond to the user's intention, combined with shared context information and the role positioning of doll agents within the group in the chat room.

[0022] Further, the agent matching algorithm includes:

[0023] In the initial stage of personalized learning topic divergent discussion, a doll agent is randomly selected from all doll agents in the chat room to speak;

[0024] According to the preset order or dynamically allocated order, the next speaker is selected from the doll agents in the chat room, so that each doll has the opportunity to express their ideas or introduce initial information in turn;

[0025] The current shared context and the role information of all dolls in the group are encapsulated into a request and sent to the core large language model for decision-making. The large language model intelligently judges and recommends the most suitable doll agent for the next speaker;

[0026] When the user's intention recognition result explicitly specifies a virtual chat room role to reply, the system will directly select the specified doll agent.

[0027] Further, after selecting the doll agent, the system will invoke the doll agent, combine the current personalized learning topic progress, the role and emotion settings of the doll agent, and the historical dialogue context to execute specific dialogue tasks and generate corresponding intelligent responses.

[0028] Further, during the dialogue generation process of the doll agent, the voice synthesis parameters of the selected doll will be dynamically adjusted according to the emotional tendency of the current dialogue content, the role setting and positioning of the doll, the stage characteristics of the learning topic, and the overall rhythm of the dialogue process.

[0029] Further, the identification of the user's intention refers to receiving the user's voice or text input, performing high-level semantic understanding and classification through a large language model, and identifying the purpose type of the user's dialogue this time.

[0030] Further, the purpose type of the user's intention includes:

[0031] The user actively expresses their views or asks questions about the current personalized learning topic;

[0032] The user specifies a specific entity doll in the virtual chat room to respond or explain;

[0033] The user expresses interest in exploring the current topic in depth, and the entity doll provides more details or related information;

[0034] The user wants to end the discussion of the current personalized learning topic and switch to a new personalized learning topic;

[0035] The user encounters difficulties in the interaction and seeks guidance or hints from the entity doll.

[0036] A device for multi-doll intelligent agent interactive learning, comprising a service processor and a distributed memory, the service processor being connected to the memory, the distributed memory storing a service self-management program configured to store machine-readable instructions, the service processor executing the service self-management program, and the instructions being executed by the processor to implement a multi-doll intelligent agent interactive learning system as described above.

[0037] Compared with the prior art, the application has the following advantages:

[0038] (1) The application realizes intelligent collaboration and role division among multiple dolls by establishing a virtual chat room, and through a chat room type multi-role interaction management mechanism, it can automatically allocate speaking order and control dialogue rhythm to simulate real group discussion experience and form an interactive atmosphere similar to group discussion, thereby improving interaction quality and learning depth;

[0039] (2) The application allows users to join the dialogue flexibly in a listening state through user intent recognition and real-time interjection mechanism, reduces the participation threshold, and uses multi-voice speech synthesis and mixing processing technology to give each role a unique voice line, thereby improving the interest and recognition of the interactive process, enhancing the sense of immersion, guiding children to naturally develop curiosity, learning "how to ask good questions", and cultivating critical thinking through conflicts, supplements and follow-up questions in multi-role dialogue;

[0040] (3) The application realizes the basis of multi-agent dialogue coherence, collaborative logic and seamless collaboration among agents according to the dialogue content of the current virtual chat room, the confirmed user intent, the user portrait data and the learning topic being discussed, and through real-time synchronization and sharing of context information, supports multiple intelligent agents to explore a specific topic in depth, and ensures the logical coherence and educational consistency of multi-role interaction content. BRIEF DESCRIPTION OF DRAWINGS

[0041] In the drawings:

[0042] Figure 1 is a system structure schematic diagram of an embodiment of the application;

[0043] Figure 2A flowchart of an embodiment of the present application;

[0044] Figure 3 A structural diagram of a multi-doll agent management platform of an embodiment of the present application;

[0045] Figure 4 A multi-modal interactive service diagram of an embodiment of the present application;

[0046] Figure 5 A one-time interrupt processing flowchart of an embodiment of the present application;

[0047] Figure 6 A data processing flowchart of an embodiment of the present application

[0048] Figure 7 A device structural diagram of an embodiment of the present application. DETAILED DESCRIPTION

[0049] To make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments of the present application. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0050] Embodiment 1

[0051] As shown in Figure 1 , a system based on multi-doll agent interactive learning includes: a physical doll, a central base and a service terminal;

[0052] The physical doll is provided with an NFC tag and is used to be placed on the central base to establish an NFC connection with the central base;

[0053] The central base is used to perform NFC induction and recognition on the physical doll, integrates a radio and a loudspeaker function, realizes collection and playing of audio signals, performs voice or touch interaction with a user, establishes a communication connection with the service terminal, and realizes real-time communication;

[0054] The service terminal is used to select a chat theme in combination with a user's portrait, establish a virtual chat room, set a role attribute of the recognized physical doll, recognize a user's intention according to collected voice information, and output interactive discussion to the central base for playing in combination with the chat theme.

[0055] The role attribute of the physical doll includes: static attributes and dynamic attributes;

[0056] The static attributes include: a name, an avatar, a tone, a language code, a world view, a role biography and a knowledge base;

[0057] Dynamic attributes include the positioning and responsibility of entity dolls in the virtual chat room, as well as changes in mood, tone, and speech speed;

[0058] Dynamic attributes are set according to learning topics and virtual chat rooms, and the same entity doll has different dynamic attributes in different chat rooms.

[0059] The establishment of the virtual chat room includes:

[0060] According to the user's historical learning profile, interest preference, knowledge mastery level and parent's pre-set education goal, through the built-in intelligent recommendation algorithm, dynamically generate or filter suitable personalized learning topics;

[0061] Each entity doll selected by the user is pre-set with a unique prompt word, each prompt word corresponds to a system prompt word in the large language model, and each entity doll has a distinctive voice expression;

[0062] The role definition of the entity doll is set by professional content teachers and copyright parties, and a role allocation mechanism is established to dynamically position the entity doll based on the role definition and personalized learning topics;

[0063] According to the progress of the current learning topic, the emotional tendency of the dialogue content and the role positioning of the entity doll, its voice parameters are dynamically adjusted in real time.

[0064] The virtual chat room is provided with a doll intelligent body selection mechanism, which is used to accurately match the doll intelligent body most suitable for responding to user requests or promoting the dialogue process in an environment where multiple entity dolls coexist;

[0065] When the user is listening, the automatic promotion and intelligent response when the user actively interrupts, when the user chooses to interrupt the dialogue between the current dolls, the intelligent body selection mechanism will additionally consider the factors based on the intention recognition results;

[0066] The doll intelligent body selection mechanism is built-in with multiple intelligent body matching algorithms, which automatically selects the doll intelligent body most suitable for processing the current task or responding to user intentions by combining shared context information and the role positioning of doll intelligent bodies within the chat room.

[0067] Intelligent body matching algorithms include:

[0068] In the initial stage of the divergent discussion of the personalized learning topic, a doll intelligent body is randomly selected from all doll intelligent bodies in the chat room to speak;

[0069] According to the preset order or dynamically allocated order, the next doll intelligent body is selected from the doll intelligent bodies in the chat room to speak, so that each doll has the opportunity to express their own ideas or introduce initial information;

[0070] The current shared context and the role information of all dolls in the group are encapsulated into a request and sent to the core large language model for decision-making. The large language model intelligently judges and recommends the doll agent most suitable for the next speaker;

[0071] When the user's intention recognition result explicitly specifies a virtual chat room role to reply, the system will directly select the specified doll agent.

[0072] After selecting the doll agent, the system will invoke the doll agent, combine the current personalized learning theme progress, the role emotion setting of the doll agent, and the historical dialogue context, execute the specific dialogue task, and generate the corresponding intelligent response.

[0073] During the dialogue generation process of the doll agent, the voice synthesis parameters of the selected doll will be dynamically adjusted according to the emotional tendency of the current dialogue content, the role setting and positioning of the doll, the stage characteristics of the learning theme, and the overall rhythm of the dialogue process.

[0074] User intent recognition refers to receiving user voice or text input, performing high-level semantic understanding and classification through a large language model, and identifying the purpose type of the user's current dialogue.

[0075] The purpose type of the user's intent includes:

[0076] The user actively expresses his / her opinion or asks questions about the current personalized learning theme;

[0077] The user specifies a specific entity doll in the virtual chat room to respond or explain;

[0078] The user expresses interest in in-depth exploration of the current topic, and the entity doll provides more details or related information;

[0079] The user wants to end the discussion of the current personalized learning theme and switch to a new personalized learning theme;

[0080] The user encounters difficulties in the interaction and seeks guidance or hints from the entity doll.

[0081] Embodiment 2

[0082] As shown in Figure 2 , the central base establishes a communication link with the service terminal. When the user (child) places the doll on the base, the service terminal establishes a multi-agent virtual chat room, selects a chat theme based on the user's profile, and conducts interactive discussions on the theme, forming a natural and educational learning atmosphere.

[0083] The central base supports multiple smart dolls (doll A, doll B, doll C, etc.) to establish a connection with the central base through NFC sensing technology. The base integrates radio and speaker functions to realize audio signal collection and playback.

[0084] The user interacts with the system through voice or touch, and the system can recognize the user's participation intention, supporting the user to join the multi-doll dialogue at any time in the listening state. The base device establishes a data transmission link with the cloud server through a network communication protocol to realize real-time communication between local hardware and cloud intelligent services.

[0085] The service terminal comprises:

[0086] Multi-modal interactive service: responsible for processing user input, multi-agent virtual chat room, agent selection, generating multi-role dialogue content according to shared context, and managing theme dialogue progress.

[0087] Multi-doll agent management platform: realizes adding, deleting, modifying and inquiring of multiple agents, and role attribute setting; the setting of each virtual agent corresponds to an entity doll of hardware.

[0088] The multi-doll agent management platform is as shown in Figure 3 and Figure 6 .

[0089] WSS / WebRTC on the left side, real-time communication service is mainly multi-modal interactive service.

[0090] Https interface service on the right side, mainly multi-doll agent management background service interface.

[0091] User interaction layer: children's toys are products provided for C-end users; the agent management background is an internal management system.

[0092] Backend application layer: as the core of the system, responsible for processing all business logic and data flow.

[0093] Protocol gateway (interface layer): the unified entrance of the system, providing services for the front-end application.

[0094] RESTful API: based on HTTPS protocol, processing non-real-time, request-response mode interaction, such as user login, data query, configuration management, etc.

[0095] WebSocket API: provides a persistent two-way communication channel, specifically handles real-time data flow, such as uploading voice stream and issuing synthesized voice, is the key to realize low-latency real-time interaction.

[0096] Real-time Services: including Chatroom Service, Dialogue Service, and Audio Service. Chatroom Service is responsible for the creation, deletion, modification, and query of chat rooms and the management of members in chat rooms. Dialogue Service is responsible for orchestrating the entire dialogue process, managing the conversation state, and interacting with downstream LLM and TTS services. Audio Service is responsible for receiving and processing real-time audio streams uploaded by the front end and working with Voice Activity Detection (VAD) and Speech-to-Text (STT) services.

[0097] Manager Services: including device management, agent management, model management, and timbre management. Each module has a single responsibility and is responsible for handling business logic in its respective field, providing API interfaces for the background.

[0098] AI Capability Abstraction Layer: This layer completely decouples the upper-layer business logic from the lower-layer specific AI technology implementation by defining unified interfaces (VadProvider for VAD Voice Activity Detection, SttProvider for STT Speech Recognition, ChatProvider for LLM Large Language Model, TtsProvider for TTS Speech Synthesis). Through the Factory Pattern, the system can dynamically select and instantiate AI service implementations from different vendors (such as OpenAI, Alibaba Cloud, Tencent Cloud, and Volcano Engine) at runtime, which makes it easy to replace or add new AI service providers without modifying the core code, with high flexibility and scalability.

[0099] Security & Monitoring Support Layer: provides basic technical support, data analysis, and security risk control for the application layer.

[0100] Security Risk Control: text risk detection for large model input and output, interfacing with cloud service vendor capabilities.

[0101] Process Audit: standardized management and supervision of business approval processes.

[0102] IOT Platform: This module provides access, management, data collection, and collaboration services for IoT devices. It is responsible for device identity authentication, connection maintenance, data transmission, and remote control, supporting the interconnection of agents and devices.

[0103] CDN & Edge: CDN accelerates user access speed by distributing content to nodes around the world. Edge computing reduces latency and reduces core network load by processing data near the source.

[0104] Network: It is responsible for providing stable, efficient and secure network connection services, including internal and external network planning, configuration, management and monitoring, ensuring smooth data transmission between service modules, and providing network security protection mechanisms (such as firewalls and intrusion detection) to ensure the overall communication quality and security of the system.

[0105] Big data analysis: Extracting valuable information and patterns from complex user data to provide data support and decision-making basis for business topics and user historical portraits, improving the intelligence and operational efficiency of the system.

[0106] Highly decoupled and scalable system architecture: Through the AI capability abstraction layer, the decoupling of business logic and specific AI service implementation is achieved, making it easy and fast to replace or add AI service providers, greatly improving the flexibility and maintainability of the system.

[0107] Multi-modal interactive services such as Figure 4 as shown.

[0108] Virtual chat room (learning interactive group) is established based on multiple physical dolls selected by the user, aiming to conduct in-depth discussions around specific learning topics.

[0109] The generation of learning topics has high personalization and adaptability. The system can dynamically generate or filter appropriate learning topics based on the user's (child's) historical learning portrait, interest preferences, knowledge mastery level, and parents' pre-set education goals through built-in intelligent recommendation algorithms, ensuring the relevance, interest and challenge of learning content, thereby maximizing the children's interest in learning.

[0110] The user (child) invites the physical doll into the virtual chat room by physically placing or recognizing (for example, through NFC technology) it on the device. This embodied interaction combines abstract agents with familiar doll images for children, greatly enhancing the immersion and closeness of the interactive experience. Users have the right to choose (for example, doll A, doll B, doll C, etc.) to participate in the chat room, which means they can freely combine different roles according to learning topics or personal preferences. Each selected doll is pre-set with unique prompt settings that directly correspond to system prompts in the large language model, ensuring that the agent always maintains its core role attributes and personality in the conversation. In addition, each doll is also given a unique timbre expression, making it visually distinct in the auditory sense.

[0111] Each doll is not just a simple dialogue carrier, but also has carefully designed role definitions set by professional content teachers and copyright parties, ensuring compliance with children's education rules and copyright requirements, and matching the doll's appearance, background story, etc. In the virtual chat room, the doll's role will be dynamically positioned according to the current learning theme to ensure the diversity and interactivity of the discussion. For example, if the learning theme is "debate", the system will intelligently assign the selected doll to "pro", "con" and "moderator" roles. If the theme is "free discussion", there may be "thinker", "supporter" and "creative mentor" positions. This dynamic role allocation mechanism, based on the doll's inherent role definition and combined with the theme characteristics, aims to simulate real multi-party discussion scenarios, thereby enhancing the interest and depth of interaction.

[0112] To further enhance the realism and emotional expression of interaction, advanced speech synthesis (TTS) capabilities of cloud service providers are fully utilized. These TTS services support multi-emotion parameter settings (e.g. happy, sad, surprised, angry, etc.) and fine-grained adjustments such as speech rate, pitch, volume, etc. The system dynamically adjusts the voice parameters in real time based on the progress of the current learning theme, the emotional tendency of the conversation content, and the role positioning of the doll. For example, when the doll plays the role of "criticizer" and asks questions, its voice may be slightly firm; when it plays the role of "supporter" and expresses praise, it may have a pleasant emotion. This dynamic voice emotion decision mechanism greatly enhances the expressiveness, naturalness and recognizability of the doll's expression, making children more immersed in interaction with the doll.

[0113] The user intent recognition module is the key to the system's understanding of user needs. This module is responsible for receiving user voice or text input and performing advanced semantic understanding and classification through large language models to accurately identify the user's purpose for this conversation. This allows the system to upgrade from a simple question-and-answer mode to a more intelligent and responsive interactive experience.

[0114] User intent recognition includes but is not limited to the following types:

[0115] Participate in chat interaction: The user wants to actively express their views, ask questions or make additions to the current topic.

[0116] Specify agent response: The user explicitly requests a specific doll agent in the chat room to respond or explain.

[0117] Topic in-depth / expansion: The user expresses interest in further exploring the current topic and wants the agent to provide more details or related information.

[0118] Topic termination / switch: The user wants to end the discussion of the current topic or switch to a new learning theme.

[0119] Seeking help / guidance: The user encounters difficulties in the interaction and seeks guidance or hints from the doll agent.

[0120] The intelligent doll agent selection mechanism can accurately match the agent most suitable for responding to user requests or promoting the progress of the conversation in a multi-doll coexistence environment. This mechanism is mainly applied in two core scenarios: automatic promotion when the user is listening, and intelligent response when the user actively interrupts. When the user chooses to interrupt the current conversation between dolls, the agent selection mechanism will additionally consider factors based on intent recognition results, combined with shared context information (i.e., the conversation history of all dolls and users before) and the specific positioning of doll agents within the chat room (e.g., moderator, creative mentor, critic, etc.). The system automatically selects the most suitable doll agent to handle the current task or respond to user intent through a series of Agent Matching algorithms.

[0121] The Agent Matching algorithm includes:

[0122] Random Matching: Randomly select a doll agent from all doll agents in the chat room to speak. This mode is suitable for the initial stage of theme divergent discussion, encouraging diverse viewpoints and thought collisions, maintaining the randomness and interest of the discussion.

[0123] Round-Robin Matching: Select the next doll agent to speak in the chat room according to a pre-set or dynamically assigned order. This mode is commonly used at the beginning of the theme to ensure that each doll has the opportunity to express their ideas or introduce initial information, ensuring the fairness and structure of the discussion.

[0124] LLM-Driven Matching: Encapsulate the current shared context and the role information of all dolls in the group (including their system prompt words and dynamic positioning) into a request and send it to the core large language model for decision-making. The large model will intelligently judge and recommend the most suitable doll agent for the next speaker based on its powerful reasoning ability. This matching method is particularly suitable for scenarios where user intent recognition expects in-depth topic development, effectively guiding the conversation to a deeper level and ensuring the professionalism and relevance of the response.

[0125] Explicit Matching: When the intent recognition result of the current step clearly indicates that the user has specified a chat room role to respond, the system will directly select the specified doll agent. This mode gives the user direct control over the conversation process.

[0126] Once the appropriate doll agent is selected, the system will invoke the agent to execute specific dialogue tasks in conjunction with the current learning topic progress, the doll's character emotional setting, and the historical dialogue context, generating appropriate intelligent responses. During the doll dialogue generation process, the tone and emotion transformation decision is the core step to achieve highly personalized interaction. The system will dynamically adjust the selected doll's speech synthesis parameters based on the following factors:

[0127] Emotional tendency of the current dialogue content: analyze the sentiment polarity (positive, negative, neutral) and specific emotions (such as surprise, confusion, affirmation) of the generated text, and match the corresponding tone and emotion mode.

[0128] Character setting and positioning of the doll: different characters (such as critics, supporters, moderators) may have different emotional expression methods when expressing the same content. The system will fine-tune according to the character characteristics.

[0129] Stage characteristics of the learning topic: when introducing new concepts, an encouraging or guiding tone may be used; when discussing controversial points, a neutral or slightly thoughtful tone may be used.

[0130] Overall rhythm of the dialogue flow: combine speech speed and pauses to ensure natural and smooth emotional expression.

[0131] Through these decision factors, the system can generate voice output that is rich in emotion, natural and lively, and consistent with the situation, greatly improving the children's interactive experience.

[0132] All historical records of the chat room dialogue will be stored as memories, which will be saved for a long time with the user's authorization, including every interaction between the user and all doll agents, as well as internal discussions among doll agents. These memories provide rich contextual information for subsequent intelligent classification selection, topic advancement, and personalized recommendations. Through continuous learning and analysis of memories, the system can better understand the user's long-term interests, knowledge gaps, and learning progress, thereby providing a more targeted and coherent learning experience.

[0133] The system pays special attention to the interaction design in the user (child) listening scenario. When the user is in a listening state without any input in the chat room, the above-mentioned doll agent selection and dialogue generation process will be executed in a loop, allowing the dolls to continuously interact around the learning topic, forming a lively "multi-doll interaction" or "group discussion" atmosphere. This design aims to provide a low-pressure language input environment for introverted or slow language development children, allowing them to become familiar with the dialogue pattern and learning content in a subtle way.

[0134] In order to avoid infinite loops and promote user participation, some hard rules and active guidance mechanisms will be added to the code logic:

[0135] Timing counter-question: After the doll's continuous interaction for X rounds (configurable parameters), the system will actively counter-question the user (child) about the topic, such as: "Xiaoming, do you have anything to say after listening to so much?" or "What do you think of this point of view?"

[0136] Guidance and suspension: If the user still has no input after the system actively counter-questions, the overall dialogue flow can enter the topic conclusion phase, with the host character making a summary speech and waiting for the user's new instructions or topic selection, which ensures the coherence of the interaction and avoids ineffective idling.

[0137] Encourage user participation: These active counter-question mechanisms aim to lower the threshold for user participation and encourage children to become active participants, thereby truly integrating into multi-doll interactive learning.

[0138] The processing flow of the user (child) interrupting once is shown in Figure 5

[0139] Precondition: During the device startup phase, the system actively establishes and maintains a long-term communication connection (WebSocket) with the backend service, laying the foundation for subsequent real-time interaction.

[0140] Audio acquisition and transmission: When the user starts voice interaction, the local microphone collects audio streams in real time, and through the established WebSocket connection, continuously and efficiently transmits them to the backend audio processing service.

[0141] Voice intelligent processing and decision-making: After the backend audio service receives the audio stream, it accurately identifies the valid speech segment through the Voice Activity Detection (VAD) module, then the Automatic Speech Recognition (ASR) module converts the speech into text, and the system intelligently analyzes the user's intention based on the recognition results and decides the exclusive agent for this round of interaction (see the doll agent selection mechanism described above).

[0142] Agent response and speech synthesis: The recognized text will drive the selected agent to generate a response, and this text will then be sent to the Text-to-Speech (TTS) module, which will synthesize it into a speech stream based on the pre-set unique voice of the agent.

[0143] Real-time voice issuance and playback: The final synthesized speech stream is pushed to the user device in real time through the network downlink channel in a low-latency manner, and the device receives and plays the speech immediately, thereby seamlessly completing a complete intelligent interaction.

[0144] Embodiment 3

[0145] As shown in Figure 7 ​As shown, an apparatus based on multi-puppet intelligent agent interactive learning, the apparatus comprises a service processor and a distributed memory, the service processor is connected with the memory, the distributed memory stores a service self-management program, which is configured to store machine readable instructions, the service processor executes the service self-management program, and the instructions are executed by the processor to realize the system based on multi-puppet intelligent agent interactive learning as described in embodiment 1.

[0146] The above-described embodiments only express the preferred embodiments of the present application, which are described in detail and specifically, but cannot be understood as the limitation of the patent scope of the present application. It should be pointed out that for ordinary skilled persons in the art, several modifications, improvements and substitutions can be made without departing from the concept of the present application, which all belong to the protection scope of the present application.

Claims

1. A system for multi-puppet agent interactive learning, characterized in that, The application relates to a virtual chat room system, which comprises: a physical doll, a central base and a service terminal; the physical doll is provided with an NFC tag and is arranged on the central base to establish an NFC connection with the central base; the central base is used for NFC induction and recognition of the physical doll, integrates a receiver and a loudspeaker function, realizes audio signal collection and playing, carries out voice or touch interaction with a user, establishes a communication connection with the service terminal, and realizes real-time communication; the service terminal is used for selecting a chat theme in combination with a user portrait, establishing a virtual chat room, setting a role attribute of the recognized physical doll, recognizing a user intention according to collected voice information, and outputting interactive discussion to the central base for playing in combination with the chat theme. 2.The system based on multi-doll agent interactive learning according to claim 1, characterized in that, The role attribute of the physical doll comprises static attributes and dynamic attributes; the static attributes comprise a name, an avatar, a tone, a language code, a world view, a role biography and a knowledge base; the dynamic attributes comprise a positioning and a responsibility of the physical doll in the virtual chat room, and changes of emotion, tone and speed; the dynamic attributes are set according to a learning theme and the virtual chat room, and the dynamic attributes of the same physical doll are different in different chat rooms. 3.The system based on multi-doll agent interactive learning according to claim 1, characterized in that, The establishment of the virtual chat room comprises: generating or screening a suitable personalized learning theme dynamically through an embedded intelligent recommendation algorithm according to a historical learning portrait, an interest preference, a knowledge mastery level of a user and an education target preset by a parent; presetting a unique prompt word for each physical doll selected by the user, with each prompt word corresponding to a system prompt word in a large language model, and giving each physical doll a distinctive tone expression; setting a role definition of the physical doll by a professional content teacher and a copyright party, establishing a role allocation mechanism, and dynamically positioning the physical doll in combination with the role definition and the personalized learning theme; dynamically adjusting a tone parameter of the physical doll in real time according to a promotion situation of a current learning theme, an emotional tendency of a dialogue content and a role positioning of the physical doll.

4. The system for multi-doll agent interactive learning according to claim 3, wherein, The virtual chat room is provided with a doll intelligent agent selection mechanism, which is used for accurately matching a doll intelligent agent most suitable for responding to a user request or promoting a dialogue process in an environment where a plurality of physical dolls coexist; automatic promotion when the user listens and intelligent response when the user actively interrupts, and the intelligent agent selection mechanism will additionally consider a factor based on an intention recognition result when the user selects to interrupt a dialogue between current dolls; the doll intelligent agent selection mechanism is embedded with a plurality of intelligent agent matching algorithms, which automatically select a doll intelligent agent most suitable for processing a current task or responding to a user intention in combination with shared context information and role positioning of doll intelligent agents in the chat room.

5. The system for multi-doll agent interactive learning according to claim 4, wherein, The intelligent agent matching algorithm comprises: randomly selecting a doll intelligent agent from all doll intelligent agents in the chat room to speak in an initial stage of a personalized learning theme divergent discussion; selecting a next doll intelligent agent from the doll intelligent agents in the chat room to speak according to a preset order or a dynamically allocated order, so that each doll has an opportunity to express its own idea or introduce initial information in turn; The current shared context and the role information of all dolls in the group are packaged into a request and sent to the core large language model for decision-making. The large language model intelligently judges and recommends the doll agent most suitable for the next speaker; When the user's intent recognition result explicitly specifies a virtual chat room role to reply, the system will directly select the specified doll agent.

6. The system for multi-doll agent interactive learning according to claim 5, wherein, After selecting the doll agent, the system will invoke the doll agent, combine the current personalized learning theme progress, the role emotional setting of the doll agent, and the historical dialogue context to execute specific dialogue tasks and generate corresponding intelligent responses.

7. The system for multi-doll agent interactive learning according to claim 3, wherein, During the dialogue generation process of the doll agent, the voice synthesis parameters of the selected doll will be dynamically adjusted according to the emotional tendency of the current dialogue content, the role setting and positioning of the doll, the stage characteristics of the learning theme, and the overall rhythm of the dialogue process.

8. The system for multi-doll agent interactive learning according to claim 3, wherein, The identification of the user's intent refers to receiving the user's voice or text input, performing high-level semantic understanding and classification through a large language model, and identifying the purpose type of the user's dialogue this time.

9. The system for multi-doll agent interactive learning according to claim 8, wherein, The purpose type of the user's intent includes: The user actively expresses views or asks questions about the current personalized learning theme; The user specifies a specific entity doll in the virtual chat room to respond or explain; The user expresses interest in in-depth exploration of the current topic, and the entity doll provides more details or related information; The user wants to end the discussion of the current personalized learning theme and switch to a new personalized learning theme; The user encounters difficulties in interaction and seeks guidance or hints from the entity doll.

10. A device for interactive learning based on multiple doll-like intelligent agents, characterized in that, The device includes a service processor and a distributed memory, the service processor is connected to the memory, the distributed memory stores a service self-management program configured to store machine-readable instructions, the service processor executes the service self-management program, and the instructions are executed by the processor to implement the system based on multi-doll agent interactive learning according to claim 1.

Citation Information

Cited By

  • Intelligent interactive tidal playing method and device based on IP exclusive dialogue generation model and NFC technology

    CN121833911A

  • An intelligent interactive trendy toy method and device based on an IP exclusive dialogue generation model and NFC technology

    CN121833911B