Children asynchronous voice social contact method supporting role voice mutual simulation and content re-expression
By using an asynchronous voice-based social approach for children, combined with a language generation model and role-playing voice simulation, the problems of missing voice roles, limited expressive abilities, and excessive control in existing systems are solved. This enables a voice-based social experience that is child-led, authentic in sound, and rich in expression, thereby enhancing children's interest in expression and role recognition.
Patent Information
- Application Number
- CN202511309282.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-02-03
AI Technical Summary
Existing voice interaction systems for children lack a voice role binding mechanism, have insufficient ability to generate expressive content, do not bind voice changes to social graphs, and have excessive system control. Children lack autonomy in their expression paths, leading to identity confusion and the risk of imitation.
This paper presents an asynchronous voice-based social interaction method for children. By combining voice acquisition and semantic recognition with language generation models and character voice simulation, it constructs content blind boxes, voice blind boxes, and theater blind boxes to ensure that children dominate the expression process. It also sets up a parent content review mechanism to enable socially restricted exchange of character voices and multi-round collaborative performances.
It enhances children's interest, creativity, and role awareness in expression, ensures safe and compliant expression, retains children's complete control over voice selection and expression process, and forms a clear role-based voice interaction system.
Smart Images

Figure CN121459801A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voice interaction and social content generation, in particular to a children's asynchronous voice social method supporting role voice mutual simulation and content re-expression, which belongs to the application field of related cross technologies such as voice recognition, natural language processing, multi-role voice synthesis and children's social interaction system. BACKGROUND
[0002] With the rapid development of voice recognition, semantic understanding and voice synthesis technologies, voice interaction devices and voice social platforms for children are increasing. The current market commonly used products mainly include the following categories:
[0003] 1. Intelligent speaker products, which usually take voice question and answer, command response as the core function, have the ability of voice wake-up and voice broadcast, but the interaction is mainly for adults, and children's use is limited, and there is lack of situational expression and role perception ability.
[0004] 2. Voice changer or role simulation toys, which have simple audio transformation function, simulate robot voice, monster voice and other effects, but do not have the ability to generate expression content, and lack the system binding mechanism with role identity, and the voice change lacks social path and content association.
[0005] 3. Voice social application or message platform, such as children's voice message board and voice chat software, which has certain social interaction ability, but usually does not embed voice role mechanism, and the expression content and voice identity are not bound, which may cause identity ambiguity and imitation boundary problems.
[0006] 4. Children's dialogue toys, such as voice dolls or educational robots with voice recognition and content triggering ability, which have voice interaction ability, but have limited control ability of content expression generation, and the voice output is mostly driven by fixed content library, lacking creativity and role replacement mechanism.
[0007] Under the above technical background, the children's voice expression system has the following main technical deficiencies:
[0008] 1. Lack of voice role identity mechanism: in existing products, voice is only used as a means of information transmission, not as a unique identity of a role, which cannot build a "voice as role" expression model.
[0009] 2. Lack of expression generation and language completion ability: children's language expression naturally has the characteristics of colloquial, fragmented and non-structural, and existing systems cannot complete expression completion and language polishing, affecting the integrity and dissemination value of the content.
[0010] 3. The lack of social map restrictions in voice changes: the exchange between voice roles is usually unrestricted, which cannot reflect the social boundaries in reality that children "can only imitate friends", and is easy to cause problems such as identity confusion, expression misplacement and improper imitation.
[0011] 4. The expression process is dominated by the system, and the children lack control: the existing system is mostly driven by program preset interaction, and lacks an expression path controlled by children's language as the only control source. The expression chain is fragmented, the control points are scattered, and there is a lack of closed-loop design.
[0012] Therefore, the current children's voice interaction technology cannot meet the comprehensive needs of "expression content generation + role voice simulation + control right to children + clear social path", and a new system method is needed to reconstruct the role mechanism, social logic and technical control structure of children's voice expression. SUMMARY
[0013] I. Invention purpose
[0014] The present application aims to solve the problems of limited expression ability, lack of voice role, chaotic social path and excessive system control in existing children's voice interaction systems, and provides a children's asynchronous voice social method that supports role voice simulation and content re-expression, to achieve the technical goals of "expression as creation", "voice as identity" and "role as social".
[0015] In the prior art, children's voice expression usually relies on single sound box devices, command voice assistants or voice changing toys, which have the following outstanding problems: first, there is a lack of voice role binding mechanism, which cannot realize the personification expression path of "voice is role"; second, the expression content is mainly based on voice commands or messages, and there is a lack of language generation and completion ability, which is difficult to stimulate children's desire to express; third, the voice change mechanism is not bound to the social map, which is easy to cause risks such as unauthorized imitation and identity confusion; fourth, the system control is too strong, and children lack the autonomy of expression path, which cannot form a real interactive closed loop.
[0016] Therefore, the present application provides an asynchronous voice social method that takes children's natural voice as the only control input source, and completes expression generation, role voice selection, task scheduling and broadcast execution through model cooperation. The method constructs a role expression mechanism with voice as an identity tag, supports social restricted exchange and multi-round cooperative performance of role voice, and integrates mechanisms such as content blind box, voice blind box and theater blind box, which effectively improves the interest, creativity and role recognition of children's expression. At the same time, the present application sets up a limited parent content review mechanism and a non-intervention feedback receipt mechanism, which ensures the safety and compliance of expression, and retains the complete control right of children on voice selection and expression process.
[0017] In summary, the present application aims to provide a new method of voice social interaction for children, which is child-led, voice authentic, expressive, role clear and feedback natural.
[0018] II. Technical solutions
[0019] 2.1 Overall structure of the method of the present application
[0020] The present application proposes a child asynchronous voice social interaction method supporting role voice mutual simulation and content re-expression. The overall process starts from voice collection, and then goes through content generation, default broadcast, voice exchange, task confirmation, multi-role cooperation and finally broadcast execution, forming a complete expression closed loop of "child language leading - AI language enhancement - role voice realization - social content achievement". The method steps include:
[0021] Step one: voice collection and semantic recognition This step is initiated by dolls and other edge agents, and the expression intention is preliminarily digitized through voice collection and semantic recognition. The system identifies social instructions, friend anchors and role calling intentions in the voice in this stage, establishes a basic semantic link for the subsequent expression process, and automatically completes the analysis task by the platform model.
[0022] Step two: content generation and language reconstruction The system calls the language generation model to complete content completion, language polishing, scene expansion and expression optimization according to the child voice input content and intention understanding result, forming a complete expression content that can be broadcast. At the same time, the "content blind box mechanism" is introduced to enhance the creativity and uncertainty of expression and encourage children to express actively.
[0023] Step three: default voice expression construction Based on content generation, the system takes the currently activated edge agent as the basis to call its registered voice fingerprint to complete the voice broadcast version generation. This stage ensures that the expression content has a baseline voice output even if there is no role exchange. At the same time, it provides a listening and confirmation interface to provide a basis for comparison for the subsequent "voice exchange" link.
[0024] Step four: voice role exchange and listening confirmation The child initiates voice role switching through voice command, and the system extracts the registered doll and corresponding voice identity after recognizing the "friend name", completes voice simulation and generates a listening version. This step integrates the "friend identity blind box mechanism" and "voice exchange blind box mechanism" to strengthen the child's exploration and recognition ability of the role voice in the social network.
[0025] Step five: expression task confirmation and scheduling setting After the expression content and sound selection are completed, the expression confirmation, receiver setting and scheduling mode configuration are completed by the child. The system supports instant reporting, timed reporting, scene triggering and other modes, and constructs a "time blind box mechanism" in the fuzzy triggering scene to improve the interest and expectation of the expression process. At the same time, the only intervention point of the parent review mechanism is retained to ensure content compliance.
[0026] Step six: multi-role collaborative expression mechanism If the child initiates a multi-role collaborative instruction, the system can perform task decomposition and role allocation on the expression content, realizing master-slave singing, chorus, theater expression and other complex modes. This step supports cross-device collaboration, theater performance, "combination command blind box" and "theater blind box" mechanisms, expands the social creation space, and constructs a high-interactive sound expression theater.
[0027] Step seven: report execution and sound expression The system wakes up the target doll according to the set triggering condition and executes content reporting. The locked sound identity is used in the reporting process to ensure expression consistency, while the "sound inheritance blind box" and "task blind box mechanism" are combined to form a social fun of sound recognition and a new round of expression triggering opportunity on the listener side, realizing a complete closed loop of expression feedback.
[0028] 2.2 Specific step description of the method of the present application
[0029] Step one: voice collection and semantic recognition
[0030] This step is the starting stage of the system to execute the child social expression process. The doll and other edge intelligent agents initiate voice collection and complete preliminary semantic analysis to establish a digital basis for expression intent.
[0031] The "doll and other edge intelligent agents" in the present application refer to interactive terminal devices with voice collection, sound expression and role simulation functions, whose identities are registered and controlled by the platform and are only used for role execution and content reporting in the child voice social expression process.
[0032] In this step:
[0033] 1. Wake-up mechanism: the child starts the doll and other edge intelligent agents by active wake-up words, specific commands or shaking, etc. to trigger the voice collection function.
[0034] 2. Voice collection: the edge intelligent agent records the child's voice through the built-in microphone and uploads the audio data to the platform processing module in real time or with delay. No physical button is needed, and it does not rely on the smart speaker wake-up mechanism to maintain a pure voice interaction experience.
[0035] 3. Semantic recognition: the system calls the voice recognition and natural language understanding module to analyze the collected voice content and extract the following key elements: Core expression text (such as stories, fragments, phrases, etc.); Intention judgment (such as telling stories, leaving messages, imitating, singing, etc.); Instruction information (such as "try it with Xiaogang" "let him listen tomorrow morning").
[0036] 4. Social anchor point recognition: The system recognizes and analyzes "friend names" appearing in the voice and extracts the unique primary key as a social path. The "friend name" in the system is the entrance to the entire social graph, used to bind the role sound, determine the content direction, and the source pointer of subsequent simulation execution.
[0037] 5. Role analysis preparation: When identifying semantic structures such as "try it with someone else", the system prepares for subsequent role sound exchange, locks the target role corresponding to the doll identity (such as "Qianqian's Gulong Ball"), and hangs the intention into the content expression link.
[0038] The core of this step is to let the expression always revolve around the child's language dominance, without complex interaction methods, ensuring that all subsequent content generation, role control, and expression paths are triggered and controlled by the child's natural voice.
[0039] In addition, all analysis tasks in this step are completed by the platform side model, and the child does not need to know the model details. The system automatically completes the conversion from "language - semantics - social intention", and builds the next expression carrier and role basis.
[0040] This step is the starting stage of the system's child social expression process, initiated by the doll and other edge agents to collect voice and complete preliminary semantic analysis, establishing a digital basis for expression intention.
[0041] Step 2: Content generation and language reconstruction
[0042] After completing semantic recognition and expression intention extraction, the system enters the content construction stage. The system, based on the child's voice input content, combines semantic understanding results, and calls the language generation model to process the expression content as follows:
[0043] 1. Expression completion and language polishing: Complete the fragments, colloquial descriptions, and inspiration fragments issued by children, and retain the original expression style of children, making the output both faithful and expressive.
[0044] 2. Scene expansion and story generation: For expressions with story tendencies or interactive scene intentions, the system automatically generates expression content with role settings, plot construction, or fairy tale style, such as "we are at the pond" which can be expanded to "we meet a duck in a suit at the pond".
[0045] 3. Language reconstruction and expression optimization: Under the premise of maintaining semantics, the language is reconstructed in terms of rhythm, prosody, and sentence structure to make it more suitable for subsequent character voice broadcasting and have child acceptance.
[0046] 4. Content blind box mechanism integration: One of the core mechanisms in this step is the "content regeneration blind box mechanism", which is: Children input a sentence or phrase, and the system reconstructs the expression into an unexpected complete sentence or scene; This mechanism emphasizes the surprise, randomness, and motivation of the expression process, and is one of the key mechanisms to encourage children to continue expressing; Content generation has randomness and style diversity, with a "light trigger → heavy feedback" blind box effect.
[0047] 5. Children's dominant control is maintained: The system does not actively broadcast after generating content, and children decide whether to adopt, rewrite, or use it with characters, ensuring that expression is always child-led and system-assisted.
[0048] The design focus of this step is to enhance children's expression ability with AI expression, turning a sentence into a fairy tale and a message into a performance, improving the creativity and expressiveness of social expression, and forming a blind box-style experience based on "expression as creation" and "expression as surprise".
[0049] Step three: Default sound expression construction
[0050] After completing content generation, the system converts the constructed expression content into the default sound broadcast format of the current active doll. This step ensures that even without sound exchange, the expression content can be fully broadcast by the original character, establishing a baseline version of sound expression.
[0051] 1. Identity sound generation: The system calls the registered character sound of the current doll and other edge intelligent agents (managed by the platform), and generates expression voice content based on it. The sound fingerprint is the unique identifier of the doll's identity and cannot be changed.
[0052] 2. Tone and context adjustment: Without changing the sound fingerprint, the system can adjust the tone, rhythm, speed, and pause of the expression content to make it contextually adaptive, such as fairy tale tone, whisper tone, and speech rhythm.
[0053] 3. Sound carrier consistency guarantee: The system does not control the sound style with "emotion tags" or "style variables", but always maintains the sound as the only carrier of expression content, with children's semantics indirectly determining the expression effect. This approach ensures that expression control and aesthetic judgment are returned to children, rather than being defined by the model as "should be emotional expression".
[0054] 4. Trial and preview support: The system will provide a trial option for the expression content generated by default sound, allowing children to confirm whether to keep the current character sound or proceed to the next step of character sound exchange process.
[0055] This step serves as the starting point for the voice implementation of content expression, preserving the integrity of the character's voice identity and providing a comparable reference basis for the subsequent "sound exchange blind box mechanism". Children can decide whether to switch character voices and whether to try the friend's voice blind box based on the trial experience, enhancing the exploration fun of voice selection.
[0056] Step four: Sound character exchange and trial confirmation
[0057] After completing the default sound expression generation, the system enters the character sound exchange and trial confirmation phase. This step is the core link of establishing the "sound exchange blind box mechanism", initiated by children to explore the diversity and mystery of friend doll world.
[0058] 1. Character sound switching command recognition: Children express sentences such as "try changing to Qianqian's" or "use Xiaogang's voice" through natural language, and the system parses the "friend name" as the unique binding primary key.
[0059] 2. Doll identity confirmation prompt: The system confirms and prompts the specific available character (such as "Qianqian's Gulong ball") from the multiple dolls registered by "friends" on the platform, enhancing the blind box exploration fun. If the name is ambiguous, the system actively prompts for clarification, stimulating children's curiosity about the "sound universe" of friends.
[0060] 3. Sound simulation and trial experience: The system generates expression voice content using the registered sound fingerprint of the target friend doll, for children to try. Children can try multiple friend doll sound versions in sequence, making horizontal comparisons and preference selections.
[0061] 4. Multiple trial and content fine-tuning: During the trial process, children can make adjustment requests such as "change another one" and "let Gulong ball speak slower", and the system supports individualized fine-tuning of expression content rhythm, speed, and tone, ensuring the best adaptation of sound characters and expression content.
[0062] 5. Sound exchange blind box mechanism: This step integrates the "friend identity blind box mechanism" and "sound exchange blind box mechanism", achieving the following experience: Children cannot predict which dolls their friends have and their naming style; Each time the character sound is switched, it is a "sound surprise"; The system retrieves the sound identity from the friend's sound world, completes the dislocation expression, and brings "possession" performance tension and social fun; All role exchanges are completed under the child's will, and the system does not actively suggest or push.
[0063] The core of this step is to make "sound a probe for social interaction", and through the transformation of sound roles, to enhance the child's perception of friend identity, role naming, and sound style, to strengthen the interactive main line of sound as identity and social interaction as exploration, and to fully realize the "blind box mechanism" of sound social interaction.
[0064] Step five: expression task confirmation and scheduling setting
[0065] After completing content generation, role sound selection, and trial listening confirmation, the system enters the final confirmation and scheduling configuration phase of the expression task. This step is completed by the child to confirm the final expression intention, lock the sound identity, and set the expression method and time, ensuring the accuracy and personalization of expression behavior.
[0066] In this step:
[0067] 1. Expression content confirmation: The system presents the expression content synthesized by the role sound in a trial listening manner, and the child makes the final confirmation. The child can choose natural language commands such as "that's it" "change the content" "do it again" to trigger re-generation or proceed to the next step.
[0068] 2. Sound identity locking: The child confirms that the role sound used for the current expression content is the final selected sound, and the system locks the sound fingerprint as the only sound identity for subsequent broadcasting, ensuring consistency between sound and social role binding.
[0069] 3. Receiver confirmation: The system again confirms which friend doll the expression content will be sent to, ensuring the accuracy of the social graph path. All receiving targets are generated by the child specifying the friend's name, and the system does not actively push or recommend friends.
[0070] 4. Scheduling method selection: Immediate sending: The expression content is broadcasted by the target doll immediately after confirmation; Timed sending: The child can set natural language commands such as "play at 8 am tomorrow" "play before lunch" and the system translates them into specific timestamps for delayed scheduling; Scene trigger sending: The system supports preset social trigger points such as "play when the other doll approaches" "play together at the party" (see theater play for details), forming a more complex linked broadcasting mechanism.
[0071] 5. Time blind box mechanism integration: If the child sets a vague or open trigger instruction (such as "early tomorrow" "listen when she comes to my house"), the system automatically schedules the expression task within a reasonable time period, making the specific broadcasting time a surprise for both the child and the receiver, forming a "time trigger blind box mechanism".
[0072] 6. Expression task mounting: The system packages expression content, target sound, receiver, and trigger mechanism together and mounts it into a task list to be executed, and executes the broadcast instruction after the trigger condition is met.
[0073] 7. Parental review mechanism interface: If the platform has set parental review permissions, this step is the only review access point left for the system. Expression content needs to pass the guardian review or automatic pass mechanism before this step (see the explanatory section for details), and can only be scheduled and broadcasted after review.
[0074] The core of this step is that all expression tasks must be confirmed by the child before execution, and there is no default sending behavior. The system supports multiple sending opportunities and trigger conditions, making the expression behavior more personalized and interesting, and forming a blind box experience in the "time uncertainty".
[0075] Step 6: Multi-role collaborative expression mechanism
[0076] After the expression content, sound role, and broadcast task are confirmed, if the child wants multiple dolls to participate in the expression content, the system will enter the multi-role collaborative expression process. This step supports multiple dolls performing in chorus, round singing, master-slave singing, or theater performance modes, and builds a more rich sound theater experience.
[0077] 1. Collaborative command recognition: Children express commands such as "let Xiao Gang's and mine say together" "I say one sentence and you say one sentence" "everyone sing together" through natural language, and the system recognizes the collaborative command and extracts the multiple doll identities involved.
[0078] 2. Role assignment and task decomposition: The system decomposes the complete sentence according to the expression content and collaborative mode, and assigns it to different roles to complete tasks such as lead singing, counter-singing, chorus, dubbing, and hosting. For example: Master-slave singing: the main doll sings the complete song, and the other dolls repeat it later; Round singing: sentence relay broadcast between dolls; Chorus: multiple dolls sing simultaneously and overlap; Insertion of supporting roles: the main actor inserts another role's lines into the broadcast; K-song mode: children lead dolls to sing a complete song.
[0079] 3. Theater blind box mechanism: If the system supports theater play, children can trigger a multi-role collaboration scene through commands such as "let's have a show tonight" "try to arrange roles", and the system randomly assigns role combinations, lines structures, or order of appearance, making the performance result have a blind box surprise.
[0080] 4. Friend co-signing mechanism: When multiple children gather, the system can recognize doll devices from different accounts and operate cooperatively, realizing cross-social collaboration expressions such as "my doll sings with Xiaoliang's doll".
[0081] 5. Instruction combination blind box mechanism expansion: In the cooperative mode, children can realize instruction combination by continuous command, such as "sing this sentence first and then change the voice" and "start with the gurgling ball and then sing together", and the system recognizes the combination intention and merges it into a complete cooperative expression plan.
[0082] The design core of this step is to fully release the imagination of children in expression form, and to extend single expression behavior into a multi-person, multi-role, and multi-mode collaborative creation scene. Cooperative expression not only improves content expressiveness, but also adds "theater blind box" like immersion, organization, and cooperation achievement to social interaction.
[0083] Step seven: broadcast execution and voice expression
[0084] When the expression task enters the execution phase, the system triggers the target doll and other edge intelligent agents to complete the content broadcast according to the mounted broadcast task list. This step is the actual output link of the entire expression process, marking the formation of a complete closed loop from voice collection to voice broadcast.
[0085] 1. Target wake-up and broadcast triggering: The system automatically wakes up the target doll and starts the broadcast process when the conditions are met according to the task setting trigger mode (immediate, timed, scene trigger, etc.).
[0086] 2. Voice expression restoration: The system uses the locked role voice in step five to generate voice output, ensuring that the broadcast content is completely consistent with the child's initial setting, and expressing their creative intention and social desire in its original form.
[0087] 3. Voice inheritance blind box mechanism: If the receiving child does not explicitly know who will broadcast the content, the system can hide the voice role information, and the receiver can identify the role identity through voice recognition, forming a "voice inheritance" blind box experience, such as "who is talking to me?" "This seems to be a gurgling ball!" and other cognitive and emotional surprises.
[0088] 4. Task receipt and state feedback: After the broadcast is completed, the system automatically generates feedback information and sends it to the original child account, such as "successfully delivered" and "the other party has finished listening", establishing a closed loop confirmation of expression behavior.
[0089] 5. Interaction opportunity extension: The receiving doll can guide the other party to respond (such as "do you want to respond?" "change your voice and say it"), forming a new round of voice interaction and activating the next expression cycle.
[0090] 6. Expanding the Task Blind Box Mechanism: After the broadcast content ends, the system can add a platform-generated "voice task challenge" or "friend interaction suggestion", such as "Try to let Gulu Ball say it again?" or "Can you imitate his speech?" Within the limits of children's permission, a "task challenge blind box" is formed to encourage continuous social interaction.
[0091] The design focus of this step is to ensure that the broadcast is authentic and controllable, the process is perceptible, and the feedback is responsive, while preserving the uncertainty of the sound, the exploratory nature of the content, and the continuity of the role perception, so as to achieve a closed-loop experience of social expression from "trigger" to "implementation".
[0092] 2.3 Description of the speech-driven expression generation process
[0093] 1. Nondeterministic generation mechanism and content exploration experience design
[0094] This invention introduces a non-deterministic content generation mechanism to enhance children's enjoyment and exploratory experience during the expression process. Based on a language generation model, this mechanism combines lightweight input information (such as keywords, short phrases, or prompts) to generate semantically reasonable outputs that differ in expression structure, plot development, and details. Specifically, it includes: Context adaptation (the system combines historical expression fragments for style matching); Enriching the plot (associating and expanding the semantic meaning of the input); Output diversity and randomness (the same input may correspond to different output versions).
[0095] This mechanism can effectively enhance the "unknown" and "surprise" of the expression process, allowing children to discover different content through multiple attempts, stimulating their creative interest, and forming an expression exploration experience based on subtle prompts.
[0096] 2. Natural Language Generation Mechanism Based on Light Input
[0097] To accommodate the varying levels of language expression ability among children, this invention supports a "light input triggered" content generation path. Children only need to utter phrases, keywords, or everyday colloquial expressions (such as "Are you coming to play today?" or "Tell me about dinosaurs"), and the system can recognize their linguistic intent. Based on this lightweight input, the system automatically invokes a language generation model to generate content that is structurally complete, semantically coherent, and emotionally expressive.
[0098] This mechanism lowers the threshold for expression, enabling children who do not yet have the ability to express themselves in complete sentences to participate in character content creation, improving ease of use and social expression participation, and reflecting the human-computer interaction concept centered on "natural interaction".
[0099] 3. The system does not dominate the natural triggering mechanism of expression decision-making.
[0100] To ensure the initiative of expression behavior completely belongs to the children, the system does not provide any form of content suggestion, role recommendation or expression guidance during the implementation of the application. That is: The system does not actively prompt expressible content sentence patterns, role voices or tone options, nor does it guide through text, voice, interface icons or animation, nor does it preset tasks or expression intentions, nor does it automatically start the expression link based on context prediction. All expression generation related model calls (including language generation models, voice synthesis models, etc.) need to be explicitly triggered by children through natural voice commands.
[0101] This mechanism effectively avoids the guiding problem of the platform "replacing children's expression", prevents the intervention of expression rights by system function logic, ensures the spontaneity, autonomy and originality of children's expression behavior, and further establishes the core technical path of the application as "children's voice commands as the only trigger".
[0102] 4. Trigger form and path specification of expression task chain
[0103] The method of the application uses "voice command" as the only trigger for expression tasks, and does not support non-natural language forms of interaction paths such as button clicks, icon selections, and recommendation guidance. That is: All expression activities, including content generation, role voice switching, chorus collaboration, and round singing control, require children to explicitly issue voice commands; The system only starts subsequent model calls, role announcements or collaborative expression processes after recognizing and semantically analyzing the voice input; If the child does not issue an instruction, the system remains in a silent listening state and does not actively make suggestions or generate content.
[0104] This mechanism establishes a natural interaction paradigm of "language as control", ensuring that the starting point of all expression activities originates from children's willful judgment and preventing the system from actively inducing expression behavior with predictive models.
[0105] 5. Process description of multi-model collaborative call
[0106] In the expression chain, the application completes the complete process "from children's language to role announcement" through the combination of voice command recognition → intention analysis → content generation → voice synthesis → announcement execution models. Among them: The language generation model is responsible for expanding the input phrase / key word into a complete structure; The role voice synthesis model realizes personalized voice reconstruction according to the current designated role identity; The announcement execution module completes voice output in the edge intelligent agent and records the announcement completion status; The whole process does not rely on the active scheduling of the central control, and all the children's voice intentions are used as the process control signals.
[0107] The mechanism supports parallel model collaboration, such as multiple edge agents simultaneously completing their own voice synthesis and broadcasting execution in a chorus / round singing task, ensuring synchronization and consistency of the role situation.
[0108] 6. Role chain expression
[0109] Children can trigger "main singer + chorus" collaborative expression through voice commands, such as the command: "Gurug ball main singer, little bear chorus". The system controls the corresponding two edge agents to complete the content segmentation broadcasting or round singing performance in sequence. All role division, broadcasting rhythm, and content control are initiated by children through voice, and the system does not default or provide preset solutions.
[0110] 7. The language generation model (such as GPT, BERT, etc.) and the speech synthesis model (such as FastSpeech, Tacotron, etc.) involved in the present invention are used as existing technical means. The present invention does not claim to protect the algorithm structure, training method, network parameters, or optimization method of such models. The role of the model in the present invention is as a callable module. The technical innovation point lies in the expression control mechanism design of its collaborative calling method and triggering path.
[0111] III. Parental supervision mechanism and social feedback channel
[0112] To ensure the children's voice social expression freedom and reasonably embed the supervision ability of parents, while avoiding system intervention in children's social will, the present invention specially designs a limited parental supervision mechanism and a non-intrusive social feedback channel, constructing a "moderate supervision + autonomous expression + natural social circulation" technical structure. This mechanism mainly reflects in the two key nodes of pre-sending audit and post-broadcasting natural feedback guidance of the expression task. The specific explanation is as follows:
[0113] (1) Parental content audit mechanism (before the sending link, text content audit)
[0114] The present invention only sets up a unique supervision node before sending the expression task, that is, after the child completes the content generation and role sound confirmation, the system will provide an optional parental audit channel for the text part of the expression content before mounting the task. The design features are as follows:
[0115] 1. The audit scope is limited to "text content": Parents do not participate in the child's listening process and do not intervene in any voice role selection or listening link; Parents do not audit voice content, and all doll voices are uniformly registered and recorded by the platform to ensure compliance, safety, and stability; The audit object is the text content itself, ensuring semantic compliance and no inappropriate expression.
[0116] 2. The audit entry method is clear and single: Parents receive audit requests through a dedicated mobile phone APP; The audit interface only displays: expression text content, target friend name (specified by the child), current doll (character name), and scheduled play time; No sound preview or multiple version content is displayed to ensure the child's freedom of sound selection.
[0117] 3. The audit authority comes from doll authentication: Only the doll currently bound to the parent's account has content audit rights; It does not have the right to audit the expression content of other children's or friend's dolls; The doll identity and sound ownership are authenticated by the platform at the time of registration and binding, and the parent only bears the responsibility for monitoring the doll's expression.
[0118] 4. The audit processing strategy is simple and clear: Parents can choose "agree", "send back" (guide the child to rewrite), or "ignore" (system default setting); The system does not recommend default full audit and encourages parents to set trust based on the child's age and content style; The audit logic is clear and cannot be expanded, and the system does not allow parents to enter the "sound review" or "character control" section.
[0119] (II) No regulatory interface is set for the sound preview and character selection process
[0120] This invention highly values the child's right to freely explore in sound expression, therefore explicitly stipulates as follows: The character sound preview process is led by the child, and the system supports multiple rounds of preview, role change, speed change, and tone change; Parents cannot intervene in this process and do not have the right to view or replace the preview; All sounds come from platform-registered and recorded edge intelligent character roles, which have been authenticated by the platform, and there is no need to review the sound content itself; During the selection of the favorite character sound, the child forms a cognitive and social exploration of friend characters, sound differences, and naming interests, which is one of the core mechanisms of this invention and should not be destroyed by regulatory behavior.
[0121] (III) Simple and natural social feedback channel (not dominated by the system)
[0122] To ensure the effective closed loop of children's voice social expression, the non-mandatory and non-guided feedback mechanism is designed, and the system is avoided to actively intervene beyond authority. The specific embodiments are as follows:
[0123] 1. The essence of the feedback mechanism is "the next round of social behavior start": The receiving children can freely choose whether to respond; If responding, a new round of expression task is formed; The invention does not provide a preset script or system suggestion, and completely leaves it to the children to make their own decisions.
[0124] 2. The system does not actively suggest feedback behavior: There is no system voice prompt such as "do you want to respond?"; There is no active push of "task challenge" or "interaction suggestion"; There is no platform set task or interaction instruction in the feedback.
[0125] 3. Feedback information is only for communication, not for suggestion: The original sender children can know the content has been broadcast through the prompt; The system can display simple status such as "delivered" and "play completed", without guiding the next action; All feedbacks are only receipts of fact status, without embedding any suggestion or guidance content.
[0126] 4. Feedback is the new social starting point: The invention believes that the real feedback should not be set by the system, but by the children through re-expression behavior; The system only needs to ensure that the feedback channel is smooth, and the real "response" is triggered freely between children; From this, a rolling asynchronous expression chain is formed, so that social interaction continues to develop in natural flow.
[0127] (Four) Mechanism principle and design boundary description
[0128] In summary, in order to achieve the balance between expression autonomy and supervision, the invention adheres to the following four principles in mechanism design:
[0129] First, the sound freedom principle. The invention ensures that all sound-related choices, listening and use are completed by children independently, and the system does not set any restrictions, and parents have no right to intervene. The platform ensures the legality and safety of all edge agent voices through unified record management, and ensures that voice content does not need to be checked by parents from the source.
[0130] Secondly, the content review limitation principle. The parents only have the right to review the text content before the expression task is officially mounted, and this right only applies to the dolls they manage, and they cannot review other people's content or interfere with the sound content itself. This right cannot be extended or transferred, ensuring that children have complete sound expression rights during the expression process.
[0131] Thirdly, the feedback is not mandatory principle. The invention only provides basic broadcast state feedback and does not embed any system-initiated suggestions, task pushing or interactive instructions. Whether to respond is determined by the child, and the system does not participate in judgment, recommendation or guidance, thereby avoiding platform intervention in the child's social rhythm and expression rhythm.
[0132] Finally, the expression closed-loop self-control principle. The starting point and endpoint of each round of expression process are controlled by the child, and the system only serves as an execution channel for expression, without assuming the role of content design or social task generation. Even in multiple rounds of expression rolling, there is no platform task chain design, ensuring that expression behavior naturally flows and freely grows under the guidance of the child.
[0133] Through the above four mechanism principles, the invention realizes the system design of parental supervision, sound non-abuse, expression closed loop, and feedback flexibility, taking into account safety, autonomy and social interest.
[0134] Four, model calling mechanism and expression control path explanation
[0135] The "asynchronous voice social method for children" proposed by the invention is based on the core technical concept of "sound as role, sound carrying emotion", which builds a complete expression technology path of expression task link + multi-model collaborative calling + role identity constraint control. The path is composed of multiple model modules, but the core technical innovation is not in the algorithm structure of the model itself, but in the collaborative calling method of the model in the expression process, the role sound identity binding mechanism and the system design logic of expression master control right.
[0136] (I) Role identity function of sound
[0137] In the expression system defined by the invention, sound is not only the physical waveform of voice output, but also the unique identifier of the role. Each "doll edge intelligent agent" binds its unique sound fingerprint when registering on the platform, and assumes the following role attributes during the entire expression process: 1. Represents the role personality, which constitutes the basis for children's cognitive role differences; 2. Transmits changes in tone, rhythm and emotion (without changing the sound fingerprint); 3. Carrying the expression identity, binding with the friend name in the social graph, and constituting the content routing path.
[0138] The binding mechanism ensures that the sound cannot be randomly spliced or mixed, preventing unauthorized expression or role confusion.
[0139] (2) Model calling path structure
[0140] The method of the present application involves the following types of model module combination calls in the expression task execution process, the order of which is as follows:
[0141] 1. Speech recognition model (ASR) The input is the original voice collected by the doll; The output is the text content for the downstream semantic model; This module supports edge collection and cloud analysis, supports fuzzy recognition and child speech optimization.
[0142] 2. Semantic understanding and intent recognition model (NLU) Perform syntax analysis and semantic intent extraction on the recognition results; Extract keywords, task types (such as leaving a message, imitation, singing), and role invocation commands; Identify social anchor points (such as friend names) and map them to unique IDs within the system.
[0143] 3. Language generation model (LLM) Used to complete, polish, transcribe, and fairy tale the original expression; Generate output sentences with rhythm and emotional tone; Support "content blind box mechanism" light trigger and heavy feedback expression generation mode; The model calling process is controlled by the child's voice, and the system does not automatically generate content.
[0144] 4. Sound synthesis model (TTS) The input is the text content generated in step three + the determined role voice fingerprint; The output is a speech waveform; The platform calls the exclusive voice model template of the role binding to generate expression audio; Unauthorized role voice synthesis is not supported, and custom synthesized sound is not allowed.
[0145] 5. Expression scheduling and dispatching module Generate dispatching parameters (immediately, timing, scene trigger) combined with semantic instructions; Mount the expression task, complete role sound encapsulation, receiver confirmation, and sending method determination; Support parent review interface and expression permission constraint interface.
[0146] The above models are general modules, and the system can call different third-party models to complete the above tasks, and the application does not limit any model architecture or training method.
[0147] (Three) technical protection boundary declaration
[0148] In order to clearly define the protection scope, the application is defined as follows:
[0149] 1. The model body is not claimed to be protected The application does not involve the algorithm, structure, training method or optimization path of the language model (such as GPT, BERT, etc.) or the voice model (such as FastSpeech, Tacotron, etc.); The specific model parameters, network structure, fine-tuning method or framework implementation form are not limited; All models are regarded as general technical modules as callable resources in the expression path.
[0150] 2. The technical focus claimed to be protected is: model cooperative calling method + expression control path design, including but not limited to: Combination path of various models in the voice social task flow; Expression control right is dominated by children and responded by the platform; Synthetic generation path under the sound identity binding mechanism; Content routing logic controlled by social graph; Expression generation does not actively, recommend or replace, and the model calling system is driven by the will of children.
[0151] 3. Emphasis on "identity constraint calling" of sound synthesis model The use of all TTS models must be bound to the role sound fingerprint authorized by the platform; Mixing output or free splicing is not allowed to ensure that the expression sound is strongly consistent with the role identity; This mechanism ensures the invention concept of "sound cannot be forged and role cannot be overpowered".
[0152] (Four) Model deployment and execution control principles
[0153] The application supports distributed deployment and flexible access of models, and the platform can select cloud deployment or edge inference according to actual business: ASR and NLU modules support edge lightweight deployment to shorten response delay; LLM and TTS modules are generally deployed in the cloud and accessed through interface calling; All module calls are triggered under the guidance of children's voice instructions, and no system-generated path is generated; Parents have no right to influence model selection, voice synthesis or character change, and only have the right to review part of the text content (see the third part above).
[0154] Through the above model calling mechanism and expression control path, the application realizes the system expression logic of voice as role, expression as identity, and model as collaborative executor, establishes the voice social behavior basis with children's voice as the only control source and voice fingerprint as the only identity tag, and forms an expression system with high originality and controllable expression without relying on model algorithm innovation.
[0155] Five, technical effects, beneficial values and improvements of the prior art of the application
[0156] In order to highlight the technical progress and use value of the application in the field of children's asynchronous voice socialization, the beneficial effects, innovative contributions and comparative advantages of the application with the prior art are described as follows in combination with the method logic and the industry status quo:
[0157] (I) Beneficial effects of the application The "children's asynchronous voice socialization method supporting role voice mutual simulation and content re-expression" proposed by the application has the following significant beneficial effects:
[0158] 1. Realize the role socialization mechanism with voice as the only expression identity The application defines voice as the only identity of the role, and binds the role personality and expression permission through voice fingerprint, so that "voice as role" becomes the main axis of expression, thereby avoiding common problems such as role confusion and expression boundary crossing.
[0159] 2. Support role voice exchange mechanism based on friend relationship The system only allows the doll role voice of the added friend to be simulated, which conforms to the reality social boundary and effectively prevents risks such as arbitrary voice change and anonymous expression, and realizes the controllable voice exchange play based on the social graph.
[0160] 3. Build content blind box type expression experience and stimulate expression interest Through the content re-expression mechanism of light sentence input + AI reconstruction, children can trigger a fairy tale with one sentence, form a positive cycle of expression as creation and creation as surprise, and improve language expression ability.
[0161] 4. Voice role exchange brings expression tension and expressiveness The change of the role voice is dominated by children, forming dislocation expression, "possession feeling" and voice cognition exploration, and giving expression content emotional intensity and performance interest.
[0162] 5. The whole process takes children's voice as the only control source to maintain expression freedom All expression processes are triggered by children's natural language commands, and the system does not provide active suggestions, does not generate content, and does not intervene in control. The expression control right belongs to the children themselves.
[0163] 6. Introduce a limited parental review mechanism to ensure compliance with expression but not overstep Only the "text content" establishes a parental review interface, and is limited to the pre-delivery link and the binding of the doll character, avoiding interference with sound selection and character exploration, ensuring that the child's expression space is not compressed.
[0164] 7. Realize the closed-loop link of expression behavior and natural rolling social cycle From expression generation, sound determination, broadcast execution to broadcast feedback, it constitutes a complete link and supports the receiver's spontaneous response to form the next round of expression task, promoting the continuous development of asynchronous social interaction.
[0165] (II) Technical contribution and innovation path of the present application
[0166] Compared with traditional speech synthesis, children's dialogue robots, smart speakers and other voice interaction products, the present application has the following originality in terms of technical architecture and expression control logic:
[0167] 1. Propose the expression paradigm of "sound as identity" Upgrade the sound from an output tool to an expression subject, and the character identity is bound by the sound and cannot be changed. All sounds are registered and authorized by the platform to avoid identity misplacement caused by free voice change.
[0168] 2. Build an expression control path based on model combination Instead of advocating the ontology of speech models, the present application focuses on how to use children's speech as the only control source to cooperatively call speech recognition, semantic understanding, language generation and sound synthesis models to complete the entire expression process. This path realizes the complete expression logic of "expression perceptible + control right verifiable + content controllable".
[0169] 3. Design multiple blind box mechanisms to stimulate social expression motivation The present application creates "sound blind box", "content blind box", "time blind box" and "theater blind box" mechanisms, which are combined with children's cognitive development and interest mechanisms to make expression not only purposeful, but also exploratory and game-like.
[0170] 4. Realize the deep binding of expression link and social graph All expressions are anchored by "friend names", and sounds can only imitate friends they have added. The control path is embedded in the social graph to prevent the system from becoming an arbitrary content / sound synthesizer, ensuring that it is essentially a "social system" rather than a "speech platform".
[0171] 5. The platform only serves as an execution channel and does not actively intervene in expression content All expressions are initiated by children, the platform does not provide emotional labels, does not inject style templates, and does not generate suggested content, maximizing children's language autonomy.
[0172] (Three) Comparison and analysis with prior art
[0173] Compared with the prior art, the present application has obvious advantages in structure design, expression mechanism and control logic:
[0174] 1. Different from traditional speech synthesis system Ordinary TTS system only realizes speech output, does not have the functions of role identity binding, role sound registration or social relationship constraint, and the expression content often lacks identity orientation and interactive feedback structure.
[0175] 2. Different from intelligent sound box and voice assistant products Intelligent sound box is mainly to execute commands, the interaction path is "one question and one answer" or "passive wake-up", which cannot form an expression generation chain, nor support multi-role collaboration or blind box exploration mechanism.
[0176] 3. Different from children's voice changing toys or voice changers Traditional voice changing products are based on audio transformation algorithm and have no expression content generation capability, do not have voice simulation rules based on social relationship, and cannot form controllable identity and role path.
[0177] 4. Different from social voice platform Conventional voice social products (such as message board, voice chat room) lack role driven mechanism, sound has no registered identity, content has no expression control path, and are prone to cross-border imitation and content interference problems.
[0178] 5. The present application first constructs a "expression is role + sound is identity + content is social" three-in-one path, realizes the whole process closed loop from expression intention collection to voice role broadcast, and binds the expression behavior within the children's voice dominance and social relationship real binding, and constructs a safe, controllable and expandable voice social system. BRIEF DESCRIPTION OF DRAWINGS
[0179] Figure 1 is a system structure block diagram of the present application; Figure 2 is an expression generation flowchart of the present application; Figure 3 is a schematic diagram of the role sound exchange mechanism in the present application; Figure 4 is a parent review process logic diagram; Figure 5 is a model collaboration calling structure diagram. DETAILED DESCRIPTION
[0180] In order to make the purpose, technical solutions and beneficial effects of the present application clearer and more explicit, the present application is further described in detail below in combination with the drawings and examples. The present application is not limited to the following specific examples, and any equivalent replacement or improvement under the spirit and principles of the present application shall be included in the protection scope of the present application.
[0181] Example one, identity exploration mechanism (friend identity blind box)
[0182] In this embodiment, when a child initiates expression through a voice command, the system only allows the registered voice of the doll edge agent of the child's "added friend" to be called as an expression carrier. This mechanism establishes a role calling boundary based on a real social graph, ensuring that safety and exploration interest coexist.
[0183] Example story:
[0184] A six-year-old child Mingming is at home talking to his doll edge agent "Guluguo Ball", and wants to try a different voice to say a riddle he made. He says, "Try it with Xiaomi's voice." The system recognizes "Xiaomi" as Mingming's friend and confirms that there are multiple registered role voices under this friend's account, but Mingming does not know the specific name or appearance of Xiaomi's doll character.
[0185] The system then responds, "Want to use Xiaomi's 'Rolling Big Bear' to say it?" At this time, Mingming hears for the first time that "Rolling Big Bear" is one of Xiaomi's doll names, and feels surprised and curious, so he answers, "Yes, this one!"
[0186] The system then broadcasts the riddle "I open my eyes at night, and close my eyes in the day" that Mingming originally intended to say through the registered role voice of "Rolling Big Bear", which is thick and deep, with a slight panting sound, making the riddle mysterious and funny. Mingming laughs and can't help but say, "Xiaomi's bear is so funny, let's try another one!"
[0187] At this time, Mingming says again, "Try a different voice of Xiaomi's." According to his social relationship and registration authority, the system prompts, "You can also use Xiaomi's 'Miaomiao Star' voice." Mingming nods in curiosity, and completes one after another role blind box exploration.
[0188] Mechanism explanation:
[0189] In this game, children cannot predict all the role information of their friends, and the platform does not directly display the role list of friends, but through the blind box exploration interaction of "try a different role", it guides children to gradually discover the "sound universe" of their friends. Each call is not only the completion of an expression behavior, but also a process of identity discovery in a social network.
[0190] The platform strictly limits the calling range of role voices to the registered voices of "friended" edge agents, ensuring the legal controllability of voice identities. In other words, even if the child knows that another child has some interesting voice roles, the system will not provide a calling entry before a friendship is established.
[0191] This mechanism reinforces the core concepts of "voice as identity" and "voice limited to social relationships," making voice socialization in the children's voice expression platform exhibit natural structural boundaries and exploratory fun, creating a voice-based social blind box system.
[0192] Example Two: Role Voice Simulation Mechanism (Voice Exchange Blind Box)
[0193] This example describes a play mechanism based on the "role voice attachment" logic: children can use voice commands to let their doll edge agents use the registered voices of other dolls to express their own content, but this "voice exchange" is limited to the "friended" roles in their social graph, and does not change the appearance of the physical doll, only simulating the role with voice.
[0194] Example Story:
[0195] A seven-year-old child, Xiaoyu, has a doll edge agent named "Doubao." She usually interacts with it in a cheerful and energetic voice. But today, Xiaoyu wants to tell a hilarious joke with a strong contrast, so she says to "Doubao": "I want to use Dingding's voice to say this."
[0196] Dingding is one of her friends, and has a doll called "Kadala Tiger" with a registered voice of a husky and deep Northeastern accent. The system recognizes that Dingding is in the friend list and has registered the voice "Kadala Tiger," and prompts: "Do you want to use Dingding's 'Kadala Tiger' voice?"
[0197] Xiaoyu nods with a smile: "Yes, just this one!" Then she says the joke: "One day, a duck got lost and went to the police station to report. The police asked its name... it said: Ga!"
[0198] The system immediately calls the registered voice of "Kadala Tiger" to express this content in Dingding's doll voice, creating a strong contrast in tone. The whole joke becomes a kind of mysterious toughness, and Xiaoyu laughs loudly and can't help saying: "I want to say it again with someone else's voice!"
[0199] In this process, Xiaoyu did not control Dingding's doll itself, but through voice commands, she attached the use of her friend's role voice to her edge agent "Doubao" to complete a content expression. After the expression is completed, the system only prompts "delivered" briefly, without guiding whether to respond or continue.
[0200] Mechanism Explanation:
[0201] This sound exchange mechanism allows children to temporarily invoke the registered character voices of their friends for content expression, creating a "I speak with your voice" gameplay experience. Sound becomes a carrier of expression and an important symbol of character personality.
[0202] However, the system sets strict boundaries - only registered character voices of "added friends" are allowed to be invoked, and characters of friends' friends or strangers are not allowed to be invoked, to avoid boundary simulation or invasion of others' expression personality.
[0203] Sound exchange does not involve the transfer of physical ownership or changes in appearance, and the platform only allows "sound" to dynamically attach within the scope of friends. This gameplay makes interactions between children have a certain "simulation-impersonation-character switching" creative color, stimulating more complex expression imagination and social tension.
[0204] Through sound exchange, children not only "borrow" voices, but also perform a role, reproduce tone, and express performance, which is a sound-centered identity expression game.
[0205] Example Three, Non-deterministic Expression Generation Mechanism (Content Blind Box)
[0206] This embodiment describes a "light prompt + system-generated" expression method: children only need to provide a keyword, phrase, verbal exclamation, or fragment language, and the system will generate a complete, interesting, and plot-rich expression content based on the current context and call the language generation model. Since the generated results are uncertain, each expression is like opening a "content blind box", with exploratory and surprise.
[0207] Example story:
[0208] Eight-year-old Lele sat on the sofa and said to his teddy bear "Huluhu": "Help me say a poem about the rainbow."
[0209] The system recognizes "rainbow" as the keyword, then automatically calls the language generation model, and combines Lele's age preference and past expression style to quickly generate content: "The rainbow hides behind the raindrops and quietly draws a bridge. A cat walks on the bridge, carrying stars and bubbles."
[0210] This content is then read out by "Huluhu's" voice. Lele's eyes lit up and he said, "Say another one that's different."
[0211] The system receives the new instruction and generates a new one:
[0212] "The rainbow is not a bridge, but a fish that laughs out loud. The fish swims through the sky and pokes a hole in the cloud."
[0213] Lele laughed and said, "I like this one, it looks like my drawing."
[0214] Then he shouted, "Say it in Little Watermelon's voice!" The system recognized "Watermelon" as his friend, who had the character "Rampaging Carrot" and a registered, lively, and exaggerated voice style. The system prompted, "Use Watermelon's 'Rampaging Carrot' voice?" After Lele confirmed, he could read the system-generated content again, but this time, the voice style and emotional expression had changed, making it even more fun.
[0215] Mechanism Explanation:
[0216] In this mechanism, children do not need to fully conceive their expression; they only need to say simple prompts such as "tell me a story," "help me say something in my sleep," or "a poem about rainbows," and the system will complete the "expression completion" through a language generation model.
[0217] The model being called must have the following characteristics: Context adaptability: It can understand the current content environment; Output diversity: The same prompt can generate different content; Completeness of expression: The generated content is logically and linguistically coherent and forms complete sentences; Emotional fit capability: Generate text that matches the tone of voice by combining the character's voice.
[0218] This mechanism is particularly beneficial for children with weaker expressive abilities or more imaginative minds, effectively lowering the barrier to expression and stimulating their motivation to express themselves. Because the generated results have a degree of "uncertainty," each generation has a "blind box" effect, keeping children interested in exploration.
[0219] Meanwhile, the mechanism is always triggered by the child, and the system will not actively push content suggestions to avoid interfering with the child's intention to express themselves, reflecting a collaborative mechanism of "autonomy of expression + gentle inspiration for generation".
[0220] Example 4: Multi-role Collaborative Expression Mechanism (Theater Blind Box)
[0221] This embodiment describes a gameplay mechanism that supports collaborative expression by multiple edge intelligence agents, such as dolls. Children trigger the "multi-person theater" mode through natural voice commands, and the system arranges different characters to take turns speaking according to their instructions, constructing a voice theater with character interaction. Children can either specify character lines or trigger the system to automatically generate character lines through prompts, forming a theater blind box gameplay with "story improvisation" and "voice combination effect".
[0222] Example story:
[0223] Nine-year-old An An invites her two friends Xiao Chu and Pang Dou to play "role-playing". The three of them each have a bound edge intelligent doll: "Little Elephant", "Clever Monkey", and "Fat Pig". An An gives voice commands: "We are going to perform a story about the forest kingdom. Little Elephant is the king, Clever Monkey is the minister, and Fat Pig is the troublemaker."
[0224] The system, according to the role setting, will involve the registered voices of the three dolls in order to perform, and prompt: "The theater is starting: the roles are the king, the minister, and the troublemaker. Do you want to automatically generate the plot?" An An answers: "Partly automatic, I will start first." She says loudly: "Today, something big happened in the forest kingdom. The king woke up and found that his crown was missing."
[0225] The system recognizes this sentence as the "beginning of the script" and then automatically generates the following lines, with the voice of "Clever Monkey": "Your Majesty, Your Majesty, I just saw Fat Pig sneaking into the treasure house and carrying a bunch of shiny things out!" Fat Pig's voice then bursts out: "Hmph, I was checking if the gems had faded, and I accidentally dropped your crown..." The three characters speak one after another, and the system automatically generates the subsequent plot content according to the role setting and plot progression, but always leaves a gap for An An and her friends to insert and rewrite instructions at any time.
[0226] For example, An An suddenly shouts: "Little Elephant becomes very angry and scolds Fat Pig in the most authoritative voice!" The system immediately switches to the "angry tone template" and says with "Little Elephant's" registered voice: "Daredevil Fat Pig! How dare you enter the treasure house without permission?" The story continues, and the ending is decided by An An: "Finally, Fat Pig apologizes and writes a poem to give to the king." The poem generated by the system is expressed by the voice of "Fat Pig", forming a complete and full of improvisational elements children's theater.
[0227] Mechanism Explanation: This mechanism supports the following functional features: Multiple roles are bound in parallel, and multiple edge intelligent agents work together to express; Children control role setting, order of appearance, and plot development; The system calls language generation models as needed to provide plot assistance; Each character's voice remains consistent with their identity and is never mixed or replaced; The theater mode supports script breakpoints, plot reconstruction, and character switching.
[0228] The theater interaction mechanism not only enhances the children's collaborative expression ability, but also forms a highly integrated expression scene of "sound + plot + social relationship". Children obtain a strong sense of creation, participation and social existence in the process of controlling characters, selecting sounds and developing plots.
[0229] Example Five, Time Delay Trigger Mechanism (Time Blind Box)
[0230] This embodiment relates to a social play mechanism combining "time setting" and "expression delay trigger". Children set the expression time for dolls and other edge intelligent agents through voice commands, and the system automatically wakes up the target doll to make a sound at the specified time, thereby realizing time blind box type interactive effects such as "delay surprise", "timed chorus" and "appointment story".
[0231] Example story:
[0232] Ten-year-old Chenchen gave voice commands to her two dolls on Friday night: "Little zebra, sing 'Happy Birthday' at 8 o'clock tomorrow morning; Little fox, say a paragraph of blessing I wrote at 8:30." This is a birthday surprise she prepared for her good friend Yueyue. Yueyue's home also has two doll devices, which are bound to "friend identity" of Chenchen. The system sets up the background to wake up and broadcast these two expressions at the specified time on Yueyue's birthday.
[0233] At 8 o'clock the next morning, little zebra sang in Yueyue's room: "Happy Birthday to you, Happy Birthday to you." At 8:30, little fox said in a slightly shy tone: "I hope you are happy every day, as happy as the day we played together! - Chenchen." Yueyue was very happy after listening and immediately told little fox: "Help me reply a paragraph and tell Chenchen that I am very touched." This reply entered another asynchronous expression process (which can also be set to play at a specified time), and a new social interaction is about to begin.
[0234] Mechanism explanation: The time blind box mechanism has the following characteristics: Timing expression setting: children can set the target time of expression through voice commands; Delay playback trigger: the system triggers playback according to time, and does not play any content before the setting; Appointment content variety: can be used for blessings, chorus, timed recitation, holiday greetings, etc.; Sound identity lock: the expression always uses the registered sound of the initiator doll; Cannot be viewed in advance: the target child cannot know the expression content in advance, and has the "time blind box" attribute.
[0235] There is no advance preview or "countdown reminder" mechanism in the system to ensure the expression of surprise while supporting long-term appointments (such as "next month's teacher's day to speak to the teacher doll") or recurring appointments (such as "every day before lunch, say 'eat'").
[0236] This mechanism enhances the ritual of expression, making time a bridge for emotional transmission, further extending the dimensions and continuity of "sound social interaction", and also reflecting the flexible control ability of the invention over asynchronous expression mechanisms.
[0237] Embodiment six, emotion tag activation mechanism (emotion blind box)
[0238] This embodiment proposes a mechanism for children to actively specify "emotional style" of expression. Through voice commands, children can set the expression emotion for the content to be broadcast, such as "speak happily", "shout angrily", "read out sadly", "pretend to be shy"… The system adjusts the tone, rhythm, intonation and other parameters of the voice synthesis model accordingly to complete the voice expression with emotional rendering.
[0239] This mechanism does not support automatic recognition or speculation of children's emotional intentions by the system, and the children must explicitly issue emotional instructions in natural language to ensure that the expression intention is completely dominated by the children.
[0240] Example story: Seven-year-old Duo received a voice story response from his friend Guagua: "Yesterday you told that alien dog, it was so funny, I'll write one too!" Duo wanted to respond with something that was "scary and funny", so he said to his doll "Little Hedgehog": "Start with a frightened tone, then laugh happily at the end." The system parses the two emotional paragraphs and coordinates the language generation and voice synthesis models.
[0241] When broadcasting, "Little Hedgehog" first says in a low and terrifying voice: "Do you think yesterday's alien dog was just that? No… Actually, its tail is alive!" Then immediately turns into a light and cheerful laugh: "Ha ha ha, actually I was just making it up, don't be afraid!" Guagua laughed out loud after listening, and then instructed her doll to say in a "whiny" tone: "I'm so scared that I won't even brush my teeth at night." The two of them thus triggered a round of "scary and funny" role-swapping game, interacting back and forth several times, constantly enriching the expression methods and tone levels.
[0242] Mechanism explanation: The emotion blind box mechanism has the following technical features: Voice command sets mood: e.g. "speak in an excited tone" "speak in a whisper" etc. Mood rendering is done automatically by the system: mood tags control the voice synthesis output; Automatic mood inference is not supported: the system does not determine mood based on text or context; Multiple mood segments can be concatenated: the same content can contain multiple mood switches; Tone changes do not change the voice identity: the original doll's voice fingerprint is still used.
[0243] This mechanism enriches the mood dimension of expression, allowing children to move from "content expression" to "mood expression", promoting emotional cognition and empathy development. At the same time, through the "mood tone blind box" setting, the interaction is also more interesting and dramatic.
[0244] Example Seven, Autonomous Expression Task Mechanism (Task Blind Box)
[0245] This example proposes a mechanism in which children set light task instructions through voice, and the recipient responds in an asynchronous manner. The task content is freely set by the initiating child, and the system does not have fixed templates, does not push pre-set tasks, and does not provide task reminders. It only serves as an intermediary for content delivery.
[0246] This mechanism emphasizes the fun, non-mandatory, and social interaction of tasks. The system does not record whether the task is completed, does not provide evaluation and feedback guidance, and all response behaviors are determined and triggered by the recipient child.
[0247] Example Story: Eight-year-old star star said to her friend's doll "Xiao Jiling": "Can you help me design an alien language? Then tomorrow you can tell a story in that language." This task instruction is sent in voice form, and the system presents its content in text to the parent of Gufu for review and approval, and then "Xiao Jiling" broadcasts: "Star star asks you to invent an alien language, and then tell a small story in it!" Gufu thought it was interesting and composed five alien words (such as "Gula" for "hello") and used her doll "Xiao Tiaofa" to tell a story in an alien language.
[0248] The next morning, "Xiao Jiling" in star star's home played Gufu's response before the alarm clock rang: "Gula star star! I am the frog prince of 'Pula star', we are picking candy on the moon..." Star star exclaimed after listening: "It's so cool! I want to invent a language too!" Two people thus entered a "task relay", even inviting other friends to join the "Star Alliance" to create a story about alien civilization.
[0249] Mechanism: Task blind box mechanism has the following core features: Task content free creation: children's voice directly set, no form and difficulty; Asynchronous response mechanism: the recipient can decide whether and when to respond; No system evaluation and tracking: the platform does not set completion marks, progress prompts, and task priorities; One-way trigger and diverse response: one task can trigger multiple rounds of free expression; Parents still only review the text content, not the sound expression.
[0250] The mechanism converts "expression behavior" into "interactive tasks", forming a light social-driven game structure, so that children can create, communicate and explore on their own initiative. Because the system never forces tasks or promotes responses, each response seems more active and surprising, forming a typical "social blind box" interactive experience.
[0251] Example eight, role sound inheritance mechanism (sound blind box)
[0252] This embodiment proposes a mechanism in which children specify "re-telling in the voice of a doll" or "imitating the speaking style of a friend's doll". The system supports converting the expression role sound in the original content into the registered sound of the specified doll for re-interpretation, realizing the inheritance of "sound identity" re-expression.
[0253] In this mechanism, what is inherited is not abstract emotion, tone style or timbre, but the voice fingerprint of the doll registered by the platform, i.e. the identity sound of the role, which is unique and traceable. The system does not provide blurred imitation or free voice change, but only supports formal inheritance and narration between registered role sounds.
[0254] Example story: Nine-year-old Qiqi let her doll "Little Lion" tell a story in the evening expression: "Today in the playground, I pretended to be a time police and punished Xiaoming to stand for 5 seconds, and he laughed." The next day, after listening to the story, her friend Lele told her doll "Little Snowman" with a voice command: "Tell it again in the voice of Little Lion!" The system identifies "Little Lion" as a doll of Qiqi's family, and on the premise that Lele has added Qiqi as a friend, it legally calls the voice fingerprint of "Little Lion" and completes a "voice inheritance narration" on "Little Snowman".
[0255] So "Little Snowman" says in the voice of Little Lion: "I am Little Lion, I am the time police - who does not brush his teeth, will be punished to stand for 3 seconds!" This "inheritance narration" made Qiqi laugh and say in surprise: "How can you use the voice of Little Lion?" Lele replied: "I'm playing 'identity inheritance', now I'm your Little Lion!" They later agreed to "inherit once" from each other every day, taking turns to tell new stories and express their ideas in the voice of the other's doll.
[0256] Mechanism explanation: The voice inheritance blind box mechanism has the following core elements: Voice inheritance is limited between friends: the system verifies social relationships and voice calling permissions; Voice is the unit of identity: registered voices have unique fingerprints and cannot be customized or imitated; Inheritance requires children to issue clear voice commands: such as "repeat once in the voice of so and so"; The system does not generate voice changing models: all voice expressions call existing character registration data; The expression content can be different: not only repetition, but also support for telling new content in that voice.
[0257] This mechanism gives "voice" inheritance and character extension, allowing children to exchange identities, share character perspectives, and perform cross-expression scenarios among friends, further enriching the boundaries of "voice as identity" technology. The blind box feature is that children cannot predict what interesting differences will result from telling stories with their friends' characters before inheritance, creating an exploratory expression experience.
[0258] Example Nine, Compound Expression Triggering Mechanism (Combined Blind Box)
[0259] This embodiment proposes a compound expression mechanism in which children continuously issue multiple expression commands in natural speech, and the system automatically analyzes and sequentially executes them. This mechanism supports multiple command combinations such as "change voice + change tone + delayed broadcast + specify object + mimic saying", allowing children to freely organize spoken expressions without relying on specific grammar or format, and the system automatically analyzes semantic intent and executes step by step through the semantic understanding module.
[0260] The system does not provide optional templates or limit combination methods, allowing children to creatively link multiple expression intents to form a three-dimensional, multi-dimensional expression task.
[0261] Example story: Seven-year-old Duoduo wants her doll "Bubu Bear" to say a motivational phrase to herself in the voice of her friend's doll "Gugu Cat" with a bit of coquettishness at 8 pm. So she issues a series of instructions: "Use Gugu Cat's voice at 8 pm, be a bit coquettish, and tell me 'Duoduo, you're the best, good night and be energetic!'" The system undergoes semantic analysis and identifies: Broadcast time: 8 pm; Use voice: Friend Duoduo's "Gugu Cat"; Emotional requirement: Coquettish; Content text: Duoduo, you're the best, good night and be energetic; Broadcast object: Duoduo (herself); Execution device: "Bubu Bear" (local doll); The system confirms that Duoduo is a friend, has the right to call Gugu Cat's voice, the emotional tone is in line with regulations, the time setting is reasonable, and the parent has approved the voice content after converting it to text.
[0262] At 8 pm, "Bubu Bear" speaks in Gugu Cat's registered voice with a bit of coquettishness, saying to Duoduo: "Duoduo, you're the best, good night and be energetic ~ gurgle gurgle." Duoduo happily jumps up and hugs Bubu Bear: "It's like Gugu Cat has possessed me! It's so much fun!"
[0263] Mechanism explanation: The combination command blind box mechanism has the following key features: Children can express multiple intentions in natural language: no need to learn instruction format; System intelligently analyzes and executes in sequence: ensures that multiple tasks do not conflict; No restrictions on instruction combination types and sequences: children determine the expression path; No active recommendation of combination options: no templates, no pre-set operation framework; Combination content has blind box characteristics: the result varies depending on the combination method and is full of unknowns.
[0264] This mechanism greatly releases the degree of freedom of children's voice expression, transforming expression behavior from linear to structured and from single context to multi-dimensional expression. It also strengthens the core path of "system not leading expression". Each combination is a self-design of the expression path by children, with creativity and exploration, which is a high-level embodiment of the "voice as identity + autonomous expression-driven" philosophy of this invention.
Claims
1. A child asynchronous voice social method supporting role voice mutual simulation and content re-expression, comprising the following steps: 1) A child initiates an expression request to a doll or other edge agent through a voice command; 2) The edge agent collects voice signals and uploads them to a cloud platform; 3) The platform identifies instructions, extracts keywords or expression intentions; 4) A language generation model is called to generate structured expression content; 5) Based on the specified role voice, a voice synthesis model is called to generate a broadcast voice; 6) The voice content is sent to the target edge agent for broadcast; 7) The system returns a delivery status after completing the broadcast task and does not guide subsequent behavior; In the method, all expression behaviors and model calls are triggered by the child's natural voice command, and the platform does not actively provide content suggestions, role recommendations or expression guidance.
2. The method of claim 1, wherein, The expression request supports a chorus control logic, and the child can specify a doll as the lead singer and other dolls as the backup singers. The lead singer broadcasts, and the backup singers follow in turn, achieving a multi-role chorus effect.
3. The method of claim 1, wherein, The first doll can call the role voice of the second doll added as a friend to broadcast when performing an expression task. The role voice is registered and recorded by the platform and is only allowed to be used in the social graph. It does not support simulating non-friend roles or real person voices.
4. The method of claim 1, wherein, Parents only have the right to audit text content, and this audit only applies to the "before sending" process. All voice selection, listening and expression content control is completely in the hands of the child, and the system does not provide a parent voice review interface.
5. The method of claim 1, wherein, The language generation model supports light input triggering. After the child inputs keywords, phrases or spoken expressions, the system automatically generates complete and semantic complete expression content.
6. The method of claim 1, wherein, The language generation model has non-deterministic generation capability and can output multiple different versions under the same input. It has the ability to expand the plot, adapt to the context and express randomness, which is used to enhance the interest and exploration of the expression content.
7. The method of claim 1, wherein, The language generation model and voice synthesis model are called in series. The calling chain is triggered by the child's natural voice command, and the system does not have the functions of pre-calling, pre-generating or default loading expression content.
8. The method of claim 1, wherein, The system does not provide a prompt feedback button or guide voice after the expression is completed. The feedback mechanism is not actively awakened by the platform, and the child can decide whether to initiate a response. The response behavior can trigger a new round of social interaction.
9. The method of claim 1, wherein, The platform does not pre-set expression templates, content sentence patterns or task goals, and does not push expression suggestions through interfaces, voices, animations, etc. The expression path is completely controlled by the child's free will, preventing the platform from dominating expression behavior.