Intelligent toy guiding method and system based on man-machine interaction
By using a human-computer interaction-based intelligent toy guidance method and system, and leveraging multimodal data perception and dynamic analysis, the shortcomings of existing intelligent toys in individual adaptation are addressed, enabling personalized and precise guidance and improving educational effectiveness and user experience.
Patent Information
- Application Number
- CN202511036136.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-26
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-07-26
AI Technical Summary
Existing smart toys lack the ability to dynamically perceive and adapt to individual user characteristics, resulting in poor guidance effects. Users are prone to getting stuck in their comfort zone, making it difficult to access new knowledge or adjust guidance strategies when encountering difficulties, leading to frustration and loss of interest.
The intelligent toy guidance method and system based on human-computer interaction dynamically adjusts guidance strategies by combining multimodal data perception (voice, gestures, posture, and text) with data storage, demand analysis, and pattern determination modules. It accurately identifies user weaknesses and flexibly switches guidance directions to achieve personalized and precise guidance.
It enhances the educational value and user engagement of smart toys, dynamically responds to user needs through multi-dimensional data analysis and multi-modal adaptation, avoids user stagnation and frustration in their comfort zone, and promotes comprehensive ability development.
Smart Images

Figure CN120929673A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of human-computer interaction and intelligent toy technology, and particularly relates to a toy intelligent guidance method and system based on human-computer interaction. Background Technology
[0002] With the rapid development of artificial intelligence technology, smart toys, as an important carrier integrating entertainment and education functions, are gradually becoming common companion tools in children's growth process. Although smart toys on the market generally have certain interactive guidance functions, they still have significant limitations in terms of the scientific nature and adaptability of the guidance mechanism, making it difficult to fully realize their educational auxiliary value.
[0003] Most existing smart toys rely on preset fixed processes or single interaction modes for their guidance functions, lacking the ability to dynamically perceive and adapt to individual user characteristics. These products are typically designed with standardized guidance content based on a general user profile, such as knowledge modules divided by age group or fixed-sequence skill training processes. While this approach can achieve basic interactive guidance, it ignores individual differences in users' cognitive levels, learning habits, and interests, resulting in significantly reduced guidance effectiveness.
[0004] From a practical application perspective, the limitations of existing technologies are mainly reflected in the following aspects. Firstly, existing smart toys generally rely excessively on users' comfort zone preferences. These products often quickly identify users' preferred areas through simple initial interest tests or single-dimensional interaction data, and then continuously push highly familiar content. For example, if a child shows interest in number games in initial interactions, the smart toy may focus on pushing mathematical content for a long time, neglecting guidance in other important ability dimensions such as language, logic, and art. While this approach may increase the frequency of user interaction in the short term, in the long run, it will cause users to remain in a low-challenge zone, making it difficult for them to access new knowledge points or weak skills, severely limiting the comprehensive development of their abilities. For children in their critical developmental period, this "comfort zone solidification" phenomenon may solidify their knowledge structure and form a one-sided ability development pattern.
[0005] On the other hand, existing technologies lack necessary flexible guidance mechanisms, exhibiting significant deficiencies in the flexibility of guidance strategies. While some smart toys attempt to introduce challenging content, they fail to adjust the guidance direction and difficulty in a timely manner when users show resistance or learning difficulties. For example, when children repeatedly make mistakes or experience delayed interactive responses in language learning modules, existing systems often continue to proceed according to the preset process, even repeatedly reinforcing more difficult content. This approach easily leads to user frustration, resulting in interrupted interaction or loss of interest. Survey data shows that over 60% of children actively terminate the learning process when using smart toys due to persistent frustration, fully reflecting the shortcomings of existing guidance mechanisms in maintaining user experience.
[0006] Further analysis reveals that existing smart toy guidance models are essentially passive, employing a "one-size-fits-all" or "one-way reinforcement" approach, lacking the ability to accurately capture and dynamically respond to users' true needs. These systems are unable to comprehensively analyze users' ability characteristics through multi-dimensional data, nor can they adjust guidance strategies based on real-time user feedback, resulting in a significant disconnect between guidance content and actual user needs. For example, in skills training scenarios, existing systems cannot distinguish whether users reduce interaction due to a lack of interest or insufficient ability, often adopting a uniform reinforcement strategy. This fails to effectively improve training outcomes in weak areas or fully leverage the potential of users' strengths.
[0007] This technological limitation not only reduces the educational value of smart toys but also restricts user stickiness and market competitiveness. As parents' demands for personalized education for their children continue to rise, traditional standardized guidance models can no longer meet user expectations. Summary of the Invention
[0008] To address the aforementioned technical problems, this invention proposes a toy intelligent guidance method and system based on human-computer interaction.
[0009] In a first aspect of the present invention, a toy intelligent guidance method based on human-computer interaction is proposed, the method being applied to intelligent toys;
[0010] The method includes:
[0011] Based on the human-computer interaction data of the smart toy during a preset historical time period, the target guidance needs of the current user are determined.
[0012] Based on the target guidance requirements, determine the current data perception mode of the smart toy;
[0013] The current data awareness mode is activated to guide the behavior of the current user.
[0014] The intelligent toy has multimodal data sensing capabilities;
[0015] The current data perception mode includes at least one of voice perception, gesture perception, posture perception, and text perception.
[0016] In one configuration setting, determining the current user's target guidance needs based on the human-computer interaction data of the smart toy over a preset historical time period includes:
[0017] Statistics are collected on the current user's behavioral preferences, training completion rate, and interaction frequency for each preset category of guidance needs within the preset historical time period.
[0018] The guidance needs with the lowest behavioral preferences, the lowest training completion rate, or the lowest interaction frequency are taken as the current user's primary guidance needs.
[0019] The guidance needs with the highest behavioral preferences, the highest training completion rate, or the highest interaction frequency are taken as the second current guidance needs of the current user.
[0020] Using the second current boot request or the first current boot request as the target boot request specifically includes:
[0021] First, the first current guidance requirement is set as the target guidance requirement. Then, when a preset condition is met, the system switches to the second current guidance requirement as the target guidance requirement.
[0022] In another configuration, determining the current user's target guidance needs based on the smart toy's human-computer interaction data over a preset historical time period includes:
[0023] Analyze user behavior preferences, training completion rate, and interaction frequency in the human-computer interaction data within the preset historical time period, and output the target guidance requirement based on the preset guidance requirement classification model;
[0024] The step of determining the current user's target guidance needs based on the human-computer interaction data of the smart toy over a preset historical time period also includes:
[0025] If the human-computer interaction data within the preset historical time period is empty, then the initial guidance requirement is determined as the current guidance requirement based on the user's initial age information and training target.
[0026] The guidance demand classification model is trained based on user age, cognitive level, and common training objectives.
[0027] Determining the current data perception mode of the smart toy based on the target guidance requirements includes:
[0028] If the target guidance requirement is language expression training, then the current data perception mode is determined to be speech perception and text perception;
[0029] If the target guidance requirement is limb coordination training, then the current data perception mode is determined to be gesture perception and posture perception;
[0030] If the target guidance requirement is comprehensive ability training, then the current data perception mode is determined to be at least two of the following: voice perception, gesture perception, posture perception, and text perception.
[0031] The step of activating the current data awareness mode and guiding the behavior of the current user includes:
[0032] The user's behavioral data is collected in real time using the current data perception mode.
[0033] The collected behavioral data is compared with preset standard behavioral data to generate comparison results;
[0034] Based on the comparison results, guidance instructions are output to the user through voice prompts, flashing lights, or mechanical action demonstrations.
[0035] Based on the comparison results, the step of outputting guidance instructions to the user through voice prompts, flashing lights, or mechanical demonstrations includes:
[0036] If the comparison result shows that the user's behavior conforms to the standard behavior data, then an encouraging guidance instruction will be output.
[0037] If the comparison result shows that the user's behavior deviates from the standard behavior data, a corrective guidance instruction will be output, and the standard behavior will be re-demonstrated.
[0038] In a second aspect of the invention, to implement the method described in the first aspect, a toy intelligent guidance system based on human-computer interaction is proposed, the system comprising:
[0039] The data storage module is used to store human-computer interaction data for a preset historical time period;
[0040] The requirements analysis module is used to determine the current user's current guidance requirements based on the human-computer interaction data in the data storage module.
[0041] The mode determination module is used to determine the current data perception mode of the smart toy based on the current guidance requirements; the guidance execution module is used to activate the current data perception mode and guide the behavior of the current user.
[0042] The intelligent toy has multimodal data perception capabilities, and the current data perception mode includes at least one of voice perception, gesture perception, posture perception, and text perception.
[0043] The toy intelligent guidance system based on human-computer interaction of the present invention demonstrates significant advantages in multiple dimensions through the coordinated operation of four major modules: data storage, demand analysis, pattern determination, and guidance execution, combined with a data-driven dynamic guidance method.
[0044] The data storage module provides a solid "memory foundation" for the system. By storing human-computer interaction data for preset historical time periods, it ensures that the demand analysis module can comprehensively trace user interaction trajectories from multiple dimensions such as behavioral preferences, training completion rate, and interaction frequency. This avoids the limitations of existing technologies that rely on single initial data or general templates, and provides complete data support for accurately identifying user characteristics.
[0045] Based on historical data, the demand analysis module uses quantitative statistics to determine the user's first current guidance need (weakness) and second current guidance need (strength). It then uses this to lock in the target guidance needs. This analysis method breaks through the traditional "one-size-fits-all" guidance model. It can accurately locate the user's weak links that need to be strengthened, and encourage them to step out of their comfort zone and explore new knowledge. At the same time, it can also take into account the deepening of their strengths, achieving a balance between "filling gaps" and "cultivating excellence". It effectively solves the problem of existing technologies over-focusing on comfort zones or blindly pushing challenging content.
[0046] The mode determination module flexibly selects at least one data perception mode, such as voice, gesture, posture, or text, based on the target guidance needs. By leveraging the multimodal perception function of smart toys, it adapts to the interaction habits and scenario needs of different users. For example, voice and gesture perception are prioritized for young children to lower the interaction threshold, while text perception is enabled for users with text ability to improve guidance efficiency, greatly enhancing the flexibility and inclusiveness of the interaction.
[0047] The guidance execution module dynamically executes guidance based on a defined perception pattern. When users encounter difficulties or show signs of resistance such as decreased interaction frequency, it can promptly switch from guidance on weaker areas to guidance on stronger areas. This flexible adjustment maintains user interest and avoids the interaction interruption problem caused by the lack of dynamic response in existing technologies. The synergy of the four modules and the combination of core methods enable the system to achieve personalized and precise guidance through data-driven approaches, while balancing challenge and sense of accomplishment through multimodal adaptation and flexible mechanisms. Ultimately, this significantly enhances the educational value and user engagement of smart toys.
[0048] Further specific advantages and implementation principles of the present invention will be further detailed in the specific embodiments section in conjunction with the accompanying drawings. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 This is a schematic diagram of the main execution flow of a toy intelligent guidance method based on human-computer interaction according to an embodiment of the present invention;
[0051] Figure 2 yes Figure 1 A further preferred embodiment of the method is illustrated in the diagram;
[0052] Figure 3 This is a schematic diagram of the hardware unit composition of a robotic arm precision positioning system according to an embodiment of the present invention;
[0053] Figure 4 This is a schematic diagram of a scenario in the practical application of a robotic arm precision positioning system according to an embodiment of the present invention; Detailed Implementation
[0054] In the specific embodiments of this application, if the embodiments of the relevant technical solutions involve user-related data, then when the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0055] Before introducing specific embodiments of the present invention, the existing technology and related problems of the present invention will be introduced first, so as to lead to the technical solution of the present invention, so as to better understand the improvement advantages and inventiveness of the present application.
[0056] Taking common English learning smart toys on the market as an example, these products typically have a fixed curriculum system preset according to age groups. 3-4 year olds are given a uniform introduction to letter recognition content, while 5-6 year olds focus on basic sentence structure training. When a 5-year-old shows great interest in letter recognition but frequently gets stuck in sentence structure training, the system continues to advance the sentence structure content at a fixed pace. This fails to maintain the child's initial enthusiasm for letter learning, nor does it specifically reduce the difficulty of the sentence structures. This "one-size-fits-all" approach completely ignores the individual differences in children's language sensitivities, causing some children to give up due to continuous frustration, while others become bored due to repetitive content.
[0057] Similar problems exist with math thinking training smart toys. One popular number calculation robot's guidance logic involves pushing questions of increasing difficulty. If a child answers the same type of question incorrectly three times in a row, the system automatically pushes the same type of question repeatedly. In actual use, if a child answers a simple addition problem incorrectly due to lack of concentration, the system will continuously reinforce addition training, failing to recognize that the real problem is a lack of focus rather than a lack of calculation ability. This one-dimensional judgment mechanism not only fails to effectively improve math skills but also makes children resistant to number calculations. Research shows that this loss of interest due to improper guidance accounts for as much as 65% of users of math smart toys.
[0058] In the field of early art education, the limitations of existing smart toys are even more pronounced. One smart drawing board for art instruction uses a camera to recognize children's scribbles and provide feedback, but it can only recognize 20 preset common patterns. When children try to create custom images, the system frequently prompts "Please draw the standard pattern" because it cannot recognize the image. This mechanical guidance severely inhibits children's creativity. More importantly, the system cannot record children's drawing preferences, consistently allocating guidance time evenly across modules such as oil painting, watercolor, and sketching. This results in children who excel at watercolor not receiving in-depth training, while children who dislike sketching are forced to repeatedly encounter content they are not interested in.
[0059] The core flaw in existing technology lies in the lack of a dynamic adaptation mechanism. A certain smart toy, which focuses on multidisciplinary early childhood education, claims to have a "personalized recommendation" function, but it can only push content based on the interest test results at the time of initial use. In actual use, if a child initially selects the science experiment module, the system will continuously push related content, even if the child clearly shows a greater interest in history stories three months later, the system will not adjust its push direction. This guidance model based on static data completely fails to capture the dynamic changes in user interests.
[0060] Existing smart toys also fall short in terms of flexible mechanisms to address user difficulties. One programming robot, when a child makes a mistake, simply displays "Program error, please rewrite," without analyzing the reason for the error. Even after a child repeatedly fails due to logical confusion, the system persists at the original difficulty level, ultimately leading 70% of children to abandon programming after two weeks. This inflexible approach equates "challenge" with "frustration," violating fundamental principles of educational guidance.
[0061] From a technical perspective, existing smart toys employ a passive, responsive design for guidance, lacking both multi-dimensional data collection capabilities and intelligent analysis engines. They cannot judge a child's emotional changes through voice tone, identify attention levels based on gestures, or analyze real needs through interaction frequency. This technological limitation results in a severe disconnect between guidance content and user needs, failing to encourage children to step out of their comfort zones and explore new knowledge, nor allowing for timely adjustments to strategies when encountering difficulties.
[0062] In response, this invention proposes a corresponding technical solution.
[0063] First see Figure 1 , Figure 1 This diagram illustrates the main execution flow of a toy intelligent guidance method based on human-computer interaction according to an embodiment of the present invention.
[0064] The method is applied to smart toys and includes the following steps (for brevity in subsequent descriptions, each step is numbered S1-S3, and the numbers are omitted in the accompanying drawings).
[0065] S1: Based on the human-computer interaction data of the smart toy during a preset historical time period, determine the current user's target guidance needs;
[0066] S2: Based on the target guidance requirements, determine the current data perception mode of the smart toy;
[0067] S3: Activate the current data awareness mode to guide the behavior of the current user;
[0068] The intelligent toy has multimodal data sensing capabilities;
[0069] The current data perception mode includes at least one of voice perception, gesture perception, posture perception, and text perception.
[0070] In step S1, based on the human-computer interaction data of the smart toy over a preset historical time period, the current user's target guidance needs are determined. A specific embodiment may be:
[0071] Analyze user behavior preferences, training completion rate, and interaction frequency in the human-computer interaction data within the preset historical time period, and output the target guidance requirement based on the preset guidance requirement classification model;
[0072] The guidance demand classification model is trained based on user age, cognitive level, and common training objectives.
[0073] Taking a 5-year-old child named Xiaoming using a certain smart early childhood education toy as an example, the preset historical time period is the most recent 30 days, and the human-computer interaction data covers Xiaoming's interaction records in four preset guidance categories: "language expression", "mathematical logic", "artistic creation" and "life habits".
[0074] First, the system performs multi-dimensional quantitative statistics on the data:
[0075] Behavioral preferences: By recording the number of times each module was actively triggered and the duration of continuous use, it was found that Xiaoming actively chose the "Language Expression" module (20 minutes per day) and the "Artistic Creation" module (15 minutes per day) the most frequently, while the "Mathematical Logic" module was only actively triggered 5 times, with a single usage time of less than 5 minutes, making it the category with the lowest behavioral preference.
[0076] Training completion rate: The system records the completion status of each module's tasks. The completion rate of the nursery rhyme reading task in the "Language Expression" module is 85%, and the completion rate of the brushing steps training in the "Life Habits" module is 90%. However, the completion rate of the addition and subtraction task within 10 in the "Mathematical Logic" module is only 40%, which is the lowest.
[0077] Interaction frequency: Statistics show that the "Artistic Creation" module has an average of 30 interactions per day, while the "Mathematical Logic" module has only 8 interactions per day, making it the module with the lowest interaction frequency.
[0078] Subsequently, a demand classification model was introduced into the analysis. This model has been trained based on the cognitive level of 5-6 year old children (such as primarily concrete thinking and emerging logical reasoning ability), common training objectives (such as basic arithmetic and language organization), and massive amounts of user data, and has the ability to match needs with user characteristics.
[0079] The model first inputs Xiaoming's quantitative data: age 5, cognitive level assessment as "logical thinking needs to be strengthened", and the core indicators mentioned above: "lowest behavioral preference in mathematical logic module, training completion rate 40% (lowest), interaction frequency 8 times / day (lowest)".
[0080] The model, through feature matching, discovered that 5-year-old children are in a critical period of logical thinking development and require targeted guidance. Xiaoming's three-dimensional data all pointed to a weakness in the mathematical logic module. Meanwhile, modules such as "artistic creation," which are considered strengths (although the interaction frequency is high, the completion rate is only 75%, indicating it's not an area requiring urgent reinforcement), were excluded. The final model output identified "mathematical logic guidance" as Xiaoming's primary current guidance need, i.e., the target guidance need.
[0081] This process avoids the one-sidedness of existing technologies that rely solely on a single dimension (such as only looking at usage time), and by incorporating age and cognitive characteristics into the model, it ensures that the target-oriented needs are aligned with the user's real weaknesses and conform to the educational patterns of their developmental stage.
[0082] In another alternative embodiment, step S1 further includes:
[0083] If the human-computer interaction data within the preset historical time period is empty, then the initial guidance requirement is determined as the current guidance requirement based on the user's initial age information and training target.
[0084] For example, when 4-year-old Lily uses a smart early education toy for the first time, the system detects that there is no interaction data record in the past 30 days, triggering the initial guidance needs assessment process. At this time, the system guides the parents to complete the initial information input through voice interaction:
[0085] Age information: Parents enter "4 years old";
[0086] Training goal: Parents choose "prioritize improving language expression skills, while also cultivating good living habits".
[0087] After receiving the information, the system invokes preset initial guidance needs matching rules, which are designed based on the core developmental goals and common training priorities for different age groups. For 4-year-old children, language expression ability is in a period of rapid vocabulary accumulation (core developmental goal), and the cultivation of living habits needs to be combined with concrete guidance (age-appropriate characteristics).
[0088] First, the system uses "4 years old" as the core parameter and matches it with the corresponding basic guidance framework—the preset guidance category weights for 4-year-old children are: language expression (35%), daily habits (30%), cognitive development (20%), and physical coordination (15%). Then, based on the training goal input by parents, "prioritizing language expression while also considering daily habits," the weights are dynamically adjusted: the weight of the language expression module is increased to 45%, the daily habits module is maintained at 30%, and the weights of other modules are reduced accordingly.
[0089] The guidance needs classification model further incorporates the cognitive characteristics of 4-year-old children (such as primarily learning through context and having an attention span of about 15 minutes). It selects suitable content such as "daily dialogue practice" and "short sentence creation" from the language expression module, and concrete tasks such as "steps for independent dressing" and "reminders for washing hands before meals" from the life habits module.
[0090] The final system outputs the initial guidance requirements: with "contextualized language expression training" as the core guidance content, and simultaneously interspersed with "life habit scenario simulation" tasks, forming an initial guidance plan that takes into account both user input goals and age development patterns.
[0091] This mechanism solves the problem of blank guidance for new users when there is no historical data. Through structured initial information input and age group adaptation rules, it ensures that users can receive scientific and reasonable guidance on the first use. It avoids the blindness of pushing general content in the absence of data in existing technologies and lays the foundation for the accumulation of subsequent interaction data.
[0092] However, the above approach may still have some problems, namely, users may easily get stuck in their comfort zone and be unable to adjust to new behaviors, unable to encourage children to break out of their comfort zone and explore new knowledge, and unable to adjust strategies in a timely manner when encountering difficulties.
[0093] In this regard, the present invention further proposes preferred embodiments, see below. Figure 2 , Figure 2 Show Figure 1 A further preferred embodiment of the method is illustrated in the diagram ( Figure 2 (The step numbers are also omitted in the text).
[0094] exist Figure 2 In this context, determining the current user's target guidance needs based on the human-computer interaction data of the smart toy over a preset historical time period includes:
[0095] Statistics are collected on the current user's behavioral preferences, training completion rate, and interaction frequency for each preset category of guidance needs within the preset historical time period.
[0096] The guidance needs with the lowest behavioral preferences, the lowest training completion rate, or the lowest interaction frequency are taken as the current user's primary guidance needs.
[0097] The guidance needs with the highest behavioral preferences, the highest training completion rate, or the highest interaction frequency are taken as the second current guidance needs of the current user.
[0098] The second current boot request or the first current boot request is taken as the target boot request.
[0099] Specifically, using the second current boot request or the first current boot request as the target boot request includes:
[0100] First, the first current guidance requirement is set as the target guidance requirement. Then, when a preset condition is met, the system switches to the second current guidance requirement as the target guidance requirement.
[0101] For details, see Figure 2 A further preferred example of the method is as follows:
[0102] S11: Statistics on the current user's behavioral preferences, training completion rate, and interaction frequency for each preset category of guidance needs within a preset historical time period;
[0103] S12: The guidance needs with the lowest behavioral preferences, lowest training completion rate, or lowest interaction frequency are taken as the first current guidance needs of the current user; the guidance needs with the highest behavioral preferences, highest training completion rate, or highest interaction frequency are taken as the second current guidance needs of the current user.
[0104] S13: Using the first current guidance requirement as the target guidance requirement, determine the current data perception mode of the smart toy;
[0105] S14: Activate the current data perception mode to guide the behavior of the current user;
[0106] S15: During the process of guiding the behavior of the current user, predict the probability value of the duration of the user's interaction in real time;
[0107] S16: Determine whether the interaction duration probability value is less than the threshold. If so, switch to the second current guidance requirement.
[0108] S17: Using the second current guidance requirement as the target guidance requirement, determine the current data perception mode of the smart toy; return to step S14.
[0109] It can be seen that satisfying the preset condition means that the interaction duration probability value is less than the threshold.
[0110] The interaction persistence probability value is used to characterize the probability that the current user will continue with the current behavior guidance mode.
[0111] certainly, Figure 2 Although not shown in the diagram, in step S17, after returning to step S14, when the user completes the second current guidance requirement, the method ends, or switches back to the first current guidance requirement; the first current guidance requirement is used as the target guidance requirement to determine the current data perception mode of the smart toy; and the process returns to step S14.
[0112] As a further preferred option, in step 12, "taking the guidance request with the lowest behavioral preference, the lowest training completion rate, or the lowest interaction frequency as the current user's first current guidance request" can be implemented based on the first sub-container; while "taking the guidance request with the highest behavioral preference, the highest training completion rate, or the highest interaction frequency as the current user's second current guidance request" can be implemented based on the second sub-container; the two processes are executed in parallel.
[0113] The first and second child containers both belong to the main container and communicate through the channel provided by the main container. This allows for rapid execution when switching between processes (step S16), avoiding resource delays caused by inter-process switching.
[0114] The above-mentioned preferred configuration is a further improvement of the present invention, and its key means and advantages mainly include:
[0115] The main container, as the global management unit, is responsible for coordinating the operation of the first and second sub-containers. Its core functions include:
[0116] Process scheduling and resource allocation: Allocate independent computing resources (such as CPU cores and memory space) to the main container and child containers to ensure that the processes of child containers do not interfere with each other when running; preset resource priority rules ensure that the resource supply of the currently active child containers is guaranteed first when switching is required for booting.
[0117] Communication channel management: Built-in shared memory space and lightweight message queue serve as communication channels. The main container uses this channel to synchronize user interaction data (such as behavioral preferences and training completion), sub-container status (such as current running mode and resource utilization) and switching instructions in real time, with communication latency controlled within 100ms.
[0118] Status monitoring and fault tolerance: Real-time monitoring of the running status of sub-containers (such as whether the process is abnormal or whether the resource usage exceeds the limit). When a sub-container failure is detected, the main container automatically triggers the restart mechanism and restores the most recent data to ensure that the boot process is not interrupted.
[0119] The first sub-container is specifically responsible for handling the "first current boot request" (user vulnerability), and its implementation includes:
[0120] Preload core resources: At startup, multimodal perception models related to weaknesses (such as gesture recognition models required for mathematical logic training), standard behavior databases (such as addition and subtraction operation step templates), and guidance instruction libraries (such as error correction voice packs and demonstration motion trajectory data) are preloaded. After the resources are loaded, the system is in a "ready state" and waits for the main container to activate the system.
[0121] Independent operation and data closed loop: In the active state, it independently executes the entire process of data acquisition (by calling the corresponding sensing hardware interface), comparison and analysis (comparing with preloaded standard data), and guidance output (generating correction instructions). The process data is written to the main container's shared memory in real time for the main container to monitor the progress.
[0122] The second sub-container is specifically designed to handle the processing of the "second current boot request" (user advantage), and implementation details include:
[0123] Lightweight resource pre-storage: Optimized perception models (such as high-sensitivity speech recognition models), high-completeness standard data (such as high-quality pronunciation templates), and encouragement resource libraries (such as diverse praise voices and animation lighting effect parameters) are pre-stored for user strengths (such as language expression). The resource consumption is reduced by 30% compared to the first sub-container, ensuring fast response.
[0124] Low-power standby and fast activation: When not activated, it is in a low-power standby state, with only the core process listening to the main container instructions; when a switching instruction is received (step S16), the resource wake-up is completed within 100ms and the boot process is taken over, and the latest user status data is read from the shared memory to achieve seamless connection.
[0125] When the system determines that it needs to switch from the first current boot request to the second current boot request, the process is as follows:
[0126] The main container sends a "pause command" to the first child container through the communication channel. The first child container writes the current boot progress (such as the training steps that the user has completed and the current comparison results) into the shared memory and then enters a sleep state.
[0127] The main container calls the resource scheduling interface to allocate the high-priority resources (such as camera computing power) occupied by the first sub-container to the second sub-container;
[0128] The main container sends an "activation command" to the second child container. The second child container reads user state data from shared memory, loads pre-stored advantageous resources, and starts the corresponding perception mode (such as voice perception + text perception) within 50ms and outputs a guiding command (such as "Let's play your best children's song creation game!").
[0129] The main container synchronizes the switching results to the data storage module in real time and updates the current boot status label.
[0130] The core advantages of adopting the above containerized architecture (with a main container containing sub-containers) are mainly reflected in:
[0131] 1. Zero-latency process switching improves boot continuity.
[0132] By preloading resources in the sub-container and communicating with the main container via shared memory, the response time for bootstrapping requests is reduced from over 500ms in traditional process switching to less than 150ms. This avoids distraction of users during the switching process, making it especially suitable for the short attention span of young children.
[0133] 2. Resource isolation and efficient utilization
[0134] The main container's fine-grained allocation of computing resources ensures that high-energy-consuming processes in the first sub-container (weak areas) (such as posture perception) and lightweight processes in the second sub-container (strong areas) (such as speech perception) do not compete for resources. This guarantees the accuracy requirements for training weak areas while reducing the energy consumption of guiding strong areas. 3. System stability is significantly enhanced.
[0135] The independent operation mechanism of sub-containers ensures that the failure of a single boot process will not affect the overall system, and the fault tolerance mechanism of the main container can quickly repair anomalies, solving the problem of "one failure interrupting the entire process" in the traditional single-process architecture and improving the continuous operation capability of smart toys.
[0136] 4. Scalability adapts to diverse needs
[0137] The main container supports flexibly increasing or decreasing the number of sub-containers (such as adding a third sub-container to handle comprehensive capability training). Each sub-container can independently upgrade the perception model or resource library without reconstructing the overall system, providing architectural support for future expansion into more guidance categories (such as programming enlightenment and social etiquette).
[0138] Next, we will use specific examples to explain the specific implementation of each step in the above preferred example.
[0139] Taking a 5-year-old child, Xiaoming, using an "intelligent early childhood education toy" as an example, the preset guidance needs are categorized as "language expression," "mathematical logic," "artistic creation," and "life habits," with the preset historical time period being the most recent 30 days. The specific steps are as follows:
[0140] S11: Core Indicators for Statistical Historical Interaction Data
[0141] The system performs multi-dimensional statistical analysis on Xiaoming's interaction data over 30 days and outputs quantitative results for each guidance category:
[0142] Language expression: 25 proactive triggers (high behavioral preference), 90% completion rate of ancient poetry recitation task (high training completion rate), 40 daily voice interactions (high interaction frequency);
[0143] Mathematical logic: 5 active triggers (low behavioral preference), 40% completion rate for addition and subtraction tasks within 10 (low training completion), 8 interactions per day (low interaction frequency);
[0144] Artistic creation: 18 proactive triggers, 75% completion rate for graffiti and coloring, and an average of 25 interactions per day;
[0145] Lifestyle habits: 12 active triggers, 85% completion rate of brushing steps training, and an average of 15 interactions per day.
[0146] S12: Determine the first and second current guidance requirements
[0147] Based on the statistical results of S11, the system selects key indicators:
[0148] The primary guiding requirement is mathematical logic (lowest behavioral preference, lowest training completion rate, and lowest interaction frequency).
[0149] The second current guidance requirement is language expression (highest behavioral preference, highest training completion rate, and highest interaction frequency).
[0150] S13: Determine the perception mode based on the first guiding requirement.
[0151] The system uses "mathematical logic" as the target to guide demand, and combines the concrete thinking characteristics of 5-year-old children to select an appropriate perception mode: voice perception (real-time response to answer voice) + gesture perception (recognizing the action of pointing fingers to count number cards).
[0152] S14: Enable mode to execute boot
[0153] The smart toy activates its voice and gesture recognition mode to begin guiding mathematical logic:
[0154] The voice prompt asks: "Xiaoming, what is 3 plus 2? Point to the answer on the card!"
[0155] Gesture recognition: The camera captures Xiaoming's finger pointing at the number card and provides real-time feedback such as "Yes, 5 is the correct answer!"
[0156] S15: Real-time prediction of interaction duration probability
[0157] During the guidance process, the system calculates the probability of interaction duration in real time using multi-dimensional data:
[0158] The system detected that Xiaoming delayed answering questions three times in a row (with intervals exceeding 10 seconds), his tone of voice changed from excited to depressed, and his gestures were slow. The system determined that the probability of the interaction lasting was 60%.
[0159] S16: Determine if the probability is below the threshold.
[0160] The preset interaction duration probability threshold is 70%. Since 60% < 70%, the system triggers a switching mechanism and decides to switch from the first guidance requirement to the second guidance requirement.
[0161] S17: Switch to the second boot requirement and cycle through the boot process.
[0162] The system uses "language expression" as a new target to guide demand and matches the perception mode: speech perception (recognizing and repeating speech) + text perception (displaying simple pinyin text);
[0163] Return to S14 to execute the guide: play the children's song "Two Tigers", display the pinyin text on the screen, recognize Xiaoming's pronunciation through voice recognition and encourage him with "Your pronunciation is really standard, try this sentence again~";
[0164] Once Xiaoming has finished reading 3 nursery rhymes (the second guiding requirement is completed), the system switches back to the first guiding requirement: "Let's play some more math games~", restarts the voice + gesture perception mode, and returns to S14 to continue math logic training.
[0165] Through the above cycle, the system not only helps Xiaoming overcome his weaknesses in mathematical logic, but also maintains his interest by using language expression in his area of strength when he becomes resistant, thus achieving a positive guidance loop of "challenge-encouragement-re-challenge".
[0166] In the above embodiments, determining the current data perception mode of the smart toy based on the target guidance requirements includes:
[0167] If the target guidance requirement is language expression training, then the current data perception mode is determined to be speech perception and text perception;
[0168] If the target guidance requirement is limb coordination training, then the current data perception mode is determined to be gesture perception and posture perception;
[0169] If the target guidance requirement is comprehensive ability training, then the current data perception mode is determined to be at least two of the following: voice perception, gesture perception, posture perception, and text perception.
[0170] Preferably, the corresponding multimodal perception combination can be accurately matched according to the type of target guidance needs;
[0171] When the target guidance requirement is language expression training (such as oral dialogue, pinyin practice), enable speech perception (real-time collection of user pronunciation and intonation) and text perception (recognition of user's handwritten pinyin or text);
[0172] When the target guidance requirement is physical coordination training (such as dance movement imitation, manual step learning), activate gesture perception (capturing hand movement trajectory) and posture perception (identifying body posture angle through camera);
[0173] When the goal is to guide comprehensive ability training (such as situational story performance or scientific experiment operation), at least two perception modes should be flexibly combined, such as voice perception (command interaction) + gesture perception (action operation) + posture perception (scene coordination).
[0174] In the above embodiments, the perception mode is accurately matched with the needs, avoiding the activation of invalid perception modes, reducing resource waste, and improving the targeting and accuracy of data collection. At the same time, based on the flexible and efficient achievement of the goal through multimodal combination, the optimal perception method is matched for the characteristics of different ability training. For example, language training focuses on "listening + writing", and physical training focuses on "movement + form", thereby improving guidance efficiency.
[0175] The step of activating the current data awareness mode and guiding the behavior of the current user includes:
[0176] The user's behavioral data is collected in real time using the current data perception mode.
[0177] The collected behavioral data is compared with preset standard behavioral data to generate comparison results;
[0178] Based on the comparison results, guidance instructions are output to the user through voice prompts, flashing lights, or mechanical action demonstrations.
[0179] As a specific example, this step specifically includes:
[0180] Real-time collection of behavioral data: Capture user dynamics through established perception patterns, such as collecting speech waveform data and text writing trajectory in language training, and collecting gesture coordinate changes and body posture angles in body training.
[0181] Data comparison results generation: The collected real-time data is compared with the system's preset standard data (such as standard pronunciation templates and correct dance movement trajectory library) using an algorithm, and the results of "meets the standard" or "deviates from the standard" are output.
[0182] Output guidance instructions: Select the corresponding feedback method based on the comparison results, such as outputting voice prompts through the built-in speaker, flashing different colored lights with LED lights, or demonstrating standard movements with mechanical joints.
[0183] This process achieves dynamic data closed-loop feedback, with real-time response throughout the entire process from data collection to comparison and guidance, ensuring that user behavior deviations can be corrected in a timely manner; and the feedback methods are intuitive and diverse, combining voice, light, mechanical actions and other multi-dimensional feedback to adapt to the perceptual preferences of different users and improve the comprehensibility of the guidance.
[0184] Based on the comparison results, the step of outputting guidance instructions to the user through voice prompts, flashing lights, or mechanical demonstrations includes:
[0185] If the comparison result shows that the user's behavior conforms to the standard behavior data, then an encouraging guidance instruction will be output.
[0186] If the comparison result shows that the user's behavior deviates from the standard behavior data, a corrective guidance instruction will be output, and the standard behavior will be re-demonstrated.
[0187] As an example, if the comparison result shows that the behavior meets the standard (e.g., pronunciation accuracy ≥ 90%, motion trajectory overlap ≥ 85%), the system outputs encouraging instructions, such as playing the voice message "Great! The pronunciation is very standard," flashing green lights, and accompanied by applause sound effects. If the comparison result shows that the behavior deviates from the standard (e.g., pronunciation error, motion angle deviation exceeding 30°), the system outputs corrective instructions, such as the voice prompt "Pay attention to the position of the tongue tip, repeat after me again," while demonstrating the standard mouth shape or motion trajectory through mechanical components, and restarting data acquisition for a second comparison.
[0188] By encouraging and reinforcing correct behaviors and clarifying the direction for improvement through correction, a virtuous learning cycle of "trial-feedback-optimization" is formed; combined with the demonstration function, abstract standards are transformed into intuitive actions, reducing the difficulty of understanding for users, which is especially suitable for the cognitive characteristics of young children.
[0189] During the process of guiding the behavior of the current user, the duration of the user's interaction interruption is monitored in real time;
[0190] If the duration of the interaction interruption exceeds a preset threshold, the current data sensing mode is turned off and switched to a low-power sensing mode.
[0191] The low-power sensing mode is defined as enabling only one of voice sensing or text sensing.
[0192] As a specific example, this step specifically includes:
[0193] Real-time monitoring of interruption duration: The system judges the user status by sensing the strength of interaction signals in the mode (such as voice input interval and gesture pause time) and accumulates the duration of no effective interaction;
[0194] Trigger low power mode: When the interruption duration exceeds a preset threshold (e.g., 5 minutes), automatically turn off high power consumption modes such as gesture perception and posture perception, and only retain voice perception (listening for wake words) or text perception (detecting touch screen input);
[0195] Low-power wake-up mechanism: When the user calls the toy name by voice or clicks the interface on the touch screen, the system quickly switches back from low-power mode to the original data perception mode and resumes the boot process.
[0196] The above improvements can reduce the continuous operation of invalid sensing modes, reduce device power consumption, and are especially suitable for battery-powered portable toys; by maintaining the basic interaction channel through low power mode, the cumbersome process of frequent power-on and power-off can be avoided, and the user's return needs can be responded to quickly, improving the continuity of use.
[0197] Based on the method implementation examples, see [link to relevant documentation]. Figure 3 , Figure 3This diagram illustrates the hardware unit composition of a robotic arm precision positioning system according to an embodiment of the present invention. The system includes:
[0198] The data storage module is used to store human-computer interaction data for a preset historical time period;
[0199] The requirements analysis module is used to determine the current user's current guidance requirements based on the human-computer interaction data in the data storage module.
[0200] The mode determination module is used to determine the current data perception mode of the smart toy based on the current guidance requirements;
[0201] The guidance execution module is used to activate the current data awareness mode and guide the behavior of the current user.
[0202] The intelligent toy has multimodal data perception capabilities, and the current data perception mode includes at least one of voice perception, gesture perception, posture perception, and text perception.
[0203] The system also includes:
[0204] The real-time monitoring module is used to collect user behavior data and interaction status in real time during the boot process;
[0205] The dynamic adjustment module is used to adjust the output frequency and content of the guidance instructions based on the data collected by the real-time monitoring module.
[0206] Preferably, the mode determination module is used to determine the current data perception mode of the smart toy based on the current guidance requirements, specifically including:
[0207] The guidance needs with the lowest behavioral preferences, the lowest training completion rate, or the lowest interaction frequency are taken as the current user's primary guidance needs.
[0208] The guidance needs with the highest behavioral preferences, the highest training completion rate, or the highest interaction frequency are taken as the second current guidance needs of the current user.
[0209] The second current boot request or the first current boot request is used as the target boot request. The mode determination module further includes a container management unit, which creates a main container and corresponding first and second sub-containers.
[0210] The main container is used to implement the main process, that is, to determine the current data perception mode of the smart toy based on the current boot requirements;
[0211] At the same time, "taking the guidance needs with the lowest behavioral preferences, the lowest training completion rate, or the lowest interaction frequency as the current user's first current guidance needs" is used as the first process, which is implemented based on the first sub-container;
[0212] The process of "taking the guidance needs with the highest behavioral preferences, the highest training completion rate, or the highest interaction frequency as the current user's second current guidance needs" is implemented as the second process, based on the second sub-container; the first and second sub-containers both belong to the main container and communicate through the channels provided by the main container.
[0213] Based on the current boot requirements, the main process determines the current data perception mode of the smart toy, specifically including:
[0214] The first process takes the first current guidance requirement as the target guidance requirement to determine the current data perception mode of the smart toy;
[0215] After the first sub-container completes the first current boot requirement determination, it synchronizes the requirement result to the main container, and the main process executes the perception mode determination based on the result;
[0216] The main process activates the current data awareness mode to guide the behavior of the current user;
[0217] During the process of guiding the behavior of the current user, the main process predicts the probability value of the duration of the user's interaction in real time;
[0218] The main process determines whether the probability value of the interaction duration is less than a threshold. If it is, it switches to the second process.
[0219] The second process takes the second current guidance requirement as the target guidance requirement, determines the current data perception mode of the smart toy, and then returns to the main process.
[0220] The second sub-container only outputs the second current boot request, while the main process is responsible for mode determination and switching execution. This maintains the specialized computing advantages of the sub-containers while clearly defining the overall decision-making role of the main process.
[0221] The aforementioned container-based boot request handling solution has the following core advantages: First, it offers efficient and low-latency process switching. By leveraging sub-container resource preloading and shared memory communication with the main container, the boot request switching response time is compressed to within 150ms, avoiding user distraction and catering to the short attention spans of young children. Second, it achieves resource isolation and efficient utilization. The main container allocates resources granularly, ensuring that high-energy-consuming processes with weaknesses do not compete with lightweight processes with strengths, balancing training accuracy and energy consumption control. Third, it significantly enhances system stability. The independent operation of sub-containers ensures that a single process failure does not affect the overall system, and the main container's fault-tolerance mechanism can quickly repair anomalies, solving the problem of complete system interruption in traditional single-process architectures. Fourth, it offers strong scalability. The main container supports flexible addition and removal of sub-containers, and each sub-container can independently upgrade its resources, expanding to more boot categories without system reconstruction to adapt to diverse needs.
[0222] In summary, this invention comprehensively breaks through the technical limitations of traditional smart toys through data-driven precise decision-making, flexible switching guidance strategies, containerized efficient architecture, and multimodal interaction adaptation, making smart toys truly personalized growth partners that combine educational value and companionship attributes.
[0223] Figure 4 The diagram illustrates a scenario of the precise positioning system for robotic arms according to an embodiment of the present invention. It depicts a warm children's room setting where the intelligent toy robot becomes the central focus of interaction. Based on the technical solution of the present invention, it can precisely adapt to children's needs: for language expression training, it activates voice + text perception to collect children's pronunciation and dialogue; for physical coordination training, it activates gesture + posture perception to capture movements. When a child reaches out to interact, the robot collects behavioral data through multimodal perception, compares it against standards, and guides the child with voice, light (such as a blue halo around the head), or mechanical movements. Simultaneously, the system monitors the interaction status; if the interruption exceeds a threshold, it switches to a low-power mode (such as retaining only voice perception), ready to respond to the child at any time, achieving personalized, intelligent, and low-power companionship and guidance, assisting in children's ability training.
[0224] In this section, the present invention provides several method or product embodiments. Each method or product embodiment constitutes an independent technical solution and makes at least one contribution to the prior art, capable of solving one or more technical problems mentioned in the background art, and having one or more of the aforementioned outstanding effects or advantages.
[0225] However, it is understood that not every embodiment is required to solve all the technical problems mentioned in the background or achieve all the technical effects (such as the five major effects / advantages mentioned later). For each individual embodiment, as long as it makes at least one contribution relative to the prior art and achieves an improved technical effect (e.g., at least one of the five major advantages mentioned later), has outstanding substantive features or significant progress, the technical solution constituted by that individual embodiment should possess novelty and inventiveness in the sense of patent law.
[0226] Compared with the prior art, the technical solution of the present invention can be summarized by at least the following five advantages:
[0227] I. Precise and personalized guidance to achieve a "one-size-fits-all" approach to growth and development.
[0228] By constructing user demand profiles through multi-dimensional data analysis, the technical solution of this invention achieves a breakthrough from "general templates" to "individual adaptation." Based on human-computer interaction data within a preset historical time period, the system quantifies user characteristics from three core dimensions: behavioral preferences, training completion rate, and interaction frequency, accurately identifying weaknesses (the first current guidance need) and strengths (the second current guidance need). For example, if a 5-year-old child, Xiaoming, exhibits "three lows" in the mathematical logic module—low preference, low completion rate, and low frequency—the system will identify this module as the key guidance direction. Meanwhile, strengths such as language expression are treated as flexible resources for maintaining interest. This demand determination mechanism based on real interaction data avoids the one-sidedness of traditional smart toys relying on age tags or initial testing, ensuring that the guidance content is highly matched to the user's ability shortcomings and growth pace. For new users, the solution quickly generates an adapted plan based on the initial age and training target input. For example, if the parents of a 4-year-old child, Lili, select "language expression priority," the system immediately loads the corresponding age-appropriate scenario dialogue resources, achieving personalized service with "zero data startup."
[0229] II. A dynamic and flexible guidance mechanism to balance challenge and sense of accomplishment.
[0230] This invention innovatively designs a dual-track guidance mode of "overcoming weaknesses + maintaining strengths," achieving a dynamic balance in the learning pace through real-time monitoring and intelligent switching. During the guidance process, the system assesses the user's state in real time through the interaction persistence probability value. When it detects resistance to weaknesses such as mathematical logic (e.g., interaction persistence probability < 60%), it immediately switches from the first sub-container to the second sub-container, initiating guidance for strengths such as language expression. For example, if Xiaoming makes multiple mistakes in addition and subtraction training, causing the interaction to be interrupted, the system switches to the nursery rhyme creation module within 150ms, rebuilding confidence through encouraging instructions such as "You read so well." This flexible mechanism solves the drawbacks of traditional toys' "one-way reinforcement"—it avoids both users staying in their comfort zone for too long, leading to stagnation, and prevents loss of interest due to continuous frustration. At the same time, the system outputs tiered guidance instructions by comparing results, providing diverse encouragement for correct behaviors and concrete corrections for deviant behaviors (e.g., mechanically demonstrating standard actions), forming a closed-loop learning cycle of "trial-feedback-optimization," ensuring that challenge and a sense of accomplishment are always dynamically balanced.
[0231] III. Containerized architecture supports efficient switching, ensuring boot continuity and response speed.
[0232] By introducing a layered architecture of main and sub-containers, this invention solves the switching latency problem of traditional single-process architectures. The main container, acting as the global management hub, constructs an efficient communication channel through shared memory and a lightweight message queue, enabling resource preloading and rapid activation of the first sub-container (the weakest process) and the second sub-container (the strongest process). The first sub-container preloads core resources such as the gesture recognition model and is in a ready state, while the second sub-container remains in a lightweight standby mode, reducing resource consumption by 30%. When a switching command is triggered, the main container completes resource reallocation and process activation within 150ms through the process scheduling interface, far exceeding the traditional process switching response speed of over 500ms. This design is particularly important for young users, as children's short attention spans necessitate keeping guidance interruptions within a perceptual threshold. The containerized architecture ensures a seamless switching process, avoiding distraction. Simultaneously, the independent operation mechanism of the sub-containers ensures that the high-energy-consuming computation of the mathematical logic module and the lightweight interaction of the language module do not interfere with each other, improving resource utilization efficiency by 40%.
[0233] IV. Multimodal perception and deep adaptation enhance the naturalness and inclusivity of interaction.
[0234] This invention dynamically matches the optimal perception combination based on guidance needs, achieving "on-demand adaptation" of interaction methods. For language expression training, it employs a dual-perception mode of voice and text, capturing pronunciation details through high-sensitivity voice recognition and displaying pinyin prompts through text perception. For body coordination training, it activates gesture and posture perception, accurately capturing movement trajectories and angle deviations. Comprehensive ability training can flexibly combine multiple modes, such as simultaneously using voice commands, gesture interaction, and posture recognition in scenario-based story performances. This multimodal strategy adapts to the interaction habits of different users—young children rely on voice and gestures, users with text skills can input via touchscreen, and users with limited physical mobility can switch to a voice-driven mode. At the guidance output level, the system integrates diverse feedback such as voice prompts, flashing lights, and mechanical demonstrations. For example, correct behavior triggers a green light and applause, while incorrect behavior demonstrates the standard posture through mechanical movements, transforming abstract guidance commands into intuitive perception signals and improving the interaction error tolerance rate by more than 50%.
[0235] V. High stability and scalability ensure long-term usability.
[0236] Containerized architecture endows the system with strong stability and scalability. The main container, through real-time status monitoring and fault tolerance mechanisms, quickly restarts abnormal processes and restores data, solving the problem of "one failure causing complete system interruption" in traditional single-process architectures. For example, when the posture perception process in the art creation module crashes, the main container restarts the sub-container and restores the current drawing progress within 50ms, without the user noticing any interruption. In terms of resource management, the system achieves intelligent energy saving through interactive interruption monitoring; after 5 minutes of inactivity, high-power posture perception is shut down, retaining only the voice wake-up function, extending battery life by 30%. Looking at long-term development, the main container supports flexible addition and removal of sub-containers, easily adding new guidance categories such as programming enlightenment and social etiquette. Each sub-container can independently upgrade its model and resource library without refactoring the entire system. This design of "stable core architecture + scalable functional modules" enables smart toys to continuously iterate their service capabilities as users grow, adapting fully from early childhood education to school age transitions, significantly enhancing the product's lifecycle value.
[0237] Although not shown in the accompanying drawings, a preferred and further embodiment of the product may also be an electronic device comprising: a memory and one or more processors. The memory stores one or more application programs adapted to be executed by the one or more processors using the aforementioned human-computer interaction-based intelligent toy guidance method.
[0238] Although not shown in the accompanying drawings, further embodiments also include a computer-readable storage medium storing a computer program that, when executed, implements the steps of the aforementioned human-computer interaction-based intelligent toy guidance method.
[0239] It is understood that the system, product, equipment, and media implementation examples and method implementations correspond to each other and can be referenced by each other, and their principles are similar or the same, so they will not be elaborated again.
[0240] Other technologies, principles, algorithms, or models not elaborated in detail in this application can be found in the prior art.
[0241] The foregoing has shown and described the method embodiments and systems of the present invention, but it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A toy intelligent guidance method based on human-computer interaction, the method being applied to intelligent toys, characterized in that, The method includes: Based on the human-computer interaction data of the smart toy during a preset historical time period, the target guidance needs of the current user are determined. Based on the target guidance requirements, determine the current data perception mode of the smart toy; The current data awareness mode is activated to guide the behavior of the current user. The intelligent toy has multimodal data sensing capabilities; The current data perception mode includes at least one of voice perception, gesture perception, posture perception, and text perception.
2. The method as described in claim 1, characterized in that, The process of determining the current user's target guidance needs based on the human-computer interaction data of the smart toy over a preset historical time period includes: Analyze user behavior preferences, training completion rate, and interaction frequency in the human-computer interaction data within the preset historical time period, and output the target guidance requirement based on the preset guidance requirement classification model; The guidance demand classification model is trained based on user age, cognitive level, and common training objectives.
3. The method according to claim 1, characterized in that, Determining the current data perception mode of the smart toy based on the target guidance requirements includes: If the target guidance requirement is language expression training, then the current data perception mode is determined to be speech perception and text perception; If the target guidance requirement is limb coordination training, then the current data perception mode is determined to be gesture perception and posture perception; If the target guidance requirement is comprehensive ability training, then the current data perception mode is determined to be at least two of the following: voice perception, gesture perception, posture perception, and text perception.
4. The method according to claim 1, characterized in that, The step of activating the current data awareness mode and guiding the behavior of the current user includes: The user's behavioral data is collected in real time using the current data perception mode. The collected behavioral data is compared with preset standard behavioral data to generate comparison results; Based on the comparison results, guidance instructions are output to the user through voice prompts, flashing lights, or mechanical action demonstrations.
5. The method according to claim 4, characterized in that, Based on the comparison results, the step of outputting guidance instructions to the user through voice prompts, flashing lights, or mechanical demonstrations includes: If the comparison result shows that the user's behavior conforms to the standard behavior data, then an encouraging guidance instruction will be output. If the comparison result shows that the user's behavior deviates from the standard behavior data, a corrective guidance instruction will be output, and the standard behavior will be re-demonstrated.
6. The method according to claim 1, characterized in that, The method further includes: During the process of guiding the behavior of the current user, the duration of the user's interaction interruption is monitored in real time; If the duration of the interaction interruption exceeds a preset threshold, the current data sensing mode is turned off and switched to a low-power sensing mode. The low-power sensing mode is defined as enabling only one of voice sensing or text sensing.
7. The method according to claim 1, characterized in that, The step of determining the current user's target guidance needs based on the human-computer interaction data of the smart toy over a preset historical time period also includes: If the human-computer interaction data within the preset historical time period is empty, then the initial guidance requirement is determined as the current guidance requirement based on the user's initial age information and training target.
8. The method according to claim 1, characterized in that, The process of determining the current user's target guidance needs based on the human-computer interaction data of the smart toy over a preset historical time period includes: Statistics are collected on the current user's behavioral preferences, training completion rate, and interaction frequency for each preset category of guidance needs within the preset historical time period. The guidance needs with the lowest behavioral preferences, the lowest training completion rate, or the lowest interaction frequency are taken as the current user's primary guidance needs. The guidance needs with the highest behavioral preferences, the highest training completion rate, or the highest interaction frequency are taken as the second current guidance needs of the current user. Using the second current boot request or the first current boot request as the target boot request specifically includes: First, the first current guidance requirement is set as the target guidance requirement. Then, when a preset condition is met, the system switches to the second current guidance requirement as the target guidance requirement.
9. A toy intelligent guidance system based on human-computer interaction, characterized in that, The system includes: The data storage module is used to store human-computer interaction data for a preset historical time period; The requirements analysis module is used to determine the current user's current guidance requirements based on the human-computer interaction data in the data storage module. The mode determination module is used to determine the current data perception mode of the smart toy based on the current guidance requirements. The guidance execution module is used to activate the current data awareness mode and guide the behavior of the current user. The intelligent toy has multimodal data perception capabilities, and the current data perception mode includes at least one of voice perception, gesture perception, posture perception, and text perception.
10. The system according to claim 9, characterized in that, The system also includes: The real-time monitoring module is used to collect user behavior data and interaction status in real time during the boot process; The dynamic adjustment module is used to adjust the output frequency and content of the guidance instructions based on the data collected by the real-time monitoring module.
Citation Information
Patent Citations
Interaction method, toy, cloud and interaction system
CN119488712A
AI intelligent training platform and method supporting meta-universe multi-mode interaction
CN120122816A
Design and development guidance method for teaching toys for children with mental disorder based on big data
CN120278046A
Computer-implemented learning method and apparatus
US20100003659A1
Adaptive, individualized, and contextualized text-to-speech systems and methods
US20240194178A1