Internet of Things toy management method and system
By using intelligent model library dynamic matching technology, the problem of managing the AI capabilities of intelligent toys has been solved, enabling personalized interactive experience and efficient operation and maintenance, thereby improving the intelligence level of toys and user satisfaction.
Patent Information
- Application Number
- CN202510826808.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-11-11
AI Technical Summary
The AI capabilities of existing smart toys are difficult to update, optimize, and manage, making it hard to adapt to rapid technological development and changing user needs, resulting in poor interactive experience and high maintenance costs.
This invention provides an IoT toy management method and system that dynamically matches automatic speech recognition, large language models, and text-to-speech models through an intelligent model library to generate personalized response voice data, and optimizes interaction strategies by combining user profiles and device information.
This has enhanced the personalized interactive experience of smart toys, reduced operation and maintenance costs, and improved the utilization rate of AI resources and interaction efficiency.
Smart Images

Figure CN120932649A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of the Internet of Things (IoT), and in particular to an IoT toy management method and system. Background Technology
[0002] While existing smart toys on the market possess some voice interaction capabilities, they still suffer from several problems. The AI capabilities built into these toys, or the specific AI services they rely on, are often fixed or deeply tied to specific service providers. Updating, optimizing, switching, or combining these AI capabilities is extremely difficult, making it hard to adapt to rapid technological advancements and changing user needs. Different toy products or different user scenarios may have varying requirements for AI capabilities. Existing solutions struggle to dynamically configure and optimize AI services in a targeted manner. Smart toys often lack a unified AI capability management and scheduling platform, leading to redundant construction or low utilization of AI resources. Furthermore, monitoring, maintaining, and upgrading AI capabilities scattered across various toy products incurs high operational costs. Due to the limitations of AI capability management, it is difficult to quickly integrate the latest AI technological advancements or flexibly adjust interaction strategies based on user feedback, resulting in a poor user experience. Summary of the Invention
[0003] This application provides an IoT toy management method and system to address the problem of poor intelligent interactive experience in smart toys.
[0004] To address the aforementioned technical problems, this application provides a technical solution: an IoT toy management method, comprising: acquiring user voice data sent by the IoT toy; dynamically matching the automatic speech recognition model used in the current dialogue within a preset intelligent model library, and performing speech recognition on the user voice data using the automatic speech recognition model to generate user voice text; the intelligent model library includes preset automatic speech recognition models, various types of large language models, and various types of text-to-speech models; dynamically matching the large language model used in the current dialogue within the preset intelligent model library, and generating a response text based on the user voice text using the large language model; dynamically matching the text-to-speech model used in the current dialogue within the preset intelligent model library, and setting the language style of the response text using the text-to-speech model to generate response voice data; and sending the response voice data to the IoT toy for playback.
[0005] In some embodiments, the dynamic matching of the automatic speech recognition model used in the current round of dialogue, and the generation of user speech text by performing speech recognition on the user speech data through the automatic speech recognition model, includes: selecting the automatic speech recognition model used in the current round of dialogue from the intelligent model library based on the device information of the IoT toy, user profile data, or the interaction context content in the user speech data; calling the automatic speech recognition model, and converting the speech data into corresponding user speech text through the automatic speech recognition model, combined with the device information of the IoT toy, the user profile data, or the interaction context content in the user speech data.
[0006] In some embodiments, the dynamic matching of the large language model used in the current round of dialogue and the generation of response text based on the user's voice text through the large language model includes: selecting the large language model used in the current round of dialogue from the intelligent model library based on the user's voice text, the device information of the IoT toy, user profile data, the interaction context content in the user's voice data, or the functional type of the IoT toy; invoking the large language model and performing semantic understanding on the user's voice text through the large language model; and generating response text based on the semantic understanding result, the user's voice text, the device information of the IoT toy, user profile data, the interaction context content in the user's voice data, or the functional type of the IoT toy through the large language model.
[0007] In some embodiments, the dynamic matching of the text-to-speech model used in the current round of dialogue, and the setting of the language style of the response text through the text-to-speech model to generate response speech data, includes: selecting the text-to-speech model used in the current round of dialogue from the intelligent model library based on the response text, the language style preferences of the user profile data, and the style information of the IoT toy; setting the timbre, speech rate, intonation, or emotional features of the response text through the text-to-speech model according to the response text, the language style preferences of the user profile data, and the style information of the IoT toy; and generating the response speech data based on the set timbre, speech rate, intonation, and emotional features.
[0008] In some embodiments, the IoT toy management method further includes: obtaining management information input by an administrator or authorized user; and dynamically updating the selection strategy of the automatic speech recognition model, the large language model, and the text-to-speech model based on the management information.
[0009] In some embodiments, the IoT toy management method further includes: displaying a graphical smart model management page; and in response to user modifications to the graphical smart model management page, constructing and testing the modified smart model interaction logic.
[0010] In some embodiments, the IoT toy management method further includes: the smart model library is pre-set with multiple types of visual smart models, which are used to perform image recognition and facial expression recognition based on the image data sent by the IoT toy.
[0011] To address the aforementioned technical problems, another technical solution adopted in this application is: providing an IoT toy management system for implementing the IoT toy management method as described in claims 1-7, comprising: a central management platform, including an automatic speech recognition model management module, a large language model management module, and a text-to-speech model management module for dynamic selection; an IoT toy, including a sound acquisition module, a sound playback module, and a wireless communication module; the IoT toy acquires user voice through the sound acquisition module and converts it into corresponding user voice data, and sends the user voice data to the central management platform through the wireless communication module; the central management platform dynamically matches the automatic speech recognition model used in the current round of dialogue through the automatic speech recognition model management module, dynamically matches the large language model used in the current round of dialogue through the large language model management module, and dynamically matches the text-to-speech model used in the current round of dialogue through the text-to-speech model management module; the central management platform sends the response voice data generated by the text-to-speech model to the wireless communication module of the IoT toy, and the IoT toy plays the response voice data through the sound playback module.
[0012] In some embodiments, the IoT toy is also equipped with a variety of lightweight smart models, which are used to assist the central management platform in conducting intelligent dialogue.
[0013] In some embodiments, the IoT toy is further equipped with a wake word recognition model, which is used to activate the IoT toy in response to a user's wake word.
[0014] The beneficial effects of this application are as follows: Unlike existing technologies, this application discloses an IoT toy management method and system. By acquiring user voice data sent by IoT toys, the system dynamically matches the automatic speech recognition model used in the current dialogue within a pre-set intelligent model library. The automatic speech recognition model then performs speech recognition on the user voice data to generate user voice text. The intelligent model library contains various types of automatic speech recognition models, various types of large language models, and various types of text-to-speech models. The system dynamically matches the large language model used in the current dialogue within the pre-set intelligent model library and generates a response text based on the user's voice text. It also dynamically matches the text-to-speech model used in the current dialogue within the pre-set intelligent model library and sets the language style of the response text to generate response voice data. Based on the positioning of different toys in the product management module and the analysis of user profiles in the device management module, the interaction strategy and scheduling module can dynamically select and configure the most suitable functional model combination for specific toys or users, thereby improving the user experience. The response voice data is then sent to the IoT toy for playback. Unified management of various functional models allows for centralized resource scheduling and improved utilization. Meanwhile, the monitoring, maintenance, and upgrades of the functional model are all completed on one side, greatly simplifying the operation and maintenance of numerous dispersed IoT toy terminals. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein: Figure 1 This is a flowchart illustrating an embodiment of the Internet of Things toy management method provided in this application; Figure 2 Is it like this? Figure 1 The flowchart of step 20 of the method shown is a schematic diagram of one embodiment; Figure 3 Is it like this? Figure 1 The flowchart of step 30 of the method shown is a schematic diagram of an embodiment; Figure 4 Is it like this? Figure 1 The flowchart of step 40 of the method shown is a schematic diagram of an embodiment. Figure 5 This is a flowchart illustrating another embodiment of the IoT toy management method provided in this application; Figure 6 This is a flowchart illustrating yet another embodiment of the IoT toy management method provided in this application; Figure 7 This is a schematic diagram of an embodiment of the Internet of Things toy management system provided in this application. Detailed Implementation
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0017] The terms "first," "second," and "third" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0018] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0019] See Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the IoT toy management method provided in this application. The IoT toy management method includes the following steps: 10: Acquire user voice data sent by IoT toys.
[0020] The IoT toy collects user-input voice content, converts the voice content into voice data or pre-processed voice features, and sends it to the central management platform via the network. The central management platform receives the user voice data sent by the IoT toy as input data for the AI recognition process.
[0021] IoT toys typically possess some voice interaction capabilities, enabling question-and-answer interactions with users through the integration of cloud-based AI services. IoT toys typically include built-in microphone arrays, speakers, and network communication modules.
[0022] 20: Dynamically match the automatic speech recognition model used in this round of dialogue from the preset intelligent model library, and perform speech recognition on the user's speech data through the automatic speech recognition model to generate user speech text. The intelligent model library has preset various types of automatic speech recognition models, various types of large language models, and various types of text-to-speech models.
[0023] After a user inputs voice data through an IoT toy terminal, the terminal uploads the voice data to the central management platform. The central management platform dynamically matches and recognizes Automatic Speech Recognition (ASR) models, and supports version management, performance monitoring, and on-demand loading of ASR models.
[0024] The intelligent model library contains various types of automatic speech recognition models, large language models (LLM), and text-to-speech (TTS) models, which are managed and dynamically scheduled by the central management platform.
[0025] Specifically, the intelligent model library stores various automatic speech recognition models, supporting system version management, performance monitoring, and on-demand loading. During interaction, the platform dynamically selects the appropriate automatic speech recognition model based on the IoT toy terminal's device information (such as model and firmware version), user profile (such as age and language preference), or interaction context (such as ambient noise and dialogue scenario).
[0026] The intelligent model library stores various large language models (covering different knowledge domains, scales, and response styles), supporting system version management, knowledge base updates, and prompt engineering optimization. During interaction, the platform combines the text output by the automatic speech recognition model, user history dialogues, and toy product positioning (such as education or entertainment) to intelligently route to the most suitable large language model, and configures context information and response strategies.
[0027] The intelligent model library stores various TTS engines / models, supporting the system in configuring voice selection, speech rate, tone adjustment, and emotional synthesis. During interaction, the platform selects the appropriate TTS model and parameters to generate speech based on the LLM-generated response text, user preferences (such as voice preferences), or toy character settings (such as cute or serious characters).
[0028] Further, see Figure 2 , Figure 2 Is it like this? Figure 1The flowchart of step 20 of the method shown is a schematic diagram of one embodiment. Step 20 further includes the following steps: 21: Based on the interactive context content in the device information, user profile data, or user voice data of IoT toys, select the automatic speech recognition model used in this round of dialogue from the intelligent model library.
[0029] The device information of IoT toys, including model number and firmware version, is registered and certified in advance in the system.
[0030] User profile data is user tags built by the device management module through long-term interaction data such as voice input and operation feedback, including age preferences, language preferences, and usage habits.
[0031] The interaction context captures dynamic factors such as environmental noise and dialogue scenarios in real time to ensure that the model selection is accurately adapted to the current interaction needs.
[0032] Different toy models may be equipped with chips with varying computing power, such as low-power chips and high-performance chips, which affects the selection of the automatic speech recognition model. For example, older devices should use lightweight models to avoid overloading computing power. Newer firmware versions may support more complex automatic speech recognition functions, such as multi-language switching, requiring a compatible automatic speech recognition model.
[0033] When there is network latency or bandwidth limitation, such as in a weak outdoor network environment, the locally deployed lightweight automatic speech recognition model should be selected first; when the network is stable, the high-precision automatic speech recognition model in the cloud should be called.
[0034] For example, children aged 3-6 have a child's voice and unclear pronunciation, so they need to be adapted to the "child's voice optimized automatic speech recognition model"; teenagers aged 7-12 may use more standard Mandarin, so they can switch to the general automatic speech recognition model.
[0035] For dialects such as Cantonese and Sichuanese, an automatic speech recognition model that supports the corresponding dialect should be used; for bilingual family users such as those using a mix of Chinese and English, a multilingual recognition model should be selected.
[0036] For users who frequently ask questions, such as "Why is the sky blue?" or "Do stars fall down?", the automatic speech recognition model needs to support "streaming recognition," which means converting speech to text as it is spoken, to improve the smoothness of the interaction.
[0037] A basic automatic speech recognition model can be used in quiet environments such as bedrooms; however, a "noise-resistant and robust automatic speech recognition model" should be used in noisy environments such as living rooms and outdoors.
[0038] Learning scenarios require accurate recognition of academic terminology; entertainment scenarios require support for colloquial expressions. Urgent questions require low-latency automatic speech recognition models, with priority given to models that respond quickly; regular questions can tolerate slightly higher latency to improve accuracy.
[0039] For example, in a scenario where a young child asks a question in a noisy living room, the device information is that the toy model is "Early Education Version A" (medium computing power), and the network is stable; the user profile is a 3-year-old child speaking Mandarin (with occasional unclear pronunciation); the interaction context includes a living room (background noise is the TV, with a noise level of 60dB); the automatic speech recognition model selection is to call the "Children's Voice Optimization + Noise-Resistant Automatic Speech Recognition Model", which uses a microphone array to suppress TV noise and specifically optimizes the recognition of the child's unclear pronunciation.
[0040] 22: Call the automatic speech recognition model, and use the automatic speech recognition model, combined with the device information of IoT toys, user profile data, or interactive context content in user voice data, to convert user voice data into corresponding user voice text.
[0041] The selected automatic speech recognition model is invoked. This model, combined with information such as the device information of the IoT toy, user profile data, and interaction context, is used to recognize the user's speech data and output the corresponding text-formatted user speech text.
[0042] By fine-tuning the automatic speech recognition model, efficient and accurate speech-to-text conversion is ensured in different scenarios, enhancing the user experience. Fine-tuning not only improves recognition accuracy but also optimizes the model based on user habits, enabling personalized services, further reducing false recognition rates, and ensuring a smooth and natural interactive experience for users in various environments.
[0043] 30: Dynamically match the large language model used in this round of dialogue from the preset intelligent model library, and generate user response text based on user voice text through the large language model.
[0044] A dynamic matching strategy for large language models is developed, accurately selecting the appropriate model based on user profiles, interaction context, and device information to ensure the relevance and accuracy of the generated response text. For example, in academic scenarios, a large language model rich in technical terminology is prioritized; in entertainment scenarios, a model adept at colloquial expressions is chosen to achieve a natural and fluent dialogue experience. Through intelligent adaptation of large language models, user interaction satisfaction and efficiency are further improved.
[0045] Further, see Figure 3 , Figure 3 Is it like this? Figure 1 The flowchart of step 30 of the method shown is a schematic diagram of one embodiment. Step 30 further includes the following steps: 31: Based on user voice text, IoT toy device information, user profile data, interactive context content in user voice data, or IoT toy function type, select the large language model to be used in this round of dialogue from the intelligent model library.
[0046] The text output by the automatic speech recognition model determines the type of content that the large language model needs to process. For example, knowledge-based texts such as scientific questions and mathematical calculations require large language models with strong knowledge specialization and broad knowledge base coverage, such as the "Science Encyclopedia Large Language Model" and the "Mathematical Reasoning Large Language Model". Entertainment-related texts require large language models with strong generation capabilities and vivid language styles, such as the "Story Creation Large Language Model" and the "Dialogue Personification Large Language Model". Command-type texts such as "Turn on music" and "Set alarm clock" require large language models with accurate intent recognition and fast response speed, such as the "Command Parsing Large Language Model".
[0047] Entry-level toys and other low-computing-power devices should choose lightweight large language models with fewer parameters and faster inference speed to avoid interaction delays due to insufficient computing power. Flagship toys and other high-performance devices should support calling large language models with a large number of parameters and complex functions, such as multimodal understanding large language models, to improve the depth of responses. When used in weak network environments such as outdoors, locally deployed large language models should be preferred to reduce cloud transmission latency. When the network is stable, high-precision large language models in the cloud should be called.
[0048] For children aged 3-6, choose a language model that simplifies and anthropomorphizes their responses, such as the "Kid-Friendly Language Model," which uses "the sun" instead of "stars." For teenagers aged 7-12, choose a language model that offers strong knowledge expansion and moderate linguistic rigor, such as the "Science Popularization and Enlightenment Language Model," which incorporates animated metaphors when explaining "Earth's rotation." Regarding language preferences (dialect / bilingualism): choose a language model that supports multilingual comprehension, such as the "Cantonese-Mandarin Bilingual Language Model," which accurately identifies dialect questions and outputs Mandarin answers.
[0049] Historical interaction records, such as previous user questions and system answers, are used to maintain the dialogue context and avoid "fragmentation" of the large language model.
[0050] In multi-turn dialogue scenarios, if a user asks consecutive questions such as "What kinds of dinosaurs are there?" → "What is the largest dinosaur?", the large language model needs to combine historical records to identify "dinosaurs" as the current topic and prioritize the "paleontological knowledge large language model". If the user has previously given feedback that "the answer is too complicated", a "simplified answer strategy" (such as reducing technical terms) needs to be configured in subsequent interactions.
[0051] The toy types defined in the product management module, such as educational, entertainment, and companion toys, determine the core capability requirements of the large language model. Educational toys need to select a large language model with high knowledge accuracy and support for knowledge point decomposition; entertainment toys need to select a large language model with strong creativity and vivid language style; and companion toys need to select a large language model with strong emotional understanding and anthropomorphic responses.
[0052] 32: Call the large language model and use the large language model to perform semantic understanding of the user's speech text.
[0053] The large language model selected by the response strategy is invoked to interpret the user's voice text.
[0054] 33: Generate response text based on semantic understanding results, user voice text, IoT toy device information, user profile data, interactive context content in user voice data, or the functional type of IoT toys through large language models.
[0055] For example, the user's response text in the input information is "Tell me a princess story," corresponding to the entertainment category; the device information is a portable toy, corresponding to a weak network environment and a local model; the user profile is 5 years old, likes pink and princesses; the history shows the user heard the story of "The Princess and the Unicorn" last week; the product positioning is entertainment and interactive. The large language model routing selects the locally deployed "Story Creation Large Language Model," which is lightweight and can generate stories quickly; the context configuration injects historical preferences "the user likes princesses and unicorns" and the product positioning is "entertainment and interactive." The answer strategy is to prompt "The story should have a pink castle, a unicorn, and end with the question 'Does the princess go to save the kitten or find the star?'" Then, the output example of the large language model is: "Once upon a time, there was a little princess named Lily in a pink castle. Her friend, the unicorn, said: 'There is a lost kitten in the forest, and there are also shining stars... Princess, do you want to save the kitten or find the star first?'" Through precise analysis of user history interactions and product positioning, the large language model not only continues the storyline but also aligns with user preferences, enhancing the interactive experience and achieving efficient generation of personalized content.
[0056] 40: Dynamically match the text-to-speech model used in this round of dialogue from the preset intelligent model library, and set the language style of the response text through the text-to-speech model to generate response speech data.
[0057] The system combines user profile data and device information to adjust the speech rate, tone, and emotion of the voice, ensuring that the voice style matches the user's age and the context, thereby improving the naturalness and friendliness of the voice interaction.
[0058] Further, see Figure 4 , Figure 4 Is it like this? Figure 1 The flowchart of step 40 of the method shown is a schematic diagram of one embodiment. Step 40 further includes the following steps: 41: Based on the language style preferences of the response text, user profile data, and style information of IoT toys, select the text-to-speech model to be used in this round of dialogue from the intelligent model library.
[0059] The response text output by the large language model is the core input of the text-to-speech model. Its content types include knowledge explanation, story generation, emotional response, and emotional tags including happy, sad, and surprised, which directly determine the parameter direction of the text-to-speech model.
[0060] Knowledge-based texts, such as "The Earth rotates because of the conservation of initial angular momentum," require clear pronunciation and a steady tone to facilitate user comprehension; story-generating texts, such as "The princess pushed open the castle gate and found a talking fox," require vivid pronunciation and varied intonation to enhance the visual experience; and emotionally responsive texts, such as "I know you are sad, let's think of a solution together," require gentle pronunciation and a warm tone to convey empathy.
[0061] User preferences include timbre preference, speech rate / tone preference, and emotional intensity preference.
[0062] Toy character design includes cartoon characters, such as "little dinosaurs," which need to be matched with cute and childlike voices, speaking at a relatively fast pace and with large intonation; science assistants, such as "doctors," need to be matched with calm and professional voices, speaking at a moderate pace and with a steady tone; and emotional companionship characters need to be matched with soft and friendly voices, speaking at a slower pace and with a gentle tone.
[0063] 42: Based on the language style preferences of the response text, user profile data, and style information of IoT toys, the timbre, speech rate, tone, or emotional characteristics of the response text are set using a text-to-speech model.
[0064] The system receives input information such as response text, language style preferences from user profile data, and style information of IoT toys. This information serves as the core basis for setting the parameters of the text-to-speech model, configuring a voice output scheme that is more suitable for the current usage environment of IoT toys. Through fine-tuning of parameters, the system ensures that every voice interaction accurately meets user needs, improving the user experience. The system also monitors the interaction effect in real time, continuously optimizing model parameters to achieve a more efficient end-to-end interaction process.
[0065] 43: Generate response voice data based on the set timbre, speech rate, tone and emotional characteristics.
[0066] For example, the use case is a companion toy for a 6-year-old child expressing "unhappiness". The input information includes: the large language model output text is "I know you must be very sad right now... It's not your fault that your classmates are bullying you. Let's tell the teacher together tomorrow, okay?", the type is emotional response, and the emotional tag is comfort; the user preference is "likes the older sister's voice"; the toy character is a "warm-hearted older sister", characterized by gentleness and kindness.
[0067] The configuration parameters for the text-to-speech model are as follows: the engine is selected as the emotional TTS engine, which supports subtle emotional adjustment; the timbre is "gentle female voice" to match user preferences and roles; the speech rate is 120 words / minute, which is 15% slower than the default and can convey a sense of comfort; the tone is generally calm with a slight drop at the end of the sentence to avoid a sense of oppression; the emotional synthesis is achieved by reducing the fundamental frequency by 5%, making the voice softer, and extending the duration of each word by 8%, making each word sound more patient.
[0068] The final output of the text-to-speech model is "I know you must be very sad right now... It's not your fault that you were bullied by your classmates. Let's tell the teacher together tomorrow, okay?" This tone is warm and gentle, meeting the need for emotional support.
[0069] By fine-tuning the text-to-speech model, the system ensures that the voice output closely matches the user's emotional needs, enhancing the companionship experience. The system records every interaction and feedback, continuously optimizing the model to achieve personalized voice customization, further strengthening the emotional connection between the user and the toy.
[0070] 50: Send a response voice data to the IoT toy for playback.
[0071] The IoT toy clearly plays the generated voice data through its built-in speaker, enabling voice interaction with the user.
[0072] Optionally, see Figure 5 , Figure 5 This is a flowchart illustrating another embodiment of the IoT toy management method provided in this application. The IoT toy management method further includes the following steps: 60: Obtain management information entered by the administrator or authorized user.
[0073] The system provides an information input interface for administrators or authorized users, supporting the collection of specific content of input subjects and access control information.
[0074] Specifically, administrators can set user permissions and define rules such as toy usage time and content filtering through the interface to ensure safety and compliance. The system records management operations in real time, facilitating traceability and adjustments, and ensuring the flexibility and security of IoT toy management.
[0075] Authorized users can view toy usage records, analyze user behavior, and provide personalized service suggestions based on their permissions. The system automatically synchronizes and updates management information to ensure real-time performance and accuracy, comprehensively improving management efficiency.
[0076] 70: Dynamically update the selection strategy for automatic speech recognition model, large language model and text-to-speech model based on management information.
[0077] After receiving management information, the system dynamically adjusts the selection strategy of the automatic speech recognition model, large language model, and text-to-speech model according to the user permissions and usage rules set by the administrator. This ensures accurate speech recognition, precise semantic understanding, and speech output that matches user preferences. The system updates model configurations in real time, optimizes the interactive experience, and guarantees the safety and personalized services of IoT toys.
[0078] Optionally, see Figure 6 , Figure 6 This is a flowchart illustrating another embodiment of the IoT toy management method provided in this application. The IoT toy management method further includes the following steps: 80: Displays the graphical intelligent model management page.
[0079] The interface allows administrators to intuitively view the running status of each model, adjust parameters in real time, and ensure efficient model collaboration.
[0080] The management console's front-end interface displays a graphical intelligent model management page to administrators / authorized users, with its core design goal being visual operation and real-time feedback.
[0081] Specifically, the integrated functional models are displayed in the form of a list or card, including model metadata, status, and operation entry points; the rules and parameters for model selection can be configured through interactive methods such as drag and drop, sliders, or drop-down menus; the relationship between user tags, toy scenes, and adapted models is displayed in the form of a topology diagram; and the statistical data of model calls are displayed in real time, supporting the filtering of specific models or scenes to view details.
[0082] 90: In response to user modifications to the graphical intelligent model management page, build and test the modified intelligent model interaction logic.
[0083] The visualization page supports adjusting priority by dragging and dropping model cards, and opening a parameter editing window by clicking on a model card; the simulation effect is displayed synchronously when the strategy is modified; and unauthorized functions are hidden according to the user's role.
[0084] The management console backend parses user operations and extracts key parameters. If the change is "Add / Delete Model", the model management interface is called to register the new model to the intelligent model library or remove the old model from the library. If the change is "Adjust Strategy Rules", the rule engine of the interaction strategy and scheduling module is updated. Based on the changes in the user group and scene mapping, the association table of "User Tag-Scene-Model" is updated.
[0085] The system ensures that model adjustments meet expectations through real-time monitoring and feedback, thereby improving user experience. Administrators can flexibly adjust model configurations and optimize AI capability combinations based on real-time data to achieve more precise personalized services.
[0086] Optionally, the intelligent model library also includes a variety of pre-set visual intelligent models, which are used for image recognition and facial expression recognition based on image data sent by IoT toys.
[0087] The visual intelligence models in the intelligent model library include image recognition models and facial expression recognition models, based on the interaction requirements of IoT toys.
[0088] Among them, the image recognition model is used to analyze the image data collected by IoT toys, such as objects and environmental scenes photographed by users, and output specific recognition results such as "This is a cat" or "This is a picture book".
[0089] Facial expression recognition models are used to analyze users' facial expression images and identify their emotional states, such as "happy," "confused," and "sad," providing emotional feedback for interaction strategies.
[0090] See Figure 7 , Figure 7 This is a schematic diagram of an embodiment of the IoT toy management system provided in this application. The IoT toy management system includes: The central management platform includes modules for automatic speech recognition model management, large language model management, and text-to-speech model management, which can be dynamically selected.
[0091] The central management platform is responsible for the dynamic scheduling and interaction logic decision-making of AI models. It includes an automatic speech recognition model management module: maintaining a speech recognition model library, dynamically selecting suitable speech recognition models based on user profiles, and converting user speech data into text; a large language model management module: managing multiple types of large language models, combining the text content output by the speech recognition model with the user's historical interaction records, and selecting the most suitable large language model to generate semantic understanding results; and a text-to-speech model management module: integrating a multi-timbre, multi-emotion text-to-speech model library, dynamically selecting a text-to-speech model based on the user's voice preferences, toy style, and the emotional content of the text generated by the large language model, and converting the text into natural speech.
[0092] Internet of Things (IoT) toys include a sound acquisition module, a sound playback module, and a wireless communication module.
[0093] IoT toys serve as both the "entry point" and the "exit point" for user interaction, using hardware modules to collect and respond to voice messages.
[0094] The sound acquisition module has a built-in microphone to acquire user voice and convert analog voice signals into digital voice data; the wireless communication module uploads voice data to the central management platform via protocols such as Wi-Fi / Bluetooth and receives response voice data returned by the platform; the sound playback module has a built-in speaker to restore the received response voice data into audible sound, completing the interactive loop.
[0095] The Internet of Things (IoT) toy acquires the user's voice through a sound acquisition module, converts it into corresponding user voice data, and sends the user voice data to the central management platform through a wireless communication module.
[0096] Users trigger the sound acquisition module by interacting with the toy; the microphone collects the voice signal, which is then converted from analog to digital to generate digital voice data. The wireless communication module encrypts the voice data and uploads it to the central management platform.
[0097] The central management platform dynamically matches the automatic speech recognition model used in the current round of dialogue through the automatic speech recognition model management module, the large language model used in the current round of dialogue through the large language model management module, and the text-to-speech model used in the current round of dialogue through the text-to-speech model management module.
[0098] After receiving the speech data, the automatic speech recognition model management module first analyzes the user profile and interaction context, and matches the appropriate automatic speech recognition model from the model library according to the preset model selection strategy. The selected automatic speech recognition model recognizes the speech data and outputs the recognition result to the large language model management module.
[0099] After receiving the text output by the automatic speech recognition model, the large language model management module combines the user's historical interaction records and toy location to match the appropriate large language model; the large language model generates a semantic response based on the text content and passes the result to the text-to-speech model management module.
[0100] The text-to-speech model management module selects an appropriate text-to-speech model based on the user profile's voice preferences, toy style, and the emotional content of the text generated by the large language model. The text-to-speech model converts the text into speech data and generates a transmittable audio file.
[0101] The central management platform sends the response voice data generated by the text-to-speech model to the wireless communication module of the IoT toy, and the IoT toy plays the response voice data through the sound playback module.
[0102] The central management platform acquires user voice data sent by IoT toys.
[0103] The central management platform dynamically matches the automatic speech recognition model used in the current dialogue from a pre-set intelligent model library, and then uses the automatic speech recognition model to perform speech recognition on the user's voice data to generate user speech text. The intelligent model library contains various types of automatic speech recognition models, various types of large language models, and various types of text-to-speech models.
[0104] The central management platform dynamically matches the large language model used in this round of dialogue from the preset intelligent model library, and generates a response text based on the user's voice text through the large language model.
[0105] The central management platform dynamically matches the text-to-speech model used in this round of dialogue from the preset intelligent model library, and sets the language style of the response text through the text-to-speech model to generate response speech data.
[0106] The central management platform sends a response voice data to the IoT toy for playback.
[0107] Optionally, the IoT toy is also equipped with a variety of lightweight smart models, which are used to assist the central management platform in conducting intelligent dialogues.
[0108] In an IoT toy interaction system integrating AI capability management, IoT toys further optimize interaction efficiency and user experience by configuring lightweight intelligent models and wake word recognition models. These two types of models complement the functions of the central management platform from the two dimensions of "localized auxiliary processing" and "active wake-up activation," respectively, making the system architecture more complete and the interaction more intelligent.
[0109] The lightweight smart model configured in the IoT toy is deployed locally on the terminal to reduce interaction latency and alleviate the load on the central platform. By preprocessing simple dialogue requirements, it enables collaborative processing between the local and cloud.
[0110] Optionally, the IoT toy is also equipped with a wake word recognition model, which is used to activate the IoT toy in response to the user's wake word.
[0111] The wake word recognition model serves as an interactive switch for IoT toys. By continuously listening for user wake words such as "Hello, little bear" or "Start conversation," it enables the toy to switch from standby to active state, balancing convenient interaction with power consumption control.
[0112] Unlike existing technologies, this application provides an IoT toy management method integrating AI capabilities. By dynamically selecting and configuring speech recognition models, large language models, and text-to-speech models, it offers a more accurate and personalized interactive experience, significantly improving user satisfaction and the toy's intelligence level. The system dynamically adjusts the configuration strategies of functional modules such as speech recognition models, large language models, and text-to-speech models based on user profiles and IoT toy types, achieving highly customized intelligent interaction. Through continuous optimization of the model matching algorithm, the system not only improves the accuracy of speech recognition but also enhances the subtlety of semantic understanding and emotional expression, ensuring that every dialogue accurately meets user needs and significantly improves the user experience of IoT toys.
[0113] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. An Internet of Things (IoT) toy management method, characterized in that, include: Acquire user voice data sent by IoT toys; The system dynamically matches the automatic speech recognition model used in the current round of dialogue from the preset intelligent model library, and performs speech recognition on the user's speech data through the automatic speech recognition model to generate user speech text; the intelligent model library has preset multiple types of automatic speech recognition models, multiple types of large language models and multiple types of text-to-speech models; The system dynamically matches the large language model used in this round of dialogue from the preset intelligent model library, and generates a response text based on the user's voice text through the large language model. The text-to-speech model used in this round of dialogue is dynamically matched in the preset intelligent model library, and the language style of the response text is set through the text-to-speech model to generate response speech data. The response voice data is sent to the IoT toy for playback.
2. The IoT toy management method according to claim 1, characterized in that, The dynamic matching of the automatic speech recognition model used in this round of dialogue, and the speech recognition of the user's speech data through the automatic speech recognition model to generate user speech text, including: Based on the device information of the IoT toy, user profile data, or the interactive context content in the user voice data, the automatic speech recognition model used in this round of dialogue is selected from the intelligent model library. The automatic speech recognition model is invoked, and the user speech data is converted into corresponding user speech text by combining the device information of the IoT toy, the user profile data, or the interactive context content in the user speech data.
3. The IoT toy management method according to claim 1, characterized in that, The dynamic matching of the large language model used in this round of dialogue, and the generation of response text based on the user's voice text through the large language model, includes: Based on the user's voice text, the device information of the IoT toy, user profile data, the interactive context content in the user's voice data, or the functional type of the IoT toy, the large language model used in this round of dialogue is selected from the intelligent model library. The large language model is invoked, and semantic understanding of the user's speech text is performed through the large language model. The large language model generates response text based on semantic understanding results, the user's voice text, the device information of the IoT toy, user profile data, the interactive context content in the user's voice data, or the functional type of the IoT toy.
4. The IoT toy management method according to claim 1, characterized in that, The dynamic matching uses the text-to-speech model employed in this round of dialogue, and sets the language style of the response text through the text-to-speech model to generate response speech data, including: Based on the response text, the language style preferences of the user profile data, and the style information of the IoT toy, the text-to-speech model used in this round of dialogue is selected from the intelligent model library; Based on the response text, the language style preferences of the user profile data, and the style information of the IoT toy, the timbre, speech rate, tone, or emotional features of the response text are set using the text-to-speech model. The response voice data is generated based on the set timbre, speech rate, tone, and emotional characteristics.
5. The Internet of Things toy management method according to any one of claims 1-4, characterized in that, The IoT toy management method also includes: Obtain management information entered by the administrator or authorized user; The selection strategy for the automatic speech recognition model, the large language model, and the text-to-speech model is dynamically updated based on the management information.
6. The Internet of Things toy management method according to any one of claims 1-4, characterized in that, The IoT toy management method also includes: Displays a graphical intelligent model management page; In response to user modifications to the graphical intelligent model management page, the modified intelligent model interaction logic is constructed and tested.
7. The Internet of Things toy management method according to any one of claims 1-4, characterized in that, The IoT toy management method also includes: The intelligent model library also includes various types of visual intelligent models, which are used for image recognition and facial expression recognition based on the image data sent by the IoT toy.
8. An Internet of Things (IoT) toy management system for implementing the IoT toy management method as described in claims 1-7, characterized in that, include: The central management platform includes modules for automatic speech recognition model management, large language model management, and text-to-speech model management, which can be dynamically selected. Internet of Things (IoT) toys, including a sound acquisition module, a sound playback module, and a wireless communication module; The IoT toy acquires the user's voice through a sound acquisition module, converts it into corresponding user voice data, and sends the user voice data to the central management platform through the wireless communication module. The central management platform dynamically matches the automatic speech recognition model used in the current round of dialogue through the automatic speech recognition model management module, dynamically matches the large language model used in the current round of dialogue through the large language model management module, and dynamically matches the text-to-speech model used in the current round of dialogue through the text-to-speech model management module. The central management platform sends the response voice data generated by the text-to-speech model to the wireless communication module of the IoT toy, and the IoT toy plays the response voice data through the sound playback module.
9. The Internet of Things toy management system according to claim 8, characterized in that, The IoT toy is also equipped with a variety of lightweight smart models, which are used to assist the central management platform in conducting intelligent dialogues.
10. The Internet of Things toy management system according to claim 8, characterized in that, The IoT toy is also equipped with a wake word recognition model, which is used to activate the IoT toy in response to the user's wake word.