Intelligent sound box system with doll identification function

By integrating doll recognition and interaction functions in the smart speaker system, the problem of single functions of existing smart speakers and poor human-computer interaction is solved, achieving a higher sense of immersion and personalized entertainment experience.

CN119996881AInactive Publication Date: 2025-05-13李佳佳
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510013210.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent
Patent Text Reader

Abstract

The invention discloses an intelligent sound box system with a doll identification function. The intelligent sound box system comprises the following steps: S1, hardware design: S1.1, a top area: designing an area for placing a doll; an inductor, such as an RFID card reader, an NFC sensor, a visual identification camera (such as a 2D bar code / two-dimensional code / image identification) and the like, can be embedded into the top; and S1.2, the sound box part is provided with a high-quality loudspeaker, a microphone array, an LED lamp, a touch screen and the like, and is used for outputting sound, displaying interactive information, displaying states and the like. By identifying different dolls and interacting with the dolls, the intelligent sound box not only is a tool for playing audio, but also becomes an intelligent partner capable of accompanying a user. The interaction between the user and different dolls can enrich daily life and increase entertainment and emotion connection, and each doll corresponds to a unique virtual character, so that the immersion of the user is improved. For example, children can select story contents or talk with their favorite roles according to their preferences, so that the dull feeling of single content is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart speakers, and in particular to a smart speaker system with a doll recognition function. Background Art

[0002] A smart speaker is a device that integrates multiple functions such as audio playback, voice recognition, and intelligent control. It usually interacts with users through built-in voice assistants (such as Alexa, Siri, Google Assistant, etc.). In addition to providing high-quality audio playback, smart speakers can also perform tasks such as controlling home appliances, checking the weather, playing music, setting reminders, and answering questions. It usually has Wi-Fi or Bluetooth connection functions, can be wirelessly connected to smartphones, home appliances, etc., supports voice commands, App control, and multi-device linkage, and provides a convenient smart life experience. With the development of technology, modern smart speakers often add a variety of innovative functions, such as sound optimization, remote voice interaction, and device self-learning capabilities, and gradually become the center of home entertainment and smart homes.

[0003] In the existing technology, the functions of smart speakers are relatively simple, the human-computer interaction is not ideal, and the personalization and entertainment are low. For this reason, we propose a smart speaker system with the function of recognizing dolls.

[0004] To this end, we propose an intelligent speaker system with the function of recognizing dolls. Summary of the invention

[0005] To achieve the above object, the present invention provides the following technical solution: an intelligent speaker system with a doll recognition function, comprising the following steps:

[0006] S1: Hardware Design

[0007] S1.1: Top area: Design an area where the doll can be placed. Sensors can be embedded in the top, such as RFID readers, NFC sensors, visual recognition cameras (such as 2D barcodes / QR codes / image recognition), etc.

[0008] S1.2: Speaker part: equipped with high-quality speakers, microphone arrays, LED lights and touch screens, etc., used to output sound, display interactive information, display status, etc.;

[0009] S1.3: Touch screen (optional): displays the doll information, the currently playing story or content, and can even display the interface for interacting with the doll;

[0010] S1.4: RFID / NFC module: Identify dolls through RFID / NFC tags. Each doll has a unique tag built into it, and the smart speaker identifies the corresponding character by reading the tag of the doll;

[0011] S1.5: Visual recognition module (optional): If image recognition technology (such as camera + AI algorithm) is used, different dolls can be identified by recognizing their appearance, color and features;

[0012] S1.6: Microphone array: used for voice recognition and receiving user commands.

[0013] S2: Doll recognition and role mapping

[0014] S2.1: Doll tags and character database: Each doll has a unique identification tag (RFID / NFC or image features). The speaker identifies the doll by reading the tag or identifying the image features. A character database is maintained inside the speaker or on a cloud server. The database contains the attributes of each doll (such as name, gender, story content, voice features, etc.);

[0015] S2.2: Doll identification process: The user places a doll on the top of the speaker, and the speaker identifies the doll through RFID / NFC or image recognition technology. The speaker queries the database to find the virtual character information (story, voice, personality, etc.) associated with the doll.

[0016] S3: Story and Voice Content Management

[0017] S3.1: Cloud server: Stories, voices, and dialogues are stored in the cloud server. Each doll corresponds to a set of resources, including sound files: voices of different dolls, audio files of stories; dialogue library: voice data, dialogue logic, and interactive content of each character; personalized stories: stories and interactive scenes customized according to the attributes of the characters and dolls;

[0018] S3.2: Content request and playback: When the doll is recognized, the speaker requests the corresponding content data (audio, dialogue script, etc.) through the network; the cloud server returns the corresponding audio file and dialogue script, and the speaker plays a specific story or starts the dialogue mode; the speaker supports switching between different dolls, and adjusts the output story content, sound effects and voice according to the currently recognized doll.

[0019] S4: Human-computer interaction and speech recognition

[0020] S4.1: Voice recognition module: The speaker integrates voice recognition technology to receive user commands or interact with the doll. Users can control the behavior of the doll through voice questions or commands, such as: "Doll tell a story" or "Talk to the doll. The microphone array of the speaker can capture the user's voice and recognize it through AI;

[0021] S4.2: Dialogue system: Dialogue engine: Each doll has an independent dialogue system that generates responses based on artificial intelligence (such as GPT or other dialogue systems), context understanding: The speaker can recognize the user's questions or requests and combine them with the doll's character background to generate interesting and personalized dialogue content. Each doll has a different tone and accent, allowing users to feel the personality differences between different characters;

[0022] S4.3: Feedback and emotion recognition: The speaker can provide feedback through multimodal information such as voice intonation and expression. Users can ask questions to the doll, and the doll will answer according to the preset dialogue library. At the same time, the doll can also make appropriate adjustments based on the user's emotional feedback (through voice emotion recognition technology).

[0023] S5: Sentiment Analysis and Tone Recognition Technology

[0024] S5.1: High-sensitivity microphone array: Similar to the original microphone array function, a new set of microphone sensors is added to support more accurate capture of changes in the user's pitch, tone, speed, etc. These microphones can provide high-quality data for the speech emotion recognition system;

[0025] S5.2: Equipped with a real-time sound processing chip, which is responsible for capturing and analyzing the user's audio signals in real time and converting them into feature data suitable for sentiment analysis for subsequent algorithm analysis and processing;

[0026] S5.3: Sentiment classification algorithm: Use deep learning technology (such as convolutional neural network CNN, long short-term memory network LSTM, etc.) to perform sentiment analysis on the user's voice signal. The system will identify the emotion category (such as happiness, sadness, anger, surprise, etc.) based on the user's tone, voice, speed, pauses and other parameters;

[0027] S5.4: Emotion conversion model: The emotion recognition model maps the speech signal to the corresponding emotion label, and then adjusts the doll's response based on the character characteristics of the doll. For example, if the user's tone is high and excited, the doll's response will be more vivid and energetic;

[0028] S5.5: The characters of the dolls also have a dedicated emotional library, and each character has a different emotional response pattern. For example, a gentle character may use a comforting tone under the "sad" emotion, while a brave character may respond in an encouraging and supportive manner. The emotional library includes: tone (such as soft, passionate, comforting), speaking rhythm, speed, emotional processing of voice (such as soft voice or strong tone);

[0029] S5.6: Identify user emotions: When a user issues a voice command to the speaker or talks to the doll, the speaker first uses voice recognition technology to convert the user's voice into text. Then, the sentiment analysis module processes the user's voice data in real time and analyzes the emotional information. At this time, the system will determine the user's current emotional state. If the user's tone is more excited or happy, the emotion classification may be happy; if the tone is low and the speed is slow, the emotion classification may be sad; if the voice is rapid and has an increased volume, the emotion may be marked as angry;

[0030] S5.7: Adjust the doll response: Based on the results of the emotion analysis, the speaker will adjust the doll's response. For example: Happy emotion: If the user is in an excited state, the doll's voice will become more lively, faster, and more positive, such as: "Wow, it sounds like you are really happy! Let me tell you a more interesting story!, Sad emotion: If the user's tone is low or sad, the doll's tone will become gentle and comforting. For example: "It sounds like you are a little unhappy. It's okay, I'm here with you, let me tell you a warm story, okay?, Angry emotion: If the user's tone has anger or dissatisfaction, the doll's voice may be more peaceful and try to calm the emotion, for example: "I can feel that you are a little angry. Want to talk about what happened? I can help you.

[0031] S5.8: Feedback and emotional evolution: The doll can continuously adjust its emotional response based on the user's feedback. Through continuous interaction, the speaker will record the user's feedback in different emotional states (for example, whether the user prefers a positive tone or a soft tone), and optimize the emotional response in future conversations. In continuous conversations, the system can track changes in user emotions. For example, when the user changes from excitement to frustration, the doll will gradually adjust the tone through the content of the conversation, or give emotional support at the right time to create a deeper emotional connection. User personalization settings: Allow users to customize the doll's voice emotional response style. For example, if some users like the doll to appear more humorous or calm in emotional interactions, users can adjust the doll's emotional response mode through settings. Character diversity: Different doll characters have different emotional response styles. For example, a robot character may be more rational in emotional response, while an animal character may appear more emotionally resonant;

[0032] S5.9: Emotion database update: All emotion analysis and response data can be stored in the cloud, and the emotion database can be optimized and updated based on user feedback. In this way, over time, the speaker's recognition and response to user emotions will become more and more accurate. Multi-doll synchronization update: When new doll characters or new interactive content are launched, the system will update the relevant emotion data to ensure that all dolls can make appropriate responses based on the user's emotional state.

[0033] S6: Cloud Server Architecture

[0034] S6.1: Cloud service: provides integrated services such as user doll recognition, content storage and downloading, voice recognition, and dialogue engine. When the doll is recognized, the speaker requests the corresponding character content from the server, such as stories, voice files, and dialogue logic;

[0035] S6.2: Data management and update: The server supports regular updates of the doll's story content, dialogue data, and character voice. Users can view and download updated content through the application (or the speaker itself), and provide a doll management background. Users can select, add, and delete dolls through the application, and personalize the story content as needed.

[0036] S7: Security and Privacy

[0037] S7.1: Data encryption: All data transmission between the speaker and the server (including doll identification data, voice commands, etc.) should be protected by encryption technology;

[0038] S7.2: Privacy protection: The user's voice data should be processed locally as much as possible to avoid unnecessary data transmission and protect the user's privacy;

[0039] S7.3: Child Protection: If the device is intended for children, provide parental control options to restrict inappropriate content and functionality.

[0040] Preferably, the RFID / NFC module in step S1 is used to identify the doll by reading the built-in RFID tag, confirming the identity of the doll, and visual recognition camera: if image recognition technology is used, the shape, color or appearance characteristics of the doll can be analyzed by AI algorithm for identification. Use deep learning algorithms (such as convolutional neural networks CNN) to enhance the recognition ability of the doll, sensor type selection: if RFID technology is used, a unique RFID tag needs to be embedded for each doll; if visual recognition is used, the camera resolution is at least 720P to ensure recognition accuracy, configure high-quality speakers (such as Dolby or Hi-Res Audio certified speakers) to ensure sound quality, configure microphone arrays, support multi-microphone arrays to achieve high-precision voice recognition and noise elimination, and LED lights are used to display status and interactive feedback.

[0041] Preferably, in step S1, the RFID / NFC module: select a reliable RFID / NFC module (such as NXP PN532), and select an appropriate reading distance according to the size of the doll and the usage scenario. The RFID tag can be embedded in the doll, and the card reader on the top of the speaker can read the tag through near-field communication to obtain a unique identifier. Visual recognition module: If image recognition technology is used, the speaker needs to be equipped with a camera (such as OmniVision OV5640), and the captured image is input into an edge computing device (such as Jetson Nano) for processing. The doll is recognized in real time through a deep learning algorithm (for example, using TensorFlow Lite). Microphone array: Up to 6-8 microphones are configured to form a microphone array, which can achieve high-precision voice positioning and noise suppression. The microphone array should support sound source localization technology to ensure that the system can accurately recognize voice input from different directions.

[0042] Preferably, the character database design in step S2 is as follows: the character database is stored in the cloud or locally, and each doll corresponds to a record. Each record contains: the unique identifier of the doll (such as RFID tag number or image feature), the name, gender, age, personality characteristics, etc. of the doll, the story, audio file, voice file, dialogue script, etc. associated with the doll, the voice characteristics of each doll (such as intonation, timbre, speaking speed, etc.) and personalized content, and the metadata of the doll is stored using a relational database (such as MySQL) or a NoSQL database (such as MongoDB). The information of each doll should include queryable fields, such as name, category, personality, etc., so that the system can quickly locate it when querying.

[0043] Preferably, in step S2, the user places the doll on the top of the speaker, the RFID / NFC module reads the doll tag and transmits its ID to the speaker host, and the speaker host queries the locally cached character database (or requests the cloud database through an API) to obtain virtual character information related to the doll. If visual recognition is used, the speaker inputs the image captured by the camera into the AI ​​algorithm, identifies the appearance features of the doll (such as color, shape, pattern, etc.), and obtains relevant virtual character information from the database.

[0044] Preferably, in the step S3, the cloud server: all audio files (stories, character voices) and interactive data are stored on the cloud server to ensure that users can get the latest content at any time. Cloud storage services (such as Amazon S3) can be used to store a large number of audio files, and multi-format support: the story content should support multiple formats, such as MP3 or WAV audio files. The dialogue content can be stored in JSON or XML format for subsequent parsing and playback. Content request process: when the doll is identified, the speaker requests the cloud server through the network to obtain the story and dialogue content related to the doll. This process includes: the doll identifies and transmits the identifier, the speaker sends a request to the server, the server returns the corresponding audio and dialogue script, the speaker plays the content, and displays relevant interactive information (such as the touch screen displays the current story or character information).

[0045] Preferably, the speech recognition technology selection in step S4: adopt mature speech recognition technology, such as Google Assistant SDK, Amazon Alexa Voice Service or iFlytek speech recognition, support voice command recognition and natural language processing, real-time speech recognition: the speaker receives and recognizes the user's voice command (such as "tell a story" or "talk to the doll") in real time, and triggers corresponding actions according to the command, dialogue engine: each doll uses an independent dialogue system (such as a chatbot engine based on GPT or Dialogflow) to generate a response. The dialogue system will generate natural dialogue content based on the user's questions and the background story of the doll, personalized tone and accent: the speech synthesis system of each doll (such as WaveNet or Tacotron 2) adjusts the speed, pitch and tone according to the character design of the doll to present different personalities, and sentiment analysis module: the speaker recognizes the user's emotions through speech sentiment analysis (using technologies such as IBM Watson Tone Analyzer) and adjusts the response of the doll. For example, when the user's tone is excited, the doll may comfort the user; when the tone is excited, the doll will respond more actively.

[0046] Preferably, in step S6, cloud platform selection: select a reliable cloud platform (such as AWS, Google Cloud or Microsoft Azure) to provide services such as doll recognition, voice recognition, content storage, and dialogue generation. All doll recognition requests, content requests, etc. are transmitted through RESTful APIs. Server-side support: The server should support load balancing and be able to maintain high-performance response under multiple concurrent requests.

[0047] Preferably, in step S6, the updates are performed regularly: the cloud server regularly pushes new content updates (such as new stories, new dialogues, character voices, etc.), and the speaker automatically downloads via Wi-Fi; the doll management background: a background management system is provided, and users can view and download updated content through mobile applications, and can also customize personalized doll stories and dialogues.

[0048] Compared with the prior art, the present invention provides an intelligent speaker system with a doll recognition function, which has the following beneficial effects:

[0049] 1. This smart speaker system with doll recognition function can recognize different dolls and interact with them. The smart speaker is not only a tool for playing audio, but also a smart companion that can accompany users. The interaction between users and different dolls can enrich daily life, increase entertainment and emotional connection. Each doll corresponds to a unique virtual character. This personalized interaction will enhance the user's immersion. For example, children can choose the story content according to their preferences or talk to their favorite characters, avoiding the boredom of a single content.

[0050] 2. The smart speaker system with doll recognition function can provide clear, layered sound output and high-precision voice recognition through high-quality speakers and microphone arrays. This not only improves the playback effect of stories and dialogues, but also enhances the recognition accuracy of voice commands. The integrated touch screen and sensors (such as RFID / NFC and image recognition modules) allow users to choose a variety of ways to interact with the speaker. Users can control the speaker by voice, or view doll information or make more detailed personalized settings through the touch screen.

[0051] 3. The smart speaker system with doll recognition function ensures accurate identification of dolls through the application of RFID / NFC and image recognition technology. The speaker can quickly identify the doll placed by the user and accurately load the relevant role and content. This greatly improves the speed and accuracy of the system response. Through the design of the role database, the system can accurately map the identity characteristics of the doll (such as name, gender, personality, etc.) and provide targeted stories and dialogue content, so that each doll is closely associated with its unique virtual role, enhancing the user experience.

[0052] 4. The smart speaker system with doll recognition function stores content and supports automatic updates through cloud servers, so the system can keep the content fresh. Users can get the latest stories, voices, and conversations without worrying about the content being outdated or unable to be updated. This method reduces the storage burden while ensuring the continuity of content. The stories and conversations corresponding to the dolls can be adjusted at any time through personalized configuration in the cloud. For example, a doll can provide stories and conversations of different difficulty or style for users of different age groups, improving entertainment and education.

[0053] 5. This smart speaker system with the function of recognizing dolls uses advanced voice recognition technology. The system can quickly and accurately recognize the user's voice commands and make reasonable responses according to different commands. Users do not need tedious manual operations, and can interact with the dolls only through voice, which greatly improves the convenience and interactive experience. After integrating emotional analysis technology, the speaker can not only recognize the user's voice commands, but also adjust the doll's response method according to the user's emotions. For example, when the user is excited, the doll can respond with a more enthusiastic and active tone; when the user is depressed, the doll will communicate in a gentle and comforting tone. This interaction is more in line with the user's emotional needs and improves user satisfaction.

[0054] 6. The smart speaker system with doll recognition function can achieve rapid data update and accurate distribution through cloud content storage and management. The cloud architecture supports large-scale concurrent requests, ensuring efficient processing and smooth experience when multiple users request different content at the same time. The cloud server can not only provide content storage, but also push customized content based on the user's interaction history and preferences. For example, if a user frequently selects a certain type of story or doll, the system can intelligently recommend related content to enhance the personalized experience. DETAILED DESCRIPTION

[0055] The technical solutions in the embodiments of the present invention are described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0056] Example

[0057] Embodiment of a smart speaker system with doll recognition function

[0058] An intelligent speaker system with a doll recognition function comprises the following steps:

[0059] S1: Hardware Design

[0060] S1.1: Top area: Design an area where the doll can be placed. Sensors can be embedded in the top, such as RFID readers, NFC sensors, visual recognition cameras (such as 2D barcodes / QR codes / image recognition), etc.

[0061] S1.2: Speaker part: equipped with high-quality speakers, microphone arrays, LED lights and touch screens, etc., used to output sound, display interactive information, display status, etc.;

[0062] S1.3: Touch screen (optional): displays the doll information, the currently playing story or content, and can even display the interface for interacting with the doll;

[0063] S1.4: RFID / NFC module: Identify dolls through RFID / NFC tags. Each doll has a unique tag built into it, and the smart speaker identifies the corresponding character by reading the tag of the doll;

[0064] S1.5: Visual recognition module (optional): If image recognition technology (such as camera + AI algorithm) is used, different dolls can be identified by recognizing their appearance, color and features;

[0065] S1.6: Microphone array: used for voice recognition and receiving user commands.

[0066] S2: Doll recognition and role mapping

[0067] S2.1: Doll tags and character database: Each doll has a unique identification tag (RFID / NFC or image features). The speaker identifies the doll by reading the tag or identifying the image features. A character database is maintained inside the speaker or on a cloud server. The database contains the attributes of each doll (such as name, gender, story content, voice features, etc.);

[0068] S2.2: Doll identification process: The user places a doll on the top of the speaker, and the speaker identifies the doll through RFID / NFC or image recognition technology. The speaker queries the database to find the virtual character information (story, voice, personality, etc.) associated with the doll.

[0069] S3: Story and Voice Content Management

[0070] S3.1: Cloud server: Stories, voices, and dialogues are stored in the cloud server. Each doll corresponds to a set of resources, including sound files: voices of different dolls, audio files of stories; dialogue library: voice data, dialogue logic, and interactive content of each character; personalized stories: stories and interactive scenes customized according to the attributes of the characters and dolls;

[0071] S3.2: Content request and playback: When the doll is recognized, the speaker requests the corresponding content data (audio, dialogue script, etc.) through the network; the cloud server returns the corresponding audio file and dialogue script, and the speaker plays a specific story or starts the dialogue mode; the speaker supports switching between different dolls, and adjusts the output story content, sound effects and voice according to the currently recognized doll.

[0072] S4: Human-computer interaction and speech recognition

[0073] S4.1: Voice recognition module: The speaker integrates voice recognition technology to receive user commands or interact with the doll. Users can control the behavior of the doll through voice questions or commands, such as: "Doll tell a story" or "Talk to the doll. The microphone array of the speaker can capture the user's voice and recognize it through AI;

[0074] S4.2: Dialogue system: Dialogue engine: Each doll has an independent dialogue system that generates responses based on artificial intelligence (such as GPT or other dialogue systems), context understanding: The speaker can recognize the user's questions or requests and combine them with the doll's character background to generate interesting and personalized dialogue content. Each doll has a different tone and accent, allowing users to feel the personality differences between different characters;

[0075] S4.3: Feedback and emotion recognition: The speaker can provide feedback through multimodal information such as voice intonation and expression. Users can ask questions to the doll, and the doll will answer according to the preset dialogue library. At the same time, the doll can also make appropriate adjustments based on the user's emotional feedback (through voice emotion recognition technology).

[0076] S5: Sentiment Analysis and Tone Recognition Technology

[0077] S5.1: High-sensitivity microphone array: Similar to the original microphone array function, a new set of microphone sensors is added to support more accurate capture of changes in the user's pitch, tone, speed, etc. These microphones can provide high-quality data for the speech emotion recognition system;

[0078] S5.2: Equipped with a real-time sound processing chip, which is responsible for capturing and analyzing the user's audio signals in real time and converting them into feature data suitable for sentiment analysis for subsequent algorithm analysis and processing;

[0079] S5.3: Sentiment classification algorithm: Use deep learning technology (such as convolutional neural network CNN, long short-term memory network LSTM, etc.) to perform sentiment analysis on the user's voice signal. The system will identify the emotion category (such as happiness, sadness, anger, surprise, etc.) based on the user's tone, voice, speed, pauses and other parameters;

[0080] S5.4: Emotion conversion model: The emotion recognition model maps the speech signal to the corresponding emotion label, and then adjusts the doll's response based on the character characteristics of the doll. For example, if the user's tone is high and excited, the doll's response will be more vivid and energetic;

[0081] S5.5: The characters of the dolls also have a dedicated emotional library, and each character has a different emotional response pattern. For example, a gentle character may use a comforting tone under the "sad" emotion, while a brave character may respond in an encouraging and supportive manner. The emotional library includes: tone (such as soft, passionate, comforting), speaking rhythm, speed, emotional processing of voice (such as soft voice or strong tone);

[0082] S5.6: Identify user emotions: When a user issues a voice command to the speaker or talks to the doll, the speaker first uses voice recognition technology to convert the user's voice into text. Then, the sentiment analysis module processes the user's voice data in real time and analyzes the emotional information. At this time, the system will determine the user's current emotional state. If the user's tone is more excited or happy, the emotion classification may be happy; if the tone is low and the speed is slow, the emotion classification may be sad; if the voice is rapid and has an increased volume, the emotion may be marked as angry;

[0083] S5.7: Adjust the doll response: Based on the results of the emotion analysis, the speaker will adjust the doll's response. For example: Happy emotion: If the user is in an excited state, the doll's voice will become more lively, faster, and more positive, such as: "Wow, it sounds like you are really happy! Let me tell you a more interesting story!, Sad emotion: If the user's tone is low or sad, the doll's tone will become gentle and comforting. For example: "It sounds like you are a little unhappy. It's okay, I'm here with you, let me tell you a warm story, okay?, Angry emotion: If the user's tone has anger or dissatisfaction, the doll's voice may be more peaceful and try to calm the emotion, for example: "I can feel that you are a little angry. Want to talk about what happened? I can help you.

[0084] S5.8: Feedback and emotional evolution: The doll can continuously adjust its emotional response based on the user's feedback. Through continuous interaction, the speaker will record the user's feedback in different emotional states (for example, whether the user prefers a positive tone or a soft tone), and optimize the emotional response in future conversations. In continuous conversations, the system can track changes in user emotions. For example, when the user changes from excitement to frustration, the doll will gradually adjust the tone through the content of the conversation, or give emotional support at the right time to create a deeper emotional connection. User personalization settings: Allow users to customize the doll's voice emotional response style. For example, if some users like the doll to appear more humorous or calm in emotional interactions, users can adjust the doll's emotional response mode through settings. Character diversity: Different doll characters have different emotional response styles. For example, a robot character may be more rational in emotional response, while an animal character may appear more emotionally resonant;

[0085] S5.9: Emotion database update: All emotion analysis and response data can be stored in the cloud, and the emotion database can be optimized and updated based on user feedback. In this way, over time, the speaker's recognition and response to user emotions will become more and more accurate. Multi-doll synchronization update: When new doll characters or new interactive content are launched, the system will update the relevant emotion data to ensure that all dolls can make appropriate responses based on the user's emotional state.

[0086] S6: Cloud Server Architecture

[0087] S6.1: Cloud service: provides integrated services such as user doll recognition, content storage and downloading, voice recognition, and dialogue engine. When the doll is recognized, the speaker requests the corresponding character content from the server, such as stories, voice files, and dialogue logic;

[0088] S6.2: Data management and update: The server supports regular updates of the doll's story content, dialogue data, and character voice. Users can view and download updated content through the application (or the speaker itself), and provide a doll management background. Users can select, add, and delete dolls through the application, and personalize the story content as needed.

[0089] S7: Security and Privacy

[0090] S7.1: Data encryption: All data transmission between the speaker and the server (including doll identification data, voice commands, etc.) should be protected by encryption technology;

[0091] S7.2: Privacy protection: The user's voice data should be processed locally as much as possible to avoid unnecessary data transmission and protect the user's privacy;

[0092] S7.3: Child Protection: If the device is intended for children, provide parental control options to restrict inappropriate content and functionality.

[0093] Specifically, the RFID / NFC module in step S1 is used to identify the doll by reading the built-in RFID tag, confirming the identity of the doll. The visual recognition camera: if image recognition technology is used, the shape, color or appearance characteristics of the doll can be analyzed by the AI ​​algorithm for identification. Use deep learning algorithms (such as convolutional neural networks CNN) to enhance the recognition ability of the doll. Sensor type selection: if RFID technology is used, a unique RFID tag needs to be embedded for each doll; if visual recognition is used, the camera resolution is at least 720P to ensure recognition accuracy, configure high-quality speakers (such as Dolby or Hi-Res Audio certified speakers) to ensure sound quality, configure a microphone array, support multi-microphone arrays to achieve high-precision voice recognition and noise elimination, and LED lights are used to display status and interactive feedback.

[0094] Specifically, in step S1, the RFID / NFC module: select a reliable RFID / NFC module (such as NXP PN532), and select an appropriate reading distance according to the size of the doll and the usage scenario. The RFID tag can be embedded in the doll, and the card reader on the top of the speaker can read the tag through near-field communication to obtain a unique identifier. Visual recognition module: If image recognition technology is used, the speaker needs to be equipped with a camera (such as OmniVision OV5640), and the captured image is input into an edge computing device (such as Jetson Nano) for processing. The doll is recognized in real time through a deep learning algorithm (for example, using TensorFlow Lite). Microphone array: Up to 6-8 microphones are configured to form a microphone array, which can achieve high-precision voice positioning and noise suppression. The microphone array should support sound source localization technology to ensure that the system can accurately recognize voice input from different directions.

[0095] Specifically, the character database design in step S2 is as follows: the character database is stored in the cloud or locally, and each doll corresponds to a record. Each record contains: the unique identifier of the doll (such as RFID tag number or image feature), the name, gender, age, personality characteristics, etc. of the doll, the story, audio file, voice file, dialogue script, etc. associated with the doll, the voice characteristics of each doll (such as intonation, timbre, speaking speed, etc.) and personalized content, and the metadata of the doll is stored using a relational database (such as MySQL) or a NoSQL database (such as MongoDB). The information of each doll should include queryable fields, such as name, category, personality, etc., so that the system can quickly locate it when querying.

[0096] Specifically, in step S2, the user places the doll on the top of the speaker, the RFID / NFC module reads the doll tag and transmits its ID to the speaker host, and the speaker host queries the locally cached character database (or requests the cloud database through an API) to obtain virtual character information related to the doll. If visual recognition is used, the speaker inputs the image captured by the camera into the AI ​​algorithm, identifies the appearance features of the doll (such as color, shape, pattern, etc.), and obtains relevant virtual character information from the database.

[0097] Specifically, in step S3, the cloud server: all audio files (stories, character voices) and interactive data are stored on the cloud server to ensure that users can get the latest content at any time. Cloud storage services (such as Amazon S3) can be used to store a large number of audio files. Multi-format support: story content should support multiple formats, such as MP3 or WAV audio files. The dialogue content can be stored in JSON or XML format for subsequent parsing and playback. Content request process: When the doll is identified, the speaker requests the cloud server through the network to obtain the story and dialogue content related to the doll. This process includes: the doll identifies and transmits the identifier, the speaker sends a request to the server, the server returns the corresponding audio and dialogue script, the speaker plays the content, and displays relevant interactive information (such as the touch screen displays the current story or character information).

[0098] Specifically, the speech recognition technology selection in step S4: adopt mature speech recognition technology, such as Google Assistant SDK, Amazon Alexa Voice Service or iFlytek speech recognition, support voice command recognition and natural language processing, real-time speech recognition: the speaker receives and recognizes the user's voice command (such as "tell a story" or "talk to the doll") in real time, and triggers corresponding actions according to the command, dialogue engine: each doll uses an independent dialogue system (such as a chatbot engine based on GPT or Dialogflow) to generate a response. The dialogue system will generate natural dialogue content based on the user's questions and the background story of the doll, personalized tone and accent: the speech synthesis system of each doll (such as WaveNet or Tacotron 2) adjusts the speed, pitch and tone according to the character design of the doll to present different personalities, and sentiment analysis module: the speaker recognizes the user's emotions through speech sentiment analysis (using technologies such as IBM Watson Tone Analyzer) and adjusts the doll's response. For example, when the user's tone is excited, the doll may comfort the user; when the tone is excited, the doll will respond more actively.

[0099] Specifically, in step S6, cloud platform selection: select a reliable cloud platform (such as AWS, Google Cloud or Microsoft Azure) to provide services such as doll recognition, voice recognition, content storage, and dialogue generation. All doll recognition requests, content requests, etc. are transmitted through RESTful APIs. Server-side support: The server should support load balancing and be able to maintain high-performance response under multiple concurrent requests.

[0100] Specifically, in step S6, the cloud server regularly updates new content (such as new stories, new dialogues, character voices, etc.), and the speaker automatically downloads via Wi-Fi. The doll management background provides a background management system, and users can view and download updated content through mobile applications, and can also customize personalized doll stories and dialogues.

[0101] Through the above technical solution, in the present invention, by identifying and interacting with different dolls, the smart speaker is not only a tool for playing audio, but also becomes an intelligent companion that can accompany the user. The interaction between users and different dolls can enrich daily life, increase entertainment and emotional connection. Each doll corresponds to a unique virtual character, and this personalized interaction will enhance the user's immersion. For example, children can choose story content or talk to their favorite characters according to their preferences, avoiding the dullness of single content; through high-quality speakers and microphone arrays, the speaker can provide clear, layered sound output, and high-precision voice recognition. This not only improves the playback effect of stories and conversations, but also enhances the recognition accuracy of voice commands. The integrated touch screen and sensors (such as RFID / NFC and image recognition modules) allow users to choose multiple ways to operate when interacting with the speaker. Users can control the speaker by voice, or view doll information or make more detailed personalized settings through the touch screen; the application of RFID / NFC and image recognition technology ensures the accurate recognition of dolls, and the speaker can quickly identify the dolls placed by the user and accurately load the roles and content related to them. This greatly improves the speed and accuracy of the system response. Through the design of the character database, the system can accurately map the identity characteristics of the doll (such as name, gender, personality, etc.), and provide targeted stories and dialogue content, so that each doll is closely associated with its unique virtual role, enhancing the user experience; by storing content on the cloud server and supporting automatic updates, the system can keep the content fresh. Users can get the latest stories, voices and dialogues without worrying about the content being outdated or unable to be updated. This method reduces the storage burden while ensuring the continuity of the content. The stories and dialogue content corresponding to the doll can be adjusted at any time through personalized configuration in the cloud. For example, a doll can provide stories and dialogues of different difficulty or style for users of different age groups, improving entertainment and education; by adopting advanced voice recognition technology, the system can quickly and accurately recognize the user's voice commands and make reasonable responses according to different commands. Users do not need tedious manual operations, and can interact with the doll only through voice, which greatly improves the convenience and interactive experience. After integrating emotional analysis technology, the speaker can not only recognize the user's voice commands, but also adjust the doll's response method according to the user's emotions. For example, when the user is excited, the doll can respond with a more enthusiastic and active tone; when the user is depressed, the doll will communicate with a gentle and comforting tone. This interaction is more in line with the user's emotional needs and improves user satisfaction; through cloud content storage and management, data can be quickly updated and accurately distributed.The cloud architecture supports large-scale concurrent requests, ensuring efficient processing and smooth experience when multiple users request different content at the same time. The cloud server can not only provide content storage, but also push customized content based on the user's interaction history and preferences. For example, if a user frequently selects a certain type of story or doll, the system can intelligently recommend related content to enhance the personalized experience.

[0102] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent speaker system with a doll recognition function, characterized in that: The following steps are involved: S1: Hardware Design S1.1: Top area: Design an area where the doll can be placed. Sensors can be embedded on the top, such as RFID readers, NFC sensors, visual recognition cameras (such as 2D barcodes / QR codes / image recognition), etc. S1.2: Speaker part: equipped with high-quality speakers, microphone arrays, LED lights and touch screens, etc., used to output sound, display interactive information, display status, etc.; S1.3: Touch screen (optional): displays the doll information, the currently playing story or content, and can even display the interface for interacting with the doll; S1.4: RFID / NFC module: Identify dolls through RFID / NFC tags. Each doll has a unique tag built into it, and the smart speaker identifies the corresponding character by reading the tag of the doll; S1.5: Visual recognition module (optional): If image recognition technology (such as camera + AI algorithm) is used, different dolls can be identified by recognizing their appearance, color and features; S1.6: Microphone array: used for voice recognition and receiving user commands. S2: Doll recognition and role mapping S2.1: Doll tags and character database: Each doll has a unique identification tag (RFID / NFC or image features). The speaker identifies the doll by reading the tag or identifying the image features. A character database is maintained inside the speaker or on a cloud server. The database contains the attributes of each doll (such as name, gender, story content, voice features, etc.); S2.2: Doll identification process: The user places a doll on the top of the speaker, and the speaker identifies the doll through RFID / NFC or image recognition technology. The speaker queries the database to find the virtual character information (story, voice, personality, etc.) associated with the doll. S3: Story and Voice Content Management S3.1: Cloud server: Stories, voices, and dialogues are stored in the cloud server. Each doll corresponds to a set of resources, including sound files: audio files of different dolls’ voices and stories; Dialogue library: voice data, dialogue logic and interactive content of each character; Personalized stories: stories and interactive scenes customized according to the attributes of characters and dolls; S3.2: Content request and playback: When the doll is recognized, the speaker requests the corresponding content data (audio, dialogue script, etc.) through the network; the cloud server returns the corresponding audio file and dialogue script, and the speaker plays a specific story or starts the dialogue mode; The speaker supports switching between different dolls, and adjusts the output story content, sound effects and voice according to the currently recognized doll. S4: Human-computer interaction and speech recognition S4.1: Voice recognition module: The speaker integrates voice recognition technology to receive user commands or interact with the doll. Users can control the behavior of the doll through voice questions or commands, such as: "Doll tell a story" or "Talk to the doll." The speaker's microphone array can capture the user's voice and recognize it through AI; S4.2: Dialogue system: Dialogue engine: Each doll has an independent dialogue system that generates responses based on artificial intelligence (such as GPT or other dialogue systems), context understanding: The speaker can recognize the user's questions or requests and combine them with the doll's character background to generate interesting and personalized dialogue content. Each doll has a different tone and accent, allowing users to feel the personality differences between different characters; S4.3: Feedback and emotion recognition: The speaker can provide feedback through multimodal information such as voice intonation and expression. Users can ask questions to the doll, and the doll will answer according to the preset dialogue library. At the same time, the doll can also make appropriate adjustments based on the user's emotional feedback (through voice emotion recognition technology). S5: Sentiment Analysis and Tone Recognition Technology S5.1: High-sensitivity microphone array: Similar to the original microphone array function, a new set of microphone sensors is added to support more accurate capture of changes in the user's pitch, tone, speed, etc. These microphones can provide high-quality data for the speech emotion recognition system; S5.2: Equipped with a real-time sound processing chip, which is responsible for capturing and analyzing the user's audio signals in real time and converting them into feature data suitable for sentiment analysis for subsequent algorithm analysis and processing; S5.3: Sentiment classification algorithm: Use deep learning technology (such as convolutional neural network CNN, long short-term memory network LSTM, etc.) to perform sentiment analysis on the user's voice signal. The system will identify the emotion category (such as happiness, sadness, anger, surprise, etc.) based on the user's tone, voice, speed, pauses and other parameters; S5.4: Emotion conversion model: The emotion recognition model maps the speech signal to the corresponding emotion label, and then adjusts the doll's response based on the character characteristics of the doll. For example, if the user's tone is high and excited, the doll's response will be more vivid and energetic; S5.5: The characters of the dolls also have a dedicated emotional library, and each character has a different emotional response pattern. For example, a gentle character may use a comforting tone under the "sad" emotion, while a brave character may respond in an encouraging and supportive manner. The emotional library includes: tone (such as soft, passionate, comforting), speaking rhythm, speed, and emotional processing of voice (such as soft voice or strong tone); S5.6: Identify user emotions: When a user issues a voice command to the speaker or talks to the doll, the speaker first uses voice recognition technology to convert the user's voice into text. Then, the sentiment analysis module processes the user's voice data in real time and analyzes the emotional information. At this time, the system will determine the user's current emotional state. If the user's tone is more excited or happy, the emotion classification may be happy; if the tone is low and the speed is slow, the emotion classification may be sad; if the voice is rapid and has an increased volume, the emotion may be marked as angry; S5.7: Adjust the doll response: Based on the results of the sentiment analysis, the speaker will adjust the doll's response. For example: Happy emotion: If the user is in an excited state, the doll's voice will become more lively, faster, and more positive, such as: "Wow, it sounds like you are really happy! Let me tell you a more interesting story!, Sad emotion: If the user's tone is low or sad, the doll's tone will become gentle and comforting. For example: "It sounds like you are a little unhappy. It's okay, I'm here with you, let me tell you a warm story, okay?, Angry emotion: If the user's tone has anger or dissatisfaction, the doll's voice may be more peaceful and try to calm the emotions, for example: "I can feel that you are a little angry. Want to talk about what happened? I can help you. S5.8: Feedback and emotional evolution: The doll can continuously adjust its emotional response based on the user's feedback. Through continuous interaction, the speaker will record the user's feedback in different emotional states (for example, whether the user prefers a positive tone or a soft tone), and optimize the emotional response in future conversations. In continuous conversations, the system can track changes in user emotions. For example, when the user changes from excitement to frustration, the doll will gradually adjust the tone through the content of the conversation, or give emotional support at the right time to create a deeper emotional connection. User personalization settings: Allow users to customize the doll's voice emotional response style. For example, if some users like the doll to appear more humorous or calm in emotional interactions, users can adjust the doll's emotional response mode through settings. Character diversity: Different doll characters have different emotional response styles. For example, a robot character may be more rational in emotional response, while an animal character may appear more emotionally resonant; S5.9: Emotion database update: All emotion analysis and response data can be stored in the cloud, and the emotion database can be optimized and updated based on user feedback. In this way, over time, the speaker's recognition and response to user emotions will become more and more accurate. Multi-doll synchronization update: When new doll characters or new interactive content are launched, the system will update the relevant emotion data to ensure that all dolls can make appropriate responses based on the user's emotional state. S6: Cloud Server Architecture S6.1: Cloud service: provides integrated services such as user doll recognition, content storage and downloading, voice recognition, and dialogue engine. When the doll is recognized, the speaker requests the corresponding character content from the server, such as stories, voice files, and dialogue logic; S6.2: Data management and update: The server supports regular updates of the doll's story content, dialogue data, and character voice. Users can view and download updated content through the application (or the speaker itself), and provide a doll management background. Users can select, add, and delete dolls through the application, and personalize the story content as needed. S7: Security and Privacy S7.1: Data encryption: All data transmission between the speaker and the server (including doll identification data, voice commands, etc.) should be protected by encryption technology; S7.2: Privacy protection: The user's voice data should be processed locally as much as possible to avoid unnecessary data transmission and protect the user's privacy; S7.3: Child Protection: If the device is intended for children, provide parental control options to restrict inappropriate content and functionality.

2. The smart speaker system with doll recognition function according to claim 1, characterized in that: The RFID / NFC module in step S1 is used to identify the doll by reading the built-in RFID tag, confirming the identity of the doll. Visual recognition camera: If image recognition technology is used, the shape, color or appearance characteristics of the doll can be analyzed by AI algorithm for identification. Use deep learning algorithms (such as convolutional neural networks (CNN)) to enhance the recognition ability of the doll. Sensor type selection: If RFID technology is used, a unique RFID tag needs to be embedded in each doll. If visual recognition is used, the camera resolution must be at least 720P to ensure recognition accuracy, high-quality speakers (such as Dolby or Hi-Res Audio certified speakers) must be configured to ensure sound quality, a microphone array must be configured, and multi-microphone arrays must be supported to achieve high-precision voice recognition and noise cancellation. LED lights are used to display status and interactive feedback.

3. The intelligent speaker system with doll recognition function according to claim 1, characterized in that: RFID / NFC module in step S1: select a reliable RFID / NFC module (such as NXP PN532), and select an appropriate reading distance according to the size of the doll and the usage scenario. The RFID tag can be embedded in the doll, and the card reader on the top of the speaker can read the tag through near-field communication to obtain a unique identifier. Visual recognition module: If image recognition technology is used, the speaker needs to be equipped with a camera (such as OmniVision OV5640), and the captured image is input into an edge computing device (such as Jetson Nano) for processing. The doll is recognized in real time through a deep learning algorithm (for example, using TensorFlow Lite). Microphone array: configure up to 6-8 microphones to form a microphone array, which can achieve high-precision voice positioning and noise suppression. The microphone array should support sound source localization technology to ensure that the system can accurately recognize voice input from different directions.

4. The intelligent speaker system with doll recognition function according to claim 1, characterized in that: The character database design in step S2: the character database is stored in the cloud or locally, and each doll corresponds to a record. Each record contains: the unique identifier of the doll (such as RFID tag number or image feature), the name, gender, age, personality characteristics, etc. of the doll, the story, audio file, voice file, dialogue script, etc. associated with the doll, the voice characteristics of each doll (such as intonation, timbre, speaking speed, etc.) and personalized content, and the metadata of the doll is stored using a relational database (such as MySQL) or a NoSQL database (such as MongoDB). The information of each doll should include queryable fields, such as name, category, personality, etc., so that the system can quickly locate it when querying.

5. The intelligent speaker system with doll recognition function according to claim 1, characterized in that: In step S2, the user places the doll on the top of the speaker, the RFID / NFC module reads the doll tag and transmits its ID to the speaker host, and the speaker host queries the locally cached character database (or requests the cloud database through an API) to obtain virtual character information related to the doll. If visual recognition is used, the speaker inputs the image captured by the camera into the AI ​​algorithm, identifies the appearance features of the doll (such as color, shape, pattern, etc.), and obtains relevant virtual character information from the database.

6. The intelligent speaker system with doll recognition function according to claim 1, characterized in that: In step S3, the cloud server: all audio files (stories, character voices) and interactive data are stored on the cloud server to ensure that users can access the latest content at any time. Cloud storage services (such as Amazon S3) can be used to store a large number of audio files. Multi-format support: the story content should support multiple formats, such as MP3 or WAV audio files. The conversation content can be stored in JSON or XML format for subsequent parsing and playback. Content request process: When the doll is recognized, the speaker requests the cloud server through the network to obtain the story and conversation content related to the doll. This process includes: the doll recognizes and transmits the identifier, the speaker sends a request to the server, the server returns the corresponding audio and conversation script, the speaker plays the content, and displays related interactive information (such as the touch screen displays the current story or character information).

7. The intelligent speaker system with doll recognition function according to claim 1, characterized in that: The speech recognition technology selected in step S4: adopt mature speech recognition technology, such as Google Assistant SDK, Amazon Alexa Voice Service or iFlytek speech recognition, support voice command recognition and natural language processing, real-time speech recognition: the speaker receives and recognizes the user's voice command (such as "tell a story" or "talk to the doll") in real time, and triggers corresponding actions according to the command, dialogue engine: each doll uses an independent dialogue system (such as a chatbot engine based on GPT or Dialogflow) to generate a response. The dialogue system generates natural dialogue content based on the user's questions and the doll's background story. Personalized tone and accent: Each doll's speech synthesis system (such as WaveNet or Tacotron 2) adjusts the speed, pitch and tone of each doll according to the doll's character design to present different personalities. Emotional analysis module: The speaker recognizes the user's emotions through speech emotion analysis (using technologies such as IBM Watson Tone Analyzer) and adjusts the doll's response. For example, when the user's tone is excited, the doll may comfort the user; when the tone is excited, the doll will respond more actively.

8. The intelligent speaker system with doll recognition function according to claim 1, characterized in that: In step S6, cloud platform selection: select a reliable cloud platform (such as AWS, Google Cloud or Microsoft Azure) to provide services such as doll recognition, voice recognition, content storage, and dialogue generation. All doll recognition requests, content requests, etc. are transmitted through RESTful APIs. Server-side support: The server should support load balancing and be able to maintain high-performance response under multiple concurrent requests.

9. The intelligent speaker system with doll recognition function according to claim 1, characterized in that: Regular updates in step S6: the cloud server regularly pushes new content updates (such as new stories, new dialogues, character voices, etc.), and the speaker automatically downloads via Wi-Fi. Doll management background: a background management system is provided, and users can view and download updated content through mobile applications, and can also customize personalized doll stories and dialogues.

Citation Information

Cited By

  • Intention understanding driven large model AI multi-role cooperative interaction method and device

    CN120704636A

  • Large model AI interaction system based on master-slave system

    CN120723198A

  • Large model ai interaction system based on master-slave system

    CN120723198B

  • An active type of interactive system and method for a robot dog

    CN122618998A