AI interaction system capable of actively sensing and conversation, AI toy and interaction method

By combining multi-dimensional sensor components and AI models, AI toys can actively perceive and engage in dialogue, solving the problem of user-driven interaction in existing technologies and enhancing the immersive experience and emotional connection of the user experience.

CN120994058APending Publication Date: 2025-11-21SHANGHAI KUQIQI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511066633.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing AI toys cannot proactively perceive users' non-instructional intentions or changes in the environment, limiting the richness and naturalness of the interaction and resulting in a lack of a lifelike and warm user experience.

Method used

It employs multi-dimensional sensor components to detect user physical operations and environmental information in real time. Combined with information processing and behavior decision-making modules, it uses AI models to identify complex physical interaction events, make autonomous decisions, and generate response behaviors or dialogue content to achieve proactive perception and dialogue.

Benefits of technology

It enhances the immersive, emotionally connected, and fun experience of the user, allowing toys to integrate more naturally and proactively into user interaction, significantly increasing the sense of life and companionship.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994058A_ABST
    Figure CN120994058A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent interaction equipment, and provides an AI interaction system capable of active perception and conversation, an AI toy and an interaction method.The AI interaction system comprises a physical sensor module at a hardware end and an information processing and behavior decision module and an interaction execution module at a cloud end, and the physical sensor module comprises at least one or more sensors with different functions; the system is used for detecting physical operation of a user on the system or physical states of the user around the system in real time. Physical behaviors and environment information of a user can be actively sensed through the integrated multi-dimensional sensor assembly, and dialogue and interaction behaviors with the user can be intelligently and actively triggered or adjusted based on the sensed information, so that the system does not only passively wait for a user instruction, but can be used as a partner with sensing ability, and the interaction with the user can be intelligently and actively triggered or adjusted. And the interaction with the user is more naturally and actively fused, so that the immersion, emotion connection and interestingness of the user experience are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent interactive devices, and particularly relates to an AI interactive system and an interactive method capable of active sensing and conversation. BACKGROUND

[0002] At present, the interactive process of an existing AI toy is to give information by a user, for example, setting a button and a touch sensing area, the user actively pressing the button, patting or touching the touch sensing area to trigger feedback, or the user needs to speak a specific wake-up word to get the feedback of the toy. Therefore, most AI toys can only respond after receiving the explicit instructions (specific voice, button) of the user, and cannot actively sense the non-instructional intention of the user or environmental changes and initiate interaction. The user is always the initiator and initiator of interaction, which limits the richness and naturalness of interaction, and the user needs to deliberately "use" the toy, rather than interacting with a "friend" that seems to have autonomous consciousness, resulting in a lack of life and warmth of the toy and insufficient emotional connection with the user. SUMMARY

[0003] The purpose of the application is to overcome the existing defects and provide an AI interactive system, an AI toy and an interactive method capable of active sensing and conversation.

[0004] In order to solve the above technical problems, the application provides the following technical solutions:

[0005] The first purpose of the application is to provide an AI interactive system capable of active sensing and conversation, comprising:

[0006] A physical sensor module comprising at least one or more sensors with different functions for detecting the physical operation of the user on the system or the physical state of the user around the system in real time;

[0007] An information processing and behavior decision module connected with the physical sensor module for receiving and processing the signals collected by the physical sensor module, identifying a predefined physical interaction event, and autonomously deciding and generating a response behavior or conversation content based on the identified physical interaction event in combination with context information and internal state data;

[0008] An interactive execution module connected with the information processing and behavior decision module for executing the response behavior or conversation content.

[0009] Further, the physical sensor module adopts one or more combinations of proximity sensors, touch sensors, motion / posture sensors and light sensors.

[0010] Further, the response behavior comprises one or more of voice output, action execution and sound and light effects.

[0011] Further, the information processing and behavior decision module is further configured to:

[0012] The combined data is processed by a multi-sensor data fusion algorithm, and the combined data is converted into structured text information understandable by the AI model in combination with the processed current behavior, historical interaction data, and environmental information;

[0013] The AI model identifies complex physical interaction events and infers user intent or emotional state according to the structured text information;

[0014] In combination with the currently perceived user behavior, environmental information, and internal state, the response strategy and active interaction willingness are dynamically adjusted.

[0015] Further, the AI model uses one or more of a rule engine, a traditional machine learning model, a deep learning model, and a reinforcement learning model.

[0016] Further, the response strategy is generated according to at least one of the type and intensity of the physical interaction event, the internal state, the historical interaction data, and the environmental information.

[0017] Another object of the present application is to provide an AI toy that can actively perceive and dialogue, which is equipped with the AI interaction system that can actively perceive and dialogue provided by the first object of the present application.

[0018] Another object of the present application is to provide an AI interaction method that can actively perceive and dialogue, comprising:

[0019] The physical operation or physical state of the user is continuously monitored by at least one or more sensors;

[0020] When the monitored sensor signal meets the preset physical interaction event trigger condition, the physical interaction event is identified;

[0021] Based on the identified physical interaction event, in the case where the user does not issue an explicit voice instruction or a key operation, a corresponding response strategy is autonomously decided and controlled to be executed or a dialogue is actively initiated;

[0022] The content of the response behavior or the actively initiated dialogue is dynamically adjusted according to the type and / or intensity of the identified physical interaction event, and optionally the current internal state of the toy.

[0023] Further, the physical interaction event trigger condition includes one or more of a distance change to the user reaching a threshold, a touch pattern of the user to a specific area meeting a specific condition, and a specific change in motion state.

[0024] Further, the response strategy corresponding to the physical interaction event is learned and optimized according to long-term interaction history data of the user, for personalized interaction.

[0025] In combination with the above technical solution, the present application has the following beneficial effects compared with the prior art:

[0026] The present application can actively perceive the physical behavior (such as approaching, touching, moving, etc.) and environmental information of the user through the integrated multi-dimensional sensor assembly, and intelligently and actively trigger or adjust the dialogue and interaction behavior with the user based on these perception information, so that the toy is no longer passively waiting for user instructions, but can be more natural and more active to integrate into the interaction with the user like a perceptive partner, thereby significantly improving the immersion, emotional connection and interest of the user experience.

[0027] Compared with the conventional scheme which usually only makes decision processing for a single sensor and has a fixed execution strategy, the present application can make decision response in combination with the data combination of the uploaded multiple types of sensors, and can be described in text form, so that a large model can be analyzed and understood, and an output result can be given.

[0028] The present application can also realize the following functions:

[0029] 1. Event-driven interaction based on physical perception: not only can respond to voice instructions, but also can be triggered by specific events of physical sensors (such as user entering a specific distance, specific mode of touch), thereby actively initiating dialogue or behavior;

[0030] 2. Context-aware and adaptive interaction: can combine the currently perceived user behavior, environmental information and internal state (such as simulated emotional state, dialogue history) to dynamically adjust its response strategy and active interaction willingness.

[0031] 3. From "tool" to "partner" transformation: through enhancing active perception and interaction ability, aiming to improve the "life sense" and "companion sense" of the toy, so that the user experience is closer to the natural communication with real organisms. BRIEF DESCRIPTION OF DRAWINGS

[0032] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, which together with the embodiments of the present application, serve to explain the present application, and do not constitute a limitation on the present application. In the drawings:

[0033] Figure 1 is a structural schematic diagram of an AI interaction system that can actively perceive and dialogue provided by an embodiment of the present application;

[0034] Figure 2 is a flowchart of an interaction method provided by an embodiment of the present application. DETAILED DESCRIPTION

[0035] The preferred embodiments of the present application will be described below in conjunction with the accompanying drawings, it should be understood that the preferred embodiments described herein are only used to explain and illustrate the present application, and are not used to limit the present application.

[0036] Embodiment 1:

[0037] As shown in FIG. 1, which is an embodiment of the AI interaction system provided by the present application, comprising: Figure 1 A physical sensor module, containing at least one or more sensors with different functions, for detecting the physical operation of the user on the system or the physical state of the user around the system in real time;

[0038] An information processing and behavior decision module, connected with the physical sensor module, for receiving and processing the signals collected by the physical sensor module, identifying the predefined physical interaction events, and based on the identified physical interaction events, combining the context information and internal state data, autonomously deciding and generating the response behavior or dialogue content;

[0039] An interaction execution module, connected with the information processing and behavior decision module, for executing the response behavior or dialogue content.

[0040] Specifically, the physical sensor module is configured on the hardware side, and the information processing and behavior decision module and the interaction execution module are configured on the cloud side, through the hardware side to transmit the collected various events to the cloud side, the cloud side makes a comprehensive decision to generate a response.

[0041] The physical state of the user around the system represents the relative positional relationship between the user and the system, such as distance, approaching or moving away, etc.

[0042] The information processing and behavior decision module collects the raw data from various sensors, pre-processes them, combines the current behavior, historical interaction data, environmental information (such as time, light) and other processed data, and converts them into structured information that can be understood by AI, i.e. using textual event content as input, allowing the AI model to actively initiate voice response, the AI model will try to understand the user's possible intent or emotional state, make behavior decisions and output text and voice content. The interaction execution module converts the text response generated by AI into natural and fluent voice data through speech synthesis TTS, and plays it through a loudspeaker.

[0043]

[0044] ​Extended sources of context information: external information access, with user authorization, access to calendar information (such as birthday reminders), weather information, user's preferred music / story list, etc., to enrich the decision basis. Multi-user identification and personalization, if the system can interact with multiple family members, it can combine voiceprint recognition or (if there is a camera) face recognition to provide personalized interaction history and response patterns for different users.

[0045] Behavior decision logic can adopt:

[0046] Behavior tree: a modular and scalable way to design complex AI behavior logic, easy to debug and combine.

[0047] Finite state machine: used to define the behavior and state transition conditions of the toy in different states, suitable for structured interaction flow.

[0048] Planning algorithm: the AI system plans a series of actions based on the current state and goals (such as improving user happiness) to achieve the goal.

[0049] Utility-based AI: calculate a "utility value" for each possible behavior based on the current context, and select the behavior with the highest utility to execute.

[0050] The voice output and audio playback of the interaction execution module can have the following functions:

[0051] High expressiveness / emotional TTS: use synthesized speech with different emotional colors (such as happy, comforting, surprised) and tones to make the output more lively.

[0052] Personalized TTS: allow users to choose their favorite tone, or mimic some of the user's speech characteristics in long-term interaction (with explicit user authorization).

[0053] Non-verbal sound feedback: in addition to TTS, a richer preset sound effect library (such as animal sounds, natural sounds, musical instrument sounds) or programmatically generated sounds (such as electronic sound effects based on the current "emotion", humming melodies) can be used to express state or emotion.

[0054] Directional sound emission technology: use small speaker arrays and beamforming technology to make the sound seem to come from a specific part of the toy (such as the mouth), or enhance the volume in a specific direction.

[0055] As preferred, the physical sensor module in the embodiments of the present application adopts one or more combinations of proximity sensors, touch sensors, motion / pose sensors, and light sensors.

[0056] Proximity sensor: Detects the distance between the user and the toy, sensing the user's approach or departure. For example, when the user enters a set range, the toy can proactively greet them. Distance detection methods can include Bluetooth signal strength, Wi-Fi signal strength (RSSI), ultra-wideband (UWB), infrared proximity sensors, ultrasonic sensors, time-of-flight sensors, capacitive proximity sensing, visual proximity detection, and RFID / NFC.

[0057] Touch sensors: Touch sensors, pressure sensors, and other similar sensors are distributed in key areas of the toy's body (such as the head, back, and hands). They are used to sense different types of touch actions from the user, such as stroking, hugging, and patting, and the intensity of those touches. For example, gently stroking the head may trigger a satisfied sound from the toy, while patting may trigger different reactions.

[0058] Motion / attitude sensors: Accelerometers and gyroscopes are used to sense whether the toy itself is picked up, shaken, rotated, or dropped. For example, being picked up can express happiness, and being shaken might indicate "a little dizzy."

[0059] Light sensor: It senses the intensity of ambient light to help toys adjust their behavior patterns. For example, it may be more active during the day and give a goodnight greeting when the lights are turned off at night, or it may say "It's so dark".

[0060] Preferably, the response behavior in the embodiments of the present invention includes one or more of voice output, action execution, and audio-visual effects.

[0061] It should be noted that the embodiments of the present invention also include functional modules that are present in traditional AI interactive devices, such as a voice playback module and a light module. The voice playback module is used to actively greet the user and play voice messages, and the light module is used to provide status prompts using light effects, such as emitting specific light effects when the AI ​​system is speaking.

[0062] Preferably, the information processing and behavior decision-making module in this embodiment of the invention is further configured as follows:

[0063] By processing combined data through multi-sensor data fusion algorithms, and combining the current behavior, historical interaction data, and environmental information obtained from the processing, the combined data is transformed into structured text information that AI models can understand.

[0064] AI models identify complex physical interaction events based on structured text information and infer user intent or emotional state.

[0065] By combining the currently perceived user behavior, environmental information, and internal state, the response strategy and willingness to proactively interact are dynamically adjusted.

[0066] As preferred, the AI model in the embodiments of the present application adopts one or more of a rule engine, a traditional machine learning model, a deep learning model, and a reinforcement learning model.

[0067] Specifically, the enhanced rule engine is an expert system combining fuzzy logic or learnable rules.

[0068] The traditional machine learning model is, for example, a support vector machine (SVM), a decision tree, a random forest, etc., and is used for classifying user behavior or emotional state.

[0069] The deep learning model is:

[0070] The recurrent neural network (RNN) / long short-term memory network (LSTM) / gated recurrent unit (GRU) is suitable for processing sequence information such as sensor time series data and dialogue history.

[0071] The convolutional neural network (CNN) is used for image feature extraction and expression recognition if there is visual input, and is also used for processing time-frequency spectrograms of certain sensor data.

[0072] The transformer model is used for deeper natural language understanding, context awareness, and generation of more coherent dialogues.

[0073] The graph neural network (GNN) is used for reasoning if a relationship graph is constructed among user behavior, environmental factors, and toy state.

[0074] Reinforcement learning: the toy autonomously learns and optimizes its behavior decision-making strategy through interaction with the user and environmental feedback (such as whether the user continues to interact or shows positive emotions as a reward signal).

[0075] As preferred, the response strategy in the embodiments of the present application is generated according to at least one of a physical interaction event type, intensity, internal state, historical interaction data, and environmental information.

[0076] Embodiment 2:

[0077] The embodiments of the present application provide an AI toy that can actively perceive and dialogue, which is equipped with the AI interaction system that can actively perceive and dialogue provided in Embodiment 1 of the present application.

[0078] The embodiments of the present application are intelligent toys based on artificial intelligence and Internet of Things technology, and specifically, when a user interacts with the AI toy, the AI toy can actively perceive the existence and behavior of the user and optimize on this basis, so that the user and the AI toy interact more naturally, smoothly, and emotionally. The AI toy is transformed from a passive instruction receiver into an "intelligent partner" that can actively perceive users, understand situations, and actively initiate and adjust interactions.

[0079] Workflow example:

[0080] When the user approaches and touches (touch recognition) the AI toy while crying (voice recognition) at night (environment recognition), the physical sensor module combines the previous historical chat data, which is processed by the information processing and decision module to obtain the appropriate output. At this time, the AI toy can better guide and comfort through the interaction execution module, rather than accepting specific prompt words and then using general rhetoric broadly.

[0081] Embodiment 3:

[0082] As shown in Figure 2 The present application provides an AI interaction method that can actively perceive and dialogue, which specifically includes the following steps:

[0083] S1: Continuously monitor the user's physical operation or physical state through at least one or more sensors;

[0084] S2: When the monitored sensor signal meets the preset physical interaction event trigger condition, identify the physical interaction event;

[0085] S3: Based on the identified physical interaction event, in the case where the user does not issue explicit voice instructions or key operations, autonomously decide and control the execution of the corresponding response strategy or actively initiate a dialogue;

[0086] S4: The content of the response behavior or the actively initiated dialogue is dynamically adjusted according to the type and / or intensity of the identified physical interaction event, and optionally the current internal state of the toy.

[0087] As a preferred embodiment, the physical interaction event trigger condition in the present application includes one or more of the following: the distance to the user changes to a threshold, the user's touch pattern on a specific area meets a specific condition, and the motion state changes in a specific way.

[0088] As a preferred embodiment, the response strategy corresponding to the physical interaction event in the present application is learned and optimized according to the long-term interaction history data with the user, for personalized interaction.

[0089] It should be understood that, although the steps in the flowcharts of the embodiments of the present application are shown in a certain order according to the arrows, the steps are not necessarily executed in the order of the arrows. Unless otherwise specified in the present application, the execution of the steps is not strictly limited in order, and the steps can be executed in other orders. Moreover, at least some of the steps in the embodiments can include a plurality of sub-steps or a plurality of stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of the sub-steps or stages is not necessarily sequential, but can be round-robin or alternately executed with at least some of the other steps or sub-steps or stages of the other steps.

[0090] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments of the methods. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0091] Finally, it should be noted that: the above only describes the preferred embodiments of the present application and is not intended to limit the present application. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent replacements to some technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An actively aware and conversational AI interaction system, characterized by, The system comprises: a physical sensor module comprising at least one or more sensors of different functions for detecting in real time the physical operation of the user on the system or the physical state of the user around the system; an information processing and behavior decision module connected with the physical sensor module for receiving and processing the signals collected by the physical sensor module, identifying a predefined physical interaction event, and based on the identified physical interaction event, combining context information and internal state data to autonomously decide and generate a response behavior or dialogue content; an interaction execution module connected with the information processing and behavior decision module for executing the response behavior or dialogue content. 2.The actively aware and conversational AI interaction system of claim 1, wherein, The physical sensor module uses one or more combinations of proximity sensors, touch sensors, motion / posture sensors, and light sensors. 3.The actively aware and conversational AI interaction system of claim 1, wherein, The response behavior includes one or more of voice output, action execution, and sound and light effects. 4.The actively aware and conversational AI interaction system of claim 1, wherein, The information processing and behavior decision module is further configured to: process the combined data through a multi-sensor data fusion algorithm, combine the processed current behavior, historical interaction data, and environmental information, and convert the combined data into structured text information understandable by an AI model; the AI model identifies complex physical interaction events and infers user intent or emotional state based on the structured text information; combine the currently perceived user behavior, environmental information, and internal state to dynamically adjust the response strategy and active interaction intention. 5.The actively aware and conversational AI interaction system of claim 4, wherein, The AI model uses one or more of a rule engine, a traditional machine learning model, a deep learning model, and a reinforcement learning model. 6.The actively aware and conversational AI interaction system of claim 4, wherein, The response strategy is generated based on at least one of the type and intensity of the physical interaction event, the internal state, the historical interaction data, and the environmental information.

7. An AI toy capable of active sensing and dialogue, characterized by, The AI toy is equipped with the AI interaction system capable of active sensing and dialogue according to any one of claims 1 to 6.

8. An AI interaction method capable of active perception and dialogue, characterized in that, The interaction method comprises: continuously monitoring the physical operation or physical state of the user through at least one or more sensors; when the monitored sensor signals meet the preset physical interaction event trigger condition, identifying the physical interaction event; based on the identified physical interaction event, autonomously deciding and controlling the execution of the corresponding response strategy or actively initiating a dialogue without explicit voice instructions or button operations from the user; the content of the response behavior or the actively initiated dialogue is dynamically adjusted according to the type and / or intensity of the identified physical interaction event and the optional current internal state of the toy.

9. The interaction method of claim 8, wherein, The physical interaction event trigger condition includes one or more of a change in distance to the user reaching a threshold, a touch pattern of the user on a specific area meeting a specific condition, and a specific change in motion state.

10. The interaction method of claim 8, wherein, The response strategy corresponding to the physical interaction event is learned and optimized based on the long-term interaction history with the user for personalized interaction.