Multi-modal emotion recognition toy robot interaction system and method

By using a multimodal emotion recognition system to perform temporal alignment and quantitative analysis of multimodal data in the toy robot interaction system, synchronous or alternating activation guidance sequences are generated. This solves the shortcomings of single-modal guidance in existing technologies, enables accurate evaluation of toy robot interaction and continuous attention maintenance, and improves the interaction effect.

CN121868880APending Publication Date: 2026-04-17DONGGUAN YONGNKIDS TOYS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DONGGUAN YONGNKIDS TOYS TECHNOLOGY CO LTD
Filing Date
2025-11-21
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing toy robot interaction systems rely on single-modal guidance, making it difficult to accurately assess the response status of the target user. The lack of multi-robot collaboration mechanisms results in guidance strategies that are not targeted or adaptable, and fail to maintain the user's sustained attention.

Method used

A multimodal emotion recognition system is adopted to acquire visual and audio feedback data from the target interactor and combine it with the guidance visual and audio data of the toy robot to achieve temporal alignment and quantitative analysis of multimodal data, generate synchronous or alternating activation guidance sequences, and dynamically adjust guidance strategies to maintain attention.

Benefits of technology

It achieves comprehensive perception and precise analysis of the interaction process, provides objective evaluation of interaction effects, ensures the pertinence and adaptability of guidance strategies, continuously maintains the attention level of the target interactor, and enriches the interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121868880A_ABST
    Figure CN121868880A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode emotion recognition toy robot interaction system and method. The system comprises a plurality of toy robots arranged in an interaction space, and further comprises a positioning module for acquiring positioning data of the toy robots in the interaction space; the identification module is arranged on each toy robot and is used for acquiring multi-modal interaction data of a target interactor in the interaction space; the data processing module obtains the multi-modal interaction data and the multi-modal instruction data from the toy robot, and the multi-modal interaction data and the multi-modal instruction data are subjected to time sequence alignment to obtain time sequence interaction data; the control module receives the time sequence interaction data and generates a control instruction to each toy robot; and the response module is arranged on each toy robot, and executes the appointed response of the toy robot in the interaction space according to the control instruction until the target interactor reaches the preset interaction area. According to the invention, multi-modal emotion quantitative identification and dynamic response calibration in the toy interaction process are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control technology, specifically to a toy robot interaction system and method with multimodal emotion recognition. Background Technology

[0002] Current toy robot interaction technologies primarily rely on single-modal guidance methods, typically involving simple interactions through visual or audio signals, lacking the ability to comprehensively process multimodal data. In practical applications, such systems struggle to accurately assess the target user's response status and cannot quantify effectiveness metrics during the interaction process, resulting in guidance strategies lacking specificity and adaptability. Furthermore, existing technologies often employ fixed guidance sequences, failing to dynamically adjust guidance methods based on real-time feedback from the target user, making it difficult to maintain sustained user attention, especially in scenarios requiring prolonged focus, such as children's education and rehabilitation training. The lack of collaborative mechanisms among multiple toy robots within the interaction space also prevents the formation of an effective multi-robot guidance network, limiting the richness of the interactive experience and the precision of the guidance effect.

[0003] Therefore, there is an urgent need for a multimodal emotion recognition interactive system and method for toy robots. Summary of the Invention

[0004] To address the above problems, this invention provides a toy robot interaction system and method for multimodal emotion recognition.

[0005] A first aspect of the present invention provides a multimodal emotion recognition toy robot interaction system, comprising multiple toy robots disposed within an interaction space, and further comprising: The positioning module is configured to acquire positioning data of multiple toy robots within the interactive space; A recognition module, located in each of the toy robots, is configured to acquire multimodal interaction data of the target interactor within the interaction space; The data processing module is configured to acquire the multimodal interaction data and the multimodal instruction data from the toy robot, and to time-align the multimodal interaction data and the multimodal instruction data to obtain time-series interaction data; The control module is configured to receive the timing interaction data and generate control commands to each of the toy robots; A response module, located in each of the toy robots, is configured to execute a specified response by the toy robot within the interaction space according to the control command, directing the target interactor to a preset interaction area.

[0006] As a preferred embodiment, the multimodal interaction data includes feedback visual data and feedback audio data of the target interactor, and the multimodal instruction data includes guidance visual data and guidance audio data of the toy robot; When the data processing module is configured to perform timing alignment, it includes the following steps: The guiding visual data, the guiding audio data, the feedback visual data, and the feedback audio data are synchronized using timestamps. Generate time-series interaction data, where each unit is an instruction event. Each instruction event includes the feedback time corresponding to the feedback visual data of the target interactor, the feedback intensity corresponding to the feedback audio data, and an interaction effectiveness index calculated based on the feedback time and feedback intensity after the occurrence of the guidance visual data and guidance audio data.

[0007] As a preferred embodiment, the guiding visual data includes guiding trajectory data, and the guiding trajectory data has a quantified first correlation value with the emotional feature data and motion posture data extracted from the feedback visual data of the target interactor; The guiding audio data includes spatial audio guiding parameters and audio guiding trajectory lines, and the spatial audio guiding parameters have a quantized second correlation quantization value with the feedback audio data of the target interactor; Both the first correlation quantization value and the second correlation quantization value represent their respective guidance strength and mapping relationship. The first correlation quantization value reflects the guiding effect of visual guidance on the attention level of the target interactor, and the second correlation quantization value reflects the guiding effect of audio guidance on the attention level of the target interactor.

[0008] As a preferred method, the spatial audio guidance parameters are generated by continuously defining the audio guidance trajectory line by the maximum audio intensity feature points obtained by each toy robot in the interactive space at the specified sound source intensity position, which varies with time. Furthermore, while maintaining its own position, the toy robot generates spatial audio guidance parameters with phase differences by using toy robots in different positions, or the spatial audio guidance parameters are obtained by adjusting the position of the toy robot.

[0009] In a preferred manner, the guiding visual data and the guiding audio data are coordinated to generate a guiding sequence, the guiding sequence including a synchronous activation sequence and an alternating activation sequence; In the synchronous activation sequence, the guiding visual data and the guiding audio data are aligned in time and overlap in space to form a multimodal guiding focus. When the product of the first association quantization value and the second association quantization value is higher than the synchronous activation threshold, the control module is configured to start the synchronous activation sequence to enhance the guidance intensity for the target interactor. The guiding visual data and the guiding audio data are temporally aligned but spatially independent; The control module is configured as follows: Based on the dynamic trend of the first associated quantization value, when the first associated quantization value is lower than the first associated quantization decrease threshold and continues to decrease, the activation interval of the guiding audio data is shortened so as to reset the attention level of the target interactor through audio guidance. Based on the dynamic trend of the second correlation quantization value, when the second correlation quantization value is lower than the second correlation quantization decrease threshold and continues to decrease, the trajectory point density of the guiding visual data is increased to enhance the action response of the target interactor through visual guidance. Based on the feedback time change rate in the time-series interaction data, the activation sequence of the guiding visual data and the guiding audio data is dynamically adjusted so that the alternation interval between visual guidance and audio guidance matches the attention recovery cycle of the target interactor. By using an alternately activated guiding sequence, the average value of the first correlated quantization value and the second correlated quantization value is increased while the fluctuation amplitude is reduced, wherein: When the feedback intensity of the visual feedback data remains below the visual intensity maintenance threshold, the control module is configured to extend the activation duration of the guiding audio data in order to utilize the indirect guiding effect of audio guidance on the action response. When the feedback time of the feedback audio data continues to exceed the time recovery threshold, the control module is configured to expand the trajectory coverage of the guiding visual data in order to utilize the indirect guiding effect of visual guidance on attention level. Continuously maintain the target user's attention level and optimize their action response intensity to guide the target user to the preset interaction area.

[0010] As a preferred embodiment, the audio guide trajectory line is defined by a first quantization parameter and a second quantization parameter, wherein: The first quantization parameter represents the minimum spatial position interval of the audio guide trajectory line; The second quantization parameter represents the minimum update interval of the audio guide trajectory line in time; The data processing module is configured as follows: When the feedback time exceeds the time recovery threshold, the second quantization parameter is reduced to improve the temporal resolution, causing the audio guidance trajectory line to switch in the alternating activation guidance sequence to maintain the target interactor's attention level; when the feedback intensity of the feedback audio data is lower than the audio intensity maintenance threshold, the first quantization parameter is reduced to improve the spatial resolution and enhance the guidance effect of the audio guidance on the target interactor.

[0011] As a preferred embodiment, the guidance trajectory data is limited by a third quantization parameter and a fourth quantization parameter, wherein: The third quantization parameter represents the minimum spatial distance between adjacent trajectory points in the guided trajectory data; The fourth quantization parameter represents the minimum time interval for updating trajectory points in the guided trajectory data; The control module is configured as follows: When the feedback intensity of the feedback visual data is lower than the visual intensity maintenance threshold, the third quantization parameter is reduced to increase the trajectory point density; when the feedback time exceeds the time recovery threshold, the fourth quantization parameter is reduced to increase the trajectory update rate. By dynamically adjusting the third and fourth quantization parameters, the guidance trajectory data works synergistically with the audio guidance trajectory line in the alternating activation guidance sequence, continuously improving the target user's attention level and action response intensity.

[0012] A second aspect of the present invention provides a toy robot interaction method for multimodal emotion recognition, comprising the following steps: S1. Acquire the positioning data of multiple toy robots in the interactive space, as well as the visual and audio feedback data of the target interactor; S2. Acquire the guidance visual data and guidance audio data of the toy robot, wherein the guidance visual data includes guidance trajectory data, and the guidance audio data includes spatial audio guidance parameters and audio guidance trajectory lines; S3. Use timestamps to synchronize and guide visual data, guide audio data, provide feedback visual data and feedback audio data, and generate time-series interaction data that includes feedback time, feedback intensity and interaction effectiveness indicators; S4. Based on the first and second correlation quantization values ​​in the time-series interaction data, generate a synchronous activation sequence or an alternating activation sequence, wherein: When the product of the first associated quantization value and the second associated quantization value is higher than the synchronization activation threshold, a time-aligned and spatially overlapping synchronization activation sequence is generated. Otherwise, generate time-aligned but spatially independent alternating activation sequences and dynamically adjust the activation timing based on feedback data; S5. Based on the guide sequence, generate control instructions to drive the toy robot to execute the specified response, continuously maintain the attention level of the target interactor and guide it to the preset interaction area.

[0013] Compared with the prior art, the present invention has the following advantages: The multimodal recognition toy robot interaction system provided by this invention achieves comprehensive perception and precise analysis of the interaction process by acquiring feedback visual and audio data from the target interactor, as well as guidance visual and audio data from the toy robot. The system synchronizes multimodal data streams with timestamps, generating time-series interaction data containing feedback time, feedback intensity, and interaction effectiveness indicators. This provides an objective and quantitative evaluation basis for interaction effects, making the formulation of guidance strategies more data-driven.

[0014] By introducing a first correlation quantization value and a second correlation quantization value, the system can separately evaluate the guiding effect of visual guidance on the target interactor's action response and the guiding effect of audio guidance on attention level, achieving refined differentiation and independent evaluation of multimodal guidance effects. This quantification mechanism enables the system to accurately identify which guidance method is more effective, providing a key decision-making basis for the subsequent generation of guidance sequences.

[0015] The system-generated synchronous activation sequences and alternating activation sequences constitute a flexible guidance strategy system. When the product of the first and second association quantification values ​​exceeds the synchronous activation threshold, the system initiates a time-aligned and spatially overlapping synchronous activation sequence, forming a multimodal guidance focus and significantly enhancing the guidance intensity for the target user. In other cases, the system employs a time-aligned but spatially independent alternating activation sequence and dynamically adjusts the activation order based on feedback data, effectively avoiding sensory fatigue and achieving intelligent switching of guidance methods. This dual-sequence mechanism ensures that the system can adapt to different interaction states and continuously maintain the target user's attention level.

[0016] The audio guidance trajectory is defined by a first and a second quantization parameter, while the guidance trajectory data is defined by a third and a fourth quantization parameter, enabling the system to precisely control the spatiotemporal characteristics of the guidance. When the feedback time exceeds the time recovery threshold, the system decreases the second quantization parameter to improve temporal resolution and accelerate the switching speed of the audio guidance; when the feedback intensity is below the threshold, the system decreases the first quantization parameter to improve spatial resolution, making the guidance more refined. This dynamic adjustment mechanism of quantization parameters ensures that the guidance strategy can be optimized based on the real-time reactions of the target interactor, continuously improving the interaction effect.

[0017] By using alternating activation of the guidance sequence, the system increases the average value of the first and second correlation quantization values ​​while reducing their fluctuations, thus improving interaction stability. When the feedback intensity of the visual data remains below the visual intensity maintenance threshold, the system extends the activation duration of the guidance audio data, utilizing the indirect guidance effect of audio on action response. When the feedback time of the audio data continuously exceeds the time recovery threshold, the system expands the trajectory coverage of the guidance visual data, utilizing the indirect guidance effect of visual guidance on attention level. This cross-compensation mechanism allows the system to effectively compensate for the decline in the effectiveness of one guidance method with another, ensuring continuous optimization of the overall guidance effect.

[0018] The spatial audio guidance parameters are continuously defined by the maximum audio intensity feature points of each toy robot within the interactive space, constrained by a specified sound source intensity location, forming an audio guidance trajectory line. This allows the system to dynamically move the virtual sound source by controlling phase difference while keeping the toy robots in a fixed position. This technology achieves audio guidance without relying on physical movement, enriching the interaction methods while reducing the latency and malfunction risks associated with mechanical movement.

[0019] The system dynamically adjusts the activation sequence of guiding visual and audio data based on the feedback time change rate in the temporal interaction data, ensuring that the alternation interval between visual and audio guidance matches the target user's attention recovery cycle. This personalized timing adjustment mechanism makes the guidance rhythm more aligned with the target user's cognitive characteristics, improving the adaptability and effectiveness of the guidance.

[0020] The system can continuously maintain the target user's attention level and optimize their action response intensity, effectively guiding them to the preset interaction area and achieving a comprehensive improvement in interaction effectiveness. Furthermore, the system can continuously optimize parameter settings based on historical interaction data, ensuring that the guidance effect improves over time and provides users with an increasingly precise and effective interactive experience. Attached Figure Description

[0021] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the system provided in an embodiment of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] In a first aspect, this embodiment provides a multimodal emotion recognition toy robot interaction system, such as... Figure 1 As shown, it includes multiple toy robots set up in the interactive space, and also includes: The positioning module is configured to acquire positioning data of multiple toy robots within the interactive space; A recognition module, located in each of the toy robots, is configured to acquire multimodal interaction data of the target interactor within the interaction space; The data processing module is configured to acquire the multimodal interaction data and the multimodal instruction data from the toy robot, and to time-align the multimodal interaction data and the multimodal instruction data to obtain time-series interaction data; The control module is configured to receive the timing interaction data and generate control commands to each of the toy robots; A response module, located in each of the toy robots, is configured to execute a specified response by the toy robot within the interaction space according to the control command, directing the target interactor to a preset interaction area.

[0025] As a preferred embodiment, the multimodal interaction data includes feedback visual data and feedback audio data of the target interactor, and the multimodal instruction data includes guidance visual data and guidance audio data of the toy robot; When the data processing module is configured to perform timing alignment, it includes the following steps: The guiding visual data, the guiding audio data, the feedback visual data, and the feedback audio data are synchronized using timestamps. Generate time-series interaction data, where each unit is an instruction event. Each instruction event includes the feedback time corresponding to the feedback visual data of the target interactor, the feedback intensity corresponding to the feedback audio data, and an interaction effectiveness index calculated based on the feedback time and feedback intensity after the occurrence of the guidance visual data and guidance audio data.

[0026] As a preferred embodiment, the guiding visual data includes guiding trajectory data, and the guiding trajectory data has a quantified first correlation value with the emotional feature data and motion posture data extracted from the feedback visual data of the target interactor; The guiding audio data includes spatial audio guiding parameters and audio guiding trajectory lines, and the spatial audio guiding parameters have a quantized second correlation quantization value with the feedback audio data of the target interactor; Both the first correlation quantization value and the second correlation quantization value represent their respective guidance strength and mapping relationship. The first correlation quantization value reflects the guiding effect of visual guidance on the attention level of the target interactor, and the second correlation quantization value reflects the guiding effect of audio guidance on the attention level of the target interactor.

[0027] As a preferred method, the spatial audio guidance parameters are generated by continuously defining the audio guidance trajectory line by the maximum audio intensity feature points obtained by each toy robot in the interactive space at the specified sound source intensity position, which varies with time. Furthermore, while maintaining its own position, the toy robot generates spatial audio guidance parameters with phase differences by using toy robots in different positions, or the spatial audio guidance parameters are obtained by adjusting the position of the toy robot.

[0028] In a preferred manner, the guiding visual data and the guiding audio data are coordinated to generate a guiding sequence, the guiding sequence including a synchronous activation sequence and an alternating activation sequence; In the synchronous activation sequence, the guiding visual data and the guiding audio data are aligned in time and overlap in space to form a multimodal guiding focus. When the product of the first association quantization value and the second association quantization value is higher than the synchronous activation threshold, the control module is configured to start the synchronous activation sequence to enhance the guidance intensity for the target interactor. In the alternating activation sequence, the guiding visual data and the guiding audio data are temporally aligned but spatially independent; The control module is configured as follows: Based on the dynamic trend of the first associated quantization value, when the first associated quantization value is lower than the first associated quantization decrease threshold and continues to decrease, the activation interval of the guiding audio data is shortened so as to reset the attention level of the target interactor through audio guidance. Based on the dynamic trend of the second correlation quantization value, when the second correlation quantization value is lower than the second correlation quantization decrease threshold and continues to decrease, the trajectory point density of the guiding visual data is increased to enhance the action response of the target interactor through visual guidance. Based on the feedback time change rate in the time-series interaction data, the activation sequence of the guiding visual data and the guiding audio data is dynamically adjusted so that the alternation interval between visual guidance and audio guidance matches the attention recovery cycle of the target interactor. By using the alternating activation sequence, the average value of the first correlated quantization value and the second correlated quantization value is increased while the fluctuation amplitude is reduced, wherein: When the feedback intensity of the visual feedback data remains below the visual intensity maintenance threshold, the control module is configured to extend the activation duration of the guiding audio data in order to utilize the indirect guiding effect of audio guidance on the action response. When the feedback time of the feedback audio data continues to exceed the time recovery threshold, the control module is configured to expand the trajectory coverage of the guiding visual data in order to utilize the indirect guiding effect of visual guidance on attention level. Continuously maintain the target user's attention level and optimize their action response intensity to guide the target user to the preset interaction area.

[0029] As a preferred embodiment, the audio guide trajectory line is defined by a first quantization parameter and a second quantization parameter, wherein: The first quantization parameter represents the minimum spatial position interval of the audio guide trajectory line; The second quantization parameter represents the minimum update interval of the audio guide trajectory line in time; The data processing module is configured as follows: When the feedback time exceeds the time recovery threshold, the second quantization parameter is reduced to improve the temporal resolution, causing the audio guidance trajectory line to switch in the alternating activation sequence to maintain the target interactor's attention level; when the feedback intensity of the feedback audio data is lower than the audio intensity maintenance threshold, the first quantization parameter is reduced to improve the spatial resolution, thereby enhancing the audio guidance effect on the target interactor.

[0030] As a preferred embodiment, the guidance trajectory data is limited by a third quantization parameter and a fourth quantization parameter, wherein: The third quantization parameter represents the minimum spatial distance between adjacent trajectory points in the guided trajectory data; The fourth quantization parameter represents the minimum time interval for updating trajectory points in the guided trajectory data; The control module is configured as follows: When the feedback intensity of the feedback visual data is lower than the visual intensity maintenance threshold, the third quantization parameter is reduced to increase the trajectory point density; when the feedback time exceeds the time recovery threshold, the fourth quantization parameter is reduced to increase the trajectory update rate. By dynamically adjusting the third and fourth quantization parameters, the guidance trajectory data works synergistically with the audio guidance trajectory line in the alternating activation guidance sequence, continuously improving the target user's attention level and action response intensity.

[0031] A second aspect of this embodiment provides a toy robot interaction method with multimodal emotion recognition, comprising the following steps: S1. Acquire the positioning data of multiple toy robots in the interactive space, as well as the visual and audio feedback data of the target interactor; S2. Acquire the guidance visual data and guidance audio data of the toy robot, wherein the guidance visual data includes guidance trajectory data, and the guidance audio data includes spatial audio guidance parameters and audio guidance trajectory lines; S3. Use timestamps to synchronize and guide visual data, guide audio data, provide feedback visual data and feedback audio data, and generate time-series interaction data that includes feedback time, feedback intensity and interaction effectiveness indicators; S4. Based on the first and second correlation quantization values ​​in the time-series interaction data, generate a synchronous activation sequence or an alternating activation sequence, wherein: When the product of the first associated quantization value and the second associated quantization value is higher than the synchronization activation threshold, a time-aligned and spatially overlapping synchronization activation sequence is generated. Otherwise, generate time-aligned but spatially independent alternating activation sequences and dynamically adjust the activation timing based on feedback data; S5. Based on the guide sequence, generate control instructions to drive the toy robot to execute the specified response, continuously maintain the attention level of the target interactor and guide it to the preset interaction area.

[0032] Specifically, in this embodiment, the system includes multiple toy robots disposed within an interactive space, and further includes a positioning module, a recognition module, a data processing module, a control module, and a response module. In this embodiment, the positioning module is configured to acquire positioning data of the multiple toy robots within the interactive space. Specifically, it employs UWB positioning technology to track the coordinate positions of each toy robot in three-dimensional space in real time, and transmits the positioning data to the data processing module via a wireless communication network, ensuring that the system can accurately grasp the relative positional relationships of each toy robot. The recognition module is disposed on each toy robot and is configured to acquire multimodal interaction data of the target interactor within the interactive space. The recognition module includes a high-definition camera and a microphone array. The camera captures the target interactor's facial expressions, body movements, and spatial position information at a frame rate of 30fps, while the microphone array collects the target interactor's voice tone and ambient sounds in real time, forming a complete multimodal data stream. The data processing module is configured to acquire multimodal interaction data and multimodal command data from the toy robots, and to align the multimodal interaction data and multimodal command data in time to obtain time-series interaction data. In its implementation, the data processing module uses a high-precision timestamp synchronization mechanism to align data streams from different sensors with microsecond-level accuracy, ensuring the accuracy of subsequent analysis. The control module is configured to receive time-series interaction data and generate control commands to each toy robot. This module adopts a distributed computing architecture, capable of handling the control needs of multiple toy robots simultaneously, ensuring the real-time performance and coordination of the system response. The response module is located on each toy robot and is configured to execute a specified response within the interaction space according to the control commands, moving the target interactor to a preset interaction area. The response module includes a servo motor, a steerable speaker, and an RGB-LED array, enabling precise positional movement, sound output, and light effect display.

[0033] In this embodiment, the multimodal interaction data includes feedback visual data and feedback audio data of the target interactor, and the multimodal instruction data includes guidance visual data and guidance audio data of the toy robot. The feedback visual data records the target interactor's facial expression changes, body movement trajectories, and eye movements. The feedback audio data includes speech tone features, breathing frequency, and environmental sound response. The guidance visual data consists of light effect trajectories, projected images, and motion postures generated by the toy robot. The guidance audio data includes directional sounds, spatial audio effects, and voice commands. When the data processing module performs time-series alignment, it specifically includes synchronizing and guiding visual data, guiding audio data, feedback visual data, and feedback audio data with timestamps to generate time-series interaction data. Each unit is a command event. Each command event includes the feedback time corresponding to the target interactor's feedback visual data and the feedback intensity corresponding to the feedback audio data after the occurrence of the guiding visual data and the guiding audio data, as well as the interaction effectiveness index calculated based on the feedback time and feedback intensity. In practical applications, the system determines the reasonable range of feedback time to be 200ms-1500ms through statistical analysis. The feedback intensity is quantified by the amplitude of the action or the volume of the voice. The interaction effectiveness index is calculated using a weighted average algorithm, and the weights are dynamically adjusted according to the target interactor's age and interaction experience.

[0034] The guiding visual data includes guiding trajectory data, which has a quantified first correlation value with the emotional feature data and motion posture data extracted from the feedback visual data of the target interactor. The guiding audio data includes spatial audio guiding parameters and audio guiding trajectory lines, which have a quantified second correlation value with the feedback audio data of the target interactor. Both the first and second correlation values ​​represent their respective guiding strength and mapping relationship. The first correlation value reflects the guiding effect of visual guidance on the target interactor's attention level, while the second correlation value reflects the guiding effect of audio guidance on the target interactor's attention level. In specific implementation, the system uses a deep learning model to extract micro-expression feature vectors and motion keypoints from the feedback visual data and calculates the correlation coefficient with the guiding trajectory data as the first correlation value. Simultaneously, through sound source localization and speech emotion analysis technology, the matching degree between the feedback audio data and the spatial audio guiding parameters is determined as the second correlation value. Both quantification values ​​range from 0 to 1, with higher values ​​indicating better guiding effects.

[0035] When generating spatial audio guidance parameters, the maximum audio intensity feature points obtained by each toy robot within the interactive space at a specified sound source intensity location, which vary over time, continuously define the audio guidance trajectory line. Furthermore, while maintaining their own positions, the spatial audio guidance parameters are obtained by creating spatial audio guidance parameters with phase differences through the toy robots in different positions, or by adjusting the positions of the toy robots. In actual deployment, the system pre-establishes a three-dimensional sound field model within the interactive space. By calculating the relative positions of each toy robot and the speaker direction, the movement trajectory of the virtual sound source in space is determined. When multiple toy robots work collaboratively, by precisely controlling the phase difference and delay of each sound source, smooth movement of the virtual sound source in space is achieved. Even if the toy robots themselves remain stationary, a dynamic audio guidance effect can be created.

[0036] When guiding visual data and guiding audio data work together, a guiding sequence is generated. This sequence includes a synchronous activation sequence and an alternating activation sequence. In the synchronous activation sequence, the guiding visual data and guiding audio data are temporally aligned and spatially overlapping to form a multimodal guiding focus. When the product of the first correlation quantization value and the second correlation quantization value is higher than the synchronous activation threshold, the control module is configured to activate the synchronous activation sequence to enhance the guiding intensity for the target user. When the guiding visual data and guiding audio data are temporally aligned but spatially independent, the control module is configured to, based on the dynamic trend of the first correlation quantization value, shorten the activation interval of the guiding audio data when the first correlation quantization value falls below the first correlation quantization decline threshold and continues to decline, thereby resetting the target user's attention level through audio guidance. Based on the dynamic trend of the second correlation quantization value, when the second correlation quantization value falls below the second correlation quantization decline threshold and continues to decline, increase the trajectory points of the guiding visual data. The system optimizes the density of visual guidance to enhance the target user's action response. Based on the feedback time change rate in the temporal interaction data, it dynamically adjusts the activation sequence of guiding visual and audio data to match the alternation interval between visual and audio guidance with the target user's attention recovery cycle. Through alternating activation sequences, the average value of the first and second correlation quantification values ​​is increased while their fluctuation amplitude decreases. Specifically, when the feedback intensity of the visual data remains below the visual intensity maintenance threshold, the control module is configured to extend the activation duration of the guiding audio data to utilize the indirect guiding effect of audio guidance on action response. When the feedback time of the audio data remains above the time recovery threshold, the control module is configured to expand the trajectory coverage of the guiding visual data to utilize the indirect guiding effect of visual guidance on attention level. The system continuously maintains the target user's attention level and optimizes their action response intensity, guiding the target user to a preset interaction area. In practice, the system monitors the target user's reaction pattern in real time. When signs of attention loss are detected, the system automatically switches guidance strategies. For example, in children's educational scenarios, when a child loses interest in visual guidance, the system increases the frequency and intensity of audio guidance to re-attract attention through music and voice prompts.

[0037] The audio guidance trajectory is defined by a first quantization parameter and a second quantization parameter. The first quantization parameter represents the minimum spatial positional interval of the audio guidance trajectory, and the second quantization parameter represents the minimum temporal update interval of the audio guidance trajectory. The data processing module is configured to decrease the second quantization parameter to improve temporal resolution when the feedback time exceeds a time recovery threshold, allowing the audio guidance trajectory to switch in an alternating activation sequence to maintain the target user's attention level. When the feedback intensity of the feedback audio data is lower than the audio intensity maintenance threshold, the first quantization parameter is decreased to improve spatial resolution, enhancing the guidance effect of the audio guidance on the target user. In practical applications, the initial value of the first quantization parameter is set to 50 cm, representing the minimum step size for the virtual sound source to move in space, and the initial value of the second quantization parameter is set to 500 milliseconds, representing the minimum time interval for updating the virtual sound source position. The system dynamically adjusts these two parameters based on the target user's real-time response. For example, when the system detects a slow response from the target user to the guidance, it will decrease the first quantization parameter to 25 cm to make the audio guidance more precise, and simultaneously decrease the second quantization parameter to 250 milliseconds to accelerate the guidance rhythm.

[0038] The guidance trajectory data is defined by a third quantization parameter and a fourth quantization parameter. The third quantization parameter represents the minimum spatial distance between adjacent trajectory points in the guidance trajectory data, and the fourth quantization parameter represents the minimum time interval for trajectory point updates. The control module is configured to decrease the third quantization parameter to increase trajectory point density when the feedback intensity of the visual feedback data is lower than the visual intensity maintenance threshold; and to decrease the fourth quantization parameter to increase the trajectory update rate when the feedback time exceeds the time recovery threshold. By dynamically adjusting the third and fourth quantization parameters, the guidance trajectory data works in conjunction with the audio guidance trajectory line in an alternately activated guidance sequence to continuously improve the target user's attention level and action response intensity. In specific implementation, the initial value of the third quantization parameter is 30 cm, representing the minimum distance between adjacent points in the visual guidance trajectory, and the initial value of the fourth quantization parameter is 600 milliseconds, representing the minimum time interval for trajectory updates. When the system detects that the target user's action response is weakening, it automatically decreases the third quantization parameter to 15 cm to make the visual guidance trajectory more continuous and smooth. At the same time, when attention is distracted, the fourth quantization parameter is decreased to 300 milliseconds to accelerate the trajectory update speed and help the target user refocus.

[0039] In this embodiment, the multimodal recognition toy robot interaction method includes the following steps: acquiring positioning data of multiple toy robots in the interaction space, as well as feedback visual data and feedback audio data of the target interactor; acquiring guidance visual data and guidance audio data of the toy robots, wherein the guidance visual data includes guidance trajectory data, and the guidance audio data includes spatial audio guidance parameters and audio guidance trajectory lines; synchronizing the guidance visual data, guidance audio data, feedback visual data, and feedback audio data with timestamps to generate temporal interaction data containing feedback time, feedback intensity, and interaction effectiveness indicators; generating a synchronous activation sequence or an alternating activation sequence based on the first correlation quantization value and the second correlation quantization value in the temporal interaction data, wherein when the product of the first correlation quantization value and the second correlation quantization value is higher than the synchronous activation threshold, a time-aligned and spatially overlapping synchronous activation sequence is generated; otherwise, a time-aligned but spatially independent alternating activation sequence is generated, and the activation sequence is dynamically adjusted according to the feedback data; generating control commands based on the guidance sequence to drive the toy robots to execute specified responses, continuously maintaining the target interactor's attention level and guiding them to a preset interaction area. In practical applications, this method first determines the position of each toy robot through a positioning system, then synchronously collects multimodal data of the target interactor, and after time alignment and quantitative analysis, the system intelligently decides whether to adopt a synchronous activation or alternating activation strategy, and finally generates precise control commands so that the toy robot can guide the target interactor to complete the interaction task in the most effective way. The whole process is completed within 200 milliseconds, ensuring the smoothness and real-time performance of the interaction.

[0040] It should be noted that, in this embodiment, the recognition module specifically includes a high-resolution camera and a microphone array. The camera uses 1080P resolution and a frame rate of 30fps, capable of accurately capturing the facial expressions and body movements of the target interactor. The microphone array consists of four omnidirectional microphones with a sampling rate of 48kHz, enabling sound source localization and voice enhancement. The data processing module adopts an edge computing architecture, performing preliminary data processing locally on the toy robot to reduce the burden on the central server, while simultaneously achieving data synchronization and coordination between the toy robots via a 5G network. The control module uses a fuzzy logic controller, capable of generating appropriate control commands based on complex interaction scenarios, avoiding mechanical responses. The directional speaker in the response module uses beamforming technology to precisely direct sound towards the target interactor, reducing environmental interference; the RGB-LED array can generate various colors and dynamic effects, enhancing the attractiveness of visual guidance.

[0041] In this embodiment, the system's interaction effectiveness index is calculated by weighting feedback time and feedback intensity, with feedback time having a weight of 0.6 and feedback intensity having a weight of 0.4. This weighting is determined based on a large amount of experimental data and can accurately reflect the target user's participation and response quality. When the interaction effectiveness index consistently falls below a set threshold, the system automatically adjusts its guidance strategy, such as increasing the diversity and interest of the guidance, to re-attract the target user's attention. The system also has learning capabilities, enabling it to optimize parameter settings based on historical interaction data, thus continuously improving the guidance effect over time.

[0042] In practical applications, this system can be used in multiple fields such as children's education, rehabilitation training, and interactive entertainment. In children's education, the system helps children complete learning tasks through multimodal guidance, such as guiding them through puzzles via visual cues while providing voice prompts and encouragement through audio guidance. In rehabilitation training, the system helps patients train their physical movements through precise guidance sequences, monitoring the completion rate of movements in real time and providing feedback. In interactive entertainment, the system creates an immersive interactive experience, allowing users to interact naturally and smoothly with multiple toy robots. Regardless of the scenario, the system can dynamically adjust its guidance strategy based on the user's real-time responses, ensuring the effectiveness and enjoyment of the interaction.

[0043] Another feature of this disclosure is that the system can simultaneously support multiple target interactors, providing a personalized guidance sequence for each target interactor through independent multimodal data processing channels. In multi-person interaction scenarios, the system first distinguishes different target interactors through facial recognition and voiceprint recognition, and then generates independent temporal interaction data and guidance sequences for each person, ensuring that each user can obtain an interactive experience suitable for them. When interaction between multiple target interactors is detected, the system can also generate a collaborative guidance sequence to promote cooperation and communication among users and enhance the social interaction effect.

[0044] In this embodiment, the intelligent switching between synchronous activation sequences and alternating activation sequences is the core innovation of the system. When the system detects that the target user responds well to multimodal guidance (i.e., the product of the first and second association quantification values ​​is higher than the synchronous activation threshold), it activates the synchronous activation sequence to precisely align visual and audio guidance in time and space, forming a multimodal guidance focus. This simultaneous stimulation of multiple senses significantly enhances the guidance effect, making it particularly suitable for tasks requiring high concentration. Conversely, when the system detects a decline in the effectiveness of a particular modality of guidance, it automatically switches to the alternating activation sequence. This temporal alternation avoids sensory fatigue and compensates for the deficiencies of one modality with another. For example, when the visual guidance effect declines, the system increases the intensity and frequency of audio guidance to help the target user refocus.

[0045] The system also possesses adaptive capabilities, dynamically adjusting parameter settings based on the target user's age, gender, emotional state, and interaction history. For example, for young children, the system uses brighter visual guidance and simpler audio prompts, while shortening the guidance interval; for users with easily distracted attention, the system increases the diversity and fun of the guidance to avoid monotony and repetition. This personalized adjustment is achieved by analyzing historical interaction data and real-time feedback, enabling the system to provide the most suitable interactive experience for each user.

[0046] In this embodiment, precise control of the audio and visual guidance trajectories is achieved through dynamic adjustment of quantization parameters. The system monitors the feedback data of the target user in real time. When it detects that the feedback time is too long or the feedback intensity is insufficient, it immediately adjusts the corresponding quantization parameters to improve the accuracy and response speed of the guidance. For example, when the system detects that the target user's response to the audio guidance is slowing down, it decreases the second quantization parameter to speed up the switching speed of the audio guidance; when it detects that the action response is not obvious enough, it decreases the third quantization parameter to make the visual guidance trajectory more continuous and smooth. This dynamic adjustment mechanism ensures the optimization of the guidance effect, enabling the system to adapt to various interaction scenarios and user needs.

[0047] It should be noted that in this embodiment, key parameters such as the system's time recovery threshold and visual intensity maintenance threshold are not fixed, but dynamically adjusted based on the individual differences and interaction history of the target user. The system continuously optimizes these threshold settings through machine learning algorithms and stored open-source data to better suit each user's specific needs. For example, for users with short attention spans, the system appropriately lowers the time recovery threshold to initiate the attention recovery mechanism earlier; for users with weak action responses, the system raises the visual intensity maintenance threshold to more actively adjust the visual guidance strategy. This adaptive threshold adjustment mechanism significantly improves the system's practicality and effectiveness.

[0048] Another innovation of this disclosure is that the system can continuously optimize the guidance sequence using feedback data. After each interaction, the system analyzes the data from the entire interaction process, including the correlation quantification values ​​of each stage, feedback time, and feedback intensity, to evaluate the effectiveness of the guidance sequence and adjust the parameter settings for subsequent interactions accordingly. Through this continuous learning and optimization, the system can continuously improve the interaction effect and provide users with an increasingly better experience. In long-term use, the system will form a personalized interaction model for each user, making guidance more accurate and effective.

[0049] In its implementation, the system's data processing module adopts a distributed architecture. Based on the data processing and control modules mounted on the server, such as... Figure 1 As shown, preliminary data processing can also be performed locally on the toy robot, reducing the burden on the central server. Simultaneously, data synchronization and coordination between devices are achieved through a high-speed wireless network. This architecture not only improves the system's response speed but also enhances its reliability and scalability, enabling it to support larger-scale interactive scenarios and a greater number of toy robots. The control module employs a fuzzy logic controller, capable of generating appropriate control commands based on complex interactive scenarios, avoiding mechanical responses and making the interaction more natural and human-like.

[0050] Finally, the system in this embodiment also has a safety protection mechanism that can monitor the safety status of the interactive environment in real time. When a potential danger is detected (such as the target user getting too close to the toy robot or the ambient light being too dim), the system will automatically adjust the guidance strategy or suspend the interaction to ensure the user's safety. This safety protection mechanism is achieved through multi-sensor fusion and intelligent analysis, providing comprehensive safety assurance without affecting the interactive experience.

[0051] The multimodal recognition toy robot interaction system provided in this disclosure achieves comprehensive perception and precise analysis of the interaction process by acquiring feedback visual and audio data from the target interactor, as well as guidance visual and audio data from the toy robot. The system synchronizes the multimodal data stream with timestamps, generating time-series interaction data containing feedback time, feedback intensity, and interaction effectiveness indicators. This provides an objective and quantitative evaluation basis for the interaction effect, making the formulation of guidance strategies more data-driven.

[0052] By introducing a first correlation quantization value and a second correlation quantization value, the system can separately evaluate the guiding effect of visual guidance on the target interactor's action response and the guiding effect of audio guidance on attention level, achieving refined differentiation and independent evaluation of multimodal guidance effects. This quantification mechanism enables the system to accurately identify which guidance method is more effective, providing a key decision-making basis for the subsequent generation of guidance sequences.

[0053] The system-generated synchronous activation sequences and alternating activation sequences constitute a flexible guidance strategy system. When the product of the first and second association quantification values ​​exceeds the synchronous activation threshold, the system initiates a time-aligned and spatially overlapping synchronous activation sequence, forming a multimodal guidance focus and significantly enhancing the guidance intensity for the target user. In other cases, the system employs a time-aligned but spatially independent alternating activation sequence and dynamically adjusts the activation order based on feedback data, effectively avoiding sensory fatigue and achieving intelligent switching of guidance methods. This dual-sequence mechanism ensures that the system can adapt to different interaction states and continuously maintain the target user's attention level.

[0054] The audio guidance trajectory is defined by a first and a second quantization parameter, while the guidance trajectory data is defined by a third and a fourth quantization parameter, enabling the system to precisely control the spatiotemporal characteristics of the guidance. When the feedback time exceeds the time recovery threshold, the system decreases the second quantization parameter to improve temporal resolution and accelerate the switching speed of the audio guidance; when the feedback intensity is below the threshold, the system decreases the first quantization parameter to improve spatial resolution, making the guidance more refined. This dynamic adjustment mechanism of quantization parameters ensures that the guidance strategy can be optimized based on the real-time reactions of the target interactor, continuously improving the interaction effect.

[0055] By using alternating activation of the guidance sequence, the system increases the average value of the first and second correlation quantization values ​​while reducing their fluctuations, thus improving interaction stability. When the feedback intensity of the visual data remains below the visual intensity maintenance threshold, the system extends the activation duration of the guidance audio data, utilizing the indirect guidance effect of audio on action response. When the feedback time of the audio data continuously exceeds the time recovery threshold, the system expands the trajectory coverage of the guidance visual data, utilizing the indirect guidance effect of visual guidance on attention level. This cross-compensation mechanism allows the system to effectively compensate for the decline in the effectiveness of one guidance method with another, ensuring continuous optimization of the overall guidance effect.

[0056] The spatial audio guidance parameters are continuously defined by the maximum audio intensity feature points of each toy robot within the interactive space, constrained by a specified sound source intensity location, forming an audio guidance trajectory line. This allows the system to dynamically move the virtual sound source by controlling phase difference while keeping the toy robots in a fixed position. This technology achieves audio guidance without relying on physical movement, enriching the interaction methods while reducing the latency and malfunction risks associated with mechanical movement.

[0057] The system dynamically adjusts the activation sequence of guiding visual and audio data based on the feedback time change rate in the temporal interaction data, ensuring that the alternation interval between visual and audio guidance matches the target user's attention recovery cycle. This personalized timing adjustment mechanism makes the guidance rhythm more aligned with the target user's cognitive characteristics, improving the adaptability and effectiveness of the guidance.

[0058] The system can continuously maintain the target user's attention level and optimize their action response intensity, effectively guiding them to the preset interaction area and achieving a comprehensive improvement in interaction effectiveness. Furthermore, the system can continuously optimize parameter settings based on historical interaction data, ensuring that the guidance effect improves over time and provides users with an increasingly precise and effective interactive experience.

[0059] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.

[0060] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to achieve the described functions, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the described devices, apparatuses, and units can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0061] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, function, and operation of possible implementations of apparatus, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than those disclosed in the description; sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based device that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. A multimodal recognition toy robot interaction system, comprising multiple toy robots disposed within an interaction space, characterized in that, Also includes: The positioning module is configured to acquire positioning data of multiple toy robots within the interactive space; A recognition module, located in each of the toy robots, is configured to acquire multimodal interaction data of the target interactor within the interaction space; The data processing module is configured to acquire the multimodal interaction data and the multimodal instruction data from the toy robot, and to time-align the multimodal interaction data and the multimodal instruction data to obtain time-series interaction data; The control module is configured to receive the timing interaction data and generate control commands to each of the toy robots; A response module, located in each of the toy robots, is configured to execute a specified response by the toy robot within the interaction space according to the control command, directing the target interactor to a preset interaction area.

2. The multimodal recognition toy robot interaction system according to claim 1, characterized in that, The multimodal interaction data includes the feedback visual data and feedback audio data of the target interactor, and the multimodal instruction data includes the guidance visual data and guidance audio data of the toy robot; When the data processing module is configured to perform timing alignment, it includes the following steps: The guiding visual data, the guiding audio data, the feedback visual data, and the feedback audio data are synchronized using timestamps. Generate time-series interaction data, where each unit is an instruction event. Each instruction event includes the feedback time corresponding to the feedback visual data of the target interactor, the feedback intensity corresponding to the feedback audio data, and an interaction effectiveness index calculated based on the feedback time and feedback intensity after the occurrence of the guidance visual data and guidance audio data.

3. The multimodal recognition toy robot interaction system according to claim 2, characterized in that, The guiding visual data includes guiding trajectory data, and the guiding trajectory data has a quantified first correlation value with the emotional feature data and motion posture data extracted from the feedback visual data of the target interactor; The guiding audio data includes spatial audio guiding parameters and audio guiding trajectory lines, and the spatial audio guiding parameters have a quantized second correlation quantization value with the feedback audio data of the target interactor; Both the first correlation quantization value and the second correlation quantization value represent their respective guidance strength and mapping relationship. The first correlation quantization value reflects the guiding effect of visual guidance on the attention level of the target interactor, and the second correlation quantization value reflects the guiding effect of audio guidance on the attention level of the target interactor.

4. The multimodal recognition toy robot interaction system according to claim 3, characterized in that, When the spatial audio guidance parameters are generated, the audio guidance trajectory line is continuously defined by the maximum audio intensity feature points obtained by each toy robot in the interactive space at the specified sound source intensity position, which varies with time. Furthermore, while maintaining its own position, the toy robot generates spatial audio guidance parameters with phase differences by using toy robots in different positions, or the spatial audio guidance parameters are obtained by adjusting the position of the toy robot.

5. The multimodal recognition toy robot interaction system according to claim 4, characterized in that, When the guiding visual data and the guiding audio data are coordinated, a guiding sequence is generated, which includes a synchronous activation sequence and an alternating activation sequence. In the synchronous activation sequence, the guiding visual data and the guiding audio data are aligned in time and overlap in space to form a multimodal guiding focus. When the product of the first association quantization value and the second association quantization value is higher than the synchronous activation threshold, the control module is configured to start the synchronous activation sequence to enhance the guidance intensity for the target interactor. In the alternating activation sequence, the guiding visual data and the guiding audio data are temporally aligned but spatially independent; The control module is configured as follows: Based on the dynamic trend of the first associated quantization value, when the first associated quantization value is lower than the first associated quantization decrease threshold and continues to decrease, the activation interval of the guiding audio data is shortened so as to reset the attention level of the target interactor through audio guidance. Based on the dynamic trend of the second correlation quantization value, when the second correlation quantization value is lower than the second correlation quantization decrease threshold and continues to decrease, the trajectory point density of the guiding visual data is increased to enhance the action response of the target interactor through visual guidance. Based on the feedback time change rate in the time-series interaction data, the activation sequence of the guiding visual data and the guiding audio data is dynamically adjusted so that the alternation interval between visual guidance and audio guidance matches the attention recovery cycle of the target interactor. By using the alternating activation sequence, the average value of the first associated quantization value and the second associated quantization value is increased while the fluctuation amplitude is reduced, wherein: When the feedback intensity of the feedback visual data remains below the visual intensity maintenance threshold, the control module is configured to extend the activation duration of the guiding audio data in order to utilize the indirect guiding effect of audio guidance on the action response. When the feedback time of the feedback audio data continues to exceed the time recovery threshold, the control module is configured to expand the trajectory coverage of the guiding visual data in order to utilize the indirect guiding effect of visual guidance on attention level. Continuously maintain the target user's attention level and optimize their action response intensity to guide the target user to the preset interaction area.

6. The multimodal recognition toy robot interaction system according to claim 5, characterized in that, The audio guide trajectory line is defined by a first quantization parameter and a second quantization parameter, wherein: The first quantization parameter represents the minimum spatial position interval of the audio guide trajectory line; The second quantization parameter represents the minimum update interval of the audio guide trajectory line in time; The data processing module is configured as follows: When the feedback time exceeds the time recovery threshold, the second quantization parameter is reduced to improve the temporal resolution, causing the audio guidance trajectory line to switch in the alternating activation sequence to maintain the target interactor's attention level; when the feedback intensity of the feedback audio data is lower than the audio intensity maintenance threshold, the first quantization parameter is reduced to improve the spatial resolution, thereby enhancing the audio guidance effect on the target interactor.

7. The multimodal recognition toy robot interaction system according to claim 6, characterized in that, The guidance trajectory data is limited by a third quantization parameter and a fourth quantization parameter, wherein: The third quantization parameter represents the minimum spatial distance between adjacent trajectory points in the guided trajectory data; The fourth quantization parameter represents the minimum time interval for updating trajectory points in the guided trajectory data; The control module is configured as follows: When the feedback intensity of the feedback visual data is lower than the visual intensity maintenance threshold, the third quantization parameter is reduced to increase the trajectory point density; when the feedback time exceeds the time recovery threshold, the fourth quantization parameter is reduced to increase the trajectory update rate. By dynamically adjusting the third and fourth quantization parameters, the guidance trajectory data works synergistically with the audio guidance trajectory line in the alternating activation guidance sequence, continuously improving the target user's attention level and action response intensity.

8. A multimodal recognition method for toy robot interaction, characterized in that, Includes the following steps: S1. Acquire the positioning data of multiple toy robots in the interactive space, as well as the visual and audio feedback data of the target interactor; S2. Acquire the guidance visual data and guidance audio data of the toy robot, wherein the guidance visual data includes guidance trajectory data, and the guidance audio data includes spatial audio guidance parameters and audio guidance trajectory lines; S3. Use timestamps to synchronize and guide visual data, guide audio data, provide feedback visual data and feedback audio data, and generate time-series interaction data that includes feedback time, feedback intensity and interaction effectiveness indicators; S4. Based on the first and second correlation quantization values ​​in the time-series interaction data, generate a synchronous activation sequence or an alternating activation sequence, wherein: When the product of the first associated quantization value and the second associated quantization value is higher than the synchronization activation threshold, a time-aligned and spatially overlapping synchronization activation sequence is generated. Otherwise, generate time-aligned but spatially independent alternating activation sequences and dynamically adjust the activation timing based on feedback data; S5. Based on the guide sequence, generate control instructions to drive the toy robot to execute the specified response, continuously maintain the attention level of the target interactor and guide it to the preset interaction area.