A dynamic sight tracking and attention guiding system and method for a toy robot
By acquiring multimodal data through the visual acquisition and behavioral perception modules of the toy robot, and combining this with the adaptive learning module to optimize the guidance strategy, the problem of monotonous user interaction in existing technologies has been solved. This has enabled precise user interaction and a dynamic interactive experience, thereby improving user engagement and market competitiveness.
Patent Information
- Application Number
- CN202510595991.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-05-09
AI Technical Summary
Existing toy robots cannot accurately perceive the user's gaze trajectory and action response, lack multi-dimensional data collection methods, resulting in a single interaction mode, making it difficult to achieve personalized and dynamic guidance, and lacking adaptive learning capabilities, thus failing to continuously improve user engagement and attention.
The system employs a visual acquisition module and a behavior perception module working together to acquire users' eye movement and limb movement parameters through multimodal data acquisition. The central processing unit performs attention analysis and strategy generation, and combines the adaptive learning module to optimize the guidance strategy, thereby achieving dynamic eye tracking and attention guidance.
It enables precise collection of user interaction data and personalized, dynamic interactive guidance, enhancing user engagement and interactive experience, adapting to different users' behavioral habits and attention characteristics, and expanding application scenarios and user groups.
Smart Images

Figure CN120495345B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot vision control technology, specifically to a dynamic gaze tracking and attention guidance system and method for a toy robot. Background Technology
[0002] With the rapid development of artificial intelligence and robotics, toy robots are increasingly widely used in children's education and entertainment. Existing toy robots mostly rely on pre-programmed instructions or simple sensor feedback to interact with users, resulting in relatively simple and fixed interaction modes. During user interaction with toy robots, it is often impossible to accurately perceive the user's attention state and motor responses, making personalized and dynamic guidance difficult. For example, traditional toy robots cannot effectively capture the user's gaze trajectory or determine whether the user's focus is within the robot's guidance area, leading to a mismatch between guidance content and user interests. When acquiring user body movement information, there is a lack of multi-dimensional data collection methods, making it impossible to fully understand the user's intentions and execution, and thus difficult to adjust guidance strategies based on the user's actual state. Furthermore, existing guidance strategies typically lack adaptive learning capabilities and cannot optimize according to changes in user behavior and attention characteristics, causing the interactive experience to gradually lose its appeal and making it difficult to maintain user engagement and attention for extended periods, thus limiting the development of toy robots in terms of improving user interaction experience and expanding functionality. Summary of the Invention
[0003] To address the above problems, the present invention provides a dynamic gaze tracking and attention guidance system and method for toy robots.
[0004] In a first aspect, the present invention provides a dynamic gaze tracking and attention guidance system for a toy robot, comprising a user interaction terminal and a central processing terminal;
[0005] The user interaction terminal includes at least:
[0006] The vision acquisition module is configured to capture the user's eye movement data in real time to generate gaze trajectory information, and is also configured to collect the user's action response data when interacting with the robot.
[0007] The behavior perception module is configured to acquire the user's limb movement parameters when following the robot's guidance commands through a multi-axis inertial sensor;
[0008] The first communication unit is configured to transmit the gaze trajectory information, action response data and limb action parameters to the central processing unit.
[0009] The central processing unit includes at least:
[0010] The second communication unit is configured to receive multimodal data from the user interaction terminal;
[0011] The attention analysis module is configured to calculate the duration and frequency of the user's gaze focus within the robot-guided area based on the gaze trajectory information, and generate an attention intensity index.
[0012] The dynamic guidance module is configured to generate an initial interaction scheme containing multi-level guidance strategies based on the attention intensity index and action response data. The guidance strategies include robot motion trajectory, audio-visual feedback mode and task difficulty gradient, wherein each strategy parameter is dynamically adjusted based on real-time user data.
[0013] The evaluation and correction module is configured to compare the deviation between the user’s actual body movement parameters when performing the initial interaction plan and the preset movement model to generate a first correction coefficient; at the same time, it generates a second correction coefficient based on the difference between the attention intensity index and the expected threshold.
[0014] The strategy optimization module is configured to integrate a first correction coefficient and a second correction coefficient to generate a composite evaluation index, and generate a differentiated guidance scheme based on the index. The scheme includes: adjusting the robot's movement speed, increasing the intensity of tactile feedback, or reconstructing the task sequence to improve the user's attention focus level.
[0015] As a preferred embodiment, the visual acquisition module integrates an infrared imaging unit and a visible light camera to work together. It is configured to automatically switch to infrared mode when the ambient light intensity is below 50 lux, and uses a pupil-corneal reflection vector algorithm to compensate for eye tracking errors caused by head movement.
[0016] As a preferred embodiment, the behavior-aware module includes:
[0017] A nine-axis MEMS sensor array is configured to capture the angular velocity and linear acceleration of the user's limb movements at a sampling rate of 100Hz.
[0018] A pressure-sensitive epidermal layer is configured to cover the robot's accessible surface, quantifying the force and frequency of user touch through piezoelectric signals;
[0019] The sound field localization unit is configured to recognize the azimuth and emotional features of user voice commands using beamforming technology based on a microphone array.
[0020] As a preferred embodiment, the dynamic boot module includes:
[0021] The motion planning submodule is configured to dynamically adjust the complexity of the robot's movement trajectory based on the attention intensity index, and trigger a spiral progressive motion mode when the index value is below a threshold.
[0022] The multimodal feedback submodule is configured to synchronously drive an RGBW four-color LED array, a micro vibration motor, and a directional speaker to generate complex sensory stimuli that match the user's attention state.
[0023] As a preferred approach, an adaptive learning module is also included, configured to perform the following operations:
[0024] Establish a user behavior database to store historical eye movement patterns, action response latency, and task completion accuracy.
[0025] Temporal convolutional neural networks are used to analyze user attention decay cycles and predict the optimal intervention time.
[0026] By iteratively optimizing the guidance strategy parameters through reinforcement learning algorithms, the second correction coefficient is reduced to a preset threshold within three consecutive interaction cycles.
[0027] As a preferred embodiment, the adaptive learning module is further configured as follows:
[0028] Construct a virtual twin model to simulate the typical attention characteristics of users in different age groups;
[0029] During the offline training phase, adversarial training is performed using real user data and virtual model data to generate a robust guidance strategy library.
[0030] In the online application phase, transfer learning technology is used to dynamically adapt the strategy library parameters to the current user characteristics.
[0031] A second aspect of the present invention provides a method for dynamic gaze tracking and attention guidance of a toy robot, comprising the following steps:
[0032] The steps include:
[0033] S1. Real-time acquisition of user's eye movement data through visual sensors to generate gaze trajectory information including gaze position and duration;
[0034] S2. Capture user limb movement parameters using an inertial sensor array, the parameters including at least the amplitude of movement, the frequency of movement, and the response delay time;
[0035] S3. Transmit the gaze trajectory information and body movement parameters to the central processing unit, and calculate the attention intensity index based on the distribution characteristics of the gaze focus in the preset guidance area;
[0036] S4. Generate an initial interactive guidance scheme based on the attention intensity index. The scheme defines the robot's movement path, multimodal sensory feedback mode, and stage task sequence, wherein the difficulty of each stage task is dynamically adjusted according to real-time user data.
[0037] S5. During the user's execution of the initial interactive guidance plan, the deviation between the actual limb movement parameters and the preset movement model is compared simultaneously to generate a first correction coefficient; at the same time, a second correction coefficient is generated based on the degree to which the attention intensity index deviates from the expected threshold.
[0038] S6. Integrate the first correction coefficient and the second correction coefficient to generate a composite evaluation index. When the index exceeds a preset threshold, reconstruct at least one parameter group in the initial interactive guidance scheme. The parameter group includes: extending the robot's motion path dwell time, increasing the acoustic and optical feedback intensity gradient, or inserting auxiliary tactile prompting tasks.
[0039] S7. Send the reconstructed interactive guidance plan to the robot for execution, and repeat steps S1 to S6 until the user's attention intensity index stabilizes within the target range.
[0040] As a preferred embodiment, the generation of the first correction coefficient in step S5 includes:
[0041] The generation of the first correction coefficient in step S5 includes:
[0042] Extract the three-dimensional spatial motion trajectory from the user's limb movement parameters and perform dynamic time warping matching on it with the preset motion model;
[0043] Calculate the weighted sum of the standard deviations of the actual trajectory and the model trajectory in the dimensions of velocity, acceleration, and orientation angle, and map the weighted sum to the range of 0-1 to generate the first correction coefficient;
[0044] The generation of the second correction coefficient in step S5 includes:
[0045] Baseline reference values are calculated by averaging the standard deviation of fixation duration and peak shift frequency in users' historical attention data.
[0046] Load a preset weight matrix based on the current task type to generate dynamic threshold upper and lower limits;
[0047] The standardized difference between the attention intensity index and the historical mean is calculated in real time. The standardized difference is the difference between the current value and the historical mean divided by the historical standard deviation.
[0048] Calculate the time series fluctuation entropy value, which is based on the distribution probability of attention intensity within a dynamic threshold range for at least ten consecutive frames;
[0049] When the standardized difference exceeds a preset range, a linear correction component proportional to the absolute value of the difference is generated;
[0050] When the fluctuation entropy value exceeds a preset threshold, a nonlinear correction component is generated using the hyperbolic tangent function;
[0051] The linear and nonlinear correction components are weighted and averaged, and a time decay factor is added to smooth the coefficient fluctuations.
[0052] When the second correction factor exceeds 0.5, the user is determined to have entered a state of persistent attention loss.
[0053] As a preferred embodiment, S3 further includes:
[0054] S3a. Within the field of view of the camera of the robot vision module, establish a three-dimensional spatial coordinate system centered on feature marker points, wherein the feature marker points include reflective markers worn by the user;
[0055] S3b. Calculate the projection coordinates of feature markers in the image plane in real time, and construct a spatial angular position mapping model based on the camera optical distortion parameters. The model includes radial distortion compensation factor and tangential distortion compensation factor.
[0056] S3c. When the feature marker point deviates from the preset tracking area, the compensation action is calculated in the following way:
[0057] Extract the elliptical fitting contour of the marked points in the current frame image, and calculate the azimuth offset Δθ based on the angle between the major axis direction and the preset standard direction;
[0058] The radial distance offset Δd is calculated based on the rate of change of the ratio of the area of the marked point to the area of the preset reference, combined with the camera focal length parameter.
[0059] S3d. Input the Δθ and Δd into the motion compensation controller to generate control commands for robot chassis movement or gimbal rotation, so that the feature marker points return to the standardized trajectory band within ±5% of the center area of the image;
[0060] S3e. When the cumulative displacement of the compensation commands in three consecutive frames exceeds a preset threshold, the dynamic recalibration mode is triggered:
[0061] The camera is driven to perform multi-angle swing scanning to re-establish the spatial positional constraint relationship between the feature markers and the robot body;
[0062] Based on the Kalman filter, the target's trajectory is predicted, and the distortion compensation factor in the spatial angular position mapping model is updated.
[0063] As a preferred approach, the feature marker expansion in step S3a includes non-reflective user-native graphical feature points, comprising the following correction calculation steps:
[0064] S3f. Based on the salient features of the user's facial contours or clothing patterns, dynamic reference points are extracted in real time through a convolutional neural network to generate a hybrid marker topology map associated with reflective markers;
[0065] S3g. Based on the actual imaging position of each feature point in the hybrid label topology map, the theoretical spatial coordinates are inferred from the camera distortion model, the coordinate deviation Δe caused by distortion is calculated, and mapped to a third correction coefficient in the range of 0-1.
[0066] S3h. In step S6, the third correction coefficient is introduced into the composite evaluation index to generate the expanded attention misbehavior coefficient:
[0067] When Δe exceeds 0.3 for five consecutive frames, it is determined to be an active behavior of attention shift, triggering the robot to move backward to compensate.
[0068] When the product of Δe and the second correction coefficient is greater than 0.2, the audio-visual feedback is enhanced and the visual sampling frequency is increased.
[0069] Compared with the prior art, the present invention has the following advantages:
[0070] The dynamic gaze tracking and attention guidance system and method for toy robots described in this invention have significant beneficial effects. Firstly, the visual acquisition module and behavior perception module at the user interaction end work together to comprehensively and accurately acquire the user's eye movement data, action response data, and limb movement parameters through multimodal data acquisition, providing a rich and accurate data foundation for subsequent attention analysis and guidance strategy generation. Specifically, the visual acquisition module integrates an infrared imaging unit and a visible light camera, enabling stable operation under different lighting conditions. Furthermore, it employs advanced algorithms to compensate for gaze tracking errors caused by head movement, greatly improving the accuracy and stability of gaze tracking.
[0071] The attention analysis module at the central processing unit generates an attention intensity index based on gaze trajectory information, accurately quantifying the user's attention state. The dynamic guidance module generates an initial interaction plan containing multi-level guidance strategies based on this index and action response data. It can dynamically adjust the parameters of each strategy according to real-time user data, achieving personalized and dynamic interactive guidance and effectively improving user engagement and experience. The evaluation and correction module generates a first correction coefficient and a second correction coefficient by comparing the deviation between the user's actual limb movement parameters and the preset action model, as well as the difference between the attention intensity index and the expected threshold. This provides a quantitative basis for strategy optimization. The strategy optimization module integrates the two correction coefficients to generate a composite evaluation index, and based on this, generates differentiated guidance plans, further improving the adaptability and effectiveness of the guidance strategy and ensuring that the user's attention is focused on interacting with the robot.
[0072] The introduction of the adaptive learning module enables the system to self-optimize. By establishing a user behavior database, analyzing user attention decay cycles using temporal convolutional neural networks, and iteratively optimizing guidance strategy parameters using reinforcement learning algorithms, the system can continuously adapt to the behavioral habits and attention characteristics of different users, thereby continuously improving the accuracy and effectiveness of the guidance strategy. Simultaneously, the adaptive learning module constructs a virtual twin model and performs adversarial training and transfer learning, further enhancing the robustness and generalization ability of the guidance strategy. This makes it applicable to users of different age groups, significantly expanding the application scenarios and user base of toy robots, and greatly enhancing the market competitiveness and user experience of toy robots. Attached Figure Description
[0073] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.
[0074] Figure 1 This is a structural block diagram of the system provided in the embodiments of the present invention. Detailed Implementation
[0075] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0076] A first aspect of this disclosure provides a dynamic gaze tracking and attention guidance system for a toy robot, such as... Figure 1 As shown, it includes a user interface and a central processing unit;
[0077] The user interaction terminal includes at least:
[0078] The vision acquisition module is configured to capture the user's eye movement data in real time to generate gaze trajectory information, and is also configured to collect the user's action response data when interacting with the robot.
[0079] The behavior perception module is configured to acquire the user's limb movement parameters when following the robot's guidance commands through a multi-axis inertial sensor;
[0080] The first communication unit is configured to transmit the gaze trajectory information, action response data and limb action parameters to the central processing unit.
[0081] The central processing unit includes at least:
[0082] The second communication unit is configured to receive multimodal data from the user interaction terminal;
[0083] The attention analysis module is configured to calculate the duration and frequency of the user's gaze focus within the robot-guided area based on the gaze trajectory information, and generate an attention intensity index.
[0084] The dynamic guidance module is configured to generate an initial interaction scheme containing multi-level guidance strategies based on the attention intensity index and action response data. The guidance strategies include robot motion trajectory, audio-visual feedback mode and task difficulty gradient, wherein each strategy parameter is dynamically adjusted based on real-time user data.
[0085] The evaluation and correction module is configured to compare the deviation between the user’s actual body movement parameters when performing the initial interaction plan and the preset movement model to generate a first correction coefficient; at the same time, it generates a second correction coefficient based on the difference between the attention intensity index and the expected threshold.
[0086] The strategy optimization module is configured to integrate a first correction coefficient and a second correction coefficient to generate a composite evaluation index, and generate a differentiated guidance scheme based on the index. The scheme includes: adjusting the robot's movement speed, increasing the intensity of tactile feedback, or reconstructing the task sequence to improve the user's attention focus level.
[0087] As a preferred embodiment, the visual acquisition module integrates an infrared imaging unit and a visible light camera to work together. It is configured to automatically switch to infrared mode when the ambient light intensity is below 50 lux, and uses a pupil-corneal reflection vector algorithm to compensate for eye tracking errors caused by head movement.
[0088] As a preferred embodiment, the behavior-aware module includes:
[0089] A nine-axis MEMS sensor array is configured to capture the angular velocity and linear acceleration of the user's limb movements at a sampling rate of 100Hz.
[0090] A pressure-sensitive epidermal layer is configured to cover the robot's accessible surface, quantifying the force and frequency of user touch through piezoelectric signals;
[0091] The sound field localization unit is configured to recognize the azimuth and emotional features of user voice commands using beamforming technology based on a microphone array.
[0092] As a preferred embodiment, the dynamic boot module includes:
[0093] The motion planning submodule is configured to dynamically adjust the complexity of the robot's movement trajectory based on the attention intensity index, and trigger a spiral progressive motion mode when the index value is below a threshold.
[0094] The multimodal feedback submodule is configured to synchronously drive an RGBW four-color LED array, a micro vibration motor, and a directional speaker to generate complex sensory stimuli that match the user's attention state.
[0095] As a preferred approach, an adaptive learning module is also included, configured to perform the following operations:
[0096] Establish a user behavior database to store historical eye movement patterns, action response latency, and task completion accuracy.
[0097] Temporal convolutional neural networks are used to analyze user attention decay cycles and predict the optimal intervention time.
[0098] By iteratively optimizing the guidance strategy parameters through reinforcement learning algorithms, the second correction coefficient is reduced to a preset threshold within three consecutive interaction cycles.
[0099] As a preferred embodiment, the adaptive learning module is further configured as follows:
[0100] Construct a virtual twin model to simulate the typical attention characteristics of users in different age groups;
[0101] During the offline training phase, adversarial training is performed using real user data and virtual model data to generate a robust guidance strategy library.
[0102] In the online application phase, transfer learning technology is used to dynamically adapt the strategy library parameters to the current user characteristics.
[0103] A second aspect of this disclosure provides a method for dynamic gaze tracking and attention guidance of a toy robot, comprising the following steps:
[0104] The steps include:
[0105] S1. Real-time acquisition of user's eye movement data through visual sensors to generate gaze trajectory information including gaze position and duration;
[0106] S2. Capture user limb movement parameters using an inertial sensor array, the parameters including at least the amplitude of movement, the frequency of movement, and the response delay time;
[0107] S3. Transmit the gaze trajectory information and body movement parameters to the central processing unit, and calculate the attention intensity index based on the distribution characteristics of the gaze focus in the preset guidance area;
[0108] S4. Generate an initial interactive guidance scheme based on the attention intensity index. The scheme defines the robot's movement path, multimodal sensory feedback mode, and stage task sequence, wherein the difficulty of each stage task is dynamically adjusted according to real-time user data.
[0109] S5. During the user's execution of the initial interactive guidance plan, the deviation between the actual limb movement parameters and the preset movement model is compared simultaneously to generate a first correction coefficient; at the same time, a second correction coefficient is generated based on the degree to which the attention intensity index deviates from the expected threshold.
[0110] S6. Integrate the first correction coefficient and the second correction coefficient to generate a composite evaluation index. When the index exceeds a preset threshold, reconstruct at least one parameter group in the initial interactive guidance scheme. The parameter group includes: extending the robot's motion path dwell time, increasing the acoustic and optical feedback intensity gradient, or inserting auxiliary tactile prompting tasks.
[0111] S7. Send the reconstructed interactive guidance plan to the robot for execution, and repeat steps S1 to S6 until the user's attention intensity index stabilizes within the target range.
[0112] As a preferred embodiment, the generation of the first correction coefficient in step S5 includes:
[0113] The generation of the first correction coefficient in step S5 includes:
[0114] Extract the three-dimensional spatial motion trajectory from the user's limb movement parameters and perform dynamic time warping matching on it with the preset motion model;
[0115] Calculate the weighted sum of the standard deviations of the actual trajectory and the model trajectory in the dimensions of velocity, acceleration, and orientation angle, and map the weighted sum to the range of 0-1 to generate the first correction coefficient;
[0116] The generation of the second correction coefficient in step S5 includes:
[0117] Baseline reference values are calculated by averaging the standard deviation of fixation duration and peak shift frequency in users' historical attention data.
[0118] Load a preset weight matrix based on the current task type to generate dynamic threshold upper and lower limits;
[0119] The standardized difference between the attention intensity index and the historical mean is calculated in real time. The standardized difference is the difference between the current value and the historical mean divided by the historical standard deviation.
[0120] Calculate the time series fluctuation entropy value, which is based on the distribution probability of attention intensity within a dynamic threshold range for at least ten consecutive frames;
[0121] When the standardized difference exceeds a preset range, a linear correction component proportional to the absolute value of the difference is generated;
[0122] When the fluctuation entropy value exceeds a preset threshold, a nonlinear correction component is generated using the hyperbolic tangent function;
[0123] The linear and nonlinear correction components are weighted and averaged, and a time decay factor is added to smooth the coefficient fluctuations.
[0124] When the second correction factor exceeds 0.5, the user is determined to have entered a state of persistent attention loss.
[0125] As a preferred embodiment, S3 further includes:
[0126] S3a. Within the field of view of the camera of the robot vision module, establish a three-dimensional spatial coordinate system centered on feature marker points, wherein the feature marker points include reflective markers worn by the user;
[0127] S3b. Calculate the projection coordinates of feature markers in the image plane in real time, and construct a spatial angular position mapping model based on the camera optical distortion parameters. The model includes radial distortion compensation factor and tangential distortion compensation factor.
[0128] S3c. When the feature marker point deviates from the preset tracking area, the compensation action is calculated in the following way:
[0129] Extract the elliptical fitting contour of the marked points in the current frame image, and calculate the azimuth offset Δθ based on the angle between the major axis direction and the preset standard direction;
[0130] The radial distance offset Δd is calculated based on the rate of change of the ratio of the area of the marked point to the area of the preset reference, combined with the camera focal length parameter.
[0131] S3d. Input the Δθ and Δd into the motion compensation controller to generate control commands for robot chassis movement or gimbal rotation, so that the feature marker points return to the standardized trajectory band within ±5% of the center area of the image;
[0132] S3e. When the cumulative displacement of the compensation commands in three consecutive frames exceeds a preset threshold, the dynamic recalibration mode is triggered:
[0133] The camera is driven to perform multi-angle swing scanning to re-establish the spatial positional constraint relationship between the feature markers and the robot body;
[0134] Based on the Kalman filter, the target's trajectory is predicted, and the distortion compensation factor in the spatial angular position mapping model is updated.
[0135] As a preferred approach, the feature marker expansion in step S3a includes non-reflective user-native graphical feature points, comprising the following correction calculation steps:
[0136] S3f. Based on the salient features of the user's facial contours or clothing patterns, dynamic reference points are extracted in real time through a convolutional neural network to generate a hybrid marker topology map associated with reflective markers;
[0137] S3g. Based on the actual imaging position of each feature point in the hybrid label topology map, the theoretical spatial coordinates are inferred from the camera distortion model, the coordinate deviation Δe caused by distortion is calculated, and mapped to a third correction coefficient in the range of 0-1.
[0138] S3h. In step S6, the third correction coefficient is introduced into the composite evaluation index to generate the expanded attention misbehavior coefficient:
[0139] When Δe exceeds 0.3 for five consecutive frames, it is determined to be an active behavior of attention shift, triggering the robot to move backward to compensate.
[0140] When the product of Δe and the second correction coefficient is greater than 0.2, the audio-visual feedback is enhanced and the visual sampling frequency is increased.
[0141] In this embodiment of the disclosure, step S1 further includes:
[0142] When the ambient light intensity is below 50 lux, the infrared supplementary light device is activated, and the pupil-corneal reflection vector compensation algorithm is used to eliminate the eye tracking error caused by the user's head deviation.
[0143] When the user's continuous blinking frequency is detected to be more than 3 times / second in this embodiment, the anti-misjudgment mechanism is triggered, the current gaze trajectory analysis is paused and the backup action guidance mode is started.
[0144] The generation of the multimodal sensory feedback pattern in step S4 includes:
[0145] Based on the numerical range of the attention intensity index, a preset feedback combination strategy is matched. When the index value is in the low attention range, the flashing frequency of the robot's LED breathing light and the tone of the buzzer are simultaneously increased.
[0146] When the user's limb movement parameters match the preset model to a preset threshold, the robot's joint vibration motors are activated to generate tactile reward signals.
[0147] The embodiments disclosed herein have the following advantages over the prior art:
[0148] The dynamic gaze tracking and attention guidance system and method for toy robots described in this disclosure have significant beneficial effects. Firstly, the visual acquisition module and behavior perception module of the user interaction terminal work together to comprehensively and accurately acquire the user's eye movement data, action response data, and limb movement parameters through multimodal data acquisition, providing a rich and accurate data foundation for subsequent attention analysis and guidance strategy generation. Specifically, the visual acquisition module integrates an infrared imaging unit and a visible light camera, enabling stable operation under different lighting conditions. Furthermore, it employs advanced algorithms to compensate for gaze tracking errors caused by head movement, greatly improving the accuracy and stability of gaze tracking.
[0149] The attention analysis module at the central processing unit generates an attention intensity index based on gaze trajectory information, accurately quantifying the user's attention state. The dynamic guidance module generates an initial interaction plan containing multi-level guidance strategies based on this index and action response data. It can dynamically adjust the parameters of each strategy according to real-time user data, achieving personalized and dynamic interactive guidance and effectively improving user engagement and experience. The evaluation and correction module generates a first correction coefficient and a second correction coefficient by comparing the deviation between the user's actual limb movement parameters and the preset action model, as well as the difference between the attention intensity index and the expected threshold. This provides a quantitative basis for strategy optimization. The strategy optimization module integrates the two correction coefficients to generate a composite evaluation index, and based on this, generates differentiated guidance plans, further improving the adaptability and effectiveness of the guidance strategy and ensuring that the user's attention is focused on interacting with the robot.
[0150] The introduction of the adaptive learning module enables the system to self-optimize. By establishing a user behavior database, analyzing user attention decay cycles using temporal convolutional neural networks, and iteratively optimizing guidance strategy parameters using reinforcement learning algorithms, the system can continuously adapt to the behavioral habits and attention characteristics of different users, thereby continuously improving the accuracy and effectiveness of the guidance strategy. Simultaneously, the adaptive learning module constructs a virtual twin model and performs adversarial training and transfer learning, further enhancing the robustness and generalization ability of the guidance strategy. This makes it applicable to users of different age groups, significantly expanding the application scenarios and user base of toy robots, and greatly enhancing the market competitiveness and user experience of toy robots.
[0151] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for descriptive purposes only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or,” as used herein, means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.
[0152] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to achieve the described functions, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the described devices, apparatuses, and units can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0153] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, function, and operation of possible implementations of apparatus, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than those disclosed in the description; sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based device that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
Claims
1. A dynamic gaze tracking and attention guidance system for a toy robot, comprising a user interaction terminal and a central processing terminal; The user interaction terminal includes at least: The vision acquisition module is configured to capture the user's eye movement data in real time to generate gaze trajectory information, and is also configured to collect the user's action response data when interacting with the robot. The behavior perception module is configured to acquire the user's limb movement parameters when following the robot's guidance commands through a multi-axis inertial sensor; The first communication unit is configured to transmit the gaze trajectory information, action response data and limb action parameters to the central processing unit. The central processing unit includes at least: The second communication unit is configured to receive multimodal data from the user interaction terminal; The attention analysis module is configured to calculate the duration and frequency of the user's gaze focus within the robot-guided area based on the gaze trajectory information, and generate an attention intensity index. The dynamic guidance module is configured to generate an initial interaction scheme containing multi-level guidance strategies based on the attention intensity index and action response data. The guidance strategies include robot motion trajectory, audio-visual feedback mode and task difficulty gradient, wherein each strategy parameter is dynamically adjusted based on real-time user data. The evaluation and correction module is configured to compare the deviation between the user’s actual body movement parameters when performing the initial interaction plan and the preset movement model to generate a first correction coefficient; at the same time, it generates a second correction coefficient based on the difference between the attention intensity index and the expected threshold. The strategy optimization module is configured to integrate a first correction coefficient and a second correction coefficient to generate a composite evaluation index, and generate a differentiated guidance scheme based on the index. The scheme includes: adjusting the robot's movement speed, increasing the intensity of tactile feedback, or reconstructing the task sequence to improve the user's attention focus level.
2. The dynamic gaze tracking and attention guidance system for a toy robot according to claim 1, characterized in that, The visual acquisition module integrates an infrared imaging unit and a visible light camera to work together. It is configured to automatically switch to infrared mode when the ambient light intensity is below 50 lux, and uses a pupil-corneal reflection vector algorithm to compensate for eye tracking errors caused by head movement.
3. The dynamic gaze tracking and attention guidance system for a toy robot according to claim 1, characterized in that, The behavior perception module includes: A nine-axis MEMS sensor array is configured to capture the angular velocity and linear acceleration of the user's limb movements at a sampling rate of 100Hz. A pressure-sensitive epidermal layer is configured to cover the robot's accessible surface, quantifying the force and frequency of user touch through piezoelectric signals; The sound field localization unit is configured to recognize the azimuth and emotional features of user voice commands using beamforming technology based on a microphone array.
4. The dynamic gaze tracking and attention guidance system for a toy robot according to claim 1, characterized in that, The dynamic boot module includes: The motion planning submodule is configured to dynamically adjust the complexity of the robot's movement trajectory based on the attention intensity index, and trigger a spiral progressive motion mode when the index value is below a threshold. The multimodal feedback submodule is configured to synchronously drive an RGBW four-color LED array, a micro vibration motor, and a directional speaker to generate complex sensory stimuli that match the user's attention state.
5. The dynamic gaze tracking and attention guidance system for a toy robot according to claim 1, characterized in that, It also includes an adaptive learning module, configured to perform the following operations: Establish a user behavior database to store historical eye movement patterns, action response latency, and task completion accuracy. Temporal convolutional neural networks are used to analyze user attention decay cycles and predict the optimal intervention time. By iteratively optimizing the guidance strategy parameters through reinforcement learning algorithms, the second correction coefficient is reduced to a preset threshold within three consecutive interaction cycles.
6. The dynamic gaze tracking and attention guidance system for a toy robot according to claim 1, characterized in that, The adaptive learning module is further configured as follows: Construct a virtual twin model to simulate the typical attention characteristics of users in different age groups; During the offline training phase, adversarial training is performed using real user data and virtual model data to generate a robust guidance strategy library. In the online application phase, transfer learning technology is used to dynamically adapt the strategy library parameters to the current user characteristics.
7. A method for dynamic gaze tracking and attention guidance of a toy robot, characterized in that, Includes the following steps: S1. Real-time acquisition of user's eye movement data through visual sensors to generate gaze trajectory information including gaze position and duration; S2. Capture user limb movement parameters using an inertial sensor array, the parameters including at least the amplitude of movement, the frequency of movement, and the response delay time; S3. Transmit the gaze trajectory information and body movement parameters to the central processing unit, and calculate the attention intensity index based on the distribution characteristics of the gaze focus in the preset guidance area; S4. Generate an initial interaction guidance scheme based on the attention intensity index, wherein the scheme defines the robot's movement path, multimodal sensory feedback mode, and stage task sequence; S5. During the user's execution of the initial interactive guidance plan, the deviation between the actual limb movement parameters and the preset movement model is compared simultaneously to generate a first correction coefficient; at the same time, a second correction coefficient is generated based on the degree to which the attention intensity index deviates from the expected threshold. S6. Integrate the first correction coefficient and the second correction coefficient to generate a composite evaluation index. When the index exceeds a preset threshold, reconstruct at least one parameter group in the initial interactive guidance scheme. The parameter group includes: extending the robot's motion path dwell time, increasing the acoustic and optical feedback intensity gradient, or inserting auxiliary tactile prompting tasks. S7. Send the reconstructed interactive guidance plan to the robot for execution, and repeat steps S1 to S6 until the user's attention intensity index stabilizes within the target range.
8. The dynamic gaze tracking and attention guidance method for a toy robot according to claim 7 further includes the following step: the generation of the first correction coefficient in step S5 includes: Extract the three-dimensional spatial motion trajectory from the user's limb movement parameters and perform dynamic time warping matching on it with the preset motion model; Calculate the weighted sum of the standard deviations of the actual trajectory and the model trajectory in the dimensions of velocity, acceleration, and orientation angle, and map the weighted sum to the range of 0-1 to generate the first correction coefficient; The generation of the second correction coefficient in step S5 includes: Baseline reference values are calculated by averaging the standard deviation of fixation duration and peak shift frequency in users' historical attention data. Load a preset weight matrix based on the current task type to generate dynamic threshold upper and lower limits; The standardized difference between the attention intensity index and the historical mean is calculated in real time. The standardized difference is the difference between the current value and the historical mean divided by the historical standard deviation. Calculate the time series fluctuation entropy value, which is based on the distribution probability of attention intensity within a dynamic threshold range for at least ten consecutive frames; When the standardized difference exceeds a preset range, a linear correction component proportional to the absolute value of the difference is generated; When the fluctuation entropy value exceeds a preset threshold, a nonlinear correction component is generated using the hyperbolic tangent function; The linear and nonlinear correction components are weighted and averaged, and a time decay factor is added to smooth the coefficient fluctuations. When the second correction factor exceeds 0.5, the user is determined to have entered a state of persistent attention loss.
9. The method for dynamic gaze tracking and attention guidance of a toy robot according to claim 8, characterized in that, S3 further includes: S3a. Within the field of view of the camera of the robot vision module, establish a three-dimensional spatial coordinate system centered on feature marker points, wherein the feature marker points include reflective markers worn by the user; S3b. Calculate the projection coordinates of feature markers in the image plane in real time, and construct a spatial angular position mapping model based on the camera optical distortion parameters. The model includes radial distortion compensation factor and tangential distortion compensation factor. S3c. When the feature marker point deviates from the preset tracking area, the compensation action is calculated in the following way: Extract the elliptical fitting contour of the marked points in the current frame image, and calculate the azimuth offset Δθ based on the angle between the major axis direction and the preset standard direction; The radial distance offset Δd is calculated based on the rate of change of the ratio of the area of the marked point to the area of the preset reference, combined with the camera focal length parameter. S3d. Input the Δθ and Δd into the motion compensation controller to generate control commands for robot chassis movement or gimbal rotation, so that the feature marker points return to the standardized trajectory band within ±5% of the center area of the image; S3e. When the cumulative displacement of the compensation commands in three consecutive frames exceeds a preset threshold, the dynamic recalibration mode is triggered: The camera is driven to perform multi-angle swing scanning to re-establish the spatial positional constraint relationship between the feature markers and the robot body; Based on the Kalman filter, the target's trajectory is predicted, and the distortion compensation factor in the spatial angular position mapping model is updated.
10. The method for dynamic gaze tracking and attention guidance of a toy robot according to claim 9, characterized in that, The feature marker expansion in step S3a includes non-reflective user-native graphics feature points, and includes the following correction calculation steps: S3f. Based on the salient features of the user's facial contours or clothing patterns, dynamic reference points are extracted in real time through a convolutional neural network to generate a hybrid marker topology map associated with reflective markers; S3g. Based on the actual imaging position of each feature point in the hybrid label topology map, the theoretical spatial coordinates are inferred from the camera distortion model, the coordinate deviation Δe caused by distortion is calculated, and mapped to a third correction coefficient in the range of 0-1. S3h. In step S6, the third correction coefficient is introduced into the composite evaluation index to generate the expanded attention misbehavior coefficient: When Δe exceeds 0.3 for five consecutive frames, it is determined to be an active behavior of attention shift, triggering the robot to move backward to compensate. When the product of Δe and the second correction coefficient is greater than 0.2, the audio-visual feedback is enhanced and the visual sampling frequency is increased.
Citation Information
Patent Citations
Virtual reality system to promote reading comprehension and writing skills in students
DE202025101195U1
Systems and methods for observing eye and head information to measure ocular parameters and determine human health status
US20220133212A1