Dynamic sight tracking and attention guiding system and method of toy robot

User data is obtained through visual acquisition and behavioral perception modules, combined with the central processing terminal attention analysis and the adaptive learning module optimization guidance strategy, the problem that toy robots cannot accurately perceive user attention and motion response is solved, and personalized and dynamic interactive experience and a wide range of application scenarios are achieved.

CN120495345AActive Publication Date: 2025-08-15DONGGUAN YONGNKIDS TOYS TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510595991.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

Existing toy robots cannot accurately perceive the user's attention state and action response, and lack adaptive learning ability, which leads to the gradually losing appeal of the interactive experience and making it difficult to achieve personalized and dynamic guidance.

Method used

The visual acquisition module and the behavior perception module work together, and the user's eye movement data and limb movement parameters are obtained through multimodal data acquisition. The attention analysis module at the central processing end generates attention intensity indicators, dynamically adjusts the guidance strategy, and optimizes the guidance strategy through the adaptive learning module to adapt to the behavioral habits and attention characteristics of different users.

Benefits of technology

It has realized personalized and dynamic interactive guidance, improved user participation and interactive experience, expanded the application scenarios and user groups of toy robots, and enhanced market competitiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495345A_ABST
    Figure CN120495345A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic sight tracking and attention guiding system and method of a toy robot. The system comprises a user interaction end and a central processing end, the user interaction end captures eye movement data and action response data of a user in real time through a visual acquisition module, and a behavior perception module obtains limb movement parameters and transmits the limb movement parameters to the central processing end through a first communication unit; after the central processing end receives the data, the attention analysis module generates an attention intensity index, the dynamic guiding module generates an initial interaction scheme according to the attention intensity index and action response data, the evaluation correction module generates a correction coefficient, and the strategy optimization module generates a differential guiding scheme based on a composite evaluation index. According to the invention, personalized and dynamic interactive guidance is realized, the user participation degree and experience are improved, and the application scene and market competitiveness of the toy robot are expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot vision control, and in particular to a dynamic sightline tracking and attention guiding system and method for a toy robot. Background Art

[0002] With the rapid development of artificial intelligence and robotics, toy robots are increasingly being used in children's education and entertainment. Existing toy robots often rely on preset programs or simple sensor feedback to interact with users, resulting in a relatively simple and fixed interaction model. During user-toy robot interaction, they often fail to accurately perceive the user's attention state and action response, making personalized and dynamic guidance difficult. For example, traditional toy robots cannot effectively capture the user's gaze trajectory and cannot determine whether the user's focus is within the robot's guidance area, resulting in a mismatch between guidance content and user interests. When acquiring user body movement information, they lack multi-dimensional data collection methods, unable to fully understand the user's action intentions and execution, making it difficult to adjust guidance strategies based on the user's actual state. Furthermore, existing guidance strategies often lack adaptive learning capabilities and cannot optimize as user behavior and attention patterns change. This makes the interactive experience gradually unappealing and difficult to maintain user engagement and attention over time, limiting the development of toy robots in terms of improving user interaction experience and expanding their functionality. Summary of the Invention

[0003] In response to the above problems, the present invention provides a dynamic sightline tracking and attention guiding system and method for a toy robot.

[0004] A first aspect of the present invention provides a dynamic eye tracking and attention guidance system for a toy robot, comprising a user interaction terminal and a central processing terminal;

[0005] The user interaction terminal at least includes:

[0006] A visual acquisition module is configured to capture the user's eye movement data in real time to generate gaze trajectory information, and is also configured to collect action response data when the user interacts with the robot;

[0007] A behavior perception module is configured to obtain body movement parameters of the user when following the robot's guidance instructions through a multi-axis inertial sensor;

[0008] A first communication unit is configured to transmit the gaze trajectory information, action response data and limb movement parameters to a central processing end;

[0009] The central processing end at least includes:

[0010] a second communication unit configured to receive multimodal data from a user interaction terminal;

[0011] an attention analysis module configured to calculate, based on the gaze trajectory information, the duration of the user's gaze focus remaining in the robot's guidance area and the frequency of its shifts, and generate an attention intensity index;

[0012] a dynamic guidance module configured to generate an initial interaction plan including a multi-level guidance strategy based on the attention intensity indicator and the action response data, wherein the guidance strategy includes a robot motion trajectory, an acoustic and optical feedback mode, and a task difficulty gradient, wherein the strategy parameters are dynamically adjusted based on the real-time user data;

[0013] An evaluation and correction module is configured to compare the deviation between the actual body movement parameters of the user when executing the initial interaction plan and the preset movement model to generate a first correction coefficient; and simultaneously generate a second correction coefficient based on the difference between the attention intensity index and the expected threshold;

[0014] The strategy optimization module is configured to fuse the first correction coefficient and the second correction coefficient to generate a composite evaluation index, and generate a differentiated guidance plan based on the index. The plan includes: adjusting the robot's movement speed, increasing the intensity of tactile feedback, or reconstructing the task sequence to improve the user's attention focus level.

[0015] As a preferred embodiment, the visual acquisition module integrates an infrared imaging unit and works in conjunction with a visible light camera. It is configured to automatically switch to infrared mode when the ambient light illumination is lower than 50 lux, and uses a pupil-corneal reflection vector algorithm to compensate for the line of sight tracking error caused by head movement.

[0016] As a preferred embodiment, the behavior perception module includes:

[0017] A nine-axis MEMS sensor array configured to capture the angular velocity and linear acceleration of the user's limb movements at a 100Hz sampling rate;

[0018] A pressure-sensitive surface layer is configured to cover the robot's touchable surface and quantify the user's touch force and frequency through piezoelectric signals;

[0019] The sound field positioning unit is configured to recognize the azimuth and emotional characteristics of the user's voice commands based on the beamforming technology of the microphone array.

[0020] As a preferred embodiment, the dynamic guidance module includes:

[0021] The motion planning submodule is configured to dynamically adjust the complexity of the robot's movement trajectory according to the attention intensity index, and trigger a spiral progressive motion mode when the index value is lower than the threshold;

[0022] The multimodal feedback submodule is configured to synchronously drive the RGBW four-color LED array, micro vibration motor and directional speaker to generate compound sensory stimulation that matches the user's attention state.

[0023] As a preferred embodiment, the system further includes an adaptive learning module configured to perform the following operations:

[0024] Establish a user behavior database to store historical gaze trajectory patterns, action response delays, and task completion accuracy;

[0025] Use a temporal convolutional neural network to analyze the user's attention decay cycle and predict the best time to intervene;

[0026] The guidance strategy parameters are iteratively optimized through the reinforcement learning algorithm, so that the second correction coefficient is reduced to the preset threshold within three consecutive interaction cycles.

[0027] As a preferred embodiment, the adaptive learning module is further configured as follows:

[0028] Build a virtual twin model to simulate the typical attention characteristics of users of different age groups;

[0029] During the offline training phase, adversarial training is performed between actual user data and virtual model data to generate a robust guidance policy library;

[0030] In the online application stage, the policy library parameters are dynamically adapted to the current user characteristics based on transfer learning technology.

[0031] A second aspect of the present invention provides a method for dynamic eye tracking and attention guidance of a toy robot, comprising the following steps:

[0032] The steps include:

[0033] S1. Use visual sensors to collect real-time eye movement data from the user and generate gaze trajectory information including gaze location and duration;

[0034] S2. Capturing user body motion parameters through an inertial sensor array, the parameters including at least motion amplitude, motion frequency, and response delay time;

[0035] S3. Transmitting the gaze trajectory information and limb movement parameters to the central processing end, and calculating the attention intensity index based on the distribution characteristics of the gaze focus in the preset guidance area;

[0036] S4. Generate an initial interactive guidance plan based on the attention intensity indicator, wherein the plan defines the robot's motion path, multimodal sensory feedback mode, and stage task sequence, wherein the difficulty of each stage task is dynamically adjusted based on the user's real-time data;

[0037] S5. During the user's initial interactive guidance program, the deviation between the actual body movement parameters and the preset movement model is compared to generate a first correction coefficient; and a second correction coefficient is generated based on the degree to which the attention intensity indicator deviates from the expected threshold.

[0038] S6. Fusing the first correction coefficient and the second correction coefficient to generate a composite evaluation index. When the index exceeds a preset threshold, reconstructing at least one parameter group in the initial interactive guidance scheme, the parameter group including: extending the robot's dwell time on the motion path, increasing the intensity gradient of the acoustic and optical feedback, or inserting an auxiliary tactile prompt task;

[0039] S7. Send the reconstructed interactive guidance plan to the robot for execution, and repeat steps S1 to S6 until the user's attention intensity index stabilizes in the target range.

[0040] As a preferred embodiment, the generation of the first correction coefficient in step S5 includes:

[0041] The generation of the first correction coefficient in step S5 includes:

[0042] Extract the three-dimensional spatial motion trajectory from the user's body movement parameters and perform dynamic time-warping matching with the preset motion model;

[0043] Calculating a weighted sum of standard deviations of the actual trajectory and the model trajectory in terms of velocity, acceleration, and angular dimensions, and mapping the weighted sum to a range of 0-1 to generate a first correction coefficient;

[0044] The generation of the second correction coefficient in step S5 includes:

[0045] The baseline reference value is calculated based on the standard deviation of the fixation duration and the peak of the shift frequency in the user's historical attention data, and the two are averaged;

[0046] Load the preset weight matrix according to the current task type and generate dynamic upper and lower thresholds;

[0047] Calculate in real time the standardized difference between the attention intensity index and the historical mean, where the standardized difference is the difference between the current value and the historical mean divided by the historical standard deviation;

[0048] Calculating a time series fluctuation entropy value, where the entropy value is calculated based on a distribution probability of at least ten consecutive frames of attention intensity within a dynamic threshold interval;

[0049] When the normalized difference exceeds a preset range, generating a linear correction component proportional to the absolute value of the difference;

[0050] When the fluctuation entropy value exceeds a preset threshold, a nonlinear correction component is generated using a hyperbolic tangent function;

[0051] Perform weighted average of the linear correction component and the nonlinear correction component, and superimpose a time decay factor to smooth coefficient fluctuations;

[0052] When the second correction coefficient exceeds 0.5, it is determined that the user has entered a state of persistent inattention.

[0053] As a preferred embodiment, the step S3 further includes:

[0054] S3a within the camera field of view of the robot vision module, establish a three-dimensional coordinate system centered on the characteristic marker point, the characteristic marker point includes a reflective marker worn by the user;

[0055] S3b real-time calculation of the projection coordinates of the feature marker points in the image plane, based on the camera optical distortion parameters to build a spatial angular position mapping model, the model includes a radial distortion compensation factor and a tangential distortion compensation factor;

[0056] S3c. When the feature marker deviates from the preset tracking area, the compensation action is calculated as follows:

[0057] Extract the ellipse fitting contour of the marked point in the current frame image, and calculate the azimuth offset Δθ based on the angle between the major axis direction and the preset standard direction;

[0058] Based on the ratio change rate of the marker area to the preset reference area and the camera focal length parameter, the radial distance offset Δd is calculated;

[0059] S3d. The Δθ and Δd are input to the motion compensation controller to generate control instructions for the robot chassis movement or pan / tilt rotation, so that the feature markers are regressed to the standardized trajectory within the range of ±5% pixels in the center area of the image;

[0060] S3e. When the cumulative displacement of the compensation instructions for three consecutive frames exceeds the preset threshold, the dynamic recalibration mode is triggered:

[0061] Drive the camera to perform multi-angle swing scanning to re-establish the spatial position constraint relationship between the feature markers and the robot body;

[0062] The target motion trajectory is predicted based on the Kalman filter, and the distortion compensation factor in the spatial angular position mapping model is updated.

[0063] As a preferred embodiment, the feature marker point extension in step S3a includes non-reflective user native graphic feature points, and includes the following correction calculation steps:

[0064] S3f. Based on the salient features of the user's facial contour or clothing pattern, dynamic reference points are extracted in real time through a convolutional neural network to generate a hybrid marker topology associated with the reflective markers;

[0065] S3g based on the actual imaging position of each feature point in the mixed marker topology, combined with the camera distortion model to reverse the theoretical space coordinates, calculate the coordinate deviation caused by the distortion Δe, and mapped to the third correction coefficient in the range of 0-1;

[0066] S3h. In step S6, the third correction coefficient is introduced into the composite evaluation index to generate the expanded attention loss behavior coefficient:

[0067] When Δe exceeds 0.3 for five consecutive frames, it is determined to be an active behavior of attention deviation, triggering the robot's reverse movement compensation;

[0068] When the product of Δe and the second correction coefficient is greater than 0.2, the acoustic and optical feedback is enhanced and the visual sampling frequency is increased.

[0069] Compared with the prior art, the present invention has the following beneficial effects:

[0070] The dynamic gaze tracking and attention guidance system and method for a toy robot described in the present invention have significant beneficial effects. First, the visual acquisition module and the behavioral perception module on the user interaction end work together to comprehensively and accurately obtain the user's eye movement data, action response data, and limb movement parameters through multimodal data acquisition, providing a rich and accurate data foundation for subsequent attention analysis and guidance strategy generation. The visual acquisition module integrates an infrared imaging unit and a visible light camera, capable of stable operation under different lighting conditions. It also uses advanced algorithms to compensate for gaze tracking errors caused by head movement, greatly improving the accuracy and stability of gaze tracking.

[0071] The central processing unit's attention analysis module generates an attention intensity index based on gaze trajectory information, accurately quantifying the user's attention state. The dynamic guidance module generates an initial interaction plan containing multi-level guidance strategies based on this index and action response data, and dynamically adjusts each strategy parameter based on the user's real-time data, achieving personalized and dynamic interactive guidance and effectively improving user engagement and interactive experience. The evaluation and correction module generates a first correction coefficient and a second correction coefficient by comparing the deviation between the user's actual body movement parameters and the preset movement model, as well as the difference between the attention intensity index and the expected threshold, providing a quantitative basis for strategy optimization. The strategy optimization module combines the two correction coefficients to generate a composite evaluation index, and accordingly generates a differentiated guidance plan, further improving the adaptability and effectiveness of the guidance strategy and ensuring that the user's attention is focused on interacting with the robot.

[0072] The introduction of the adaptive learning module enables the system to self-optimize. By establishing a user behavior database, employing a temporal convolutional neural network to analyze user attention decay cycles, and applying a reinforcement learning algorithm to iteratively optimize guidance strategy parameters, the system is able to continuously adapt to the behavioral habits and attentional characteristics of different users, continuously improving the accuracy and effectiveness of the guidance strategy. Furthermore, the adaptive learning module constructs a virtual twin model and conducts adversarial training and transfer learning, further enhancing the robustness and generalization of the guidance strategy. This makes it applicable to users of different age groups, significantly expanding the application scenarios and user base of toy robots, and greatly improving the market competitiveness and user experience of toy robots. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] The present invention is further described with reference to the accompanying drawings. However, the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative effort.

[0074] Figure 1 It is a structural block diagram of the system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0075] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0076] In a first aspect of the disclosed embodiments, a dynamic sight tracking and attention guiding system for a toy robot is provided. Figure 1 As shown, it includes a user interaction end and a central processing end;

[0077] The user interaction terminal at least includes:

[0078] A visual acquisition module is configured to capture the user's eye movement data in real time to generate gaze trajectory information, and is also configured to collect action response data when the user interacts with the robot;

[0079] A behavior perception module is configured to obtain body movement parameters of the user when following the robot's guidance instructions through a multi-axis inertial sensor;

[0080] A first communication unit is configured to transmit the gaze trajectory information, action response data and limb movement parameters to a central processing end;

[0081] The central processing end at least includes:

[0082] a second communication unit configured to receive multimodal data from a user interaction terminal;

[0083] an attention analysis module configured to calculate, based on the gaze trajectory information, the duration of the user's gaze focus remaining in the robot's guidance area and the frequency of its shifts, and generate an attention intensity index;

[0084] a dynamic guidance module configured to generate an initial interaction plan including a multi-level guidance strategy based on the attention intensity indicator and the action response data, wherein the guidance strategy includes a robot motion trajectory, an acoustic and optical feedback mode, and a task difficulty gradient, wherein the strategy parameters are dynamically adjusted based on the real-time user data;

[0085] An evaluation and correction module is configured to compare the deviation between the actual body movement parameters of the user when executing the initial interaction plan and the preset movement model to generate a first correction coefficient; and simultaneously generate a second correction coefficient based on the difference between the attention intensity index and the expected threshold;

[0086] The strategy optimization module is configured to fuse the first correction coefficient and the second correction coefficient to generate a composite evaluation index, and generate a differentiated guidance plan based on the index. The plan includes: adjusting the robot's movement speed, increasing the intensity of tactile feedback, or reconstructing the task sequence to improve the user's attention focus level.

[0087] As a preferred embodiment, the visual acquisition module integrates an infrared imaging unit and works in conjunction with a visible light camera. It is configured to automatically switch to infrared mode when the ambient light illumination is lower than 50 lux, and uses a pupil-corneal reflection vector algorithm to compensate for the line of sight tracking error caused by head movement.

[0088] As a preferred embodiment, the behavior perception module includes:

[0089] A nine-axis MEMS sensor array configured to capture the angular velocity and linear acceleration of the user's limb movements at a 100Hz sampling rate;

[0090] A pressure-sensitive surface layer is configured to cover the robot's touchable surface and quantify the user's touch force and frequency through piezoelectric signals;

[0091] The sound field positioning unit is configured to recognize the azimuth and emotional characteristics of the user's voice commands based on the beamforming technology of the microphone array.

[0092] As a preferred embodiment, the dynamic guidance module includes:

[0093] The motion planning submodule is configured to dynamically adjust the complexity of the robot's movement trajectory according to the attention intensity index, and trigger a spiral progressive motion mode when the index value is lower than the threshold;

[0094] The multimodal feedback submodule is configured to synchronously drive the RGBW four-color LED array, micro vibration motor and directional speaker to generate compound sensory stimulation that matches the user's attention state.

[0095] As a preferred embodiment, the system further includes an adaptive learning module configured to perform the following operations:

[0096] Establish a user behavior database to store historical gaze trajectory patterns, action response delays, and task completion accuracy;

[0097] Use a temporal convolutional neural network to analyze the user's attention decay cycle and predict the best time to intervene;

[0098] The guidance strategy parameters are iteratively optimized through the reinforcement learning algorithm, so that the second correction coefficient is reduced to the preset threshold within three consecutive interaction cycles.

[0099] As a preferred embodiment, the adaptive learning module is further configured as follows:

[0100] Build a virtual twin model to simulate the typical attention characteristics of users of different age groups;

[0101] During the offline training phase, adversarial training is performed between actual user data and virtual model data to generate a robust guidance policy library;

[0102] In the online application stage, the policy library parameters are dynamically adapted to the current user characteristics based on transfer learning technology.

[0103] A second aspect of the present disclosure provides a method for dynamic eye tracking and attention guidance of a toy robot, comprising the following steps:

[0104] The steps include:

[0105] S1. Use visual sensors to collect real-time eye movement data from the user and generate gaze trajectory information including gaze location and duration;

[0106] S2. Capturing user body motion parameters through an inertial sensor array, the parameters including at least motion amplitude, motion frequency, and response delay time;

[0107] S3. Transmitting the gaze trajectory information and limb movement parameters to the central processing end, and calculating the attention intensity index based on the distribution characteristics of the gaze focus in the preset guidance area;

[0108] S4. Generate an initial interactive guidance plan based on the attention intensity indicator, wherein the plan defines the robot's motion path, multimodal sensory feedback mode, and stage task sequence, wherein the difficulty of each stage task is dynamically adjusted based on the user's real-time data;

[0109] S5. During the user's initial interactive guidance program, the deviation between the actual body movement parameters and the preset movement model is compared to generate a first correction coefficient; and a second correction coefficient is generated based on the degree to which the attention intensity indicator deviates from the expected threshold.

[0110] S6. Fusing the first correction coefficient and the second correction coefficient to generate a composite evaluation index. When the index exceeds a preset threshold, reconstructing at least one parameter group in the initial interactive guidance scheme, the parameter group including: extending the robot's dwell time on the motion path, increasing the intensity gradient of the acoustic and optical feedback, or inserting an auxiliary tactile prompt task;

[0111] S7. Send the reconstructed interactive guidance plan to the robot for execution, and repeat steps S1 to S6 until the user's attention intensity index stabilizes in the target range.

[0112] As a preferred embodiment, the generation of the first correction coefficient in step S5 includes:

[0113] The generation of the first correction coefficient in step S5 includes:

[0114] Extract the three-dimensional spatial motion trajectory from the user's body movement parameters and perform dynamic time-warping matching with the preset motion model;

[0115] Calculating a weighted sum of standard deviations of the actual trajectory and the model trajectory in terms of velocity, acceleration, and angular dimensions, and mapping the weighted sum to a range of 0-1 to generate a first correction coefficient;

[0116] The generation of the second correction coefficient in step S5 includes:

[0117] The baseline reference value is calculated based on the standard deviation of the fixation duration and the peak of the shift frequency in the user's historical attention data, and the two are averaged;

[0118] Load the preset weight matrix according to the current task type and generate dynamic upper and lower thresholds;

[0119] Calculate in real time the standardized difference between the attention intensity index and the historical mean, where the standardized difference is the difference between the current value and the historical mean divided by the historical standard deviation;

[0120] Calculating a time series fluctuation entropy value, where the entropy value is calculated based on a distribution probability of at least ten consecutive frames of attention intensity within a dynamic threshold interval;

[0121] When the normalized difference exceeds a preset range, generating a linear correction component proportional to the absolute value of the difference;

[0122] When the fluctuation entropy value exceeds a preset threshold, a nonlinear correction component is generated using a hyperbolic tangent function;

[0123] Perform weighted average of the linear correction component and the nonlinear correction component, and superimpose a time decay factor to smooth coefficient fluctuations;

[0124] When the second correction coefficient exceeds 0.5, it is determined that the user has entered a state of persistent inattention.

[0125] As a preferred embodiment, the step S3 further includes:

[0126] S3a within the camera field of view of the robot vision module, establish a three-dimensional coordinate system centered on the characteristic marker point, the characteristic marker point includes a reflective marker worn by the user;

[0127] S3b real-time calculation of the projection coordinates of the feature marker points in the image plane, based on the camera optical distortion parameters to build a spatial angular position mapping model, the model includes a radial distortion compensation factor and a tangential distortion compensation factor;

[0128] S3c. When the feature marker deviates from the preset tracking area, the compensation action is calculated as follows:

[0129] Extract the ellipse fitting contour of the marked point in the current frame image, and calculate the azimuth offset Δθ based on the angle between the major axis direction and the preset standard direction;

[0130] Based on the ratio change rate of the marker area to the preset reference area and the camera focal length parameter, the radial distance offset Δd is calculated;

[0131] S3d. The Δθ and Δd are input to the motion compensation controller to generate control instructions for the robot chassis movement or pan / tilt rotation, so that the feature markers are regressed to the standardized trajectory within the range of ±5% pixels in the center area of the image;

[0132] S3e. When the cumulative displacement of the compensation instructions for three consecutive frames exceeds the preset threshold, the dynamic recalibration mode is triggered:

[0133] Drive the camera to perform multi-angle swing scanning to re-establish the spatial position constraint relationship between the feature markers and the robot body;

[0134] The target motion trajectory is predicted based on the Kalman filter, and the distortion compensation factor in the spatial angular position mapping model is updated.

[0135] As a preferred embodiment, the feature marker point extension in step S3a includes non-reflective user native graphic feature points, and includes the following correction calculation steps:

[0136] S3f. Based on the salient features of the user's facial contour or clothing pattern, dynamic reference points are extracted in real time through a convolutional neural network to generate a hybrid marker topology associated with the reflective markers;

[0137] S3g based on the actual imaging position of each feature point in the mixed marker topology, combined with the camera distortion model to reverse the theoretical space coordinates, calculate the coordinate deviation caused by the distortion Δe, and mapped to the third correction coefficient in the range of 0-1;

[0138] S3h. In step S6, the third correction coefficient is introduced into the composite evaluation index to generate the expanded attention loss behavior coefficient:

[0139] When Δe exceeds 0.3 for five consecutive frames, it is determined to be an active behavior of attention deviation, triggering the robot's reverse movement compensation;

[0140] When the product of Δe and the second correction coefficient is greater than 0.2, the acoustic and optical feedback is enhanced and the visual sampling frequency is increased.

[0141] In the embodiment of the present disclosure, step S1 further includes:

[0142] When the ambient light intensity is lower than 50 lux, the infrared fill light device is activated, and the pupil-corneal reflection vector compensation algorithm is used to eliminate the line of sight tracking error caused by the user's head deviation;

[0143] When it is detected that the user blinks continuously at a preset frequency, which in the embodiment of the present disclosure exceeds 3 times per second, the anti-misjudgment mechanism is triggered, the current eye trajectory analysis is paused and the backup action guidance mode is started.

[0144] The generation of the multimodal sensory feedback pattern in step S4 includes:

[0145] The preset feedback combination strategy is matched according to the numerical range of the attention intensity index. When the index value is in the low attention range, the flashing frequency of the robot's LED breathing light and the buzzer prompt tone are synchronously enhanced;

[0146] When the matching degree between the user's limb movement parameters and the preset model exceeds a preset threshold, the robot joint vibration motor is activated to generate a tactile reward signal.

[0147] Compared with the prior art, the embodiments of the present disclosure have the following beneficial effects:

[0148] The dynamic eye tracking and attention guidance system and method for a toy robot described in the embodiments of the present disclosure have significant beneficial effects. First, the visual acquisition module and the behavior perception module at the user interaction end work together to comprehensively and accurately obtain the user's eye movement data, action response data, and limb movement parameters through multimodal data acquisition, providing a rich and accurate data foundation for subsequent attention analysis and guidance strategy generation. Among them, the visual acquisition module integrates an infrared imaging unit and a visible light camera, can work stably under different lighting conditions, and uses advanced algorithms to compensate for the eye tracking error caused by head movement, greatly improving the accuracy and stability of eye tracking.

[0149] The central processing unit's attention analysis module generates an attention intensity index based on gaze trajectory information, accurately quantifying the user's attention state. The dynamic guidance module generates an initial interaction plan containing multi-level guidance strategies based on this index and action response data, and dynamically adjusts each strategy parameter based on the user's real-time data, achieving personalized and dynamic interactive guidance and effectively improving user engagement and interactive experience. The evaluation and correction module generates a first correction coefficient and a second correction coefficient by comparing the deviation between the user's actual body movement parameters and the preset movement model, as well as the difference between the attention intensity index and the expected threshold, providing a quantitative basis for strategy optimization. The strategy optimization module combines the two correction coefficients to generate a composite evaluation index, and accordingly generates a differentiated guidance plan, further improving the adaptability and effectiveness of the guidance strategy and ensuring that the user's attention is focused on interacting with the robot.

[0150] The introduction of the adaptive learning module enables the system to self-optimize. By establishing a user behavior database, employing a temporal convolutional neural network to analyze user attention decay cycles, and applying a reinforcement learning algorithm to iteratively optimize guidance strategy parameters, the system is able to continuously adapt to the behavioral habits and attentional characteristics of different users, continuously improving the accuracy and effectiveness of the guidance strategy. Furthermore, the adaptive learning module constructs a virtual twin model and conducts adversarial training and transfer learning, further enhancing the robustness and generalization of the guidance strategy. This makes it applicable to users of different age groups, significantly expanding the application scenarios and user base of toy robots, and greatly improving the market competitiveness and user experience of toy robots.

[0151] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to also include plural forms. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of one or more associated listings. In addition, when used in this application, the term "comprise" and its variations "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups of these. In the absence of further restrictions, an element defined by the sentence "comprising a..." does not exclude the presence of other identical elements in the process, method or device that includes the element. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments can be referenced to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can be found in the description of the method part.

[0152] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these effects are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods for each specific application to implement the described effects, but such implementations should not be considered to exceed the scope of the embodiments of the present disclosure. The technicians will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0153] The flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of the code, and the module, program segment or part of the code contains one or more executable instructions for implementing the specified logical functions. In some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by dedicated hardware-based devices that perform the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A dynamic gaze tracking and attention guidance system for a toy robot, comprising a user interaction terminal and a central processing terminal; The user interaction terminal at least includes: A visual acquisition module is configured to capture the user's eye movement data in real time to generate gaze trajectory information, and is also configured to collect action response data when the user interacts with the robot; A behavior perception module is configured to obtain body movement parameters of the user when following the robot's guidance instructions through a multi-axis inertial sensor; A first communication unit is configured to transmit the gaze trajectory information, action response data and limb movement parameters to a central processing end; The central processing end at least includes: a second communication unit configured to receive multimodal data from a user interaction terminal; an attention analysis module configured to calculate, based on the gaze trajectory information, the duration of the user's gaze focus remaining in the robot's guidance area and the frequency of its shifts, and generate an attention intensity index; a dynamic guidance module configured to generate an initial interaction plan including a multi-level guidance strategy based on the attention intensity indicator and the action response data, wherein the guidance strategy includes a robot motion trajectory, an acoustic and optical feedback mode, and a task difficulty gradient, wherein the strategy parameters are dynamically adjusted based on the real-time user data; An evaluation and correction module is configured to compare the deviation between the actual body movement parameters of the user when executing the initial interaction plan and the preset movement model to generate a first correction coefficient; and simultaneously generate a second correction coefficient based on the difference between the attention intensity index and the expected threshold; The strategy optimization module is configured to fuse the first correction coefficient and the second correction coefficient to generate a composite evaluation index, and generate a differentiated guidance plan based on the index. The plan includes: adjusting the robot's movement speed, increasing the intensity of tactile feedback, or reconstructing the task sequence to improve the user's attention focus level.

2. The dynamic sight tracking and attention guiding system for a toy robot according to claim 1, characterized in that: The visual acquisition module integrates an infrared imaging unit and works in conjunction with a visible light camera. It is configured to automatically switch to infrared mode when the ambient light illumination is lower than 50 lux, and uses a pupil-corneal reflection vector algorithm to compensate for the line of sight tracking error caused by head movement.

3. The dynamic sight tracking and attention guiding system for a toy robot according to claim 1, characterized in that: The behavior perception module includes: A nine-axis MEMS sensor array configured to capture the angular velocity and linear acceleration of the user's limb movements at a 100Hz sampling rate; A pressure-sensitive surface layer is configured to cover the robot's touchable surface and quantify the user's touch force and frequency through piezoelectric signals; The sound field positioning unit is configured to recognize the azimuth and emotional characteristics of the user's voice commands based on the beamforming technology of the microphone array.

4. The dynamic sight tracking and attention guiding system for a toy robot according to claim 1, characterized in that: The dynamic guidance module includes: The motion planning submodule is configured to dynamically adjust the complexity of the robot's movement trajectory according to the attention intensity index, and trigger a spiral progressive motion mode when the index value is lower than the threshold; The multimodal feedback submodule is configured to synchronously drive the RGBW four-color LED array, micro vibration motor and directional speaker to generate compound sensory stimulation that matches the user's attention state.

5. The dynamic sight tracking and attention guiding system for a toy robot according to claim 1, characterized in that: Also included is an adaptive learning module configured to perform the following operations: Establish a user behavior database to store historical gaze trajectory patterns, action response delays, and task completion accuracy; Use a temporal convolutional neural network to analyze the user's attention decay cycle and predict the best time to intervene; The guidance strategy parameters are iteratively optimized through the reinforcement learning algorithm, so that the second correction coefficient is reduced to the preset threshold within three consecutive interaction cycles.

6. The dynamic sight tracking and attention guiding system for a toy robot according to claim 1, characterized in that: The adaptive learning module is further configured to: Build a virtual twin model to simulate the typical attention characteristics of users of different age groups; During the offline training phase, adversarial training is performed between actual user data and virtual model data to generate a robust guidance policy library; In the online application stage, the policy library parameters are dynamically adapted to the current user characteristics based on transfer learning technology.

7. A method for dynamic sight tracking and attention guidance of a toy robot, characterized in that: The steps include: S1. Use visual sensors to collect real-time eye movement data from the user and generate gaze trajectory information including gaze location and duration; S2. Capturing user body motion parameters through an inertial sensor array, the parameters including at least motion amplitude, motion frequency, and response delay time; S3. Transmitting the gaze trajectory information and limb movement parameters to the central processing end, and calculating the attention intensity index based on the distribution characteristics of the gaze focus in the preset guidance area; S4. generating an initial interactive guidance plan based on the attention intensity indicator, the plan defining the robot motion path, multimodal sensory feedback mode and stage task sequence; S5. During the user's initial interactive guidance program, the deviation between the actual body movement parameters and the preset movement model is compared to generate a first correction coefficient; and a second correction coefficient is generated based on the degree to which the attention intensity indicator deviates from the expected threshold. S6. Fusing the first correction coefficient and the second correction coefficient to generate a composite evaluation index. When the index exceeds a preset threshold, reconstructing at least one parameter group in the initial interactive guidance scheme, the parameter group including: extending the robot's dwell time on the motion path, increasing the intensity gradient of the acoustic and optical feedback, or inserting an auxiliary tactile prompt task; S7. Send the reconstructed interactive guidance plan to the robot for execution, and repeat steps S1 to S6 until the user's attention intensity index stabilizes in the target range.

8. The method for dynamic eye tracking and attention guidance of a toy robot according to claim 7, further comprising the following steps: generating the first correction coefficient in step S5 comprises: Extract the three-dimensional spatial motion trajectory from the user's body movement parameters and perform dynamic time-warping matching with the preset motion model; Calculating a weighted sum of standard deviations of the actual trajectory and the model trajectory in terms of velocity, acceleration, and angular dimensions, and mapping the weighted sum to a range of 0-1 to generate a first correction coefficient; The generation of the second correction coefficient in step S5 includes: The baseline reference value is calculated based on the standard deviation of the fixation duration and the peak of the shift frequency in the user's historical attention data, and the two are averaged; Load the preset weight matrix according to the current task type and generate dynamic upper and lower thresholds; Calculate in real time the standardized difference between the attention intensity index and the historical mean, where the standardized difference is the difference between the current value and the historical mean divided by the historical standard deviation; Calculating a time series fluctuation entropy value, where the entropy value is calculated based on a distribution probability of at least ten consecutive frames of attention intensity within a dynamic threshold interval; When the normalized difference exceeds a preset range, generating a linear correction component proportional to the absolute value of the difference; When the fluctuation entropy value exceeds a preset threshold, a nonlinear correction component is generated using a hyperbolic tangent function; Perform weighted average of the linear correction component and the nonlinear correction component, and superimpose a time decay factor to smooth coefficient fluctuations; When the second correction coefficient exceeds 0.5, it is determined that the user has entered a state of persistent inattention.

9. The method for dynamic sight tracking and attention guidance of a toy robot according to claim 8, characterized in that: Said S3 further comprises: S3a within the camera field of view of the robot vision module, establish a three-dimensional coordinate system centered on the characteristic marker point, the characteristic marker point includes a reflective marker worn by the user; S3b real-time calculation of the projection coordinates of the feature marker points in the image plane, based on the camera optical distortion parameters to build a spatial angular position mapping model, the model includes a radial distortion compensation factor and a tangential distortion compensation factor; S3c. When the feature marker deviates from the preset tracking area, the compensation action is calculated as follows: Extract the ellipse fitting contour of the marked point in the current frame image, and calculate the azimuth offset Δθ based on the angle between the major axis direction and the preset standard direction; Based on the ratio change rate of the marker area to the preset reference area and the camera focal length parameter, the radial distance offset Δd is calculated; S3d. The Δθ and Δd are input to the motion compensation controller to generate control instructions for the robot chassis movement or pan / tilt rotation, so that the feature markers are regressed to the standardized trajectory within the range of ±5% pixels in the center area of the image; S3e. When the cumulative displacement of the compensation instructions for three consecutive frames exceeds the preset threshold, the dynamic recalibration mode is triggered: Drive the camera to perform multi-angle swing scanning to re-establish the spatial position constraint relationship between the feature markers and the robot body; The target motion trajectory is predicted based on the Kalman filter, and the distortion compensation factor in the spatial angular position mapping model is updated.

10. The method for dynamic sight tracking and attention guidance of a toy robot according to claim 9, characterized in that: The feature marker point extension in step S3a includes non-reflective user native graphic feature points, and includes the following correction calculation steps: S3f. Based on the salient features of the user's facial contour or clothing pattern, dynamic reference points are extracted in real time through a convolutional neural network to generate a hybrid marker topology associated with the reflective markers; S3g based on the actual imaging position of each feature point in the mixed marker topology, combined with the camera distortion model to reverse the theoretical space coordinates, calculate the coordinate deviation caused by the distortion Δe, and mapped to the third correction coefficient in the range of 0-1; S3h. In step S6, the third correction coefficient is introduced into the composite evaluation index to generate the expanded attention loss behavior coefficient: When Δe exceeds 0.3 for five consecutive frames, it is determined to be an active behavior of attention deviation, triggering the robot's reverse movement compensation; When the product of Δe and the second correction coefficient is greater than 0.2, the acoustic and optical feedback is enhanced and the visual sampling frequency is increased.

Citation Information

Patent Citations

  • Automatic tracking system based on XR space

    CN119722746A

  • Virtual reality system to promote reading comprehension and writing skills in students

    DE202025101195U1

  • Systems and methods for observing eye and head information to measure ocular parameters and determine human health status

    US20220133212A1

Cited By

  • Guide AR glasses interaction method and system based on gesture recognition

    CN121050583A