Gaze-guided teleoperation of robotic manipulators
The use of gaze-tracking smart glasses with filtering and inverse kinematics algorithms enhances direct control over robotic manipulators, improving user interaction efficiency and accuracy by addressing the limitations of existing gaze-controlled systems.
Patent Information
- Application Number
- PCT/US2025/046432
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-13
- Filing Date
- 2025-09-15
- Publication Date
- 2026-03-19
AI Technical Summary
Existing gaze-controlled robotic systems face challenges in accurately and directly controlling robotic manipulators due to coupled sensing and actuation systems, limited gaze degrees of freedom, and instability from rapid eye movements, leading to inefficient and confusing user interactions.
A method and system utilizing a wearable device like smart glasses for gaze tracking, combined with gaze filtering techniques and inverse kinematics algorithms, enabling direct control over robotic manipulators through gaze and voice commands, allowing for smoother and safer robot operation.
The proposed solution reduces user workload by 37.51% and lowers gripper positioning error by 39.09%, providing precise and intuitive control over robotic manipulators without the need for additional visual cues or external object detection.
Smart Images

Figure US2025046432_19032026_PF_FP_ABST
Abstract
Description
PATENT APPLICATIONDocket No.: 055632-00044GAZE-GUIDED TELEOPERATION OF ROBOTIC MANIPULATORSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 694,325, titled “GAZE-GUIDED TELEOPERATION OF ROBOTIC MANIPULATORS” and filed September 13, 2024, the contents of which are incorporated by reference herein in their entirety.FIELD
[0002] The technology described herein generally relates to gaze-guided teleoperation of robotic manipulators, and more particularly to modeling methods and model-based control for direct gaze guided teleoperation to command a robot.BACKGROUND
[0003] Owing to the rapid advancements in the field of Robotics and Artificial Intelligence (Al), assistive robotic systems have proven to be instrumental in addressing several healthcare problems for the elderly, children, and persons with disabilities. Examples of such robotic systems are involved in rehabilitation, assistance in drinking and walking, detection and prevention of falls, and telepresence caregiving. Object manipulation is one task required in certain assistive systems, including for aiding people with Parkinson's disease, quadriplegia, and muscular dystrophy. Gaze is a suitable candidate that can be continuously tracked to eliminate the need for arms and can provide an interface for users with the aforementioned disabilities to interact with a robotic manipulator and command the robot to pick up a desired object in the environment and place it at a target location. However, using gaze to control robots can be very challenging because the user's sensing and actuation systems are coupled, and gaze data has lower degrees of freedom than the task space of the robot. Moreover, rapid movements of the eyes (known as a “saccade”) can easily destabilize the robot controllers.
[0004] While conventional techniques and systems attempt to provide solutions, they eliminate direct control of the robot's end effector (e.g., position-to-position matching between the gaze commands and the robot). In such implementations, gaze fixations, operations zones, external camera systems for object detection, and multi-step task methods execution are often used.
[0005] There is a need for improved systems and methods that enables accurate teleoperation of a robot.PATENT APPLICATIONDocket No.: 055632-00044SUMMARY
[0006] At a high level, embodiments of the technology are generally directed towards methods and systems related to teleoperation of a robot or robot components. In some aspects, systems and methods for providing a modeling and control strategy for the gaze-guided teleoperation of a robot or robotic manipulators are provided. In some aspects, gaze-guided teleoperation incorporates the use of a wearable device, such as a head-mounted device / display (HMD), for example smart glasses. In some aspects, systems and methods are provided for model identification and validation for gaze tracking in the task space of a robot via gaze tracking devices. In some aspects, systems and methods are provided for gaze filtering techniques that enable a user direct control over a robotic manipulator, such as position. In some aspects, systems and methods are provided for the direct control over a robotic manipulator based on one or more inputs (e.g. gaze, voice command) via a wearable device (e g., a headset). In some aspects, direct control over a robot in a task space can implement one or more of a user interface, a gaze tracker, hand-eye coordination, gaze command filtering, inverse kinematics, and a low-level controller.
[0007] In accordance with a first aspect of the disclosure, a method for gaze-guided control of a robot is provided. An example method includes receiving, from a headset, data associated with a gaze of a wearer of the headset. The example method further includes generating at least one motion profile based at least in part on the data associated with the gaze of the wearer and a gaze filter selection. The example method further includes generating at least one command by applying an inverse kinematics algorithm to the at least one motion profile. The example method further includes transmitting the at least one command to at least one manipulator of the gaze- controlled robot. The at least one command triggers the at least one manipulator to move based at least in part on the at least one command.
[0008] In some embodiments of the example method, the gaze filter selection includes a selection of a no filter method, a selection of a moving average filter method, or a selection of a low-pass filter method.
[0009] In some embodiments of the example method, the headset includes at least one sensor, where the at least one sensor includes at least a first sensor that captures at least one voice command, and the method further includes receiving the voice command from the headset, where the at least one command is generated based at least in part on the at least one voicePATENT APPLICATIONDocket No.: 055632-00044 command.
[0010] In some embodiments of the example method, the gaze-controlled robot includes a gripper. In some such embodiments, the at least one command triggers the at least one manipulator to move to position the gripper to a position based at least in part on the data associated with the gaze of the wearer.
[0011] In some embodiments of the example method, generating the at least one motion profile includes applying an algorithm that generates motion data based at least in part on the data associated with the gaze of the wearer; and applying the gaze filter selection to the motion data to generate the at least one motion profile.
[0012] In some embodiments of the example method, the method further includes outputting, to the headset, an interface includes a gaze position indicator positioned in the interface based at least in part on gaze position indicator data.
[0013] In accordance with a second aspect of the disclosure, an apparatus for gaze-guided control of a robot is provided. An example apparatus includes at least one processor and at least one memory having computer-coded instructions stored thereon. The computer-coded instructions, upon execution by the at least one processor, cause the apparatus to perform the steps of any one of the example methods discussed herein.
[0014] In accordance with a third aspect of the disclosure, a system for gaze-guided control of a robot is provided. An example system includes a headset comprising at least one sensor, wherein the headset is configured to output data associated with a gaze of a wearer of the headset based at least in part on the at least one sensor. The system further includes a gaze-controlled robot, wherein the gaze-controlled robot includes at least one manipulator configured receive at least one command. The system further includes a gaze-controlled robot-instructing system. The gaze-controlled robot-instructing system is configured to receive the data associated with the gaze of the wearer from the headset; generate at least one motion profile based at least in part on the data associated with the gaze of the wearer and a gaze filter selection; generate the at least one command by applying an inverse kinematics algorithm to the at least one motion profile; and transmitting the at least one command to the at least one manipulator of the gaze-controlled robot, wherein the at least one command triggers the at least one manipulator to move based at least in part on the at least one command.
[0015] In accordance with a fourth aspect of the disclosure, a non-transitory computer-PATENT APPLICATIONDocket No.: 055632-00044 readable storage medium for gaze-guided control of a robot is provided. An example non- transitory computer-readable storage medium includes computer-coded instructions stored thereon that, when executed by at least one processor, are configured for performing the steps of any one of the example methods discussed herein.
[0016] Additional features and embodiments are further described in the detailed description which follows.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Many aspects of the present disclosure will be better understood with reference to the following drawings. The components in the drawings are not necessarily to scale, with emphasis instead being placed upon clearly illustrating the principles of the disclosure. Moreover, in the drawings, like reference numerals designate corresponding parts throughout the several views. It should be recognized that these implementations and embodiments are merely illustrative of the principles of the present disclosure. Therefore, in the drawings:
[0018] FIG. 1 A illustrates a block diagram of an example gaze-guided robotic control strategy in accordance with at least some embodiments of the present disclosure;
[0019] FIG. IB illustrates a block diagram of an example apparatus in accordance with at least some embodiments of the present disclosure;
[0020] FIG. 2 illustrates a diagram of a hand-eye coordination algorithm and communication between a headset and a robot in accordance with at least some embodiments of the present disclosure;
[0021] FIG. 3 illustrates a diagram of an algorithm for processing an identification target and the controlled reference patterns in accordance with at least some embodiments of the present disclosure;
[0022] FIGS. 4A, 4B, and 4C illustrates example motion profiles of the robot and the tracked gaze positions, specifically for various references, in accordance with at least some embodiments of the present disclosure;
[0023] FIG. 5 illustrates components of an example reference tracking model for a gaze in accordance with at least some embodiments of the present disclosure;
[0024] FIG. 6 illustrates example model validation profiles for a gaze model in accordance with at least some embodiments of the present disclosure;
[0025] FIG. 7 illustrates depictions of a gaze-guided robot at various stages throughoutPATENT APPLICATIONDocket No.: 055632-00044 performance of an example pick-up and place task in accordance with at least some embodiments of the present disclosure;
[0026] FIGS. 8A-8D illustrates a depiction of results of a performed case study in accordance with at least some embodiments of the present disclosure;
[0027] FIG. 9 illustrates example robot trajectories controlled by gaze using a particular filter method in accordance with at least some embodiments of the present disclosure; and
[0028] FIG. 10 illustrates an example process for gaze-guided robot control in accordance with at least some embodiments of the present disclosure.DETAILED DESCRIPTION
[0029] The presently disclosed subject matter now will be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all embodiments of the presently disclosed subject matter are shown. Like numbers refer to like elements throughout. The presently disclosed subject matter may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements.
[0030] Indeed, many modifications and other embodiments of the presently disclosed subject matter set forth herein will come to mind to one skilled in the art to which the presently disclosed subject matter pertains having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the presently disclosed subject matter is not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims.
[0031] Throughout this specification and the claims, the terms “comprise,” “comprises”, and “comprising” are used in a non-exclusive sense, except where the context requires otherwise. Likewise, the term “includes” and its grammatical variants are intended to be non-limiting, such that recitation of items in a list is not to the exclusion of other like items that can be substituted or added to the listed items.
[0032] The present invention can be understood in conjunction with embodiments thereof as illustrated in the attached Appendix. However, it should be evident that many alternatives, modifications and variations can be produced. In particular, these embodiments are merely illustrative of the principles of the present invention. Accordingly, this disclosure is not intended to embrace all such alternatives, modifications and variations that fall within the spirit and broadPATENT APPLICATIONDocket No.: 055632-00044 scope of the Appendix and in view of the claims.
[0033] All publications, patents and patent applications mentioned in this specification are herein incorporated in their entirety by reference into the specification, to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated herein by reference. In addition, citation or identification of any reference in this application shall not be construed as an admission that such reference is available as prior art to the present invention. To the extent that section headings are used, they should not be construed as necessarily limiting.
[0034] Embodiments described herein can be understood more readily by reference to the following detailed description, examples, and figures. Elements, apparatus, and methods described herein, however, are not limited to the specific embodiments presented in the detailed description, examples, and figures. It should be recognized that the exemplary embodiments herein are merely illustrative of the principles of the invention. Numerous modifications and adaptations will be readily apparent to those of skill in the art without departing from the scope of the disclosure.
[0035] In addition, all ranges disclosed herein are to be understood to encompass any and all subranges subsumed therein. For example, a stated range of "1.0 to 10.0" should be considered to include any and all subranges beginning with a minimum value of 1.0 or more and ending with a maximum value of 10.0 or less, e.g., 1.0 to 5.3, or 4.7 to 10.0, or 3.6 to 7.9.
[0036] All ranges disclosed herein are also to be considered to include the end points of the range, unless expressly stated otherwise. For example, a range of "between 5 and 10," "from 5 to 10," or "5-10" should generally be considered to include the end points 5 and 10.
[0037] Disjunctive language such as the phrase "at least one of X, Y, or Z," unless specifically stated otherwise, is otherwise understood with the context as used in general to present that an item, term, etc., may be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y, or at least one of Z to each be present.
[0038] Further, when the phrase "up to" is used in connection with an amount or quantity, it is to be understood that the amount is at least a detectable amount or quantity. For example, a material present in an amount "up to" a specified amount can be present from a detectable amount and up to and including the specified amount.I. Example Use CasePATENT APPLICATIONDocket No.: 055632-00044
[0039] In one example use case, embodiments are provided for gaze-guided control of a robot. An example system in accordance with the present disclosure is provided. The example system includes a wearable device, for example a headset. The headset includes at least one sensor and, via the sensor, tracks a user’s (e.g., the wearer’s gaze). In this regard, the headset may determine, and be able to output, data representing or otherwise associated with the wearer’s gaze, for example data that indicates where in a plane of an environment a user is looking at a particular time. Data representing the gaze, or that may be used to derive a user’s gaze, may be referred to as “gaze data.”
[0040] The example system further includes a gaze-controlled robot. The gaze-controlled robot may be any robot that is moveable based on commands (e.g., instructions, messages, requests, and / or the like) transmitted to the robot or a subcomponent thereof based on gaze data from the headset. The gaze-controlled robot may include at least one manipulator, which moves and / or repositions at least part of the robot. For example, a command sent to the gaze-controlled robot may trigger the manipulator to move in a particular manner that repositions a component of the robot (or the whole robot) to a location represented by the gaze data.
[0041] The example system further includes a gaze-controlled robot-instructing system. The gaze-controlled robot-instructing system serves as an intermediary processing system for controlling the gaze-controlled robot based on gaze data collected via a headset. For example, the gaze-controlled robot-instructing system may be embodied by a specially configured computer, terminal, server, and / or the like, such as any combination of hardware, software, and / or firmware that executes one or more software applications that are configured to provide the functionality described herein. In some embodiments, the gaze-controlled robot-instructing system receives data associated with the gaze of the wearer from the headset (e.g., “gaze data”). The gaze- controlled robot-instructing system generates at least one motion profile based at least in part on the data and a gaze filter selection, for example a determined or predetermined selection for using no filter on the data, a moving average filter, and / or a low-pass filter, as discussed herein. The at least one motion profile may then be utilized to generate at least one command by applying an inverse kinematics algorithm to the at least one motion profile. The inverse kinematics algorithm generates low-level commands for controlling the gaze-controlled robot, or a particular component (e g., at least one manipulator) thereof, in accordance with the gaze of the wearer of the headset. The gaze-controlled robot-instructing system then transmits the at leastPATENT APPLICATIONDocket No.: 055632-00044 one command to the at least one manipulator of the gaze-controlled robot (e g., directly or via another component of the gaze-controlled robot) to trigger the at least one manipulator to move in accordance with the user’s gaze, for example based on the at least one command. In this regard, the user’s gaze is utilized to control the gaze-controlled robot, for example to reposition one or more components thereof to particular positions and / or to accomplish any one of a myriad of tasks described herein.II. Systems and Methods
[0042] An example system is provided for gaze-guided control of a robot. The system includes a headset, for example that may be worn by a user. The headset may have any number of sensors utilized to detect and / or provide any number of data parameters defining a user’s gaze, and / or positioning data associated with a gaze-controlled robot. The system further includes a gaze-controlled robot. The gaze-controlled robot is controlled based at least in part on gaze data outputted and / or derived based on data from the headset. Additionally, in some embodiments the gaze-controlled robot is controlled based at least in part on voice data and / or other command data captured via the headset (e g., voice commands spoken by a user during operation. The system further includes a gaze-controlled robot-instructing system. The gaze-controlled robot-instructing system includes hardware, software, firmware, and / or any combination thereof, to perform the modeling and / or model-based gaze-filtering of gaze data (or other data) received from the headset. Additionally, or alternatively, the gaze-controlled robot-instructing system in some embodiments includes hardware, software, firmware, and / or any combination thereof, to generate instructions that command a gaze-controlled robot to move in accordance with the gaze data from the headset.
[0043] Embodiments of the present disclosure are configured to perform one or more methods described herein. An example method in some embodiments includes receiving, at a gaze- controlled robot-instructing system, headset data associated with a gaze of a user. The method further includes applying, by the gaze-controlled robot-instructing system, an algorithm that generates motion data based at least in part on the headset data. The method further includes applying, by the gaze-controlled robot-instructing system, a gaze filter selection to the motion data to generate at least one motion profile. The gaze filter selection may be one of no filter, a moving average filter, or a low-pass filter. The method further includes applying, by the gaze-controlled robot-instructing system, an inverse kinematics algorithm to the at least one motion profile to generate at least one command to control at least one manipulator of a gaze-guided robot. ThePATENT APPLICATIONDocket No.: 055632-00044 method further includes transmitting, by the gaze-controlled robot-instructing system, the at least one command to the gaze-guided robot, wherein transmitting the at least one command triggers movement of the at least one manipulator of the gaze-guided robot in accordance with the at least one command.
[0044] FIG. 10 depicts an example process for gaze-guided robot control. Specifically, FIG. 10 depicts an example process 1000. In some embodiments, the process 1000 is performed by a specially configured device, system, and / or apparatus. For example, in some embodiments the process 1000 is performed by a gaze-controlled robot-instructing system embodied by the apparatus 150 as depicted and described herein with respect to FIGS. 1A and IB.
[0045] At operation 1002, process 1000 includes receiving, from a headset, data associated with a gaze of a wearer of the headset. In some embodiments, the headset includes one or more sensors that detect the data associated with the gaze of the wearer. The data may embody “gaze data” representing or otherwise indicating a location within an environment that the wearer is looking at a given time. Additionally, or alternatively, in some embodiments, the headset is configured to receive voice data (e.g., voice commands) that are further processable for controlling at least a component of the gaze-controlled robot.
[0046] At operation 1004, process 1000 includes generating at least one motion profile based at least in part on the data associated with the gaze of the wearer and a gaze filter selection. In some embodiments, the data associated with the gaze of the wearer is processed utilizing a handeye coordination algorithm and / or the gaze filter selection. The gaze filter selection may be a predetermined or determined indicator of a filter to be applied to the data. For example, in some embodiments the gaze filter selection indicates one of applying no filter, a moving average filter, or a low-pass filter.
[0047] At operation 1006, process 1000 includes generating at least one command by applying an inverse kinematics algorithm to the at least one motion profile. The inverse kinematics algorithm may generate movement data that is interpretable or otherwise processable by the gaze- controlled robot, or manipulator thereof, to effectuate a movement in accordance with the user’s gaze. For example, the manipulator may be triggered via the at least one command to move at least one component of the gaze-controlled robot to a position indicated by or otherwise consistent with the gaze data.
[0048] At operation 1008, process 1000 includes transmitting the at least one command to atPATENT APPLICATIONDocket No.: 055632-00044 least one manipulator of the gaze-controlled robot. The at least one command triggers the at least one manipulator to move based at least in part on the at least one command. The at least one command may be transmitted wirelessly over a network (e.g., over Wi-Fi, Bluetooth, radio frequency transmission, and / or other wireless protocol), via a wired connection (e.g., USB, Ethernet, and / or the like), or the like. The command may trigger the manipulator to move in accordance with the at least one command, for example to be consistent with repositioning the gaze-controlled robot (or a component thereof) to a particular position consistent with the wearer’ s gaze.III. With Reference to the Figures
[0049] Object manipulation is a high-frequency task required in assistive systems in order to aid the elderly or those with disabilities that impact motor control. In the instance where arms cannot be used to command a robot, gaze-tracking via smart glasses is a suitable option. Embodiments are described herein that use gaze-tracking as a solution, specifically including a modeling method and model-based filtering and control strategy for direct gaze-guided teleoperation of robotic manipulators. An object manipulation case study with six participants was performed, with results discussed further herein. The results indicate that a model-based gaze filtering and control strategy produces smoother and safer commands for the robot that are easy for the participants to use. These methods can reduce the workload of the user, for example determined to be by 37.51%, and / or lower the gripper positioning error, for example determined to be by 39.09%, compared to using unfiltered gaze data in other implementations.
[0050] In the field of Robotics and Artificial Intelligence (Al), assistive robotic systems have proven to be instrumental in addressing several healthcare problems for the elderly, children, and persons with disabilities. Examples of such robotic systems include rehabilitation, assistance in drinking and walking, detection and prevention of falls, and telepresence caregiving. Object manipulation is one of the high-frequency tasks required in assistive systems, including for aiding people with Parkinson’s disease, quadriplegia, and muscular dystrophy (e.g. for the elderly population that will be doubled by 2050 according to the United State Census Bureau). Gaze may be continuously tracked to eliminate the need for arms and can provide an interface for users with the aforementioned disabilities to interact with a robotic manipulator and command the robot to pick up a desired object in the environment and place it at a target location. However, using gaze to control robots can be very challenging because the user’s sensing and actuation systems arePATENT APPLICATIONDocket No.: 055632-00044 coupled, and gaze data has lower degrees of freedom than the task space of the robot. Moreover, rapid movements of the eyes, known as a saccade, can easily destabilize the robot controllers. Implementations are desired that resolve existing problems with gaze-controlled systems, and specifically that allows the user to freely move the desired objects to any location in the feasible task space, does not include the extra delays for gaze fixation and zone selection, enables a higher resolution for motion control without the need for cursor and additional visual cues, and does not rely on external object detection algorithms since the user is directly choosing the desired objects by moving the robot
[0051] In some systems, a user can fixate their gaze on an object and the target point will be fed to an underlying control algorithm that moves the end effector of the manipulator arm to the desired position. While existing implementations allow the user to operate the arm for grasping tasks, such implementations do not provide the user with any direct control over the robot’s motion. Another common approach is the use of command zones to complete teleoperation tasks. This approach provides zones, sometimes centered around the robot’s end effector, that generate movement commands once a user’s gaze falls in the zone. Once the gaze is in the zone, the end effector of the manipulator arm moves in the corresponding direction by a certain pre-defined measure (e.g., 2cm). Systems using this approach allow the user to generate their own path for robot motion but do not allow for diagonal motions via gaze input. Further in this case, a user must switch their gaze to a different command zone to send the robot in a different direction.
[0052] Additionally, or alternatively, in some embodiments, velocity thresholds on the eye movement are used to filter the gaze commands. Some implementations also provide a mode that allows the user to teleoperate a manipulator directly with their gaze. The user can look in a direction and have the end effector mimic the motion in its task space. In this control scheme, the velocity of the manipulator arm is determined by the magnitude of the corresponding gaze vector. In some implementations, to filter out the saccades, moving averages, weighted moving averages, and / or Kalman filters are used. For the user to stop the motion, they must return their gaze back to the center of their view. While this design allows for diagonal motions, it can cause unnecessary confusion for the user. If the user continues to fixate on the target object after the end effector arrives at the target, the end effector will continue to move.
[0053] Part of using gaze as an input for continuous teleoperation is the filtering out of saccades. Depending on the control scheme, some implementations use either a moving average,PATENT APPLICATIONDocket No.: 055632-00044 weighted moving average, or Kalman Filter to filter out these motions. Some systems use a velocity threshold on the eye movement. Using gaze positions to send velocity commands to a robot with limited workspace requires command resetting to prevent drift, which can make the usage difficult. Furthermore, the velocity map regions used in these strategies may cause a mismatch between the gaze and the actual position of the robot, which can be confusing to the user.
[0054] The use of augmented reality devices in robotics may be used in some implementations. Such devices enable a portable head and gaze tracking device in the task space, which removes the need for a fixed screen and workstation. Some systems have made robot programming more accessible. Furthermore, they provide the use of holographic manipulator arms to plan out their starting and ending configurations. The user can then see the proposed motion plan and determine if they wish to accept it. These devices have also been used to teleoperate robots. Some implementations provide the user information about the state of a manipulator arm and allow the implementations to pick a target location via a head-gaze cursor and then execute the plan via voice commands.
[0055] Although these implementations provide various solutions to assist the user with a gaze-guided motion control method, they eliminate direct control of the robot’s end effector (e.g., position-to-position matching between the gaze commands and the robot). In such solutions, gaze fixations, operations zones, external camera systems for object detection, and multi-step task methods execution are used, while the raw gaze commands are not studied and utilized deeply. Despite being technically more challenging, direct teleoperation allows the user to freely move the desired objects to any location in the feasible task space, does not include extra delays for gaze fixation and zone selection, enables a higher resolution for motion control without the need for cursor and additional visual cues, and does not rely on external object detection algorithms since the user is directly choosing the desired objects by moving the robot. Furthermore, by further identifying the strengths and weaknesses of pure teleoperation, higher levels of autonomy (e.g., assisted teleoperation and shared control with human initiative) can be implemented more efficiently and accurately.
[0056] Embodiments discussed further below include a modeling method and a model-based control strategy for direct gaze-guided teleoperation of robotic manipulators to command the robot to pick up a desired object in the environment and place it at a target location. Discussions belowPATENT APPLICATIONDocket No.: 055632-00044 include development of embodiments including a model identification and validation method for gaze tracking in the task space of the robot via smart glasses gaze-tracking devices. Additionally, discussed below, in some embodiments such models are utilized to design gaze filtering and control techniques that allow a user to directly control the position of a robotic manipulator in a manner that is easy to use by the user. A case study of human subjects to examine the effectiveness of the developed strategies is further discussed herein.
[0057] FIG. 1 A depicts the block diagram of an example gaze-guided robotic control strategy. A robot equipped with a gripper is provided as an example robot for gaze-guided control, where an augmented or mixed reality headset is utilized to provide gaze-guided control and output. As depicted, FIG. 1 includes a robot in a task space 102. The robot in the task space 102 is depicted via a user interface 104, for example where the user interface 104 is displayed on a headset. The headset and user interface in some embodiments are configured to implement gaze tracker 106. The gaze tracking is utilized to perform hand-eye coordination algorithm 108. The gaze-guided robotic control further includes gaze command filtering 110, which is performed based on the results from the hand-eye coordination algorithm 108 (e.g., based on gaze data from the headset associated with the user interface 104). Further inverse kinematics 112 (e.g., an “inverse kinematics algorithm”) converts the high-level motion profiles from the gaze command filtering 110 to low-level robot motions. The low-level robot motions are processed by low-level controller 114, for example to control the robot in the task space 102. This cycle may continue for any number of movements of the robot in the task space 102.
[0058] Specifically, a Universal Robot UR5e (Universal Robots, Odense, Denmark) robot (e.g., a manipulator) equipped with a Robotiq 2F-85 gripper (Robotiq, Le vis, QC) is used for object manipulation. A Microsoft HoloLens 2 headset (Microsoft, Redmond, WA) is used for the user interface and commanding the robot, for example via gaze-guided control. This mixed-reality device is equipped with several sensors, including four visible light cameras and an Inertial Measurement Unit (IMU) for head tracking, two Infrared (IR) cameras for eye tracking, RGB cameras, and a 5-channel Microphone array. It should be appreciated that in other implementations and / or for other headsets, other sensors and / or sensor arrangements may be leveraged, which obtain the same and / or different data for processing. These sensors and features allow us to track the gaze and head, as well as receive voice commands for controlling the robot operating state (e.g., movement, opening, and / or closing the gripper). Moreover, the holographic display in somePATENT APPLICATIONDocket No.: 055632-00044 embodiments has a gaze pointer that is shown to the user. The gaze pointer provides a means of feedback to the user for calibration, identification, and control of the robot. In some embodiments, the headset (and / or associated control device) is connected to the robot (and / or associated control device) via a wireless access point. A gaze command filtering algorithm along with a hand-eye coordination logic algorithm produces the high-level motion profiles that are converted to low- level robot motions via inverse kinematics and robot drivers. The control strategies were realized via several computer vision and control programs, for example implemented in C++ language and implemented at a 30 Hz control loop frequency. The overall system is integrated via a gaze- controlled robot-instructing system, for example a Robot Operating System (ROS) Noetic on Ubuntu 20.04 LTS.
[0059] FIG. IB depicts a block diagram of an example apparatus, specifically apparatus 150, for example embodying a gaze-controlled robot-instructing system. In this regard, the example apparatus depicted and described with respect to FIG. IB may be the gaze-controlled robotinstruction system described with respect to FIG. 1A, for example where the apparatus receives data, and / or transmits data to, a headset and / or a robot under control. The apparatus 150 may be specially configured to perform the operations of one or more processes for gaze-guided robot control, as described herein.
[0060] As illustrated in FIG. 1A, the apparatus 150 includes a processor 152, a memory 154, input / output (“I / O”) circuitry 156, communications circuitry 158, data sensing circuitry 160, gaze processing circuitry 162, and robot instructing circuitry 164. The apparatus 150 may be configured to execute some or all of the operations described herein with respect to gaze-guided robot control, specifically to use gaze-guided data from a headset to control and operate a connected robot. Although the circuitry 152-164 are described with respect to functional limitations, it should be understood that the particular implementations necessarily include the use of particular hardware. It should also be understood that certain of these components 152-164 may include similar or common hardware. For example, two modules may both leverage use of the same processor, network interface, storage medium, or the like, to perform their associated functions, such that the duplicate hardware is not required for each individually named circuitry.
[0061] The use of the term “circuitry” as used herein with respect to the components of the apparatus will be understood to include particular hardware configured to perform the functions associated with the particular module circuitry depicted and described. The term “circuitry”PATENT APPLICATIONDocket No.: 055632-00044 should be understood broadly to include hardware, software that configures the hardware, firmware that configures the hardware, and / or any combination thereof. For example, in some embodiments, “circuitry” may include processing circuitry, storage media, network interfaces, input and / or output devices, and the like. In some embodiments, other elements of the apparatus 150 may provide or supplement the functionality of particular circuitry. For example, in some embodiments, the processor 152 provides processing functionality, the memory 154 provides storage functionality, the communications circuitry 158 provides network interface functionality, and the like.
[0062] It should be appreciated that, in some embodiments, some or all of the circuitry may be associated with a separate device, server, and / or associated computing hardware, which may be in communication with one or more of the other circuitry components of the apparatus 150. For example, in some embodiments, the data sensing circuitry 160, gaze processing circuitry 162, and / or robot instructing circuitry 164 are included in and / or embodied by a separate computing apparatus. The separate computing apparatus in some embodiments may include a separate processor, memory, I / O circuitry, and / or communications circuitry.
[0063] In some embodiments, the processor 152 (and / or co-processors in some embodiments) is in communication with the memory 154 via a bus for passing information among components of the apparatus 150. The memory 154 may be non-transitory and in some embodiments includes one or more volatile and / or non-volatile memories. In other words, for example, the memory 154 in some embodiments is an electronic storage device (e.g., a computer-readable storage medium). The memory 154 in some embodiments is configured to store and / or provide access to data maintained by the apparatus 150, for example to enable the apparatus 150 to carry out various functions utilizing such data as described herein.
[0064] The processor 152 may be embodied in any of a myriad of different ways. For example, in some embodiments, the processor 152 includes one or more processing devices and / or sub-processors configured to perform independently. Additionally, or alternatively, in some embodiments the processor 152 may include one or more processors configured to operate in tandem via a bus, for example to enable independent execution of instructions, pipelining, and / or multithreading. The use of the terms “processing device,” “processor,” and / or “processing circuitry” may be understood to include a single core processor, a multi-core processor, multiple processors internal to the apparatus 150, and / or one or more separate, remote, and / or “cloud”PATENT APPLICATIONDocket No.: 055632-00044 processors.
[0065] In some embodiments, the processor 152 is configured to execute computer-coded instructions stored in the memory 154, or otherwise accessible to the processor 152.Alternatively, or additionally, the processor 152 in some embodiments is configured to execute hard-coded functionality. As such, whether configured by hardware or software, or by a combination of hardware with software, the processor 152 in some embodiments represents an entity (e.g., physically embodied in the circuitry) capable of performing operations in accordance with one or more embodiments of the present disclosure when configured accordingly. Alternatively, as another example, when the processor is embodied as an executor of software instructions, the computer-coded instructions in some embodiments specifically configure the processor to perform steps described here, for example embodying one or more algorithms and / or operations thereof.
[0066] In some embodiments, the apparatus 150 includes I / O circuitry 156 that may, in turn, be in communication with the processor 152 to provide output to the user associated with the apparatus 150. Additionally, or alternatively, in some embodiments, the I / O circuitry 156 is in communication with the processor 152 to receive input from a user. In some embodiments, the I / O circuitry 156 comprises a headset, for example an augmented or mixed-reality headset. Additionally, or alternatively, in some embodiments the I / O circuitry 156 in some embodiments comprises a user interface, for example a device display, web interface, mobile application, client device, and / or the like. Additionally, or alternatively, in some embodiments, the I / O circuitry 156 includes one or more input devices, for example a keyboard, a mouse, a joystick, a touch screen, a microphone, and / or input / output mechanisms. The I / O circuitry 156, alone or together with the processor 152, in some embodiments controls one or more functions of the user interface, for example through executing computer-coded instructions stored on the memory 154 or otherwise accessible to the processor 152 (e.g., embodied in software and / or firmware).The communications circuitry 158 in embodied in hardware, or a combination of both hardware and software, that is configured for data receiving and / or data transmission, for example over a network. The communications circuitry 158 may transmit data from the apparatus 150, and / or receive data at the apparatus 150 from another device, system, and / or the like. In some embodiments, the communications circuitry 158 includes at least one network card, at least one antenna, at least one switch, at least one router, at least one modem, at least one bus connectingPATENT APPLICATIONDocket No.: 055632-00044 components, and / or supporting hardware and / or software of any such components. In some embodiments, the communications circuitry 158 is a separate device that is configured to enable the apparatus 150 to perform such data receiving and data transmitting. For example, in some embodiments, the communications circuitry 158 is configured to interact with at least one antenna to facilitate signal transmission, and / or facilitate signal reception via the at least one antenna. In some embodiments, the apparatus 150 is configured to communicate via the communications circuitry 158 utilizing any communications protocol, or a combination of multiple communications protocol. Non-limiting examples of such communications protocols include Bluetooth Low Energy, infrared wireless communication, ultra-wideband communication, Wi-Fi, Near Field Communication, Worldwide Interoperability for Microwave Access, and / or the like.
[0067] The apparatus 150 in some embodiments includes data sensing circuitry 160. In some embodiments, the data sensing circuitry 160 includes hardware, software, firmware, and / or any combination thereof, configured to receive and / or collect data from a headset (e.g., a mixed- reality or other augmented reality headset). For example, in some embodiments, the data sensing circuitry 160 is specially configured to receive and / or store sensor data received from one or more sensors of a headset. Additionally, or alternatively, in some embodiments, the data sensing circuitry 160 is specially configured to receive and / or store processed data outputted from a headset. For example, in some embodiments, the data sensing circuitry 160 is configured to receive and / or store data representing a gaze of a user as outputted via a headset, where the headset processes sensor data to generate the data representing the gaze of the user. Additionally, or alternatively, in some embodiments, the data sensing circuitry 160 is specially configured to receive and / or store data representing relative and / or absolute movement of a headset, for example as a user moves their head to update their gaze within an environment. In some embodiments, the data sensing circuitry 160 includes one or more field programmable gate arrays, application-specific integrated circuits, specially-configured processors, and / or the like.
[0068] The apparatus 150 in some embodiments includes gaze processing circuitry 162. In some embodiments, the gaze processing circuitry 162 includes hardware, software, firmware, and / or any combination thereof, configured to update gaze data associated with a headset and corresponding robot. For example, in some embodiments, the gaze processing circuitry 162 is specially configured to maintain a virtual environment, for example within which an indicator orPATENT APPLICATIONDocket No.: 055632-00044 representation associated with a user’s gaze is depicted. Additionally, or alternatively, in some embodiments, the gaze processing circuitry 162 is configured to maintain a virtual environment including a virtual representation of a robot and / or other objects in an environment. Additionally, or alternatively, in some embodiments, the gaze processing circuitry 162 is configured to generate and / or otherwise maintain an augmented or mixed-reality interface for displaying via the headset. In some embodiments, the gaze processing circuitry 162 includes one or more field programmable gate arrays, application-specific integrated circuits, specially-configured processors, and / or the like.
[0069] The apparatus 150 in some embodiments includes robot instructing circuitry 164. In some embodiments, the robot instructing circuitry 164 includes hardware, software, firmware, and / or any combination thereof, configured to facilitate control of a particular robot. For example, in some embodiments, the robot instructing circuitry 164 is specially configured to generate robot control commands based on gaze data from the headset or derived based on data from the headset. Additionally, or alternatively, in some embodiments, the robot instructing circuitry 164 is configured to receive and / or monitor feedback representing an operational state of the robot. The robot instructing circuitry 164 in some embodiments receives and / or stores data associated with the robot for use in generating subsequent gaze-guided commands for the robot, and / or corresponding interfaces for displaying to a user (e.g., via a headset). In some embodiments, the robot instructing circuitry 164 includes one or more field programmable gate arrays, application-specific integrated circuits, specially-configured processors, and / or the like.
[0070] FIG. 2 depicts an example of human coordination and the communication between the headset and the robot via the system. Specifically, FIG. 2 depicts an example headset 206 that communicates with a robot 208. The headset 206 communicates with the robot 208 via the system 202. In some embodiments, the system 202 is embodied by the apparatus 150, as depicted and described herein. The system 202 maintains a first environment 204 A (e.g., maintained by one or more software applications) for interacting with the headset 206 (e.g., a HoloLens 2 headset) and a second environment 204B (e.g., maintained by one or more software applications) for interacting with the robot 208 and / or processing gaze data associated with the headset 206. In some embodiments, for example, the first environment 204A is used to collect and / or gather voice commands and / or gaze data that is provided to the second environment 204B for processing (e.g., to generate corresponding instructions for direct robot control). In some embodiments, the data isPATENT APPLICATIONDocket No.: 055632-00044 utilized to generate and / or provide (e.g., by rendering) an interface associated with control of the robot 208 via the headset 206. For example, as illustrated, the interface 250 includes a gaze position indicator 212 corresponding to the location of the user’s gaze via the headset 206. Additionally, or alternatively, as depicted, the interface 250 includes a digital representation or depiction 210, which corresponds to and depicts the robot 208. The gaze position indicator 212 and / or the digital representation 210 is depicted, positioned, oriented, and / or otherwise configured based on data captured via the headset 206 (e.g., headset orientation data, gaze data, and / or the like, and / or detected position, orientation, or other data associated with the robot 208 within the environment of the headset 206) and / or data captured via the robot 208 (e.g., headset state data, position data, orientation data, and / or the like).
[0071] In some embodiments, the gaze-controlled robot-instructing system embodied by the apparatus 150 executes one or more applications (e.g., embodied by computer-coded instructions stored to one or more memory and executable by one or more corresponding processor) that perform the gaze-guided robot control processes described herein. For instance, in some embodiments, for integration with the Microsoft HoloLens 2, an application using the Unity game engine and Microsoft’s Mixed Reality Toolkit (MRTK) may be used. The position of the user’s gaze as well as the position and orientation of the user’s head is extracted. The MRTK’s ability to process voice input is used to send voice commands to the robot. For example, in some embodiments, the user can say ’’Release” and ’’Grab” to control the opening and closing, respectively, of the Robotiq gripper. In some embodiments, to move the robot after actuating the gripper, the user says ’’Move”. This data is then packed and sent to the robot control machine using the ROS TCP Connector and ROS TCP Endpoint packages from the Unity Robotics Hub. These calculations happen within Unity’s ’’Update Loop” which is called with each frame of the application. In order to ensure that our data publishing remains frame-independent, we rely on the built-in Time.deltaTime in some implementations, which tracks the time in seconds since the last frame. This allows us to publish data (e.g., from the Unity engine) at a rate of lOhz.
[0072] To prevent singularities (e.g., overextended arm configurations) and collision with the table and robot’s body, in some embodiments a safety bounding box is defined to constrain the motion of the robot within -0.65m< y < -0.25m and 0.03m < z < 0.4m (e.g., along the y-z plane, with x = 0.13m). To track the gaze and convert it to the robot’s base frame, homogeneous transformations that are realized via the ROS transformation package may be used (e.g., Tj whichPATENT APPLICATIONDocket No.: 055632-00044 represents the transformation of coordinate frame / in z). Therefore, the gaze frame is converted to the base frame via the total transformation Tg — T^ Tg , where b, h, and g represent the base, head (e.g., HoloLens 2.0), and gaze frames respectively.
[0073] This transformation is used to determine the unit vector of the gaze frame in the robot’ s base frame and find the intersection of this vector with the task plane of the robot (e.g., along the y-z plane). This is done since the output of the device is a unit vector indicating the direction of the gaze (e.g., with no depth information). The projected points (e.g., pgyand pgz) are then used as the reference position of the gripper within the bounding box, which is then (with or without filtering) sent to the inverse kinematic solvers to control the robot, for example as depicted and discussed above with respect to FIG. 1 and FIG. 3. To track the head frame, a predefined position and orientation from the robot may be utilized as a starting point, and a measurement of the user height may be utilized to initialize the frame. In some embodiments, and in one or more of the experiments described herein, the relative motion of the head was then tracked during the experiment via the built-in sensors and the MRTK package (e.g., an IMU and the visible light cameras, and / or the like).
[0074] A reference tracking experiment designed to identify the dynamic models of the user gaze when using the MS HoloLens 2 smart glasses is discussed. The identified models reveal how the raw gaze data will command the robot and whether command shaping / filtering is advantageous before sending the final commands to the robot to move it in the task space. In order to accurately measure the position of the desired point for gaze tracking, we use an identification target which is held by the robot gripper, for example as depicted in FIG. 3. Specifically, FIG. 3 depicts a diagram of an algorithm for processing an identification target and the controlled reference patterns. As illustrated, an identification target 302 is presented, for example within a particular frame at least a portion of a robot (e.g., a gripper) that holds the identification target. At least one reference motion profile 304 is generated, for example based on position feedback data associated with the robot as described herein. A trajectory generator 306 in some embodiments processes the reference motion profile 304. The results of such processing may be processed via inverse kinematics 308, resulting in low-level commands that are implementable via the low-level controller 310 to move and / or otherwise operate the robot.
[0075] In some embodiments, a measurement of the distance from the center of the identification target (e.g., a crosshair / plus mark) to the robot flange is taken and included in thePATENT APPLICATIONDocket No.: 055632-00044 kinematic chain to track the position of the target. One or more motion profiles may then be used for the gaze model identification, the data collection process, and / or the identified models, as discussed further herein.
[0076] In some embodiments, two signals are used to complete the identification process. For example, in some embodiments, step-like and sinusoidal motion profiles are created for the identification target so that a human user can track them via the developed setup. FIGS. 4A, 4B, and 4C provide such example motion profiles. In FIGS. 4A, 4B, and 4C, the top and middle rows show the y and z coordinates versus time, respectively, and the bottom rows show the corresponding planar motion. For example, FIG. 4A depicts step-like motion profiles in the y- direction 402A, the z-direction 402B, and the z-y plane 402C. FIG. 4B depicts fast sinusoidal motion profiles in the y-direction 404A, the x-direction 404B, and the z-y plane 404C. FIG. 4C depicts slow sinusoidal motion profiles in the y-direction 406A, the x-direction 406B, and the z-y plane 406C. The sinusoidal commands in some embodiments represent cyclical and smooth changes in the reference point, and the step commands in some embodiments simulate rapid movements to desired locations. The sinusoidal commands are created based on the following model:
[0077] r, = a, sin (y tj + o,
[0078] where rt(i G y, z) is the reference position of the robot along y or z axis, Atis the amplitude, T is the period of oscillations (co =is the frequency), and o(- is the offset for thereference position to keep the robot motion symmetric compared to the center of the bounding box. Here we use Ay= 0.15m, Az= 0.15m, oy= -0.45m, and oz= 0.25m. We use three different values for T to produce, slow (T = 12.57 seconds, co = 0.5 rad / sec), moderate (T = 6.28 seconds, co = 1.0 rad / sec), and fast (T = 2.51 seconds, o = 2.5 rad / sec) reference profiles. In the fast condition, the speed of the end effector of the robot can reach up to 500 mm / sec, which is a considerably large speed for the bounding box and the task space. For the step-like references, a series of eight reference positions near the edges of the task space bounding box may be sent. The robot begins from the center of the bounding box (yc= -0.45m, zc= 0.25m) and travels to the other positions (e.g., 0.15m above, below, left, and right of the center) in a random order. For this step, in some embodiments the Reflexxes Motion library is used to produce the motion profiles for the robot. When the robot reaches the reference position, it pauses for a period of time (e g., fivePATENT APPLICATIONDocket No.: 055632-00044 seconds) before moving to the next position. For each reference profile, participants complete several 1-minute data collection tests. For the step response, in an example experiment participants did three rounds. The users did four tests for each sin wave period (e.g., only on y-motion, z- motion, diagonal motion, and so on). The robot drivers are able to reproduce these profiles with a repeatability of less than 0.1 mm.
[0079] Using the identification target, motion profiles, and the closed-loop controllers, a series of data collection tests may be performed to model the motion of the user’s gaze during the target tracking tasks.
[0080] In an example study, five participants (four males and one female between the ages of 20 and 26) took part in the identification tests. After a quick training about the goal of the identification task and the user interface, an initial calibration process was performed. Each participant was asked to sit comfortably on a chair positioned 2 meters away from the robot gripper. Once the user was seated, they were given the HoloLens 2 device and instructed on how to complete an eye-tracking calibration (e.g., the built-in eye-tracking calibration for the headset). Upon completion of the built-in calibration, the user was instructed to focus on an identification target (e.g., as depicted and described with respect to FIG. 3, for example)! to help remove any initial gaze offsets. The relative height of the headset visor was then measured as compared to the robot’s base and was used as the start position of the HoloLens 2 coordinate frame. The user was then allowed to start the Unity application, and the positioning of the robot was checked via the control machine (e.g., embodied by the apparatus 150 as depicted and described herein). Then, the UR5e robot moved in the desired pattern based on one or more motion profiles, for example as discussed above.
[0081] A total of 3,784.6 seconds of data, including the y and z coordinates of the target and the projected gaze on the task plane, was recorded during the identification tests. The correlation between the variables was checked to identify any coupling between the gaze motions. A strong correlation between the y coordinate of the target and the y coordinate of the gaze was identified (0.92). A similar pattern was observed for the z coordinate (0.91). The absolute values of the correlation between the y and z components were less than 0.005, indicating linear independence.
[0082] FIG. 5 depicts components of a reference tracking model for a gaze. Specifically, as illustrated in FIG. 5, a reference tracking model 500 (“model 500”) is depicted. With respect to FIG. 5, G represents gaze and R represents the reference.PATENT APPLICATIONDocket No.: 055632-00044
[0083] In some embodiments, the data (e.g., gaze and reference data) into training and validation categories (e.g., 80% for training and 20% for validation). Matlab System Identification Toolbox may be used to identify the dynamic model of the gaze motion when tracking the identification target. In some embodiments, a series of first and second-order transfer functions were used to identify dynamic models that explain the relationship between the position of the target as input and the gaze position as output. During the model identification, in some embodiments the best fits are obtained via the following first-order transfer function and continuous-time structure:
[0085] where G(s) represents the gaze (e.g., represented as gaze data) as output and R(s) represents the reference (e.g., the position of the identification target, represented by position data) as input. In the model 500, kprepresents a gain, idis the processing delay, and rpis a time constant for the first order model. The continuous time counterpart of this model is according to the following Equation 1 :
[0086] Equation 1 : G(t) = — (kpR(t — rd) — G(t))Tp
[0087] The continuous time counterpart of this model shows the closed-loop kinematic behavior of the gaze. This structure and the collected data obtained a model-fitting accuracy of 77.6% and 71.2% and a validation accuracy of 71.6% and 68.1% for the y and z axis, respectively.
[0088] FIG. 6 depicts an example of the model validation profiles for the gaze model. Specifically, FIG. 6 depicts a y model validation profile 602A and a z model validation profile 602B. The identified parameters include: kpy= 1.002, Tpy= 0.19, Tdy= 0.05, kpz= 1.003, Tpz= 0.24, and Tdz= 0.06. As an example for the y axis, with these parameters, equation 1 becomes G'y(t)~ 5.19 x (R(t - 0.05) - Gy(t)), which is a proportional control action done by an average user to control the velocity of the gaze to track the target. Based on this model, a user perceives the target’s motion (e.g., reference point shown by R(t)) with a small delay. The corresponding motion of the gaze (G' (t)) is then proportional to the difference between the perceived reference and the current gaze (G(t)). In a circumstance where R(t) reaches a steady state level, the gaze finally converges to the target point. Using the time constant Tpy= 0.19 as an example, it takes an average user in our identification test 0.19 seconds to reach 63% of the desired gaze distance. For example, for a 0.5-meter motion in the bounding box, the user gaze can travel 0.315 meters in 0.19PATENT APPLICATIONDocket No.: 055632-00044 seconds, an average speed of 1.66 meters per second. These linear models provide a reliable tool for predicting users’ behavior when freely controlling the robot via the gaze commands, including when the outputs and dynamics are not captured completely. Filters for filtering the gaze commands before sending them to the robot may be designed accordingly.
[0089] In the model identification step, the raw gaze commands can reach considerable speed levels, which include high-frequency components. In various contexts, when these commands are sent directly to the robot without any pre-filtering, the direct teleoperation of the robot becomes a difficult task. Two different filters for processing the raw gaze commands and sending them to the inverse kinematic solver for the robot to follow may be designed. In some embodiments, the estimated time constants Tpy= 0.19 and Tpz= 0.24 are used as a guideline to design filters that are an order of magnitude slower so that the speed of filtered commands are usable within the task space of the robot. The first filter uses a moving average method (“MA”) (0( / c) =—d)), where f0(k) I fi( ) is the filter output / input, k is the current time step, and m is the length of the observation window). In some embodiments, a 2-second observation window (e.g., a> = 60 based on the 30 Hz rate of the control loop). The second filter uses a low-pass filtering method (e.g., the transfer functionwhere o)cis the cut-off frequency). In somes+cocembodiments a cutoff frequency of 0.5 rad / sec is selected, which results in a time-constant of 2.0 seconds for this filter. In some embodiments, the filter is discretized via Matlab’s “c2d” command and is implemented it in the control loop.CASE STUDY
[0090] Three different gaze processing methods and analysis of their effectiveness are discussed further below. The effectiveness and discussion below is based on a case study performed for the three different gaze processing methods.
[0091] A pick-up and place task was performed in the case study. The pick-up and place task includes use of a gaze-controlled robot to move an object (e.g., a “target” object) from a starting position to an ending position (e.g., a “drop zone”), while avoiding an obstacle within the environment. FIG. 7 depicts a gaze-guided robot at various stages throughout performance of the pick-up and place task. Specifically, FIG. 7 depicts a gaze-guided robot in a first stage 702A (e.g., the gaze-guided robot is in a starting position in an environment with a target object, obstacle, andPATENT APPLICATIONDocket No.: 055632-00044 drop zone), a second stage 702B (e.g., the gaze guided robot is positioned above the target object with a gripper opened), a third stage 702C (e.g., the gaze-guided robot is positioned to close the gripper and engage the target object), a fourth stage 702D (e.g., the gaze-guided robot is positioned to be over the obstacle while holding the target object in the closed gripper), a fifth stage 702E (e.g., the gaze-guided robot is positioned at the drop zone with the gripper open to “drop” the target object at the drop zone), and a sixth and final stage 702F (e.g., the gaze-guided robot returns to the starting position, without holding the target object). This explains the overall procedure of the pick-up and place task via the system discussed herein, for example embodying a testbed for this task and the case study performed.
[0092] To remove the directional bias in the robot motion and commands, two pick-up and place locations around an obstacle may be placed. An obstacle is added to make sure the participants cannot move the target object horizontally to the drop-off location (e.g., to prevent a pure ID motion). For this task, a human user teleoperates the robot via the gaze commands to pick up an object from left or right, carry it over the obstacle, and finally drop it at the drop location right or left marked on the table. A tennis ball may be chosen for the target object, as was in the performed case study, and may similarly be chosen for the obstacle, so that both represent a uniform object of the same size, in the case study, the coordinates of each location measured in the robot base frame are pick-up / drop left: y = -55.66 cm, z = 3.55 cm, obstacle: y = -45.6 cm, z = 3.55 cm, and pick-up / drop right: y = -35.7 cm, z = 3.55 cm. Motion in the y - z plane (x= 13 cm) is controlled in order to move it safely and accurately to the target point in two directions (e.g., laterally towards the drop zone, and vertically to avoid collision with the obstacle). In each experiment, the start position of the robot tool was at y = -45 cm, z = 25 cm.
[0093] Three different control methods were tested in the performed case study. In the first method, no filters were used to control the robot (briefly named “NF”). Two different filters for processing the raw gaze commands and sending them to the inverse kinematic solver for the robot to follow were also tested. The first filter is the moving average method (“MA”) and the second filter is the low-pass filtering method (“LP”). Data regarding the right-to-left and left-to-right directions was included for each control method to remove the directional bias (e.g., resulting in 6 conditions total). To further remove forming bias, learning effects, and fatigue from user performance as the experiment reaches the end, a balanced Latin square may be used to determine a different experiment order for each participant. In the performed case study, a total of 36 trialsPATENT APPLICATIONDocket No.: 055632-00044 were conducted (e.g., 6 conditions per participant).
[0094] Objective data collected during the experiment includes i) a task completion time, ii) a combined pick-up and place positioning error (measured via the known locations of the points and the robot’s tool position when closing and opening the gripper), and iii) the velocity profiles of the robot. After each trial, the participants filled out the NASA task load index (“TLX”) survey to assess their perceived workload.
[0095] Six subjects (four males and two females between the ages of 18 and 38) participated in the study. All of the participants received a 15-minute orientation to understand the task at hand and the overall goal and procedure. Before a subject began using the real system for gaze-guided robot control, they completed a series of training steps. First, they were instructed on how to properly wear a particular headset (e.g., the HoloLens 2 device) and were guided through the built- in eye tracking calibration. After calibration, the participants practiced with a simulated version of the robot (via RVIZ in ROS) with no motion filtering to learn how the overall system and the directional commands work. To complete the training, the subjects were instructed on how to use the voice commands to open and close the gripper on the real system. Subjects were also informed on where to position the robot after completing an experiment.
[0096] Once the subject was finished with training, the real system tests began. The subject was instructed to sit in the testing position 2 meters away from the task plane, and system calibration was completed (e.g., user height measurement and frame initialization). System calibration includes updating the system’s user height parameter by measuring the height of the HoloLens 2 on the user’s head with respect to the base of the robot. In addition, the user is instructed to look in the middle of the open gripper so the delta between the gaze projection and tool position can be calculated and used in the control algorithms (around 0.055± 0.04m). Once the subject gave the signal to begin, control of the robot was handed over to them to perform the pick-and-place task. Upon completion of the pick and place, the subject signaled that they were finished, and the test administrator stopped the control algorithms. After each test condition, the subject was then given the NASA TLX survey to complete. The test process was repeated for all six test conditions.
[0097] A one-way analysis of variance (ANOVA) with Tukey post-hoc test was completed. FIGS. 8A-8D depict results of the performed case study. Specifically, FIGS. 8A-8D depict a completion time graph 802A for the three methods (non-filtered / NF, moving average / MA, andPATENT APPLICATIONDocket No.: 055632-00044 low-pass filtered / LP), a positioning error graph 802B for the three methods, a speed graph 802C for the three methods, and a perceived task load graph for the three methods 802D. The objective measurements of the NF test conditions of the third participant were discarded since this participant was not able to complete the task due to difficulties in using this specific control strategy.
[0098] Task completion time (seconds): The results did not indicate a significant difference between the task completion times under different control modes (F (2, 31) = 1.08 and p = 0.35). The overall task completion time is shorter in NF condition (M = 54.23, SD = 10.32) as compared to the filtered motion conditions MA (M = 61.74, SD = 17.45), and LP (M = 62.15, SD = 11.87) since the fast-moving components of the gaze motions are not filtered.
[0099] Pick-up and place positioning error (mm): A significant difference exists between different modes for the overall robot tool positioning error (F (2, 31) = 4.86 and p =< 0.05). The LP control method resulted in a lower positioning error (M = 12.28, SD = 5.08) as compared to NF (M = 20.16, SD = 6.2, p < 0.05) and MA (M = 20.20, SD = 9.12, p < 0.05).
[0100] Speed (mm / sec): The results show a significant difference between different modes (F (2, 61007) = 258.69 and p =< 0.001). The LP control method resulted in more controlled motion profiles via removing the higher frequency components of the motion (M = 18.4, SD = 0.04) as compared to MA (M = 21.06, SD = 0.06, p < 0.001) and NF (M = 33.26, SD = 0.1, p < 0.001). The average values of speed for NF and MA conditions were also significantly different (p < 0.001).
[0101] Perceived workload: The results indicated a significant difference between the perceived workload under different control modes (F (2, 33) = 5.36 and p < 0.01). The LP method resulted in a lower overall workload (M = 42.22, SD = 20.23) than the NF method (M = 67.56, SD = 13.52, p < 0.001). However, the workload of using the MA method (M = 48.33, SD = 24.1) was not significantly different than LP and NF.
[0102] FIG. 9 depicts the robot trajectories controlled by gaze using the low-pass (LP) condition, where a pick-up location is on the right. Specifically, FIG. 9 depicts a y-position graph 902A, a z-position graph 902B, and a z-y position graph 902C. The graphs depict various offsets between the gaze position and corresponding robot position.
[0103] In that the results indicate that control of a robot may be performed without using any filters. However, the perceived task load are higher for the user is generally higher in suchPATENT APPLICATIONDocket No.: 055632-00044 embodiments since it is easier to lose focus and cause rapid motions for the robot. The use of filters can lower the task load for the user by 37.51% in the case of the LP method in some embodiments. As another example, the LP filtering method can result in a 39.09% lower positioning error as compared to the no filter case. The MA method produces a larger filtering lag than the LP method since the measured samples are equally weighed. The NF method suffers from rapid motions of the robot tool, which makes the fine-tuning of the final positions harder. It should be appreciated that the user might lose focus during the ’’grab” and ’’release” steps while they wait for the robot to close / open the fingers. In these cases, the users’ gaze can drift, which results in a quick lateral or upward motion for the robot until the filters capture this part of the motion. Further modeling for the gaze drift may be used when the identification target is stationary for a long period of time.
[0104] In this regard, embodiments may implement any of the modeling method for tracking the user gaze in the task space of the robot as depicted and described herein, and dynamic models were discussed herein that explain the overall behavior of the test participants in target-tracking tasks. Gaze filters for commanding the robot to complete a task (e.g., a pick-and-place teleoperation task) are described herein, where such filters are developed using this model and the control strategy discussed herein. The results discussed herein indicate the feasibility and effectiveness of the methods. Furthermore, to make the methods more portable, an external positioning system or a marker detection algorithm (such as QR codes) may be added to track the user head more accurately over extended periods of time. Additionally, including switched dynamic models of the system in some embodiments helps reduce or eliminate the drifts during the grab and release steps for safer manipulation.IV. Implementations
[0105] Certain implementations of systems and methods consistent with the present disclosure are provided as follows:
[0106] Clause 1. A system comprising: a headset comprising at least one sensor, wherein the headset is configured to output data associated with a gaze of a wearer of the headset based at least in part on the at least one sensor; a gaze-controlled robot, wherein the gaze-controlled robot includes at least one manipulator configured receive at least one command; and a gaze-controlled robot-instructing system configured to: receive the data associated with the gaze of the wearer from the headset; generate at least one motion profile based at least in part on the data associatedPATENT APPLICATIONDocket No.: 055632-00044 with the gaze of the wearer and a gaze filter selection; generate the at least one command by applying an inverse kinematics algorithm to the at least one motion profile; and transmitting the at least one command to the at least one manipulator of the gaze-controlled robot, wherein the at least one command triggers the at least one manipulator to move based at least in part on the at least one command.
[0107] Clause 2. The system according to Clause 1, wherein the gaze filter selection comprises a selection of a no filter method, a selection of a moving average filter method, or a selection of a low-pass filter method.
[0108] Clause 3. The system according to any one of Clauses 1-2, wherein the at least one sensor comprises at least a first sensor that captures at least one voice command and transmits the voice command to the gaze-controlled robot-instructing system, and wherein the gaze- controlled robot-instructing system further generates the at least one command based at least in part on the at least one voice command.
[0109] Clause 4. The system according to any one of Clauses 1-3, wherein the gaze- controlled robot comprises a gripper.
[0110] Clause 5. The system according to Clause 4, wherein the at least one command triggers the at least one manipulator to move to position the gripper to a position based at least in part on the data associated with the gaze of the wearer.[0U1] Clause 6. The system according to any one of Clauses 1-5, wherein to generate the at least one motion profile, the gaze-controlled robot-instructing system is configured to: apply an algorithm that generates motion data based at least in part on the data associated with the gaze of the wearer; and apply the gaze filter selection to the motion data to generate the at least one motion profile.
[0112] Clause 7. The system according to any one of Clauses 1-6, wherein the headset is further configured to: output an interface comprising a gaze position indicator positioned in the interface based at least in part on gaze position indicator data.
[0113] Clause 8. A method comprising: receiving, from a headset, data associated with a gaze of a wearer of the headset; generating at least one motion profile based at least in part on the data associated with the gaze of the wearer and a gaze filter selection; generating at least one command by applying an inverse kinematics algorithm to the at least one motion profile; and transmitting the at least one command to at least one manipulator of the gaze-controlled robot,PATENT APPLICATIONDocket No.: 055632-00044 wherein the at least one command triggers the at least one manipulator to move based at least in part on the at least one command.
[0114] Clause 9. The method according to Clause 8, wherein the gaze filter selection comprises a selection of a no filter method, a selection of a moving average filter method, or a selection of a low-pass filter method.
[0115] Clause 10. The method according to any one of Clauses 8-9, wherein the headset comprises at least one sensor, wherein the at least one sensor comprises at least a first sensor that captures at least one voice command, the method further comprising: receiving the voice command from the headset, wherein the at least one command is generated based at least in part on the at least one voice command.
[0116] Clause 11. The method according to any one of Clauses 8-10, wherein the gaze- controlled robot comprises a gripper.
[0117] Clause 12. The method according to Clause 11, wherein the at least one command triggers the at least one manipulator to move to position the gripper to a position based at least in part on the data associated with the gaze of the wearer.
[0118] Clause 13. The method according to any one of Clauses 8-12, wherein generating the at least one motion profde comprises: applying an algorithm that generates motion data based at least in part on the data associated with the gaze of the wearer; and applying the gaze filter selection to the motion data to generate the at least one motion profile.
[0119] Clause 14. The method according to any one of Clauses 8-13, further comprising: outputting, to the headset, an interface comprising a gaze position indicator positioned in the interface based at least in part on gaze position indicator data.
[0120] Clause 15. An apparatus comprising at least one processor and at least one memory having computer-coded instructions stored thereon, wherein the computer-coded instructions, in execution with the at least one processor, cause the apparatus to perform the method according to any one of the Clauses 8-14.
[0121] Clause 16. A non-transitory computer-readable storage medium having computer- coded instructions stored thereon that, upon execution with at least one processor, are configured for performing the method according to any one of the Clauses 8-14.CONCLUSIONPATENT APPLICATIONDocket No.: 055632-00044
[0122] It should be emphasized that the above-described embodiments of the present disclosure are merely possible examples of implementations set forth for a clear understanding of the principles of the disclosure. Many variations and modifications may be made to the abovedescribed embodiment(s) without departing substantially from the spirit and principles of the disclosure. All such modifications and variations are intended to be included herein within the scope of this disclosure and protected by the following claims.
[0123] The present technology may be embodied as, among other things, a system, method, or computer product. Accordingly, embodiments may take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware. In one embodiment, the present invention takes the form of a computer program product that includes the computer-usable instructions embodied on one or more computer-readable media and executed by one or more processors. Some embodiments can be embodied as one or more, or combination of, computer software, a computer program, an application, one or more engines, one or more modules, and / or one or more hardware and / or software components. In some aspects, components of systems described herein can be employed as a distributed system or centralized system.
[0124] Computer-readable media includes volatile and nonvolatile media, removable and nonremovable media, and media readable by a database, a switch, and various other network devices. Network switches, routers, access points, and related components, in some instances, act as a means of communication within the scope of the technology. By way of example, computer- readable media comprise computer storage media and communications media.
[0125] Computer storage media or machine-readable media can include media implemented in any method or technology for storing and / or transmitting information or data. Examples of such information include computer-usable instructions, data elements, data structures, programs and program modules, and other data representations.
[0126] Communications media generally store computer-usable or readable instructions, including data structures and program modules in a modulated data signal. A modulated data signal, in some instances, can be understood to be a propagated signal that has one or more of its characteristics set or changed to encode information in the signal. Communications media include any information-delivery media. By way of example and not limitation, communications media include wired media, such as a wired network or direct-wired connection, and wirelessPATENT APPLICATIONDocket No.: 055632-00044 media, such as radio, cellular, spread-spectrum, and other wireless media technologies. Combinations of the above are included with the scope of computer-readable media and communications media.
[0127] Embodiments described herein can be understood more readily by reference to the examples described above. Elements, apparatus, and methods described herein, however, are not limited to any specific embodiment presented in the Examples. It should be recognized that these are merely illustrative of some principles of this disclosure and are non-limiting. Numerous modifications and adaptations will be readily apparent without departing from the spirit and scope of the disclosure.
[0128] Many different arrangements of the various components and / or steps depicted and described, as well as those not shown, are possible without departing from the scope of the claims below. Embodiments of the present technology have been described with the intent to be illustrative rather than restrictive. Alternative embodiments will become apparent from reference to this disclosure. Alternative means of implementing the aforementioned can be completed without departing from the scope of the claims below. Certain features and sub-combinations are of utility and can be employed without reference to other features and sub-combinations and are contemplated within the scope of the claims.
Claims
PATENT APPLICATIONDocket No.: 055632-00044CLAIMSTherefore, the following is claimed:
1. A system comprising: a headset comprising at least one sensor, wherein the headset is configured to output data associated with a gaze of a wearer of the headset based at least in part on the at least one sensor; a gaze-controlled robot, wherein the gaze-controlled robot includes at least one manipulator configured receive at least one command; and a gaze-controlled robot-instructing system configured to: receive the data associated with the gaze of the wearer from the headset; generate at least one motion profile based at least in part on the data associated with the gaze of the wearer and a gaze filter selection; generate the at least one command by applying an inverse kinematics algorithm to the at least one motion profile; and transmitting the at least one command to the at least one manipulator of the gaze- controlled robot, wherein the at least one command triggers the at least one manipulator to move based at least in part on the at least one command.
2. The system according to claim 1, wherein the gaze filter selection comprises a selection of a no filter method, a selection of a moving average filter method, or a selection of a low-pass filter method.
3. The system according to claim 1, wherein the at least one sensor comprises at least a first sensor that captures at least one voice command and transmits the voice command to the gaze- controlled robot-instructing system, and wherein the gaze-controlled robot-instructing system further generates the at least one command based at least in part on the at least one voice command.
4. The system according to claim 1, wherein the gaze-controlled robot comprises a gripper.
5. The system according to claim 4, wherein the at least one command triggers the at least one manipulator to move to position the gripper to a position based at least in part on the data associated with the gaze of the wearer.PATENT APPLICATIONDocket No.: 055632-000446. The system according to claim 1, wherein to generate the at least one motion profile, the gaze-controlled robot-instructing system is configured to: apply an algorithm that generates motion data based at least in part on the data associated with the gaze of the wearer; and apply the gaze filter selection to the motion data to generate the at least one motion profile.
7. The system according to claim 1, wherein the headset is further configured to: output an interface comprising a gaze position indicator positioned in the interface based at least in part on gaze position indicator data.
8. A method comprising: receiving, from a headset, data associated with a gaze of a wearer of the headset; generating at least one motion profile based at least in part on the data associated with the gaze of the wearer and a gaze filter selection; generating at least one command by applying an inverse kinematics algorithm to the at least one motion profile; and transmitting the at least one command to at least one manipulator of the gaze-controlled robot, wherein the at least one command triggers the at least one manipulator to move based at least in part on the at least one command.
9. The method according to claim 8, wherein the gaze filter selection comprises a selection of a no filter method, a selection of a moving average filter method, or a selection of a low-pass filter method.
10. The method according to claim 8, wherein the headset comprises at least one sensor, wherein the at least one sensor comprises at least a first sensor that captures at least one voice command, the method further comprising: receiving the voice command from the headset, wherein the at least one command is generated based at least in part on the at least one voice command.PATENT APPLICATIONDocket No.: 055632-0004411. The method according to claim 8, wherein the gaze-controlled robot comprises a gripper.
12. The method according to claim 11, wherein the at least one command triggers the at least one manipulator to move to position the gripper to a position based at least in part on the data associated with the gaze of the wearer.
13. The method according to claim 8, wherein generating the at least one motion profile comprises: applying an algorithm that generates motion data based at least in part on the data associated with the gaze of the wearer; and applying the gaze filter selection to the motion data to generate the at least one motion profile.
14. The method according to claim 8, further comprising: outputting, to the headset, an interface comprising a gaze position indicator positioned in the interface based at least in part on gaze position indicator data.
15. A non-transitory computer-readable storage medium having computer-coded instructions stored thereon that, in execution with at least one processor, are configured for: receiving, from a headset, data associated with a gaze of a wearer of the headset; generating at least one motion profile based at least in part on the data associated with the gaze of the wearer and a gaze filter selection; generating at least one command by applying an inverse kinematics algorithm to the at least one motion profile; and transmitting the at least one command to at least one manipulator of the gaze-controlled robot, wherein the at least one command triggers the at least one manipulator to move based at least in part on the at least one command.
16. The non-transitory computer-readable storage medium according to claim 15, wherein the gaze filter selection comprises a selection of a no filter method, a selection of a movingPATENT APPLICATIONDocket No.: 055632-00044 average filter method, or a selection of a low-pass filter method.
17. The non-transitory computer-readable storage medium according to claim 15, wherein the headset comprises at least one sensor, wherein the at least one sensor comprises at least a first sensor that captures at least one voice command, and the non-transitory computer-readable storage medium is further configured for: receiving the voice command from the headset, wherein the at least one command is generated based at least in part on the at least one voice command.
18. The non-transitory computer-readable storage medium according to claim 15, wherein the gaze-controlled robot comprises a gripper.
19. The non-transitory computer-readable storage medium according to claim 18, wherein the at least one command triggers the at least one manipulator to move to position the gripper to a position based at least in part on the data associated with the gaze of the wearer.
20. The non-transitory computer-readable storage medium according to claim 15, wherein generating the at least one motion profile comprises: applying an algorithm that generates motion data based at least in part on the data associated with the gaze of the wearer; and applying the gaze filter selection to the motion data to generate the at least one motion profile.