Method and system for gaze assisted interaction

By receiving and amplifying the gaze area, generating an interactive area, and assisting the interaction of the pointing device, the problem of the pointing device being difficult to interact quickly and with high precision on a large screen in the prior art, improving user experience and accuracy.

CN120390916APending Publication Date: 2025-07-29HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380087764.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-23
Filing Date
2023-12-18
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art is difficult to achieve fast and high-precision operations when interacting with the display screen using a pointing device, especially at large screen distances, and the transition from rough movement to fine control is difficult to achieve, and the accuracy of eye shaking and poor light conditions will affect the accuracy of gaze-assisted interaction.

Method used

By receiving the user's gaze point, extracting and amplifying the gaze area, generating interactive areas, and mapping to the cursor position on the display in the system hook, assisting in pointing to the device's interaction, reducing the impact of eye jitter and light conditions on accuracy.

Benefits of technology

It realizes fast and high-precision cursor movement at large screen distances, reduces user fatigue and error rates, and improves interaction accuracy under various lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390916A_ABST
    Figure CN120390916A_ABST
Patent Text Reader

Abstract

Methods and systems for gaze assisted interaction with a pointing device on a display screen. In one embodiment, in response to receiving an activation input, a point of gaze (POG) of a user on a display screen is received, and a gaze area of the display screen corresponding to the POG is extracted, magnified, and transposed on the display screen according to a first cursor position, thereby generating an interaction area on the display screen. User interaction with the pointing device at a second cursor position on the display screen associated with the interaction area is intercepted in a system hook, mapped to a position on the display screen corresponding to the gaze area, and passed to an application. The disclosed methods and systems may enable improved GUI interaction with a pointing device on a display screen while overcoming challenges associated with the accuracy of eye gaze assist interaction, including the impact of eye jitter on gaze estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This application claims the benefit and priority of U.S. Patent Application No. 18 / 146,183, filed on December 23, 2022, the content of which is incorporated herein by reference. Technical Field

[0002] The present invention relates to the field of human - computer interaction, and more particularly to methods and systems for gaze - assisted interaction. Background Art

[0003] Human - computer interaction explores how people interact with computers and computing devices. Devices that enable interaction between humans and computing devices are called human - machine interfaces. Pointing devices are commonly used human - machine interface devices that allow users to input spatial (i.e., continuous and multi - dimensional) data to a computer, such as by controlling a cursor on a display screen. As modern software and graphical user interfaces (GUIs) become more complex, for example, with the addition of an increasing variety of user - interface (UI) widgets, it becomes more challenging to manage the display - screen space for efficient human - computer interaction.

[0004] For common pointing devices (such as computer mice, touchpads, touchscreens, styli, joysticks, etc.), the motor input required to perform an interaction is proportional to the distance covered by the cursor on the screen. Over time, the physical size, resolution, and pixel density of display screens have increased, while the size of the content (such as text and icons) on the screen has become smaller. In addition, multi - screen configurations with extended display screens are becoming increasingly common. Therefore, it can be difficult to manipulate the cursor or pointer at a large display - screen distance, typically requiring multiple swipes on a touchpad or the user to lift the mouse and re - position it on a mouse pad between several movements. In other environments, such as in a vehicle that includes one or more display screens, it may be difficult for an operator to safely reach out to interact with the screen. In addition, the relatively smaller size of text or icons on the screen also poses challenges in terms of interaction precision and accuracy, for example, in cases where pixel - level positioning is required. For example, challenges arise when a user performs a quick and rough movement to rapidly manipulate the cursor of a pointing device at a large screen distance and then performs a slower and more precise movement to interact precisely with a smaller content item or object. In this regard, it can be difficult for users to perform both fast and precise interactions.

[0005] Gaze-assisted interaction adds the input of user eye gaze information to at least one other input mode (e.g., pointing device, keyboard, gesture recognition system, etc.) for interacting with a computing system. For example, assisting in controlling GUI elements, such as the cursor on a display screen; or assisting in viewing content items on the screen, such as by magnifying elements on the screen.

[0006] For example, U.S. Patent 6,204,828 (the entire content of which is incorporated herein by reference) discloses an interaction scenario including gaze tracking, where upon receiving a mechanical input (e.g., a mouse click), the cursor is repositioned on the screen to correspond to the gaze area on the screen. However, moving the cursor to the screen area where the user is gazing does not solve the problem of precisely controlling small GUI elements.

[0007] In another example, U.S. Patent 10,444,831 (the entire content of which is incorporated herein by reference) discloses a system that includes gaze and head tracking for hands-free device control, e.g., mode switching or cursor control via head movement. However, the drawbacks of this method include: using head tracking can cause jitter problems in pointer control. Additionally, performance may degrade in poor lighting conditions.

[0008] In another example, the following article discusses a method for content magnification based on gaze when interacting with a virtual environment: "Magnification Vision – a Novel Gaze-Directed User Interface" by Agledahl, S. and Steed, A. (the entire content of which is incorporated herein by reference), Proceedings of the IEEE Virtual Reality Conference and 3D User Interface Abstracts and Workshops (VRW), IEEE, 2021. However, the described method is limited to virtual environments and may not be generally applicable to real-world environments using common computing devices (e.g., laptops, desktops, or tablets) without including specific hardware. Additionally, the method does not meet the need for fast interaction at larger distances.

[0009] In another example, methods of text magnification are discussed in the following article: "Evaluation of a gaze-controlled vision enhancement system for reading in visually impaired people" by Aguilar, C. and Castet, E., which is incorporated herein by reference in its entirety, PLoS ONE 12.4 (2017): e0174910. However, the methods described are limited to a specific application where the user reads content on a screen and cannot be generally used for other GUI element interaction scenarios.

[0010] One drawback of current gaze-assisted interaction methods is that they do not account for small and rapid eye movements known as jitter. Although the human eye can move at high speeds, during such movement, the user's gaze direction constantly changes, even during periods of fixation, making it difficult to precisely control a pointing device through gaze tracking.

[0011] Accordingly, it would be useful to provide methods and systems for improving a user's interaction with a pointing device on a display screen. SUMMARY OF THE INVENTION

[0012] In various examples, the present invention describes methods and systems for improving user interaction with a pointing device on a display screen, e.g., using multiple input modes. Specifically, the user's interaction with the pointing device on the display screen can be assisted by input of eye gaze information. In response to receiving an activation input, a point of gaze (POG) of the user on the display screen is received. Based on a first cursor position on the display screen, a gaze region corresponding to the user's point of gaze is extracted, magnified, and transposed on the display screen to generate an interaction region. User interaction with a second cursor position on the display screen associated with the interaction region and the pointing device is intercepted in a system hook, mapped to a position on the display screen corresponding to the gaze region, and passed to an application. The disclosed methods and systems can achieve improved GUI interaction in a real-world environment (e.g., using multiple extended screens, or under a range of lighting conditions that may affect gaze tracking accuracy), while overcoming challenges associated with the accuracy of gaze-assisted interaction with a pointing device on a display screen, e.g., the impact of eye jitter on gaze estimation.

[0013] In various examples, the present invention provides the following technical effects: Performing operations using a pointing device on a computing system incorporates input from an additional source, namely eye gaze information. In this regard, gaze-assisted interaction aims to improve a user's interaction with a GUI on a display screen by addressing existing challenges of using a pointing device and a display screen.

[0014] In an example, instead of moving a pointing device to maneuver a cursor to desired display content, a user can gaze at the desired display content and activate a separator operation to initiate a gaze-assisted operation. In some examples, the present invention provides the following technical advantages: Spatial movement and interaction with a pointing device can be performed quickly and with high precision to manipulate a pointer (e.g., a cursor) in a GUI. For example, in a situation where a user needs to quickly move a pointing device over a large distance, such as moving a cursor to a screen target location far from the cursor origin and then slowing down the movement speed to precisely operate the pointing device when interacting with GUI elements (e.g., icons, buttons, text) at the screen target location. In this regard, quickly and precisely transitioning from a coarse movement to a fine movement can be useful in time-sensitive scenarios or real-time scenarios, such as when operating a machine or driving a vehicle, or when interacting with a video game, etc.

[0015] In some examples, the present invention provides the following technical advantages: Larger distance pointer movements or other manual operations (typically performed with a pointing device) can be performed more quickly and efficiently by leveraging the speed of eye gaze movement. Additionally, using eye gaze to assist with coarse movements rather than precise interaction with GUI objects can reduce challenges associated with eye jitter or gaze estimation accuracy.

[0016] In some examples, the present invention provides the following technical advantages: Reduces operator fatigue associated with coarse motor movements of a pointing device (e.g., maneuvering a cursor over a large screen distance). For example, gaze-assisted interaction enables the pointing device to be spared from making coarse movements or repetitive motor movements to maneuver a pointer over a large display screen distance, such as in an interaction where operation in a distal region of the screen is required.

[0017] In some examples, the present invention provides the following technical advantages: Improves the input accuracy of the pointing device by magnifying the content on the screen before initiating an input operation of the pointing device. In this regard, it may make interaction with the content on the screen easier and can reduce errors caused by a user selecting an incorrect object (e.g., due to difficulty in precisely manipulating the pointing device with a motor or inability to read small-sized content on the screen, etc.).

[0018] In some aspects, the present invention describes a method for gaze-assisted interaction with a pointing device on a display screen. The method includes a plurality of steps. The method includes: receiving an activation input for initiating a gaze-assisted operation corresponding to an application; in response to receiving the activation input, receiving a gaze point of a user on the display screen; in response to receiving the activation input, further receiving a first cursor position of the pointing device on the display screen; extracting a gaze region based on the gaze point on the display screen; generating an interaction region based on the gaze region and the first cursor position; receiving an interaction event corresponding to the interaction region at a second cursor position of the pointing device on the display screen; and in a system hook, generating a replacement interaction event for passing to the application, the replacement interaction event being generated in response to receiving the interaction event at the second cursor position.

[0019] In the above exemplary aspect of the method, generating the replacement interaction event includes: receiving an interaction mapping for associating a position of the gaze region on the display screen with a position of the interaction region on the display screen; mapping the second cursor position to a replacement cursor position based on the interaction mapping; and generating the replacement interaction event based on the replacement cursor position.

[0020] In the above exemplary aspect of the method, the method further includes: processing the replacement interaction event to perform a command operation of the application.

[0021] In the exemplary aspect of the method, extracting the gaze region includes: generating an image of a portion of the content of the display screen based on the gaze point.

[0022] In the above exemplary aspect of the method, generating the interaction region includes: magnifying the image of the portion of the content of the display screen to generate a magnified image; and generating the interaction region based on the magnified image and the first cursor position.

[0023] In the exemplary aspect of the method, a size of the image of the portion of the content of the display screen depends on an accuracy of the gaze point.

[0024] In the exemplary aspect of the method, receiving the activation input includes: receiving a delimiter operation input, the delimiter operation input including at least one of the following: pointing device input; keyboard event input; gesture input; or audio input.

[0025] In an exemplary aspect of the method, receiving the activation input includes: receiving a delimiter operation input, the delimiter operation input including: a fixation point at a first position on the display screen; a mouse event at a second position on the display screen; wherein the distance between the first position and the second position on the display screen exceeds a threshold value.

[0026] In an exemplary aspect of the method, obtaining a fixation point of a user on the display screen includes: obtaining a face image of the user; calculating eye fixation information based on the face image; calculating the fixation point on the display screen based on the eye fixation information.

[0027] In an exemplary aspect of the method, the method further includes: receiving a request to terminate the fixation assistance operation.

[0028] In some aspects, the present invention describes a system, including: a pointing device; a display screen; one or more processor devices; one or more memories storing machine-executable instructions, when the machine-executable instructions are executed by the one or more processor devices, causing the system to: receive an activation input for initiating a fixation assistance operation corresponding to an application program; in response to receiving the activation input, receive a fixation point of a user on the display screen; in response to receiving the activation input, further receive a first cursor position of the pointing device on the display screen; extract a fixation area based on the fixation point on the display screen; generate an interaction area based on the fixation area and the first cursor position; receive an interaction event corresponding to the interaction area at a second cursor position of the pointing device on the display screen; in a system hook, generate a replacement interaction event for passing to the application program, the replacement interaction event being generated in response to receiving the interaction event at the second cursor position.

[0029] In the above exemplary aspect of the system, when the machine-executable instructions are executed by the one or more processors, causing the system to generate a replacement interaction event by: receiving an interaction mapping for associating a position of the fixation area on the display screen with a position of the interaction area on the display screen; mapping the second cursor position to a replacement cursor position based on the interaction mapping; generating the replacement interaction event based on the replacement cursor position.

[0030] In the above exemplary aspect of the system, when the machine-executable instructions are executed by the one or more processors, further causing the system to: process the replacement interaction event to execute a command operation of the application program.

[0031] In an exemplary aspect of the system, when the machine-executable instructions are executed by the one or more processors, the system extracts the region of regard by generating an image of a portion of the content of the display screen based on the point of regard.

[0032] In the above exemplary aspect of the system, when the machine-executable instructions are executed by the one or more processors, the system generates an interaction region by magnifying the image of the portion of the content of the display screen to generate a magnified image; and generating the interaction region based on the magnified image and the first cursor position.

[0033] In an exemplary aspect of the system, the size of the image of the portion of the content of the display screen depends on the accuracy of the point of regard.

[0034] In an exemplary aspect of the system, when the machine-executable instructions are executed by the one or more processors, the system receives the activation input by receiving a delimiter operation input, the delimiter operation input including at least one of the following: pointing device input; keyboard event input; gesture input; or audio input.

[0035] In an exemplary aspect of the system, when the machine-executable instructions are executed by the one or more processors, the system receives the activation input by receiving a delimiter operation input, the delimiter operation input including: a point of regard at a first position on the display screen; a mouse event at a second position on the display screen; wherein the distance between the first position and the second position on the display screen exceeds a threshold.

[0036] In an exemplary aspect of the system, when the machine-executable instructions are executed by the one or more processors, the system obtains the point of regard of the user on the display screen by obtaining an image of the user's face; calculating eye gaze information based on the face image; and calculating the point of regard on the display screen based on the eye gaze information.

[0037] In some exemplary aspects, the present invention describes a computer-readable medium storing instructions thereon. When the instructions are executed by one or more processor units of a computing system, the computing system is caused to perform the following operations: receive an activation input for initiating a gaze assist operation corresponding to an application; in response to receiving the activation input, receive a gaze point of a user on a display screen; in response to receiving the activation input, further receive a first cursor position of a pointing device on the display screen; extract a gaze region based on the gaze point on the display screen; generate an interaction region based on the gaze region and the first cursor position; receive an interaction event corresponding to the interaction region at a second cursor position of the pointing device on the display screen; and in a system hook, generate a replacement interaction event for passing to the application, the replacement interaction event being generated in response to receiving the interaction event at the second cursor position. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The following is a reference by way of example to the drawings showing exemplary embodiments of the present application, wherein:

[0039] Figure 1 is a block diagram of an exemplary computing system that can be used to implement examples of the present invention;

[0040] Figure 2 is applicable to the embodiments described herein and is Figure 1 a schematic diagram of a graphical user interface (GUI) presented on the display screen of a computing system;

[0041] Figure 3 is a block diagram of an exemplary architecture of a gaze assist interaction system according to an example of the present invention;

[0042] Figure 4 is according to an exemplary embodiment and is Figure 1 a schematic diagram of a GUI presented on the display screen of a computing system;

[0043] Figure 5 is a block diagram of an exemplary architecture of an interception module according to an example of the present invention;

[0044] Figure 6 is a flowchart of an exemplary method for gaze assist interaction on a display screen according to an example of the present invention.

[0045] Like reference numerals may be used in different drawings to represent like components. DETAILED DESCRIPTION

[0046] The following describes exemplary technical solutions of the present invention with reference to the drawings.

[0047] To facilitate understanding of the present invention, some related terms that may be relevant to the examples disclosed herein are described below.

[0048] In the present invention, a "pointing device" may refer to: a human-machine interface device that enables a user to input spatial data to a computer. In an example, the pointing device may be a hand-held input device, including a mouse, a touchpad, a touch screen, a stylus, a joystick, or a trackball, etc. In an example, the pointing device may be used to control a cursor or a pointer in a GUI, for pointing, moving, or selecting text or an object on a display screen, etc. In an example, the spatial data may be continuous and / or multi-dimensional data.

[0049] In the present invention, "display content" or "content of the display screen" may refer to: an image displayed on a display screen of a computing system. For example, the display screen shows a desktop or an active application window or any number of application windows, and the positions of these application windows may be simultaneously visible on the display screen.

[0050] In the present invention, "gaze tracking" or "gaze estimation" may refer to: a method of tracking eye movements and estimating a gaze vector (for example, the estimated gaze vector may include two angles describing the gaze direction, which are the yaw angle and the pitch angle) or a point of gaze (POG) on a display screen or in the surrounding environment. A commonly used method for gaze estimation is video-based eye tracking, such as using a camera to capture a face or eye image and calculating a gaze vector or POG based on the face or eye image.

[0051] In the present invention, "point of gaze (POG)" may refer to: the object or location that a person is gazing at in a scene of interest, or more specifically, the geometric intersection between the gaze vector and the scene of interest. In other examples, the POG may correspond to the position where the visual axis intersects the 2D display screen on the display screen. In an example, the POG on the display screen may be described by a set of 2D coordinates (x, y), and this set of coordinates corresponds to the position relative to the display screen coordinate system on the display screen.

[0052] In the present invention, "eye gaze information" may include information representing the gaze direction of a user, for example, a gaze vector or a point of gaze (POG) on a display screen or in the surrounding environment, etc.

[0053] In the present invention, the "fixed state" may refer to: the state in which the user maintains fixation at a single position. When the user is in the fixed state, the user's eyes do not remain completely stationary and may exhibit jitter. In the present invention, "jitter" may refer to: the slight involuntary movement of the eyes during fixation. In an example, the jitter may not significantly affect the user's vision, but the jitter may pose a challenge to the gaze tracking system.

[0054] In the present invention, "saccade" may refer to: the rapid, simultaneous movement of both eyes between two or more fixed states in order to direct the point of fixation from one position to another. Saccades are voluntary and should not be mistaken for jitter.

[0055] In the present invention, "separator operation" may refer to: an indication that marks the boundary between two different regions in a data stream, text stream, or expression stream. In the present invention, a separator operation may be generated in response to an action performed by the user in order to invoke or activate a gaze assistance interaction system function, for example, using device input (such as a mouse click, keyboard, voice command, air gesture, or body movement, etc.) or a combination of device inputs, or using other separator operation logics.

[0056] In the present invention, "gaze region" or "region of interest (ROI) of gaze concern" may refer to: the region on the display screen that includes the user's point of fixation, for example, a rectangular region defined by width (w) and height (h) centered on the user's POG (x,y).

[0057] In the present invention, "enlarged gaze region" or "enlarged region of interest (eROI)" may refer to: an image that serves as an enlarged copy of the ROI including the user's POG, for example, a rectangular region defined by width (s) and height (t) centered on the cursor position (a,b).

[0058] In the present invention, "interaction region" may refer to:

[0059] In the present invention, a "cursor" may refer to: a movable object for indicating a position on a display screen. In an example, the cursor may be a pointer associated with a mouse or another pointing device, or the cursor may be any movable indicator on the display screen corresponding to an input device, and indicates a position on the display screen where an input operation can be directed. For example, the cursor of a pointing device may indicate a position on the display screen corresponding to an interaction event, such as a mouse event (e.g., a mouse click, double - click, mouse up - slide or mouse down - slide), a touchpad event (e.g., a tap action or a slide action), or a keyboard event (e.g., a text cursor may identify the position where typed text will be inserted or deleted), etc. In the present invention, the position of the cursor on the display screen can be described by coordinates relative to the display screen coordinate system.

[0060] In the present invention, a "GUI object" or "GUI element" may refer to: an object or element visible within a GUI, which is associated with control operations within an application window, such as icons, buttons, folders, menus, etc. that require user interaction (e.g., mouse click, tap, slide, etc.) to perform an operation.

[0061] Other terms used in the present invention may be introduced and defined in the following description.

[0062] Figure 1 is a block diagram of a simplified exemplary implementation of a computing system 100 suitable for implementing the embodiments described herein. Examples of the present invention may be implemented in other computing systems, which may include components different from those discussed below. The computing system 100 may be used to execute instructions for gaze - assisted pointing device interaction using any of the examples described herein.

[0063] Although Figure 1 a single instance of each component is shown, there may be multiple instances of each component in the computing system 100. Further, although the computing system 100 is shown as a single block, the computing system 100 may be a single physical machine or device (e.g., implemented as a single computing device, such as a single workstation, a single end - user device, a single server, etc.), and may include a mobile communication device (smartphone), a laptop computer, a tablet computer, a desktop computer, an automotive driving assistance system, a smart appliance, a wearable device, an assistive technology device, a virtual reality device, an augmented reality device, an Internet of Things (IoT) device, an interactive kiosk, etc.

[0064] The computing system 100 includes at least one processor 102, such as a central processing unit, a microprocessor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), dedicated logic circuitry, a dedicated artificial intelligence processor unit, a graphics processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU), a hardware accelerator, or a combination thereof.

[0065] The computer system 100 may include an input / output (I / O) interface 104 that may support connections to input device(s) 106 and / or output device(s) 114. In the example shown, input device 106 (e.g., a keyboard, a touch screen, a keypad, etc.) may further include a camera 108 (e.g., an RGB camera or an infrared (IR) camera), a pointing device 110 (e.g., a mouse, a stylus, a touchpad, a trackball, and / or a joystick), or an optional microphone 112. In the example shown, output device 114 may include a display screen 116 and other output devices (e.g., speakers and / or a printer). In other exemplary embodiments, there may be no input device 106 and output device 114, in which case the I / O interface 104 may not be required.

[0066] The computing system 100 may include an optional communication interface 118 for wired or wireless communication with other computing systems (e.g., other computing systems in a network). The communication interface 118 may include a wired link (e.g., an Ethernet cable) and / or a wireless link (e.g., one or more antennas) for intra-network communication and / or inter-network communication.

[0067] The computing system 100 may include one or more memories 120 (collectively referred to as "memory 120"), which may include volatile or non-volatile memory (e.g., flash memory, random access memory (RAM), and / or read-only memory (ROM)). The non-transitory memory 120 may store instructions executed by the processor 102, e.g., to execute the examples described in the present invention. For example, the memory 120 may store instructions for implementing any of the networks and methods disclosed herein. The memory 120 may include other software instructions, such as for implementing an operating system (OS) and other applications 124 or functions. The instructions may include instructions 300-I for implementing and operating the gaze-assisted interaction system 300 described below with reference to Figure 3 The memory 120 may also store other data 122, information, rules, policies, and machine-executable instructions described herein, e.g., including instructions 310-I for implementing the gaze tracking system 310.

[0068] In some examples, the computing system 100 may also include one or more electronic storage units (not shown), such as solid state drives, hard disk drives, disk drives, and / or optical disk drives. In some examples, data and / or instructions may be provided by an external memory (e.g., an external drive that communicates with the computing system 100 wired or wirelessly), or by a transitory or non-transitory computer-readable medium. Examples of non-transitory computer-readable media include RAM, ROM, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, CD-ROM, or other portable memory. The storage unit and / or the external memory may be used in conjunction with the memory 120 to implement the data storage, retrieval, and caching functions of the computing system 100. For example, the components of the computing system 100 may communicate with each other via a bus.

[0069] Figure 2 is applicable to the embodiments described herein and is a schematic diagram of a graphical user interface (GUI) 200 presented on the display screen 116 of the computing system in Figure 1 According to an example of the present invention, the GUI 200 is an illustrative example of an interface to which the systems, methods, and processor-readable media described herein may be applied.

[0070] In Figure 2In the example shown, the application window 205 corresponding to the software application 124 is active in the GUI 200, and the cursor 210 is visible within the application window 205. In the example, the cursor 210 can be associated with a pointing device 110 (e.g., a mouse), or the cursor 210 can be a text cursor, etc. In the example, the cursor 210 can be positioned within the application window 205 corresponding to the initial cursor position 215, e.g., described by a set of coordinates (a, b) relative to the coordinate space of the display screen 116. For illustrative purposes, the application window 205 can represent a drawing application, e.g., having a drawing canvas 220 and a control panel 225 including a plurality of GUI objects 230. However, it should be understood that any application can be used.

[0071] In a common scenario of the application 124, the user can manipulate the pointing device 110 to move the cursor 210 within the application window 205 to contact a desired GUI object 231, e.g., by selecting an icon or a button, or interacting with a menu, etc. In the example, before moving the cursor 210 to contact the desired GUI object 231, the user can move their point of gaze (POG) 240 on the display screen 116, and can fix their POG 240 on the desired GUI object 231 before manipulating the pointing device 110 to move the cursor 210 to the desired GUI object 231 and contacting the pointing device 110 to interact with the desired GUI object 231. In the example, the user's POG 240 can be determined by the gaze tracking system 310 of the computing system 100. In some embodiments, e.g., the user's POG 240 can be described by a set of coordinates (x, y) relative to the coordinate space of the display screen 116, as described in the discussion below. However, as described above, in a scenario where the user has to quickly manipulate the cursor at a large screen distance, it may be difficult to effectively transition from a rough movement of the pointing device 110 to precise control of the pointing device 110 to interact with the desired GUI object 231. Figure 3 As an example according to the present invention, FIG. 9 is a block diagram of an example architecture of a gaze-assisted interaction system 300 that can be used to implement a method for gaze-assisted pointing device interaction.

[0072] Figure 3 As an example according to the present invention, FIG. 9 is a block diagram of an example architecture of a gaze-assisted interaction system 300 that can be used to implement a method for gaze-assisted pointing device interaction.

[0073] In some embodiments, for example, a gaze tracking system 310 external to the gaze assistive interaction system 300 may continuously capture images of a user (e.g., a face image 305) using a camera 108. In an example, the gaze tracking system 310 may generate eye gaze information by estimating a point of gaze (POG) 240 of the user based on the face image 305. In other embodiments, the eye gaze information may include a gaze vector, and the POG 240 may be estimated based on the gaze vector and the known position of the display screen 116. In an example, the POG 240 calculated by the gaze tracking system 310 may be input to a gaze region extractor 320 of the gaze assistive interaction system 300.

[0074] In an example, typical hardware configurations of the gaze tracking system 310 include remote systems and head-mounted or wearable systems. In a remote gaze tracking system, the hardware components including the camera 108 are placed away from the user, while in a head-mounted gaze tracking system, the hardware components are placed inside a head-mounted device (e.g., an augmented reality (AR) or virtual reality (VR) device in the form of a helmet, a head-mounted headset, or a pair of glasses), positioning the camera 108 in a position very close to the eyes. The cameras used in the eye tracking system may include infrared (IR) cameras (capturing IR data) or RGB cameras (capturing visible spectrum data). In an example, the quality of the gaze tracking hardware may vary on different devices. For example, remote tracking systems typically use RGB cameras built into devices (e.g., mobile communication devices, tablets, laptops, etc.), while external eye tracking devices may include IR cameras and may be placed near the display screen 116.

[0075] In some embodiments, for example, the gaze assistive interaction system 300 may receive other inputs in addition to the POG 240, including an activation signal 315 for initiating a gaze assist operation; an initial cursor position 215 of a cursor visible in an application window 205 on the display screen 116 (before initiating the gaze assist operation); an interaction event 350 performed by a pointing device 110, e.g., a mouse event performed by the user during the gaze assist operation. In an example, the gaze assistive interaction system 300 may output a mapped interaction event 370 to the application 124.

[0076] In some embodiments, for example, an activation signal 315 is provided to the gaze assist interaction system 300 to initiate a gaze assist operation. In an example, the activation input 315 can be triggered using an input device, for example, as a delimiter operation. In an example, the delimiter operation can be generated by a voluntary or intentional input action performed by a user to indicate to the gaze assist interaction system 300 to initiate a gaze assist operation. In some embodiments, for example, the input action can be configured to be impossible to accidentally perform or cause the user to accidentally initiate a gaze assist operation. In an example, the delimiter operation can be caused by a device input (such as a mouse click, keyboard, etc.), an audio input (such as a voice command), a gesture input (such as an air gesture or body movement, etc.), or a combination thereof, or using other delimiter operation logics. In some embodiments, for example, the activation signal 315 can be triggered by a keyboard input (such as by pressing a specific key on the keyboard) or by a mouse input (such as using a programmable key on the mouse) or by a combined mouse-keyboard input. In other embodiments, the activation signal 315 can be triggered by a gesture input performed by the user that is captured by the camera 108 and processed by the gaze tracking system 310, such as a gesture, a facial expression, a blink, a nod, or a combination of gesture inputs. In other embodiments, for example, the activation signal 315 can include a combination of a gaze fixation at the POG 240 and a mouse input at an initial cursor position 215 that is away from the POG 240 (for example, greater than a threshold distance from the POG 240).

[0077] In an example, in response to receiving the activation signal 315, the gaze assist interaction system 300 can initiate a gaze assist operation. In an example, the gaze assist operation can include: receiving an initial cursor position 215 of the pointing device 110; extracting a gaze region 325 of the application window 205 corresponding to at least one GUI object 230 in the user's point of gaze (POG) 240 and the application window 205; magnifying the gaze region 325 to generate a magnified gaze region 335; transposing the magnified gaze region 345 to a new position in the application window 205 to generate an interaction region 345, the new position corresponding to the initial cursor position 215; capturing a user interaction event 350 using the pointing device 110 at a second cursor position associated with the interaction region 345; and outputting a mapped interaction event 370 to the application 124, the mapped interaction event 370 corresponding to at least one of the GUI objects 230 in the GUI object 230 in the application window 205. In this regard, the gaze assist interaction system 300 can solve the problem of cursor movement at a large distance by benefiting from the rapid movement of the gaze while addressing the jitter associated with gaze tracking. For illustrative purposes, examples of gaze assist operations will be described below Figure 4 Describe an example of a gaze assist operation.

[0078] In an example, in response to receiving the activation signal 315, the gaze region extractor 320 of the gaze-assisted interaction system 300 may receive the POG 240 and may generate a gaze region 325. In an example, the gaze region 325 may be configured as a single layered window having dimensions (w×h) that captures an image of a portion of the displayed content including the POG 240. In an example, the gaze region 325 may cover the application window 205 or any other open application windows. In an example, the size of the gaze region 325 may be predetermined, or the size of the gaze region 325 may depend on the accuracy of the gaze tracking system 310. For example, in response to receiving a POG 240 generated by a gaze tracking system 310 displaying lower accuracy (e.g., a gaze tracking system 310 using an RGB camera or operating under poor lighting conditions), the gaze region extractor 320 may generate a gaze region 325 larger than the predetermined size to compensate for errors in the POG 240, for example, to ensure that the gaze region 325 adequately captures the portion of the application window 205 corresponding to the user's gaze.

[0079] Figure 4 According to an exemplary embodiment of the present invention, during gaze assist operation, Figure 1 FIG2 is a diagram of a graphical user interface (GUI) 200 presented on the display screen 116 of a computing system of FIG. GUI 200 is an illustrative example of an interface to which the systems, methods, and processor-readable media described herein may be applied, according to an example of the present invention.

[0080] refer to Figure 4 The gaze area 325 may be a transparent window having a window image that displays a portion of the display content corresponding to the user's gaze. In an example, the gaze area 325 may be rectangular or square in shape and described by a width w 250 and a height h 255. The center of the gaze area 325 is described by a set of coordinates (x, y) corresponding to the POG 240. In an example, the coordinates representing each corner of the gaze area 325 may define the boundaries of the window. For example, the upper left corner and lower right corner of the gaze area 325 may be described as follows:

[0081]

[0082]

[0083] See again Figure 3, the fixation region 325 can be input to the amplifier 330, and the size of the window of the fixation region 325 and the corresponding window image can be enlarged to generate an enlarged fixation region 335 with a size of (s×t). In an example, the size of the enlarged fixation region 335 can be predetermined, or the size of the enlarged fixation region 335 can depend on the size or resolution of the display screen 116 or the quality of the gaze tracking system 310. In an example, the enlarged fixation region 335 can then be input to the transposer 340 to position the enlarged fixation region 335 on the display screen 116 based on the initial cursor position 215, thereby generating an interaction region 345. In an example, the interaction region 345 can be a single-layer window that has a window image representing a part of the display content with a size of (s×t) and including the POG 240. In an example, the interaction region 345 can cover the application window 205 or any other open application window. The transposer 340 can also output an interaction map 355 that correlates the position of the fixation region 325 with the position of the interaction region 345 on the display screen 116.

[0084] In an example, the size of the interaction region 345 can depend on the quality of the gaze tracking system 310. For example, in response to receiving the POG 240 generated by a gaze tracking system 310 with lower-quality hardware or exhibiting noisy or low-precision performance, the interaction region 345 can include an optional panning operation that enables the user to manipulate the interaction region 345 through a panning motion (e.g., using a specific mouse button for performing the panning operation). In an example, in a case where the gaze tracking system 310 cannot accurately capture the desired POG 240 in the interaction region 345, the panning operation can be used to slide the interaction window to bring the desired POG 240 into view.

[0085] Reference Figure 4 , the interaction region 345 can be a transparent window with a window image that displays a part of the display content corresponding to the user's gaze. In an example, the interaction region 345 can be in the shape of a rectangle or a square and is described by a width s 260 and a height t 265, and the center of the interaction region 345 is described by a set of coordinates (a,b) corresponding to the initial cursor position 215. In an example, the coordinates representing each corner of the interaction region 345 can define the boundaries of the window. For example, the upper left corner and the lower right corner of the interaction region 345 can be described as follows:

[0086]

[0087]

[0088] In some embodiments, for example, the user may manipulate the pointing device 110 to move the cursor 210 from the initial cursor position 215 on the display screen 116 to the assistive cursor position 270 on the display screen 116. In an example, the assistive cursor position 270 may be described by a set of coordinates (a’, b’) with respect to the coordinate space of the display screen 116. In an example, the interaction area 345 may include, in its window image, an enlarged GUI object image 280 corresponding to the GUI object 230 of the control panel 225. In an example, when the user manipulates the cursor to the assistive cursor position 270, the assistive cursor position 270 may correspond to a desired one of the GUI object images in the GUI object image, for example, the desired enlarged GUI object image 281, as Figure 4 shown. Additionally, the user may interact with the pointing device 110 associated with the cursor 210 to initiate an interaction event 350 at the assistive cursor position 270, for example, by clicking a mouse button or otherwise contacting the pointing device 110 to initiate the interaction event 350.

[0089] Referring again to Figure 3 , the interception module 360 of the gaze assist interaction system 300 may receive the interaction map 355 and the interaction event 350, and may output the mapped interaction event 370, as described in the discussion below Figure 5 . In an example, the interception module 360 may be software implemented in the computing system 100, where the processor 102 is used to execute the instructions 300-I of the gaze assist interaction system 300 stored in the memory 120.

[0090] Figure 5 is a block diagram of an exemplary interception module architecture 360 according to an example of the present invention, and this example interception module architecture 360 may be used to intercept the system-level function 540 called during the gaze assist operation. In some examples, the interception module 360 utilizes a hooking mechanism to intercept the system-level function 540 called during the gaze assist operation. In the case of intercepting the system-level function 540, the hooking mechanism may intercept the call between two processes and call a customized function (e.g., the hooking process 520) between them.

[0091] In an example, during the gaze assist operation, when the interaction event 350 is executed, the interception module 360 may call the system library 330 (and the associated system-level function 340) on behalf of the application 124, for example, to generate a replacement event (e.g., the mapped interaction event 370) including the replacement cursor position (x’, y’) and pass it back to the application 124. In an example, the coordinates of the replacement cursor position in the mapped interaction event 370 may be calculated based on the interaction map 355 in the following manner:

[0092]

[0093]

[0094] Where (x, y) are the coordinates of the POG 240, (a, b) are the coordinates of the initial cursor position 215, (a', b') are the coordinates of the assist cursor position 270, w is the width of the fixation region 325, h is the height of the fixation region 325, s is the width of the interaction region 345, and h is the height of the interaction region 345. In an example, in response to receiving the mapped interaction event 370, the application 124 may perform a control operation as if the user directly performed an interaction event (e.g., a mouse click) at the location of the mapped interaction event 370.

[0095] Before initiating the gaze assist operation, the software code may first be initialized to activate the gaze assist interaction system 300 and load the instructions associated with the interception module 360 into the dynamic library 510. In an example, the instructions may include one or more hook procedures 520 to intercept system-level functions 540 called by the application 124 during a gaze assist operation (e.g., a mouse event). For example, in a typical mouse operation within the Windows™ operating system (OS), when a mouse event occurs (e.g., cursor movement, pressing or releasing a mouse button, etc.), the OS posts the input as a message to the queue of the appropriate thread, which determines the thread based on the window over which the mouse hovered during the mouse event, or the window that captured the mouse input. If a hook is used to intercept the mouse event before the mouse event message is posted to the window's queue, the hook procedure can process the mouse event using its own defined subroutine and then pass the mouse input message to the next application in the mouse hook chain, or if the hook chain is empty, pass the mouse input message to the target window. Although a mouse is used as the input device to describe the exemplary implementation, it should be understood that other input devices may also be used.

[0096] In an example, using system-level programming, the option to enable or disable the gaze assist interaction system 300 can be configured as a toggle in the system settings of the computing system 100. By implementing the gaze assist interaction system 300 using system-level programming, the interception module 360 can be used to: during the activation of the gaze assist interaction system 300, intercept any interaction event 350 performed using the pointing device 110, where these interaction events 350 correspond to any program running on the OS of the computing system 100 (e.g., Windows™, Linux™, MacOS™, Android™, etc.).

[0097] Figure 6FIG. 600 is a flow diagram of an exemplary method 600 for performing a gaze assist operation using a pointing device 110 according to an example of the present invention. Method 600 may be executed by computing system 100. For example, processor 102 may execute computer-readable instructions 300-I (which may be stored in memory 120) to cause computing system 100 to execute method 600. Method 600 may be executed using a single physical machine (e.g., a workstation or a server), multiple physical machines working together (e.g., a server cluster), or cloud-based resources (e.g., using virtual resources on a cloud computing platform).

[0098] Method 600 begins at step 602, where gaze assist interaction system 300 receives an activation input 315 for initiating a gaze assist operation corresponding to application 124. In an example, activation input 315 may be a delimiter operation caused by a user performing an input action to indicate to gaze assist interaction system 300 to initiate a gaze assist operation. In an example, device input (e.g., a mouse click, a keyboard, etc.), audio input (e.g., a voice command), gesture input (e.g., an air gesture or a body movement, etc.), or a combination thereof may be used to generate the input action, such as including one or more input modes or using other delimiter operation logics. In an example, the input action may be defined as requiring a user-intended action to avoid accidentally triggering activation input 315. In one embodiment, for example, the input action may include a combination of a POG 240 at a first location on display screen 116 and a mouse interaction at a second location on display screen 116, where the distance between the first location and the second location on display screen 116 exceeds a threshold.

[0099] In the example, before receiving the activation input 315, the user can view the display screen 116 and interact with the display content of the application window 205, where the user's eyes can be fixed on the desired content on the display screen 116. At step 604, in response to receiving the activation input 315, the gaze-assisted interaction system 300 can receive the point of gaze (POG) 240 of the user from the gaze tracking system 310 of the computing system 100, and the POG 240 corresponds to the gaze fixation of the user on the display screen 116. In the example, the gaze tracking system 310 can capture a face image 305 corresponding to the user, and can calculate the point of gaze (POG) 240 for the user on the display screen 116 based on the face image 305. In some embodiments, for example, the gaze-assisted interaction system 300 can request the POG 240 from the gaze tracking system 310 in response to receiving the activation input 315. In other embodiments, for example, the gaze tracking system 310 can continuously capture a sequence of face images 305 to calculate a corresponding sequence of POGs 240, or can repeatedly capture the face image 305 at a predetermined frequency to calculate the corresponding POGs 240 at a predetermined frequency for providing to the gaze-assisted interaction system 300. In the example, the face image 305 can be captured by the camera 108 on the computing system 100, or can be a digital image captured by another camera on another electronic device and transmitted to the computing system 100. In some embodiments, for example, image recognition techniques known in the art can be used to detect facial feature points in the face image 305, including the eyes. In addition, gaze tracking techniques known in the art can be used to calculate the POG 240.

[0100] At step 606, further in response to receiving the activation input 315, the first cursor position (e.g., the initial cursor position 215) of the pointing device 110 on the display screen 116 can be received.

[0101] At step 608, a portion of the display content can be extracted as the gaze region 325 based on the POG 240 on the display screen 116. In the example, the gaze region 325 can be configured as a single layered window with dimensions (w×h), and the window captures an image of a portion of the display content including the POG 240.

[0102] In step 610, an interaction area 345 can be generated based on the fixation area 325 and the first cursor position 215. In an example, the interaction area 345 can be an enlarged copy of the fixation area 325, e.g., a single layered window configured to have dimensions (s×t) that captures an enlarged image of a portion of the display content including the POG 240. Additionally, the interaction area 345 can be positioned on the display screen 116 based on the first cursor position, and the center of the interaction area 345 is described by a set of coordinates (a,b) corresponding to the initial cursor position 215.

[0103] Optionally, in step 612, a cancellation input that operates as a separator can be received to end the fixation assistance operation and remove the interaction area 345 from the display screen 116. In an example, if the interaction area 345 has erroneously captured the POG 240, or if the user decides to interact with a different area of the display screen 116, the user can choose to end the fixation assistance operation. In an example, after step 612, if a cancellation input is received, the method can return to step 602. Otherwise, step 614 is executed.

[0104] In step 614, an interaction event 350 corresponding to the interaction area 345 can be received at a second cursor position (e.g., the assist cursor position 270) on the display screen 116. In an example, the pointing device 110 can be used to locally position the cursor 210 on the enlarged GUI object image 281 on the interaction area 345, and the interaction event 350 (e.g., a mouse click) can be performed by the user.

[0105] In step 616, in response to receiving the interaction event 350 at the second cursor position, a replacement mapped interaction event 370 can be generated in the system hook for passing to the application 124. In an example, the system hook can intercept the interaction event 350 and can use its own subroutine to process the interaction event 350 to replace the coordinates associated with the interaction event 350 based on the interaction mapping 355 before passing it to the application 124. In an example, the interaction mapping 355 can map the position (a’,b’) corresponding to the assist cursor position 270 within the interaction area 345 to the corresponding replacement cursor position (x’,y’) within the fixation area 325.

[0106] Optionally, in step 618, the mapped interaction event 370 can be passed to the application 124. When the mapped interaction event 370 is passed to the application 124, the application 124 can perform a control operation corresponding to the GUI object 230 associated with the mapped coordinates (x’,y’) as if the interaction event 350 actually occurred at the mapped coordinates (x’,y’) of the mapped interaction event 370 on the display screen 116.

[0107] Optionally, at step 620, after the control operation is performed, one or more post-processing steps may be performed. For example, the initial cursor position 215 may be updated to reflect the coordinates of the assist cursor position 270. In other examples, the interaction area 345 may be removed from the display screen, and the gaze assist operation may end. In other examples, for example, if further input selection is required to perform the control operation, a subsequent menu may be displayed at the initial cursor position 215. For example, a drop-down menu or a dialog box may be displayed, and the user may use the pointing device 110 to make the required selection as a further interaction event 350. In an example, the system hook may further intercept any other interaction event 350a and may pass the further mapped interaction event 370a to the application 124.

[0108] In an exemplary embodiment of the present invention, the display screen 116 may include the display screen of a laptop or a desktop computer, and the pointing device 110 may include a mouse or a touchpad. In an example, the display screen 116 may include multiple display screens, such as in an extended screen configuration. In another embodiment, the display screen 116 may include a single large display screen or multiple large touch display screens. For example, the display screen is too large or the elements on the display screen are too far away for the user to reach. In one embodiment, for example, the gaze assist interaction system may be integrated into an in-vehicle computing system. For example, the area of interest on the display screen may be made closer to the operator's hand by gazing.

[0109] In another exemplary embodiment of the present invention, the cursor 210 may be configured to be used in a text editor. In an example, for example, a typical cursor 210 associated with a pointing device for selecting GUI objects may be positioned at any pixel on the display screen. In contrast, when the cursor 210 is configured to be used in a text editor, an alternative cursor (e.g., a text cursor 211) may indicate the position of text input (e.g., from a keyboard or a numeric keypad), where the text cursor is restricted to the position between two characters in the text editor. In an example, the text cursor 211 may be repositioned by the pointing device or the arrow keys on the keyboard. In an example, when generating the interaction area 345 corresponding to the text editor, the text cursor 211 may be automatically generated and placed at the assist cursor position 270 associated with the interaction event 350.

[0110] The various embodiments of the present invention have been described in detail by way of examples. For those skilled in the art, changes and modifications can be made to these embodiments without departing from the scope of the present invention. The present invention includes all such changes and modifications that fall within the scope of the appended claims.

[0111] Although the present invention describes methods and processes by steps performed in a certain order, one or more steps in the methods and processes may be appropriately omitted or changed. Where appropriate, one or more steps may be performed in an order other than the described order.

[0112] Although the present invention is described at least in part in terms of methods, those of ordinary skill in the art will understand that the present invention also relates to various components for performing at least some aspects and features of the described methods, whether hardware components, software, or any combination of the two. Accordingly, the technical solution of the present invention may be implemented in the form of a software product. A suitable software product may be stored in a pre-recorded storage device or other similar non-volatile or non-transitory computer-readable medium, including DVDs, CD-ROMs, USB flash drives, removable hard disks, or other storage media. The software product includes instructions tangibly stored thereon that enable a processing device (e.g., a personal computer, a server, or a network device) to execute examples of the methods disclosed herein. The machine-executable instructions may be in the form of a code sequence, configuration information, or other data, which, when executed, cause a machine (e.g., a processor or other processing device) to perform the steps in the methods provided by the examples of the present invention.

[0113] Without departing from the subject matter of the claims, the present invention may be implemented in other specific forms. The described exemplary embodiments are illustrative in all respects and not restrictive. Features selected from one or more of the above embodiments may be combined to create alternative embodiments not explicitly described, and features suitable for such combinations may be understood to be within the scope of the present invention.

[0114] All values and sub-ranges within the disclosed range are disclosed. In addition, although the systems, devices, and processes disclosed and illustrated herein may include a specific number of elements / components, these systems, devices, and components may be modified to include more or fewer such elements / components. For example, although any disclosed element / component may be a single quantity, the embodiments disclosed herein may be modified to include multiple such elements / components. The subject matter described herein is intended to cover and encompass all suitable technical variations.

Claims

1. A computer-implemented method, characterized in that, Comprising: Receiving an activation input for initiating a gaze assistance operation corresponding to an application; In response to receiving the activation input, receiving a gaze point of a user on a display screen; In response to receiving the activation input, further receiving a first cursor position of a pointing device on the display screen; Extracting a gaze region based on the gaze point on the display screen; Generating an interaction region based on the gaze region and the first cursor position; Receiving an interaction event corresponding to the interaction region at a second cursor position of the pointing device on the display screen; In a system hook, generating a replacement interaction event for passing to the application, the replacement interaction event being generated in response to receiving the interaction event at the second cursor position.

2. The method according to claim 1, wherein Generating the replacement interaction event includes: Receiving an interaction mapping for associating a position of the gaze region on the display screen with a position of the interaction region on the display screen; Mapping the second cursor position to a replacement cursor position based on the interaction mapping; Generating the replacement interaction event based on the replacement cursor position.

3. The method according to claim 2, wherein Further comprising: Processing the replacement interaction event to perform a command operation of the application.

4. The method according to any one of claims 1 to 3, characterized in that, Extracting the gaze region includes: Generating an image of a part of the content of the display screen based on the gaze point.

5. The method according to claim 4, wherein Generating the interaction region includes: Magnifying the image of the part of the content of the display screen to generate a magnified image; Generating the interaction region based on the magnified image and the first cursor position.

6. The method according to claim 4, wherein The size of the image of the part of the content of the display screen depends on the accuracy of the gaze point.

7. The method according to any one of claims 1 to 6, characterized in that, Receiving the activation input includes: Receiving a delimiter operation input, the delimiter operation input including at least one of the following: Pointing device input; Keyboard event input; Gesture input; or Audio input.

8. The method according to any one of claims 1 to 6, characterized in that, Receiving the activation input includes: Receiving a delimiter operation input, the delimiter operation input including: A gaze point at a first position on the display screen; A mouse event at a second position on the display screen; Wherein, the distance between the first position and the second position on the display screen exceeds a threshold.

9. The method according to any one of claims 1 to 6, characterized in that, Obtaining a gaze point of a user on the display screen includes: Obtaining a face image of the user; Calculating eye gaze information based on the face image; Calculating the gaze point on the display screen based on the eye gaze information.

10. The method according to any one of claims 1 to 9, characterized in that, Further comprising: Receiving a request to terminate the gaze assistance operation.

11. A system, characterized in that, Comprising: A pointing device; A display screen; One or more processor devices; One or more memories storing machine-executable instructions, which when executed by the one or more processor devices cause the system to perform the following operations: Receiving an activation input for initiating a gaze assistance operation corresponding to an application; In response to receiving the activation input, receiving a gaze point of a user on the display screen; In response to receiving the activation input, further receiving the first cursor position of the pointing device on the display screen; Extracting a gaze region based on the gaze point on the display screen; Generating an interaction region based on the gaze region and the first cursor position; Receiving an interaction event corresponding to the interaction area at a second cursor position of the pointing device on the display screen; In a system hook, generating a replacement interaction event for passing to the application, the replacement interaction event being generated in response to receiving the interaction event at the second cursor position.

12. The system according to claim 11, wherein When the machine-executable instructions are executed by the one or more processors, causing the system to generate a replacement interaction event by: Receiving an interaction mapping for associating a position of the gaze area on the display screen with a position of the interaction area on the display screen; Mapping the second cursor position to a replacement cursor position based on the interaction mapping; Generating the replacement interaction event based on the replacement cursor position.

13. The system according to claim 12, wherein When the machine-executable instructions are executed by the one or more processors, further causing the system to: Process the replacement interaction event to perform a command operation of the application.

14. The system according to any one of claims 11 to 13, characterized in that When the machine-executable instructions are executed by the one or more processors, causing the system to extract the gaze area by: Generating an image of a part of the content of the display screen based on the gaze point.

15. The system according to claim 14, wherein When the machine-executable instructions are executed by the one or more processors, causing the system to generate an interaction area by: Magnifying the image of the part of the content of the display screen to generate a magnified image; Generating the interaction area based on the magnified image and the first cursor position.

16. The system according to claim 14, wherein The size of the image of the part of the content of the display screen depends on the accuracy of the gaze point.

17. The system according to any one of claims 11 to 16, characterized in that When the machine-executable instructions are executed by the one or more processors, causing the system to receive the activation input by: Receiving a delimiter operation input, the delimiter operation input including at least one of the following: Pointing device input; Keyboard event input; Gesture input; or Audio input.

18. The system according to any one of claims 11 to 16, characterized in that When the machine-executable instructions are executed by the one or more processors, causing the system to receive the activation input by: Receiving a delimiter operation input, the delimiter operation input including: A gaze point at a first position on the display screen; A mouse event at a second position on the display screen; Wherein, the distance between the first position and the second position on the display screen exceeds a threshold.

19. The system according to any one of claims 11 to 16, characterized in that, When the machine-executable instructions are executed by the one or more processors, causing the system to obtain a gaze point of the user on the display screen by: Obtaining a face image of the user; Calculating eye gaze information based on the face image; Calculating the gaze point on the display screen based on the eye gaze information.

20. A non-transitory computer-readable medium, characterized in that, The non-transitory computer-readable medium stores machine-executable instructions that, when executed by one or more processors of a computing system, cause the computing system to: Receive an activation input for initiating a gaze assistance operation corresponding to an application; In response to receiving the activation input, receive a gaze point of the user on the display screen; In response to receiving the activation input, further receive a first cursor position of the pointing device on the display screen; Extract a fixation area based on the fixation point on the display screen; Generate an interaction area based on the fixation area and the first cursor position; Receive an interaction event corresponding to the interaction area at a second cursor position of the pointing device on the display screen; In the system hook, generate a replacement interaction event for transmission to the application, the replacement interaction event being generated in response to receiving the interaction event at the second cursor position.

Citation Information

Patent Citations

  • User-input apparatus, method and program for user-input

    US10444831B2

  • Integrated gaze / manual cursor positioning system

    US6204828B1