Line-of-sight classification

JP2024527316A5Pending Publication Date: 2025-05-16GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023580480
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-01
Filing Date
2022-07-01
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Conventional gaze direction trackers for augmented reality systems are resource-intensive and prone to errors when identifying pixels on transparent displays, leading to complex mapping tasks that interfere with user experience.

Method used

The system determines the region of the display where the user's gaze is directed using a classification engine with a convolutional neural network, reducing computational resources and errors by identifying regions rather than individual pixels.

Benefits of technology

This approach reduces computing resources and errors, allowing for more efficient and accurate activation of user interface elements based on gaze direction, enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A technique for tracking a user's gaze includes identifying an area of ​​a display to which the user's gaze is directed, the area including a plurality of pixels. By determining an area rather than a point, if the area corresponds to an element of a user interface, the improved technique allows the system to activate the element to which the determined area is selected. In some implementations, the system makes the determination using a classification engine including a convolutional neural network, such engine taking as input an image of the user's eyes and outputting a list of probabilities that the gaze is directed to each of a plurality of areas.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of, and claims the benefit of, U.S. Application No. 17 / 305,219, filed July 1, 2021, the disclosure of which is incorporated herein by reference in its entirety.

[0002] Technical Field This description relates to determining the area of ​​a display to which a user's gaze is directed. [Background technology]

[0003] background Some augmented reality (AR) systems track gaze direction, i.e., the direction in which a user's eyes are pointed. For example, an AR system may include smart glasses for displaying content to a user on a transparent display. Some smart glasses include a camera on the eyeglass frame configured to generate an image of the user's eyes and track the gaze direction.

[0004] Such an AR system allows a user interface on a transparent display to track the gaze direction. For example, there may be a first content and a second content depicted on the transparent display. By determining the user's gaze direction, the AR system can infer whether the user is looking at the first content or the second content. Summary of the Invention

[0005] overview The embodiments disclosed herein provide an improved technique for tracking a user's gaze on a display. In some embodiments, the display is a transparent display, such as a display embedded in smart glasses used in an AR system. In some embodiments, the display is a display used in a mobile computing system, such as a smartphone, a tablet computer, and the like. However, rather than tracking the user's gaze on a particular point on the display, the improved technique includes determining an area of ​​the display to which the user's gaze is directed. By determining an area, rather than a point, if the area corresponds to an element of a user interface, the improved technique allows the system to activate the element to which the determined area is selected. In some embodiments, the system makes the determination using a classification engine including a convolutional neural network, such engine taking an image of the user's eye as input and outputting a list of probabilities that the gaze is directed to each of a plurality of areas.

[0006] In one general aspect, a method may include receiving image data representing at least one image of a user's eye looking at a display configured to operate with an augmented reality (AR) application, the display including a plurality of regions, each of the plurality of regions including a plurality of pixels, corresponding to a respective element of a user interface. The method may also include identifying, based on the image data, a region of the plurality of regions of the display to which the user's gaze is directed at the moment, the identifying including inputting the at least one image of the user's eye to a classification engine configured to classify the gaze as being directed at one of the plurality of regions. The method may further include activating an element of the user interface to which the identified region corresponds.

[0007] In another general aspect, a computer program product comprises a non-transitory storage medium, the computer program product including code that, when executed by processing circuitry of a computing device, causes the processing circuitry to implement a method. The method may include receiving image data representing at least one image of a user's eye looking at a display at a moment in time, the display including a plurality of regions and configured to operate in an augmented reality (AR) application, each of the plurality of regions including a plurality of pixels and corresponding to a respective element of a user interface. The method may also include identifying, based on the image data, a region of the plurality of regions of the display to which the user's gaze is directed at the moment in time, the identifying including inputting the at least one image of the user's eye to a classification engine configured to classify the gaze as being directed at one of the plurality of regions. The method may further include activating an element of the user interface to which the identified region corresponds.

[0008] In another general aspect, an electronic device includes a memory and control circuitry coupled to the memory. The control circuitry can be configured to receive image data representing at least one image of a user's eye looking at a display at a moment in time, the display including a plurality of regions and configured to operate with an augmented reality (AR) application, each of the plurality of regions including a plurality of pixels and corresponding to a respective element of a user interface. The control circuitry can also be configured to identify, based on the image data, a region of the plurality of regions of the display to which the user's gaze is directed at the moment in time, the identifying including inputting the at least one image of the user's eye to a classification engine configured to classify the gaze as being directed at one of the plurality of regions. The control circuitry can be further configured to activate an element of the user interface to which the identified region corresponds.

[0009] One or more embodiments are set forth in detail in the accompanying drawings and description below. Other features will be apparent from the description and drawings, and from the claims. [Brief description of the drawings]

[0010] [Figure 1A] FIG. 1 illustrates an exemplary pair of smart glasses for use in an augmented reality (AR) system. [Figure 1B] FIG. 1 illustrates an exemplary electronic environment in which the improved techniques described herein may be implemented. [Figure 2A] FIG. 2 illustrates an example region on a display, including an activated region. [Figure 2B] FIG. 13 shows exemplary regions on a display, including activated regions, where the regions may not be contiguous. [Figure 3A] FIG. 1 illustrates an example convolutional neural network (CNN) configured to classify images of a user's eyes as having gaze directed at a particular region. [Figure 3B] FIG. 1 illustrates an example convolutional neural network (CNN) configured to classify an image of a user's eye and determine the specific point to which gaze is directed. [Figure 4] FIG. 1 illustrates example convolutional layers forming different branches of a CNN configured to adapt to different domain or tile geometries. [Diagram 5] 1 is a flow chart illustrating an exemplary method of implementing improved techniques within an electronic environment. [Figure 6] 1A-1C illustrate examples of computing devices and mobile computing devices that can be used to implement the described techniques. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Detailed Description Conventional eye gaze trackers are configured to predict which pixels of a transparent display a user's eyes are most likely to look at. For example, a conventional eye gaze tracker can derive a pixel-based heat map for a transparent display, where each pixel has a color based on the probability that the user is looking at that pixel.

[0012] A technical problem with the above-described conventional approach for tracking a user's gaze is that the output of a conventional gaze tracker is targeted to identifying pixels of a transparent display, and thus the conventional gaze tracker is resource intensive and can be error-prone. For example, a conventional gaze tracker can identify pixels that a user is most likely to look at, but cannot identify such pixels with any content depicted on the transparent display. Thus, a system using such a gaze tracker may also need to map the pixels to the displayed content, which may consume computing resources needed for other tasks and may disrupt the user's experience.

[0013] According to the embodiments described herein, the technical solution to the above-described technical problem includes identifying an area of ​​a display to which a user's gaze is directed, the area including a plurality of pixels. By determining an area rather than a point, if the area corresponds to an element of a user interface, the improved technique allows the system to activate the element to which the determined area is selected. In some embodiments, the system makes the determination using a classification engine including a convolutional neural network, such engine taking an image of the user's eye as an input and outputting a list of probabilities that the gaze is directed to each of a plurality of areas.

[0014] A technical advantage of the disclosed implementations is that such implementations use fewer computing resources and are less prone to error. For example, in some implementations, a region may be associated with a user interface element, e.g., a window containing content depicted on a display, and such association uses fewer computing resources to activate the window on the display than does mapping an identified pixel to such a window, as conventional gaze trackers do.

[0015] Note that in contrast to the conventional approach described above, the output of the improved technique is a vector of possibilities corresponding to regions, rather than individual pixels, and therefore the size of this output is much smaller than that of the conventional approach.

[0016] FIG. 1A illustrates an exemplary smart glasses 110 for use in an augmented reality (AR) system as a head. FIG. 1A illustrates the world side 112(a) of the transparent display 112 of the smart glasses 110. The smart glasses 110 can be used as a head-mounted display (HMD) in the AR system. The smart glasses 110 include a frame 111 to which the transparent display 112 is coupled. In some implementations, an audio output device 113 is coupled to the frame 111. In some implementations, a user can control the input of the smart glasses 110 by a touch surface 114, and so on. The smart glasses 110 can include a perception system 116 including various perception system devices, and a control system 117 including various control system devices that facilitate operation of the smart glasses 110. The control system 117 can include a processor 119 operably coupled to components of the control system 117, and a communication module 115 that provides communication with external devices and / or networks. The smart glasses 110 can also include an image sensor 118 (i.e., a camera 118), a depth sensor, a light sensor, and other such sensory devices. In some implementations, the image sensor 118 or camera 118 can capture still and / or moving images, patterns, features, lights, and the like.

[0017] It should be noted that in some implementations, the smart glasses 110 can be replaced with any type of HMD that includes a transparent display, and the HMD does not necessarily have to be in the form of wearable glasses or goggles. For example, one such HMD can take the form of a camera with a viewfinder configured to display AR content and allow observation of the world-side environment.

[0018] 1B illustrates an exemplary electronic environment 100 in which the technical solutions described above can be implemented. A computer 120 is configured to determine an area of ​​a display to which a user's gaze is directed.

[0019] The computer 120 includes a network interface 122, one or more processing units 124, and a memory 126. The network interface 122 includes, for example, an Ethernet adapter, etc., for converting electronic and / or optical signals received from the network 150 into an electronic form for use by the computer 120. The set of processing units 124 includes one or more processing chips and / or assemblies. The memory 126 includes both volatile memory (e.g., RAM) and non-volatile memory, such as one or more ROMs, disk drives, solid state drives, etc. The set of processing units 124 and the memory 126 together form control circuitry that is configured and arranged to perform the various methods and functions described herein.

[0020] In some implementations, one or more of the components of computer 120 may be or include a processor (e.g., processing unit 124) configured to process instructions stored in memory 126. Examples of such instructions shown in Figure 1 include an input manager 130, a classification manager 140, and a boot manager. Additionally, as shown in Figure 1, memory 126 is configured to store various data, which will be described in conjunction with respective managers that use such data.

[0021] The input manager 130 is configured to receive input data, such as image data 132, area data 134, slippage data 136, and user data 138. In some implementations, the various input data are captured via hardware coupled to a display, such as the transparent display 112 (FIG. 1A). For example, the hardware may include a camera, such as the camera 118, configured to capture an image of the user's eye. In some implementations, the hardware may include any of a gyroscope, a magnetometer, a GPS receiver, and the like, for acquiring input data, such as the slippage data 136 and the user data 138. In some implementations, the input manager 130 is configured to receive input data via the network interface 122.

[0022] Image data 132 represents at least one image of a user's eye. Image data 132 is adapted to be input into classification manager 140. In some implementations, image data 132 represents a series of images of a user's eye for tracking gaze direction. In some implementations, the series of images are frames of a video that tracks the user's eye movements.

[0023] Region data 134 represents regions of the display. Each of the multiple regions comprises multiple pixels of the display. In some embodiments, each region corresponds to a respective element of the user interface, such as a window that contains content viewed by a user. In some embodiments, each region has a rectangular shape that includes an array of pixels. In some embodiments, at least one region is non-contiguous and includes multiple rectangles. In some embodiments, region data 134 includes an identifier that identifies each region.

[0024] The slippage data 136 represents parameter values ​​corresponding to the degree of slippage of the smart glasses from a nominal location on the user's face for an embodiment in which the display is a transparent display used for smart glasses. In a nominal configuration of the gaze tracking system, the eyes are in a nominal location with a known (designed) pose relative to the display. During wear, the position of the glasses can be changed from this nominal location (slip on the nose, adjustment by the user, etc.). When slippage occurs, the pose between the eye and the eye tracking camera changes, and the image of the user's eye appears different than in the nominal configuration. Furthermore, the position of the display changes depending on the slippage of the smart glasses, so the gaze angle will also be different. Thus, in some embodiments, the slippage data 136 includes as parameter values ​​a prediction of the position of the eye relative to the camera. In some embodiments, such relative eye positions are represented as three-dimensional vectors. In some embodiments, such relative eye positions are represented as angular coordinates on a sphere.

[0025] User data 138 represents parameter values ​​describing physical differences between users that may affect the determination of the area to which the user's gaze is directed. User differences such as eye appearance, visual axis, head shape, etc., may all affect the accuracy of the determination. Thus, in some implementations, user data 138 represents values ​​of parameters defining eye appearance, visual axis, and head shape. In some implementations, such parameter values ​​are inferred directly from image data 134.

[0026] The classification manager 140 is configured to determine an area of ​​the display to which the user's gaze is directed, thereby generating classification data 144 based on at least the image data 132. In some implementations, the classification data 144 is based on any of the area data 134, the slippage data 136, and the user data 138. The classification manager 140, in some implementations, includes at least one branch of a convolutional neural network (CNN) that acts as the classification engine. The classification manager 140 includes a training manager 141 and a discrimination manager 144. The classification data 144 includes training data 145 and discrimination data 146.

[0027] In some implementations, the CNN receives image data 132 as input. In some implementations, the CNN also receives any of region data 134, slippage data 136, and user data 138 as input. In some implementations, the CNN has a defined number of layers, one of which is an output layer that generates an output. Such a neural network is configured to generate a classification result, where the region of the display is the region to which the user's gaze is directed. In some implementations, the output includes a vector of values ​​indicating the likelihood that the user's gaze is directed to each region.

[0028] The training manager 141 is configured to generate a classification engine, i.e., a neural network, based on training data 145. For example, in some implementations, the training data 145 represents an image of a user's eye along with a corresponding identifier of the area where the user's gaze was directed when the image was taken. The training manager 141 then adjusts weights of nodes in a hidden layer of the neural network to optimize a prescribed loss function. In some implementations, the loss function includes a classification-specific cross-entropy suitable for a multi-class classification engine such as a CNN described above. In some implementations, the loss function includes a Kullback-Leibler divergence loss. In some implementations, the classification engine learns to calibrate to different area layouts based on the area data 134, further details of which are shown in connection with FIG. 4. The weights adjusted by the training manager 141 and other data representing the architecture of the neural network are included in the classification data 144.

[0029] Identification manager 142 is configured to classify image data 132 to identify an area of ​​the display to which the user's gaze is directed. As shown in FIG. 1B, classification data 144 includes identification data 146 output from identification manager 142. In some implementations, identification data 146 represents a probability value, e.g., a vector of probabilities that the user's gaze is directed to each of a plurality of areas of the display.

[0030] In some implementations, the discrimination manager 142 is configured to generate classification data based on the loss function described above that is used to train the classification engine. In some implementations, the discrimination manager 142 is configured to accept as input any of the region data 134, the slippage data 136, and the user data 138. In some implementations, the training data 145 includes the region data, the slippage data, and the user data combined with an identifier of the region to which the user's gaze is directed. In some implementations, each of the above types of data is used to generate additional branches of the classification engine, and further details are provided in connection with FIG. 4.

[0031] In some implementations, each of the regions described in region data 134 corresponds to an element of a user interface, e.g., a window. The launch manager 150 is configured to launch a user interface element on the display in response to the user interface element corresponding to a region determined to be the region to which the user's gaze is directed. For example, if it is determined that the user's gaze is directed to a region of the display corresponding to a window, e.g., the window is included in the region, the window is the region, etc., the launch manager 150 is configured to launch the window, i.e., activate the window by highlighting its title bar and dimming the title bars of other windows. In an implementation that requires the display to be embedded in smart glasses, the user can perform operations on the contents displayed in the window, e.g., using voice commands. When the user directs his or her gaze to another region and the classification manager 140 identifies the other region as the region to which the user's gaze is directed, the launch manager 150 launches the window corresponding to the other region.

[0032] The components of the user device 120 (e.g., modules, processing unit 124) can be configured to operate based on one or more platforms (e.g., one or more similar or different platforms), which can include one or more types of hardware, software, firmware, operating systems, runtime libraries, and / or the like. In some implementations, the components of the computer 120 can be configured to operate within a cluster of devices (e.g., a server farm). In such implementations, the functionality and processing of the components of the computer 120 can be distributed across several devices of the cluster of devices.

[0033] The components of computer 120 may be or include any type of hardware and / or software configured to process attributes. In some implementations, one or more portions of the components depicted in the components of computer 120 of FIG. 1B may be or include hardware-based modules (e.g., digital signal processor (DSP), field programmable gate array (FPGA), memory), firmware modules, and / or software-based modules (e.g., modules of computer code, sets of computer-readable instructions that can be executed by a computer). For example, in some implementations, one or more portions of the components of computer 120 may be or include software modules configured to be executed by at least one processor (not shown). In some implementations, the functionality of the components may be included in different modules and / or components than those depicted in FIG. 1B, including combining functionality depicted as two components into a single component.

[0034] Although not shown, in some implementations, the components of computer 120 (or portions thereof) may be configured to operate within, for example, a data center (e.g., a cloud computing environment), a computer system, one or more server / host devices, and / or the like. In some implementations, the components of computer 120 (or portions thereof) may be configured to operate within a network. Thus, the components of computer 120 (or portions thereof) may be configured to function within various types of network environments that may include one or more devices and / or one or more server devices. For example, the network may be or may include a local area network (LAN), a wide area network (WAN), and / or the like. The network may be or may include a wireless network, and / or a wireless network implemented using, for example, gateway devices, bridges, switches, and / or the like. The network may include one or more segments, and / or may have portions based on various protocols, such as Internet Protocol (IP) and / or proprietary protocols. The network may include at least a portion of the Internet.

[0035] In some implementations, one or more of the components of computer 120 may be or include a processor configured to process instructions stored in memory. For example, input manager 130 (and / or portions thereof), classification manager 140 (and / or portions thereof), and launch manager 150 (and / or portions thereof) may be a combination of a processor and memory configured to execute process-related instructions to implement one or more functions.

[0036] In some implementations, the memory 126 may be any type of memory, such as random-access memory, disk drive memory, flash memory, and / or the like. In some implementations, the memory 126 may be implemented as two or more memory components (e.g., two or more RAM components or disk drive memory) associated with the components of the VR server computer 120. In some implementations, the memory 126 may be a database memory. In some implementations, the memory 126 may be or include a non-local memory. For example, the memory 126 may be or include a memory shared by multiple devices (not shown). In some implementations, the memory 126 may be associated with a server device (not shown) in a network and configured to serve the components of the computer 120. As shown in FIG. 1, the memory 126 is configured to store various data, including image training data 131, reference object data 136, and prediction engine data 150.

[0037] 2A illustrates exemplary regions 220(1-4) on a display 200, including activated region 220(1). In some implementations, display 200 is a transparent display embedded in smart glasses for AR applications. In some implementations, display 200 is a display on a portable computing device, such as a smartphone or tablet computer.

[0038] Display 200 includes an array of pixels 210, each of which represents a color or grayscale level that is a building block of the content to be displayed. Regions 220(1-4) each include a respective array of pixels. As shown in FIGURE 2A, regions 210(1-4) have rounded corners, although this is by no means a requirement.

[0039] The depicted region 220(1) has been identified as the region toward which the user's gaze is directed, and the other regions 220(2-4) have not been so identified. This means that identification manager 142 has determined that the user's gaze is directed toward region 220(1). Classification engine 140 is configured to perform region identification in real time, such that regions are identified with little latency when the user changes the direction of his or her gaze. Thus, the arrangement of regions shown in FIG. 2A implies that the user is currently looking at region 220(1).

[0040] FIG. 2B illustrates exemplary regions 260(1-4) on a display, including activated region 260(1), where the regions may not be contiguous. For example, as shown in FIG. 2B, region 260(1) includes two separate sub-regions that are separated from one another. Such non-contiguous regions may arise, for example, when the sub-regions correspond to different windows in a single application. In some implementations, region 260(1) may not have a rectangular shape but is contiguous; in such implementations, the region may be broken down into several rectangles, or the region may be defined as a polygon on a grid of pixels 210.

[0041] Note that the identification manager 142 does not need an eye tracker to provide the exact pixel of the gaze, but simply one that finds the region in the direction in which the gaze is directed. In the example shown in Figures 2A and 2B, if the gaze is on region 220(1-4), the camera would have to provide input data. By formulating the problem in this way, the accuracy requirements for tracking the user's gaze can be relaxed to only the accuracy required for a given user interface design. Furthermore, doing so can reduce computation time, power consumption, memory, and possibly make gaze tracking more robust.

[0042] FIG. 3A illustrates an exemplary convolutional neural network (CNN) 300 configured to classify an image of a user's eye as having gaze directed to a particular region. This is a convolutional neural network (CNN) consisting of a CNN layer, a pooling layer, and a Dense layer. The input is an image acquired from a camera attached to a smart eyeglass frame and imaging the user's eye. The output of the network is a vector of length N, where N is the number of regions. The index of the largest value in the output vector 350 gives the identified region of gaze. Optionally, a softmax layer 340 can be added after the last Dense layer, which normalizes each value inside the output vector to be between 0 and 1, and also sums all the values ​​to 1. In this case, each value represents the probability that the gaze is directed to a particular region. In some implementations, each CNN layer includes an activation function such as a rectified linear unit, a sigmoid, or a hyperbolic tangent.

[0043] In some implementations where there are two classes or regions with similar probabilities (i.e., the difference in probabilities is less than a specified threshold difference), computer 120 may delineate and display only these two regions, in some implementations computer 120 may request the user to manually select a region from the displayed two.

[0044] In some implementations, the region data 134 includes an identifier for the space outside the display, i.e., defining a new region that is not contained within any region and does not contain any pixels. In such implementations, it can be determined that the user is not looking at the display.

[0045] As shown in Figure 3A, input data 305 is introduced into a 2D convolutional layer 310(1), which is followed by a pooling layer 320(1). In the example shown in Figure 3, there are four 2D convolutional layers 310(1-4), each followed by a respective pooling layer 320(1-4). Each of the four convolutional layers 310(1-4) shown has a kernel size of 3x3 with a stride of 2. The output sizes from the 2D convolutional layers 310(1-4) are 16, 32, 64, and 128, respectively.

[0046] These convolutional layers 310(1-4) and their respective pooling layers 320(1-4) are followed by dense layers 330(1) and 330(2). In some implementations, the dense layers 330(1) and 330(2) are used to merge other branches of the CNN into the branch defined as having the convolutional layers 310(1-4) and their respective pooling layers 320(1-4). For example, the second branch of the CNN can be trained to generate output values ​​based on the region data 134, the slippage data 136, and / or the user data 138. The dense layers 330(1) and 330(2) can then provide adjustments to the classification model defined by the first branch of the CNN based on the placement of the regions, any slippage of the smart glasses sliding down the user's face, or other user characteristics.

[0047] The classification model is generated by the training manager 141 using training data 145, which represents a training dataset and identifiers for regions of gaze represented in the training dataset. The training dataset includes images of the user's eyes over time. In some implementations, the images are generated periodically, for example every 1 second, every 0.5 seconds, every 0.1 seconds, etc. These images and corresponding region identifiers are input to a loss function, and values ​​for the layer nodes are generated by optimizing the loss function.

[0048] In some implementations, the region data 134 is expressed in terms of pixel coordinates or angles. In such implementations, the training manager 141 converts the coordinates or angles to a region identifier. In some implementations, this conversion can be accomplished using a lookup table. In some implementations, this conversion is accomplished by computing the closest tile center according to, for example, a Euclidean or cosine distance between the gaze vector and a vector representing the center of the region.

[0049] The loss function, in some embodiments, includes a class-specific cross-entropy loss suitable for multi-class classification problems. Such a cross-entropy loss can be expressed as:

[0050]

number

[0051] where C is the number of classes (e.g., the number of regions across the display), p i is the label for class i, q i is the output of the network, and in some configurations is the output of a softmax layer. In some implementations, the labels for the classes are represented in a one-hot representation, and p is the label for only the class to which the example belongs. i =1, otherwise p i = 0. The above formula represents the loss per example (i.e., an image from the training data 145), and the total loss is obtained by summing the cross-entropy losses over all examples in a batch and dividing by the batch size.

[0052] In some implementations, the loss function includes a Kullback-Leibler divergence loss. Such a Kullback-Leibler divergence loss is

[0053]

number

[0054] where C is the number of classes (e.g., the number of regions across the display), p i is the label for class i, q i is the output of the network, and in some configurations the output of a softmax layer. The above formula represents the loss per example (i.e., an image from the training data), and the total loss is obtained by summing the Kullback-Leibler divergence losses over all examples in the batch and dividing by the batch size.

[0055] In some implementations, the loss function may include a triplet loss that is used to optimize a metric space defined by the area encompassing the gaze space, i.e., all the places the user can see, in which image clusters may be defined. These image clusters in turn define anchor points, and the loss function may be based in part on the distance from the anchor points. Such a triplet loss may be

[0056]

number

[0057] where f(x) represents the neural network transformation and A k represents the anchor input (i.e., image), and P k are positive examples, and N k are negative examples. The sum is the total possible triplets of (anchor, positive example, negative example). α represents a small number that introduces a margin between the positive and negative examples, used to avoid trivial solutions of all zeros.

[0058] FIG. 3B illustrates an example convolutional neural network (CNN) 360 configured to classify an image of a user's eye as having gaze directed at a particular point. CNN 360 is similar to CNN 300 (FIG. 3A) except for the presence of a new Dense layer 370(1-4). Dense layer 370(1,2) is similar to Dense layer 330(1,2) of FIG. 3A. However, Dense layer 370(3,4) is configured to generate the coordinates of the most likely point on the display where the user's gaze is directed.

[0059] As mentioned above, dense layers 330(1) and 330(2) of FIG. 3A, or 370(1-4) of FIG. 3B, allow for tuning of the classification model based on other input data, such as domain data 134, slippage data 136, and user data 138. An exemplary branch providing output to a dense layer, such as dense layer 330(1), is shown in FIG.

[0060] 4 illustrates an example convolutional layer forming another branch of the CNN 400, in this case configured to adapt to different domain or tile geometries. Note that similar branches can be defined for slippage and user adjustments. The output of the other branch is fed into a Dense layer that uses only image data 132 to adjust the output from it.

[0061] The CNN 400 takes input 405 into a convolutional layer 410 of a first branch and provides a first output through a first dense layer to a concatenation layer 440. In a second branch, the CNN inputs a tile design 420, i.e., an arrangement of regions, to a trained embedding layer 430, whose output is also input to a concatenation layer 440. The concatenation layer 440 concatenates the output of the first dense layer with the output of the embedding layer 430. The concatenated output is used to determine the region to which attention is directed.

[0062] There are several alternative approaches to determining the region to which attention is directed other than the approach shown in FIG. 4. In some embodiments, multiple models are trained, each model corresponding to a different configuration. In some embodiments, training is performed on small regions and the output is combined for each configuration. In some embodiments, the network is trained such that the network has fixed convolutional layers for all configurations, but different Dense layers for each configuration.

[0063] Similar approaches can be used for other input data. For example, for user calibration based on slippage data 136, the eyes are in a nominal location with a known (designed) pose relative to the display. During wear, the position of the glasses can be changed from this nominal location (slip on the nose, adjustment by the user, etc.). When slippage occurs, the pose between the eye and the eye tracking camera changes, and the eye image appears different than in the nominal configuration. Furthermore, the display position changes depending on the slippage of the glasses, so the gaze angle will be different. This can lead to the neural network generating erroneous classifications.

[0064] In some implementations, the eye pose (position and orientation) relative to the camera is estimated. The following approach can be taken to do this: Train a network that learns to decouple gaze classification and also learns to predict the eye position relative to the camera (or display). For predicting the eye position relative to the display, parts of the eye image that do not change with gaze (such as the corners of the eye) can be used. The selection of these image parts can be performed, for example, using an attention neural network model. These image parts can be compared in two sets of images, the first set being from a calibration phase and the second set being from images captured during gaze classification runtime.

[0065] Another alternative to using slippage data 136 to train the classification model includes assuming a finite set of possible slippage positions and performing eye position classification based on the finite set. It is also possible to use user cues to detect whether slippage has occurred. Finally, it is also possible to use a brute force approach in which the CNN is trained to be invariant to position changes due to slippage.

[0066] With respect to user data 138, the CNN can be calibrated to account for differences in eye appearance, visual axis, head shape, etc. For example, a calibration scheme may involve having a user look at a particular target on a display with known region relationships. During calibration, a camera image of the eye is recorded and the region identifier of the target is saved as a label as part of user data 138. User data 138 may then be used in the following ways:

[0067] Fine-tune an existing pre-trained neural network Training additional network branches or additional structures for calibration among others During training, which takes only the calibration data as input, we train the encoder and predict the user-specific embedding layer. Align the images so that the eye landmarks in the user data 138 are aligned with the eye landmarks in the training data 141 In some implementations, smooth and stable temporal gaze is desired, in which case temporal filtering can be used to achieve smooth and stable temporal gaze, as follows.

[0068] Compute the mean or median score from the softmax layer output from successive video frames in the image data 132. Incorporating an inductive neural network (RNN) layer configured to take the output of the convolutional layer and train the top RNN cells (e.g. long- and short-period memory cells, GRU cells, etc.) In some embodiments, classifying eye movements as fixation, saccadic eye movement or pursuit. In some implementations, the output of the CNN includes a region identifier as well as a predicted location within the display where gaze is directed.

[0069] 5 is a flow chart illustrating an example method 500 for determining an area to which a user's gaze is directed. Method 500 may be implemented by a number of software structures described in connection with FIG. 1B , which may be resident in memory 126 of computing circuitry 120 and executed by set of processing units 124, or may be implemented by software structures resident in memory of computing circuitry 120.

[0070] At 502, computer 120 receives image data (e.g., image data 132) representing at least one image of a user's eye viewing a display at a given moment, the display including a plurality of regions (e.g., regions 220(1-4)) configured to operate in an augmented reality (AR) application (e.g., in smart glasses 110), each of the plurality of regions including a plurality of pixels (e.g., pixel 210) corresponding to a respective element of a user interface.

[0071] At 504, the computer 120 identifies, based on the image data, one of a plurality of regions of the display to which the user's gaze is directed at that moment, which identification includes inputting at least one image of the user's eye to a classification engine configured to classify the gaze as being directed at one of the plurality of regions.

[0072] At 506, the computer 120 activates the user interface element to which the identified region corresponds.

[0073] 6 illustrates an example of a general purpose computing device 600 and a general purpose mobile computing device 650 that may be used with the techniques described herein. Computing device 600 is an exemplary configuration of computer 120 of FIGS. 1 and 2.

[0074] As shown in Figure 6, computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 650 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown, their connections and relationships, and their functions are meant to be merely exemplary and are not meant to limit the implementation of the invention described and / or claimed herein.

[0075] Computing device 600 includes a processor 602, a memory 604, a storage device 606, a high-speed interface 608 connecting to memory 604 and a high-speed expansion port 610, and a low-speed interface 612 connecting to a low-speed bus 614 and storage device 606. Each of these components 602, 604, 606, 608, 610, and 612 are interconnected using various buses and may be mounted on a common motherboard or may be mounted in other suitable manners. Processor 602 may process instructions for execution within computing device 600, including instructions stored in memory 604 or on storage device 606 for displaying graphical information for a GUI on an external input / output device such as a display 616 coupled to high-speed interface 608. In other implementations, multiple processors and / or multiple buses may be used along with multiple memories and multiple types of memories, as appropriate. Also, multiple computing devices 600 may be connected, each providing a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).

[0076] The memory 604 stores information within the computing device 600. In one implementation, the memory 604 is one or more volatile memory units. In another implementation, the memory 604 is one or more non-volatile memory units. The memory 604 may also be another form of computer-readable medium, such as a magnetic disk or an optical disk.

[0077] The storage device 606 can provide mass storage for the computing device 600. In one embodiment, the storage device 606 can be or include a computer-readable medium such as a floppy disk device, a hard disk device, an optical disk device or a tape device, a flash memory or other similar solid-state memory device, or an array of devices including a storage area network or other configuration of devices. The computer program product can be tangibly embedded in an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or machine-readable medium, such as the memory 604, the storage device 606, or a memory on the processor 602.

[0078] The high-speed controller 608 manages bandwidth-intensive operations for the computing device 500, while the low-speed controller 612 manages less bandwidth-intensive operations. This allocation of functionality is merely exemplary. In one embodiment, the high-speed controller 608 is coupled to the memory 604, a display 616 (e.g., via a graphics processor or accelerometer), and a high-speed expansion port 610 that can accept various expansion cards (not shown). In an embodiment, the low-speed controller 612 is coupled to the storage device 506 and a low-speed expansion port 614. The low-speed expansion port, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device, such as a switch or a router, for example, via a network adapter.

[0079] The computing device 600 can be implemented in many different forms as shown in the figure. For example, the computing device 600 can be implemented as a standard server 620, or can be formed as multiple times a group of such servers. The computing device 600 can also be implemented as part of a rack server system 624. Furthermore, the computing device 600 can be implemented in a personal computer, such as a laptop computer 622. Alternatively, components from the computing device 600 can be combined with other components in a mobile device (not shown), such as device 650. Each such device can include one or more of the computing devices 600, 650, and an entire system can also be built with multiple computing devices 600, 650 in communication with each other.

[0080] Various implementations of the systems and techniques described herein may be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or translatable on a programmable system including at least one programmable processor, which may be a special purpose or general purpose programmable processor coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0081] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages ​​and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus and / or device used to provide machine instructions and / or data to a programmable processor (e.g., magnetic disks, optical disks, memory, programmable logic devices (PLDs)), including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0082] To provide for user interaction, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) monitor or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide for user interaction, for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0083] The systems and techniques described herein may be implemented in a computing system including back-end components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including front-end components (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an embodiment of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.

[0084] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of clients and servers arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0085] A number of embodiments have been described above. It will be understood, however, that various modifications can be made without departing from the spirit and scope of the present disclosure.

[0086] It will also be understood that when an element is referred to as being on, connected to, electrically connected to, coupled to, or electrically coupled to another element, it may be directly on, directly connected to, or directly coupled to the other element, or there may be one or more intervening elements. On the other hand, when an element is referred to as being directly on, directly connected to, or directly coupled to another element, there are no intervening elements present. Although the terms directly on, directly connected to, or directly coupled to are not used in some cases throughout the detailed description, elements shown as being directly on, directly connected to, or directly coupled to may be referenced as such. The claims of this application may be amended to describe the exemplary relationships described herein or shown in the figures.

[0087] While certain features of the described embodiments have been illustrated as described herein, many modifications, substitutions, changes, and equivalents will occur to those skilled in the art at this time. It is therefore to be understood that the appended claims are intended to cover all such modifications and changes within the scope of the embodiments. It is to be understood that they are presented merely by way of example, not by way of limitation, and that various changes in form and details may be made. Any part of the apparatus and / or method described herein may be combined in any combination, except in mutually exclusive combinations. The embodiments described herein may include various combinations and / or subcombinations of the functions, components, and / or features of the different embodiments described.

[0088] Moreover, the logic flows depicted in the figures do not require the particular order shown or sequential order to achieve desired results. Moreover, other steps could be provided or steps could be removed from the described flows, and other components could be added to or removed from the described systems. Accordingly, other implementations are within the scope of the following claims.

[0089] Some examples will be described below. Example 1 A method includes receiving image data representing at least one image of a user's eye viewing a display at a given moment. The display includes a plurality of regions and is configured to operate with an augmented reality (AR) application, each of the plurality of regions including a plurality of pixels and corresponding to a respective element of a user interface. The method further includes identifying, based on the image data, a region of the plurality of regions of the display to which the user's gaze is directed at the given moment, the identifying including inputting the at least one image of the user's eye to a classification engine configured to classify the gaze as being directed at one of the plurality of regions. The method further includes activating the element of the user interface to which the identified region corresponds.

[0090] Example 2 The method described in Example 1, wherein the classification engine includes a first branch representing a convolutional neural network (CNN).

[0091] Example 3: A method as described in claim 2, wherein the classification engine is configured to generate as output a vector having a number of elements equal to the number of regions in the plurality of regions, each element of the vector including a number corresponding to a respective region of the plurality of regions, the number representing the likelihood that the user's gaze is directed at the region to which the number corresponds.

[0092] EXAMPLE 4 The method described in Example 3, wherein the classification engine includes a softmax layer configured to generate, as an output of the classification engine, a probability between zero and unity as a possibility corresponding to each of the plurality of regions, and wherein identifying the region further includes selecting, as the identified region, a region from the plurality of regions having a higher probability than the probability of each of the other regions from the plurality of regions.

[0093] Example 5. The method described in Example 3, wherein identifying the regions further includes generating image cluster data representing a set of image clusters to which the plurality of regions on the display correspond, and the classification engine includes a loss function based on distance from the set of image clusters.

[0094] Example 6: The method described in Example 1, further comprising training the classification engine, the training being based on a mapping between an image of a user's eye and a region identifier that identifies a region of the plurality of regions to which the user's gaze is directed.

[0095] Example 7 The method as described in Example 1, wherein the display is a transparent display embedded in smart glasses.

[0096] EXAMPLE 8: The method described in Example 7, wherein the classification engine further includes a second branch representing a neural network, and the method further includes outputting, from the second branch, based on the image data, a pose of the user's eyes relative to a camera mounted on the smart glasses.

[0097] Example 9. The method described in Example 8, wherein the classification engine includes an attention layer, and identifying the region further includes causing the attention layer to adjust a probability of gaze being directed to the region of the display based on the output eye pose.

[0098] EXAMPLE 10. The method described in Example 1, wherein the user is a first user and the classification engine further includes a second branch representing a neural network, the method further including inputting a parameter value indicative of a difference between the first user and the second user into the second branch, and causing the second branch to adjust a probability of a gaze being directed to a region of the display based on the parameter value.

[0099] Example 11. The method described in Example 1, wherein the user is a first user and the classification engine further includes a second branch representing a neural network, and the method further includes inputting a parameter value indicative of a geometric configuration of the plurality of regions into the second branch, and causing the second branch to adjust a probability of gaze being directed to the region of the display based on the parameter value.

[0100] EXAMPLE 12 The method described in Example 1, wherein the user is a first user and the classification engine further includes a second branch representing a neural network, the method further including inputting a parameter value indicative of temporal smoothness of the image data into the second branch, and causing the second branch to adjust a probability of gaze being directed to a region of the display based on the parameter value.

[0101] Example 13. A computer program product comprising a non-transitory storage medium, the computer program product including code that, when executed by processing circuitry of a computer, causes the processing circuitry to perform a method, the method including receiving image data representing at least one image of a user's eye looking at a display at a certain moment, the display including a plurality of regions, the display being configured to operate with an augmented reality (AR) application, each of the plurality of regions including a plurality of pixels and corresponding to a respective element of a user interface, and identifying, based on the image data, a region of the plurality of regions of the display to which the user's gaze is directed at the moment, the identifying including inputting the at least one image of the user's eye to a classification engine configured to classify the gaze as being directed at one of the plurality of regions, and activating an element of the user interface to which the identified region corresponds.

[0102] Example 14. The computer program product of Example 13, wherein the classification engine includes a first branch representing a convolutional neural network (CNN).

[0103] Example 15. The computer program product of Example 13, wherein the classification engine is configured to generate as an output a number corresponding to each of the plurality of regions, the number representing a likelihood that the user's gaze is directed toward the region to which the number corresponds.

[0104] Example 16: The computer program product of Example 13, wherein the method further includes training the classification engine, the training being based on a mapping between an image of the user's eye and a region identifier that identifies a region of the plurality of regions to which the user's gaze is directed.

[0105] Example 17. The computer program product of Example 13, wherein the display is a transparent display embedded in smart glasses.

[0106] EXAMPLE 18. The computer program product of Example 17, wherein the classification engine further includes a second branch representing a neural network, and the method further includes outputting, from the second branch, a position and orientation of the user's eyes relative to a camera mounted on the smart glasses based on the image data.

[0107] Example 19: The computer program product of Example 18, wherein the classification engine includes an attention layer, and wherein identifying the region further includes causing the attention layer to adjust a probability of gaze being directed to the region of the display based on the output eye position and orientation.

[0108] Example 20. An electronic device, the electronic device comprising: a memory; and processing circuitry coupled to the memory. the processing circuitry is configured to receive image data representing at least one image of a user's eye looking at a display at a given moment, the display including a plurality of regions, configured to operate with an augmented reality (AR) application, each of the plurality of regions including a plurality of pixels and corresponding to a respective element of a user interface, and configured to identify, based on the image data, a region of the plurality of regions of the display to which the user's gaze is directed at the given moment, the identifying including inputting the at least one image of the user's eye to a classification engine configured to classify the gaze as being directed at one of the plurality of regions, and configured to activate an element of the user interface to which the identified region corresponds.

Claims

1. 1. A method comprising: receiving at least one image of a user's eye observing the display at a given time; the display includes a plurality of regions and is configured to operate with an augmented reality (AR) application, the plurality of regions corresponding to a user interface; The method comprises: and identifying, based on the at least one image of the eye, one of the plurality of regions of the display to which gaze is directed at the time; the identifying includes inputting the at least one image of the eye of the user to an engine configured to classify the gaze as being directed at the region; The method comprises: and activating an element of the user interface to which the region corresponds.

2. The method of claim 1 , wherein the engine includes a first branch representing a convolutional neural network (CNN).

3. 2. The method of claim 1, wherein the engine is configured to generate as output a vector having a number of elements equal to the number of regions in the plurality of regions, the elements of the vector including a number corresponding to each region of the plurality of regions, the numbers representing a likelihood that the user's gaze is directed toward the region to which the number corresponds.

4. the engine includes a softmax layer configured to generate, as an output of the engine, a probability between zero and unity as a likelihood corresponding to a region of the plurality of regions; The identification of the region comprises: selecting as the identified region from among the plurality of regions a region having a probability higher than the probabilities of other regions from the plurality of regions. The method of claim 3 further comprising:

5. The identification of the region comprises: generating image cluster data representing a set of image clusters to which the plurality of regions on the display correspond; Further comprising: the engine includes a loss function based on distance from the set of image clusters; The method according to claim 3.

6. The method comprises: training the engine; The method of claim 1 , wherein the training is based on a mapping between an image of the user's eye and a region identifier that identifies a region of the plurality of regions to which the user's gaze is directed.

7. The method of claim 1 , wherein the display is a transparent display embedded in smart glasses.

8. the engine further includes a second branch representing a neural network; The method comprises: and from the second branch, outputting an eye pose of the user relative to a camera mounted on the smart glasses based on the at least one image of the eye of the user. Further comprising: The method of claim 7.

9. the engine includes an attention layer; The identification of the region comprises: adjusting the probability of the gaze being directed to the plurality of regions of the display based on the pose of the eye to the attention layer. Further comprising: The method according to claim 8.

10. the user is a first user, the engine further includes a second branch representing a neural network; The method comprises: inputting a parameter value into the second branch indicative of a difference between the first user and a second user; adjusting a probability of the gaze being directed to the plurality of regions of the display based on the parameter value to the second branch; Further comprising: The method of claim 1.

11. The engine further includes a second branch representing a neural network; The method comprises: inputting parameter values ​​into the second branch that are indicative of a geometric configuration of the plurality of regions; adjusting a probability of the gaze being directed to the plurality of regions of the display based on the parameter value to the second branch; Further comprising: The method of claim 1.

12. The engine further includes a second branch representing a neural network; The method comprises: inputting a parameter value indicative of the temporal smoothness of said at least one image into said second branch; adjusting a probability of the gaze being directed to the plurality of regions of the display based on the parameter value to the second branch; Further comprising: The method of claim 1.

13. A method comprising: a method for determining whether a first bit of a signal is received from a processor; receiving at least one image of a user's eye observing a display at a time, the display including a plurality of regions and configured to operate with an augmented reality (AR) application, the plurality of regions corresponding to a user interface; identifying a region of the plurality of regions of the display to which gaze is directed at the time based on the at least one image of the eye, the identifying comprising inputting the at least one image of the user's eye to an engine configured to classify the gaze as being directed to the region; activating an element of the user interface to which the region corresponds.

14. 14. The computer program product of claim 13, wherein the engine includes a first branch representing a convolutional neural network (CNN).

15. 14. The computer program product of claim 13, wherein the engine is configured to generate as output a number corresponding to the plurality of regions, the number representing a likelihood that the gaze is directed at the region to which the number corresponds.

16. The method comprises:

14. The computer program product of claim 13, further comprising training the engine, the training being based on a mapping between an image of the user's eye and a region identifier that identifies a region of the plurality of regions to which the user's gaze is directed.

17. The computer program product of claim 13 , wherein the display is a transparent display embedded in smart glasses.

18. the engine further includes a second branch representing a neural network; The method comprises: and from the second branch, outputting an eye pose of the user relative to a camera mounted on the smart glasses based on the at least one image of the eye of the user. Further comprising:

18. A computer program product as claimed in claim 17.

19. the engine includes an attention layer; The identification of the region comprises: and adjusting the probability of the gaze being directed to the region of the display based on the output eye pose to the attention layer. Further comprising:

20. A computer program product as claimed in claim 18.

20. An electronic device, the electronic device comprising: Memory, control circuitry coupled to the memory; The control circuitry comprises: configured to receive at least one image of an eye of a user observing a display at a given time, said display including a plurality of regions and configured to operate with an augmented reality (AR) application, said plurality of regions corresponding to a user interface; configured to identify, based on the at least one image, one of the plurality of regions of the display to which gaze is directed at the time, the identifying including inputting the at least one image of the eye of the user to a classification engine configured to classify the gaze as being directed to the region; the region is configured to activate a corresponding element of the user interface; electronic equipment.

21. The electronic device of claim 20, wherein the classification engine includes a first branch representing a convolutional neural network (CNN).