Eye Tracking Device, Eye Tracking Method, and Computer-Readable Medium

The proposed gaze tracking system employs an event-based optical sensor and neural networks to generate inference frames, addressing the limitations of frame-based cameras and hand-designed accumulation schemes, resulting in improved reliability and efficiency.

JP7699839B2Active Publication Date: 2025-06-30INITIATION ARE GAME
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022580817
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-03
Filing Date
2021-06-30
Publication Date
2025-06-30
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

Existing gaze tracking systems rely on frame-based cameras, which are slow and generate large amounts of data, and even event-based systems often require hand-designed accumulation schemes that result in noisy and artifact-prone images.

Method used

A gaze tracking device utilizing an event-based optical sensor and a controller that processes the event stream to generate an inference frame using a first artificial neural network, which is then input to a machine learning module to estimate eye fixation parameters without relying on conventional image frames.

Benefits of technology

This approach enables high-quality inference frames, improving the estimation of eye fixation parameters and reducing the need for hand-designed accumulation schemes, thus enhancing the reliability and efficiency of gaze tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699839000001
    Figure 0007699839000001
  • Figure 0007699839000002
    Figure 0007699839000002
  • Figure 0007699839000003
    Figure 0007699839000003
Patent Text Reader

Abstract

The present invention relates to an eye-tracking device, an eye-tracking method, and a computer-readable medium. The eye tracking device includes an event-based optical sensor (1) configured to receive radiation (12) reflected from a user's eye (2) and produce a signal stream (3) of events (31), each event (31) corresponding to a detection of a temporal change in the received radiation at one or more pixels of the optical sensor (1); and a controller (4) connected to the optical sensor (1), the controller (4) configured to: a) receive the signal stream (3) of events (31) from the optical sensor (1); b) generate an inference frame (61) based on at least a portion of the stream (3) of events (31); c) use the inference frame (61) as input to a machine learning module (6) and operate the machine learning module (6) to obtain output data; and e) extract information related to the user's eye (2) from the output data, the controller (4) configured to generate the inference frame (61) using a first artificial neural network (5).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a gaze tracking device, a gaze tracking method, and a computer-readable medium.

Background Art

[0002] Gaze tracking generally refers to monitoring the movement or fixation of the eyes of a human, generally called a user. However, the user may of course be any other human or animal having eyes that can change the direction in which they are looking in their eye sockets.

[0003] One possible technique for tracking a user's fixation is to use a conventional video camera or photographic camera that periodically acquires full frames or conventional frames of an eye image. Then, in order to determine the position of the pupil at the time the frame is captured, a controller connected to the camera analyzes each of those image frames, thereby making it possible to estimate the direction in which the user is looking. This method requires the use of a frame-based camera, such as a video camera or photographic camera, that acquires an image of the eye that the controller analyzes. Such conventional cameras or frame-based cameras are often slow. Also, they generate a large amount of data that needs to be transferred between the camera and the controller.

[0004] The gaze tracking process can be accelerated by using an event-based camera, or an event-based sensor also known as a Dynamic Vision Sensor (DVS). EP3598274A1 describes a system with multiple cameras, where one of the cameras is an event-based camera or DVS. However, this known system also relies on a second frame-based camera. Similarly, the publication "Event Based, Near Eye Gaze Tracking Beyond 10,000Hz", Angelopoulos, Anastasios N. et al., preprint arXiv:2004.03577 (2020) uses ellipse detection on conventional image frames in addition to event-based DVS data for the task of gaze tracking. Although conventional computer vision approaches are used, the authors state that a deep learning-based extraction method would be an easy extension of their technique. Thus, also in this case, the gaze tracking process at least partially depends on conventional image frames acquired by a frame-based camera. The dependence on the availability of eye image frames requires the gaze tracking system to acquire full frames before it can accurately predict the position of the eye. Some systems can utilize interpolation to predict future states, but the time taken to acquire full frames defines the worst-case delay.

[0005] US10466779A1, which describes a gaze tracking system using DVS data and outlines a method for converting the received DVS data into an intensity image, similar to a conventional frame, follows a different approach. This algorithm uses the mathematical properties of the DVS stream. Conventional computer vision approaches are used to predict fixation and pupil characteristics from the intensity image acquired as described above.

[0006] For gaze tracking, a method that combines the acquisition of the output of a fully event-based sensor with a machine learning approach using a convolutional neural network is described in WO2019147677A1. In WO2019147677A1, a system is described in which events from an event camera are accumulated to create either an intensity image, a frequency image, or a timestamp image, and those images are then fed into a neural network algorithm to predict various fixation parameters. The described system uses a static accumulation regime designed by hand, which is a common and well-known technique for creating an approximation of an intensity image from event data. A drawback of this approach is that the images are noisy and tend to exhibit artifacts from past pupil positions. Downstream frame-based convolutional neural networks, such as those described in WO2019147677A1, can handle noisy data and temporal artifacts such as artifacts that are unavoidable when accumulating DVS events, and thus require more complex neural networks.

Prior Art Documents

Patent Documents

[0007]

Patent Document 1

Patent Document 2

Patent Document 3

Non-Patent Documents

[0008]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0009] An object of the present invention is to propose a device and a method for more reliably tracking the movement of a user's eyes.

Means for Solving the Problems

[0010] This object is achieved by the present invention by providing a gaze tracking device having the features of claim 1, a gaze tracking method having the features of claim 14, and a computer-readable medium having the features of claim 15. Further advantageous embodiments of the present invention are the subject matter of the dependent claims.

[0011] According to the present invention, the gaze tracking device includes an event-based optical sensor and a controller connected to the sensor. Radiation reflected from the user's eyes is received by an event-based optical sensor configured to create a signal stream of events in response to the radiation. This signal stream is sent to the controller, which performs various processes on the signal stream to obtain the results of the gaze tracking process. Thus, the controller may comprise at least a processing unit and a memory for performing the analysis described below. Hereinafter, the event-based optical sensor is simply referred to as an event sensor.

[0012] Sensors, in particular, dynamic vision sensors, comprise a number of individual pixels arranged in an array, and each pixel has a photosensitive cell or photosensitive region. When detecting a temporal change in the incident light hitting the photosensitive cell, an event signal, simply referred to herein as an "event", is generated. Thus, each event in the signal stream of events created by the sensor corresponds to the detection of a temporal change in the received radiation in one or more pixels of the photosensor. Each event may in particular include the position of the corresponding pixel in the array, and an indicator indicating the polarity, and optionally the magnitude of the temporal change and the time at which the change occurred. The events are sent to the controller as part of the signal stream for further processing.

[0013] The controller receives the event signal stream, generates an inference frame using a first artificial neural network to be used as an input to the machine learning module, operates the machine learning module to obtain output data, and is configured to extract the required information related to the user's eye from the output data. Advantageously, the output data generated by the machine learning module is the required information such as the position / orientation of the pupil.

[0014] The inference frame can be defined as the output of the first artificial neural network and the frame that is the input to the machine learning module. The term inference frame may refer to a 3D tensor of dimensions width, height, channels. The channels are a collection of various representations of the data. These various representations may include linear or non-linear intensities such as in particular logarithms, scales, spatial or temporal derivatives, intensities and / or phases of frequency components.

[0015] The first neural network is trained using corresponding input data and output data. This input data and output data can be generated using simulation software. For the training data, while the components of the inference frame, which is the output of the network, are created using standard image processing techniques, the event input stream is calculated using a mathematical model of the event sensor. The selection of the representation is advantageously made so as to optimize the performance of the second neural network. The first neural network is preferably trained to create the best possible approximation of those representations.

[0016] By having the first neural network directly create all representations, the system can achieve better reconstruction performance than using a standard image processing approach for a single learned representation. By having multiple representations as input to the second neural network, the second neural network can achieve better performance in estimating eye fixation parameters than having only a single representation.

[0017] The present invention is based on the concept of using a first artificial neural network to generate an inference frame before passing the generated inference frame to a machine learning module. Thus, while it is known from the prior art to use a static accumulation regime designed by hand, a common and well-known technique for creating an inference frame from event data and inputting it into a machine learning module, the present invention proposes a system that does not utilize a hand-designed accumulation scheme and instead assigns the generation of the inference frame to a neural network. By using this approach, a very high-quality inference frame can be realized, leading to good estimation by a subsequent machine learning module or process. Furthermore, according to the present invention, the pupil position is determined based only on the signal stream of the events without the need to access the image frames collected by a conventional frame-based camera.

[0018] In order for the machine learning module to be able to process the input data, the input data needs to be provided in an appropriate form. The first artificial neural network exists to convert a stream of events into an inference frame and can thus be handled by the machine learning module. Preferably, the inference frame has the same number of pixels as the event sensor. However, the inference frame needs to be distinguished from the conventional image frame of the eye that can be provided by a conventional camera. The inference frame includes a plurality of frame pixels arranged in an array and can further be an approximation of the eye's image, but depending on the parameters and the response of the first artificial neural network being used, it is not necessarily intended as such an approximation. In particular, the first artificial neural network does not need to be configured such that the output it provides is an approximation of the monitored eye. Rather, beneficially, the first artificial neural network is configured to create an inference frame in a form that improves or maximizes the performance of the subsequent machine learning module. A suitable inference frame can be any kind of frame that contains the information necessary for the machine learning module to process. A suitable inference frame can include, for example, an approximate intensity on a linear or non-linear scale, a first-order spatial derivative of the approximate intensity, and / or a higher-order spatial derivative of the approximate intensity.

[0019] According to an advantageous embodiment, the controller is configured to convert a portion of the event stream into a sparse tensor and use the sparse tensor as an input for a first artificial neural network. The tensor may in particular have the dimensions of the event sensor, i.e., W×H×1, where W and H are the width and height of the sensor in pixels. The tensor contains all zeros except for the coordinates x, y at which the sensor reported an event in the corresponding pixel. If the event is positive, the tensor contains a 1 at coordinates x, y, while for a negative event it contains a -1. Here, positive and negative indicate the polarity of the change in light intensity recorded as an event. If the event sensor is configured to notify not only about the polarity but also about the magnitude of the change in light intensity of the pixel, the value of the tensor at the coordinates x, y of the pixel is the value of this signed magnitude. If multiple events occur at the same pixel in a batch of events corresponding to the same tensor, only the first event is considered. On the other hand, multiple events at different pixels will be included separately in the tensor.

[0020] The controller may split the event stream into one or more portions of events, in particular consecutive events. Each such portion may contain a predetermined number of events. Or the portion may represent a predetermined time interval and may include all events occurring within that time interval, time slot or time duration. In an implementation with a sparse tensor, the controller may be configured to generate the sparse tensor based on a predetermined number of events or based on events occurring within a predetermined time interval or predetermined time duration.

[0021] Advantageously, the first artificial neural network is a recurrent neural network, i.e., an RNN. An RNN means that, in particular, in, for example, the last layer, the first layer, and / or some layer in between, the last output from the RNN is fed back or supplied to the RNN in some way. Advantageously, the output of the RNN is fed back to the input of the RNN. In particular, the output of the RNN after one execution of the RNN algorithm can be used as one of a plurality of inputs to the RNN during successive executions of the RNN algorithm, together with, for example, other tensors to be processed.

[0022] The first artificial neural network includes one layer of it, or a plurality of layers where one or more of the layers can also be convolutional layers. Therefore, when the first artificial neural network is an RNN, it can also be called a convolutional recurrent neural network. The first artificial neural network can also include a concatenation layer, in particular as the first layer, to combine or concatenate the output after the execution of the neural network algorithm with new inputs for successive executions. Furthermore, the first neural network can be equipped with one or more non-linear activation functions, in particular a rectifier and / or a normalization layer.

[0023] Preferably, one or more of the layers of the first artificial neural network are memorization layers. The memorization layer stores the results of that layer during the latest pass, i.e., during the latest execution of the neural network algorithm. Implementation by the memorization layer enables, during all passes, only the values stored in the memorization layer to be updated according to the non-zero tensor elements in the input sparse tensor. This technique can significantly accelerate the neural network inference speed and can lead to a higher-quality inference framework for the continuous machine learning module in this device.

[0024] The idea behind the use of one or more layers of memoization layers is that when the changes in the previous layer are extremely small, it is sufficient to only update the internal values / states of the affected neural network. This can save the processing power when updating the state in the neural network. In addition to convolutional layers, non - linear activation functions and / or normalization layers can also be memoized. Advantageously, all convolutional layers and / or all non - linear activation functions are of the type that can be memoized. In this case, in all layers, only the values directly affected by the changes in the input are updated. This input can be both the sparse tensor of the neural network and the latest result. Therefore, only the values directly affected by the sparse matrix input are updated. These values are recalculated taking all inputs into account. Memoization can be applied to any layer of any type of artificial neural network, but here, advantageously, it is applied to a first neural network which can in particular be an RNN.

[0025] According to a preferred embodiment, the machine learning module comprises a second artificial neural network. In other words, the controller is configured to use the inference frame as an input to the second artificial neural network and operate the second artificial neural network to obtain output data. In particular, this second artificial neural network can be a neural network trained by backpropagation such as a convolutional neural network. In an alternative embodiment, a part of the second artificial neural network can be a convolutional neural network, in particular its common backend, as will be described in detail later.

[0026] Advantageously, the second artificial neural network is already trained using known methods for training neural networks, in particular the "adam" or "stochastic gradient descent (SGD)" optimizer. This training can be performed using a large amount of annotated data that is recorded and manually annotated. Some select layers can be retrained or fine-tuned by the user of the device. In this case, the user may need to perform tasks such as looking at a specific location on a computer screen, for example for calibration purposes. The device then collects data from the sensors and fine-tunes the last layer of the second artificial neural network, in particular the last trainable layer, to better suit the individual characteristics and behavior of the user. The first and second artificial neural networks can be trained either individually or simultaneously as one system. The second neural network can also be preferably trained as one neural network using various losses applied to various front-ends when it comprises a common backend and multiple front-ends.

[0027] Preferably, the convolutional neural network comprises a common backend and one or more front-ends. The common backend performs a bulk analysis of the input, but the output is then analyzed by a front-end that is configured and / or trained to estimate particular attributes that are to be produced, in particular, by the device. These attributes can include the user's gaze direction, the pupil center position of the user's eye, the pupil contour of the user's eye, the pupil diameter of the user's eye, the eyelid position of the user's eye, the pupil shape of the user's eye, person identification information related to the user, and / or prediction of the pupil movement of the user's eye.

[0028] These can be important attributes that may be of interest when the eye-tracking device is acquired. Thus, the controller is advantageously configured to extract one or more of these attributes from the output data information even when the machine learning method does not include a convolutional neural network having a common backend and one or more frontends. Advantageously, this backend is a full convolutional encoder / decoder system. The above-described person identification information may be valid when the device is used by multiple users, and in that case, the determined person identification information may be useful in identifying which of the users is currently using the device.

[0029] According to an advantageous embodiment, the convolutional neural network, or a part of the second neural network implemented as a convolutional neural network, in particular the common backend part, is realized using at least in part an encoder / decoder scheme comprising one or more encoder blocks and one or more decoder blocks. Advantageously, the common backend may comprise two, four, or six encoder blocks and / or two, four, or six decoder blocks. Advantageously, the convolutional neural network or its common backend is a full convolutional encoder / decoder system. Such an encoder / decoder neural network enables the implementation of feature learning or representation learning. Each of the encoder blocks and / or decoder blocks may in particular comprise at least two convolutional layers. The convolutional neural network may further comprise an identity skip connection between the encoder block and the decoder block. Such a skip connection or shortcut, also called a residual connection, is used to skip one or more layers of the neural network in order to enable the training of deeper neural networks and helps with the vanishing gradient problem. Advantageously, the skip connection connects the first encoder of the encoder / decoder system to the last decoder, and / or the second encoder to the second last decoder, and so on.

[0030] The event sensor of the gaze tracking device may be provided with an optical filter, particularly an infrared band-pass filter, so that the optical filter detects only radiation from a specific wavelength range such as infrared (IR) rays. Although the radiation reflected from the eye can be ambient light, such an approach has the drawback that it can create parasitic signals due to low radiation levels or light disturbances in some cases. Thus, beneficially, a radiation source is provided that is configured to send radiation to the user's eye such that the radiation is reflected from the eye and received by the event sensor. In order for the radiation source not to interfere with the user, the radiation created by the radiation source needs to be good outside the visible regime. Preferably, the radiation source is an infrared (IR) emitter.

[0031] Beneficially, the gaze tracking device comprises a body-worn device, particularly a head-mounted device, for mounting the gaze tracking device on the user's body, particularly his or her head. The application fields of such a device may include virtual reality or augmented reality and may support the implementation of foveated rendering.

[0032] According to a further aspect of the invention, a gaze tracking method and a computer-readable medium are provided. Any features described above in connection with the gaze tracking device may be used alone or in suitable combination in the gaze tracking method or the computer-readable medium.

[0033] Some examples of embodiments of the invention are described in more detail below with reference to the accompanying schematic diagrams.

Brief Description of the Drawings

[0034]

Figure 1

Figure 2

Figure 3a

Figure 3b

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0035] FIG. 1 is a schematic diagram of the setup of a gaze tracking device according to the prior art. The radiation source 10 emits radiation 12, which is reflected from the user's eye 2 and tracked. The reflected radiation 12 is incident on a conventional camera, i.e., a frame-based camera 1'. The frame-based camera 1' detects the incident radiation 12 and generates a sequence 11 of video or image frames, which is transmitted to a conventional controller 4'. The controller 4' can analyze the video or image frames and determine various parameters of the eye 2 under monitoring, particularly the fixation direction.

[0036] A schematic diagram of the setup of the gaze tracking device according to a preferred embodiment of the present invention is shown in FIG. 2. Similar to the prior art case shown in FIG. 1, the radiation source 10 emits radiation 12, and the radiation 12 is reflected from the user's eye 2. The reflected radiation 12 is incident on an event-based sensor or event sensor 1. The radiation source 10, the event sensor 1, and an optical lens (not shown) for converging the radiation are attached to a head-mounted device (not shown) such as glasses, a virtual reality (VR) or augmented reality (AR) device. The event sensor 1 is equipped with an infrared bandpass filter. The movement of the eye causes a change in the light intensity of the radiation 12 reflected from the user's eye 2. These light intensity changes or fluctuations are captured by the event sensor 1. In response, the event sensor 1 generates a stream 3 of light change events, and the stream 3 is sent to the controller 4 for processing. This processing includes preprocessing the event stream 3 to obtain a suitable input for a recurrent neural network (RNN) as described below, performing the RNN on the preprocessed data to obtain an inference frame, and performing a convolutional neural network (CNN) to estimate the desired attributes.

[0037] An event is a 4-tuple defined as (p, x, y, t), where p is either the polarity of the light change (positive means an increase in light intensity and negative means a decrease in light intensity), or the magnitude of the signed change in linear, logarithmic, or other scaling of the light intensity change. x and y are the pixel coordinates of the event, and t is the exact timestamp of the observed event. Such events are shown in FIGS. 3a and 3b. FIGS. 3a and 3b visualize the preprocessing steps executed by the controller 4 upon receipt of the events. One or more events shown on the left side of the arrow are accumulated and converted into the sparse tensor or sparse matrix shown on the right side of the arrow. FIG. 3a shows a single event that is each converted into a sparse tensor. The sparse tensor is filled with zeros except for the (x, y) position corresponding to the (x, y) coordinates of the corresponding event containing the value p. In contrast, FIG. 3b shows that the events are converted into a sparse tensor in pairs.

[0038] There may be more events that are gathered into one sparse tensor. Further, as an alternative to creating each sparse tensor based on a predetermined number of events, it is also possible to be based on the events present within a predetermined time interval or time duration.

[0039] The various processing stages of the line-of-sight tracking device and the types of data transferred between these stages are shown in FIG. 4. The event sensor 1 generates a stream 3 of events 31. These events 31 are transferred to the controller 4 and processed by the preprocessing module 41 to obtain the sparse tensor 51 as described above with reference to FIGS. 3a / 3b. The tensor 51 is used as an input for a first neural network, here a recurrent neural network (RNN) 5, to generate an inference frame 61. Finally, the inference frame 61 is fed into a convolutional neural network (CNN) 6, and the convolutional neural network (CNN) 6 estimates the pupil parameters, particularly the fixation direction.

[0040] Figure 5 shows an overview of a possible setup of an RNN-based algorithm. The RNN algorithm is triggered based on the availability of data. Thus, each time a new sparse tensor is generated, the RNN algorithm is called to generate a new inference frame. As a first step, the sparse tensor obtained from the DVS event stream is input into the RNN 501. The RNN maintains, along with the internal state of the last generated inference frame, optionally other intermediate activation maps. The RNN network estimates a new state based on the last state and the sparse input tensor. To achieve high performance, the sparsity of the input tensor is utilized to update only the deeper values, which are changes affected in the input tensor or input matrix. Figure 6 shows this sparse update regime, which will be further explained below.

[0041] The concatenation layer 502 concatenates or joins, in the channel dimension, the sparse input tensor and the inference frame generated during the previous processing of the RNN. Next, the first convolutional layer 503 performs a convolution on this concatenation. Then, the second convolutional layer 505 operates on the output of the first convolutional layer 503. Next, the output of the first convolutional layer 503 is normalized (batch normalization) 507. The RNN further includes two non-linear activation functions 504, 506. This layer structure generates an inference frame 508, which is used as one of the inputs for the concatenation layer 502.

[0042] Each of the RNN layers 503, 504, 505, 506, 507 is memoized. A "memoized" layer stores the results of the latest path. In this embodiment, in all paths, only the values in the RNN that depend on the tensor elements that are non-zero in the sparse input tensor are updated. This technique significantly accelerates the RNN inference speed and enables the use of higher-quality inference frames for successive CNN estimators. FIG. 6 shows this approach using a 3×3 convolutional kernel. The previous inference frame 61 and the sparse input tensor 51, or short sparse tensor 51, are concatenated as shown on the left. On the right, it is shown that only a subset 602 of the activations in the convolutional layer are updated due to the sparsity of the input tensor 51.

[0043] FIG. 7 is a diagram showing an overview of the concept of the architecture of a CNN that receives an inference frame 701 from an RNN. This CNN has a common backend 702 that generates an abstract feature vector 703 or an abstract feature map. This feature vector is used as an input for various frontend modules 704 that follow the common backend 702. The exemplary frontend shown in FIG. 7 includes a pupil position estimation module 704a that outputs a resulting pupil position 704c and a blink classification module 704b that outputs information 704d regarding whether the eye is open or closed. Other modules 704e of the frontend 703 can be provided for other attributes to be determined.

[0044] FIG. 8 shows a possible implementation of the backend in more detail, while possible implementations of the two frontends shown in FIG. 7 are presented in more detail in FIGS. 9 and 10.

[0045] The common backend shown on the left side of FIG. 8 is based on an encoder / decoder scheme and includes two encoder blocks 802, 803 and two decoder blocks 804, 805 following those encoder blocks. At the end of the backend, there is a synthesis layer 806 that produces the result of the abstract feature map 808 to be further processed by the frontend. As can be seen from the arrow connecting from the second encoder block 803 to the synthesis layer 806, the abstract feature map 808 includes information from the outputs of the second encoder block 803 and the second / last decoder block 805. Further, as can be seen from the two arrows on the left side, there are two identity skip connections between the encoder stage and the decoder stage. These skip connections connect from the first encoder block 802 to the second decoder block 805 and further from the second encoder block 803 to the first decoder block 804.

[0046] On the right side of FIG. 8, both the encoder block 802 and the decoder block 804 are shown in more detail. All of the two encoder blocks 802, 803 can have, in particular, the same or very similar architectures. This is also the case for the two decoder blocks 804, 805. The encoder block 802 includes max-pooling 817 together with two convolutional layers 811, 814, two non-linear activation functions 812, 815, and batch normalization 813, 816. The decoder block 804 includes an upsampling layer 821, a concatenation layer 822, two convolutional layers 823, 826, two non-linear activation functions 824, 827, and batch normalization 825, 828. Further, as described above, the common backend can include a different number of encoder blocks and / or decoder blocks from the system shown in this specification, but has the same architecture encoder blocks and / or decoder blocks. For example, two, four, or six encoder blocks and / or two, four, or six decoder blocks can be provided.

[0047] Two exemplary front-ends shown in FIG. 7 are shown in more detail by FIGS. 9 and 10. The front-end of FIG. 9 is for pupil position identification or pupil position estimation. This front-end receives the feature vector 901 from the back-end and applies the feature selection mask 902. Subsequently, a convolutional layer 903, a non-linear activation function 904, and batch normalization 905 follow. The result after batch normalization 905 is sent to the spatial softmax layer 906 to estimate the pupil position 907 and further to the fully connected layer 908 to estimate the pupil diameter 909. The front-end of FIG. 10 is for blink detection. This front-end receives the feature vector 911 from the back-end and applies a different feature selection mask 912. This front-end comprises a first fully connected layer 913, a non-linear activation layer 914, a second fully connected layer 915, and a softmax activation function 916 to provide information 917 regarding whether the eye is open or closed.

[0048] The entire CNN with the back-end and front-ends is trained as one neural network, and different losses are applied to different front-ends. It is also possible to first train a CNN with only one front-end attached, then freeze one or more layers, and then train a CNN with one or more front-ends, possibly all front-ends attached. Another training scheme where training between different front-ends is done alternately is also possible. Different from the RNN network for inference frame generation, the CNN is very, more complex and requires more processing power to generate results. This is the reason why it is preferred that the CNN is not triggered by the availability of data but instead by the requirements of an application for new predictions.

Explanation of Signs

[0049] 1’ Frame-based camera, conventional camera 11 Sequence of video frames 10 Radiation source, IR emitter 1 Event-based optical sensor, event sensor, DVS sensor 12 Radiation incident on and reflected from the eye 2 User's eye 3 Event signal stream 31 Event 4’ Prior art controller 4 Controller 41 Input processing module 5 First artificial neural network, recurrent neural network, RNN 51 Sparse tensor 6 Machine learning module, second artificial neural network, convolutional neural network, CNN 61 Inference frame

Claims

1. - An event-based optical sensor (1) configured to receive radiation (12) reflected from a user's eye (2) and create a signal stream (3) of events (31), each event (31) corresponding to the detection of a temporal change in the received radiation at one or more pixels of the optical sensor (1), the optical sensor (1); - A controller (4) connected to the optical sensor (1), a) receiving a signal stream (3) of events (31) from the optical sensor (1), b) generating an inference frame (61) based on at least a portion of the stream (3) of events (31), the inference frame (61) comprising a plurality of frame pixels arranged in an array, c) operating the machine learning module (6) to obtain output data using the inference frame (61) without using the stream (3) of events (31) as an input to the machine learning module (6), e) extracting information related to the user's eye (2) from the output data a controller (4) configured as such; A gaze tracking device comprising: The gaze tracking device, characterized in that the controller (4) is configured to generate the inference frame (61) using a first artificial neural network (5).

2. The gaze tracking device according to claim 1, characterized in that the controller (4) is configured to convert the portion of the stream (3) of events (31) into a sparse tensor (51) and use the sparse tensor (51) as an input for the first artificial neural network (5).

3. The gaze tracking device according to claim 2, characterized in that the controller (4) is configured to generate a sparse tensor (51) based on a predetermined number of events (31) or based on events (31) occurring within a predetermined time interval or time duration.

4. The gaze tracking device according to any one of claims 1 to 3, characterized in that the controller is configured such that the first artificial neural network (5) is a regression neural network.

5. The line-of-sight tracking device according to any one of claims 1 to 4, characterized in that the controller is configured such that the first artificial neural network (5) has at least one memorization layer.

6. The controller is configured to extract, from the output data, the user's gazing direction, the pupil center position of the user's eye, the pupil contour of the user's eye, the pupil diameter of the user's eye, the eyelid position of the user's eye, the pupil shape of the user's eye, person identification information related to the user, and / or prediction of pupil movement of the user's eye, the line-of-sight tracking device according to any one of claims 1 to 5.

7. The line-of-sight tracking device according to any one of claims 1 to 6, characterized in that the machine learning module (6) comprises a second artificial neural network, and the controller is configured to utilize the inference frame (61) as an input to the second artificial neural network and operate the second artificial neural network to obtain the output data.

8. The line-of-sight tracking device according to claim 7, characterized in that the controller is configured such that the second artificial neural network (6) comprises a common backend and one or more frontends.

9. The line-of-sight tracking device according to claim 7 or 8, characterized in that the second artificial neural network (6) is a convolutional neural network.

10. The line-of-sight tracking device according to claim 9, characterized in that the controller is configured such that the convolutional neural network is realized at least partially using an encoder / decoder scheme, and comprises one or more encoder blocks and one or more decoder blocks.

11. The line-of-sight tracking device according to claim 10, characterized in that the controller is configured such that the convolutional neural network comprises an identity skip connection between the encoder block and the decoder block.

12. The controller is configured such that each of the encoder block and / or the decoder block includes at least two convolutional layers, the line-of-sight tracking device according to claim 10 or 11.

13. A radiation source configured to send radiation (12) to the user's eye (2) such that the radiation (12) is reflected from the user's eye (2) and received by the event-based optical sensor (1), the line-of-sight tracking device according to any one of claims 1 to 12.

14. - Receiving a signal stream (3) of events (31) created by the event-based optical sensor (1) due to radiation (12) reflected from the user's eye (2) and received by the event-based optical sensor (1), each event (31) corresponding to the detection of a temporal change in the received radiation in one or more pixels of the optical sensor (1), - Generating an inference frame (61) based on at least a portion of the stream (3) of events (31), the inference frame (61) comprising a plurality of frame pixels arranged in an array, - Operating the machine learning module (6) to obtain output data using the inference frame (61) without using the stream (3) of events (31) as an input to the machine learning module (6), - Extracting information related to the user's eye (2) from the output data, a line-of-sight tracking method comprising: A line-of-sight tracking method characterized by using a first artificial neural network (5) to generate the inference frame (61).

15. When executed by a computer or a microcontroller, the computer or the microcontroller - Receiving a signal stream (3) of events (31) created by the event-based optical sensor (1) due to radiation (12) reflected from the user's eye (2) and received by the event-based optical sensor (1), each event (31) corresponding to the detection of a temporal change in the received radiation in one or more pixels of the optical sensor (1), - Generating an inference frame (61) based on at least a portion of the stream (3) of the event (31), the inference frame (61) comprising a plurality of frame pixels arranged in an array; - Operating the machine learning module (6) to obtain output data by using the inference frame (61) without using the stream (3) of the event (31) as an input to the machine learning module (6); - A computer-readable medium comprising instructions for causing execution of a step of extracting information related to the eye (2) of the user from the output data; A computer-readable medium, characterized in that a first artificial neural network (5) is used to generate the inference frame (61).

Citation Information

Patent Citations

  • System and method for hybrid eye tracker

    EP3598274A1

  • Event camera for eye tracking

    US10466779B1

  • Method and device for eye tracking using event camera data

    WO2019067731A1

  • Event camera-based gaze tracking using neural networks

    WO2019147677A1

  • Systems and methods for generating and transmitting image sequences based on sampled color information

    WO2020068140A1