Event camera based gaze tracking using neural networks

By using event cameras and neural network technology, the problems of high bandwidth requirements and high power consumption in existing gaze tracking systems are solved, and more efficient and faster gaze tracking effects are achieved.

CN119964225APending Publication Date: 2025-05-09APPLE INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510045717.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-01-24
Filing Date
2019-01-23
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing gaze tracking systems rely on shutter-based camera images, resulting in a high bandwidth communication link requiring increased heat and power consumption of the device.

Method used

An event camera is used to receive a pixel event stream and generate gaze characteristics through a neural network to perform gaze tracking based on the event camera. The system derives images by accumulating pixel events of multiple event camera pixels and uses these images as input to the neural network to determine the gaze characteristics.

Benefits of technology

By reducing the amount of data transmission and the use of computing resources, faster gaze tracking is achieved, reducing the heat and power consumption of the equipment, while improving the efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964225A_ABST
    Figure CN119964225A_ABST
Patent Text Reader

Abstract

The invention relates to event camera-based gaze tracking using neural networks. One embodiment relates to an apparatus that receives a stream of pixel events output by an event camera. The device derives an input image by accumulating pixel events for a plurality of event camera pixels. The device generates a gaze characteristic using the derived input image as an input to a neural network trained to determine the gaze characteristic. The neural network is configured in multiple stages. A first stage of the neural network is configured to determine an initial gaze characteristic, such as an initial pupil center, using a reduced resolution input. A second stage of the neural network is configured to determine adjustments to the initial gaze characteristic using input in a location set, such as using only small input images centered on the initial pupil center. Thus, the determination at each stage is performed efficiently using a relatively compact neural network configuration. The device tracks a gaze of the eye based on the gaze characteristic.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with the application date of January 23, 2019, named “Event-based Gaze Tracking Using Neural Networks” and application number 201980010044.3. Technical Field

[0002] The present disclosure relates generally to gaze tracking, and in particular to systems, methods, and devices for gaze tracking using event camera data. Background Art

[0003] Existing gaze tracking systems determine the user's gaze direction based on a shutter-based camera image of the user's eye. Existing gaze tracking systems often include a camera that transmits an image of the user's eye to a processor that performs gaze tracking. Transmitting images at a frame rate sufficient to enable gaze tracking requires a communication link with considerable bandwidth, and using such a communication link increases the heat generated and power consumption of the device. Summary of the invention

[0004] Various embodiments disclosed herein include devices, systems, and methods for using a neural network to perform gaze tracking based on an event camera. An exemplary embodiment involves performing operations at a device having one or more processors and a computer-readable storage medium. The device receives a stream of pixel events output by an event camera. The event camera has a pixel sensor positioned to receive light from the surface of an eye. In response to a change in light intensity exceeding a comparator threshold of light at a corresponding event camera pixel detected by the corresponding pixel sensor, each corresponding pixel event is generated. The device derives an image from the pixel event stream by accumulating pixel events of multiple event camera pixels. The device uses the derived image as an input to a neural network to generate a gaze characteristic. A training data set of training images that identify the gaze characteristic is used to train the neural network to determine the gaze characteristic. The device tracks the gaze of the eye based on the gaze characteristic generated using the neural network.

[0005] Various implementations configure the neural network to efficiently determine the gaze characteristic. For example, efficiency is achieved by using a multi-stage neural network. The first stage of the neural network is configured to determine an initial gaze characteristic (e.g., an initial pupil center) using a reduced resolution input. The second stage of the neural network is configured to determine an adjustment to the initial gaze characteristic using a position-focused input (e.g., using only a small input image centered around the initial pupil center). Thus, the determination at each stage is efficiently computed using a relatively compact neural network configuration.

[0006] In some implementations, a recurrent neural network (such as a long / short term memory (LSTM) or a gated recurrent unit (GRU) based network) is used to determine the gaze characteristics. Using a recurrent neural network can provide efficiency. The neural network maintains internal states that are used to refine the gaze characteristics over time and produce smoother output results. During transient blurred scenes (such as occlusion due to eyelashes), the internal states are used to ensure temporal consistency of the gaze characteristics.

[0007] According to some specific implementations, a device includes one or more processors, non-volatile memory, and one or more programs; the one or more programs are stored in the non-volatile memory and are configured to be executed by one or more processors, and the one or more programs include instructions for performing or causing the execution of any of the methods described herein. According to some specific implementations, a non-volatile computer-readable storage medium stores instructions that, when executed by one or more processors of the device, cause the device to perform or cause the execution of any of the methods described herein. According to some specific implementations, a device includes: one or more processors, non-volatile memory, and means for performing or causing the execution of any of the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] So that the present disclosure may be understood by those of ordinary skill in the art, a more detailed description may be obtained with reference to aspects of some exemplary implementations, some of which are illustrated in the accompanying drawings.

[0009] Figure 1 is a block diagram of an exemplary operating environment according to some implementations.

[0010] Figure 2 is a block diagram of an example controller according to some implementations.

[0011] Figure 3 is a block diagram of an exemplary head mounted device (HMD) according to some implementations.

[0012] Figure 4 is a block diagram of an exemplary head mounted device (HMD) according to some implementations.

[0013] Figure 5 A block diagram of an event camera according to some implementations is shown.

[0014] Figure 6 is a flowchart representation of a method for event-camera based gaze tracking according to some specific implementations.

[0015] Figure 7 A functional block diagram is shown illustrating an event camera-based gaze tracking process according to some specific implementations.

[0016] Figure 8 A functional block diagram is shown illustrating a system for gaze tracking using a convolutional neural network according to some specific implementations.

[0017] Fig. 9 A functional block diagram is shown illustrating a system for gaze tracking using a convolutional neural network according to some specific implementations.

[0018] Fig.10 shows a functional block diagram showing Fig. 9 The convolutional layer of a convolutional neural network.

[0019] Fig.11 A functional block diagram is shown illustrating a system for gaze tracking using an initialization network and a refinement network according to some implementations.

[0020] As is common practice, the various features shown in the drawings may not be drawn to scale. Therefore, the sizes of the various features may be arbitrarily expanded or reduced for clarity. In addition, some drawings may not depict all components of a given system, method, or device. Finally, throughout the specification and drawings, similar reference numerals may be used to represent similar features. DETAILED DESCRIPTION

[0021] Many details are described in order to provide a thorough understanding of the exemplary implementations shown in the accompanying drawings. However, the accompanying drawings only illustrate some example aspects of the present disclosure and should not be considered limiting. One of ordinary skill in the art will appreciate that other effective aspects and / or variations do not include all of the specific details described herein. In addition, well-known systems, methods, components, devices, and circuits are not described in detail in order to avoid obscuring more relevant aspects of the exemplary implementations described herein.

[0022] In various embodiments, gaze tracking is used to enable user interaction, provide foveal rendering, or mitigate geometric distortion. A gaze tracking system includes a camera and a processor that performs gaze tracking on data received from a camera about light from a light source reflected from a user's eyes. In various embodiments, the camera includes an event camera having multiple light sensors at multiple corresponding positions, and the event camera generates an event message indicating a specific position of the specific light sensor in response to a specific light sensor detecting a change in light intensity. The event camera may include or be referred to as a dynamic vision sensor (DVS), a silicon retina, an event-based camera, or a frameless camera. Therefore, the event camera generates (and transmits) data about changes in light intensity, rather than a larger amount of data about the absolute intensity at each light sensor. In addition, because data is generated when the intensity changes, in various embodiments, the light source is configured to emit light with a modulated intensity.

[0023] In various implementations, asynchronous pixel event data from one or more event cameras is accumulated to generate one or more inputs to a neural network configured to determine one or more gaze characteristics (e.g., pupil center, pupil outline, glint location, gaze direction, etc.). The accumulated event data can be accumulated over time to generate one or more input images to the neural network. A first input image can be created by accumulating event data over time to generate an intensity reconstruction image that uses the event data to reconstruct the image intensity at each pixel location. A second input image can be created by accumulating event data over time to generate a timestamp image that encodes the age of camera events at each event camera pixel (e.g., the time since the most recent event). A third input image can be created by accumulating glint-specific event camera data over time to generate a glint image. These input images are used alone or in combination with each other and / or with other inputs to the neural network to generate gaze characteristics. In other implementations, event camera data is used as input to the neural network in other forms (e.g., individual events, events within a predetermined time window (e.g., 10 milliseconds)).

[0024] In various implementations, the neural network for determining the gaze characteristic is configured to do so efficiently. For example, efficiency is achieved by using a multi-stage neural network. The first stage of the neural network is configured to determine an initial gaze characteristic (e.g., an initial pupil center) using a reduced resolution input. For example, instead of using a 400×400 pixel input image, the resolution of the input image at the first stage can be reduced to 50×50 pixels. The second stage of the neural network is configured to use a position-focused input (e.g., using only a small input image centered on the initial pupil center) to determine adjustments to the initial gaze characteristic. For example, instead of using a 400×400 pixel input image, a selected portion of the input image at the same resolution (e.g., 80×80 pixels centered on the pupil center) can be used as an input at the second stage. Thus, the determination at each stage is performed using a relatively compact neural network configuration. Since the corresponding inputs (e.g., a 50×50 pixel image and an 80×80 pixel image) are less than the full resolution of the entire image of the data received from the event camera (e.g., a 400×400 pixel image), the corresponding neural network configuration is relatively small and efficient.

[0025] Figure 11 is a block diagram of an exemplary operating environment 100 according to some implementations. Although relevant features are shown, one of ordinary skill in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the exemplary implementations disclosed herein. To this end, as a non-limiting example, the operating environment 100 includes a controller 110 and a head mounted device (HMD) 120.

[0026] In some implementations, the controller 110 is configured to manage and coordinate the user's augmented reality / virtual reality (AR / VR) experience. In some implementations, the controller 110 includes a suitable combination of software, firmware, and / or hardware. Figure 2 The controller 110 is described in more detail. In some implementations, the controller 110 is a computing device that is located locally or remotely relative to the scene 105. In one example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside the scene 105. In some implementations, the controller 110 is communicatively coupled to the HMD 120 via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.).

[0027] In some implementations, the HMD 120 is configured to present an AR / VR experience to a user. In some implementations, the HMD 120 includes a suitable combination of software, firmware, and / or hardware. Figure 3 HMD 120 is described in further detail. In some implementations, the functionality of controller 110 is provided by and / or combined with HMD 120.

[0028] According to some implementations, the HMD 120 provides an augmented reality / virtual reality (AR / VR) experience to the user while the user is virtually and / or physically present within the scene 105. In some implementations, while presenting an augmented reality (AR) experience, the HMD 120 is configured to present AR content and enable optical see-through of the scene 105. In some implementations, while presenting a virtual reality (VR) experience, the HMD 120 is configured to present VR content and enable video see-through of the scene 105.

[0029] In some implementations, the user wears the HMD 120 on his head. Thus, the HMD 120 includes one or more AR / VR displays provided for displaying AR / VR content. For example, the HMD 120 surrounds the user's field of view. In some implementations, the HMD 120 is replaced with a handheld electronic device (e.g., a smartphone or tablet) configured to present AR / VR content to the user. In some implementations, the HMD 120 is replaced with an AR / VR room, enclosure, or room configured to present AR / VR content, in which the user does not wear or hold the HMD 120.

[0030] Figure 2 is a block diagram of an example of a controller 110 according to some implementations. While some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the implementations disclosed herein. To this end, as a non-limiting example, in some specific implementations, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., a universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), infrared (IR), Bluetooth, ZIGBEE and / or similar type interfaces), one or more programming (e.g., I / O) interfaces 210, a memory 220, and one or more communication buses 204 for interconnecting these components and various other components.

[0031] In some implementations, the one or more communication buses 204 include circuits for interconnecting system components and controlling communications between system components. In some implementations, the one or more I / O devices 206 include at least one of a keyboard, a mouse, a touch pad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and the like.

[0032] The memory 220 includes a high-speed random access memory, such as a dynamic random access memory (DRAM), a static random access memory (SRAM), a double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some specific implementations, the memory 220 includes a non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 220 optionally includes one or more storage devices located away from the one or more processing units 202. The memory 220 includes a non-transitory computer-readable storage medium. In some specific implementations, the memory 220 or the non-transitory computer-readable storage medium of the memory 220 stores the following programs, modules, and data structures or their subsets, including an optional operating system 230 and an augmented reality / virtual reality (AR / VR) experience module 240.

[0033] The operating system 230 includes processes for handling various basic system services and for performing hardware-related tasks. In some specific implementations, the AR / VR experience module 240 is configured to manage and coordinate one or more AR / VR experiences for one or more users (e.g., a single AR / VR experience for one or more users, or multiple AR / VR experiences for corresponding groups of one or more users). To this end, in various specific implementations, the AR / VR experience module 240 includes a data acquisition unit 242, a tracking unit 244, a coordination unit 246, and a rendering unit 248.

[0034] In some implementations, the data acquisition unit 242 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the HMD 120. To this end, in various implementations, the data acquisition unit 242 includes instructions and / or logic for instructions as well as heuristics and metadata for the heuristics.

[0035] In some implementations, the tracking unit 244 is configured to map the scene 105 and to track at least the position / location of the HMD 120 relative to the scene 105. To this end, in various implementations, the tracking unit 244 includes instructions and / or logic for instructions and heuristics and metadata for the heuristics.

[0036] In some implementations, the coordination unit 246 is configured to manage and coordinate the AR / VR experience presented to the user by the HMD 120. To this end, in various implementations, the coordination unit 246 includes instructions and / or logic components for instructions and heuristics and metadata for the heuristics.

[0037] In some implementations, the rendering unit 248 is configured to render content for display on the HMD 120. To this end, in various implementations, the rendering unit 248 includes instructions and / or logic for instructions as well as heuristics and metadata for the heuristics.

[0038] Although the data acquisition unit 242, tracking unit 244, coordination unit 246, and rendering unit 248 are illustrated as residing on a single device (e.g., controller 110), it should be understood that in other implementations, any combination of the data acquisition unit 242, tracking unit 244, coordination unit 246, and rendering unit 248 may be located in separate computing devices.

[0039] also, Figure 2 It is more of a functional description of various features present in a particular implementation, as opposed to a schematic diagram of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 2 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software and / or firmware selected for a particular embodiment.

[0040] Figure 3 1 is a block diagram of an example of a head mounted device (HMD) 120 according to some implementations. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the implementations disclosed herein. To this end, as a non-limiting example, in some specific implementations, the HMD 120 includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE, SPI, I2C, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more AR / VR displays 312, one or more internal-facing and / or external-facing image sensor systems 314, memory 320, and one or more communication buses 304 for interconnecting these components and various other components.

[0041] In some implementations, the one or more communication buses 304 include circuits for interconnecting and controlling communications between system components. In some implementations, the one or more I / O devices and sensors 306 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, and / or one or more depth sensors (e.g., structured light, time of flight, etc.), etc.

[0042] In some implementations, the one or more AR / VR displays 312 are configured to present an AR / VR experience to the user. In some implementations, one or more AR / VR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light emitting field effect transistor (OLET), organic light emitting diode (OLED), surface conduction electron emitter display (SED), field emission display (FED), quantum dot light emitting diode (QD-LED), microelectromechanical system (MEMS) and / or similar display types. In some implementations, one or more AR / VR displays 312 correspond to diffraction, reflection, polarization, holographic and other waveguide displays. For example, HMD 120 includes a single AR / VR display. In another example, HMD 120 includes an AR / VR display for each eye of the user. In some implementations, one or more AR / VR displays 312 are capable of presenting AR and VR content. In some implementations, one or more AR / VR displays 312 are capable of presenting AR or VR content.

[0043] In some implementations, the one or more image sensor systems 314 are configured to acquire image data corresponding to at least a portion of a user's face including the user's eyes. For example, the one or more image sensor systems 314 include one or more RGB cameras (e.g., with complementary metal oxide semiconductor (CMOS) image sensors or charge coupled device (CCD) image sensors), monochrome cameras, IR cameras, event-based cameras, etc. In various implementations, the one or more image sensor systems 314 also include an illumination source that emits light on the portion of the user's face, such as a flash or strobe light source.

[0044] The memory 320 includes a high-speed random access memory, such as DRAM, SRAM, DDR RAM or other random access solid-state memory devices. In some specific implementations, the memory 320 includes a non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices or other non-volatile solid-state storage devices. The memory 320 optionally includes one or more storage devices located away from the one or more processing units 302. The memory 320 includes a non-transitory computer-readable storage medium. In some specific implementations, the memory 320 or the non-transitory computer-readable storage medium of the memory 320 stores the following programs, modules and data structures or their subsets, including an optional operating system 330, an AR / VR rendering module 340, and a user data storage device 360.

[0045] The operating system 330 includes processes for handling various basic system services and for performing hardware-related tasks. In some implementations, the AR / VR rendering module 340 is configured to present AR / VR content to the user via the one or more AR / VR displays 312. To this end, in various implementations, the AR / VR rendering module 340 includes a data acquisition unit 342, an AR / VR rendering unit 344, a gaze tracking unit 346, and a data transmission unit 348.

[0046] In some implementations, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110. To this end, in various implementations, the data acquisition unit 342 includes instructions and / or logic components for instructions and heuristics and metadata for the heuristics.

[0047] In some implementations, the AR / VR rendering unit 344 is configured to present AR / VR content via the one or more AR / VR displays 312. To this end, in various implementations, the AR / VR rendering unit 344 includes instructions and / or logic for instructions and heuristics and metadata for the heuristics.

[0048] In some implementations, the gaze tracking unit 346 is configured to determine the user's gaze tracking characteristics based on the event messages received from the event camera. To this end, in various implementations, the gaze tracking unit 346 includes instructions and / or logic components for instructions, a configured neural network, and heuristics and metadata for the heuristics.

[0049] In some implementations, the data transfer unit 348 is configured to transfer data (e.g., presentation data, location data, etc.) to at least the controller 110. To this end, in various implementations, the data transfer unit 348 includes instructions and / or logic for instructions as well as heuristics and metadata for the heuristics.

[0050] Although the data acquisition unit 342, the AR / VR rendering unit 344, the gaze tracking unit 346, and the data transmission unit 348 are shown as residing on a single device (e.g., the HMD 120), it should be understood that in other specific implementations, any combination of the data acquisition unit 342, the AR / VR rendering unit 344, the gaze tracking unit 346, and the data transmission unit 348 may be located in separate computing devices.

[0051] also, Figure 3 It is more of a functional description of various features present in a particular implementation, as opposed to a schematic diagram of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 3 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software and / or firmware selected for a particular embodiment.

[0052] Figure 4 A block diagram of a head mounted device 400 according to some implementations is shown. The head mounted device 400 includes a housing 401 (or enclosure) that houses various components of the head mounted device 400. The housing 401 includes (or is coupled to) an eye pad 405 disposed at a proximal (relative to the user 10) end of the housing 401. In various implementations, the eye pad 405 is a plastic or rubber piece that comfortably and snugly holds the head mounted device 400 in place on the face of the user 10 (e.g., around the eyes of the user 10).

[0053] Housing 401 houses display 410, which displays images, emits light toward or onto the eyes of user 10. In various implementations, display 410 emits light through an eyepiece (not shown), which refracts the light emitted by display 410, so that the display appears to user 10 to be at a virtual distance farther than the actual distance from the eye to display 410. In order for the user to focus on display 410, in various implementations, the virtual distance is at least greater than the minimum focal length of the eye (e.g., 7 cm). In addition, in order to provide a better user experience, in various implementations, the virtual distance is greater than 1 meter.

[0054] Although Figure 4 A head mounted device 400 is shown including a display 410 and an eye pad 405 , but in various implementations, the head mounted device 400 does not include the display 410 or includes an optical see-through display without the eye pad 405 .

[0055] The housing 401 also houses a gaze tracking system that includes one or more light sources 422, a camera 424, and a controller 480. The one or more light sources 422 emit light onto the eyes of the user 10, which is reflected into a light pattern (e.g., a flashing circle) that can be detected by the camera 424. Based on the light pattern, the controller 480 can determine the gaze tracking characteristics of the user 10. For example, the controller 480 can determine the gaze direction and / or blink state (eyes open or closed) of the user 10. As another example, the controller 480 can determine the pupil center, pupil size, or focus point. Therefore, in various specific implementations, light is emitted by the one or more light sources 422, reflected from the eyes of the user 10, and detected by the camera 424. In various specific implementations, the light from the eyes of the user 10 is reflected from the hot mirror or passes through the eyepiece before reaching the camera 424.

[0056] The display 410 emits light within a first wavelength range, and the one or more light sources 422 emit light within a second wavelength range. Similarly, the camera 424 detects light within the second wavelength range. In various implementations, the first wavelength range is a visible wavelength range (e.g., a wavelength range of approximately 400 nm-700 nm within the visible spectrum), and the second wavelength range is a near infrared wavelength range (e.g., a wavelength range of approximately 700 nm-1400 nm within the near infrared spectrum).

[0057] In various embodiments, gaze tracking (or specifically, the determined gaze direction) is used to enable user interaction (e.g., user 10 selecting an option on display 410 by looking at it), provide foveated rendering (e.g., presenting a higher resolution in the area of ​​display 410 that user 10 is looking at and a lower resolution elsewhere on display 410), or reduce geometric distortion (e.g., in 3D rendering of objects on display 410).

[0058] In various implementations, the one or more light sources 422 emit light toward the user's eyes, which is reflected in the form of multiple flashes.

[0059] In various implementations, one or more light sources 422 emit light with a modulated intensity toward the user's eye. Thus, at a first time, a first light source of the plurality of light sources projects light onto the user's eye at a first intensity, and at a second time, the first light source of the plurality of light sources projects light onto the user's eye at a second intensity different from the first intensity (which may be zero, e.g., off).

[0060] Multiple flickers can be generated by light emitted toward the user's eye (and reflected by the cornea) with modulated intensity. For example, at a first time, the first flicker and the fifth flicker of the multiple flickers are reflected by the eye at a first intensity. At a second time later than the first time, the intensity of the first flicker and the fifth flicker is modulated to a second intensity (e.g., zero). In addition, at the second time, the second flicker and the sixth flicker of the multiple flickers are reflected by the user's eye at a first intensity. At a third time later than the second time, the third flicker and the seventh flicker of the multiple flickers are reflected by the user's eye at a first intensity. At a fourth time later than the third time, the fourth flicker and the eighth flicker of the multiple flickers are reflected by the user's eye at a first intensity. At a fifth time later than the fourth time, the intensity of the first flicker and the fifth flicker is modulated back to the first intensity.

[0061] Thus, in various implementations, each of the multiple flashes flashes on and off at a modulation frequency (e.g., 600 Hz). However, the phase of the second flash is offset relative to the phase of the first flash, the phase of the third flash is offset relative to the phase of the second flash, etc. The flashes may be configured in this manner to appear to rotate around the cornea.

[0062] Thus, in various implementations, the intensities of different ones of the plurality of light sources are modulated in different ways. Thus, when analyzing the glint reflected by the eye and detected by the camera 424, the identity of the glint and the corresponding light source (e.g., which light source produced the detected glint) can be determined.

[0063] In various implementations, one or more light sources 422 are differentially modulated in various ways. In various implementations, a first light source in the plurality of light sources is modulated at a first frequency with a first phase offset (e.g., a first flicker), and a second light source in the plurality of light sources is modulated at the first frequency with a second phase offset (e.g., a second flicker).

[0064] In various implementations, the one or more light sources 422 modulate the intensity of the emitted light at different modulation frequencies. For example, in various implementations, a first light source among the plurality of light sources is modulated at a first frequency (e.g., 600 Hz), and a second light source among the plurality of light sources is modulated at a second frequency (e.g., 500 Hz).

[0065] In various implementations, the one or more light sources 422 modulate the intensity of the emitted light according to different orthogonal codes, such as those that can be used in CDMA (code division multiple access) communications. For example, rows or columns of a Walsh matrix can be used as orthogonal codes. Thus, in various implementations, a first light source in the plurality of light sources is modulated according to a first orthogonal code, and a second light source in the plurality of light sources is modulated according to a second orthogonal code.

[0066] In various implementations, the one or more light sources 422 modulate the intensity of the emitted light between a high intensity value and a low intensity value. Thus, at various times, the intensity of the light emitted by the light source is either a high intensity value or a low intensity value. In various implementations, the low intensity value is zero. Thus, in various implementations, the one or more light sources 422 modulate the intensity of the emitted light between an on state (at a high intensity value) and an off state (at a low intensity value). In various implementations, the number of light sources in the plurality of light sources that are in the on state is constant.

[0067] In various implementations, the one or more light sources 422 modulate the intensity of emitted light within an intensity range (e.g., between 10% of maximum intensity and 40% of maximum intensity). Thus, at various times, the intensity of the light source is a low intensity value, a high intensity value, or a value in between. In various implementations, the one or more light sources 422 are differentially modulated such that a first light source of the plurality of light sources is modulated within a first intensity range, and a second light source of the plurality of light sources is modulated within a second intensity range that is different from the first intensity range.

[0068] In various implementations, the one or more light sources 422 modulate the intensity of the emitted light based on the gaze direction. For example, if the user is looking in the direction where a particular light source is reflected by the pupil, the one or more light sources 422 change the intensity of the emitted light based on this knowledge. In various implementations, the one or more light sources 422 reduce the intensity of the emitted light to reduce the amount of near infrared light entering the pupil as a safety precaution.

[0069] In various implementations, the one or more light sources 422 modulate the intensity of the emitted light based on the user's biometrics. For example, if the user blinks more than normal, has an elevated heart rate, or is registered as a child, the one or more light sources 422 reduce the intensity of the emitted light (or the total intensity of all the light emitted by the multiple light sources) to reduce the strain on the eyes. For another example, the one or more light sources 422 modulate the intensity of the emitted light based on the user's eye color, because the spectral reflectivity may be different for blue eyes compared to brown eyes.

[0070] In various implementations, the one or more light sources 422 modulate the intensity of the emitted light according to the presented user interface (e.g., the content displayed on the display 410). For example, if the display 410 is abnormally bright (e.g., a video of an explosion is being displayed), the one or more light sources 422 increase the intensity of the emitted light to compensate for potential interference from the display 410.

[0071] In various implementations, camera 424 is a frame / shutter based camera that generates images of the eyes of user 10 at a particular point in time or multiple points in time at a frame rate. Each image includes a matrix of pixel values ​​corresponding to pixels of the image, the pixels corresponding to positions of the camera's light sensor matrix.

[0072] In various implementations, camera 424 is an event camera comprising multiple light sensors (e.g., a light sensor matrix) at multiple corresponding positions, which generates an event message indicating a specific position of a particular light sensor in response to the particular light sensor detecting a change in light intensity.

[0073] Figure 5 A functional block diagram of an event camera 500 according to some implementations is shown. The event camera 500 includes a plurality of light sensors 515 respectively coupled to a message generator 532. In various implementations, the plurality of light sensors 515 are arranged in a matrix 510 of rows and columns, and thus, each of the plurality of light sensors 515 is associated with a row value and a column value.

[0074] Each of the plurality of light sensors 515 includes Figure 55 . The light sensor 520 is shown in detail in FIG. The light sensor 520 includes a photodiode 521 connected in series with a resistor 523 between a source voltage and a ground voltage. The voltage across the photodiode 521 is proportional to the intensity of light incident on the light sensor 520. The light sensor 520 includes a first capacitor 525 connected in parallel with the photodiode 521. Therefore, the voltage across the first capacitor 525 is the same as the voltage across the photodiode 521 (e.g., proportional to the intensity of light detected by the light sensor 520).

[0075] The light sensor 520 includes a switch 529 coupled between a first capacitor 525 and a second capacitor 527. The second capacitor 527 is coupled between the switch and a ground voltage. Thus, when the switch 529 is closed, the voltage across the second capacitor 527 is the same as the voltage across the first capacitor 525 (e.g., proportional to the intensity of light detected by the light sensor 520). When the switch 529 is open, the voltage across the second capacitor 527 is fixed at the voltage across the second capacitor 527 when the switch 529 was last closed.

[0076] The voltage on the first capacitor 525 and the voltage on the second capacitor 527 are fed to a comparator 531. When the difference 552 between the voltage on the first capacitor 525 and the voltage on the second capacitor 527 is less than a threshold amount, the comparator 531 outputs a "0" voltage. When the voltage on the first capacitor 525 is higher than the voltage on the second capacitor 527 by at least a threshold amount, the comparator 531 outputs a "1" voltage. When the voltage on the first capacitor 525 is lower than the voltage on the second capacitor 527 by at least a threshold amount, the comparator 531 outputs a "-1" voltage.

[0077] When the comparator 531 outputs a “1” voltage or a “−1” voltage, the switch 529 is closed, and the message generator 532 receives the digital signal and generates a pixel event message.

[0078] For example, at a first time, the intensity of light incident on the light sensor 520 is a first light value. Therefore, the voltage on the photodiode 521 is a first voltage value. Similarly, the voltage on the first capacitor 525 is a first voltage value. For this example, the voltage on the second capacitor 527 is also a first voltage value. Therefore, the comparator 531 outputs a "0" voltage, the switch 529 remains closed, and the message generator 532 does not perform any operation.

[0079] At the second time, the intensity of the light incident on the light sensor 520 increases to a second light value. Therefore, the voltage on the photodiode 521 is a second voltage value (higher than the first voltage value). Similarly, the voltage on the first capacitor 525 is a second voltage value. Because the switch 529 is disconnected, the voltage on the second capacitor 527 is still the first voltage value. Assuming that the second voltage value is at least the threshold value higher than the first voltage value, the comparator 531 outputs a "1" voltage, closes the switch 529, and the message generator 532 generates an event message based on the received digital signal.

[0080] When the switch 529 is closed due to the "1" voltage from the comparator 531, the voltage on the second capacitor 527 changes from the first voltage value to the second voltage value. Therefore, the comparator 531 outputs a "0" voltage, and the switch 529 is turned off.

[0081] At the third time, the intensity of the light incident on the light sensor 520 increases (again) to the third light value. Therefore, the voltage on the photodiode 521 is the third voltage value (higher than the second voltage value). Similarly, the voltage on the first capacitor 525 is the third voltage value. Because the switch 529 is disconnected, the voltage on the second capacitor 527 is still the second voltage value. Assuming that the third voltage value is at least the threshold value higher than the second voltage value, the comparator 531 outputs a "1" voltage, closes the switch 529, and the message generator 532 generates an event message based on the received digital signal.

[0082] When the switch 529 is closed due to the "1" voltage from the comparator 531, the voltage on the second capacitor 527 changes from the second voltage value to the third voltage value. Therefore, the comparator 531 outputs a "0" voltage, and the switch 529 is turned off.

[0083] At the fourth time, the intensity of the light incident on the light sensor 520 decreases back to the second light value. Therefore, the voltage on the photodiode 521 is the second voltage value (less than the third voltage value). Similarly, the voltage on the first capacitor 525 is the second voltage value. Because the switch 529 is disconnected, the voltage on the second capacitor 527 is still the third voltage value. Therefore, the comparator 531 outputs a "-1" voltage, closes the switch 529, and the message generator 532 generates an event message based on the received digital signal.

[0084] When the switch 529 is closed due to the "-1" voltage from the comparator 531, the voltage on the second capacitor 527 changes from the third voltage value to the second voltage value. Therefore, the comparator 531 outputs a "0" voltage, and the switch 529 is turned off.

[0085] The message generator 532 receives a digital signal from each of the plurality of light sensors 510 at different times, the digital signal indicating an increase in light intensity (a "1" voltage) or a decrease in light intensity (a "-1" voltage). In response to receiving a digital signal from a particular light sensor in the plurality of light sensors 510, the message generator 532 generates a pixel event message.

[0086] In various implementations, each pixel event message indicates a specific location of a specific light sensor in a location field. In various implementations, the event message indicates a specific location in pixel coordinates, such as a row value (e.g., in a row field) and a column value (e.g., in a column field). In various implementations, the event message further indicates the polarity of the light intensity change in a polarity field. For example, the event message may include a "1" in the polarity field to indicate an increase in light intensity, and may include a "0" in the polarity field to indicate a decrease in light intensity. In various implementations, the event message further indicates the time at which the light intensity change was detected in a time field (e.g., the time at which the digital signal was received). In various implementations, the event message indicates a value representing the intensity of the detected light in an absolute intensity field (not shown), as an alternative to or in addition to the polarity.

[0087] Figure 6 60 is a flowchart representation of a method 60 for event-based camera gaze tracking according to some specific implementations. In some specific implementations, the method 600 is performed by a device (e.g., Figure 1 and Figure 2 The method 600 may be performed on a device having a screen for displaying 2D images and / or a screen for viewing stereoscopic images (e.g., Figure 1 and Figure 3 The method 600 is performed on a HMD 120 (e.g., a virtual reality (VR) display (e.g., a head mounted display (HMD)) or an augmented reality (AR) display. In some implementations, the method 600 is performed by a processing logic component (including hardware, firmware, software, or a combination thereof). In some implementations, the method 600 is performed by a processor executing code stored in a non-transitory computer readable medium (e.g., a memory).

[0088] At block 610, method 600 receives a pixel event stream output by an event camera. The pixel event data may be in various forms. The pixel event stream may be received as a series of messages identifying pixel events at one or more pixels of the event camera. In various implementations, pixel event messages are received, each pixel event message including a position field, a polarity field, a time field, and / or an absolute intensity field for a particular position of a particular light sensor.

[0089] At block 620, method 600 derives one or more images from the pixel event stream. The one or more images are derived to provide a combined input to the neural network. In alternative implementations, the pixel event data is fed directly into the neural network as individual input items (e.g., one input per event), input batches (e.g., 10 events per input), or otherwise conditioned into a suitable form for input into the neural network.

[0090] In a specific implementation of deriving an input image, the information in the input image is represented in an event camera grid (e.g., Figure 5 The event camera data at the corresponding position in the grid 510 of the input image. Therefore, the value of the pixel in the upper right corner of the input image corresponds to the event camera data of the upper right event camera sensor in the pixel grid of the event camera. The pixel event data of multiple events are accumulated and used to generate an image compiling the event data of multiple pixels. In one example, an input image representing pixel events occurring within a specific time window (e.g., within the last 10ms) is created. In another example, an input image representing pixel events occurring until a specific point in time is created, for example, identifying the most recent pixel event occurring at each pixel. In another example, pixel events are accumulated over time to track or estimate the absolute intensity value of each pixel.

[0091] The location field associated with a pixel event included in the pixel event stream can be used to identify the location of the corresponding event camera pixel. For example, the pixel event data in the pixel event stream can identify a pixel event that occurs in the upper right pixel, and this information can be used to assign a value in the corresponding upper right pixel of the input image. Figure 7 Examples of deriving input images such as intensity reconstructed images, time stamp images, and scintillation images are described.

[0092] At box 630, method 600 uses a neural network to generate a gaze characteristic. One or more input images derived from the pixel event stream of the event camera are used as inputs to the neural network. A training data set of training images that identify gaze characteristics is used to train the neural network to determine the gaze characteristics. For example, shutter-based images of the eyes of multiple subjects (e.g., 25 subjects, 50 subjects, 100 subjects, 1,000 subjects, 10,000 subjects, etc.) or images derived from the event camera can be used to create a training set. The training data set can include multiple images of each subject's eyes (e.g., 25 images, 50 images, 100 images, 1,000 images, 10,000 images, etc.). The training data can include ground truth gaze characteristic recognition, for example, identifying position or direction information of the pupil center position, pupil contour shape, pupil dilation, blinking position, gaze direction, etc. For pupil contour shape, a neural network can be trained with data indicating a set of points around the pupil perimeter (e.g., five points sampled in a repeatable manner around the pupil and fitted to an ellipse). The training data can be additionally or alternatively labeled with sentiment characteristics (e.g., "interested" as indicated by a relatively large pupil size, and "not interested" as indicated by a relatively small pupil size). Ground truth data can be manually identified, semi-automatically determined, or automatically determined. The neural network can be configured to use event camera data, shutter-based camera data, or a combination of both types of data. The neural network can be configured to be independent of whether the data comes from a shutter-based camera or an event camera.

[0093] At block 640, method 600 tracks fixation based on fixation characteristics. In some implementations, fixation is tracked based on pupil center position, glint position, or a combination of these features. In some implementations, fixation is tracked by tracking pupil center and glint position to determine and update a geometric model of the eye that is then used to reconstruct the fixation direction. In some implementations, fixation is tracked by comparing the current fixation characteristic with the previous fixation characteristic. For example, an algorithm may be used to track fixation by comparing the pupil center position as it changes over time.

[0094] In some implementations, gaze is tracked based on additional information. For example, when a user selects a UI item, the correspondence between the selection of the UI item displayed on the screen of the HMD and the position of the pupil center can be determined. This assumes that the user is looking at it when they select the UI item. Based on the position of the UI element on the display, the position of the display relative to the user, and the current pupil position, the gaze direction associated with the direction from the eye to the UI element can be determined. Such information can be used to adjust or calibrate gaze tracking performed based on event camera data.

[0095] In some embodiments, tracking the gaze of an eye involves updating a gaze characteristic in real time upon receiving subsequent pixel events in an event stream. The pixel events are used to derive additional images, and the additional images are used as inputs to a neural network to generate updated gaze characteristics. The gaze characteristics can be used for a variety of purposes. In one example, the determined or updated gaze characteristics are used to identify items displayed on a display, for example, to identify which button, image, text, or other user interface item a user is looking at. In another example, the determined or updated gaze characteristics are used to display movement of a graphical indicator (e.g., a cursor or other user-controlled icon) on a display. In another example, the determined or updated gaze characteristics are used to select an item displayed on a display (e.g., via a cursor selection command). For example, a specific gaze movement pattern can be identified and interpreted as a specific command.

[0096] Event camera-based gaze tracking techniques (such as Figure 6 The method 600 shown provides many advantages over techniques that rely solely on shutter-based camera data. Event cameras may capture data at very high sampling rates and therefore allow input images to be created at a faster rate than using shutter-based cameras. The input images (e.g., intensity reconstructed images) that have been created can simulate data from extremely fast shutter-based cameras without the high energy and data requirements of such cameras. Since the event camera does not collect / send an entire frame for each event, it produces relatively sparse data. However, the sparse data is accumulated over time to provide a dense input image that is used as input in gaze characteristic determination. The result is faster gaze tracking that is achieved using less data and computing resources.

[0097] Figure 7 A functional block diagram is shown illustrating an event camera based gaze tracking process 700 according to some implementations. The gaze tracking process 700 outputs a user's gaze direction based on event messages received from an event camera 710.

[0098] The event camera 710 includes a plurality of light sensors at a plurality of corresponding positions. In response to a particular light sensor detecting a change in light intensity, the event camera 710 generates an event message indicating a particular position of the particular light sensor. Figure 6 In various embodiments, the specific location is indicated by pixel coordinates. In various embodiments, the event message also indicates the polarity of the light intensity change. In various embodiments, the event message also indicates the time when the light intensity change was detected. In various embodiments, the event message also indicates a value representing the intensity of the detected light.

[0099] The event message from the event camera 710 is received by the separator 720. The separator 720 divides the event message into a target frequency event message (associated with a frequency band centered on the modulation frequency of one or more light sources) and a deviation from the target frequency event message (associated with other frequencies). The deviation from the target frequency event message is a pupil event 730 and is fed to the intensity reconstruction image generator 750 and the timestamp image generator 760. The target frequency event message is a flicker event 740 and is fed to the flicker image generator 770. In various specific implementations, the separator 720 determines that the event message is a target frequency event message (or a deviation from the target frequency event message) based on the timestamp of the time when the light intensity change is detected in the time field. For example, in various specific implementations, if the event message is one of a set of multiple event messages within a set range that indicates a specific position within a set amount of time, the separator 720 determines that the event message is a target frequency event message. Otherwise, the separator 720 determines that the event message is a deviation from the target frequency event message. In various specific implementations, the set range and / or the set time amount is proportional to the modulation frequency of the modulated light emitted toward the user's eyes. For example, in various specific implementations, if the time between consecutive events with similar or opposite polarity is within the set time range, the separator 720 determines that the event message is a target frequency event message.

[0100] The intensity reconstruction image generator 750 accumulates pupil events 730 for the pupils over time to reconstruct / estimate the absolute intensity value of each pupil. As additional pupil events 730 accumulate, the intensity reconstruction image generator 750 changes the corresponding values ​​in the reconstructed image. In this way, it generates and maintains an updated image of values ​​for all pixels of the image, even if only some pixels may have recently received events. The intensity reconstruction image generator 750 can adjust pixel values ​​based on additional information (e.g., information about nearby pixels) to improve the clarity, smoothness, or other aspects of the reconstructed image.

[0101] In various specific implementations, the intensity reconstructed image includes an image having a plurality of pixel values ​​at a corresponding plurality of pixels corresponding to corresponding positions of the light sensor. Upon receiving an event message indicating a specific position and a positive polarity (indicating that the light intensity has increased), a quantity (e.g., 1) is added to the pixel value at the pixel corresponding to the specific position. Similarly, upon receiving an event message indicating a specific position and a negative polarity (indicating that the light intensity has decreased), the quantity is subtracted from the pixel value at the pixel corresponding to the specific position. In various specific implementations, the intensity reconstructed image is filtered, such as blurred.

[0102] The timestamp image generator 760 encodes information about the timing of events. In one example, the timestamp image generator 760 creates an image with values ​​representing the length of time since the corresponding pixel event was received for each pixel. In such an image, pixels with more recent events may have higher intensity values ​​than pixels with less recent events. In one specific implementation, the timestamp image is a positive timestamp image with multiple pixel values ​​indicating when the corresponding light sensor triggered the most recent corresponding event with a positive polarity. In one specific implementation, the timestamp image is a negative timestamp image with multiple pixel values ​​indicating when the corresponding light sensor triggered the most recent corresponding event with a negative polarity.

[0103] The flicker image generator 770 determines the event associated with a particular flicker. In one example, the flicker image generator 770 identifies flickers based on associated frequencies. In some implementations, the flicker image generator 770 accumulates flicker events 740 over a period of time and generates a flicker image identifying the locations of all flicker events received within the period of time (e.g., within the last 10 ms, etc.). In some implementations, the flicker image generator 770 modulates the intensity of each pixel, depending on how well the flicker frequency or time since the most recent event at that pixel matches an expected value (derived from a target frequency), such as by evaluating a Gaussian function with a given standard deviation centered on the expected value.

[0104] The intensity reconstruction image generator 750, the timestamp image generator 760, and the scintillation image generator 770 provide images that are input to a neural network 780, which is configured to generate gaze characteristics. In various implementations, the neural network 780 includes a convolutional neural network, a recurrent neural network, and / or a long / short term memory (LSTM) network.

[0105] Figure 8 8 shows a functional block diagram illustrating a system 800 for gaze tracking using a convolutional neural network 830 according to some implementations. The system 800 uses an input image 810, such as an intensity reconstructed image, a timestamp image, and / or a scintillation image, as described above with respect to Figure 7Discussed. The input image 810 is resized to become a resized input image 820. Resizing the input image 810 may include downsampling to reduce the resolution of the image and / or cropping portions of the image. The resized input image 820 is input to a convolutional neural network 830. The convolutional neural network 830 includes one or more convolutional layers 840 and one or more fully connected layers 850, and generates an output 860. The convolutional layer 840 is configured to apply a convolution operation to its corresponding input and pass its result to the next layer. Each convolutional neuron in each layer of the convolutional layer 840 may be configured to process data of a receiving field, such as a portion of the resized input image 820. The fully connected layer 850 connects each neuron of one layer to each neuron of another layer.

[0106] Fig. 9 1 shows a functional block diagram illustrating a system 900 for gaze tracking using a convolutional neural network 920 according to some implementations. The system 900 uses an input image 910, such as an intensity reconstructed image, a timestamp image, and / or a scintillation image, as described above with respect to Figure 7 Discussed. The input image 910 is resized to become a resized input image 930. In this example, the 400×400 pixel input image 910 is resized by downsampling to become a resized input image 930 of 50×50 pixels. The resized input image 930 is input to the convolutional neural network 920. The convolutional neural network 920 includes three 5×5 convolutional layers 940a, 940b, 940c and two fully connected layers 950a, 950b. The convolutional neural network 920 is configured to generate an output 960 that identifies the x-coordinate and y-coordinate of the center of the pupil. For each glint, the output 960 also includes an x-coordinate, a y-coordinate, and a visible / invisible indication. The convolutional neural network 920 includes regression of pupil performance (e.g., mean square error (MSE)) with respect to the x-coordinate, y-coordinate, and radius and regression of glint performance (e.g., MSE) with respect to the x-coordinate and y-coordinate. The convolutional neural network 920 also includes a classification for flicker visibility / invisibility (e.g., flicker softmax cross entropy on visible / invisible logits). The data used by the convolutional neural network 920 (during training) can be augmented by random translation, rotation, and / or scaling.

[0107] Fig.10 shows a functional block diagram showing Fig. 9 The convolutional layers 940a-c of the convolutional neural network include a 2×2 average pool 1010, a convolution W×W×N 1020, a batch normalization layer 1030, and a rectified linear unit (ReLu) 1040.

[0108] Fig.11 A functional block diagram is shown, showing a system 1100 using an initialization network 1115 and a refinement network 1145 for gaze tracking according to some specific implementations. The use of two subnetworks or two stages can enhance the efficiency of the entire neural network 1110. The initialization network 1115 processes a low-resolution image of the input image 1105 and therefore does not need to have as many convolutions as otherwise. The pupil position output from the initialization network 1115 is used to crop the original input image 1105 input to the refinement network 1145. The refinement network 1145 therefore also does not need to use the entire input image 1105 and can accordingly include fewer convolutions than otherwise. However, the refinement network 1145 refines the pupil position to a more accurate position. Compared to using a single neural network, the initialization network 1105 and the refinement network 1145 together produce pupil position results more efficiently and more accurately.

[0109] The initialization network 1115 receives an input image 1105, which is 400×400 pixels in this example. The input image 1105 is resized to become a resized input image 1120, which is 50×50 pixels in this example. The initialization network 1115 includes five 3×3 convolutional layers 1130a, 1130b, 1130c, 1130d, 1130e and a fully connected layer 1135 at the output. The initialization network 1115 is configured to produce an output 1140 that identifies the x-coordinate and y-coordinate of the center of the pupil. For each glint, the output 1140 also includes an x-coordinate, a y-coordinate, and a visible / invisible indication.

[0110] The pupil x,y 1142 is input to the refinement network 1145 along with the input image 1105. The pupil x,y 1142 is used to perform a 96×96 pixel crop around the pupil location at full resolution (i.e., no downsampling) to produce a cropped input image 1150. The refinement network 1145 includes five 3×3 convolutional layers 1155a, 1155b, 1155c, 1155d, 1155e, a concatenation layer 1160, and a fully connected layer 1165 at the output. The concatenation layer 1160 concatenates the convolved features with the features from the initialization network 1115. In some implementations, the features from the initialization network 1115 encode global information (e.g., eye geometry / layout / eyelid position, etc.), and the features from the refinement network 1145 encode only local states (i.e., content that can be derived from the cropped image). By concatenating features from the initialization network 1115 and the refinement network 1145, the final fully connected layer in the refinement network 1145 can combine both global and local information and thereby generate a better estimate of the error than would be possible using only local information.

[0111] Refinement network 1145 is configured to produce output 1165 that identifies an estimated error 1175 that can be used to determine the x- and y-coordinates of the center of the pupil.

[0112] The refinement network 1145 thus acts as an error estimator for the initialization network 1115 and produces an estimated error 1175. This estimated error 1175 is used to adjust the initial pupil estimate 1180 from the initialization network 1115 to produce a refined pupil estimate 1185. For example, the initial pupil estimate 1180 may identify the pupil center at x, y: 10, 10, and the estimated error 1175 from the refinement network 1145 may indicate that the x of the pupil center should be greater by 2, thereby producing a refined pupil estimate of x, y: 12, 10.

[0113] In some implementations, the neural network 1110 is trained to avoid overfitting by introducing random noise. During the training phase, random noise (e.g., normally distributed noise with zero mean and small sigma) is added. In some implementations, random noise is added to the initial pupil position estimate. Without such random noise, the refinement network 11145 may otherwise learn to be overly optimistic. To avoid this, during training, the output of the initialized network 1115 is artificially made worse by adding random noise to simulate the situation where the initialized network 1115 is faced with data it has not seen before.

[0114] In some implementations, a stateful machine learning / neural network architecture is used. For example, the neural network used may be an LSTM or other recursive network. In such implementations, the event data used as input may be provided as a labeled sequential event stream. In some implementations, the recursive neural network is configured to remember previous events and learn the dynamic movement of the eyes based on the history of events. This stateful architecture can learn which eye movements are natural (e.g., eye twitches, blinks, etc.) and suppress these fluctuations.

[0115] In some implementations, gaze tracking is performed for both eyes of the same individual simultaneously. A single neural network receives input images of event camera data for both eyes and jointly determines the gaze characteristics of the eyes. In some implementations, one or more event cameras capture one or more images of a facial portion including both eyes. In implementations where images of both eyes are captured or derived, the network may determine or generate an output of a convergence point that helps determine the gaze direction of the two eyes. The network may additionally or alternatively be trained to account for special cases such as misaligned optical axes.

[0116] In some implementations, an ensemble of multiple networks is used. By combining the results of multiple neural networks (convolutional, recurrent, and / or other types), the variation in the output can be reduced. Each neural network can be trained with different hyperparameters (learning rate, batch size, architecture, etc.) and / or different data sets (e.g., using random subsampling).

[0117] In some implementations, post-processing of the gaze characteristics is employed. Filtering and prediction methods (e.g., using a Kalman filter) can be used to reduce noise in the tracking points. These methods can also be used to interpolate / extrapolate the gaze characteristics over time. For example, if the state of the gaze characteristic is required at a different timestamp than the recorded state, the method can be used.

[0118] Numerous specific details are set forth herein to provide a comprehensive understanding of the claimed subject matter. However, those skilled in the art will appreciate that the claimed subject matter may be practiced without these specific details. In other instances, methods, devices, or systems known to those of ordinary skill are not described in detail so as not to obscure the claimed subject matter.

[0119] Unless otherwise specifically noted, it should be understood that throughout the specification, discussions utilizing terms such as "process," "compute," "calculate," "determine," and "identify" refer to the actions or processes of a computing device, such as one or more computers or similar electronic computing devices, that manipulate or transform data represented as physical electronic or magnetic quantities within a memory, register, or other information storage device, transmission device, or display device of a computing platform.

[0120] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include multi-purpose microprocessor-based computer systems that access stored software that programs or configures the computing system from a general-purpose computing device to a dedicated computing device that implements one or more specific implementations of the subject matter of the present invention. Any suitable programming, scripting, or other type of language or combination of languages ​​may be used to implement the teachings contained herein in software for programming or configuring a computing device.

[0121] The specific implementation of the method disclosed herein can be performed in the operation of such a computing device. The order of the blocks presented in the above examples can be changed, for example, the blocks can be reordered, combined and / or divided into sub-blocks. Some blocks or processes can be executed in parallel.

[0122] The use of "suitable for" or "configured to" herein is meant to be open and inclusive language that does not exclude devices that are suitable for or configured to perform additional tasks or steps. In addition, the use of "based on" is meant to be open and inclusive, as a process, step, calculation, or other action "based on" one or more of the stated conditions or values ​​may in practice be based on additional conditions or values ​​beyond those stated. The headings, lists, and numbers included herein are for ease of explanation only and are not intended to be limiting.

[0123] It will also be understood that, although the terms "first", "second", etc. may be used to describe various elements in this article, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. For example, a first node may be referred to as a second node, and similarly, a second node may be referred to as a first node, which changes the meaning of the description, as long as all occurrences of the "first node" are consistently renamed and all occurrences of the "second node" are consistently renamed. Both the first node and the second node are nodes, but they are not the same node.

[0124] The terms used herein are only for describing specific implementations and are not intended to limit the claims. As used in the description of this specific implementation and the appended claims, the singular forms of "a" and "the" are intended to also cover the plural forms, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" used herein refers to and covers any and all possible combinations of one or more items in the associated listed items. It will also be understood that the term "comprising" when used in this specification specifies the presence of stated features, integers, steps, operations, elements and / or parts, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts, and / or their grouping.

[0125] As used herein, the term “if” may be interpreted to mean “when the antecedent is true” or “when the antecedent is true” or “in response to determining” or “upon determining” or “in response to detecting” that the antecedent is true, depending on the context. Similarly, the phrase “if it is determined that [the antecedent is true]” or “if [the antecedent is true]” or “when [the antecedent is true]” is interpreted to mean “upon determining that the antecedent is true” or “in response to determining” or “upon determining” that the antecedent is true or “when detecting that the antecedent is true” or “in response to detecting” that the antecedent is true, depending on the context.

[0126] The foregoing description and summary of the present invention should be understood to be illustrative and exemplary in every aspect, rather than restrictive, and the scope of the present invention disclosed herein is determined not only by the detailed description of the exemplary specific implementations, but by the full breadth allowed by the patent law. It should be understood that the specific implementations shown and described herein are only illustrative of the principles of the present invention, and that various modifications can be implemented by those skilled in the art without departing from the scope and essence of the present invention.

Claims

1. A system comprising: a non-transitory computer-readable storage medium; and One or more processors, the one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions, and when the program instructions are executed on the one or more processors, the electronic device performs operations including the following: receiving a stream of pixel events output by an event camera, the event camera comprising a plurality of pixel sensors positioned to receive light from a surface of an eye, each respective pixel event being generated in response to the respective pixel sensor detecting a change in light intensity of light at a respective event camera pixel that exceeds a comparator threshold; generating a first input from the pixel event stream; generating an initial pupil characteristic based on inputting the first input to the first neural network; determining a second input comprising a subset of the first input based on the initial pupil characteristic; generating a correction to the pupil characteristic based on inputting the second input to the second neural network; as well as A gaze characteristic is determined by adjusting the initial pupil characteristic using the correction.

2. The system of claim 1, wherein the initial gaze characteristic is an initial pupil center. 3 . The system of claim 1 , wherein generating a first input from the pixel event stream comprises deriving an image from the pixel event stream. The system of claim 3 , wherein generating the first input comprises generating a lower resolution image from the image. 5 . The system of claim 3 , wherein deriving an image comprises accumulating pixel events in the pixel event stream for a plurality of event camera pixels. The system of claim 3 , wherein the subset of the first input comprises a portion of the image.

7. The system of claim 6, wherein the subset of the first input comprises a small input image comprising a portion of the image centered at a location of the initial gaze characteristic, wherein the initial gaze characteristic is an initial pupil center.

8. The system of claim 6, wherein the portion of the image and the image have the same resolution, and the portion of the image has fewer pixels than the image.

9. A non-transitory computer-readable storage medium storing program instructions, the program instructions being executable by one or more processors to perform operations comprising: receiving a stream of pixel events output by an event camera, the event camera comprising a plurality of pixel sensors positioned to receive light from a surface of an eye, each respective pixel event being generated in response to the respective pixel sensor detecting a change in light intensity of light at a respective event camera pixel that exceeds a comparator threshold; generating a first input from the pixel event stream; generating an initial pupil characteristic based on inputting the first input to the first neural network; determining a second input comprising a subset of the first input based on the initial pupil characteristic; generating a correction to the pupil characteristic based on inputting the second input to the second neural network; as well as A gaze characteristic is determined by adjusting the initial pupil characteristic using the correction.

10. The non-transitory computer-readable storage medium of claim 9, wherein: Generating a first input from the pixel event stream includes deriving an image by accumulating pixel events in the pixel event stream for a plurality of event camera pixels; The subset of the first input is a portion of the image that is smaller than the entire image, centered at the location of the initial fixation feature; and The initial gaze characteristic is the initial pupil center.

11. A method for gaze tracking based on an event camera, the method comprising: At a device having one or more processors and a computer-readable storage medium: receiving a stream of pixel events output by an event camera, the event camera comprising a plurality of pixel sensors positioned to receive light from a surface of an eye, each respective pixel event being generated in response to the respective pixel sensor detecting a change in light intensity of light at a respective event camera pixel that exceeds a comparator threshold; generating a first input from the pixel event stream; generating an initial pupil characteristic based on inputting the first input to the first neural network; determining a second input comprising a subset of the first input based on the initial pupil characteristic; generating a correction to the pupil characteristic based on inputting the second input to the second neural network; as well as A gaze characteristic is determined by adjusting the initial pupil characteristic using the correction.

12. The method of claim 11, wherein the initial pupil characteristic is an initial pupil center.

13. The method of claim 11, wherein generating a first input from the pixel event stream comprises deriving an image from the pixel event stream. The method of claim 13 , wherein generating a first input comprises generating a lower resolution image from the image.

15. The method of claim 13, wherein deriving an image comprises accumulating pixel events in the pixel event stream for a plurality of event camera pixels. The method of claim 13 , wherein the subset of the first input comprises a portion of the image.

17. The method of claim 16, wherein the subset of the first input is a portion of the image centered about the location of the initial gaze characteristic, wherein the initial gaze characteristic is an initial pupil center.

18. The method of claim 16, wherein the portion of the image and the image have the same resolution and the portion of the image has fewer pixels than the image.

19. The method of claim 11, wherein the first neural network is a recurrent neural network.

20. The method of claim 11, further comprising identifying an item displayed on a display based on the gaze characteristic.

Citation Information

Patent Citations

  • Method of carrying out iris recognition through neural networks

    CN106778567A

  • Method of Determining Reflections of Light

    US20140002349A1

  • Eye tracker system and methods for detecting eye parameters

    US20170188823A1

  • Method and apparatus for event sampling of dynamic vision sensor on image formation

    US20170213105A1

  • Multiplexed temporal calibration for event-based cameras

    US9489735B1