Eye tracking system for head-mounted display devices

By using multiple eye tracking components and machine learning models in head-mounted display devices, the problems of insufficient eye tracking accuracy and waste of computing resources in existing technologies are solved, efficient gaze direction determination and dynamic adjustment of the display are achieved, and user experience and resource utilization efficiency are improved.

CN114730217BActive Publication Date: 2025-09-26VALVE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180006626.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-27
Filing Date
2021-01-20
Publication Date
2025-09-26
Estimated Expiration
2041-01-20

AI Technical Summary

Technical Problem

Existing head-mounted display devices have problems with insufficient accuracy and waste of computing resources in eye tracking technology. Especially in virtual reality and augmented reality systems, it is difficult to efficiently determine the user's gaze direction and dynamically adjust the displayed content.

Method used

It uses multiple eye tracking components, including light sources and light detectors, combined with polarizers and machine learning models, to determine the user's gaze direction through optical detection and machine learning algorithms, and dynamically adjust the output of the display based on this information, using the interpupillary distance adjustment component to optimize the display alignment.

Benefits of technology

Improves the accuracy and efficiency of eye tracking, reduces the waste of computing resources, dynamically adjusts display content to enhance user experience, and saves power and bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114730217B_ABST
    Figure CN114730217B_ABST
Patent Text Reader

Abstract

The present disclosure relates to systems and methods for eye tracking for use in various applications, such as virtual reality or augmented reality applications including head-mounted display devices. An eye tracking subsystem may be provided that includes multiple components, each component including a light source (e.g., an LED), a light detector (e.g., a silicon photodiode), and a polarizer. The polarizer is configured to prevent light reflected via specular reflection from reaching the light detector, such that the light detector detects only scattered light. Machine learning or other techniques may be used to track or otherwise determine a user's gaze direction, which may be used by one or more components of the HMD device to improve its functionality in various ways.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The following disclosure relates generally to techniques for eye tracking and, more particularly, to techniques for eye tracking in head-mounted display devices. Background Art

[0002] A head-mounted display (HMD) device or system is an electronic device that is worn on a user's head and, when so worn, fixes at least one electronic display within the visible area of ​​at least one of the user's eyes, regardless of the position or orientation of the user's head. HMD devices used to implement virtual reality systems typically completely surround the wearer's eyes and replace the actual view (or actual reality) in front of the wearer with a "virtual" reality, while HMD devices used for augmented reality systems typically provide a translucent or transparent overlay of one or more screens in front of the wearer's eyes, thereby augmenting the actual view with additional information. For augmented reality systems, the "display" component of the HMD device is either transparent or peripheral to the user's field of view so that it does not completely block the user from seeing their external environment. Summary of the Invention

[0003] A head-mounted display (HMD) system can be generally described as comprising a support structure wearable on a user's head; an eye-tracking subsystem connected to the support structure, the eye-tracking subsystem comprising a plurality of eye-tracking components, each including a light source for emitting light; a light detector for detecting the light; and a polarizer located proximate to at least one of the light source and the light detector, the polarizer configured to prevent light reflected via specular reflection from being received by the light detector; at least one processor; and a memory storing a set of instructions or data that, as a result of execution, causes the HMD system to: selectively cause the light sources of the plurality of eye-tracking components to emit light; receive light detection information captured by the light detectors of the plurality of eye-tracking components; provide the received light detection information as input to a prediction model; receive a determined gaze direction of the user's eyes from the prediction model in response to providing the light detection information; and provide the determined gaze direction to components associated with the HMD system for use. The light detection information may include a characteristic radiation pattern of each light source after light from the light source has reflected, scattered, or absorbed off the user's face or eyes. Each light source may be directed toward an expected location of the user's pupil. The light source may include a light emitting diode (LED) that emits light with a wavelength between 780 nm and 1000 nm. The light detector may include a photodiode. The eye tracking subsystem may include four eye tracking components positioned to determine the gaze direction of the user's left eye, and four eye tracking components positioned to determine the gaze direction of the user's right eye. The polarizer may include two crossed linear polarizers. For each of the multiple eye tracking components, the polarizer may include a first polarizer located in the light emission path of the light source and a second polarizer located in the light detection path of the light detector. The polarizer may include at least one of a circular polarizer and a linear polarizer. For each of the multiple eye tracking components, the light source may be positioned away from the optical axis of the user's eye to provide darkfield illumination of the pupil. The prediction model may include a machine learning model or other type of function or model (e.g., a polynomial, a lookup table). The machine learning model may include a mixture density network (MDN) model. The machine learning model may include a recurrent neural network (RNN) model. The machine learning model may utilize past input information or eye movement information to determine gaze direction. The machine learning model may be a model trained during field operation of the multiple HMD systems. The HMD system may include at least one display, and at least one processor may cause the at least one display to present a user interface element; may selectively cause light sources of multiple eye tracking components to emit light; may receive light detection information captured by light detectors of the multiple eye tracking components; and may update a machine learning model based at least in part on the received light detection information and known or inferred gaze direction information associated with the received light detection information. The user interface element may include a static user interface element or a moving user interface element.The HMD system may include at least one display, and the at least one processor may dynamically change a rendered output of the at least one display based at least in part on a determined gaze direction. The HMD system may include an interpupillary distance (IPD) adjustment component, and the at least one processor may cause the IPD adjustment component to align at least one component of the HMD system for a user based at least in part on the determined gaze direction.

[0004] A method for operating a head-mounted display (HMD) system. The HMD system may include an eye-tracking subsystem connected to a support structure, the support structure including a plurality of eye-tracking components, each eye-tracking component including a light source, a light detector, and a polarizer. The method may be summarized as including selectively illuminating the light sources of the plurality of eye-tracking components; receiving light detection information captured by the plurality of light detectors; providing the received light detection information as input to a trained machine learning model; receiving a determined gaze direction of a user's eyes from the machine learning model in response to providing the light detection information; and providing the determined gaze direction to a component associated with the HMD system for use. Providing the received light detection information as input to the trained machine learning model may include providing the received light detection information as input to a mixture density network (MDN) model. Providing the received light detection information as input to the trained machine learning model may include providing the received light detection information as input to a recurrent neural network (RNN) model. Providing the received light detection information as input to the trained machine learning model may include providing the received light detection information as input to a machine learning model that utilizes past input information or eye movement information to determine gaze direction.

[0005] The method may also include training the machine learning model during field operation of multiple HMD systems. The HMD system may include at least one display, and the method may include causing the at least one display to present a user interface element; selectively causing light sources of multiple eye tracking components to emit light; receiving light detection information captured by multiple light detectors; and updating the machine learning model based at least in part on the received light detection information and known or inferred gaze direction information associated with the received light detection information. Causing the at least one display to present the user interface element may include causing the at least one display to present a static user interface element or a moving user interface element. The HMD system may include at least one display, and the method may include dynamically changing an output of the at least one display based at least in part on the determined gaze direction.

[0006] The method may also include mechanically aligning at least one component of the HMD system for the user based at least in part on the determined gaze direction.

[0007] A head-mounted display (HMD) system can be generally described as comprising a support structure wearable on a user's head; an eye-tracking subsystem coupled to the support structure, the eye-tracking subsystem comprising a plurality of eye-tracking components, each eye-tracking component comprising: a light-emitting diode; a photodiode; and a polarizer positioned proximate to at least one of the light-emitting diode and the photodiode, the polarizer configured to prevent light reflected via specular reflection from being received by the photodiode; at least one processor; and a memory storing a set of instructions or data. The instructions or data, when executed, cause the HMD system to: selectively illuminate the light-emitting diodes of the plurality of eye-tracking components; receive light detection information captured by the photodiodes of the plurality of eye-tracking components; provide the received light detection information as input to a trained machine learning model; receive a determined gaze direction of the user's eyes from the machine learning model in response to providing the light detection information; and dynamically alter operation of components associated with the HMD system based at least in part on the determined gaze direction. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a schematic diagram of a networked environment that includes one or more systems suitable for performing at least some of the techniques described in this disclosure, including an embodiment of an eye tracking subsystem.

[0009] Figure 2 is a diagram illustrating an example environment in which at least some of the described techniques are used with an example head-mounted display device that is tethered to a video rendering computing system and provides a virtual reality display to a user.

[0010] Figure 3 is a front view of an HMD device with a binocular display subsystem.

[0011] Figure 4 A top view of an HMD device with a binocular display subsystem and various sensors is shown according to an exemplary embodiment of the present disclosure.

[0012] Figure 5 An example of using a light source and a light detector to determine pupil position is shown, such as for eye tracking in an HMD device according to the techniques described in this disclosure.

[0013] Figure 6 is a schematic diagram of an environment in which machine learning techniques may be used with an eye tracking subsystem of an embodiment HMD device, according to one non-limiting illustrative embodiment.

[0014] Figure 7 is a flow chart of a method of operating an HMD device including eye tracking capabilities, according to one non-limiting, illustrative embodiment. DETAILED DESCRIPTION

[0015] In the following description, certain specific details are set forth to provide a thorough understanding of the various disclosed embodiments. However, one skilled in the relevant art will recognize that the embodiments can be practiced without one or more of these specific details, or with other methods, components, materials, etc. In other instances, well-known structures associated with computer systems, server computers, and / or communication networks are not shown or described in detail to avoid unnecessarily obscuring the description of the embodiments.

[0016] Throughout the specification and following claims, the word "comprising" is used synonymously with "including" and is inclusive or open-ended (ie, does not exclude other unrecited elements or methodological acts) unless the context requires otherwise.

[0017] Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0018] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should also be noted that the term "or" is generally employed in its sense including "and / or" unless the context clearly dictates otherwise.

[0019] The titles and abstracts of the disclosure provided herein are for convenience only and do not interpret the scope or meaning of the embodiments.

[0020] Eye tracking is a process by which the position, orientation, or movement of the eye can be measured, detected, sensed, determined, or monitored (collectively, "measured"). In many applications, this is done to determine the direction of a user's gaze. The position, orientation, or movement of the eye can be measured in a variety of different ways, the least intrusive of which can employ one or more optical detectors or sensors to optically track the eye. Some techniques can include illuminating or flooding the entire eye with infrared light at once and measuring the reflection with at least one optical sensor tuned to be sensitive to infrared light. Information about how the infrared light is reflected from the eye is analyzed to determine the position, orientation, and / or movement of one or more eye features (e.g., the cornea, pupil, iris, or retinal blood vessels).

[0021] Eye tracking functionality is highly advantageous in wearable head-mounted display system applications. Some examples of the utility of eye tracking in head-mounted display systems include affecting the position of displayed content in a user's field of view by changing the display of content outside the user's field of view (e.g., viewpoint rendering), saving power, bandwidth, or computing resources, and affecting what content is displayed to the user, determining where the user is looking or gazing, determining whether the user is looking at content displayed on the display, and providing a method by which the user can control or interact with the displayed content, among other applications.

[0022] The present disclosure generally relates to techniques for eye tracking. Such techniques can be used, for example, in head-mounted display ("HMD") devices for VR or AR applications. Some or all of the techniques described herein can be performed via automatic operation of an embodiment of an eye tracking subsystem, such as by one or more configured hardware processors and / or other configured hardware circuits. The one or more hardware processors or other configured hardware circuits of such a system or device may include, for example, one or more GPUs ("graphics processing units") and / or CPUs ("central processing units") and / or other microcontrollers ("MCUs") and / or other integrated circuits, e.g., the hardware processor is part of an HMD device or other device that includes one or more display panels to display image data, or is part of a computing system that generates or otherwise prepares image data to be sent to a display panel for display, as discussed further below. More generally, such hardware processors or other configured hardware circuits may include, but are not limited to, one or more application-specific integrated circuits (ASICs), standard integrated circuits, controllers (e.g., by executing appropriate instructions, and including microcontrollers and / or embedded controllers), field programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), digital signal processors (DSPs), programmable logic controllers (PLCs), and the like. Additional details are included elsewhere in this document, including regarding the Figure 1 .

[0023] Technical benefits in at least some embodiments of the described techniques include addressing and mitigating increased media transmission bandwidth for image encoding by reducing image data size, improving the speed of controlling display panel pixels (e.g., based at least in part on the corresponding reduced image data size), improving foveated image systems and other techniques for reflecting subsets of display panels and / or images of particular interest, and more. Foveated image encoding systems exploit specific aspects of the human visual system (which can provide detailed information only at and around focal points), but typically use specialized computational processing to avoid visual artifacts to which peripheral vision is particularly sensitive (e.g., artifacts related to motion and contrast in video and image data). In the case of certain VR and AR displays, both bandwidth and computational usage for processing high-resolution media are exacerbated because certain display devices involve two separate display panels (i.e., one for each eye) with two separately addressable pixel arrays, each associated with an appropriate resolution. Thus, the described techniques can be used, for example, to reduce transmission bandwidth for local and / or remote display of video frames or other images, while maintaining resolution and detail in a viewer's "region of interest" within the image, while also minimizing computational usage for processing such image data. Additionally, the use of lenses and other displays in a head-mounted display device can provide greater focus or resolution on a subset of the display panel, such that when such techniques are used in such embodiments, using such techniques to display lower resolution information in other portions of the display panel can provide further benefits.

[0024] For illustrative purposes, some embodiments are described below in which specific types of information are acquired and used in specific types of structures in specific ways and using specific types of devices. However, it should be understood that the techniques described in this manner can be used in other ways in other embodiments, and therefore the present disclosure is not limited to the exemplary details provided. As a non-exclusive example, the various embodiments discussed herein include the use of images as video frames. However, although many of the examples described herein refer to "video frames" for convenience, it should be understood that the techniques described with reference to these examples can be used with respect to one or more images of various types, including the non-exclusive examples of a continuous plurality of video frames (e.g., at 30, 60, 90, 180 or some other amount of frames per second), other video content, photographs, computer-generated graphics content, other visual media articles, or combinations thereof. In addition, various details are provided in the drawings and text for illustrative purposes, but are not intended to limit the scope of the present disclosure. In addition, as used herein, "pixel" refers to the smallest addressable image element of a display that can be activated to provide all possible color values ​​for the display. In many cases, a pixel includes separate corresponding sub-elements (in some cases referred to as separate "sub-pixels") for separately producing red, green, and blue light for perception by a human viewer, with separate color channels being used to encode pixel values ​​for the differently colored sub-pixels. A pixel "value," as used herein, refers to a data value corresponding to a respective stimulus level of one or more of those respective RGB elements of a single pixel.

[0025] Figure 1 is a schematic diagram of a networked environment 100 that includes a local media presentation (LMR) system 110 (e.g., a gaming system) that includes a local computing system 120 and a display device 180 (e.g., an HMD device with two display panels) that is suitable for performing at least some of the techniques described herein. Figure 1 In the embodiment shown, the local computing system 120 is connected to the network via a transmission link 115 (which may be wired or tethered, for example, via a Figure 2 The local computing system 120 is communicatively connected to the display device 180 via one or more cables (cable 220) as shown, or alternatively may be wireless. In other embodiments, the local computing system 120 may provide encoded image data for display to a panel display device (e.g., a TV, console, or monitor) via a wired or wireless link, whether in addition to or in place of the HMD device 180, and the display devices each include one or more addressable pixel arrays. In various embodiments, the local computing system 120 may include a general-purpose computing system, a game console, a video streaming device, a mobile computing device (e.g., a cell phone, PDA, or other mobile device), a VR or AR processing device, or other computing system.

[0026] In the illustrated embodiment, the local computing system 120 has components including one or more hardware processors (e.g., a centralized processing unit or "CPU") 125, memory 130, various I / O ("input / output") hardware components 127 (e.g., a keyboard, a mouse, one or more game controllers, a speaker, a microphone, an IR transmitter and / or receiver, etc.), a video subsystem 140 including one or more dedicated hardware processors (e.g., a graphics processing unit or "GPU") 144 and video memory (VRAM) 148, computer-readable memory 150, and a network connection 160. Also in the illustrated embodiment, an embodiment of an eye-tracking subsystem 135 executes in the memory 130 to perform at least some of the described techniques, such as by using the CPU 125 and / or GPU 144 to perform automated operations that implement those described techniques, and the memory 130 may optionally further execute one or more other programs 133 (e.g., to generate video or other images for display, such as a game program). As part of automated operations implementing at least some of the techniques described herein, the eye tracking subsystem 135 and / or program 133 executing in memory 130 may store or retrieve various types of data, including in an example database data structure in memory 150, in which example the data used may include various types of image data information in a database ("DB") 154, various types of application data in DB 152, various types of configuration data in DB 157, and may include additional information, such as system data or other information.

[0027] In the depicted embodiment, the LMR system 110 is also communicatively connected via one or more computer networks 101 and network links 102 to an exemplary network-accessible media content provider 190, which may further provide content to the LMR system 110 for display, either in addition to or in lieu of the image generation program 133. The media content provider 190 may include one or more computing systems (not shown), each of which may have components similar to those of the local computing system 120, including one or more hardware processors, I / O components, local storage devices, and memory, although for the sake of brevity, some details are not shown for the network-accessible media content providers.

[0028] It should be understood that although Figure 1In the illustrated embodiment, the display device 180 is depicted as distinct and separate from the local computing system 120 , but in certain embodiments, some or all components of the local media rendering system 110 may be integrated or housed within a single device, such as a mobile gaming device, a portable VR entertainment system, an HMD device, etc. In such embodiments, the transmission link 115 may, for example, include one or more system buses and / or video bus architectures.

[0029] As an example involving operations performed locally by local media rendering system 120, assume that the local computing system is a gaming computing system, such that application data 152 includes one or more gaming applications executed by CPU 125 using memory 130, and various video frame display data is generated and / or processed by image generation program 133, e.g., in conjunction with GPU 144 of video subsystem 140. To provide a high-quality gaming experience, a large amount of video frame data (corresponding to a high image resolution for each video frame, and a high "frame rate" of approximately 60-180 such video frames per second) is generated by local computing system 120 and provided to display device 180 via wired or wireless transmission link 115.

[0030] It will also be understood that computing system 120 and display device 180 are only illustrative and are not intended to limit the scope of the present invention. Computing system 120 can alternatively include multiple interactive computing systems or devices, and can be connected to other devices not shown, including through one or more networks such as the Internet, via the Web, or via a dedicated network (e.g., a mobile communication network, etc.). More generally, a computing system or other computing node can include any combination of hardware or software that can interact and perform the functions of the type, including but not limited to desktop or other computers, game systems, database servers, network storage devices and other network devices, PDAs, cellular phones, wireless phones, pagers, electronic organizers, Internet appliances, systems based on television (e.g., using set-top boxes and / or personal / digital video recorders) and various other consumer products including appropriate communication capabilities. Display device 180 can similarly include one or more devices with one or more display panels of various types and forms, and optionally include various other hardware and / or software components.

[0031] 135 or its components) and / or data structures (e.g., by executing software instructions of one or more software programs and / or by storing such software instructions and / or data structures). Some or all of the components, systems, and data structures may also be stored (e.g., as software instructions or structured data) on a non-transitory computer-readable storage medium, such as a hard disk or flash drive or other non-volatile storage device, volatile or non-volatile memory (e.g., RAM), a network storage device, or a portable media product (e.g., a DVD, CD, optical disk, etc.) to be read by an appropriate drive (e.g., RAM) or via an appropriate connection. In some embodiments, the systems, components, and data structures may also be transmitted as a generated data signal (e.g., as part of a carrier wave or other analog or digital propagation signal) on various computer-readable transmission media, including wireless and wired / cable-based media, and may take various forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). In other embodiments, such a computer program product may also take other forms. Therefore, the present invention may be implemented using other computer system configurations.

[0032] Figure 2An exemplary environment 200 is shown in which at least some of the described techniques are used with an exemplary HMD device 202, which is connected to a video rendering computing system 204 via a tethered connection 220 (or a wireless connection in other embodiments) to provide a virtual reality display to a human user 206. The user wears the HMD device 202 and receives display information of a simulated environment, distinct from the actual physical environment, from the computing system 204 via the HMD device, where the computing system acts as an image rendering system that provides images of the simulated environment to the HMD device for display to the user, such as images generated by a game program and / or other software program executing on the computing system. In this example, the user is also able to move within a tracked volume 201 of the actual physical environment 200 and may also have one or more I / O ("input / output") devices to allow the user to further interact with the simulated environment, which in this example includes handheld controllers 208 and 210.

[0033] In the example shown, the environment 200 may include one or more base stations 214 (two are shown, labeled base stations 214a and 214b), which can help track the HMD device 202 or controllers 208 and 210. As the user moves the position of the HMD device 202 or changes the orientation of the HMD device 202, the position of the HMD device is tracked to allow the corresponding portion of the simulated environment to be displayed to the user on the HMD device, and the controllers 208 and 210 can also use similar techniques to track the position of the controllers (and optionally use this information to help determine or verify the position of the HMD device). After the tracked position of the HMD device 202 is known, the corresponding information is sent to the computing system 204 via the tether 220 or wirelessly, which uses the tracked position information to generate one or more subsequent images of the simulated environment for display to the user.

[0034] A variety of different position tracking methods may be used in various embodiments of the present disclosure, including but not limited to acoustic tracking, inertial tracking, magnetic tracking, optical tracking, combinations thereof, and the like.

[0035] In at least some embodiments, the HMD device 202 may include one or more optical receivers or sensors that can be used to implement tracking functionality or other aspects of the present disclosure. For example, the base stations 214 can each scan for optical signals on the tracked volume 201. Each base station 214 can generate more than one optical signal, depending on the requirements of each particular embodiment. For example, while a single base station 214 is generally sufficient for six-degree-of-freedom tracking, in some embodiments, multiple base stations (e.g., base stations 214a, 214b) may be needed or desired to provide robust room-scale tracking for the HMD device and peripheral devices. In this example, the optical receivers are incorporated into the HMD device 202 and / or other tracked objects, such as controllers 208 and 210. In at least some embodiments, the optical receivers can be paired with accelerometers and gyroscopes on each tracked device. Inertial measurement units ("IMUs") can be used to support low-latency sensor fusion.

[0036] In at least some embodiments, each base station 214 includes two rotors that scan linear beams across the tracked volume 201 on orthogonal axes. At the beginning of each scanning cycle, the base station 214 can transmit an omnidirectional light pulse (referred to as a "synchronization signal") that is visible to all sensors on the tracked object. Thus, each sensor calculates a unique angular position in the scanned volume by timing the duration between the synchronization signal and the beam signal. Sensor range and direction can be resolved using multiple sensors fixed to a single rigid body.

[0037] One or more sensors located on the tracked object (e.g., HMD device 202, controllers 208 and 210) can include optoelectronic devices capable of detecting modulated light from the rotor. For visible light or near-infrared (NIR) light, a silicon photodiode and appropriate amplifier / detector circuitry can be used. Because the environment 200 may contain static and time-varying signals (optical noise) having wavelengths similar to those of the base station 214 signal, in at least some embodiments, the base station light can be modulated in such a way that it is easily distinguishable from any interfering signals and / or the sensor can be filtered from radiation of any wavelength other than the base station signal.

[0038] Inside-out tracking is also a type of positional tracking that can be used to track the position of the HMD device 202 and / or other objects (e.g., controllers 208 and 210, tablet computers, smartphones). Inside-out tracking differs from outside-in tracking in the location of the camera or other sensor used to determine the position of the HMD. With inside-out tracking, the camera or sensor is located on the HMD or on the object being tracked, while with outside-out tracking, the camera or sensor is placed at a fixed location in the environment.

[0039] An HMD that utilizes inside-out tracking utilizes one or more cameras to “look outside” to determine how its position has changed relative to the environment. As the HMD moves, the sensors readjust their position in the room, and the virtual environment responds accordingly in real time. This type of positional tracking can be implemented with or without markers placed in the environment. Cameras placed on the HMD observe features of the surrounding environment. When markers are used, the markers are designed to be easily detected by the tracking system and are placed in specific areas. With “markerless” inside-out tracking, the HMD system uses unique features that are initially present in the environment (e.g., natural features) to determine position and orientation. The HMD system’s algorithms recognize specific images or shapes and use them to calculate the device’s position in space. Data from accelerometers and gyroscopes can also be used to increase the accuracy of positional tracking.

[0040] Figure 3 Information 300 shows a front view of an exemplary HMD device 344 when worn on the head of a user 342. The HMD device 344 includes a front-facing structure 343 supporting a front-facing or forward-facing camera 346 and a plurality of sensors 348a-348d (collectively, 348) of one or more types. As one example, some or all of the sensors 348 can help determine the position and orientation of the device 344 in space, such as light sensors, to detect and use information from one or more external devices (not shown, e.g., Figure 2 214). As shown, the forward-facing camera 346 and sensors 348 are directed forward toward the actual scene or environment (not shown) in which the user 342 is operating the HMD device 344. The actual physical environment may include, for example, one or more objects (e.g., walls, ceilings, furniture, stairs, cars, trees, tracking markers, or any other type of object). The specific number of sensors 348 may be less than or greater than the number of sensors shown. The HMD device 344 may further include one or more additional components not attached to the forward-facing structure (e.g., within the HMD device), such as an IMU (inertial measurement unit) 347 electronic device that measures and reports specific forces, angular rates, and / or magnetic fields surrounding the HMD device 344 (e.g., using a combination of accelerometers, gyroscopes, and, optionally, magnetometers). The HMD device may also include additional components not shown, including one or more display panels and optical lens systems that are oriented toward the user's eyes (not shown) and optionally have one or more internal motors connected to change the alignment or other positioning of the one or more optical lens systems and / or display panels within the HMD device, as described below with respect to Figure 4 discussed in more detail.

[0041] The illustrated example of an HMD assembly 344 is based, at least in part, on the head of a user 342, supported by one or more straps 345 connected to the housing of the HMD assembly 344 and extending in whole or in part around the user's head. Although not shown here, the HMD assembly 344 may also have one or more external motors, for example, connected to one or more straps 345, and the automatic corrective action may include using such motors to adjust such straps to change the alignment or other positioning of the HMD assembly on the user's head. It should be understood that the HMD assembly may include other support structures (e.g., a nose piece, a chin strap, etc.) not shown here, either in addition to or in place of the illustrated straps, and some embodiments may include motors connected to one or more of these other support structures to similarly adjust their shape and / or position, thereby changing the alignment or other positioning of the HMD assembly on the user's head. Other display devices not secured to the user's head may similarly be connected to one or more structures, or portions thereof, that affect the positioning of the display devices, and in at least some embodiments may include motors or other mechanical actuators to similarly change their shape and / or position, thereby changing the alignment or other positioning of the display devices relative to one or more pupils of one or more users of the display devices.

[0042] Figure 4 A simplified top plan view 400 of an HMD device 405 is shown that includes a pair of near-eye display systems 402 and 404. The HMD device 405 may be, for example, Figure 1-3 The same or similar HMD device or a different HMD device as shown in , and the HMD devices discussed herein can be further used in the examples discussed below. Figure 4 The near-eye display systems 402 and 404 of FIG. 4 include display panels 406 and 408 (e.g., OLED microdisplays), respectively, and optical lens systems 410 and 412, each having one or more optical lenses. Display systems 402 and 404 can be mounted to or otherwise positioned within a housing (or frame) 414 that includes a front-facing portion 416 (e.g., Figure 3 The two display systems 402 and 404 can be fixed to the housing 414 in an eyeglass assembly, which can be worn on the head 422 of the wearing user 424, with the left and right eyeglass arms 418 and 420 resting on the user's ears 426 and 428, respectively, and the nose assembly 492 can rest on the user's nose 430. Figure 4In the example of FIG, the HMD device 405 can be partially or fully supported on the user's head by the nose display and / or the left and right earpieces, although in some embodiments a strap (not shown) or other structure can be used to secure the HMD device to the user's head, for example Figure 2 and 3 . The housing 414 can be shaped and sized to position each of the two optical lens systems 410 and 412 in front of one of the user's eyes 432 and 434, respectively, so that the target position of each pupil 494 is centered vertically and horizontally in front of the corresponding optical lens system and / or display panel. Although the housing 414 is shown in a simplified manner similar to glasses for illustrative purposes, it should be understood that more complex structures (e.g., goggles, integrated headband, helmet, straps, etc.) can actually be used to support and position the display systems 402 and 404 on the head 422 of the user 424.

[0043] Figure 4 The HMD device 405 and other HMD devices discussed herein are capable of presenting a virtual reality display to a user, for example via corresponding video presented at a display rate such as 30 or 60 or 90 frames (or images) per second, while other embodiments of similar systems may present an augmented reality display to a user. Figure 4 Each of displays 406 and 408 can generate light that is transmitted through corresponding optical lens systems 410 and 412 and focused onto eyes 432 and 434 of user 424, respectively. The pupil 494 of each eye, through which light enters the eye, typically has a pupil size ranging from 2 mm (millimeters) in diameter under very bright conditions to as much as 8 mm under dark conditions, while the larger iris containing the pupil can have a size of approximately 12 mm - the pupil (and closed iris) can also typically move several millimeters in the horizontal and / or vertical directions within the visible part of the eye under open eyelids. When the eyeball rotates about its center (resulting in a three-dimensional volume in which the pupil can move), this orientation will also cause the pupil to move to different depths from the optical lenses or other physical elements of the display for different horizontal and vertical positions. The light entering the user's pupil is perceived by user 424 as an image and / or video. In some embodiments, the distance between each of optical lens systems 410 and 412 and the user's eyes 432 and 434 can be relatively short (e.g., less than 30 mm, less than 20 mm), which advantageously makes the HMD device appear lighter to the user because the weight of the optical lens system and display system is relatively close to the user's face, and can also provide the user with a larger field of view. Although not shown here, some embodiments of such an HMD device can include various additional internal and / or external sensors.

[0044] In the embodiment shown, Figure 4The HMD device 405 also includes hardware sensors and additional components, such as one or more accelerometers and / or gyroscopes 490 (e.g., as part of one or more IMU units). As discussed in more detail elsewhere herein, the values ​​from the accelerometers and / or gyroscopes can be used to locally determine the orientation of the HMD device. Furthermore, the HMD device 405 can include one or more front-facing cameras, such as camera 485 on the exterior of the front portion 416, and information from these cameras can be used as part of the operation of the HMD device, such as to provide AR functionality or positioning functionality. Furthermore, the HMD device 405 can also include other components 475 (e.g., electronic circuitry to control the display of images on the display panels 406 and 408, internal memory, one or more batteries, location tracking devices that interact with external base stations, etc.), as discussed in more detail elsewhere herein. Other embodiments may not include one or more of components 475, 485, and / or 490. Although not shown here, some embodiments of such HMD devices can include various additional internal and / or external sensors to track various other types of motion and position of the user's body, eyes, controllers, etc.

[0045] In the embodiment shown, Figure 4 The HMD device 405 also includes hardware sensors and additional components that can be used by the disclosed embodiments as part of the described techniques for determining the user's pupils or gaze direction, which can be provided to one or more components associated with the HMD device for use, as discussed elsewhere herein. The hardware sensors in this example include one or more eye tracking components 472 of the eye tracking subsystem, which are mounted on or near the display panels 406 and 408 and / or on the inner surface 421 located near the optical lens systems 410 and 412, for obtaining information about the actual position of the user's pupils 494, such as, for example, separately for each pupil in this example.

[0046] Each eye tracking assembly 472 may include one or more light sources (e.g., IR LEDs) and one or more light detectors (e.g., silicon photodiodes). Figure 4 Only four full eye tracking assemblies 472 are shown, but it should be understood that a different number of eye tracking assemblies may be provided. In some embodiments, a total of eight eye tracking assemblies 472 are provided, four for each eye of user 424. Furthermore, in at least some embodiments, each eye tracking assembly includes a light source directed toward one of eyes 432 and 434 of user 424, a light detector positioned to receive light reflected by the user's corresponding eye, and a polarizer positioned and configured to prevent light reflected via specular reflection from being applied to the light detector.

[0047] As discussed in greater detail elsewhere herein, information from the eye tracking assembly 472 can be used to determine and track the user's gaze direction during use of the HMD device 405. Furthermore, in at least some embodiments, the HMD device 405 can include one or more internal motors 438 (or other movement mechanisms) that can be used to move 439 the alignment and / or other positioning (e.g., in vertical, horizontal left-right, and / or horizontal front-to-back directions) of one or more optical lens systems 410 and 412 and / or display panels 406 and 408 within the housing of the HMD device 405, such as to personalize or otherwise adjust the target pupil position of one or both of the near-eye display systems 402 and 404 to correspond to the actual position of one or both of the pupils 494. Such motors 438 can be controlled, for example, by user manipulation of one or more controls 437 on the housing 414 and / or by user manipulation of one or more associated separate I / O controllers (not shown). In other embodiments, HMD device 405 can control the alignment and / or other positioning of optical lens systems 410 and 412 and / or display panels 406 and 408 without requiring such a motor 438, such as by using an adjustable positioning mechanism (e.g., a screw, a slider, a ratchet, etc.) that is manually changed by a user using controller 437. Figure 4 Motor 438 is shown for only one of the near-eye display systems, but in some embodiments, each near-eye display system can have its own one or more motors, and in some embodiments, one or more motors can be used to control (e.g., independently) each of multiple near-eye display systems.

[0048] While the described techniques may be used in some embodiments with a display system similar to that shown, other types of display systems may be used in other embodiments, including with a single optical lens and display device, or with a plurality of such optical lenses and display devices. Non-exclusive examples of other such devices include cameras, telescopes, microscopes, binoculars, spotting scopes, mapping scopes, and the like. In addition, the described techniques may be used with various display panels or other display devices that emit light to form an image, with one or more users viewing the image through one or more optical lenses. In other embodiments, a user may view one or more images through one or more optical lenses that are produced in a manner other than through a display panel, such as on a surface that partially or fully reflects light from another light source.

[0049] Figure 5 An example of using multiple eye tracking components, each including a light source and a light detector, to determine the user's gaze position in a particular manner according to the described techniques is shown. Specifically, Figure 5Information 500 is included to illustrate the operation of an example display panel 510 and associated optical lens 508 in providing image information to a user's eye 504, e.g., focusing the information on the eye's pupil 506. In the illustrated embodiment, four eye tracking components 511a-511d (collectively, 511) of the eye tracking subsystem are mounted near an edge of the optical lens 508, and each eye tracking component 511a-511d is generally oriented toward the pupil 506 of the eye 504 for emitting light toward the eye 504 and capturing light reflected from part or all of the pupil 506 or the surrounding iris 502.

[0050] In the example shown, each eye tracking assembly 511 includes a light source 512, a light detector 514, and a polarizer 516, the polarizer 516 being positioned and configured to provide scattered light (diffuse reflection) to the light detector while substantially inhibiting light reflected by specular reflection from reaching the detector 514. In this example, the eye tracking assembly 511 is positioned near the top of the optical lens 508 along the central vertical axis, near the bottom of the optical lens along the central vertical axis, near the left side of the optical lens along the central horizontal axis, and near the right side of the display panel along the central horizontal axis. In other embodiments, the eye tracking assembly 511 can be located in other locations, and fewer or more eye tracking assemblies can be used.

[0051] The polarizer 516 of each eye assembly 511 can include one or more polarizers positioned in front of the light source 512 and / or the light detector 514 to reduce or eliminate specularly reflected light from reaching the light detector 514. In one example, the polarizer 516 of a particular eye tracking assembly 511 can include two crossed linear polarizers, a first linear polarizer positioned in front of the light source 512 and a second linear polarizer oriented 90 degrees relative to the first polarizer and positioned in front of the light detector 514 to block specularly reflected light. In other embodiments, one or more circular or linear polarizers can be used to prevent light reflected via specular reflection from reaching the light detector 514, such that the light detector receives substantially all light reflected via diffuse reflection.

[0052] It should also be understood that the light sources and light detectors are shown for illustrative purposes only, and other embodiments may include more or fewer light sources or detectors, and the light sources or detectors may be located in other locations. In addition, although not shown here, in some embodiments, additional hardware components may be used to help obtain data from one or more light detectors. For example, an HMD device or other display device may include various light sources (e.g., infrared, visible light, etc.) at different locations to illuminate the iris and pupil to reflect back to one or more light detectors, such as a light source mounted on or near the display panel 510, or alternatively, elsewhere (e.g., between the optical lens 508 and the eye 504, such as on an inner surface of the HMD device including the display panel 510 and the optical lens 508, not shown). In some such embodiments, light from such an illumination source may be further reflected from the display panel before passing through the optical lens 508 to illuminate the iris and pupil.

[0053] Figure 6 is a schematic diagram of an environment 600 according to one non-limiting illustrated embodiment in which machine learning techniques can be used to implement a device's eye-tracking subsystem, such as the eye-tracking subsystem discussed herein. Environment 600 includes a model training portion 601 and an inference portion 603. In the training portion 601, training data 602 is fed into a machine learning algorithm 604 to generate a trained machine learning model 606. The training data can include, for example, labeled data from light detectors specifying gaze locations. As a non-limiting example, in an embodiment including four light detectors directed toward a user's eyes, each training sample can include output from each of the four light detectors and a known or inferred gaze direction. In at least some embodiments, a user's gaze direction can be learned or inferred by directing the user to gaze at a specific user interface element (e.g., a word, a dot, an "X," another shape, or an object) on the display of an HMD device, which can be static or movable across the display. The training data can also include samples in which one or more light sources or light detectors are obscured, for example, by a blink, eyelashes, glasses, a hat, or other obstructions. For such training samples, the label can be “unknown” or “occluded” instead of a specified gaze direction.

[0054] Training data 602 can be obtained from multiple users and / or from a single user of the HMD system. Training data 602 can be obtained in a controlled environment and / or during actual user usage ("field training"). Furthermore, in at least some embodiments, model 606 can be updated or calibrated from time to time (e.g., periodically, continuously, after certain events) to provide accurate gaze direction predictions.

[0055] In the inference portion 603, runtime data 608 is provided as input to a trained machine learning model 606, which generates a gaze direction prediction 610. Continuing with the example above, the output data of the light detector can be provided as input to the trained machine learning model 606, which can process the data to predict the gaze location. The gaze direction prediction 610 can then be provided to one or more components associated with the HMD device, such as one or more VR or AR applications executing on the HMD device, one or more display or rendering modules, one or more mechanical controls, one or more position tracking subsystems, and the like.

[0056] The machine learning techniques used to implement the features discussed herein can include any type of appropriate structure or technology. As non-limiting examples, the machine learning model 606 can include one or more of a decision tree, a statistical hierarchical model, a support vector machine, an artificial neural network (ANN) such as a convolutional neural network (CNN) or a recurrent neural network (RNN) (e.g., a long short-term memory (LSTM) network), a mixture density network (MDN), a hidden Markov model, or other models that can be used. In at least some embodiments, such as embodiments utilizing RNNs, the machine learning model 606 can utilize past input (memory, feedback) information to predict gaze direction. Such embodiments can advantageously utilize sequence data to determine motion information or previous gaze direction predictions, which can provide more accurate real-time gaze direction predictions.

[0057] Figure 7 is a flow chart of an exemplary embodiment of a method 700 of operating an eye tracking subsystem of an HMD device. The method 700 may be performed by, for example Figure 1 The method 700 may be performed by the eye tracking subsystem 135 or other systems discussed elsewhere herein. Although the illustrated embodiment 700 discusses the execution of operations to determine gaze direction for a single eye, it should be understood that the operations of the method 700 may be applied to both eyes of the user simultaneously to track gaze direction substantially in real time. It should also be understood that the illustrated embodiment of the method 700 may be implemented in software and / or hardware, as appropriate, and may be executed by, for example, a computing system associated with an HMD device.

[0058] As described above, an HMD device may include a support structure wearable on a user's head, an eye-tracking subsystem connected to the support structure, at least one processor, and a memory storing a set of instructions or data. The eye-tracking subsystem may include multiple eye-tracking components, each of which includes a light source, a light detector, and a polarizer.

[0059] A polarizer can be positioned proximate to at least one of the light source and the light detector and can be configured to prevent light reflected via specular reflection from being received by the light detector. In at least some embodiments, the polarizer can include two crossed linear polarizers. In at least some embodiments, for each of the plurality of eye tracking assemblies, the polarizer includes a first polarizer positioned in a light emission path of the light source and a second polarizer positioned in a light detection path of the light detector. More generally, the polarizer can include at least one of a circular polarizer or a linear polarizer.

[0060] Each light source can be directed to a target location of the user's pupil, which can allow the use of lower power light sources because the energy is focused on the target location. In addition, various optical devices (e.g., lenses) or light source types (e.g., IR lasers) can also be used to focus the light on the target area. In at least some embodiments, the light source can include a light emitting diode that emits light having a wavelength of, for example, between 780nm and 1000nm. The light source can be positioned away from the optical axis of the user's eye to provide dark field illumination of the pupil. In at least some embodiments, the light detector can include a silicon photodiode that provides an output signal that depends on the power of the incident light.

[0061] The illustrated embodiment of method 700 begins at 702, where at least one processor of an HMD device may selectively illuminate light sources of multiple eye-tracking components. The at least one processor may illuminate the light sources simultaneously, sequentially, in another pattern, or any combination thereof. At 704, the at least one processor may receive light detection information captured by light detectors of the multiple eye-tracking components. For example, the at least one processor may store output data received from the multiple light detectors over one or more time periods.

[0062] At 706, at least one processor may provide the received light detection information as input to a trained machine learning model, such as that described above with respect to Figure 6 Model 606 discussed above. As described above, the machine learning model can include an RNN, an MDN, or any other type of machine learning model suitable for providing accurate gaze direction prediction based on light data input received from multiple light detectors. As discussed elsewhere herein, in other embodiments, a prediction model or function other than a machine learning model can be used, such as a 1D or 2D polynomial, one or more lookup tables, etc.

[0063] At least one processor may receive, from the machine learning model, a determined gaze direction of the user's eyes in response to providing the light detection information, at 708. The determined gaze direction may be provided in any suitable format.

[0064] To calibrate or update the machine learning model, at least one processor may cause at least one display of the HMD to present a user interface element, selectively cause light sources of multiple eye tracking components to emit light, and receive light detection information captured by light detectors of the multiple eye tracking components. The light detection information and corresponding known or inferred gaze direction information may be used to update the machine learning model. The model may be updated multiple times as needed to provide accurate gaze direction predictions. The user interface elements may include static user interface elements or moving user interface elements.

[0065] At 710, at least one processor provides a determined gaze direction to a component associated with the HMD system for use. For example, the determined gaze direction can be provided to an image rendering subsystem of the HMD system to provide viewpoint rendering based on the determined gaze direction, as described above. As another example, an eye tracking subsystem can determine that the user's eyes are rapidly scanning and can dynamically change image presentation to utilize rapid scanning masking or rapid scanning suppression. For example, one or more characteristics of image rendering can be changed during a rapid scan, such as the resolution of all or part of the image, the spatial frequency of the image, the frame rate, or any other characteristics that can allow for reduced bandwidth, reduced computing requirements, or other technical benefits.

[0066] As another example, in at least some embodiments, an HMD device may include an interpupillary distance (IPD) adjustment component for automatically adjusting one or more components of the HMD to account for a variable IPD. In this example, the IPD adjustment component may receive a gaze direction and may align at least one component of the HMD system for the user based at least in part on the determined gaze direction. For example, when the user is viewing a nearby object, the IPD may be relatively short, and the IPD adjustment may align one or more components of the HMD device accordingly. As another non-limiting example, the HMD device may automatically adjust the focus of a lens based on the determined gaze direction.

[0067] While the above examples utilize machine learning techniques to determine gaze direction from light detection information, it should be understood that features of the present disclosure are not limited to the use of machine learning techniques. Generally, any type of prediction model or function can be used. For example, in at least some embodiments, the system may not derive gaze direction directly from light detection information, but rather work in reverse, predicting light detection information given an input gaze direction. This approach can find a prediction function or model that performs this prediction, mapping gaze direction to predicted light readings for that direction. This function can be user-specific, so the system can implement a calibration process to find or customize it. Once the prediction function has been determined or generated, it can be inverted (e.g., using a numerical solver) to produce real-time predictions during use. Specifically, given a sample of light detection information from a real sensor, the solver can be used to find a gaze direction that minimizes the error between the real reading and the predicted reading from the generated prediction function. In such embodiments, the output is the solved gaze direction plus a residual error, which can be used to judge the quality of the solution. In at least some embodiments, additional corrections can be applied to address various issues, such as the HMD system sliding across the user's face during operation. The prediction function can be any type of function. As an example, given a dataset of points captured from the user, a set of 2D polynomials can be used to map gaze angles to photodiode readings. In at least some other embodiments, lookup tables or other methods can also be used, including adapting the ML system to output predictions as described above.

[0068] It should be understood that in some embodiments, the functionality provided by the above routines can be provided in an alternative manner, for example, split between more routines, or merged into fewer routines. Similarly, in some embodiments, the illustrated routines can provide more or less functionality than described, for example when other illustrated routines lack or include such functionality, or when the number of functions provided changes. In addition, although various operations can be shown as being performed in a specific manner (for example, serially or in parallel) and / or in a specific order, it will be understood by those skilled in the art that in other embodiments, operations can be performed in other orders and in other ways. It will be similarly understood that the above data structures can be constructed in different ways, including for databases or user interface screens / pages or other types of data structures, for example, by dividing a single data structure into multiple data structures or by merging multiple data structures into a single data structure. Similarly, in some embodiments, the data structures shown can store more or less information than described, for example when other illustrated data structures lack or include such information, or when the amount or type of stored information changes.

[0069] In addition, the sizes and relative positions of the elements in the drawings are not necessarily drawn to scale, including the shapes of various elements and angles, some of which are enlarged and positioned to improve the readability of the drawings, and the specific shapes of at least some elements are selected to facilitate identification without necessarily conveying information about the actual shapes or proportions of those elements. In addition, some elements may be omitted for clarity and emphasis. In addition, reference numerals that are repeated in different figures may represent the same or similar elements.

[0070] As will be appreciated from the foregoing, although specific embodiments have been described herein for illustrative purposes, various modifications may be made without departing from the spirit and scope of the present invention. Furthermore, while certain aspects of the present invention may sometimes be presented in certain claim forms, or may sometimes not be embodied in any claim form, the inventors contemplate various aspects of the present invention in any applicable claim form. For example, while some aspects of the present invention may at certain times be embodied in computer-readable media, other aspects may also be embodied.

[0071] This application claims priority to U.S. Patent Application No. 16 / 773,840, filed January 27, 2020, which is hereby incorporated by reference in its entirety.

Claims

1. A head-mounted display (HMD) system, comprising: A support structure capable of being worn on a user's head; an eye tracking subsystem coupled to the support structure and comprising a plurality of eye tracking components, each eye tracking component comprising: a light source, operated to emit light; a light detector operative to detect light; and a polarizer positioned proximate to at least one of the light source and the light detector, the polarizer configured to prevent light reflected via specular reflection from being received by the light detector and to provide light reflected via diffuse reflection to the light detector, wherein the light reflected via specular reflection includes light reflected from a surface at an angle equal to an angle of incidence and the light reflected via diffuse reflection includes light scattered from a surface at a plurality of angles, and wherein each of the plurality of eye tracking assemblies is mounted proximate to an edge of an optical lens positioned facing a display panel, wherein for each of the plurality of eye tracking assemblies, the light source is positioned away from an optical axis of an eye of the user to provide darkfield illumination of a pupil of the eye; at least one processor; and A memory storing a set of instructions or data that, as a result of execution, causes the HMD system to: selectively causing the light sources of the plurality of eye tracking assemblies to emit light; receiving light detection information captured by the light detectors of the plurality of eye tracking assemblies, the light detection information comprising information associated with diffusely reflected light detected by the light detectors; providing the received light detection information as input to a prediction model; In response to providing the light detection information, receiving a determined gaze direction of the user's eyes from the predictive model; and The determined gaze direction is provided to components associated with the HMD system for use thereby.

2. The HMD system according to claim 1, wherein: The light detection information includes a characteristic radiation pattern of each of the light sources after light from the light sources has reflected, scattered, or absorbed off the user's face or eyes.

3. The HMD system according to claim 1, wherein: Each of the light sources is directed toward an expected location of the user's pupil.

4. The HMD system according to claim 1, wherein: The light source includes a light emitting diode that emits light with a wavelength between 780 nm and 1000 nm.

5. The HMD system according to claim 1, wherein: The light detector includes a photodiode.

6. The HMD system according to claim 1, wherein: The eye tracking subsystem includes four eye tracking components positioned to determine a gaze direction of the user's left eye and four eye tracking components positioned to determine a gaze direction of the user's right eye.

7. The HMD system according to claim 1, wherein: The polarizer includes two crossed linear polarizers.

8. The HMD system according to claim 1, wherein: For each of the plurality of eye tracking assemblies, the polarizer includes a first polarizer located in a light emission path of the light source and a second polarizer located in a light detection path of the light detector.

9. The HMD system according to claim 1, wherein: The polarizer includes at least one of a circular polarizer or a linear polarizer.

10. The HMD system according to claim 1, wherein: The predictive model includes a machine learning model.

11. The HMD system according to claim 10, wherein: The machine learning model includes a mixture density network (MDN) model.

12. The HMD system according to claim 10, wherein: The machine learning model includes a recurrent neural network (RNN) model.

13. The HMD system according to claim 10, wherein: The machine learning model utilizes past input information or eye movement information to determine the gaze direction.

14. The HMD system according to claim 10, wherein: The machine learning model is a model trained during field operation of multiple HMD systems.

15. The HMD system according to claim 1, wherein: The prediction model includes a polynomial or a lookup table.

16. The HMD system according to claim 1, wherein: The HMD system includes at least one display, and the at least one processor: causing the at least one display to present a user interface element; selectively illuminating the light sources of the plurality of eye tracking assemblies; receiving light detection information captured by the light detectors of the plurality of eye tracking assemblies; as well as The prediction model is updated based at least in part on the received light detection information and known or inferred gaze direction information associated with the received light detection information.

17. The HMD system according to claim 16, wherein: The user interface element includes a static user interface element or a mobile user interface element.

18. The HMD system according to claim 1, wherein: The HMD system includes at least one display, and the at least one processor: A rendering output of the at least one display is dynamically changed based at least in part on the determined gaze direction.

19. The HMD system according to claim 1, wherein: The HMD system includes an interpupillary distance (IPD) adjustment component, and the at least one processor: The IPD adjustment component is caused to align at least one component of the HMD system for the user based at least in part on the determined gaze direction.

20. A method of operating a head-mounted display (HMD) system, the HMD system comprising an eye tracking subsystem connected to a support structure, the support structure comprising a plurality of eye tracking components, each eye tracking component comprising a light source, a light detector, and a polarizer, the polarizer being configured to prevent light reflected via specular reflection from being received by the light detector and to provide light reflected via diffuse reflection to the light detector, wherein The light reflected via specular reflection includes light reflected from the surface at an angle equal to an incident angle, and the light reflected via diffuse reflection includes light scattered from the surface at a plurality of angles, and wherein each of the plurality of eye tracking assemblies is mounted near an edge of an optical lens positioned facing the display panel, wherein, for each of the plurality of eye tracking assemblies, the light source is positioned away from an optical axis of an eye of the user to provide darkfield illumination of a pupil of the eye, the method comprising: selectively causing the light sources of the plurality of eye tracking assemblies to emit light; preventing, by each of the respective polarizers of the plurality of eye tracking components, light reflected via specular reflection from being received by a plurality of light detectors while allowing light reflected via diffuse reflection to be received by the plurality of light detectors; receiving light detection information captured by a plurality of light detectors, the light detection information including information associated with diffusely reflected light detected by the light detectors; providing the received light detection information as input to a trained machine learning model; In response to providing the light detection information, receiving a determined gaze direction of a user's eyes from the machine learning model; and The determined gaze direction is provided to components associated with the HMD system for use thereby.

21. The method according to claim 20, wherein The light detection information includes a characteristic radiation pattern of each of the light sources after light from the light sources has reflected, scattered, or absorbed off the user's face or eyes.

22. The method according to claim 20, wherein Providing the received light detection information as input to the trained machine learning model includes providing the received light detection information as input to a mixture density network (MDN) model.

23. The method according to claim 20, wherein Providing the received light detection information as input to the trained machine learning model includes providing the received light detection information as input to a recurrent neural network (RNN) model.

24. The method according to claim 20, wherein Providing the received light detection information as input to a trained machine learning model includes providing the received light detection information to a machine learning model that utilizes past input information or eye movement information to determine the gaze direction.

25. The method of claim 20, further comprising training the machine learning model during field operation of a plurality of HMD systems.

26. The method according to claim 20, wherein The HMD system includes at least one display, and the method includes: causing the at least one display to present a user interface element; selectively illuminating the light sources of the plurality of eye tracking assemblies; receiving light detection information captured by a plurality of light detectors; and The machine learning model is updated based at least in part on the received light detection information and known or inferred gaze direction information associated with the received light detection information.

27. The method according to claim 26, wherein Causing the at least one display to present the user interface element includes causing the at least one display to present a static user interface element or a moving user interface element.

28. The method according to claim 20, wherein The HMD system includes at least one display, and the method includes: An output of the at least one display is dynamically changed based at least in part on the determined gaze direction.

29. The method of claim 20, further comprising: Mechanical alignment of at least one component of the HMD system is performed for the user based at least in part on the determined gaze direction.

30. A head-mounted display (HMD) system, comprising: A support structure capable of being worn on a user's head; an eye tracking subsystem coupled to the support structure and comprising a plurality of eye tracking components, each eye tracking component comprising: light-emitting diodes; photodiode; and a polarizer positioned proximate to at least one of the light emitting diode and the photodiode, the polarizer configured to prevent light reflected via specular reflection from being received by the photodiode and to provide light reflected via diffuse reflection to a light detector, wherein the light reflected via specular reflection includes light reflected from a surface at an angle equal to an angle of incidence, and the light reflected via diffuse reflection includes light scattered from a surface at a plurality of angles, and wherein each of the plurality of eye tracking assemblies is mounted proximate to an edge of an optical lens positioned facing a display panel, and wherein, for each of the plurality of eye tracking assemblies, a light source is positioned away from an optical axis of an eye of the user to provide darkfield illumination of a pupil of the eye; at least one processor; and A memory storing a set of instructions or data that, as a result of execution, causes the HMD system to: selectively causing the light emitting diodes of the plurality of eye tracking assemblies to illuminate; receiving light detection information captured by the photodiodes of the plurality of eye tracking assemblies, the light detection information comprising information associated with diffusely reflected light detected by the photodetectors; providing the received light detection information as input to a prediction model; In response to providing the light detection information, receiving a determined gaze direction of the user's eyes from the predictive model; and Operation of components associated with the HMD system is dynamically changed based at least in part on the determined gaze direction.

Citation Information

Patent Citations

  • Polarized gaze tracking

    CN106062776A

  • Gaze tracking device and a head mounted device embedding said gaze tracking device

    US20170004363A1