Optical tracking including a high-sensitivity angle detector based on an image
Gaze tracking in HMD devices addresses the challenges of bandwidth and computational resource management in VR and AR applications by focusing processing on the area of interest, resulting in improved performance and efficiency.
Patent Information
- Application Number
- JP2024570961
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-03
- Filing Date
- 2023-05-09
- Publication Date
- 2025-06-19
AI Technical Summary
Current head-mounted display (HMD) devices face challenges in efficiently managing media transmission bandwidth and computational resources, particularly in virtual reality (VR) and augmented reality (AR) applications, due to the need for high-resolution displays and real-time image processing.
The implementation of gaze tracking techniques in HMD devices, which utilize optical detectors and sensors to track the user's gaze, allows for foveated rendering and reduced image data transmission. This approach focuses computational resources on the area of interest, reducing bandwidth and processing demands for other areas of the display.
By implementing gaze tracking in HMD devices, the described techniques effectively reduce media transmission bandwidth, improve processing efficiency, and enhance the overall performance of VR and AR applications, while maintaining high resolution and detail in the area of interest.
Smart Images

Figure 2025518791000001_ABST
Abstract
Description
Technical Field
[0001] The following disclosure generally relates to techniques for tracking, and in at least some implementations, to techniques for gaze tracking in a head-mounted display device.
Background Art
[0002] A head-mounted display (HMD) device or system is an electronic device that is worn on a user's head and, when worn, fixes at least one electronic display within the visible field of at least one of the user's eyes regardless of the position or orientation of the user's head. An HMD device used to implement a virtual reality system typically completely surrounds the wearer's eyes, replacing the wearer's actual view (or actual reality) in front with a "virtual" reality, while an HMD device for an augmented reality system typically provides a translucent or transparent overlay of one or more screens in front of the wearer's eyes such that the actual view is augmented with additional information. In the case of an augmented reality system, the "display" component of the HMD device is either transparent or located at the periphery of the user's field of view so as not to completely prevent the user from being able to view their external environment.
Brief Description of the Drawings
[0003]
Figure 1
[0004]
Figure 2
[0005]
Figure 3
[0006]
Figure 4
[0007]
Figure 5
[0008]
Figure 6
[0009]
Figure 7
[0010]
Figure 8
[0011]
Figure 9
[0012]
Figure 10A
[0013]
Figure 10B
[0014] In the following description, specific specific details are set forth in order to achieve a thorough understanding of the various disclosed implementations. However, one of ordinary skill in the art will recognize that the implementation may be practiced without one or more of these specific details, or using other methods, components, materials, etc. In other instances, well-known structures related to computer systems, server computers, and / or communication networks are not illustrated or described in detail in order to avoid unnecessarily obscuring the description of the implementation.
[0015] Unless the context requires a different interpretation, throughout this specification and the following claims, the word "comprising" is synonymous with "including" and is inclusive or open-ended (i.e., it does not exclude additional, unrecited elements or method acts).
[0016] References throughout this specification to "one implementation" or "an implementation" mean that a particular feature, structure, or characteristic described in connection with that implementation is included in at least one implementation. Thus, the appearances of the phrases "in one implementation" or "in an implementation" in various places throughout this specification are not necessarily all referring to the same implementation. Further, the particular features, structures, or characteristics may be combined in any suitable manner in one or more implementations.
[0017] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Note also that the term "or" is generally used in the sense of "and / or" unless the context clearly dictates otherwise.
[0018] The headings and abstracts presented in this specification are for convenience only and do not interpret the scope or meaning of the implementation forms.
[0019] Eye tracking is a process by which the position, orientation, or movement of an eye can be measured, detected, sensed, determined, or monitored (collectively referred to as "measured"). In many applications, this is done for the purpose of determining the direction of a user's gaze. The position, orientation, or movement of an eye may be measured in a variety of different ways, and the least invasive of those methods may utilize one or more optical detectors or sensors to optically track the eye. Some techniques may involve using infrared light to illuminate or project light onto the entire eye at once and measuring the reflection using at least one optical sensor that is adjusted to be sensitive to that infrared light. Information regarding how the infrared light is reflected from the eye is analyzed to determine the position, orientation, and / or movement of one or more eye features such as the cornea, pupil, iris, or retinal blood vessels.
[0020] The eye tracking function is very advantageous in the application of wearable head-mounted display systems. Some examples of the usefulness of eye tracking in a head-mounted display system include affecting where content is displayed within the user's field of view, modifying the display of content outside the user's field of view (e.g., foveated rendering) to conserve power, bandwidth, or computational resources, affecting the content presented to the user, determining where the user is looking or gazing, determining whether the user is looking at the content displayed on the display, providing a way for the user to control or interact with the presented content, and other applications.
[0021] The present disclosure generally relates to techniques for object tracking, such as gaze tracking or tracking of other objects. Such techniques can be used, for example, in head-mounted display ("HMD") devices used in VR or AR applications. Some or all of the techniques described herein may be performed via automated operations of embodiments of a gaze tracking subsystem, such as being implemented by one or more configured hardware processors or other configured hardware circuits. One or more hardware processors or other configured hardware circuits of such a system or device may include, for example, one or more GPUs ("graphics processing units") and / or CPUs ("central processing units") and / or other microcontrollers ("MCUs") and / or other integrated circuits. For example, as further discussed below, a hardware processor may be part of an HMD device or other device that incorporates one or more display panels on which image data is to be displayed, or part of a computing system that generates or otherwise prepares image data to be sent to a display panel for display. More generally, such a hardware processor or other configured hardware circuit may include, but is not limited to, one or more application specific integrated circuits (ASICs), standard integrated circuits, controllers (e.g., by executing appropriate instructions, and including microcontrollers and / or embedded controllers), field programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), digital signal processors (DSPs), programmable logic controllers (PLCs), etc. Further details are included elsewhere in this specification, including those discussed below with respect to FIG. 1.
[0022] Technical benefits in at least some embodiments of the described techniques include addressing and reducing increased media transmission bandwidth for image encoding by reducing image data size, improving the speed of controlling display panel pixels (e.g., based on at least a partially corresponding reduced image data size), improving foveal image systems, and other techniques that reflect subsets of the display panel and / or images of particular interest. Foveal image encoding systems utilize specific aspects of the human visual system (which can provide detailed information only at the point of fixation and its surroundings), but often use special computational processing to avoid visual artifacts (e.g., artifacts related to motion and contrast in video and image data) to which the peripheral visual field is very susceptible. In the case of certain VR and AR displays, certain display devices include two pixel arrays that are separately addressable, and each pixel array includes two separate display panels (i.e., one display panel for each eye) with appropriate resolutions, so both the bandwidth usage and computing usage for processing high-resolution media are amplified. Therefore, the described techniques can be used, for example, to reduce the transmission bandwidth for local and / or remote displays of video frames or other images, while at the same time retaining the resolution and detail in the "area of interest" of the viewer within the image, and at the same time minimizing the computing usage for processing such image data. Further, by using lenses in head-mounted display devices and with other displays, a higher focus or resolution can be provided for subsets of the display panel, and thus using such techniques to display lower-resolution information in other parts of the display panel can further provide advantages when using such techniques in such embodiments.
[0023] For illustrative purposes, several embodiments are described below, in which certain types of information are obtained and used in a particular type of method for a particular type of structure and by using a particular type of device. However, it will be understood that the techniques so described may be used in other ways in other embodiments and, therefore, the present disclosure is not limited to the exemplary details provided. As one non-exclusive example, the various embodiments discussed herein include the use of images that are video frames; however, while many of the examples described herein refer to "video frames" for convenience, the techniques described with reference to such examples may be utilized for one or more images of various types, including, by way of non-exclusive example, a plurality of sequential video frames (e.g., at 30, 60, 90, 180 or some other number of frames per second), other video content, photographs, computer-generated graphical content, other visual media products, or some combination thereof. Additionally, various details are provided in the drawings and text for illustrative purposes and are not intended to limit the scope of the present disclosure. Further, as used herein, a "pixel" refers to the smallest addressable image element of a display device that can be operated to provide all possible values of color to that display device. In many cases, a pixel includes individual sub-elements (in some cases separate "sub-pixels") that separately generate the red, green, and blue light perceived by a human viewer, and separate color channels are used to encode the pixel values of the different color sub-pixels. As used herein, a pixel "value" refers to a data value corresponding to the respective level of stimulation for one or more of those respective RGB elements of a single pixel.
[0024] FIG. 1 is a schematic diagram of a network connection environment 100 including a local media rendering (LMR) system 110 (e.g., a gaming system), and the LMR system 110 includes a local computing system 120 and a display device 180 (e.g., an HMD device having two display panels) suitable for performing at least some of the techniques described herein. In the illustrated embodiment of FIG. 1, the local computing system 120 is communicatively connected to the display device 180 via a transmission link 115 (the display device may be wired or tethered, such as via one or more cables (cable 220) as illustrated in FIG. 2, or alternatively wirelessly connected). In other embodiments, the local computing system 120 may provide encoded display image data, whether in addition to or instead of the HMD device 180, to a panel display device (e.g., a TV, console, or monitor), each of which includes one or more addressable pixel arrays, via a wired or wireless link. In various embodiments, the local computing system 120 may include a general-purpose computing system, a gaming console, a video stream processing device, a mobile computing device (e.g., a mobile phone, PDA®, or other mobile device), a VR processing device or an AR processing device, or other computing systems.
[0025] In the illustrated embodiment, the local computing system 120 includes components having one or more hardware processors (e.g., a central processing unit, or “CPU”) 125, a memory 130, various I / O (“input / output”) hardware components 127 (e.g., a keyboard, a mouse, one or more gaming controllers, speakers, microphones, IR transmitters and / or receivers, etc.), a video subsystem 140 including one or more dedicated hardware processors (e.g., a graphics processing unit, or “GPU”) 144 and video memory (VRAM) 148, a computer-readable storage 150, and a network connection 160. Also, in the illustrated embodiment, an embodiment of the gaze tracking subsystem 135 is executed in the memory 130 using the CPU 125 and / or the GPU 144 to perform automated operations implementing at least some of the described techniques for performing at least some of the described techniques, and the memory 130 may optionally further execute one or more other programs 133 (e.g., a game program that generates a displayed video or other image). As part of the automated operations implementing at least some of the techniques described herein, the gaze tracking subsystem 135 and / or the program 133 executed in the memory 130 may store or retrieve various types of data including those within the data structure of an exemplary database of the storage 150. In this example, the data used may include various types of image data information within a database (“DB”) 154, various types of application data within the DB 152, various types of configuration data within the DB 157, and may also include additional information such as system data or other information.
[0026] In the illustrated embodiment, the LMR system 110 is communicatively connected via one or more computer networks 101 and network links 102 to an exemplary network-accessible media content provider 190 that can further provide content to the LMR system 110 for display, whether in addition to or instead of the image generation program 133. The media content provider 190 can include one or more computing systems (not shown) having components similar to those of the local computing system 120, including one or more hardware processors, I / O components, local storage devices, and memory. For simplicity, some details of the network-accessible media content provider are not illustrated.
[0027] In the embodiment shown in FIG. 1, the display device 180 is shown as being distinct and separate from the local computing system 120. However, in certain embodiments, some or all of the components of the local media rendering system 110 may be integrated or housed within a single device, such as a mobile gaming device, a portable VR entertainment system, an HMD device, etc. In such embodiments, the transmission link 115 may include, for example, one or more system bus or video bus architectures.
[0028] As an example involving operations locally executed by the local media rendering system 120, assume that the local computing system is a gaming computing system, whereby application data 152 includes one or more gaming applications executed via the CPU 125 using the memory 130, and that various video frame display data is generated and / or processed by an image generation program 133 in combination with the GPU 144 of the video subsystem 140 and the like. To provide a high-quality gaming experience, a large amount of video frame data (corresponding to a high image resolution for each video frame and a high "frame rate" of about 60 to 180 per second for such video frames) is generated by the local computing system 120 and provided to the display device 180 via a wired or wireless transmission link 115.
[0029] The computing system 120 and the display device 180 are merely examples and are not intended to limit the scope of the present disclosure. It will also be understood that the computing system 120 may instead include a plurality of interacting computing systems or devices and may be connected to other devices not illustrated, such as via one or more networks such as the Internet, via the web, or via a private network (such as a mobile communication network, etc.). More generally, a computing system or other computing node may include any combination of hardware or software that can interact to perform the functions of the type described, including, but not limited to, a desktop or other computer, a gaming machine, a database server, a network storage device and other network devices, a PDA (registered trademark), a cellular phone, a wireless phone, a pager, an electronic organizer, an Internet appliance, a TV-based system (such as using a set-top box and / or a personal / digital video recorder), and various other consumer products including appropriate communication capabilities. The display device 180 may similarly include one or more devices having one or more display panels of various types and configurations and may optionally include various other hardware and / or software components.
[0030] In addition, the functionality provided by the gaze tracking subsystem 135 may in some embodiments be distributed across one or more components, and in some embodiments, some of the functionality of the gaze tracking subsystem 135 may not be provided, and / or other additional functionality may be available. Although various items are shown as being stored in memory or on storage during use, it will also be understood that these items or portions thereof may be transferred between memory and other storage devices for memory management or data integrity purposes. Accordingly, in some embodiments, some or all of the techniques described may be implemented by one or more software programs (e.g., by execution of software instructions of one or more software programs and / or by storage of such software instructions and / or data structures) by and / or data structures, e.g., by the gaze tracking subsystem 135 or components thereof, and / or may be implemented by hardware including one or more processors or other configured hardware circuits or memory or storage. Some or all of the components, systems, and data structures may also be stored on a non-transitory computer-readable storage medium, e.g., a hard disk or flash drive or other non-volatile storage device, volatile memory or non-volatile memory (e.g., RAM), network storage device, or portable media article (e.g., DVD disk, CD disk, optical disk, etc.) readable by an appropriate drive or via an appropriate connection, e.g., as software instructions or structured data. The systems, components, and data structures may also, in some embodiments, be transmitted in various computer-readable transmission media including wireless-based media and wired / cable-based media as generated data signals (e.g., as part of a carrier wave or other analog or digital propagated signal) and may take various forms (e.g., as part of one or multiplexed analog signals or as multiple individual digital packets or frames). Such computer program products may take other forms in other embodiments.Accordingly, the present invention may be practiced using other computer system configurations.
[0031] FIG. 2 shows an exemplary environment 200 that is used with an exemplary HMD device 202 coupled to a video rendering computing system 204 via a connected connection 220 (or a wireless connection in other embodiments) to provide a virtual reality display to a human user 206 for at least some of the techniques being described. The user wears the HMD device 202 and receives display information of a simulated environment different from the actual physical environment from the computing system 204 via the HMD device, and the computing system serves as an image rendering system that supplies an image of the simulated environment (e.g., an image generated by a game program and / or other software program running on the computing system) to the HMD device for display to the user. The user can further move around within a tracked volume 201 of the actual physical environment 200 in this example and may further have one or more I / O ("input / output") devices that further enable the user to interact with the simulated environment. This includes, in this example, handheld controllers 208 and 210.
[0032] In the illustrated example, environment 200 may include one or more base stations 214 (two are shown, labeled base stations 214a and 214b) that may facilitate tracking of HMD device 202 or controllers 208 and 210. As the user moves locations or changes the orientation of HMD device 202, the position of the HMD device is tracked, such as to enable the corresponding portion of the simulated environment to be presented to the user wearing the HMD device, and controllers 208 and 210 may further utilize similar techniques for use in tracking the position of the controllers (and optionally, for using information useful in determining or verifying the position of the HMD device). After the tracked position of HMD device 202 is known, corresponding information is transmitted to computing system 204 via tethering 220 or wirelessly. This uses the tracked position information to generate and present to the user one or more next images of the simulated environment.
[0033] Without limitation, there are a number of different methods of position tracking that may be used in various implementations of the present disclosure, including acoustic tracking, inertial tracking, magnetic tracking, optical tracking, combinations thereof, and the like.
[0034] In at least some implementations, the HMD device 202 may include one or more light receivers or sensors that can be used to implement the tracking features or other aspects of the present disclosure. For example, each base station 214 may sweep an optical signal across the entire tracking volume 201. Depending on the requirements of each particular implementation, each base station 214 may generate more than one optical signal. For example, while a single base station 214 is typically sufficient for 6 degrees of freedom tracking, multiple base stations (e.g., base stations 214a, 214b) may be required or desirable in some embodiments to provide robust room-scale tracking of the HMD device and peripheral devices. In this example, the light receivers are incorporated into the HMD device 202 and / or other objects to be tracked (such as controllers 208 and 210). In at least some implementations, on each object to be tracked, the light receivers may be combined with accelerometers and gyroscope inertial measurement units ("IMUs") to support low-latency sensor fusion.
[0035] In at least some implementations, each base station 214 includes two rotors that scan a linear beam across the tracking volume 201 on axes that are orthogonal to each other. At the start of each sweep cycle, the base station 214 may emit an omnidirectional optical pulse (referred to as a "synchronization signal") that is visible to all sensors on the object to be tracked. Thus, each sensor calculates its unique angular position in the sweep amount by measuring the period between the synchronization signal and the beam signal. The distance and orientation of the sensors may be derived using multiple sensors fixed to a single rigid body.
[0036] One or more sensors positioned on a tracking object (e.g., HMD device 202, controllers 208 and 210) may include optoelectronic devices capable of detecting modulated light from a rotor. In the case of visible light or near-infrared (NIR) light, silicon photodiodes and suitable amplifier / detector circuits may be used. Since environment 200 may include static and time-varying signals (optical noise) having a wavelength similar to that of the base station 214 signal, in at least some implementations, the base station light may be modulated in a manner that facilitates discrimination from any interfering signals and / or filters the sensor from any emission wavelengths other than those of the base station signal.
[0037] Inside-out tracking is also a type of position tracking that can be used to track the position of HMD device 202 and / or other objects (e.g., controllers 208 and 210, tablet computer, smartphone). Inside-out tracking differs from outside-in tracking by the location of the camera or other sensors used to determine the position of the HMD. In inside-out tracking, the camera or sensor is placed on the HMD or on the object being tracked, whereas in outside-in tracking, the camera or sensor is placed at a fixed position within the environment.
[0038] HMDs that utilize inside-out tracking use one or more cameras to "look out" to determine how their position changes relative to the environment. When the HMD moves, the sensors re-adjust their locations within the room, and the virtual environment responds in real time accordingly. This type of position tracking can be achieved with or without markers placed in the environment. The cameras placed on the HMD observe the characteristics of the surrounding environment. When using markers, the markers are designed to be easily detected by the tracking system and are placed in specific areas. In "markerless" inside-out tracking, the HMD system uses prominent features (e.g., natural features) that originally exist in the environment to determine position and orientation. The algorithms of the HMD system identify specific images or shapes and use them to calculate the position of the device in space. Data from accelerometers and gyroscopes may also be used to improve the accuracy of position tracking.
[0039] FIG. 3 shows information 300 that is a front view of an exemplary HMD device 344 when worn on the head of user 342. The HMD device 344 includes a front structure 343 that supports a front or forward camera 346 and a plurality of sensors 348a - 348d (collectively 348) of one or more types. As one example, some or all of the sensors 348, such as optical sensors that detect and use optical information emitted from one or more external devices (not shown, e.g., base station 214 of FIG. 2), can help determine the location and orientation of the device 344 within the space. As shown, the forward camera 346 and sensors 348 are directed forward toward the actual scene or environment (not shown) in which user 342 operates the HMD device 344. The actual physical environment can include, for example, one or more objects (e.g., walls, ceilings, furniture, stairs, vehicles, trees, tracking markers, or any other type of object). A specific number of sensors 348 may be fewer or more than the number of sensors shown. The HMD device 344 may further include an IMU (inertial measurement unit) electronic device 347 that measures and reports (e.g., using a combination of an accelerometer and a gyroscope, and optionally a magnetometer) certain forces, angular velocities, and / or magnetic fields surrounding the HMD device 344, which are not attached to the front structure (e.g., are inside the HMD device). The HMD device may further include additional components not shown, including one or more display panels and optical lens systems that face the user's both eyes (not shown), and optionally have one or more built-in motors for changing the alignment or other positioning of one or both of the optical lens system and / or display panel within the HMD device. This will be described in more detail below with respect to FIG. 4.
[0040] The illustrated example of the HMD device 344 is supported on the head of the user 342, at least in part, based on one or more straps 345 that are attached to the housing of the HMD device 344 and extend wholly or partially around the user's head. Although not illustrated here, the HMD device 344 may further include one or more external motors attached to one or more of the straps 345, etc., and adjusting such a strap using such a motor to correct the alignment or other positioning of the HMD device on the user's head may be included in the auto-correction operation. Whether in addition to or instead of the illustrated straps, the HMD device may include other support structures (e.g., nose piece, chin strap, etc.) not shown here, and it will be understood that some embodiments may include motors attached to one or more such other support structures that similarly adjust their shape and / or location to correct the alignment or other positioning of the HMD device on the user's head. Other display devices not fixed to the user's head may similarly be attached to or be part of one or more structures that affect the positioning of the display device, and in at least some embodiments, may include motors or other mechanical actuators to similarly modify their shape and / or location to correct the alignment or other positioning of the display device with respect to one or more pupils of one or more users of the display device.
[0041] Figure 4 shows a simplified top view 400 of an HMD device 405 with a pair of near-to-eye display systems 402 and 404. The HMD device 405 may be the same as or similar to the HMD devices illustrated in FIGS. 1 - 3, or it may be a different HMD device, and the HMD devices described herein may further be used in the examples further described below. The near-to-eye display systems 402 and 404 of FIG. 4 each include a display panel 406 and 408 (e.g., an OLED microdisplay), and each of their respective optical lens systems 410 and 412 having one or more optical lenses. The display systems 402 and 404 may be installed in a housing (or frame or support structure) 414 or positioned therein in another way. The housing includes a front portion 416 (e.g., the same as or similar to the front surface 343 of FIG. 3), a left temple 418, a right temple 420, and an inner surface 421 that contacts or is proximate to the face of the user 424 who is the wearer when the HMD device is worn. The two display systems 402 and 404 that can be worn on the head 422 of the user 424 who is the wearer may be fixed to the housing 414 in a glasses configuration. The left temple 418 and the right temple 420 may each be placed over the user's ears 426 and 428, and the nose rest 492 may be placed over the user's nose 430. A strap (not shown) or other structure may be used in some embodiments, such as the embodiments shown in FIGS. 2 and 3, to secure the HMD device to the user's head, but in the example of FIG. 4, the HMD device 405 may be supported, in part or in whole, on the user's head by the nose display and / or the right and left ear-hanging temples. The housing 414 may be made in a shape and size such that each of the two optical lens systems 410 and 412 is disposed in front of one of the user's eyes 432 and 434, whereby the target location of each respective pupil 494 will be centered both vertically and horizontally in front of each respective optical lens system and / or display panel.The housing 414 is shown in a simplified manner similar to glasses for illustrative purposes, but it should be understood that in practice, more sophisticated structures (e.g., goggles, integrated headbands, helmets, straps, etc.) can be used to support the display systems 402 and 404 and position them on the head 422 of the user 424.
[0042] The HMD device 405 of FIG. 4, and other HMD devices discussed herein, can present a virtual reality display to a user via corresponding video presented at a display rate such as 30 or 60 or 90 frames (or images) per second, etc. On the other hand, other embodiments of a similar system may present an augmented reality display to the user. The displays 406 and 408 of FIG. 4 may each generate light that is focused onto the eyes 432 and 434 of the user 424 by the respective optical lens systems 410 and 412 after passing through those optical lens systems. The aperture of the pupil 494 of each eye allows light to enter the eye through it. Usually, the size of the pupil ranges from 2 millimeters (mm) in a very bright state to 8 mm in a dark state, and the size of the larger iris including the pupil can be about 12 mm. The pupil (and the surrounding iris) can also typically move a few millimeters horizontally and / or vertically within the visible portion of the eye with the eyelids open, which, when the eyeball rotates about its center (resulting in a three-dimensional volume in which the pupil can move), will also move the pupil to different depths with respect to different horizontal and vertical positions from the optical lens of the display or other physical elements. The light entering the user's pupil is seen by the user 424 as an image or video. In some implementations, the distance between each of the optical lens systems 410 and 412 and the user's eyes 432 and 434 can be relatively short (e.g., less than 30 mm, less than 20 mm), whereby the weight of the optical lens system and the display system is relatively close to the user's face, which is advantageous in that the HMD device will feel lighter to the user and can also provide a larger field of view to the user. Although not illustrated here, some embodiments of such HMD devices may include various additional built-in or external sensors.
[0043] In the illustrated embodiment, the HMD device 405 of FIG. 4 further includes a hardware sensor and additional components such as one or more accelerometers and / or gyroscopes 490 (e.g., as part of one or more IMU units). As will be described in more detail elsewhere herein, values from the accelerometers and / or gyroscopes may be used to locally determine the orientation of the HMD device. Further, the HMD device 405 may include one or more front cameras, e.g., camera 485 on the outer surface of the front portion 416, and that information may be used as part of the operation of the HMD device, such as to provide an AR function or a positioning function. Further, the HMD device 405 may further include other components 475 (e.g., electronic circuitry for controlling the display of images to the display panels 406 and 408, built-in storage, one or more batteries, a position tracking device for communicating with an external base station, etc.), which will be described in more detail elsewhere herein. Other embodiments may not include one or more of the components 475, 485, and / or 490. Although not illustrated here, some embodiments of such HMD devices may include various additional built-in and / or external sensors for tracking various other types of movement and position of the user's body, eyes, controller, etc.
[0044] In the illustrated embodiment, the HMD device 405 of FIG. 4 further comprises hardware sensors and additional components that can be used by an embodiment disclosed as part of a technique for determining a user's pupil or gaze direction, and the user's pupil or gaze direction may be provided to one or more components associated with the HMD device for use by the HMD system, as discussed elsewhere in this specification. The hardware sensors in this example include one or more gaze tracking assemblies 472 of a gaze tracking subsystem that are installed on or near the display panels 406 and 408, and / or on the inner surface 421 near the optical lens systems 410 and 412, for use in obtaining information about the actual position of the user's pupil 494, for example, separately for each pupil in this example.
[0045] Each of the gaze tracking assemblies 472 may include one or more light detectors (e.g., silicon photodiodes, quadrant photodetectors having four active regions), and optionally one or more light sources (e.g., IR LEDs). Further, for clarity, FIG. 4 shows only a total of four gaze tracking assemblies 472, but it should be understood that in practice, a different number (e.g., one, three, six) of gaze tracking assemblies may be provided. In some embodiments, a total of eight gaze tracking assemblies 472 are provided, i.e., four gaze tracking assemblies for each of the user's 424 eyes. Further, in at least some implementations, each gaze tracking assembly 472 includes a light source directed at one of the user's 424 eyes 432 and 434, and a light detector positioned to receive light reflected by each of the user's eyes.
[0046] As discussed in more detail elsewhere in this specification, the information from the gaze tracking assembly 472 may be used to determine and track the user's gaze direction during use of the HMD device 405. Further, in at least some embodiments, the HMD device 405 may include one or more built-in motors 438 (or other movement mechanisms) that can be used to move (e.g., in the vertical, horizontal left-right, and / or horizontal front-back directions) the alignment and / or other positioning of one or more of the optical lens systems 410 and 412 and / or the display panels 406 and 408 within the housing of the HMD device 405, which is for personalizing or otherwise adjusting the target pupil positions of one or both of the near-to-eye display systems 402 and 404 corresponding to one or both of the actual positions of the pupils 494. Such a motor 438 may be controlled, for example, by the user operating one or more control buttons 437 on the housing 414 and / or by the user operating one or more associated separate I / O controllers (not shown). In other embodiments, the HMD device 405 may control the alignment and / or other positioning of the optical lens systems 410 and 412 and / or the display panels 406 and 408 using an adjustable positioning mechanism (e.g., a screw, slider, ratchet, etc.) that can be manually changed by the user using the control buttons 437 without using such a motor 438. Further, although only one motor 438 is illustrated in FIG. 4 for only one of the near-to-eye display systems, in some embodiments, each near-to-eye display system may have its own one or more motors, and in some embodiments, one or more motors may be used to control each of the plurality of near-to-eye display systems (e.g., independently).
[0047] The techniques described may be used with a display system similar to that illustrated in some embodiments, but in other embodiments, other types of display systems may be used, including those having a single optical lens and a display device, or those having a plurality of such optical lenses and display devices. Non-exclusive examples of other such devices include cameras, telescopes, microscopes, binoculars, spotting scopes, surveying scopes, etc. Further, the techniques described may be used with a variety of display panels or other display devices that emit light to form an image, and one or more users may view these images through one or more optical lenses. In other embodiments, one or more images created by a method other than through a display panel, such as an image created on a surface that reflects some or all of the light from another light source, may be viewed by a user through one or more optical lenses.
[0048] FIG. 5 shows an example of using a plurality of gaze tracking assemblies each including a light source and a light detector for determining a user's gaze position in a particular manner in a particular embodiment according to the techniques described. In particular, FIG. 5 includes information 500 for showing the operation of an exemplary display panel 510 and associated optical lens 508 when providing image information to the user's eye 504 and focusing that information, for example, on the pupil 506 of the eye. In the embodiment shown, four gaze tracking assemblies 511a-511d (collectively 511) of the gaze tracking subsystem are installed proximate to the edge of the optical lens 508 and are each generally directed toward the pupil 506 of the eye 504 to emit light toward the eye 504 and capture the light reflected from some or all of the pupil 506 or the surrounding iris 502.
[0049] In the illustrated example, each of the gaze tracking assemblies 511 includes a light source 512 and a photodetector 514. However, in other implementations, the photodetector 514 may be independent of the light source (e.g., four photodetectors and one light source, or two photodetectors and one light source, etc.). In this example, the gaze tracking assemblies 511 are disposed at positions including near the top of the optical lens 508 along the central vertical axis, near the bottom of the optical lens along the central vertical axis, near the left of the optical lens along the central horizontal axis, and near the right of the display panel along the central horizontal axis. In other embodiments, the gaze tracking assemblies 511 may be positioned at other locations, and fewer or more gaze tracking assemblies may be used.
[0050] It will also be understood that the light source and photodetector are shown for illustrative purposes only, and that other embodiments may include more or fewer light sources or detectors, and that the light source or detector may be placed at other locations. Additionally, although not shown here, in some embodiments, additional hardware components may be used to assist in obtaining data from one or more of the photodetectors. For example, an HMD device or other display device may include various light sources (e.g., infrared light, visible light, etc.) at different positions for irradiating light onto the iris and pupil, reflecting it, and returning it to one or more photodetectors, such as a light source installed on or near the display panel 510 or, alternatively, elsewhere (e.g., on the inner surface (not shown) of the HMD device including the display panel 510 and the optical lens 508, etc., between the optical lens 508 and the eye 504). In some such embodiments, light from such an illumination source may bounce further off the display panel before passing through the optical lens 508 to illuminate the iris and pupil.
[0051] FIG. 6 is a perspective view 600 of an exemplary high-sensitivity angular optical detector 603 that may be used in one or more implementations of the present disclosure. FIG. 7 is a cross-sectional view 700 of the high-sensitivity angular optical detector 603 shown in FIG. 6, and FIG. 8 is a top view 800 of the high-sensitivity angular optical detector. In this example, the optical detector 603 includes a quadrant photodetector 604 disposed on a common substrate 602. The quadrant photodetector 604 includes four active regions or cells 604a-604d separated by small gaps. It should be understood that other types of high-sensitivity angular detectors, such as photodiode detectors having fewer or more cells, high-sensitivity position detectors, etc., may also be used.
[0052] The high-sensitivity angular detector 603 includes an opaque screen or cover 606 positioned at a distance 612 above or in front of the photodetector 604. The opaque screen 606 may be coupled to the substrate 602 or integrated therewith, for example, as part of a housing. The opaque screen 606 has an aperture 608 therein that allows light 616 from the object 614 to pass therethrough. In this example, the object 614 is the user's eye, including the dark pupil 615 of the eye, but the embodiments described herein may be used to track other objects (e.g., a controller, a headset, other objects). The optical detector 603 further includes an imaging lens 610 positioned within or proximate to the aperture 608 of the opaque screen 606. Advantageously, the imaging lens 610 is configured to focus an image 618 of the pupil 615 of the user's eye 614 onto the quadrant photodetector. That is, the imaging lens 610 may be designed to have an object plane substantially in the same plane as the predicted position of the pupil 615 of the user's eye 614 with respect to the imaging lens 610, and an image plane substantially in the same plane as the quadrant photodetector 604. The imaging lens 610 may include a single lens or may include multiple lenses. Further, the imaging lens 610 may include one or more optical components, such as one or more filters, coatings, etc.
[0053] In a non-limiting illustrated example, the active regions (e.g., anodes) of each of the elements 604a - 604d of the photodetector can be individually available such that the light illuminating a single quadrant is electrically characterized as being only in that quadrant. As the light translates across the high-sensitivity angle detector 603, the energy of the light is dispersed between the adjacent elements 604a - 604d, and the difference in the electrical contribution to each element defines the relative position of the light. The relative intensity profiles of the elements 604a - 604d can be used to determine the position of the light applied to the cell.
[0054] As shown, the light 616 reflected from the user's eye 614 and passing through the lens 610 forms an image 618 of the eye 614, which includes a dark spot surrounded by a brighter region due to the dark pupil 615. This is because the dark pupil 615 reflects a relatively small amount of light compared to the light reflected by a part of the eye 614 and the user's face surrounding the pupil. Since the dark pupil 615 will affect the intensity of the light on the cell, the image 618 (including the dark pupil) formed on the quadrant photodetector 604 can be electrically characterized to determine the position of the pupil 615 relative to the high-sensitivity angle detector 603. As discussed elsewhere in this specification, the position information may be used for various other techniques that can be implemented using gaze tracking and pupil location information. As discussed below, the systems and methods of the present disclosure may utilize multiple light sources and high-sensitivity angle detectors to determine the position of a user's eye or other components such as components of an HMD system.
[0055] FIG. 9 is a simplified diagram 900 of an imaging lens 610 and an optical detector 603 of an object tracking system according to one non-limiting illustrated implementation. As shown, the imaging lens 610 is configured to focus an image 906 of an object 904 to be tracked onto a quadrant photodetector 604. That is, the imaging lens 610 may be designed to have an object plane or space substantially in the same plane as the predicted position of the object 904 to be tracked relative to the imaging lens 610, and an image plane or space substantially in the same plane as the quadrant photodetector 604. In this way, cells 604a - 604d may be used to determine the position of the object 904 in the object space so that the object can be tracked as discussed above. As noted above, the imaging lens 610 may include a single lens, may include multiple lenses, and may include one or more other types of optical components, such as one or more filters, coatings, etc.
[0056] FIG. 10A is a perspective view 1000 of a high-sensitivity angular optical detector 603 and an object 1002 to be detected, where the object is placed at a first position. The object 1002 may include a determined pattern 1004 thereon. FIG. 10B is a perspective view of the high-sensitivity angular optical detector 603 and the object 1002 to be detected of FIG. 10A, where the object 1002 is placed at a second position. As shown, the imaging lens 610 receives light 1006 reflected from the object 1002 and generates an image 1008 of the pattern 1004 onto the photodetector cells 604a - 604b. As the object 1004 moves from the first position shown in FIG. 10A to the second position shown in FIG. 10B, the image 1008 of the pattern 1004 also moves accordingly. Control circuitry may be used to process the detector data received from the photodetector cells 604a - 604b to track the position of the object 1002 in space along one or more dimensions.
[0057] In the illustrated example, pattern 1004 may be designed to have alternating light and dark portions (e.g., a series of spaced dark bars), which can help provide a more discrete signal in photodetector cells 604a - 606b. In at least some implementations, by providing a discrete signal or step, as an object moves in space along one or more dimensions, the ability of optical detector 603 to track object 1002 can be improved.
[0058] In at least some implementations, machine learning techniques may be used to implement a gaze tracking subsystem of an HMD device, such as the gaze tracking subsystem discussed herein, according to one non - limiting illustrated implementation. For example, a model training portion and an inference portion of a machine learning system may be provided. In the training portion, training data is fed into a machine learning algorithm to generate a trained machine learning model. The training data may include, for example, labeled data from photodetectors that specify fixation positions. As a non - limiting example, in one embodiment including four photodetectors directed at a user's eye, each training sample may include the output from each of the four photodetectors and a known or inferred fixation direction. In at least some implementations, the user's fixation direction may be known, or may be inferred by instructing the user to fixate on a particular user interface element (e.g., a word, dot, "X", another shape or object, etc.) on the display of the HMD device, and the user interface element may be static or movable on the display. The training data may also include samples where one or more of the light sources or photodetectors are blocked, due to the user's blink, eyelashes, glasses, hat, or other obstructions, etc. For such training samples, the label may be "unknown" or "blocked" instead of the specified fixation direction.
[0059] Training data may be obtained from multiple users of the HMD system and / or from a single user. The training data may be obtained in a controlled environment and / or during actual use by the user (“field training”). Further, in at least some implementations, the model may be updated or calibrated occasionally (e.g., periodically, continuously, after a particular event) to provide accurate gaze direction prediction.
[0060] In the inference unit, runtime data is provided as input to the trained machine learning model, and the trained machine learning model generates a gaze direction prediction. Continuing with the above example, the output data of the photodetector may be provided as input to the trained machine learning model, and the trained machine learning model may process the data to predict the gaze position. Next, the gaze direction prediction may be provided to one or more components associated with the HMD device, such as one or more VR or AR applications running on the HMD device, one or more display or rendering modules, one or more mechanical control units, one or more position tracking subsystems, etc.
[0061] The machine learning techniques utilized to implement the features discussed herein may include any type of suitable structure or technique. By way of non-limiting example, the machine learning model may include a decision tree, a statistical hierarchical model, a support vector machine, an artificial neural network (ANN), such as a convolutional neural network (CNN) or a recurrent neural network (RNN) (e.g., a long short-term memory (LSTM) network), a mixture density network (MDN), one or more of a hidden Markov model, or others may also be used. In at least some implementations, such as those utilizing an RNN, the machine learning model may utilize past input (memory, feedback) information to predict the gaze direction. Such implementations may advantageously utilize sequential data to determine motion information or previous gaze direction predictions, thereby providing a more accurate real-time gaze direction prediction.
[0062] As another example, in at least some implementations, the HMD device may include an interpupillary distance (IPD) adjustment component, which is operable to automatically adjust one or more components of the HMD to account for a variable IPD. In this example, the IPD adjustment component may receive the gaze direction and may align at least one component of the HMD system for the user based at least in part on the determined gaze direction. For example, when the user is looking at a nearby object, the IPD may be relatively short, and the IPD adjustment may accordingly align one or more components of the HMD device. As yet another non-limiting example, the HMD device may automatically adjust the focus of the lens based on the determined gaze direction.
[0063] Although the above example uses machine learning techniques to determine the gaze direction from the light detection information, it should be understood that the features of the present disclosure are not limited to using machine learning techniques. Generally, any type of prediction model or function may be used. For example, in at least some implementations, rather than proceeding directly from the light detection information to the gaze direction, the system may do the reverse, i.e., predict the light detection information given an input gaze direction. Such a method may discover a prediction function or model that maps from the gaze direction to the predicted light reading values for that direction to perform this prediction. Since this function may be user-specific, the system may implement a calibration process to discover or customize the function. Once the prediction function is determined or generated, the prediction function may then be inverted (e.g., using a numerical solver) to generate real-time predictions during use. Specifically, given a sample of light detection information from the actual sensor, the solver is operable to find the gaze direction that minimizes the error between this actual reading and the predicted reading from the generated prediction function. In such an implementation, the output is the determined gaze direction plus a residual error, which can be used to judge the quality of the solution. In at least some implementations, some additional correction may be applied to address various problems, such as the HMD system slipping around the user's face during operation. The prediction function may be any type of function. As an example, given a dataset of points captured from the user, a set of 2D polynomials may be used to map the gaze angle to the photodiode readings. In at least some other implementation manners, a lookup table, or other techniques including fitting an ML system to output predictions as discussed above, may also be used.
[0064] In some embodiments, it will be understood that the functionality provided by the routines discussed above may be provided in alternative ways, such as being divided among more routines or integrated into fewer routines. Similarly, in some embodiments, the illustrated routines may provide more or fewer functions than described, such as when other illustrated routines instead lack or include such functionality respectively, or when the amount of functionality provided is changed. Additionally, although various operations may be shown as being performed in a particular manner (e.g., sequentially or in parallel) and / or in a particular order, one of ordinary skill in the art will understand that in other embodiments, the operations may be performed in other orders and in other ways. Similarly, the data structures discussed above may be structured in different ways, such as by dividing a single data structure into multiple data structures or integrating multiple data structures into a single data structure, and may include databases or screens / pages of user interfaces or other types of data structures. Similarly, in some embodiments, the illustrated data structures may store more or less information than described, such as when other illustrated data structures instead lack or include such information respectively, or when the amount or type of information stored is changed.
[0065] In addition, the sizes and relative positions of elements within the drawings are not necessarily drawn to scale, including the shapes and angles of various elements, and some elements are enlarged and positioned to improve the readability of the drawings, and the particular shapes of at least some elements are selected to facilitate recognition without conveying information regarding their actual shapes or scales. Additionally, some elements may be omitted for clarity and emphasis. Further, reference numerals that are repeated in different drawings may denote the same or similar elements.
[0066] From the foregoing, although specific embodiments have been described herein for purposes of illustration, it will be understood that various modifications can be made without departing from the spirit and scope of the invention. In addition, although specific aspects of the invention may sometimes be presented in specific claim forms or may not be embodied in any claim, the inventors contemplate various aspects of the invention in any claim form that may be available. For example, although some aspects of the invention may be enumerated at a particular time as being embodied only in a computer-readable medium, other aspects may be so embodied as well.
Claims
1. A gaze tracking system, comprising: A support structure that can be arranged on a user's head; and A plurality of optical detectors held by the support structure, each of the plurality of optical detectors comprising: A quadrant photodetector including four optically active regions; An opaque screen spaced apart from the quadrant photodetector and positioned in front of the quadrant photodetector, the opaque screen including an aperture therein; and An imaging lens positioned proximate to the aperture of the opaque screen, the imaging lens being configured to focus an image of the user's eye onto the quadrant photodetector. A gaze tracking system comprising the above.
2. Receiving detector data from each of the plurality of optical detectors; Processing the detector data received from the plurality of optical detectors; and Tracking the position of the user's eye based at least in part on the processing of the received detector data. A control circuit configured as above. The gaze tracking system according to claim 1, further comprising the above.
3. A light source held by the support structure, the light source comprising a light emitting diode that emits light having a wavelength between 780 nm and 1000 nm. The gaze tracking system according to claim 1, further comprising the above.
4. Causing the light source to emit light; Receiving detector data from each of the plurality of optical detectors; Processing the detector data received from the plurality of optical detectors; and Tracking the position of the user's eye based at least in part on the processing of the received detector data. A control circuit configured as above. The gaze tracking system according to claim 3, further comprising the above.
5. For each of the plurality of optical detectors, the imaging lens is: An object plane substantially coplanar with the predicted position of the pupil of the user's eye with respect to the imaging lens; and An image plane substantially coplanar with the quadrant photodetector The gaze tracking system according to claim 1, which is designed to have.
6. The gaze tracking system according to any one of claims 1 to 5, wherein the imaging lens has at least two lenses.
7. A support structure that can be placed on the user's head; A display held by the support structure; A light source held by the support structure, the light source being configured to emit light toward the user's eyes; A plurality of optical detectors held by the support structure, each of the plurality of optical detectors being: A quadrant photodetector having four optically active regions; An opaque screen spaced apart from the quadrant photodetector and located in front of the quadrant photodetector, the opaque screen including an opening therein; and An imaging lens located adjacent to the opening of the opaque screen, the imaging lens being configured to focus an image of the user's eye onto the quadrant photodetector, having; and Cause the light source to emit light; Receive detector data from each of the plurality of optical detectors; Process the detector data received from the plurality of optical detectors; and Track the position of the user's eye based at least in part on the processing of the received detector data A control circuit configured as A head-mounted display system comprising.
8. The head-mounted display system according to claim 7, wherein the light source includes a light-emitting diode that emits light having a wavelength between 780 nm and 1000 nm.
9. For each of the plurality of optical detectors, the imaging lens is: An object plane substantially on the same plane as the predicted position of the pupil of the user's eye with respect to the imaging lens; and An image plane substantially on the same plane as the quadrant photodetector The head-mounted display system according to claim 7, which is designed to have.
10. The head-mounted display system according to any one of claims 7 to 9, wherein the imaging lens has at least two lenses.
11. A quadrant photodetector having four optically active regions; An opaque screen spaced apart from the quadrant photodetector and located in front of the quadrant photodetector, the opaque screen including an opening therein; and An imaging lens located close to the opening of the opaque screen, the imaging lens being configured to focus an image of an object to be tracked onto the quadrant photodetector An optical detector comprising.
12. The optical detector according to claim 11, wherein the imaging lens has at least two lenses.
13. A method for tracking an object, comprising: Providing a plurality of optical detectors, each of the plurality of optical detectors: A quadrant photodetector including four optically active regions; An opaque screen spaced apart from the quadrant photodetector and located in front of the quadrant photodetector, the opaque screen including an opening therein; and An imaging lens positioned close to the opening of the opaque screen, the imaging lens being configured to focus an image of an object to be tracked onto the quadrant photodetector; Receiving detector data from each of the plurality of optical detectors; Processing the detector data received from the plurality of optical detectors; and Tracking the position of the object based at least in part on the processing of the received detector data A method comprising.
14. Causing a light source to emit light having a wavelength between 780 nm and 1000 nm The method according to claim 13, further comprising.
15. For each of the plurality of optical detectors, the imaging lens is: An object plane substantially in the same plane as the predicted position of the pupil of the user's eye with respect to the imaging lens; and An image plane substantially in the same plane as the quadrant photodetector The method according to claim 13, designed to have.
16. The method according to any one of claims 13 to 15, wherein the imaging lens has at least two lenses.
17. The method according to any one of claims 13 to 15, wherein the object includes a human eye.