Coherent light speckles for secure authentication and liveness detection

Capillary pattern detection using a coherent light source and eye tracking camera in AR/VR devices addresses vulnerabilities in conventional methods by offering secure and efficient user authentication and liveness verification, resistant to spoofing and daily variations.

WO2025155939A1PCT designated stage expired Publication Date: 2025-07-24META PLATFORMS TECHNOLOGIES LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/012241
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-18
Filing Date
2025-01-18
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Conventional authentication methods for augmented and virtual reality (AR/VR) devices, such as passwords, retina ID, and voice ID, are vulnerable to circumvention by 3D-printed phantoms or recordings, and iris authentication requires high-resolution cameras or additional hardware, lacking robustness in user verification and liveness detection.

Method used

User authentication and liveness detection are achieved through capillary pattern detection using a coherent light source and eye tracking camera, generating a speckle contrast video to construct a capillary map, which is unique to each user and resistant to spoofing, with algorithms for improving map quality and verifying consciousness.

Benefits of technology

Enhances the accuracy and efficiency of user authentication and liveness detection in AR/VR devices without increasing manufacturing or computational complexity, providing robust security against phantom attacks and daily variations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025012241_24072025_PF_FP_ABST
    Figure US2025012241_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Augmented and / or virtual reality (AR / VR) near-eye display devices implementing adjustment of rendering quality based on saccade detection to preserve computational and power resources are disclosed. In examples, a wearable augmented reality / virtual reality (AR / VR) display device comprises a display to provide AR / VR content, an eye tracking system comprising a coherent light source and a camera, and a processor coupled to the display and the eye tracking system. The processor may capture a raw video of an eye, generate a speckle contrast video by applying speckle contrast computation on the raw video, construct a capillary map of the eye from the speckle contrast video, and authenticate a user based on the capillary map.
Need to check novelty before this filing date? Find Prior Art

Description

COHERENT LIGHT SPECKLES FOR SECURE AUTHENTICATION AND LIVENESS DETECTIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims benefit of and priority to U.S. provisional patent application Ser. No. 63 / 622320 filed January 18, 2024.TECHNICAL FIELD

[0002] This disclosure relates generally to augmented and / or virtual reality (AR / VR) devices, and in particular, to user authentication and liveness detection through capillary pattern detection in. AR / VR eye tracking systems.BACKGROUND

[0003] With recent advances in technology, prevalence and proliferation of content creation and delivery has increased greatly in recent years. In particular, interactive content such as virtual reality (VR) content, augmented reality (AR) content, mixed reality (MR) content, and content within and associated with a real and / or virtual environment (e.g., a “metaverse”) has become appealing to consumers.

[0004] To facilitate delivery of this and other related content, service providers have endeavored to provide various forms of wearable display systems. One such example may be a head-mounted display (HMD) device, such as a wearable eyewear, a wearable headset, or eyeglasses. In some examples, the head-mounted display (HMD) device may project or direct light to may display virtual objects or combine images of real objects with virtual objects, as in virtual reality (VR), augmented reality (AR), or mixed reality (MR) applications. For example, in an AR system, a user may view both images of virtual objects (e.g., computer-generated images (CGIs)) and the surrounding environment. Head-mounted display (HMD) devices may also present interactive content, where a user’s (wearer’s) gaze may be used as input for the interactive content.SUMMARY

[0005] According to a first aspect of the disclosure, there is provided a method for user authentication in an augmented reality I virtual reality (AR / VR) display device, the method comprising: capturing a raw video of an eye; generating a speckle contrast video by applying speckle contrast computation on the raw video; constructing a capillary map of the eye from the speckle contrast video; and authenticating a user based on the capillary map.

[0006] In some embodiments, the method further comprises: verifyingconsciousness of the user based on a gaze trajectory determined using gaze prediction; and authorizing the user based on the consciousness verification.

[0007] In some embodiments, the method further comprises: detecting a liveness of the user by generating a speckle frame measurement employing crosscorrelation analysis; and authorizing the user based on the detecting the liveness.

[0008] In some embodiments, constructing the capillary map of the eye from the speckle contrast video comprises: generating a three-dimensional point cloud using structure from motion; and constructing the capillary map through texture mapping.

[0009] In some embodiments, the method further comprises: presenting one or more stimuli to the user while capturing the raw video.

[0010] In some embodiments, the method further comprises: performing a quality assessment on the capillary map to identify one or more missing regions; generating new stimuli based on the identified one or more missing regions; and reconstructing the capillary map using the new stimuli.

[0011] According to a further aspect of the disclosure, there is provided a wearable augmented reality I virtual reality (AR / VR) display device comprising: a display to provide ARA / R content; an eye tracking system comprising a coherent light source and a camera; and a processor coupled to the display and the eye tracking system, the processor to: capture a raw video of an eye; generate a speckle contrast video by applying speckle contrast computation on the raw video; construct a capillary map of the eye from the speckle contrast video; and authenticate a user based on the capillary map.

[0012] In some embodiments, the processor is further to: verify consciousness of the user based on a gaze trajectory determined using gaze prediction; and authorize the user based on the consciousness verification.

[0013] In some embodiments, the processor is further to: detect a liveness of the user by generating a speckle frame measurement employing cross-correlation analysis; and authorize the user based on the liveness detection.

[0014] In some embodiments, detecting the liveness of the user further includes analyzing pupil dilation in response to stimuli.

[0015] In some embodiments, the processor is further to: generate a three- dimensional point cloud using structure from motion; and construct the capillary map through texture mapping.

[0016] In some embodiments, the processor is further to: present one or morestimuli to the user while capturing the raw video.

[0017] In some embodiments, the processor is further to: perform a quality assessment on the capillary map to identify one or more missing regions; generate new stimuli based on the identified one or more missing regions; and reconstruct the capillary map using the new stimuli.

[0018] According to a further aspect of the disclosure, there is provided a non- transitory computer readable medium configured to store program code instructions, when executed by a processor, cause the processor to perform steps comprising: capture a raw video of an eye; generate a speckle contrast video by applying speckle contrast computation on the raw video; construct a capillary map of the eye from the speckle contrast video; and authenticate a user based on the capillary map.

[0019] In some embodiments, the instructions, when executed by the processor, cause the processor to: verify consciousness of the user based on a gaze trajectory determined using gaze prediction; and authorize the user based on the consciousness verification.

[0020] In some embodiments, the instructions, when executed by the processor, cause the processor to: detect a liveness of the user by generating a speckle frame measurement employing cross-correlation analysis; and authorize the user based on the liveness detection.

[0021] In some embodiments, to detect a liveness of the user, the instructions, when executed by the processor, cause the processor to: analyze pupil dilation in response to stimuli.

[0022] In some embodiments, to construct the capillary map of the eye from the speckle contrast video, the instructions, when executed by the processor, cause the processor to: generate a three-dimensional point cloud using structure from motion; and construct the capillary map through texture mapping.

[0023] In some embodiments, the instructions, when executed by the processor, cause the processor to: present one or more stimuli to the user while capturing the raw video.

[0024] In some embodiments, the instructions, when executed by the processor, cause the processor to: perform a quality assessment on the capillary map to identify one or more missing regions; generate new stimuli based on the identified one or more missing regions; and reconstruct the capillary map using the new stimuli. BRIEF DESCRIPTION OF DRAWINGS

[0025] Features of the present disclosure are illustrated by way of example and not limited in the following figures, in which like numerals indicate like elements. One skilled in the art will readily recognize from the following that alternative examples of the structures and methods illustrated in the figures can be employed without departing from the principles described herein.

[0026] Figure 1 illustrates a block diagram of an artificial reality system environment including a near-eye display, according to an example.

[0027] Figures 2A-2C illustrate various views of a near-eye display device in the form of a head-mounted display (HMD) device, according to examples.

[0028] Figure 3 illustrates a perspective view of a near-eye display in the form of a pair of glasses, according to an example.

[0029] Figure 4A illustrates major components of an eye tracking system in an ARA / R near-eye display device, according to an example.

[0030] Figure 4B illustrates capillary pattern detection through an eye tracking system of an ARA / R near-eye display device, according to an example.

[0031] Figures 5A and 5B illustrate unique capillary patterns in different eyes, according to an example.

[0032] Figure 6 illustrates a comparison of spatial-temporal change or lack thereof of speckles in human iris and inanimate objects.

[0033] Figure 7 illustrates how speckle contrast can be used to authenticate a person overcoming phantom eye attack, according to an example.

[0034] Figure 8 illustrates a flow diagram for construction of a capillary map in an ARA / R display device, according to some examples.

[0035] Figure 9 illustrates a flow diagram for user authentication using a capillary map in an AR / VR display device, according to some examples.DETAILED DESCRIPTION

[0036] For simplicity and illustrative purposes, the present application is described by referring mainly to examples thereof. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. It will be readily apparent, however, that the present application may be practiced without limitation to these specific details. In other instances, some methods and structures readily understood by one of ordinary skill in the art have not been described in detail so as not to unnecessarily obscure the present application. As used herein, the terms “a” and “an” are intended to denote atleast one of a particular element, the term “includes” means includes but not limited to, the term “including” means including but not limited to, and the term “based on” means based at least in part on.

[0037] User authentication and liveness detection are crucial for ensuring secure access and multi-user management on AR / VR devices. Conventional authentication methods like passwords, retina ID, or voice ID may be circumvented using 3D-printed phantoms or recordings. Moreover, conventional solutions such as iris authentication require high resolution cameras or additional hardware.

[0038] In some examples of the present disclosure, approaches are described for user authentication in AR / VR devices through capillary pattern detection. Liveness detection and consciousness verification may also be performed together with eye tracking. A coherent light source (e.g., laser, VCSEL) and a detection device (eye tracking camera, scanning detector) may be used for laser speckle contrast imaging (LSCI) to map blood flow within the eye blood vessels (e.g., capillaries). Different computational approaches may be used to improve the quality of the capillary map through identification of missing regions and new screen stimuli. Initial authentication may be performed during calibration and continued during use.

[0039] While some advantages and benefits of the present disclosure are apparent, other advantages and benefits may include enhanced accuracy and efficiency of user authentication and liveness detection in AR / VR display devices without an increase in manufacturing or computational complexity.

[0040] Figure 1 illustrates a block diagram of an artificial reality system environment 100 including a near-eye display, according to an example. As used herein, a “near-eye display” may refer to a device (e.g., an optical device) that may be in close proximity to a user’s eye. As used herein, “artificial reality” may refer to aspects of, among other things, a “metaverse” or an environment of real and virtual elements and may include use of technologies associated with virtual reality (VR), augmented reality (AR), and / or mixed reality (MR). As used herein a “user” may refer to a user or wearer of a “near-eye display.”

[0041] As shown in Figure 1 , the artificial reality system environment 100 may include a near-eye display 120, an optional external imaging device 150, and an optional input / output interface 140, each of which may be coupled to a console 110. The console 110 may be optional in some instances as the functions of the console 110 may be integrated into the near-eye display 120. In some examples, the near-eye display 120 may be a head-mounted display (HMD) that presents content to a user.

[0042] In some instances, for a near-eye display system, it may generally be desirable to expand an eye box, reduce display haze, improve image quality (e.g., resolution and contrast), reduce physical size, increase power efficiency, and increase or expand field of view (FOV). As used herein, “field of view” (FOV) may refer to an angular range of an image as seen by a user, which is typically measured in degrees as observed by one eye (for a monocular head-mounted display (HMD)) or both eyes (for binocular head-mounted displays (HMDs)). Also, as used herein, an “eye box” may be a two-dimensional box that may be positioned in front of the user’s eye from which a displayed image from an image source may be viewed.

[0043] In some examples, in a near-eye display system, light from a surrounding environment may traverse a “see-through” region of a waveguide display (e.g., a transparent substrate) to reach a user’s eyes. For example, in a near-eye display system, light of projected images may be coupled into a transparent substrate of a waveguide, propagate within the waveguide, and be coupled or directed out of the waveguide at one or more locations to replicate exit pupils and expand the eye box.

[0044] In some examples, the near-eye display 120 may include one or more rigid bodies, which may be rigidly or non-rigidly coupled to each other. In some examples, a rigid coupling between rigid bodies may cause the coupled rigid bodies to act as a single rigid entity, while in other examples, a non-rigid coupling between rigid bodies may allow the rigid bodies to move relative to each other.

[0045] In some examples, the near-eye display 120 may be implemented in any suitable form-factor, including a head-mounted display (HMD), a pair of glasses, or other similar wearable eyewear or device. Examples of the near-eye display 120 are further described below with respect to Figures 2 and 3. Additionally, in some examples, the functionality described herein may be used in a head-mounted display (HMD) or headset that may combine images of an environment external to the near- eye display 120 and artificial reality content (e.g., computer-generated images). Therefore, in some examples, the near-eye display 120 may augment images of a physical, real-world environment external to the near-eye display 120 with generated and / or overlaid digital content (e.g., images, video, sound, etc.) to present an augmented reality to a user.

[0046] In some examples, the near-eye display 120 may include any number ofdisplay electronics 122, display optics 124, and an eye tracking unit 130. In some examples, the near-eye display 120 may also include one or more locators 126, one or more position sensors 128, and an inertial measurement unit (IMU) 132. In some examples, the near-eye display 120 may omit any of the eye tracking unit 130, the one or more locators 126, the one or more position sensors 128, and the inertial measurement unit (IMU) 132, or may include additional elements.

[0047] In some examples, the display electronics 122 may display or facilitate the display of images to the user according to data received from, for example, the optional console 110. In some examples, the display electronics 122 may include one or more display panels. In some examples, the display electronics 122 may include any number of pixels to emit light of a predominant color such as red, green, blue, white, or yellow. In some examples, the display electronics 122 may display a three- dimensional (3D) image, e.g., using stereoscopic effects produced by two-dimensional panels, to create a subjective perception of image depth.

[0048] In some examples, the near-eye display 120 may include a projector (not shown), which may form an image in angular domain for direct observation by a viewer’s eye through a pupil. The projector may employ a controllable light source (e.g., a laser source) and a micro-electromechanical system (MEMS) beam scanner to create a light field from, for example, a collimated light beam. In some examples, the same projector or a different projector may be used to project a fringe pattern on the eye, which may be captured by a camera and analyzed (e.g., by the eye tracking unit 130) to determine a position of the eye (the pupil), a gaze, etc.

[0049] In some examples, the display optics 124 may display image content optically (e.g., using optical waveguides and / or couplers) or magnify image light received from the display electronics 122, correct optical errors associated with the image light, and / or present the corrected image light to a user of the near-eye display 120. In some examples, the display optics 124 may include a single optical element or any number of combinations of various optical elements as well as mechanical couplings to maintain relative spacing and orientation of the optical elements in the combination. In some examples, one or more optical elements in the display optics 124 may have an optical coating, such as an anti-reflective coating, a reflective coating, a filtering coating, and / or a combination of different optical coatings.

[0050] In some examples, the display optics 124 may also be designed to correct one or more types of optical errors, such as two-dimensional optical errors,three-dimensional optical errors, or any combination thereof. Examples of two- dimensional errors may include barrel distortion, pincushion distortion, longitudinal chromatic aberration, and / or transverse chromatic aberration. Examples of three- dimensional errors may include spherical aberration, chromatic aberration field curvature, and astigmatism.

[0051] In some examples, the one or more locators 126 may be objects located in specific positions relative to one another and relative to a reference point on the near-eye display 120. In some examples, the optional console 110 may identify the one or more locators 126 in images captured by the optional external imaging device 150 to determine the artificial reality headset’s position, orientation, or both. The one or more locators 126 may each be a light-emitting diode (LED), a corner cube reflector, a reflective marker, a type of light source that contrasts with an environment in which the near-eye display 120 operates, or any combination thereof.

[0052] In some examples, the external imaging device 150 may include one or more cameras, one or more video cameras, any other device capable of capturing images including the one or more locators 126, or any combination thereof. The optional external imaging device 150 may be configured to detect light emitted or reflected from the one or more locators 126 in a field of view of the optional external imaging device 150.

[0053] In some examples, the one or more position sensors 128 may generate one or more measurement signals in response to motion of the near-eye display 120. Examples of the one or more position sensors 128 may include any number of accelerometers, gyroscopes, magnetometers, and / or other motion-detecting or errorcorrecting sensors, or any combination thereof.

[0054] In some examples, the inertial measurement unit (IMU) 132 may be an electronic device that generates fast calibration data based on measurement signals received from the one or more position sensors 128. The one or more position sensors 128 may be located external to the inertial measurement unit (IMU) 132, internal to the inertial measurement unit (IMU) 132, or any combination thereof. Based on the one or more measurement signals from the one or more position sensors 128, the inertial measurement unit (IMU) 132 may generate fast calibration data indicating an estimated position of the near-eye display 120 that may be relative to an initial position of the near-eye display 120. For example, the inertial measurement unit (IMU) 132 may integrate measurement signals received from accelerometers over time toestimate a velocity vector and integrate the velocity vector over time to determine an estimated position of a reference point on the near-eye display 120. Alternatively, the inertial measurement unit (IMU) 132 may provide the sampled measurement signals to the optional console 110, which may determine the fast calibration data.

[0055] The eye tracking unit 130 may include one or more eye tracking systems. As used herein, “eye tracking” may refer to determining an eye’s position or relative position, including orientation, location, and / or gaze of a user’s eye. In some examples, an eye tracking system may include an imaging system that captures one or more images of an eye and may optionally include a light emitter, which may generate light (e.g., a fringe pattern) that is directed to an eye such that light reflected by the eye may be captured by the imaging system (e.g., a camera). In other examples, the eye tracking unit 130 may capture reflected radio waves emitted by a miniature radar unit. These data associated with the eye may be used to determine or predict eye position, orientation, movement, location, and / or gaze.

[0056] In some examples, the near-eye display 120 may use the orientation of the eye to introduce depth cues (e.g., blur image outside of the user’s main line of sight), collect heuristics on the user interaction in the virtual reality (VR) media (e.g., time spent on any particular subject, object, or frame as a function of exposed stimuli), some other functions that are based in part on the orientation of at least one of the user’s eyes, or any combination thereof. In some examples, because the orientation may be determined for both eyes of the user, the eye tracking unit 130 may be able to determine where the user is looking or predict any user patterns, etc.

[0057] In some examples, the input / output interface 140 may be a device that allows a user to send action requests to the optional console 110. As used herein, an “action request” may be a request to perform a particular action. For example, an action request may be to start or to end an application or to perform a particular action within the application. The input / output interface 140 may include one or more input devices. Example input devices may include a keyboard, a mouse, a game controller, a glove, a button, a touch screen, or any other suitable device for receiving action requests and communicating the received action requests to the optional console 110. In some examples, an action request received by the input / output interface 140 may be communicated to the optional console 110, which may perform an action corresponding to the requested action.

[0058] In some examples, the optional console 1 10 may provide content to thenear-eye display 120 for presentation to the user in accordance with information received from one or more of external imaging device 150, the near-eye display 120, and the input / output interface 140. For example, in the example shown in Figure 1 , the optional console 110 may include an application store 112, a headset tracking module 114, a virtual reality engine 116, and an eye tracking module 118. Some examples of the optional console 110 may include different or additional modules than those described in conjunction with Figure 1 . Functions further described below may be distributed among components of the optional console 110 in a different manner than is described here.

[0059] In some examples, the optional console 110 may include a processor and a non-transitory computer-readable storage medium storing instructions executable by the processor. The processor may include multiple processing units executing instructions in parallel. The non-transitory computer-readable storage medium may be any memory, such as a hard disk drive, a removable memory, or a solid-state drive (e.g., flash memory or dynamic random access memory (DRAM)). In some examples, the modules of the optional console 110 described in conjunction with Figure 1 may be encoded as instructions in the non-transitory computer-readable storage medium that, when executed by the processor, cause the processor to perform the functions further described below. It should be appreciated that the optional console 110 may or may not be needed or the optional console 1 10 may be integrated with or separate from the near-eye display 120.

[0060] In some examples, the application store 112 may store one or more applications for execution by the optional console 110. An application may include a group of instructions that, when executed by a processor, generates content for presentation to the user. Examples of the applications may include gaming applications, conferencing applications, video playback application, or other suitable applications.

[0061] In some examples, the headset tracking module 114 may track movements of the near-eye display 120 using slow calibration information from the external imaging device 150. For example, the headset tracking module 114 may determine positions of a reference point of the near-eye display 120 using observed locators from the slow calibration information and a model of the near-eye display 120. Additionally, in some examples, the headset tracking module 114 may use portions of the fast calibration information, the slow calibration information, or any combinationthereof, to predict a future location of the near-eye display 120. In some examples, the headset tracking module 1 14 may provide the estimated or predicted future position of the near-eye display 120 to the virtual reality engine 116.

[0062] In some examples, the virtual reality engine 116 may execute applications within the artificial reality system environment 100 and receive position information of the near-eye display 120, acceleration information of the near-eye display 120, velocity information of the near-eye display 120, predicted future positions of the near-eye display 120, or any combination thereof from the headset tracking module 114. In some examples, the virtual reality engine 116 may also receive estimated eye position and orientation information from the eye tracking module 118. Based on the received information, the virtual reality engine 1 16 may determine content to provide to the near-eye display 120 for presentation to the user.

[0063] In some examples, a location of a projector of a display system may be adjusted to enable any number of design modifications. For example, in some instances, a projector may be located in front of a viewer’s eye (i.e. , “front-mounted” placement). In a front-mounted placement, in some examples, a projector of a display system may be located away from a user’s eyes (i.e., “world-side”). In some examples, a head-mounted display (HMD) device may utilize a front-mounted placement to propagate light towards a user’s eye(s) to project an image.

[0064] As mentioned herein, user authentication in AR / VR devices may be accomplished through capillary pattern detection. Liveness detection and consciousness verification may also be performed together with eye tracking. A coherent light source (e.g., laser, VCSEL) and a detection device (eye tracking camera, scanning detector) may be used for laser speckle contrast imaging (LSCI) to map blood flow within the eye blood vessels (e.g., capillaries). Different algorithms may be used to improve the quality of the capillary map through identification of missing regions and new screen stimuli. Initial authentication may be performed during calibration and continued during use.

[0065] Figures 2A-2C illustrate various views of a near-eye display device in the form of a head-mounted display (HMD) device 200, according to examples. In some examples, the head-mounted device (HMD) device 200 may be a part of a virtual reality (VR) system, an augmented reality (AR) system, a mixed reality (MR) system, another system that uses displays or wearables, or any combination thereof. As shown in diagram 200A of Figure 2A, the head-mounted display (HMD) device 200may include a body 220 and a head strap 230. The front perspective view of the headmounted display (HMD) device 200 further shows a bottom side 223, a front side 225, and a right side 229 of the body 220. In some examples, the head strap 230 may have an adjustable or extendible length. In particular, in some examples, there may be a sufficient space between the body 220 and the head strap 230 of the head-mounted display (HMD) device 200 for allowing a user to mount the head-mounted display (HMD) device 200 onto the user’s head. For example, the length of the head strap 230 may be adjustable to accommodate a range of user head sizes. In some examples, the head-mounted display (HMD) device 200 may include additional, fewer, and / or different components such as a display 210 to present a wearer augmented reality (AR) I virtual reality (VR) content and a camera to capture images or videos of the wearer’s environment.

[0066] As shown in the bottom perspective view of diagram 200B of Figure 2B, the display 210 may include one or more display assemblies and present, to a user (wearer), media or other digital content including virtual and / or augmented views of a physical, real-world environment with computer-generated elements. Examples of the media or digital content presented by the head-mounted display (HMD) device 200 may include images (e.g., two-dimensional (2D) or three-dimensional (3D) images), videos (e.g., 2D or 3D videos), audio, or any combination thereof. In some examples, the user may interact with the presented images or videos through eye tracking sensors enclosed in the body 220 of the head-mounted display (HMD) device 200. The eye tracking sensors may also be used to adjust and improve quality of the presented content.

[0067] In some examples, the head-mounted display (HMD) device 200 may include various sensors (not shown), such as depth sensors, motion sensors, position sensors, and / or eye tracking sensors. Some of these sensors may use any number of structured or unstructured light patterns for sensing purposes. In some examples, the head-mounted display (HMD) device 200 may include an input / output interface for communicating with a console communicatively coupled to the head-mounted display (HMD) device 200 through wired or wireless means. In some examples, the headmounted display (HMD) device 200 may include a virtual reality engine (not shown) that may execute applications within the head-mounted display (HMD) device 200 and receive depth information, position information, acceleration information, velocity information, predicted future positions, or any combination thereof of the head-mounted display (HMD) device 200 from the various sensors.

[0068] In some examples, the information received by the virtual reality engine may be used for producing a signal (e.g., display instructions) to the display 210. In some examples, the head-mounted display (HMD) device 200 may include locators (not shown), which may be located in fixed positions on the body 220 of the headmounted display (HMD) device 200 relative to one another and relative to a reference point. Each of the locators may emit light that is detectable by an external imaging device. This may be useful for the purposes of head tracking or other movement / orientation. It should be appreciated that other elements or components may also be used in addition or in lieu of such locators.

[0069] It should be appreciated that in some examples, a projector mounted in a display system may be placed near and / or closer to a user’s eye (i.e., “eye-side”). In some examples, and as discussed herein, a projector for a display system shaped like eyeglasses may be mounted or positioned in a temple arm (i.e., a top far comer of a lens side) of the eyeglasses. It should be appreciated that, in some instances, utilizing a back-mounted projector placement may help to reduce size or bulkiness of any required housing required for a display system, which may also result in a significant improvement in user experience for a user.

[0070] In some examples, behind-the-lens eye tracking (ET) systems, in-frame ET systems, glass-embedded ET cameras, waveguide-coupled ET cameras, and similar ones may be used to employ speckle-generated contrast as a way to authenticate a user based on capillary pattern in the eye.

[0071] Figure 3 is a perspective view of a near-eye display 300 in the form of a pair of glasses (or other similar eyewear), according to an example. In some examples, the near-eye display 300 may be a specific example of near-eye display 120 of Figure 1 and may be configured to operate as a virtual reality display, an augmented reality (AR) display, and / or a mixed reality (MR) display.

[0072] In some examples, the near-eye display 300 may include a frame 305 and a display 310. In some examples, the display 310 may be configured to present media or other content to a user. In some examples, the display 310 may include display electronics and / or display optics, similar to components described with respect to Figures 1 and 2A-2C. For example, as described above with respect to the near- eye display 120 of Figure 1 , the display 310 may include a liquid crystal display (LCD) display panel, a light-emitting diode (LED) display panel, or an optical display panel(e.g., a waveguide display assembly). In some examples, the display 310 may also include any number of optical components, such as waveguides, gratings, lenses, mirrors, etc. In other examples, the display 310 may include a projector, or in place of the display 310 the near-eye display 300 may include a projector.

[0073] In some examples, the near-eye display 300 may further include various sensors on orwithin a frame 305. In some examples, the various sensors may include any number of depth sensors, motion sensors, position sensors, inertial sensors, and / or ambient light sensors, as shown. In some examples, the various sensors may include any number of image sensors configured to generate image data representing different fields of views in one or more different directions. In some examples, the various sensors may be used as input devices to control or influence the displayed content of the near-eye display, and / or to provide an interactive virtual reality (VR), augmented reality (AR), and / or mixed reality (MR) experience to a user of the near- eye display 300. In some examples, the various sensors may also be used for stereoscopic imaging or other similar applications.

[0074] Figure 4A illustrates major components of an eye tracking system in an ARA / R near-eye display device, according to an example. Diagram 400A shows an eye 402 being illuminated with coherent illumination 410 from a light source 404, and a camera 406 capturing reflections of the projected pattern from the surface of the eye 402.

[0075] By capturing a user’s gaze or other eye properties through reflection of coherent illumination from a surface of the eye, eye tracking allows adjustment of displayed content, interactivity, and other functionalities associated with AR / VR display devices. For example, power consumption reduction and computational resource preservation may be accomplished by adjusting displayed content characteristics (e.g., resolution, frame rate, brightness level, etc.) depending on the user’s gaze. Eye tracking results may also be used as input for various functionalities of an AR / VR display device such as selection of a displayed action item, change of displayed content, etc.

[0076] In some cases, eye tracking may offer potential alternatives in user authentication and liveness detection. For example, unique patterns in which users move their eyes, known as eye movement dynamics, may serve as a distinctive authentication metric. Additionally, analyzing pupil dilation in response to certain stimuli may further confirm a user's liveness.

[0077] Figure 4B illustrates capillary pattern detection through an eye tracking system of an ARA / R near-eye display device, according to an example. Diagram 400B shows the eye 402 with details such as capillaries 412. The light source 404 may provide the coherent illumination 410, and capillary structure may be captured through a scan of the eye in reflected images with a waveguide aperture for image collection 420.

[0078] In some examples, a speckle-generated contrast may be used as a way to authenticate a user based on capillary pattern in the eye. The speckle-ID may be captured by monitoring how coherent light interacts with blood flowing in the capillaries (sub-surface blood vessels). The capillary map may be generated by evaluating spatio-temporal speckle dynamics in two or more subsequent frames from an eye tracking camera (single-frame solution may also be possible but may be less robust against forgery). Like fingerprints, the capillary map is unique for each user.

[0079] The reference map may be captured during initial calibration step, and then a few subsequent frames from eye tracking camera may be used for one-off or continuous user authentication. Thus, eye capillaries may be highlighted through speckle contrast by leveraging the coherent light source and camera in eye tracking systems. The 3D capillary map may be constructed from the videos captured during the calibration process as personal authentication ID. Furthermore, the frame stability of speckle patterns may be analyzed to verify the liveness of the imaged surface.

[0080] Figures 5A and 5B illustrate unique capillary patterns in different eyes, according to an example. Diagrams 500A and 500B show unique patterns of capillaries in the eyes of two different human subjects. The images have been modified to protect privacy of the human subjects. Thus, the illustrated patterns are modified representations of actual patterns, not the real patterns.

[0081] A system according to some examples, a coherent light source (e.g., laser, VCSEL) and a camera are leverages to capture the eye biospeckles and capillaries, which may be used for personal authentication ID, liveness detection, and consciousness verification. Described approaches are compatible with eye tracking (ET) systems using coherent light sources, and do not require additional hardware. An initial setup of personal authentication ID may be seamlessly conducted with the eye tracking calibration, improving the user experience.

[0082] Using a coherent light source to visualize capillaries adds another layer of security protection: the liveness detection (temporal evolution property of eyebiospeckles), making the technique more robust to phantom eye attacks, where an unchanging image of the eye is used. Furthermore, a capillary map may be less affected by daily activities. In comparison, a fingerprint can be affected by wet or greasy fingers, a face ID may be affected by face coverings and makeup, and an iris ID may be affected and / or attacked by contact lens, for example.

[0083] The capillary map may work seamlessly with capillary-based eye tracking. Combining the two systems may allow a simplified setup procedure (equivalent to the calibration of eye tracking). Eye tracking allows consciousness verification. Moreover, the capillary-based eye tracking may ensure that only the authorized user can operate the device at any time of the operation eliminating the need for multiple authentication actions when using an AR / VR device (e.g., authorization of purchase after unlocking the device).

[0084] Figure 6 illustrates a comparison of spatial-temporal change or lack thereof of speckles in human iris and inanimate objects. Diagram 600 shows images of speckle patterns derived from various surfaces including three living tissue examples: iris, sclera, and skin, as well as an inanimate object: cardboard.

[0085] A speckle pattern derived from living tissue (e.g., skin, eye) differs from that observed from an inanimate surface in that the speckles show a spatio-temporal change. Thus, the temporal evolution property of biospeckles in human eyes may be used to authenticate users in AR / VR systems and similar ones.

[0086] Figure 7 illustrates how speckle contrast can be used to authenticate a person overcoming phantom eye attack, according to an example. Diagram 700 shows the difference between speckle contrasts of a phantom eye and an actual human eye illustrating how phantom eye attacks cannot overcome a capillary map based authentication mechanism due to temporal evolution properties of eye biospeckles.

[0087] The temporal evolution property of eye biospeckles caused by constant physiological activities on and under the surface of living human eyes can prove the liveness of the eye, making it more robust to phantom and cadaver eye attacks.

[0088] Example authentication mechanisms employ a coherent light source and a camera from the eye tracking system to visualize capillaries. Videos of capillaries captured during the calibration of eye tracking may be used to reconstruct the capillary map as personal authentication ID. The spatio-temporal evolution property of eye biospeckles may be used as an indicator of liveness (as an “anti-spoof” technique). In some examples, the capillary ID may be combined with other authenticationmechanisms (e.g., voice ID, eye movement patterns (dynamic speckle patterns and gaze traces), iris ID, etc.) to enhance the security level. The capillary ID may also be used for multi-user management. A headset may automatically and securely detect a user based on capillary ID and provide user-specific content.

[0089] While the examples herein are presented for the eye, the discussed authentication mechanism may be extended for skin images or bare sensor array that is pressed against the skin. In this case the illumination source and sensor may be adjacent to each other and the wavelength may be selected to optimize the amount of light that enters below the skin surface, interacts with skin and blood vessels, and is detected. Speckle contrast capillary map is specific to each person. The contrast comes from motion, which is hard to simulate using fake eyes. Liveness detection based on the biospeckle is simple, fast, and inherently immune to print attacks, cadaver attacks.

[0090] Figure 8 illustrates a flow diagram for construction of a capillary map in an AR / VR display device, according to some examples. The method 800 is provided by way of example, as there may be a variety of ways to carry out the method described herein. Although the method 800 is primarily described as being performed by the components of Figures 4A and 4B, the method 800 may be executed or otherwise performed by one or more processing components of another system or a combination of systems. Each block shown in Figure 8 may further represent one or more processes, methods, or subroutines, and one or more of the blocks may include machine readable instructions stored on a non-transitory computer readable medium and executed by a processor or other type of processing circuit to perform one or more operations described herein.

[0091] The method 800 for construction of a capillary map in an AR / VR display device may begin with the user being asked to view different stimuli on the ARA / R display. A raw video of the eye with the capillaries may be captured followed by speckle contrast computation, or speckle contrast calculation, providing temporal evolution (indicated in Fig. 8 by arrow t) speckle contrast video. From the speckle contrast video, 3D eye point cloud may be obtained using structure from motion (SfM). Subsequently, the 3D capillary map may be constructed through texture mapping.

[0092] A capillary map quality assessment may be performed on the constructed 3D capillary map. If the quality assessment result is negative, or failure, (e.g., missing regions), missing region identification may be performed. Screen stimulilocation may be computed, or calculated, from the identified missing capillary map regions. New display, or screen, stimuli may be determined from the screen stimuli location computation, or screen stimuli location calculation, and the process may begin again with the capture of raw video of the eye using the new stimuli. Each time the raw video is captured a temporal evolution (indicated in Fig. 8 by arrow t) of the eye may be captured, e.g., evolution through time on a frame-by-frame basis of the video.

[0093] Figure 9 illustrates a flow diagram for user authentication using a capillary map in an AR / VR display device, according to some examples. The method 900 is provided by way of example, as there may be a variety of ways to carry out the method described herein. Although the method 900 is primarily described as being performed by the components of Figures 4A and 4B, the method 900 may be executed or otherwise performed by one or more processing components of another system or a combination of systems. Each block shown in Figure 9 may further represent one or more processes, methods, or subroutines, and one or more of the blocks may include machine readable instructions stored on a non-transitory computer readable medium and executed by a processor or other type of processing circuit to perform one or more operations described herein.

[0094] The method 900 for user authentication, liveness detection, and consciousness verification through eye tracking may begin with capture of raw video of the user’s eye. The raw video capture may include capture of a temporal evolution (indicated in Fig. 9 by arrow t) of the eye, e.g., evolution through time on a frame-by- frame basis of the video. In one example, the user may be asked to look at displayed stimuli moving with a certain trajectory (for strict consciousness verification). In some examples, this may include determining if a gaze may be moving and / or following a stimulus trajectory or looking within a certain area. For example, as in Fig. 9, consciousness verification includes performing gaze prediction and measuring a gaze trajectory to verify consciousness of the user. An output of the consciousness verification method may be a determination that the user is, or is not, conscious. In another example, the user may be asked to look at the display without stimuli for better user experience.

[0095] For capillary map pattern matching, speckle contrast computation, or speckle contrast calculation, on the raw video and resulting speckle contrast video may be used (through comparison with a reference 3D capillary map data) for pattern matching. The speckle contrast video may include temporal evolution (indicated inFig. 9 by arrow t) speckle contrast video of the eye. If the pattern is a match with the saved 3D capillary map data, a first step of authentication may be completed. As a result, a match, or no match, output may be generated by the capillary map pattern matching method. The absolute eye position resulting from the pattern matching may be used in consciousness verification for gaze prediction, as discussed above. Resulting gaze trajectory may be used for consciousness verification, which is the second step in user authentication. For liveness detection, a cross-correlation analysis of frames (e.g., across 10 ms) may be performed to obtain speckle frame stability measurement, which may be used for liveness detection, the third step in user authentication. As an output from the liveness detection method, a human or robot output may be generated. If all three steps are successfully completed, that is the capillary map pattern is matched, consciousness of the user is verified and liveness of the user is detected, secure authentication may be provided confirming the user’s unique identity, the user being conscious, and live.

[0096] Secure authentication within AR / VR devices may include requiring 3D capillary map matching; liveness detection; and consciousness verification according to examples. An eye tracking (ET) camera and a coherence light source (e.g., laser, VCSEL) of the AR / VR device may be leveraged to map blood flow within the eye blood vessels (e.g., capillaries) of a user of the AR / VR device using laser speckle contrast imaging (LSCI). A 3D reference map of the capillaries and biospeckles (e.g., primary speckles) of the user’s eyes may be generated (e.g., during initial calibration of the AR / VR device). The 3D reference map may be generated by using an algorithm that employs different screen stimuli; speckle contrast calculation for generating a speckle contrast video; structure from motion (SfM); and texture mapping. The quality of the 3D reference map may be assessed and improved by using a feedback loop, which identifies missing regions and applies new screen stimuli, when the quality assessment fails.

[0097] A new map may be generated from few subsequent frames during operation of the AR / VR device when authentication is required. The new map may be compared against the 3D reference map and the user granted access when a match occurs. The method includes determining an absolute eye position from the new map. Further, gaze tracing and temporal evolution of speckles may be used for consciousness verification. Cross-correlation analysis of the frames (e.g., across 10 milliseconds) may be employed for determining a spatio-temporal evolution of thebiospeckles. The spatio-temporal evolution of the biospeckles may be used for liveness detection as an anti-spoof technique.

[0098] According to examples, a method of making an AR / VR display device with eye tracking system that employs capillary structures for authentication and liveness detection is described herein. A system of making ARA / R display device with eye tracking system is also described herein. A non-transitory computer-readable storage medium may have an executable stored thereon, which when executed instructs a processor to perform the methods described herein.

[0099] In the foregoing description, various examples are described, including devices, systems, methods, and the like. For the purposes of explanation, specific details are set forth in order to provide a thorough understanding of examples of the disclosure. However, it will be apparent that various examples may be practiced without these specific details. For example, devices, systems, structures, assemblies, methods, and other components may be shown as components in block diagram form in order not to obscure the examples in unnecessary detail. In other instances, well- known devices, processes, systems, structures, and techniques may be shown without necessary detail in order to avoid obscuring the examples.

[0100] The figures and description are not intended to be restrictive. The terms and expressions that have been employed in this disclosure are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof. The word "example" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "example1is not necessarily to be construed as preferred or advantageous over other embodiments or designs.

[0101] Although the methods and systems as described herein may be directed mainly to digital content, such as videos or interactive media, it should be appreciated that the methods and systems as described herein may be used for other types of content or scenarios as well. Other applications or uses of the methods and systems as described herein may also include social networking, marketing, content-based recommendation engines, and / or other types of knowledge or data-driven systems.

Claims

CLAIMS1 . A method for user authentication in an augmented real ity / virtual reality, AR / VR, display device, the method comprising: capturing a raw video of an eye; generating a speckle contrast video by applying speckle contrast computation on the raw video; constructing a capillary map of the eye from the speckle contrast video; and authenticating a user based on the capillary map.

2. The method of claim 1 , further comprising: verifying consciousness of the user based on a gaze trajectory determined using gaze prediction; and authorizing the user based on the consciousness verification; and / or preferably detecting a liveness of the user by generating a speckle frame measurement employing cross-correlation analysis; and authorizing the user based on the detecting the liveness.

3. The method of claim 1 or claim 2, wherein constructing the capillary map of the eye from the speckle contrast video comprises: generating a three-dimensional point cloud using structure from motion; and constructing the capillary map through texture mapping.

4. The method of any preceding claim, further comprising: presenting one or more stimuli to the user while capturing the raw video.

5. The method of any preceding claim, further comprising: performing a quality assessment on the capillary map to identify one or more missing regions; generating new stimuli based on the identified one or more missing regions; and reconstructing the capillary map using the new stimuli.

6. A wearable augmented reality / virtual reality, AR / VR, display device comprising: a display to provide AR / VR content; an eye tracking system comprising a coherent light source and a camera; and a processor coupled to the display and the eye tracking system, the processor to: capture a raw video of an eye;generate a speckle contrast video by applying speckle contrast computation on the raw video; construct a capillary map of the eye from the speckle contrast video; and authenticate a user based on the capillary map.

7. The wearable AR / VR display device of claim 6, wherein the processor is further to: verify consciousness of the user based on a gaze trajectory determined using gaze prediction; and authorize the user based on the consciousness verification; and / or preferably detect a liveness of the user by generating a speckle frame measurement employing cross-correlation analysis; and authorize the user based on the liveness detection; and further preferably wherein detecting the liveness of the user further includes analyzing pupil dilation in response to stimuli.

8. The wearable AR / VR display device of claim 6 or claim 7, wherein the processor is further to: generate a three-dimensional point cloud using structure from motion; and construct the capillary map through texture mapping.

9. The wearable AR / VR display device of any one of claims 6 to 8, wherein the processor is further to: present one or more stimuli to the user while capturing the raw video.

10. The wearable AR / VR display device of any one of claims 6 to 9, wherein the processor is further to: perform a quality assessment on the capillary map to identify one or more missing regions; generate new stimuli based on the identified one or more missing regions; and reconstruct the capillary map using the new stimuli.

11. A non-transitory computer readable medium configured to store program code instructions, when executed by a processor, cause the processor to perform steps comprising: capture a raw video of an eye; generate a speckle contrast video by applying speckle contrast computation on the raw video;construct a capillary map of the eye from the speckle contrast video; and authenticate a user based on the capillary map.

12. The non-transitory computer readable medium of claim 11 , wherein the instructions, when executed by the processor, cause the processor to: verify consciousness of the user based on a gaze trajectory determined using gaze prediction; and authorize the user based on the consciousness verification; and / or preferably detect a liveness of the user by generating a speckle frame measurement employing cross-correlation analysis; and authorize the user based on the liveness detection; and further preferably wherein to detect a liveness of the user, the instructions, when executed by the processor, cause the processor to: analyze pupil dilation in response to stimuli.

13. The non-transitory computer readable medium of claim 11 or claim 12, wherein to construct the capillary map of the eye from the speckle contrast video, the instructions, when executed by the processor, cause the processor to: generate a three-dimensional point cloud using structure from motion; and construct the capillary map through texture mapping.

14. The non-transitory computer readable medium of any one of claims 11 to 13, wherein the instructions, when executed by the processor, cause the processor to: present one or more stimuli to the user while capturing the raw video.

15. The non-transitory computer readable medium of any one of claims 11 to 14, wherein the instructions, when executed by the processor, cause the processor to: perform a quality assessment on the capillary map to identify one or more missing regions; generate new stimuli based on the identified one or more missing regions; and reconstruct the capillary map using the new stimuli.