Eye and / or face tracking based on a dynamic vision sensor

By employing DVS with IMUs and machine learning algorithms to detect changes in light intensity from multiple sources, the challenges of tracking game controllers in VR and AR systems are addressed, achieving high update rates and efficient data processing.

JP2025516615AActive Publication Date: 2025-05-30SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024566447
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-10
Filing Date
2023-04-12
Publication Date
2025-05-30
Estimated Expiration
2043-04-12

AI Technical Summary

Technical Problem

Existing VR and AR systems face challenges in accurately and efficiently tracking game controllers, particularly in dynamic environments with head-mounted displays, due to limitations in frame rate, data processing requirements, and the need for multiple sensors and light sources.

Method used

The use of Dynamic Vision Sensors (DVS) with single or dual sensor arrays, combined with inertial measurement units (IMUs) and machine learning algorithms, to track game controllers by detecting changes in light intensity from multiple light sources, reducing the need for high frame rate cameras and simplifying data processing.

Benefits of technology

This approach provides a high update rate for motion tracking, reduces irrelevant data output, and allows for more accurate and efficient tracking of game controllers in dynamic environments, enhancing the performance of VR and AR systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516615000001_ABST
    Figure 2025516615000001_ABST
Patent Text Reader

Abstract

The tracking system uses one or more light sources to direct light towards one or more of the user's eyes. A dynamic vision sensor (DVS) operably coupled to a processor is configured to visually recognize the user's eyes. The DVS has an array of photosensitive elements within a known configuration. The DVS outputs a signal corresponding to an event at a corresponding photosensitive element within the array in response to a change in the light from the light source reflected from a portion of the user's eye. The output signal includes information corresponding to the position of the corresponding photosensitive element within the array. The processor determines an association between each event and a corresponding particular light source and adapts the orientation of the user's eyes using the determined association, the relative position of the light source with respect to the user's eyes, and the position of the corresponding photosensitive element within the array.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Aspects of the present disclosure relate to the tracking of game controllers. Specifically, aspects of the present disclosure relate to the tracking of game controllers using dynamic vision sensors.

Background Art

[0002] Latest virtual reality (VR) and augmented reality (AR) implementations rely on accurate and fast movement tracking for user interaction with devices. AR and VR often rely on information regarding the position and orientation of the controller relative to other objects. Many VR and AR implementations rely on a combination of inertial measurements obtained by accelerometers or gyroscopes within the controller and visual detection of the controller by an external camera to determine the position and orientation of the controller.

[0003] Some of the earliest implementations use detected infrared light by an infrared camera with a defined detection radius on a game controller oriented towards the screen. The camera captures images at a moderately high rate of 200 frames per second and determines the position of the infrared light. The distance between the infrared lights is from a predetermined one and from the relative position of the infrared lights in the camera image that can calculate the position of the controller relative to the screen. An accelerometer may also be used to provide information regarding relative three-dimensional changes in the position or orientation of the controller. These conventional implementations rely on a fixed position of the screen and a controller oriented towards the screen. In the latest VR and AR implementations, the screen can be placed near the user's face within a head-mounted display that moves with the user. Therefore, having absolute light positions (also called light houses) is undesirable as it requires extra setup time and the user has to set up independent light house points that limit the user's range of motion. Furthermore, even the moderately high frame rate of 200 frames per second of the infrared camera was not fast enough to provide smooth motion feedback. Additionally, since this setup is too simplistic, it is not suitable for more recent inside-out detection methods such as room mapping and hand detection.

[0004] In more recent implementations, cameras and accelerometers are used in combination with trained machine learning algorithms trained to detect both hands, controllers, and / or other body parts. To achieve smoothness in motion detection, a high frame rate camera is needed to generate image frames for body part / controller detection. This generates a large amount of data that needs to be processed quickly for smoothness of the update rate. Therefore, expensive hardware is required to process the frame data. Furthermore, much of the frame data within each frame is discarded as unnecessary as it has no relation to motion tracking.

[0005] Aspects of the present disclosure arise in this context. SUMMARY OF THE INVENTION

[0006] The teachings of the present invention can be easily understood by considering the following detailed description in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 11A

Figure 11B

Figure 11C

Figure 11D

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16A

Figure 16B

Figure 16C

Figure 16D

Figure 16E

Figure 17

Figure 18A

Figure 18B

Figure 19

DETAILED DESCRIPTION OF THE INVENTION

[0008] The following detailed description includes many specific details for illustrative purposes, but those skilled in the art will recognize that many variations and modifications to the following details are within the scope of the present invention. Therefore, the exemplary embodiments of the invention described below are presented without loss of generality to the claimed invention and without imposing limitations on the claimed invention.

[0009] Preamble A new type of vision system called a Dynamic Vision Sensor (DVS) has recently been developed. This DVS uses only changes in the light intensity of a photosensitive pixel array to resolve changes in a scene. The DVS has a very high update rate and, instead of delivering a stream of image frames, provides a nearly continuous stream of the locations of changes in pixel intensity. Each change in pixel intensity is sometimes called an event. This has the added advantage of significantly reducing irrelevant data output.

[0010] Two or more light sources can provide continuous updates regarding the position of a DVS camera relative to a position indicator at an update rate determined by the blinking rate of the light. In some implementations, the two or more light sources may be infrared light sources and the DVS may use infrared-sensitive pixels. Alternatively, the DVS may be sensitive to the visible light spectrum and the two or more light sources may be multiple visible light sources or one visible light at a known wavelength. In implementations with a DVS sensitive to visible light, the DVS may also be sensitive to movement occurring within its field of view (FOV). The DVS can detect changes in light intensity caused by the reflection of light from a moving surface. In implementations using an infrared-sensitive DVS, the light from an infrared illuminator can be used to detect movement within the FOV by reflection.

[0011] Implementation Figure 1 shows an example of the implementation of tracking a game controller using a DVS101 with a single sensor array according to one aspect of the present disclosure. In the implementation shown, the DVS is attached to a headset 102 that can be part of a head-mounted display. A controller 103 including two or more light sources is within the field of view of the DVS101. In the example shown, the controller 103 includes four light sources 104, 105, 106, and 107. These light sources have a known configuration with respect to each other and with respect to the controller 103. Here, there is one DVS with a single photosensitive array. Using such four light sources, the position and orientation of the controller 103 with respect to the DVS101 can be accurately determined. The known information regarding the light sources can include the distance between each of the other light sources with respect to each of the light sources, and the position of each of the light sources on the controller 103. As shown, three light sources 104, 105, 106 can draw a plane, and the light source 107 may be out of the plane with respect to the plane drawn by the three light sources 104, 105, 106. The light sources here have a known configuration. For example, but not limited to, the first light source 104 is located at the upper left front, the second light source 105 is located at the upper right front, the third light source 106 is located on the left side away from the top, and the fourth light source 107 is located at the lower center front of the controller. With the four light sources, a DVS having a single photosensitive array may be able to determine the movement of the controller along the X, Y, and Z axes. Further, an inertial measurement unit (IMU) 108 can be coupled to the controller 103. By way of example, the IMU 108 can include an accelerometer configured to measure acceleration with respect to one axis, two axes, or three axes. Alternatively, the IMU can include a gyroscope configured to sense changes in rotation with respect to one axis, two axes, or three axes. In some implementations, the IMU may include both an accelerometer and a gyroscope. The IMU 108 can be used to refine the determination of motion, position, and orientation based on the information from the DVS101 using a processor. The processor may be located within the headset 102, a game console, or other computing device (not shown).DVS101, headset 102, and IMU 108 can be operably coupled to a processor 110, which may be located on a separate device such as the headset 102, controller 103, or a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor can implement tracking as described herein, for example, as described hereinafter with respect to FIGS. 4, 8, and 9. Further, the processor 110 can control the blinking of light sources 104, 105, 106.

[0012] During operation, the DVS 101 having a photosensitive array can detect the movement of light sources 104, 105, 106, 107 with the photosensitive array, and can be transmitted to the processor when a change in light is detected by the photosensitive array. In some implementations, the light sources can be configured to turn on and off in a predetermined pattern, for example, but not limited to, using signals from a circuit and / or a processor. The processor can use the predetermined pattern to determine the identity of each light source. The identity of the light source can include a known position relative to the controller and other light sources. In other implementations, each light source can be configured to turn off and on in a predetermined pattern, and the pattern can be used to determine the identity of that particular light source. In some implementations, the processor can adapt the known configuration of the light source relative to the controller to the events detected by the photosensitive array.

[0013] The DVS can have a nearly continuous update rate that can be discretely approximated to about one million updates per second. A DVS with a high update rate may be able to resolve a very high-speed blinking pattern of a light source. The blinking rate is mainly limited by the Nyquist frequency, that is, half of the sample rate of the DVS. The light source can blink with a duty cycle suitable for detection of the blinking by the DVS. Generally speaking, the "on" time of the blinking needs to be long enough to be consistently detected by the DVS. Furthermore, due to the high update rate, slight differences in the blinking rate may be detectable.

[0014] The light sources 104, 105, 106, 107 may be broad visible spectrum light such as incandescent lamps or white light emitting diodes. Alternatively, the light sources 104, 105, 106, 107 may be infrared light, or the light source may have a specific light spectrum profile detectable by the DVS101. The DVS101 can include a photosensitive array configured to detect the emission spectrum of the light sources 104, 105, 106, 107. For example, but not limited to, if the light source is infrared light, the photosensitive array of the DVS may be sensitive to infrared light, or if the light source has a specific emission spectrum, the photosensitive array may be configured to enhance sensitivity to the specific emission spectrum of the light source. Additionally, for example, but not by way of limitation, the photosensitive array of the DVS may be insensitive to light of other wavelengths not emitted by the light source, or may exclude such light. For example, if the light source is infrared light, the photosensitive array may be configured to detect only infrared light.

[0015] Figure 2 shows an example of the implementation of tracking of a game controller using a DVS with a dual sensor array according to one aspect of the present disclosure. In this implementation, the headset 203 includes a first DVS 201 and a second DVS 202. Alternatively, the headset 203 may include a DVS having a first photosensitive array 201 and a second photosensitive array 202. The general functions of the light source and the DVS are the same as those described above with respect to FIG. 1. Information from the second DVS or the second array can be integrated with information from the first array to provide a better fit to the orientation of the controller and some depth information. The two DVSs or photosensitive arrays can have partially overlapping fields of view that allow the use of binocular disparity.

[0016] The two DVSs or two photosensitive arrays provide binocular vision for depth perception. Thereby, the number of light sources can be further reduced. The first DVS and the second DVS or the first photosensitive array and the second photosensitive array may be separated by a known distance, for example, but not limited to, about 50 to 100 millimeters or more than 100 millimeters. More generally, this spacing is large enough to provide sufficient disparity for the desired depth sensitivity, but not so large as to have no overlap between the fields of view. As shown, the controller 207 can include a first light source 204, a second light source 205, and a third light source 206. Since information from the two DVSs or two arrays provides sufficient information for determining the position and orientation of the controller, a fourth light source coupled to the controller may not be necessary. The controller can include an IMU 208 that can provide additional inertial information used to refine the determination of position and orientation.

[0017] The first DVS 201, the second DVS 202, the headset 203, and the IMU 208 may be operably coupled to a processor 210, which may be located on the headset 203, the controller 207, or a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor can implement tracking as described herein, for example, as will be described later with respect to FIGS. 4, 8, and 9. Further, the processor 210 can control the flashing of the light sources 204, 205, 206.

[0018] FIG. 2 shows two DVSs or two photosensitive arrays, but aspects of the present disclosure are not so limited. The device can include any number of DVSs or photosensitive arrays. For example, but not limited to, a DVS having three DVSs or three separate photosensitive arrays may allow the use of only two light sources coupled to the controller. The third DVS or photosensitive array may not be collinear with the other two DVSs or arrays. Similar to binocular disparity, each of the additional DVSs or photosensitive arrays can be separated by a known distance and can have a partially overlapping field of view, which can increase the parallax effect used. Further, some implementations may include multiple DVSs each having multiple photosensitive arrays. For example, but not limiting, there may be two DVSs each having two separate photosensitive arrays.

[0019] Figure 3 shows an example of the implementation of the tracking of a game controller using a DVS in combination with a single sensor array and a camera, according to one aspect of the present disclosure. In this implementation, a camera 302 is added to the DVS 301. The DVS 301 and the camera 302 can be coupled to a headset 303. The DVS 301 and the camera 302 may have partially overlapping fields of view, or may share the same field of view. The controller may include three or more light sources 304, 305, 306 on the controller 307. The DVS 301 and the camera 302 can be used together to determine the position and orientation of the controller. Frames from the camera 302 can be interpolated using events from the DVS 301. Also, using the image frames, position identification and mapping can be performed simultaneously to improve the determination of the orientation and position of the controller. Further, using the image frames from the camera, inside-out tracking of the user using a machine learning algorithm, for example, hand tracking or foot tracking can be performed. The IMU 308 can provide additional inertial information used to further refine the determination of the position and orientation of the controller.

[0020] The first DVS 301, the second DVS 302, the headset 303, and the IMU 308 can be operably coupled to a processor 310, which may be located on the headset 303, the controller 307, or a separate device such as a personal computer, a laptop computer, a tablet computer, a smartphone, or a game console. The processor can implement the tracking as described herein, for example, as described later with respect to FIGS. 4, 8, and 9. Further, the processor 310 can control the flashing of the light sources 304, 305, 306.

[0021] Figure 5 shows an example of the implementation of head tracking or other device tracking using a controller that includes a DVS with a single sensor array according to one aspect of the present disclosure. Here, the DVS 507 is attached to the controller 508. The headset 501 within the field of view of the DVS 507 is tracked using a light source attached to the headset. The headset 501 can include two or more light sources, here four light sources 502, 503, 504, and 505. The four light sources provide accurate information for three-dimensional position determination by a single DVS. Using additional DVSs or cameras can reduce the number of light sources used. The position of each light source relative to the headset 501 is known to the system and can be somewhat important for providing relevant information to the DVS. Here, the light sources 502, 503, and 504 represent a plane. The light source 505 is located outside the plane of the other light sources 502, 503, and 504. This facilitates the three-dimensional detection of the position, orientation, or movement of the headset. The headset 501 can include an IMU 506, and this IMU can use the information from the DVS 507 to improve the estimation of position and orientation. Further, the controller 508 can include an IMU 509. Using the information from the controller IMU 509, the determination of the position and orientation of the headset can be further refined. For example, but not limited to, using the IMU information to determine whether the controller is moving relative to the headset and to determine the speed or acceleration of that movement.

[0022] The headset 501, IMU 506, DVS 507, and controller IMU 509 can be operably coupled to a processor 510, and this processor can be located on a separate device such as the headset 501, controller 508, or a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor can implement the tracking as described herein, for example, as described later with respect to FIGS. 4, 8, and 9. Further, the processor 510 can control the blinking of the light sources 502, 503, 504, 505.

[0023] Figure 6 shows an example of the implementation of head tracking or other device tracking using a game controller that includes a DVS with a dual sensor array according to one aspect of the present disclosure. Here, the controller 605 is coupled to two DVSs 606, 607, or a single DVS having two photosensitive arrays 606 and 607. As described above, the two DVSs or two photosensitive arrays can be separated by an appropriate distance, for example, between 500 and 1000 millimeters, and can have a partially overlapping field of view. This enables the use of the parallax effect for depth determination. Further, the headset 601 can include three light sources 602, 603, 604 and an IMU 608. The use of two DVSs or two separated arrays 606, 607 may enable the use of fewer than four light sources, for example, but not limited to, three light sources. Two light sources 602, 603 can draw a line, and the third light source 604 can be outside the line formed by the other two light sources 602, 603. Information from the IMU 605 coupled to the headset 601 can be used to refine the determination of position and orientation.

[0024] The headset 601, DVSs 606, 607, and IMU 608 can be operably coupled to a processor 610, which can be located on the headset 601, the controller 605, or a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor can implement tracking as described herein, for example, as described later with respect to FIGS. 4, 8, and 9. Further, the processor 610 can control the flashing of the light sources 602, 603, 604.

[0025] FIG. 7 shows an example implementation of head tracking or other device tracking using a controller that includes a DVS in combination with a single sensor array and an image camera, according to one aspect of the present disclosure. Here, DVS 707 and camera 708 are coupled to controller 706. Camera 708 and DVS 707 may have partially overlapping fields of view, or may have completely overlapping fields of view. In some implementations, the pixels of the camera and the photosensitive elements of the DVS may share the same photosensitive array and thus function as an integrated DVS and camera.

[0026] Three or more light sources 702, 703, 704 can be coupled to the headset 701. For example, without limitation, the light sources may be integrated within the headset housing, and each light source may be an LED, incandescent, halogen, or fluorescent emitter mounted on a circuit board within or on the headset housing. In some implementations, a single emitter may use a plastic or glass light pipe or optical fiber that splits the light from a single emitter into two or more light sources on the headset housing to create multiple light sources.

[0027] Three or more light sources 702, 703, 704 can be configured to turn on and off in response to an electronic signal. In some implementations, three or more light sources can turn on and off in a predetermined sequence. Each time the light sources 702, 703, 704 move or blink, DVS 707 can generate an event. Camera 708 generates an image frame of its field of view at a set frame rate. Because of the high update rate of the DVS, it may be possible to interpolate the image frames generated by the camera with DVS events.

[0028] In some implementations, the IMU705 can also be coupled to the headset 701. The headset 701, IMU705, DVS707, camera 708, and IMU608 can be operably coupled to a processor 710, which may be located on a separate device such as the headset 701, the controller 706, or a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor can implement tracking as described herein, for example, as will be described later with respect to FIGS. 4, 8, and 9. Further, the processor 710 can control the blinking of the light sources 702, 703, 704.

[0029] In some further implementations, the camera 708 may be a depth camera, such as a depth time-of-flight (DTOF) sensor. A DToF camera acquires a depth image by measuring the time it takes for light to travel from a light source to an object in the scene and back to the pixel array. By way of example and not limitation, a DToF camera can operate using continuous wave (CW) modulation, which is an example of an indirect time-of-flight (ToF) sensing method. In a CW ToF camera, light from an amplitude-modulated light source is backscattered by objects within the camera's field of view (FOV), and the phase shift between the emitted waveform and the reflected waveform is measured. By measuring the phase shift at multiple modulation frequencies, a depth value for each pixel can be calculated. The phase shift is obtained by measuring the correlation between the emitted and received waveforms at different relative delays using pixel-integrated photon mixing demodulation.

[0030] A DTOF system generally includes an illumination module and an imaging module. The illumination module consists of a light source, a driver that drives the light source at a high modulation frequency, and a diffuser that projects a light beam from the light source onto a designed field of illumination (FOI). The DToF illumination module can include one or more light sources, and these one or more light sources can be amplitude modulation light emitters such as, for example, but not limited to, vertical cavity surface emitting lasers (VCSELs) or edge emitting lasers (EELs). The imaging module can include an imaging lens assembly, a bandpass filter (BPF), a microlens array, and an array of photosensitive elements that convert incident photon energy into an electrical signal. The microlens array increases the amount of light reaching the photosensitive elements, and the BPF decreases the amount of ambient light reaching the light elements and the microlens array.

[0031] Operation Figure 4 shows a DVS that tracks the movement of a game controller having two or more light sources, according to one aspect of the present disclosure. The DVS 401 has the controller 402 within its field of view. As shown, the controller 402 includes a plurality of light sources coupled to the controller body. The plurality of light sources can be configured to turn off and then turn on again at a predetermined rate. Each blinking of the light sources within the field of view of the DVS 401 can generate one or more events 403 in the DVS. In the events 403 shown, since all the light sources have changed from the off state to the on state, the events indicate that all the light sources are in the on state. Alternatively, depending on the sensitivity of the DVS, a change in the brightness of the light sources may be sufficient to trigger the events 403. Note that in other implementations, since the light may turn off and on at different rates or at different times, each event may correspond to fewer light sources than all the lit light sources. The time of the events generated by the blinking light and their corresponding positions within the array can be recorded in a memory (not shown). In some implementations, the position and orientation of the controller 402 can be reconstructed from one or more events from the DVS.

[0032] In some implementations, each light source can blink in a predetermined time sequence. The DVS can output the time at which each event occurred, for example, but not limited to, as a timestamp along with each event. A predetermined time sequence can be used at the time an event occurs to determine the identity of each light within the event, for example, which event corresponds to which light source position. In the example shown, one or more events 403 output by the DVS 401 indicate a first light 406, a second light 407, a third light 408, and a fourth light 409 detected by the photosensitive array. As described above, the identity of each light source can be determined from the information output by the DVS and the predetermined blinking sequence of the light source. For example, but not limited to, the photosensitive array of the DVS 401 can detect a light event 403 at time T+1, and the predetermined sequence can provide that the light source 406 turns on at T+1, so the light event is determined to correspond to the light source 406. The predetermined sequence can be stored in memory, for example, as a table listing the timing and position of the sequence of each light source. In some implementations, the predetermined sequence may be encoded in the blinking of the light itself. For example, but not limited to, each light may blink in a sequence indicating its identity. For example, a light source labeled 1 may blink in a Morse code sequence indicating the number 1. The identity of the light event can then be recovered through analysis of the light event. Alternatively, the sequence information can come from the light source itself or the driver of the light source, indicating when the light source turns on or off. Alternatively, the light sources can be turned on and off simultaneously, and a machine learning algorithm can be applied to the detected light events 403 to fit the posture 402 of the controller to the events and the known configuration of the light sources.

[0033] Furthermore, the orientation and position of the light source can be determined using information from events such as the size, intensity, and separation of the light. The system may have information defining the size, position on the controller body, and intensity of each light source. The position and orientation may be determined with higher accuracy from the differences between the detected size, intensity, and separation. Additionally, if one or more additional DVSs or photosensitive arrays are present, parallax information can be used to further enhance the determination of position and orientation.

[0034] During operation, the user can change the position and orientation of the controller 410. Because the update rate of the DVS is relatively high, as the light event moves at 412 during a change in position and orientation, the DVS can capture the event sequence at 411. The movement of the light event here is shown in FIG. 4 by the vector arrow 412. As the detected light event moves, the determined position and orientation of the controller 410 may be updated, or the determined position and orientation may be updated at regular intervals. Due to the blinking of the first light 415, second light 416, third light 417, and fourth light 418, the position and time of the events detected by the photosensitive elements of the DVS may conform to the new position and orientation of the controller. Additionally, inertial information from the IMU can be used to refine the movement and position and orientation of the controller 410.

[0035] The flowchart shown in FIG. 8 shows a method of tracking movement by DVS801 using one or more light sources and one light source configuration adaptation model according to one aspect of the present disclosure. In this implementation, the light sources (shown as LEDs) can be turned on and off simultaneously, independently, or in a predetermined sequence. Events 802 are generated from the DVS801 due to the movement of the light sources or the blinking of one or more of the light sources. Generally, an event relays an electronic signal containing information such as the time interval during which the event occurred, the position within the array of photosensitive elements of the DVS801 where the event occurred, and binary data (e.g., 1 or 0) corresponding to a change in light intensity exceeding some detection threshold. Each event is processed at 803 and can be formatted into a usable form, for example, but not limited to, placing the event position within a data array, associating a timestamp with the event, and aggregating multiple events. As an example of aggregating events, all events occurring at individual elements within the array within a predetermined time interval can be integrated into a single data structure for analysis.

[0036] The processed events are then analyzed to associate the detected DVS events with the corresponding LED pulses 804. In some implementations, since each light source can be turned off and on at a unique predetermined time interval, the timestamp of the aggregated events can be used to determine a predetermined time interval from the event and associate a particular light source with a particular event. For example, but not limited to, analyzing the spatial pattern of the aggregated events occurring within a predetermined time interval or sequence of time intervals to determine whether this pattern is consistent with the pulses of the LEDs. Event patterns that are too large, too small, or too irregular in shape may be excluded as LED events. Additionally, the timing of the events can also be analyzed, and events that are too short or too long may also be excluded.

[0037] The trained machine learning model 805 can be applied to the processed events. The model 805 can include information regarding the configuration of the light sources, such as the size of the light sources and their relative positions with respect to the controller body. The machine learning model can be trained by training event data having corresponding masked positions and orientations of the controller, as described in a later section. Apply the trained machine learning model to the processed event data to determine the correspondence 806 between the detected pulse 804 and the pose 808. The trained machine learning model can adapt the pose 808 of the controller, e.g., the position and orientation, to one or more processed events, e.g., the detected LED pulse 804. Alternatively, instead of the trained machine learning model 805, a fitting algorithm can be applied to the processed events. The fitting algorithm can adapt the position and orientation of the controller to the processed events using a manually developed light source model. Alternatively, the fitting algorithm can be a hypothesis and test type algorithm, which tries all possible permutations of the light correspondence and seeks the best fit using redundant light sources. Thereafter, a tracking / prediction algorithm can be applied to continue tracking the light source. Further, at 809, the predicted current pose can be used to predict the next pose. At 810, the inertial data from the IMU 807 can be fused with the predicted pose 808 to generate the final predicted position and orientation of the controller. The fusion may be performed by a trained machine learning algorithm trained to refine the position and orientation of the controller using the inertial data. Alternatively, the fusion may be performed, for example, but not limited to, by a Kalman filter or non-linear optimization.

[0038] FIG. 9 shows a motion tracking method by DVS using light source position information with timestamps according to one aspect of the present disclosure. In this implementation, the time of the events output by the DVS 901 is used to determine the position and orientation of the controller. Here, the light source is depicted as an LED, and each LED may turn on at different times as shown in the graph. Each time the LED turns on or off, the DVS having the LED within its field of view can generate an event 902. Each event may include an electronic signal corresponding to information such as the time interval during which the event occurred, the position within the array of photosensitive elements of the DVS 901 where the event occurred, and binary data (e.g., 1 or 0) corresponding to a change in light intensity exceeding some detection threshold. At 903, each event is processed to format a plurality of events by aggregating them, for example, but not limited to, placing the event positions in a data array as described above and integrating the multiple events occurring at individual elements within the array within a predetermined time interval into a single data structure. Then, the processed events are analyzed to detect the corresponding LED pulses at 904. For example, but not limited to, the shape and size of the events may be analyzed for the regularity and conformity of the light source, and events that are too large, too small, or too irregular in shape may be excluded as LED events, and noise suppression such as the average of the events may be performed to remove random events. Further, the timing of the events may also be analyzed, and events that are too short or too long may also be excluded, or multiple events may be condensed into a single event by excluding events that co-occur with the initial event and end within the time but after the initial event.

[0039] When an event is processed, at 905 the individual LED positions can be determined. Determining the individual LED positions can be performed by using the time sequence in which the LEDs turn on and off. For example, but not limited to, the time of one or more events can be compared with a known time sequence of LED blinking. The known time sequence can be, for example, a table having the on and off times of the LEDs, and the positions on the controller body for each LED, or timestamps from the LED driver when each LED is on or off. When timestamps are used, the timestamps can be correlated with the LED positions on the controller body. From the timing sequence, the LED position information, and the processed event information, the matching position and orientation of the controller can be determined. The IMU data 907 can be integrated with the previously determined LED positions and the inertial data from the IMU by the Kalman filter 908. The Kalman filter can predict the position of the light source based on the inertial information from the IMU, and integrate this prediction with the position information determined from the time sequence of the LEDs to refine the motion data, generate the final pose at 909, and refine the future estimation.

[0040] Training of a general neural network According to aspects of the present disclosure, a tracking system can use machine learning by a neural network (NN). For example, the trained model 805 discussed above can use machine learning as discussed below. Machine learning algorithms can use a training dataset, which can include inputs from a DVS such as events, or processed events with known positions and orientations of a controller as labels. Further, machine learning algorithms using an NN can perform a fusion between the position and orientation of a controller determined from DVS information and inertial information from an IMU. A training set for the fusion can be, for example, but not limited to, inertial data with potential positions and orientations of a controller, as well as final positions and orientations. In some implementations, the machine learning algorithm may be trained to perform SLAM (simultaneous localization and mapping) using a training set that includes objects such as the ground, landmarks, and body parts with hidden labels. Hidden labels may include the identity of the objects and their relative positions. As generally understood by those skilled in the art, SLAM technology generally solves the problem of continuously tracking the position of an agent within an unknown environment while constructing or updating a map of that environment.

[0041] The NN may include one or more of several different types of neural networks and may have many different layers. By way of example, and not limitation, the neural network may be composed of one or more convolutional neural networks (CNNs), recurrent neural networks (RNNs), and / or dynamic neural networks (DNNs). The motion detection neural network can be trained using the general training methods disclosed herein.

[0042] By way of example and not limitation, FIG. 10A shows a basic form of an RNN that can be used, e.g., in a trained model 805. In the illustrated example, the RNN has a layer of nodes 1020, each of which is characterized by an activation function S, one input weight U, a recurrent hidden node transition weight W, and an output transition weight V. The activation function S may be any non-linear function known in the art and is not limited to the hyperbolic tangent (tanh) function. For example, the activation function S may be a sigmoid function or a ReLu function. Unlike other types of neural networks, the RNN has one set of activation functions and weights for the entire layer. As shown in FIG. 10B, the RNN may be considered as a series of nodes 1020 having the same activation function over times T and T+1. Thus, the RNN maintains historical information by feeding the results from the previous time T to the current time T+1.

[0043] In some implementations, a convolutional RNN may be used. Another type of RNN that can be used is the long short-term memory (LSTM) neural network, which adds a memory block to the RNN node along with an input gate activation function, an output gate activation function, and a forget gate activation function that, as described by Hochreiter & Schmidhuber, "Long Short-term memory", Neural Computation 9(8):1735-1780 (1997), which is incorporated herein by reference, enables the network to hold information for a longer period of time.

[0044] FIG. 10C shows an exemplary layout of a convolutional neural network, such as a CRNN, that can be used for a trained model 805 and the like according to an aspect of the present disclosure. In this representation, the convolutional neural network is generated for an input 1032 having a size of 4 units in height and 4 units in width, which gives a total area of 16 units. The convolutional neural network represented has a filter 1033 having a size of 2 units in height and 2 units in width, with a skip value of 1 and 136 channels of size 9. Only the connections 1034 between the first column of channels and their filter windows are shown in FIG. 10C for clarity. However, aspects of the present disclosure are not limited to such implementations. In accordance with aspects of the present disclosure, the convolutional neural network may have any number of additional neural network node layers 1031 and may include such layer types as additional convolutional layers, fully connected layers, pooling layers, max pooling layers, local contrast normalization layers, etc. of any size.

[0045] As seen in FIG. 10D, training a neural network (NN) begins with the initialization of the weights of the NN at 1041. Generally, the initial weights should be randomly distributed. For example, an NN having a tanh activation function should have random values distributed between -1 / √n and 1 / √n, where n is the number of inputs to the node.

[0046] After initialization, the activation function and optimizer are defined. Then, at 1042, a feature vector or input dataset is provided to the NN. Each of the different feature vectors may be generated by the NN from inputs with known labels. Similarly, the NN may be provided with feature vectors corresponding to inputs with known labelings or classifications. The NN then predicts, at 1043, the label or classification for the feature or input. The predicted label or class is compared, at 1044, to the known label or class (also known as the ground truth), and the loss function measures the total error between the prediction and the ground truth across all training samples. By way of example, and not limitation, the loss function may be a cross-entropy loss function cost, triplet contrastive function, exponential cost, etc. Multiple different loss functions may be used depending on the purpose. By way of example, and not limitation, a cross-entropy loss function may be used to train a classifier, whereas a triplet contrastive function may be employed to learn a pre-trained embedding. The NN is then optimized and trained, as shown at 1045, using the result of the loss function and using known methods for training neural networks such as error backpropagation by adaptive gradient descent. At each training epoch, the optimizer attempts to select model parameters (i.e., weights) that minimize the training loss function (i.e., the total error). The data is partitioned into training samples, validation samples, and test samples.

[0047] During training, the optimizer minimizes the loss function with respect to the training samples. After each training epoch, the model is evaluated with respect to the validation samples by computing the validation loss and accuracy. If there is no significant change, training may be stopped, and the resulting trained model may be used to predict the labels of the test data.

[0048] Thus, a neural network may be trained from those inputs to identify and classify inputs having known labels or classifications. Similarly, an NN may be trained using the described method to generate feature vectors from inputs having known labels or classifications. The above discussion relates to RNNs and CRNNs, but these discussions can be applied to NNs that do not include a recurrent or hidden layer.

[0049] Hybrid sensor FIG. 11A is a diagram illustrating a hybrid DVS having a plurality of co-located sensor types, according to an aspect of the present disclosure. In some implementations, a hybrid DVS can be used to combine multiple sensor types in one device. The multiple sensor types can be, for example, but not limited to, DVS infrared photosensors, DVS visible light photosensors, DVS wavelength-specific photosensors, visible light camera pixels, infrared camera pixels, DTOF camera pixels. A second sensor type 1102 can be interspersed with a first sensor type 1102 on the same array 1101. For example, but not limited to, DVS visible light photosensors 1103 may surround visible light camera pixels 1102, or DVS infrared photosensors 1103 may surround DVS visible light photosensors 1102, or visible light camera pixels 1103 may surround DVS infrared photosensors 1102, or any combination thereof. Although a single element 1102 is shown surrounded by other elements 1103, aspects of the present disclosure are not so limited. A single element can include a cluster of multiple DVS photosensors or camera pixels. For example, but not limited to, a cluster of four camera pixels may be surrounded by eight DVS photosensors, or one camera pixel may be surrounded by eight pairs, triplets, or quadruplets of DVS photosensors.

[0050] Furthermore, as shown in FIG. 11B, the hybrid DVS can have a plurality of sensor types arranged in a moiré pattern within the array. Here, blocks of the second sensor type 1103 are evenly distributed among blocks of the first sensor type 1102. Each of the first sensor type 1102 and the second sensor type 1103 may be different. For example, but not limited to, the first sensor type 1102 may be a DVS infrared photosensor, and the second sensor type 1103 may be a DVS visible light photosensor.

[0051] Alternatively, sensor types may be distinguished by filtering. In these implementations, one or more filters selectively transmit light to photosensors located behind one or more filters. The one or more filters can, for example, but not limited to, selectively transmit one or more specific wavelengths of light or a specific polarization. Also, the one or more filters can selectively block one or more specific wavelengths of light or a specific polarization. The photosensors behind the one or more filters can be configured to be used as different sensor types. For example, but not limited to, an infrared pass filter that allows only infrared light to pass through 1102 can cover one or more sensor elements within the array 1101, and the other sensor elements may not be filtered or may be an infrared cut filter 1103. In another alternative implementation, the one or more filters can, for example, but not limited to, be an optical notch filter, which allows only a specific wavelength of light to pass through 1102, and the other filters can block that specific wavelength but allow other wavelengths to pass through 1103. This filtering may reduce the likelihood of false light source detection by enabling the use of the wavelength of light from the illuminator for DTOF and the specific wavelength for the DVS photosensor. Here, the sensor element can be any type of DVS photosensor or any type of camera pixel.

[0052] The patterned sensor or filter element can be incorporated into a hybrid imaging unit, as shown, for example, in FIGS. 11C and 11D. FIG. 11C shows an example of a hybrid DVS imaging unit 1112C having one or more lens elements 1114, an optional microlens array 1116, a bandpass filter 1118, and a patterned hybrid sensor array 1120. The sensor array includes DVS sensor elements 1122 and conventional imaging sensor elements 1124 that can be arranged in a pattern, as shown, for example, in FIGS. 11A or 11B. Such imaging units can be used with an illumination unit (not shown) within a DTOF sensor. The hybrid DVS imaging unit 1112D shown in FIG. 11D has a DVS sensor array 1126 and a patterned filter element 1118 that includes bandpass regions 1118A and bandcut regions 1118B arranged in a pattern, as shown, for example, in FIGS. 11A or 11B.

[0053] FIG. 12 is a diagram illustrating a hybrid DVS including a plurality of sensor types having inputs separated by an optical splitter according to an aspect of the present disclosure. In this implementation, the optical splitter 1204 can filter light based on wavelength, and the array 1201 can include a plurality of sensor element types physically separated based on the wavelength or polarization of the light for which detection is desired. For example, without limitation, unpolarized white light 1205 (a mixture of at least all visible light wavelengths and, in most cases, some infrared wavelengths) can enter the optical splitter 1204, which can be, for example, a dispersive prism, a diffraction grating, or a dichroic mirror. As shown, the white light 1205 entering the optical splitter 1204 can be separated by wavelength or polarization, with some light 1206 incident on a first portion of the array 1202 and other light 1207 incident on a second portion of the array 1203. Here, the optical splitter 1204 can be considered a filter that changes the diffraction angle based on wavelength or polarization. The array shown depicts the array as a single unit with a thick separation line 1203, but aspects of the present disclosure are not so limited. The first portion of the array 1201 may have a separation of up to 1 millimeter between the second portion of the array 1202, and while the array is shown as being separated vertically, other implementations may have portions separated horizontally, diagonally, or circumferentially. Further, the optical splitter here can be combined with different filtering or different sensor configurations, such as those shown in FIGS. 11A and 11B, to provide additional optical wavelength separation for different sensor types.

[0054] FIG. 13 is a diagram illustrating a hybrid DVS having inputs separated by a plurality of sensor types by a microelectromechanical (MEMS) mirror. Here, the MEMS mirror 1304 can vibrate between different positions at a set time so as to reflect light to the first part 1301 or the second part 1302 of the array according to the time when the light 1305 reaches the MEMS mirror 1304. The first part 1301 and the second part 1302 of the array can be physically separated from each other at 1303 based on the incident angle of the light reflected by the MEMS mirror 1304. In this way, light can be temporarily filtered between different sensor types. Such temporary filtering can be timed according to a known blinking pattern of a light source on a controller or a headset. For example, but not limited to, the light source on the controller or the headset can be turned on at regular intervals for a predetermined duration. For example, the light source can be turned on for 60 microseconds every 100 microseconds. In such a case, the MEMS mirror 1304 can be synchronized to reflect the light 1305 to the DVS part of the array 1301 for more than 60 microseconds every 100 microseconds or less to capture the change in the light source. Other reflected light 1307 is detected by the second part of the array 1302 during the time when the light source is off, and ambient light can be captured for image tracking or DTOF.

[0055] Alternatively, the MEMS mirror 1304 can filter the light 1305 based on the wavelength. The MEMS mirror in these implementations can be, for example, but not limited to, a MEMS Fabry-Perot filter or a diffraction grating. The MEMS mirror can diffract light in the first wavelength range 1306 to at least the first part 1301 of the array or light in the second wavelength range 1307 to the second part 1302 of the array.

[0056] FIG. 14 is a diagram illustrating a hybrid DVS having inputs where multiple sensor types are temporarily filtered. Here, the filter enables selectively passing light through the array based on time. In some implementations, the filter may be an optical waveguide. In a first time step, the filter can pass light of a first wavelength or polarization through array 1401 and block light of a second wavelength or polarization, or other wavelengths or polarizations. In a second time step, the filter enables a second wavelength or polarization 1402 but can block light of the first wavelength or polarization or other wavelengths of polarization. In this way, light can be temporarily filtered, which can be useful for tracking with different sensor types. For example, but not limited to, a light source coupled to a controller or headset may be infrared light or may have a specific wavelength. The light source can be configured to turn on and off at specific intervals. The specific intervals at which the light source turns on and off can be a sequence or a coded pattern. The temporary filtering can be activated to enable passing a specific wavelength through the sensor and blocking other wavelengths during specific intervals. Further, the switching interval of the filtering may be longer than the specific interval of the light source, taking into account the propagation time of light to the sensor.

[0057] Body tracking FIG. 15 is an illustration showing body tracking by a DVS according to an aspect of the present disclosure. As shown, user 1501 can wear a headset 1504 having a DVS 1503. Here, the DVS is shown with two arrays or DVS units or cameras and the DVS. User 1501 holds two controllers 1502 having two or more light sources 1505. DVS 1503 has a controller 1502 with a light source 1505 corresponding to within the field of view (FOV). Further, DVS 1503 can have a user's appendage such as a user's hand or arm 1507, or a leg or foot 1508 within its FOV. The DVS can also have a ground or other landmark 1509 within its FOV. For example, but not limited to, the photosensitive element of the DVS can detect the reflection of light corresponding to the user's appendage or the ground or other landmark when the user moves or the light changes. The camera detects the reflection of light from the field of view at the frame rate of the camera. From the detected light reflection, at 1506, the user's appendage or the ground or landmark can be determined. In some alternative implementations, an external DVS 1510 can be used to track the user's appendage. The external DVS 1510 may be at a distance away from the user selected such that the user's appendage fits within the field of view of the external DVS 1510. For example, but not limited to, the external DVS can be located on or under the top surface of a television or computer monitor, or other stand-alone or wall-mounted display.

[0058] The machine learning algorithm is trained to determine the user's body, appendages, ground or landmarks, and their relative positions and orientations from data such as events or frames. The machine learning algorithm may be a neural network, and the training may be similar to the method described in the section on training of the general neural network of FIGS. 10A-10D above. The neural network can be trained using a training set that includes labeled events or frames, or both. The labeled events or frames, or both, can include, for example, but not limited to, labels of the user's body, appendages, ground, landmarks, and the relative positions and orientations of the user's body, appendages, ground, and landmarks. In addition to determining the position and orientation of the controller 1502, the position and orientation of the user's body, appendages, ground or landmarks, and their relative positions and orientations can be determined, for example, using SLAM. Alternatively, the position and orientation of the controller can be determined using the same trained neural network that determines the labels of the user's body, appendages, ground or landmarks and their relative positions and orientations. As indicated by element 1506, a model can be fitted to the determined user's body and appendages to improve the determination of relative positions and locations.

[0059] Safety shutter As described above, in combination with the determination of the position and orientation of the controller, body tracking can be used to trigger the safety shutter on a VR or AR headset. FIG. 16A shows a headset with a safety shutter door according to an aspect of the present disclosure. The headset 1601 may include a head strap 1604, eye pieces 1603, and a display screen 1602. The eye pieces 1603 may include one or more lenses configured to focus on the display screen 1602. The one or more lenses may be, for example, Fresnel lenses or prescription lenses. The display screen 1602 may be transparent, and the hole 1607 within the headset body may enable vision through the display screen 1602 when the safety shutter is open.

[0060] In this implementation, the safety shutter 1606 is a door that swings away from the hole 1607 when the safety system becomes active. Here, the system-operated clasp 1605 interacts with a clasp 1608 on the safety shutter door 1606 to close and secure the door over the hole 1607. The spring-loaded hinge 1609 can ensure that the safety shutter door 1606 opens quickly when the system-operated clasp 1605 opens. The spring-loaded hinge may have, for example, but not limited to, a clock spring wound around the hinge and fixed to the door, and the spring is wound when the door closes and unwound when the door opens. Alternatively, a leaf spring may push the safety shutter door 1606 when closing the door.

[0061] The system-operated clasp 1605 can be configured to open the clasp when the safety system becomes active. The safety system-operated clasp 1605 can include an electric motor or a linear actuator that moves the clasp. The safety system may be activated when a ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendage. The safety system can use the determination of the user's body, appendage, ground or landmark, and their relative positions and orientations as described above. When the safety system becomes active, a signal to open the clasp can be sent to the safety system-operated clasp 1605. When the clasp opens, the spring-loaded hinge 1609 pushes open the safety shutter door 1606, allowing the user to see through the display 1602 and avoid the risk of activating the safety system.

[0062] Other implementations of the safety shutter may be used. For example, FIG. 16B shows an alternative headset with a sliding safety shutter according to an aspect of the present disclosure. Here, when the safety system becomes active, the safety shutter slide 1616 slides so as not to obstruct the hole 1607. The safety shutter slide 1616 can move on a spring-loaded rail 1619. Alternatively, the sliding safety shutter 1616 can include a tab that is inserted into a slot 1619 within the headset body, and a spring within the slot can also push the sliding safety shutter. In some additional alternative implementations, the spring may be omitted and the sliding safety shutter 1616 may operate by gravity. The spring-loaded rail 1619 can push the shutter when closing the sliding safety shutter 1616 and ensure that the safety shutter opens quickly when the safety system becomes active. The system-operated clasp 1605 interacts with a clasp 1618 on the sliding safety shutter 1616 to close and secure the slide.

[0063] If the ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendage, the safety system may be activated. As described above, the safety system can use the determination of the user's body, appendage, ground or landmark, and their relative positions and orientations. When the safety system is activated, a signal to open the clasp can be sent to the system-operated clasp 1605. When the clasp opens, the spring 1619 pushes open the sliding safety shutter 1616, allowing the user to see through the display 1602 and avoid the risk of activating the safety system.

[0064] Figure 16C shows a headset with a louvered safety shutter according to an aspect of the present disclosure. In this implementation, the slats of the louvered safety shutter 1626 are longer in a first dimension than in a second dimension. In the closed position, the longitudinal dimension of the slats is substantially parallel to the optical system, and each slat overlaps either another slat or the headset body, blocking light passing through the display screen 1602. In the open position, the slats change position so that the short dimension is parallel to the optical system 1603, allowing light to pass through the slats and reach the display screen. The actuator rod 1628 can hinge each of the slats 1626. The safety system-controlled actuator 1629 can push or pull the actuator rod 1628 to open the louvered safety shutter when the safety system is activated. In some other implementations, the safety system-controlled actuator may be spring-loaded, and the clasp connected to the actuator rod 1628 and the system-controlled actuator 1629 may include a clasp that interlocks with the clasp of the actuator rod. When the safety system is activated, the clasp of the system-controlled actuator can open, allowing the actuator rod to move under spring pressure and the louvers to open.

[0065] If the ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendage, the safety system may be activated. As described above, the safety system can use the determination of the user's body, appendage, ground or landmark, and their relative positions and orientations. When activated, a signal to move the actuator rod may be sent to the system-operated actuator. When the actuator rod 1629 moves, it pushes open the slats of the louvered safety shutter 1626, enabling the user to see through the display 1602 and avoid the risk of activating the safety system.

[0066] Figure 16D shows a fabric safety shutter according to an aspect of the present disclosure. In this implementation, the safety shutter 1636 is composed of an opaque fabric, for example, but not limited to, a densely woven cotton fabric, polyester, vinyl, or a densely woven wool fabric. The fabric safety shutter 1636 may be coupled to a fabric roller 1639. The fabric roller 1639 may be configured to wind up the fabric shutter when the safety system is activated. The fabric roller may be, for example, but not limited to, spring-loaded using a clock spring, so that the clock spring is tensioned when the safety shutter closes, or alternatively, an electric motor may be used to wind up the fabric safety shutter. The system-operated clasp 1605 is interlocked with a clasp 1638 coupled to the fabric safety shutter 1636 to ensure that the fabric safety shutter does not open unintentionally.

[0067] During operation, the fabric safety shutter may be in the closed position. If a ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendage, the safety system may be activated. The safety system can use, as described above, the determination of the user's body, appendage, ground or landmark, and their relative positions and orientations. When the safety system is activated, a signal to open the clasp can be sent to the system-operated clasp 1605. When the clasp opens the fabric roller 1619, the fabric safety shutter 1616 is wound up, enabling the user to view through the display 1602 and avoid the risk of activating the safety system.

[0068] Figure 16E shows a headset with a liquid crystal safety shutter according to an aspect of the present disclosure. In this implementation, the liquid crystal screen 1646 is integrated into the headset 1601 and otherwise covers a hole in the headset. The safety system-controlled liquid crystal screen driver 1649 is communicatively coupled to the liquid crystal screen 1646.

[0069] The liquid crystal screen 1646 may be, for example, but not limited to, a liquid crystal shutter having a first polarizer and a second polarizer. The first polarizer has a polarization difference of 90 degrees from the second polarizer and has a cavity filled with a fluid. The cavity filled with the fluid can contain liquid crystals, and these liquid crystals are configured to have a first orientation when there is no electric field that changes the polarization to allow light to pass from the first polarizer to the second polarizer. The liquid crystals can be further configured to align to a second orientation under an electric field. Since the second orientation of the liquid crystals does not change the polarization of the light, the light passing through the first polarizer is blocked by the second polarizer. Electrodes may be arranged along the surface of the cavity filled with the liquid to enable control of the liquid crystals. If the safety system control type liquid crystal screen driver can be communicatively coupled to these electrodes, the electrodes can control the liquid crystals in the cavity filled with the fluid. As used herein, being communicatively coupled means that an electrical signal representing a message or instruction can be transmitted and / or received from one coupling element to the other coupling element, and these signals can pass through intermediate elements, and although their formats may change, the messages contained therein remain unchanged.

[0070] During operation, the safety system control driver 1649 can send a signal to the liquid crystal safety shutter 1646 to make it opaque while the display screen 1602 is active. When the safety system becomes active, the driver 1649 can make the liquid crystal safety shutter 1646 transparent. For example, but not limited to, if the driver can reduce the voltage supplied to the liquid crystal safety shutter and return the liquid crystal to their initial orientation, the polarization of light changes, allowing light to pass through the second polarizer. The safety system may be activated when a ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendages. The safety system can use the determination of the user's body, appendages, ground or landmark, and their relative positions and orientations as described above.

[0071] Tracking of finger position Aspects of the present disclosure can be applied to finger tracking. FIG. 17 shows finger tracking by a DVS and a controller according to an aspect of the present disclosure. Here, the controller 1701 includes two or more light sources and one or more buttons 1705. The two or more light sources include one or more light sources 1706 proximate to the one or more buttons 1705 and two or more other tracking light sources 1703. As described above, the two or more other tracking light sources 1703 can generate events at the DVS 1702 for determining the position and orientation of the controller 1701. The DVS having two DVSs 1702, or two arrays and three other tracking light sources 1703 shown here is used for determining the position and orientation of the controller.

[0072] One or more light sources 1706 proximate to one or more buttons 1705 can be used for finger tracking. For example, without limitation, finger tracking can be achieved at the DVS 1702 by using the occlusion of one or more light sources 1706 proximate to the button 1705. One or more light sources 1706 proximate to the button can be turned off and on at a predetermined interval. The DVS 1702 can generate an event for each blink. The events can be analyzed to determine the occlusion of one or more light sources 1706 proximate to the button 1705. By knowing the configuration of one or more light sources proximate to the button, for example, without limitation, when occluded by a finger or palm, the pattern of light detected in the events generated by the DVS will be different from when the light source is not occluded. As described above with respect to determining the position and orientation of the controller, here the timing of the blinking is used to determine which light is occluded, and thus the position of the corresponding finger or palm can be determined. If the light source 1706 proximate to the button 1705 has a reduced intensity or no intensity during the interval, it can be determined that the light source proximate to the button is "on". The light source is determined to be occluded. Similarly, when the detected intensity of the light known to be "on" changes, an event can be generated, and from that event, it may be determined that the user's finger or hand has moved and the button has been left incomplete. The occlusion of one or more of the light sources proximate to one or more buttons can be correlated to the position of the finger or palm based on their positions. For example, without limitation, the position of the user's hand 1704 can be determined using a light source located near the user's palm when the controller 1701 is held.

[0073] The light source 1706 can be positioned around each button 1705 and can use the button configuration and design of the controller to determine the finger position. For example, but not limited to, the controller 1701 may be designed such that each finger of the user is placed near the button 1705 when held. Next, the occlusion pattern of the light source determined from the DVS event may be used to determine when the user's finger is hovering in the air over an inactive button, and moreover, may be used to determine when the user's finger has passed over the button. This can help provide additional interaction options for the user, such as a half-press of the button, or an insufficient press, or other button options. By having multiple light sources surround each button, it becomes possible to refine the determination of the position of the user's finger or palm. For example, but not limited to, in some implementations, 10 or more light sources may surround each button, and in other implementations, a single light source may shine light through a translucent diffuser around the button and use interruptions in the diffused light profile to determine the finger position.

[0074] In some implementations, the button 1705 itself can also be a light source. One or more buttons 1705 can be turned off and on at different intervals from one or more light sources 1706 or other light tracking light sources 1703 proximate to the button. Alternatively, the light source of the button 1705 may have a different wavelength or polarization from one or more light sources 1706 or other light tracking light sources 1703 proximate thereto.

[0075] Also, the power saving mode of the light source can be enabled using buttons and tracking. For example, if it is determined that a button is pressed, one or more light sources proximate to the button may be dimmed or turned off. Further, if the controller 1701 determines that it is outside the field of view of the DVS 1702, one or more lights proximate to the button may be dimmed or turned off. In some implementations, data from the IMU can be used to determine whether the controller is being held by the user. For example, but not limited to, if no changes in IMU data such as acceleration, angular velocity, etc. are detected over a threshold period, the light source may be dimmed or turned off. When a change in the IMU data is detected, the light source can be turned back on.

[0076] In an alternative implementation, finger tracking can be performed without using one or more light sources. The machine learning model can be trained using a machine learning algorithm to detect the position of a finger from events generated from changes in ambient light due to the movement of the finger. The machine learning model may be a general machine learning model such as the CNN, RNN, or DNN described above. In some implementations, for example, but not limited to, a special machine learning model such as a spiking (or spark) neural network (SNN) can be trained with a special machine learning algorithm. The SNN mimics biological NNs by having activation thresholds and weights, which are adjusted according to the relative spike times within an interval, also called spike-timing-dependent plasticity (STDP). When the activation threshold is reached, the SNN is said to spike and send its weights to the next layer. The SNN can be trained via STDP and supervised or unsupervised learning techniques. Further information on SNNs can be found in Tavanei, Amirhossein et al. "Deep Learning in Spiking Neural Networks" Neural Networks (2018) arXiv:1804.08150, the content of which is hereby incorporated by reference for all purposes.

[0077] Alternatively, an HDR (high dynamic range) image may be constructed using events aggregated from ambient data. The machine learning model is trained to recognize the position or the position and orientation of the hand or the controller from the HDR image. The trained machine learning model is applied to the HDR image generated from the events to determine the position or the position and orientation of the hand / finger or the controller. The machine learning model may be a general machine learning model trained with a supervised learning technique as described in the section on training of general neural networks.

[0078] Visual target tracking Aspects of the present disclosure can be applied to gaze tracking. Generally, gaze tracking image analysis determines the gaze direction from an image by utilizing characteristics specific to how light is reflected from the eyes. For example, an image can be analyzed to identify the position of the eyes based on corneal reflections in the image data, and further analyzed to determine the gaze direction based on the relative position of the pupils in the image.

[0079] Two common gaze tracking techniques for determining the gaze direction based on the position of the pupils are known as the bright pupil method and the dark pupil method. The bright pupil method involves illuminating the eyes with a light source that is substantially aligned with the optical axis of the DVS, which reflects the emitted light off the retina and back to the DVS through the pupil. The pupil appears in the image as a distinguishable bright spot at the position of the pupil, similar to the red-eye effect that occurs in an image during conventional flash photography. In this gaze tracking method, when the contrast between the pupil and the iris is not sufficient, the bright reflection from the pupil itself helps the system to locate the pupil.

[0080] The dark pupil method involves illuminating with a light source that is substantially offset from the optical axis of the DVS, which reflects the light directed through the pupil away from the optical axis of the DVS, resulting in a distinguishable dark spot at the position of the pupil for the event. In another dark pupil method system, an infrared light source and a camera directed at the eyes can see the corneal reflection. Such a DVS-based system tracks the positions of the reflections of the pupil and the cornea, thereby obtaining a parallax due to the different depths of the reflections and improving the accuracy.

[0081] FIG. 18A shows an example of a dark pupil eye tracking system 1800 that may be used in the context of the present disclosure. The eye tracking system tracks the orientation of the user's eye E with respect to a display screen 1801 on which a visible image is presented. Although a display screen is utilized in the exemplary system of FIG. 18A, certain alternative embodiments may utilize an image projection system that can project an image directly onto the user's eye. In these embodiments, the user's eye E is tracked with respect to the image projected onto the user's eye. In the example of FIG. 18A, the eye E collects light from the screen 1801 through a variable iris I; the lens L projects an image onto the retina R. The opening of the iris is known as the pupil. Muscles control the rotation of the eye E in response to nerve impulses from the brain. The upper and lower eyelid muscles ULM, LLM control the upper and lower eyelids UL, LL, respectively, in response to other nerve impulses.

[0082] The photosensitive cells on the retina R generate electrical impulses that are sent via the optic nerve ON to the user's brain (not shown). The visual field of the brain interprets the impulses. Not all parts of the retina R have equal photosensitivity. Specifically, the photosensitive cells are concentrated in a region known as the fovea.

[0083] The illustrated image tracking system includes one or more infrared light sources 1802, for example, light emitting diodes (LEDs) that direct non-visible light (e.g., infrared light) towards the eye E. A portion of the non-visible light is reflected by the cornea C of the eye and a portion is reflected by the iris. The reflected non-visible light is directed by a wavelength selective mirror 1806 towards a DVS 1804 that is sensitive to infrared light. The mirror transmits visible light from the screen 1801 but reflects non-visible light reflected from the eye.

[0084] The DVS 1804 generates events of the eye E that can be analyzed to determine the line of sight direction GD from the relative position of the pupil. This event can be generated by a processor 1805. The DVS 1804 is advantageous in this implementation because the very high update rate of the events provides near real-time information regarding changes in the user's line of sight.

[0085] As can be seen in FIG. 18B, by analyzing event 1811 indicating the user's head H, the line-of-sight direction GD can be determined from the relative position of the pupil. For example, the analysis may determine the two-dimensional offset of the pupil P from the center of the eye E in the image. The position of the pupil relative to the center can be converted into the line-of-sight direction with respect to the screen 1801 by simple geometric calculations of a three-dimensional vector based on the known size and shape of the eyeball. The determined line-of-sight direction GD can indicate the rotation and acceleration of the eye E when the eye E moves with respect to the screen 1801.

[0086] As can also be seen in FIG. 18B, the event may include reflections 1807 and 1808 of non-visible light from the cornea C and the lens L, respectively. Since the depths of the cornea and the lens are different, the parallax and refractive index between the reflections can be used to improve the accuracy in determining the line-of-sight direction GD. An example of this type of eye gaze tracking system is a dual Purkinje tracker, where the corneal reflection is the first Purkinje image and the lens reflection is the fourth Purkinje image. When the user wears glasses, a reflection 1808 from the user's glasses 1809 may also exist.

[0087] The performance of the eye gaze tracking system depends on a number of factors including the arrangement of the light source (IR, visible light, etc.) and the DVS, whether the user wears glasses or contacts, the headset optics, the latency of the tracking system, the speed of eye movement, the shape of the eye (which may change during the day or as a result of movement), the state of the eye, such as amblyopia, the stability of the line of sight, fixation on a moving object, the scene presented to the user, as well as the movement of the user's head. The DVS reduces the output of extra information to the processor and provides a very high update rate for events. This enables faster processing and faster determination of the state of eye gaze tracking and error parameters.

[0088] Error parameters that can be determined from the eye-tracking data can include, but are not limited to, rotational speed and prediction error, fixation error, confidence intervals regarding current and / or future eye positions, and smooth pursuit error. State information regarding the user's eye includes the individual state of the user's eye and / or gaze. Thus, exemplary state parameters that can be determined from the eye-tracking data can include, but are not limited to, blink metrics, saccade metrics, depth of field response, color vision anomalies, gaze stability, and eye movements as precursors to head movements.

[0089] In certain implementations, the eye-tracking error parameter can include a confidence interval regarding the current eye position. The confidence interval can be determined by examining whether there is a change in the rotational speed and acceleration of the user's eye from the last position. In an alternative embodiment, the eye-tracking error and / or state parameter can include a prediction of a future eye position. The future eye position can be determined by examining the rotational speed and acceleration of the eye and extrapolating the possible future positions of the user's eye. Generally speaking, the DVS update rate of the eye-tracking system can result in a smaller error between the determined future position and the actual future position for users with high rotational speed and acceleration values because the DVS update rate is very high, and this small error can potentially be significantly smaller than that of existing camera-based systems.

[0090] In yet another alternative implementation, the gaze tracking error parameter can include a measurement of the eye velocity, such as the rotational velocity. In certain alternative embodiments, determining the gaze tracking state parameter includes measuring a measurement criterion for the user's blink. During a normal blink, typically a period of 150 milliseconds (ms) elapses and the user's vision is not focused on the presented image. Thus, depending on the frame rate of the display device, the user's vision may not be focused on the presented image for up to 20 - 30 frames. However, when the blink ends, the user's line of sight direction may not correspond to the last measured line of sight direction determined by the acquired gaze tracking data. Thus, the measurement criterion for the user's line of sight can be determined from the acquired gaze tracking data. These measurement criteria can include, but are not limited to, the measured start time and end time of the user's blink, as well as the predicted end time.

[0091] In yet an additional alternative implementation, determining the gaze tracking state parameter includes measuring a measurement criterion for the user's saccade. During a normal saccade, typically a period of 20 - 200 ms elapses and the user's vision is not focused on the presented image. Thus, depending on the frame rate of the display device, the user's vision may not be focused on the presented image for up to 40 frames. However, due to the nature of the saccade, when the saccade ends, the user's line of sight direction moves to another region of interest. Thus, the gaze tracking data can be used to establish the measurement criterion for the user's saccade based on the actual or predicted time elapsed during the saccade. These measurement criteria can include, but are not limited to, the measured start time and end time of the user's saccade, as well as the predicted end time.

[0092] In certain alternative embodiments, determining the gaze tracking state parameter includes determining a transition in the user's line of sight direction between regions of interest as a result of a change in depth of field between the presented images. This is because when a transition between regions of interest of the presented images is provided, the user will experience a saccade.

[0093] In further additional alternative implementations, the determined gaze-tracking state parameters can be adapted to color vision deficiencies. For example, regions of interest may be present in an image presented to a user such that these regions are not recognized by a user having a particular form of color vision deficiency. The resulting gaze-tracking data determines, for example, whether the user's gaze has identified or responded to a region of interest as a result of a change in the user's gaze direction. Thus, as a gaze-tracking error parameter, it can be determined whether a user is color vision deficient with respect to a particular color or spectrum.

[0094] In a particular alternative implementation, the determined gaze-tracking state parameters include a measure of the user's gaze stability. Determining gaze stability can be done by measuring the radius of the user's eye microsaccades, and the smaller the overshoot and undershoot of fixation, the more stable the user's gaze.

[0095] In further additional alternative implementations, the determined gaze-tracking error and / or state parameters include the user's ability to fixate on a moving object. These parameters may include a measure of the user's eye's ability to perform smooth pursuit and the maximum object-tracking speed of the eyeball. Typically, the jitter of eye movements experienced by a user with excellent smooth pursuit ability is reduced.

[0096] In a particular alternative implementation, the determined gaze-tracking error and / or state parameters include determining eye movements as precursors of head movements. The offset between the head and eye orientations can affect certain error and / or state parameters, such as in smooth pursuit or fixation, as described above.

[0097] Further information regarding the determination of gaze-tracking and error parameters can be found in U.S. Patent No. 10,192,528, the content of which is hereby incorporated by reference for all purposes.

[0098] System FIG. 19 is a block system diagram of a tracking system by DVS according to an aspect of the present disclosure. By way of example and not limitation, according to an aspect of the present disclosure, the system 1900 may be an embedded system, a mobile phone, a personal computer, a tablet computer, a portable game device, a workstation, a game console, etc.

[0099] The system 1900 generally includes a central processing unit (CPU) 1903 and a memory 1904. The system 1900 may also include a well-known support function 1906 that can communicate with other components of the system via, for example, a data bus 1905. Such support functions may include, but are not limited to, input / output (I / O) elements 1907, a power supply (P / S) 1911, a clock (CLK) 1912, and a cache 1913.

[0100] The system 1900 may include a display device 1931 for presenting rendered graphics to a user. In an alternative implementation, the display device is a separate component that functions in conjunction with the system 1900. The display device 1931 may be in the form of a flat panel display, a head-mounted display (HMD), a cathode ray tube (CRT) screen, a projector, or other device capable of displaying visible text, numbers, graphic symbols, or images.

[0101] Here, the display device 1931 is coupled to the DVS 1901A, and the controller 1902 includes two or more light sources 1932A that may be in any configuration described herein. In an alternative implementation, the DVS can be coupled to a game controller, and the display device can alternatively include two or more light sources. In yet another alternative implementation, the DVS is a separate unit decoupled from either the display device or the controller, in which case both the controller and the display device may include two or more light sources for tracking.

[0102] In some implementations where the display device is part of a head-mounted display (HMD), such an HMD can include an inertial measurement unit (IMU), such as an accelerometer or a gyroscope. Also, as discussed above in this specification, such an HMD can include light sources 1932B, and these light sources can be tracked using a DVS that is separate from the display device 1901 and coupled to the CPU 1903. As an example, a separate DVS 1901B can be attached to the controller 1902.

[0103] In some implementations, DVS 1901A or DVS 1901B can be part of a hybrid sensor, as described above with respect to FIGS. 11A, 11B, 11C, or 11D. Such a hybrid sensor can include a depth sensor, such as a DTOF sensor, and in this case, the hybrid sensor can include an illumination unit (not shown).

[0104] Furthermore, when the display device 1931 is part of an HMD, an optional safety shutter 1933 can be externally fitted to this device, which can be operably coupled to a processor such as the CPU 1903 and operate as described above with respect to FIGS. 16A, 16B, 16C, 16D, or 16E. Alternatively, the safety shutter can be controlled by a separate processor attached to the HMD.

[0105] Also, the system 1900 includes a mass storage device 1915, such as a disk drive, CD-ROM drive, flash memory, solid state drive (SSD), tape drive, etc., to provide non-volatile storage of programs and / or data. The system 1900 may also optionally include a user interface unit 1916 to facilitate interaction between the system 1900 and the user. The user interface 1916 may include a keyboard, mouse, joystick, light pen, or other device that can be used with a graphical user interface (GUI). The system 1900 may also include a network interface 1914 to enable the device to communicate with other devices through a network 1920. The network 1920 may be, for example, a local area network (LAN), a wide area network such as the Internet, a personal area network such as a Bluetooth® network, or other type of network. These components may be implemented in hardware, software, or firmware, or any combination of two or more of these.

[0106] Each of the CPUs 1903 may include one or more processor cores, for example, a single core, two cores, four cores, eight cores, or more. In some implementations, the CPU 1903 may include GPU cores or multiple cores of the same APU (Accelerated Processing Unit). The memory 1904 may be in the form of an integrated circuit that provides addressable memory, such as random access memory (RAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), etc. The main memory 1904 may include application data 1923 used by the processor 1903 during processing. The main memory 1904 may also include event data 1909 received from the DVS 1901. As described in FIG. 9, the trained neural network (NN) 1910 can be loaded into the memory 1904 to determine the position and orientation data. Further, the memory 1904 may include a machine learning algorithm 1921 for training or tuning the NN 1910. A database 1922 can be included in the memory 1904. The database can include information regarding the light source configuration, the respective predetermined blinking intervals of one or more light sources, etc. The memory may also include the output from an IMU coupled to the controller 1902 or the display device 1931. In some implementations, the memory 1904 can include the output from one or more light sources, such as a timestamp when each of the light sources is on or off.

[0107] According to aspects of the present disclosure, processor 1903 can execute methods for determining the position and orientation of a controller or user, as described with respect to FIGS. 8 and 9, and these methods can be loaded into memory 1904 as application 1923. As a result of the processor executing the methods described with respect to FIGS. 8 and 9 and further described with respect to FIG. 15, the processor can generate the orientation and configuration of one or more of a controller, a headset, a user's body or appendage, the ground, or a landmark. These positions and orientations can be stored in database 1922 and can be used for successive iterations of the methods of FIGS. 8 and 9. In some implementations, processor 1903 can utilize such positions and / or orientations in a machine learning algorithm trained to perform SLAM (simultaneous localization and mapping).

[0108] Mass storage 1915 can include applications or programs 1917 that are loaded into main memory 1904 when processing is initiated in application 1923. Further, mass storage 1915 can include data 1918 used by the processor during processing of application 1923, NN 1910, machine learning algorithm 1921, and while populating database 1922.

[0109] As used herein and as generally understood by those skilled in the art, an application specific integrated circuit (ASIC) is an integrated circuit that is customized for a particular use, rather than for general use.

[0110] As used herein and as generally understood by those skilled in the art, a field programmable gate array (FPGA) is an integrated circuit designed to be configured by a customer or designer after manufacture - and thus "field programmable". FPGA configurations are generally specified using a hardware description language (HDL) similar to that used for ASICs.

[0111] As used herein and as generally understood by those skilled in the art, a system on a chip, i.e., a system-on-chip (SoC or SOC), is an integrated circuit (IC) that integrates all components of a computer or other electronic system onto a single chip. This can include digital, analog, mixed-signal, and often radio-frequency functions - all on a single chip substrate. Typical applications are in the field of embedded systems.

[0112] Common SoCs include the following hardware components: One or more processor cores (e.g., a microcontroller, microprocessor, or digital signal processor (DSP) core). Memory blocks, such as read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), and flash memory. A timing source, such as an oscillator or a phase-locked loop. Peripherals, such as a counter timer, real-time timer, or power-on reset generator. External interfaces, such as industry-standard specifications like the Universal Serial Bus (USB), FireWire (registered trademark), Ethernet, USART (universal asynchronous receiver / transmitter), Serial Peripheral Interface (SPI) bus, etc. Analog interfaces (including analog-to-digital converters (ADC) and digital-to-analog converters (DAC)). Voltage regulators and power management circuits.

[0113] These components are connected either by a proprietary bus or an industry-standard bus. A direct memory access (DMA) controller improves the data throughput of the SoC by directly routing data between the external interface and memory, bypassing the processor core.

[0114] A typical SoC includes both the above-described hardware components and executable instructions (e.g., software or firmware) that control the processor core(s), peripheral devices, and interfaces.

[0115] Aspects of the present disclosure provide image-based tracking that features a higher sample rate than is possible with conventional image-based tracking systems, resulting in improved tracking fidelity. Further advantages include cost reduction, weight reduction, reduction of irrelevant data generation, and reduction of processing requirements when using a DVS-based tracking system. These advantages enable, among other applications, the improvement of virtual reality (VR) and augmented reality (AR) systems.

[0116] The foregoing is a complete description of preferred embodiments of the invention, but various alternatives, modifications, and equivalents may be used. Accordingly, the scope of the invention should not be determined with reference to the foregoing, but rather should be determined with reference to the appended claims, which are to be accorded the full scope of equivalents. Whether or not preferred, any feature described herein may be combined with any other feature described herein, whether or not preferred. In the following claims, the indefinite articles "A" or "An" refer to one or more of the items following the article, unless expressly stated otherwise. The appended claims are not to be construed as including means-plus-function limitations, unless such a limitation is expressly recited in a given claim using the phrase "means for".

Claims

1. A tracking system, a processor, two or more light sources configured to direct light towards one or more of the user's eyes from a relative position with respect to the one or more of the user's eyes, a dynamic vision sensor (DVS) operably coupled to the processor and configured to visually recognize the one or more of the user's eyes, the DVS having an array of photosensitive elements in a known configuration, the dynamic vision sensor being configured to output a signal corresponding to one or more events at one or more corresponding photosensitive elements within the array in response to a change in the light output from the two or more light sources reflected from a portion of the one or more of the user's eyes, the output signal having information corresponding to the position of the one or more corresponding photosensitive elements within the array, the DVS, comprising, the processor is configured to determine an association between each of the one or more events and a corresponding specific light source of the one or more light sources, and to use the determined association, the relative position of the one or more light sources with respect to the one or more of the user's eyes, and the position of the one or more corresponding photosensitive elements within the array to adapt the orientation of the one or more of the user's eyes. A system.

2. The system according to claim 1, wherein the one or more light sources and the DVS are attached to a headset configured to be worn on the user's head.

3. The system according to claim 1, wherein the one or more light sources are configured to turn on and off within a predetermined time sequence.

4. The system according to claim 1, wherein the output signal includes information corresponding to the time of the one or more events.

5. The system according to claim 1, wherein the one or more light sources are configured to turn on and off within a predetermined time sequence, and the output signal includes information corresponding to the time of the one or more events.

6. The one or more light sources are configured to turn on and off within a predetermined time sequence, the output signal includes information corresponding to the time of the one or more events, and the processor determines the association between each of the one or more events and the corresponding particular light source of the one or more light sources, and the determined association, the relative position of the one or more light sources with respect to the one or more of the user's eyes, the position of the one or more corresponding photosensitive elements in the array, the predetermined time sequence, and the information corresponding to the time of the one or more events are used to configure the one or more of the user's eyes to adapt the orientation. The system according to claim 1.

7. The system further includes an image sensor configured to visually recognize the one or more of the user's eyes and acquire one or more images of the one or more of the user's eyes, the image sensor is operably coupled to the processor, and the processor determines an association between each of the one or more events and the corresponding particular light source of the one or more light sources, and the one or more images, the determined association, the relative position of the one or more light sources with respect to the one or more of the user's eyes, and the position of the one or more corresponding photosensitive elements in the array are used to configure the one or more of the user's eyes to adapt the orientation. The system according to claim 1.

8. The system further includes an image sensor configured to visually recognize the user's face and acquire one or more images of the user's face, the image sensor is operably coupled to the processor, the DVS is configured to visually recognize the user's face, and the processor is configured to determine an additional association between one or more face characteristics of the user's face, the one or more images, and one or more additional events at one or more corresponding additional photosensitive elements in the array. The system according to claim 1.

9. The DVS and the image sensor are further configured to visually recognize the external environment of the user. The system according to claim 8.

10. When the user wears the headset, it further includes a display screen configured to be visible to the user, and the processor is configured to cause a virtual reality representation to be presented on the display screen. The DVS and the image sensor are further configured to view the external environment of the user. The system according to claim 8, wherein the processor is configured to use the events generated by the DVS according to the light from the environment to interpolate the image of the environment from the image sensor to generate an interpolated image.

11. When the user wears the headset, it further includes a display screen configured to be visible to the user. The DVS and the image sensor are further configured to view the external environment of the user. The processor is configured to use the events generated by the DVS according to the light from the environment to interpolate the image of the environment from the image sensor to generate an interpolated image of the environment. The system according to claim 8, wherein the processor is configured to present the interpolated image of the environment on the display screen.

12. When the user wears the headset, it further includes a display screen configured to be visible to the user, and the processor is configured to cause a virtual reality representation to be presented on the display screen. The DVS and the image sensor are further configured to view the external environment of the user. The processor is configured to use the events generated by the DVS according to the light from the environment to interpolate the image of the environment from the image sensor to generate an interpolated image of the environment. The system according to claim 8, wherein the processor presents the interpolated image of the environment on the display screen and is configured to merge the virtual reality representation with the presented interpolated image.

13. A tracking method, comprising: Receiving, in one or more corresponding photosensitive elements in an array of photosensitive elements, a signal corresponding to one or more events in response to a change in the light output from two or more light sources reflected from one or more portions of a user's eye, wherein the output signal includes information corresponding to the positions of the one or more corresponding photosensitive elements in the array, the array of photosensitive elements is part of a dynamic vision sensor (DVS), the array of photosensitive elements has a known configuration, and the one or more light sources are configured to direct light towards the one or more of the user's eyes from a relative position with respect to the one or more of the user's eyes, the receiving, Determining an association between each of the one or more events and a corresponding particular light source of the one or more light sources, Adapting the orientation of the one or more of the user's eyes using the determined association, the relative position of the one or more light sources with respect to the one or more of the user's eyes, and the position of the one or more corresponding photosensitive elements in the array, A method comprising.

14. Further comprising determining an association between each of the one or more events and the corresponding particular light source of the one or more light sources, and adapting the orientation of the one or more of the user's eyes using the determined association, the relative position of the one or more light sources with respect to the one or more of the user's eyes, the position of the one or more corresponding photosensitive elements in the array, a predetermined time sequence, and information corresponding to the time of the one or more events, wherein the one or more light sources are configured to be turned on and off within the predetermined time sequence, and the output signal includes information corresponding to the time of the one or more events, the method according to claim 13.

15. Obtaining one or more images of the user's eyes using an image sensor configured to visually recognize the eyes of the one or more users; determining an association between each of the one or more events and a corresponding specific light source of the one or more light sources; and using the one or more images, the determined association, the relative positions of the one or more light sources with respect to the one or more of the user's eyes, and the positions of the one or more corresponding photosensitive elements in the array to adapt the orientation of the one or more of the user's eyes. The method according to claim 13, further comprising.

16. Obtaining one or more images of the user's face by an image sensor configured to visually recognize the user's face, wherein the DVS is configured to visually recognize the user's face, said obtaining; Determining an additional association between the one or more images and one or more additional events at one or more corresponding additional photosensitive elements in the array with respect to one or more facial characteristics of the user's face; The method according to claim 13, further comprising.

17. When the user wears a headset, displaying a virtual reality representation on a display screen configured to be visually recognized by the user; Interpolating an image of the user's external environment from the image sensor using events generated by the DVS in response to light from the environment to generate an interpolated image, wherein the DVS and the image sensor are further configured to visually recognize the user's external environment, said generating; The method according to claim 16, further comprising.

18. Interpolating an image of the user's external environment from the image sensor using events generated by the DVS in response to light from the environment to generate an interpolated image of the environment, wherein the DVS and the image sensor are further configured to visually recognize the user's external environment, said generating; Presenting the interpolated image of the environment on a display screen configured to be visually recognized by the user when the user wears a headset; The method according to claim 16, further comprising.

19. Presenting a virtual reality representation on a display screen configured to be visible to the user when the user wears the headset; Interpolating an image of the user's external environment from the image sensor using events generated by the DVS according to the light from the environment to generate an interpolated image of the environment, wherein the DVS and the image sensor are further configured to view the external environment of the user, said generating; Presenting the interpolated image of the environment on the display screen and merging the virtual reality representation with the presented interpolated image; The method according to claim 16, further comprising.

Citation Information

Patent Citations

  • Detection device and method, and program

    JP2016002353A

  • Display unit, learning device, and method for controlling display unit

    JP2020077271A

  • Information processing apparatus, information processing method, and program

    JP2021060627A

  • Calibration of sight-line detection device

    WO2021153577A1