Dynamic Visual Sensor Tracking Based on Light Source Occlusion

By employing DVS with single or dual sensor arrays and IMUs, the challenges of tracking game controllers in VR and AR systems are addressed, achieving accurate and efficient tracking with reduced hardware costs and complexity.

JP2025516617AActive Publication Date: 2025-05-30SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024566450
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-10
Filing Date
2023-04-12
Publication Date
2025-05-30
Estimated Expiration
2043-04-12

AI Technical Summary

Technical Problem

Existing VR and AR systems face challenges in accurately tracking game controllers, especially in dynamic environments where absolute light positions are undesirable and high frame rates are required for smooth motion feedback, leading to increased hardware costs and complexity.

Method used

The use of Dynamic Vision Sensors (DVS) with single or dual sensor arrays, combined with cameras and inertial measurement units (IMUs), to track game controllers by detecting changes in light intensity and adapting to the configuration of light sources, thereby reducing the need for absolute light positions and high frame rates.

Benefits of technology

This approach enables accurate and efficient tracking of game controllers with reduced hardware costs and complexity, while providing smooth motion feedback and adaptability to dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516617000001_ABST
    Figure 2025516617000001_ABST
Patent Text Reader

Abstract

The tracking system includes a processor, a controller, two or more light sources, and a Dynamic Vision Sensor (DVS). The light sources are of a known configuration relative to each other and to the controller, and turn on and off in a predetermined sequence. The DVS includes an array of photosensitive elements of a known configuration. The DVS outputs signals corresponding to events at corresponding photosensitive elements within the array in response to changes in the light from the light sources. The signals indicate the time of the event and the position of the corresponding photosensitive element. The processor determines an association between each event and one or more of the light sources, and from that association determines an occlusion of one or more of the light sources. The processor uses the determined occlusion, the known configuration of the light sources, and the position of the corresponding photosensitive element within the array to estimate the position of the object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Aspects of the present disclosure relate to the tracking of game controllers, and more particularly, aspects of the present disclosure relate to the tracking of game controllers using dynamic vision sensors.

Background Art

[0002] Latest virtual reality (VR) and augmented reality (AR) implementations rely on accurate and fast movement tracking for user interaction with devices. AR and VR often rely on information regarding the position and orientation of controllers relative to other objects. Many VR and AR implementations rely on a combination of inertial measurements obtained by accelerometers or gyroscopes within the controller and visual detection of the controller by an external camera to determine the position and orientation of the controller.

[0003] Some of the earliest implementations use detected infrared light by an infrared camera with a defined detection radius on a game controller oriented towards the screen. The camera captures images at a moderately high rate of 200 frames per second and determines the position of the infrared light. The distance between the infrared lights is from a predetermined one and from the relative position of the infrared lights in the camera image from which the position of the controller relative to the screen can be calculated. An accelerometer may also be used to provide information regarding relative three-dimensional changes in the position or orientation of the controller. These conventional implementations rely on a fixed position of the screen and a controller oriented towards the screen. In the latest VR and AR implementations, the screen can be placed near the user's face within a head-mounted display that moves with the user. Thus, having absolute light positions (also called lighthouse points) is undesirable as it requires extra setup time and the user has to set up independent lighthouse points that limit the user's range of motion. Furthermore, even the moderately high frame rate of 200 frames per second of the infrared camera was not fast enough to provide smooth motion feedback. Additionally, since this setup was too simplistic, it was not suitable for more recent inside-out detection methods such as room mapping and hand detection.

[0004] In more recent implementations, cameras and accelerometers are used in combination with trained machine learning algorithms trained to detect both hands, controllers, and / or other body parts. For smoothness of motion detection, a high frame rate camera is needed to generate image frames for body part / controller detection. This generates a large amount of data that needs to be processed quickly for smoothness of the update rate. Thus, expensive hardware is needed to process the frame data. Furthermore, much of the frame data within each frame is discarded as unnecessary as it is not related to motion tracking.

[0005] Aspects of the present disclosure arise in this context. SUMMARY OF THE INVENTION

[0006] The teachings of the present invention can be easily understood by considering the following detailed description in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 11A

Figure 11B

Figure 11C

Figure 11D

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16A

Figure 16B

Figure 16C

Figure 16D

Figure 16E

Figure 17

Figure 18A

Figure 18B

Figure 19

DETAILED DESCRIPTION OF THE INVENTION

[0008] The following detailed description includes many specific details for illustrative purposes, but those skilled in the art will recognize that many variations and modifications to the following details are within the scope of the present invention. Accordingly, the exemplary embodiments of the invention described below are presented without loss of generality to the claimed invention and without imposing limitations on the claimed invention.

[0009] Preamble A new type of vision system called a Dynamic Vision Sensor (DVS) has recently been developed. This DVS resolves changes in a scene by utilizing only changes in the light intensity of a photosensitive pixel array. The DVS has a very high update rate and, instead of delivering a stream of image frames, provides a nearly continuous stream of the locations of changes in pixel intensity. Each change in pixel intensity is sometimes called an event. This has the added advantage of significantly reducing irrelevant data output.

[0010] Two or more light sources can provide continuous updates regarding the position of the DVS camera relative to the position indicator at an update rate determined by the blinking rate of the light. In some implementations, the two or more light sources may be infrared light sources, and the DVS may use infrared-sensitive pixels. Alternatively, the DVS may be sensitive to the visible light spectrum, and the two or more light sources may be multiple visible light sources or one visible light at a known wavelength. In implementations with a DVS sensitive to visible light, the DVS may also be sensitive to movement occurring within its field of view (FOV). The DVS can detect changes in light intensity caused by the reflection of light from a moving surface. In implementations using an infrared-sensitive DVS, the light from an infrared illuminator can be used to detect movement within the FOV by reflection.

[0011] Implementation Figure 1 shows an example of the implementation of tracking a game controller using a DVS101 with a single sensor array according to one aspect of the present disclosure. In the implementation shown, the DVS is attached to a headset 102 that can be part of a head-mounted display. A controller 103 including two or more light sources is within the field of view of the DVS101. In the example shown, the controller 103 includes four light sources 104, 105, 106, and 107. These light sources have a known configuration with respect to each other and with respect to the controller 103. Here, there is one DVS with a single photosensitive array. Using such four light sources, the position and orientation of the controller 103 relative to the DVS101 can be accurately determined. The known information regarding the light sources can include the distance between each of the other light sources for each of the light sources, and the position of each of the light sources on the controller 103. As shown, three light sources 104, 105, 106 can form a plane, and the light source 107 may be out of the plane with respect to the plane formed by the three light sources 104, 105, 106. The light sources here have a known configuration. For example, but not limited to, the first light source 104 is located at the upper left front, the second light source 105 is located at the upper right front, the third light source 106 is located on the left side away from the top, and the fourth light source 107 is located at the lower center front of the controller. With the four light sources, a DVS having a single photosensitive array may be able to determine the movement of the controller along the X, Y, and Z axes. Further, an inertial measurement unit (IMU) 108 can be coupled to the controller 103. By way of example, the IMU 108 can include an accelerometer configured to measure acceleration with respect to one axis, two axes, or three axes. Alternatively, the IMU can include a gyroscope configured to sense changes in rotation with respect to one axis, two axes, or three axes. In some implementations, the IMU may include both an accelerometer and a gyroscope. The IMU 108 can be used to refine the determination of motion, position, and orientation based on information from the DVS101 using a processor. The processor may be located within the headset 102, a game console, or other computing device (not shown).DVS101, headset 102, and IMU108 can be operably coupled to a processor 110, which may be located on a separate device such as the headset 102, controller 103, or a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor can implement tracking as described herein, for example, as will be described later with respect to FIGS. 4, 8, and 9. Further, the processor 110 can control the blinking of the light sources 104, 105, 106.

[0012] During operation, the DVS101 having a photosensitive array can detect the movement of the light sources 104, 105, 106, 107 with the photosensitive array, and when a change in light is detected by the photosensitive array, it can be transmitted to the processor. In some implementations, the light sources can be configured to turn on and off in a predetermined pattern, for example, but not limited to, using signals from a circuit and / or a processor. The processor can use the predetermined pattern to determine the identity of each light source. The identity of the light source can include its known position relative to the controller and other light sources. In other implementations, each light source can be configured to turn off and on in a predetermined pattern, and that pattern can be used to determine the identity of that particular light source. In some implementations, the processor can adapt the known configuration of the light source relative to the controller to the events detected by the photosensitive array.

[0013] The DVS can have a nearly continuous update rate that can be discretely approximated to about one million updates per second. A DVS with a high update rate may be able to resolve a very high-speed blinking pattern of a light source. The blinking rate is mainly limited by the Nyquist frequency, that is, half of the sample rate of the DVS. The light source can blink with a duty cycle suitable for detection of the blinking by the DVS. Generally speaking, the "on" time of the blinking needs to be long enough to be consistently detected by the DVS. Furthermore, due to the high update rate, slight differences in the blinking rate may be detectable.

[0014] The light sources 104, 105, 106, 107 may be broad visible spectrum light such as incandescent lamps or white light emitting diodes. Alternatively, the light sources 104, 105, 106, 107 may be infrared light, or the light sources may have a specific light spectrum profile detectable by the DVS101. The DVS101 can include a photosensitive array configured to detect the emission spectra of the light sources 104, 105, 106, 107. For example, but not limited to, if the light source is infrared light, the photosensitive array of the DVS may be sensitive to infrared light, or if the light source has a specific emission spectrum, the photosensitive array may be configured to enhance the sensitivity to the specific emission spectrum of the light source. Additionally, for example, but not limited to, the photosensitive array of the DVS may not be sensitive to light of other wavelengths not emitted by the light source, or may exclude such light. For example, if the light source is infrared light, the photosensitive array may be configured to detect only infrared light.

[0015] Figure 2 shows an example of the implementation of the tracking of a game controller using a DVS with a dual sensor array according to one aspect of the present disclosure. In this implementation, the headset 203 includes a first DVS 201 and a second DVS 202. Alternatively, the headset 203 may include a DVS having a first photosensitive array 201 and a second photosensitive array 202. The general functions of the light source and the DVS are the same as those described above with respect to FIG. 1. Information from the second DVS or the second array can be integrated with information from the first array to provide a better fit with the orientation of the controller and some depth information. The two DVSs or photosensitive arrays can have partially overlapping fields of view that allow the use of binocular disparity.

[0016] The two DVSs or the two photosensitive arrays provide binocular vision for depth perception. Thereby, the number of light sources can be further reduced. The first DVS and the second DVS or the first photosensitive array and the second photosensitive array may be separated by a known distance, for example, but not limited to, about 50 to 100 millimeters or more than 100 millimeters. More generally, this spacing is large enough to provide sufficient disparity for the desired depth sensitivity, but not so large as to have no overlap between the fields of view. As shown, the controller 207 can include a first light source 204, a second light source 205, and a third light source 206. Since information from the two DVSs or the two arrays provides sufficient information for determining the position and orientation of the controller, a fourth light source coupled to the controller may not be necessary. The controller can include an IMU 208 that can provide additional inertial information used to refine the determination of the position and orientation.

[0017] The first DVS 201, the second DVS 202, the headset 203, and the IMU 208 can be operably coupled to a processor 210, which may be located on the headset 203, the controller 207, or a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor can implement tracking, as described herein and as will be described later with respect to FIGS. 4, 8, and 9, for example. Further, the processor 210 can control the flashing of the light sources 204, 205, 206.

[0018] FIG. 2 shows two DVSs or two photosensitive arrays, but aspects of the disclosure are not so limited. The device can include any number of DVSs or photosensitive arrays. For example, but not limited to, having three DVSs, or a DVS having three separate photosensitive arrays, may allow the use of only two light sources coupled to the controller. The third DVS or photosensitive array may not be collinear with the other two DVSs or arrays. Similar to binocular disparity, each of the additional DVSs or photosensitive arrays can be separated by a known distance and can have a partially overlapping field of view, which can increase the parallax effect used. Further, some implementations can include multiple DVSs, each having multiple photosensitive arrays. For example, but not limiting, there may be two DVSs, each having two separate photosensitive arrays.

[0019] Figure 3 shows an example of the implementation of the tracking of a game controller using a DVS in combination with a single sensor array and a camera, according to one aspect of the present disclosure. In this implementation, a camera 302 is added to the DVS 301. The DVS 301 and the camera 302 can be coupled to a headset 303. The DVS 301 and the camera 302 may have partially overlapping fields of view or may share the same field of view. The controller may include three or more light sources 304, 305, 306 on the controller 307. The DVS 301 and the camera 302 can be used together to determine the position and orientation of the controller. Frames from the camera 302 can be interpolated using events from the DVS 301. Also, using the image frames, position identification and mapping can be performed simultaneously to improve the determination of the orientation and position of the controller. Further, using the image frames from the camera, inside-out tracking of the user using a machine learning algorithm, for example, hand tracking or foot tracking can be performed. The IMU 308 can provide additional inertial information used to further refine the determination of the position and orientation of the controller.

[0020] The first DVS 301, the second DVS 302, the headset 303, and the IMU 308 can be operably coupled to a processor 310, which may be located on the headset 303, the controller 307, or a separate device such as a personal computer, a laptop computer, a tablet computer, a smartphone, or a game console. The processor can implement the tracking as described herein, for example, as described later with respect to FIGS. 4, 8, and 9. Further, the processor 310 can control the blinking of the light sources 304, 305, 306.

[0021] Figure 5 shows an example of the implementation of head tracking or other device tracking using a controller that includes a DVS with a single sensor array, according to one aspect of the present disclosure. Here, the DVS 507 is attached to the controller 508. The headset 501 within the field of view of the DVS 507 is tracked using a light source attached to the headset. The headset 501 can include two or more light sources, here four light sources 502, 503, 504, and 505. The four light sources provide accurate information for three-dimensional position determination by a single DVS. Using additional DVSs or cameras can reduce the number of light sources used. The position of each light source relative to the headset 501 is known to the system and can be somewhat important for providing relevant information to the DVS. Here, the light sources 502, 503, and 504 represent a plane. The light source 505 is located outside the plane of the other light sources 502, 503, and 504. This facilitates the three-dimensional detection of the position, orientation, or movement of the headset. The headset 501 can include an IMU 506, which can use information from the DVS 507 to improve the estimation of position and orientation. Additionally, the controller 508 can include an IMU 509. Using information from the controller IMU 509, the determination of the position and orientation of the headset can be further refined. For example, but not limited to, using IMU information to determine whether the controller is moving relative to the headset and to determine the speed or acceleration of that movement.

[0022] The headset 501, IMU 506, DVS 507, and controller IMU 509 can be operably coupled to a processor 510, which can be located on the headset 501, the controller 508, or a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor can implement tracking as described herein, for example, as described later with respect to FIGS. 4, 8, and 9. Further, the processor 510 can control the flashing of the light sources 502, 503, 504, 505.

[0023] Figure 6 shows an example of the implementation of head tracking or other device tracking using a game controller including a DVS with a dual sensor array according to one aspect of the present disclosure. Here, the controller 605 is coupled to two DVSs 606, 607, or a single DVS having two photosensitive arrays 606 and 607. As described above, the two DVSs or the two photosensitive arrays can be separated at an appropriate distance, for example, between 500 and 1000 millimeters, and can have partially overlapping fields of view. This enables the use of the parallax effect for depth determination. Further, the headset 601 can include three light sources 602, 603, 604 and an IMU 608. The use of two DVSs or two separate arrays 606, 607 may enable the use of fewer than four light sources, for example, but not limited to, three light sources. The two light sources 602, 603 can draw a line, and the third light source 604 can be outside the line formed by the other two light sources 602, 603. The information from the IMU 605 coupled to the headset 601 can be used to refine the determination of position and orientation.

[0024] The headset 601, DVSs 606, 607, and IMU 608 can be operably coupled to a processor 610, which can be located on the headset 601, the controller 605, or a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor can implement tracking as described herein, for example, as described later with respect to FIGS. 4, 8, and 9. Further, the processor 610 can control the blinking of the light sources 602, 603, 604.

[0025] Figure 7 shows an example of the implementation of head tracking or other device tracking using a controller that includes a DVS in combination with a single sensor array and an image camera, according to one aspect of the present disclosure. Here, DVS 707 and camera 708 are coupled to controller 706. Camera 708 and DVS 707 may have partially overlapping fields of view, or may have completely overlapping fields of view. In some implementations, the pixels of the camera and the photosensitive elements of the DVS may share the same photosensitive array, and thus function as an integrated DVS and camera.

[0026] Three or more light sources 702, 703, 704 can be coupled to the headset 701. For example, but not limited to, the light sources may be integrated within the headset housing, and each light source may be an LED, incandescent, halogen, or fluorescent emitter mounted on a circuit board within or on the headset housing. In some implementations, a single emitter may use a plastic or glass light pipe or optical fiber that splits the light from a single emitter into two or more light sources on the headset housing to create multiple light sources.

[0027] Three or more light sources 702, 703, 704 can be configured to turn on and off in response to an electronic signal. In some implementations, three or more light sources can turn on and off in a predetermined sequence. Each time the light sources 702, 703, 704 move or blink, the DVS 707 can generate an event. The camera 708 generates an image frame of its field of view at a set frame rate. Because of the high update rate of the DVS, it may be possible to interpolate the image frames generated by the camera with DVS events.

[0028] In some implementations, the IMU705 can also be coupled to the headset 701. The headset 701, IMU705, DVS707, camera 708, and IMU608 can be operably coupled to a processor 710, which can be located on a separate device such as the headset 701, the controller 706, or a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor can implement tracking as described herein, for example, as will be described later with respect to FIGS. 4, 8, and 9. Further, the processor 710 can control the blinking of the light sources 702, 703, 704.

[0029] In some further implementations, the camera 708 can be a depth camera such as a depth time-of-flight (DTOF) sensor. A DToF camera acquires a depth image by measuring the time it takes for light to travel from a light source to an object in a scene and back to the pixel array. By way of example and not limitation, a DToF camera can operate using continuous wave (CW) modulation, which is an example of an indirect time-of-flight (ToF) sensing method. In a CW ToF camera, light from an amplitude-modulated light source is backscattered by an object within the camera's field of view (FOV), and the phase shift between the emitted waveform and the reflected waveform is measured. By measuring the phase shift at multiple modulation frequencies, the depth value for each pixel can be calculated. The phase shift is obtained by measuring the correlation between the emitted and received waveforms with different relative delays using pixel-integrated photon mixing demodulation.

[0030] A DTOF system generally includes an illumination module and an imaging module. The illumination module consists of a light source, a driver that drives the light source at a high modulation frequency, and a diffuser that projects a light beam from the light source onto a designed field of illumination (FOI). A DToF illumination module can include one or more light sources, and these one or more light sources can be amplitude modulation light emitters such as, for example, but not limited to, vertical cavity surface emitting lasers (VCSELs) or edge emitting lasers (EELs). The imaging module can include an imaging lens assembly, a bandpass filter (BPF), a microlens array, and an array of photosensitive elements that convert incident photon energy into an electrical signal. The microlens array increases the amount of light reaching the photosensitive elements, and the BPF reduces the amount of ambient light reaching the light elements and the microlens array.

[0031] Operation FIG. 4 shows a DVS that tracks the motion of a game controller having two or more light sources, according to one aspect of the present disclosure. The DVS 401 has the controller 402 within its field of view. As shown, the controller 402 includes a plurality of light sources coupled to the controller body. The plurality of light sources can be configured to turn off and then turn on again at a predetermined rate. Each blinking of a light source within the field of view of the DVS 401 can generate one or more events 403 in the DVS. In the events 403 shown, since all the light sources have changed from the off state to the on state, the events indicate that all the light sources are in the on state. Alternatively, depending on the sensitivity of the DVS, a change in the brightness of the light source may be sufficient to trigger the event 403. Note that in other implementations, since the light may turn off and on at different rates or at different times, each event may correspond to fewer light sources than all the lit light sources. The time of the events generated by the blinking light and their corresponding positions within the array can be recorded in a memory (not shown). In some implementations, the position and orientation of the controller 402 can be reconstructed from one or more events from the DVS.

[0032] In some implementations, each light source can blink in a predetermined time sequence. The DVS can output the time at which each event occurred, for example, but not limited to, as a timestamp along with each event. A predetermined time sequence can be used at the time an event occurs to determine the identity of each light within the event, for example, which event corresponds to which light source position. In the example shown, one or more events 403 output by the DVS 401 indicate a first light 406, a second light 407, a third light 408, and a fourth light 409 detected by the photosensitive array. As described above, the identity of each light source can be determined from the information output by the DVS and the predetermined blinking sequence of the light source. For example, but not limited to, the photosensitive array of the DVS 401 can detect a light event 403 at time T+1, and the predetermined sequence can provide that the light source 406 is turned on at T+1, so the light event is determined to correspond to the light source 406. The predetermined sequence can be stored in memory, for example, as a table listing the timing and position of the sequence of each light source. In some implementations, the predetermined sequence may be encoded in the blinking of the light itself, for example, but not limited to, each light may blink in a sequence indicating its identity. For example, a light source labeled 1 may blink in a Morse code sequence indicating the number 1. The identity of the light event can then be recovered through analysis of the light event. Alternatively, the sequence information can come from the light source itself or the driver of the light source and indicate when the light source turns on or off. Alternatively, the light sources can be turned on and off simultaneously, and a machine learning algorithm can be applied to the detected light events 403 to fit the posture 402 of the controller to the events and the known configuration of the light sources.

[0033] Furthermore, the orientation and position of the light source can be determined using information from events such as the size, intensity, and separation of the light. The system may have information defining the size, position on the controller body, and intensity of each light source. The position and orientation may be determined with higher accuracy from the differences between the detected size, intensity, and separation. Additionally, if one or more additional DVSs or photosensitive arrays are present, parallax information can be used to further enhance the determination of the position and orientation.

[0034] During operation, the user can change the position and orientation of the controller 410. Since the update rate of the DVS is relatively high, the DVS can capture the event sequence at 411 as the light event moves at 412 during changes in position and orientation. The movement of the light event here is shown in FIG. 4 by the vector arrow 412. As the detected light event moves, the determined position and orientation of the controller 410 may be updated, or the determined position and orientation may be updated at regular intervals. Due to the blinking of the first light 415, second light 416, third light 417, and fourth light 418, the position and time of the events detected by the photosensitive elements of the DVS may be adapted to the new position and orientation of the controller. Additionally, inertial information from the IMU can be used to refine the movement and position and orientation of the controller 410.

[0035] The flowchart shown in FIG. 8 shows a method of motion tracking by a DVS801 using one or more light sources and one light source configuration adaptation model according to one aspect of the present disclosure. In this implementation, the light sources (shown as LEDs) can be turned on and off simultaneously, independently, or in a predetermined sequence. Events 802 are generated from the DVS801 due to movement of the light sources or blinking of one or more of the light sources. Generally, an event relays an electronic signal containing information such as the time interval during which the event occurred, the position within the array of photosensitive elements of the DVS801 where the event occurred, and binary data (e.g., 1 or 0) corresponding to a change in light intensity exceeding some detection threshold. Each event is processed at 803 and can be formatted into a usable form, e.g., but not limited to, placing the event position within a data array, associating a timestamp with the event, and aggregating multiple events. As an example of aggregating events, all events occurring at individual elements within the array within a predetermined time interval can be integrated into a single data structure for analysis.

[0036] The processed events are then analyzed to associate the detected DVS events with the corresponding LED pulses 804. In some implementations, since each light source can be turned off and on at a unique predetermined time interval, the time stamp of the aggregated events can be used to determine a predetermined time interval from the events and associate a particular light source with a particular event. For example, but not limited to, analyzing the spatial pattern of the aggregated events occurring within a predetermined time interval or sequence of time intervals to determine whether this pattern is consistent with the pulses of the LEDs. Event patterns that are too large, too small, or too irregular in shape may be excluded as LED events. Additionally, the timing of the events can also be analyzed, and events that are too short or too long may also be excluded.

[0037] The trained machine learning model 805 can be applied to the processed events. The model 805 can include information regarding the configuration of the light sources, such as the size of the light sources and their relative positions with respect to the controller body. The machine learning model can be trained by training event data having corresponding masked positions and orientations of the controller, as will be described in a later section. Apply the trained machine learning model to the processed event data to determine the correspondence 806 between the detected pulse 804 and the pose 808. The trained machine learning model can adapt the pose 808 of the controller, e.g., the position and orientation, to one or more processed events, e.g., the detected LED pulse 804. Alternatively, instead of the trained machine learning model 805, a fitting algorithm can be applied to the processed events. The fitting algorithm can use a manually developed light source model to fit the position and orientation of the controller to the processed events. As an alternative, the fitting algorithm can be a hypothesis and test type algorithm, which tries all possible permutations of the light correspondence and uses redundant light sources to find the best fit. Thereafter, a tracking / prediction algorithm can be applied to continue tracking the light source. Further, at 809, the predicted current pose can be used to predict the next pose. At 810, the inertial data from the IMU 807 can be fused with the predicted pose 808 to generate the final predicted position and orientation of the controller. The fusion may be performed by a trained machine learning algorithm trained to refine the position and orientation of the controller using the inertial data. Alternatively, the fusion may be performed by, for example, but not limited to, a Kalman filter, or non-linear optimization.

[0038] FIG. 9 shows a motion tracking method by DVS using timestamped light source position information according to one aspect of the present disclosure. In this implementation, the time of the events output by DVS 901 is used to determine the position and orientation of the controller. Here, the light source is depicted as an LED, and each LED may turn on at different times as shown in the graph. Each time an LED turns on or off, the DVS having the LED in its field of view can generate an event 902. Each event may include an electronic signal corresponding to information such as the time interval in which the event occurred, the position within the array of photosensitive elements of DVS 901 where the event occurred, and binary data (e.g., 1 or 0) corresponding to a change in light intensity exceeding some detection threshold. At 903, each event is processed to format a plurality of events by aggregating them, for example, but not limited to, placing the event positions in a data array as described above and integrating the multiple events occurring at individual elements within the array within a predetermined time interval into a single data structure. Then, the processed events are analyzed to be able to detect the corresponding LED pulses at 904. For example, but not limited to, the shape and size of the events may be analyzed for the regularity and conformity of the light source, and events that are too large, too small, or too irregular in shape may be excluded as LED events, and noise suppression such as the average of the events may be performed to remove random events. Further, the timing of the events may also be analyzed, and events that are too short or too long may also be excluded, or multiple events may be condensed into a single event by excluding events that co-occur with the initial event and end within the time but after the initial event.

[0039] When an event is processed, at 905 the individual LED positions can be determined. Determining the individual LED positions can be performed by using the time sequence in which the LEDs turn on and off. For example, but not limited to, the times of one or more events can be compared to a known time sequence of LED blinking. The known time sequence can be, for example, a table having the on and off times of the LEDs, and the positions on the controller body for each LED, or timestamps from the LED driver when each LED is on or off. When timestamps are used, the timestamps can be correlated with the LED positions on the controller body. From the timing sequence, the LED position information, and the processed event information, the matching position and orientation of the controller can be determined. The IMU data 907 can be integrated with the previously determined LED positions and the inertial data from the IMU by the Kalman filter 908. The Kalman filter can predict the position of the light source based on the inertial information from the IMU, and integrate this prediction with the position information determined from the time sequence of the LEDs to refine the motion data, and at 909 generate a final pose and refine future estimations.

[0040] Training of a general neural network According to aspects of the present disclosure, the tracking system can use machine learning by a neural network (NN). For example, the trained model 805 discussed above can use machine learning as discussed below. The machine learning algorithm can use a training dataset, which can include inputs from a DVS such as events, or processed events with known positions and orientations of controllers as labels. Further, the machine learning algorithm using the NN can perform a fusion between the position and orientation of the controller determined from the DVS information and the inertial information from the IMU. The training set for the fusion can be, for example, but not limited to, inertial data with potential positions and orientations of the controller, as well as the final positions and orientations. In some implementations, the machine learning algorithm may be trained to perform SLAM (simultaneous localization and mapping) using a training set that includes objects such as the ground, landmarks, and body parts with hidden labels. The hidden labels may include the identity of the objects and their relative positions. As generally understood by those skilled in the art, SLAM technology generally solves the problem of continuously tracking the position of an agent within an environment while constructing or updating a map of the unknown environment.

[0041] The NN may include one or more of several different types of neural networks and may have many different layers. By way of example and not limitation, the neural network may be composed of one or more convolutional neural networks (CNNs), recurrent neural networks (RNNs), and / or dynamic neural networks (DNNs). The motion detection neural network can be trained using the general training methods disclosed herein.

[0042] By way of example and not limitation, FIG. 10A shows a basic form of an RNN that can be used, for example, with a trained model 805. In the illustrated example, the RNN has a layer of nodes 1020, each of which is characterized by an activation function S, one input weight U, a recurrent hidden node transition weight W, and an output transition weight V. The activation function S can be any non-linear function known in the art and is not limited to the hyperbolic tangent (tanh) function. For example, the activation function S can be a sigmoid function or a ReLu function. Unlike other types of neural networks, the RNN has one set of activation functions and weights for the entire layer. As shown in FIG. 10B, the RNN can be considered as a series of nodes 1020 having the same activation function moving through times T and T+1. Thus, the RNN maintains historical information by feeding the results from the previous time T to the current time T+1.

[0043] In some implementations, a convolutional RNN may be used. Another type of RNN that can be used is a long short-term memory (LSTM) neural network, which adds a memory block to the RNN node along with an input gate activation function, an output gate activation function, and a forget gate activation function that, as described by Hochreiter & Schmidhuber, "Long Short-term memory", Neural Computation 9(8):1735-1780 (1997), which is incorporated herein by reference, allows the network to hold some information for a longer period of time.

[0044] Figure 10C shows an exemplary layout of a convolutional neural network, such as a CRNN, that can be used for a trained model 805 and the like according to an aspect of the present disclosure. In this representation, the convolutional neural network is generated for an input 1032 having a size of 4 units in height and 4 units in width, which gives a total area of 16 units. The convolutional neural network represented has a filter 1033 having a size of 2 units in height and 2 units in width, with a skip value of 1 and 136 channels of size 9. Only the connections 1034 between the first column of channels and their filter windows are shown in FIG. 10C for clarity. However, aspects of the present disclosure are not limited to such implementations. According to aspects of the present disclosure, the convolutional neural network may have any number of additional neural network node layers 1031 and may include such layer types as additional convolutional layers, fully connected layers, pooling layers, max pooling layers, local contrast normalization layers, etc. of any size.

[0045] As seen in FIG. 10D, training a neural network (NN) begins with the initialization of the weights of the NN at 1041. Generally, the initial weights should be randomly distributed. For example, an NN with a tanh activation function should have random values distributed between -1 / √n and 1 / √n, where n is the number of inputs to the node.

[0046] After initialization, the activation function and optimizer are defined. Then, at 1042, the NN is provided with a feature vector or an input dataset. Each of the different feature vectors may be generated by the NN from inputs with known labels. Similarly, the NN may be provided with a feature vector corresponding to an input with a known labeling or classification. The NN then predicts, at 1043, the label or classification for the feature or input. The predicted label or class is compared, at 1044, to the known label or class (also known as the ground truth), and the loss function measures the total error between the prediction and the ground truth across all training samples. By way of example, and not limitation, the loss function may be a cross-entropy loss function cost, a triplet contrastive function, an exponential cost, etc. Multiple different loss functions may be used depending on the purpose. By way of example, and not limitation, a cross-entropy loss function may be used to train a classifier, whereas a triplet contrastive function may be employed to learn a pre-trained embedding. The NN is then optimized and trained, as shown at 1045, using the result of the loss function and a known method for training neural networks such as error backpropagation by adaptive gradient descent. At each training epoch, the optimizer attempts to select the model parameters (i.e., weights) that minimize the training loss function (i.e., the total error). The data is partitioned into training samples, validation samples, and test samples.

[0047] During training, the optimizer minimizes the loss function with respect to the training samples. After each training epoch, the model is evaluated with respect to the validation samples by computing the validation loss and accuracy. If there is no significant change, training may be stopped, and the resulting trained model may be used to predict the labels of the test data.

[0048] Therefore, a neural network may be trained from those inputs to identify and classify inputs having known labels or classifications. Similarly, an NN may be trained using the described method to generate feature vectors from inputs having known labels or classifications. The above discussion relates to RNNs and CRNNs, but these discussions can be applied to NNs that do not include a recurrent or hidden layer.

[0049] Hybrid sensor FIG. 11A is a diagram illustrating a hybrid DVS having a plurality of co-located sensor types, according to an aspect of the present disclosure. In some implementations, a hybrid DVS can be used to combine multiple sensor types into one device. The multiple sensor types can be, for example, but not limited to, a DVS infrared photosensor, a DVS visible light photosensor, a DVS wavelength-specific photosensor, a visible light camera pixel, an infrared camera pixel, a DTOF camera pixel. The second sensor type 1102 can be interspersed with the first sensor type 1102 on the same array 1101. For example, but not limited to, the DVS visible light photosensor 1103 can surround the visible light camera pixel 1102, or the DVS infrared photosensor 1103 can surround the DVS visible light photosensor 1102, or the visible light camera pixel 1103 can surround the DVS infrared photosensor 1102, or any combination thereof. Although a single element 1102 is shown surrounded by other elements 1103, aspects of the present disclosure are not so limited. A single element can include a cluster of multiple DVS photosensors or camera pixels. For example, but not limited to, a cluster of four camera pixels can be surrounded by eight DVS photosensors, or one camera pixel can be surrounded by eight pairs, triplets, or quadruplets of DVS photosensors.

[0050] Furthermore, as shown in FIG. 11B, the hybrid DVS can have a plurality of sensor types arranged in a moiré pattern within the array. Here, blocks of the second sensor type 1103 are evenly distributed among blocks of the first sensor type 1102. Each of the first sensor type 1102 and the second sensor type 1103 may be different. For example, but not limited to, the first sensor type 1102 may be a DVS infrared photosensor, and the second sensor type 1103 may be a DVS visible light photosensor.

[0051] Alternatively, the sensor types may be distinguished by filtering. In these implementations, one or more filters selectively transmit light to photosensors located behind one or more filters. The one or more filters can, for example, but not limited to, selectively transmit one or more specific wavelengths of light or a specific polarization. Also, the one or more filters can selectively block one or more specific wavelengths of light or a specific polarization. The photosensors behind the one or more filters can be configured to be used as different sensor types. For example, but not limited to, an infrared pass filter that allows only infrared light to pass through 1102 can cover one or more sensor elements within the array 1101, and the other sensor elements may not be filtered, or may be an infrared cut filter 1103. In another alternative implementation, the one or more filters can, for example, but not limited to, be an optical notch filter, which allows only a specific wavelength of light to pass through 1102, and the other filters can block that specific wavelength, but may allow other wavelengths to pass through 1103. This filtering enables the use of the wavelength of the light source for DTOF and the specific wavelength for the DVS photosensor, which may reduce the likelihood of false light source detection. Here, the sensor element can be any type of DVS photosensor or any type of camera pixel.

[0052] The patterned sensor or filter element can be incorporated into a hybrid imaging unit, as shown, for example, in FIGS. 11C and 11D. FIG. 11C shows an example of a hybrid DVS imaging unit 1112C having one or more lens elements 1114, an optional microlens array 1116, a bandpass filter 1118, and a patterned hybrid sensor array 1120. The sensor array includes DVS sensor elements 1122 and conventional imaging sensor elements 1124 that can be arranged in a pattern, as shown, for example, in FIGS. 11A or 11B. Such imaging units can be used with an illumination unit (not shown) within a DTOF sensor. The hybrid DVS imaging unit 1112D shown in FIG. 11D has a DVS sensor array 1126 and a patterned filter element 1118 that includes bandpass regions 1118A and bandcut regions 1118B arranged in a pattern, as shown, for example, in FIGS. 11A or 11B.

[0053] FIG. 12 is a diagram illustrating a hybrid DVS including a plurality of sensor types having inputs separated by an optical splitter, according to an aspect of the present disclosure. In this implementation, the optical splitter 1204 can filter light based on wavelength, and the array 1201 can include a plurality of sensor element types physically separated based on the wavelength or polarization of the light for which detection is desired. For example, without limitation, unpolarized white light 1205 (a mixture of at least all visible light wavelengths and, in most cases, some infrared wavelengths) can enter the optical splitter 1204, which can be a dispersive prism, diffraction grating, dichroic mirror, or the like. As shown, the white light 1205 entering the optical splitter 1204 can be separated by wavelength or polarization, with some light 1206 incident on a first portion of the array 1202 and other light 1207 incident on a second portion of the array 1203. Here, the optical splitter 1204 can be considered a filter that changes the diffraction angle based on wavelength or polarization. The array shown depicts the array as a single unit with a thick separation line 1203, but aspects of the present disclosure are not so limited. The first portion of the array 1201 may have a separation of up to 1 millimeter between the second portion of the array 1202, and while the array is shown as being separated vertically, other implementations may have portions separated horizontally, diagonally, or circumferentially. Further, the optical splitter here can be combined with different filtering or different sensor configurations, such as those shown in FIGS. 11A and 11B, to provide additional optical wavelength separation for different sensor types.

[0054] Figure 13 is an illustration of a hybrid DVS having inputs separated by a plurality of sensor types by a microelectromechanical (MEMS) mirror. Here, the MEMS mirror 1304 can vibrate between different positions at set times so as to reflect light to the first part 1301 or the second part 1302 of the array according to the time when the light 1305 reaches the MEMS mirror 1304. The first part 1301 and the second part 1302 of the array can be physically separated from each other at 1303 based on the incident angle of the light reflected by the MEMS mirror 1304. In this way, light can be temporarily filtered between different sensor types. Such temporary filtering can be timed according to a known blinking pattern of a light source on a controller or a headset. For example, but not limited to, the light source on the controller or the headset can be turned on at regular intervals for a predetermined duration. For example, the light source can be turned on for 60 microseconds every 100 microseconds. In such a case, the MEMS mirror 1304 can be synchronized to reflect the light 1305 to the DVS part of the array 1301 for more than 60 microseconds every 100 microseconds or less to capture changes in the light source. Other reflected light 1307 is detected by the second part of the array 1302 during the time when the light source is off, and ambient light can be captured for image tracking or DTOF.

[0055] Alternatively, the MEMS mirror 1304 can filter the light 1305 based on wavelength. The MEMS mirror in these implementations can be, for example, but not limited to, a MEMS Fabry-Perot filter or a diffraction grating. The MEMS mirror can diffract light in the first wavelength range 1306 to at least the first part 1301 of the array or light in the second wavelength range 1307 to the second part 1302 of the array.

[0056] FIG. 14 is an illustration of a hybrid DVS having inputs in which multiple sensor types are temporarily filtered. Here, the filter enables selective passage of light through the array based on time. In some implementations, the filter may be an optical waveguide. In a first time step, the filter can pass light of a first wavelength or polarization through array 1401 and block light of a second wavelength or polarization, or other wavelengths or polarizations. In a second time step, the filter enables a second wavelength or polarization 1402 but can block light of the first wavelength or polarization or other wavelengths of polarization. In this way, light can be temporarily filtered to aid in tracking with different sensor types. For example, without limitation, a light source coupled to a controller or headset may be infrared light or may have a specific wavelength. The light source can be configured to turn on and off at specific intervals. The specific intervals at which the light source turns on and off can be a sequence or a coded pattern. The temporary filtering can be activated to allow a specific wavelength to pass through the sensor and block other wavelengths during specific intervals. Further, the switching interval of the filtering may be longer than the specific interval of the light source, taking into account the propagation time of light to the sensor.

[0057] Body tracking FIG. 15 is an illustration showing body tracking by a DVS according to an aspect of the present disclosure. As shown, user 1501 can wear a headset 1504 having a DVS 1503. Here, the DVS is shown with two arrays or DVS units or cameras and DVS. User 1501 holds two controllers 1502 having two or more light sources 1505. DVS 1503 has a controller 1502 with a light source 1505 corresponding to within the field of view (FOV). Further, DVS 1503 can have a user appendage such as a user's hand or arm 1507, or leg or foot 1508 within its FOV. The DVS can also have a ground or other landmark 1509 within its FOV. For example, but not limited to, the photosensitive element of the DVS can detect the reflection of light corresponding to the user's appendage or the ground or other landmark when the user moves or the light changes. The camera detects the reflection of light from the field of view at the frame rate of the camera. From the detected light reflection, at 1506, the user's appendage or the ground or landmark can be determined. In some alternative implementations, an external DVS 1510 can be used to track the user's appendage. The external DVS 1510 can be at a distance away from the user selected such that the user's appendage fits within the field of view of the external DVS 1510. For example, but not limiting, the external DVS can be located on or under the top surface of a television or computer monitor, or other stand-alone or wall-mounted display.

[0058] The machine learning algorithm is trained to determine the user's body, appendages, ground or landmarks, and their relative positions and orientations from data such as events or frames. The machine learning algorithm may be a neural network, and the training may be similar to the method described in the section on training general neural networks of FIGS. 10A-10D above. The neural network can be trained using a training set that includes labeled events or frames, or both. The labeled events or frames, or both, can include, for example, but not limited to, labels of the user's body, appendages, ground, landmarks, and the relative positions and orientations of the user's body, appendages, ground, and landmarks. In addition to determining the position and orientation of the controller 1502, the position and orientation of the user's body, appendages, ground or landmarks, and their relative positions and orientations can be determined using, for example, SLAM. Alternatively, the position and orientation of the controller can be determined using the same trained neural network that determines the labels of the user's body, appendages, ground or landmarks and their relative positions and orientations. As shown by element 1506, a model can be fitted to the determined user's body and appendages to improve the determination of relative positions and locations.

[0059] Safety shutter As described above, body tracking can be used in conjunction with the determination of the position and orientation of the controller to trigger a safety shutter in a VR or AR headset. FIG. 16A shows a headset with a safety shutter door according to an aspect of the present disclosure. Headset 1601 may include a head strap 1604, eye pieces 1603, and a display screen 1602. Eye pieces 1603 may include one or more lenses configured to focus on display screen 1602. The one or more lenses may be, for example, Fresnel lenses or prescription lenses. Display screen 1602 may be transparent, and hole 1607 within the headset body may enable vision through display screen 1602 when the safety shutter is open.

[0060] In this implementation, safety shutter 1606 is a door that swings away from hole 1607 when the safety system is activated. Here, system-operated clasp 1605 interacts with clasp 1608 on safety shutter door 1606 to close and secure the door over hole 1607. Spring-loaded hinge 1609 can ensure that safety shutter door 1606 opens quickly when system-operated clasp 1605 opens. The spring-loaded hinge may have, for example, but not limited to, a clock spring wound around the hinge and secured to the door, the spring being wound when the door closes and unwound when the door opens. Alternatively, a leaf spring may push on this door when safety shutter door 1606 is closed.

[0061] The system-operated clasp 1605 can be configured to open the clasp when the safety system is activated. The safety system-operated clasp 1605 can include an electric motor or a linear actuator to move the clasp. The safety system may be activated when a ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendage. The safety system can use the determination of the user's body, appendage, ground or landmark, and their relative positions and orientations as described above. When the safety system is activated, a signal to open the clasp can be sent to the safety system-operated clasp 1605. When the clasp opens, the spring-loaded hinge 1609 pushes open the safety shutter door 1606, enabling the user to view through the display 1602 and avoid the risk of activating the safety system.

[0062] Other implementations of the safety shutter may be used. For example, FIG. 16B shows an alternative headset with a sliding safety shutter according to an aspect of the present disclosure. Here, when the safety system is activated, the safety shutter slide 1616 slides so as not to obstruct the hole 1607. The safety shutter slide 1616 can move on a spring-loaded rail 1619. Alternatively, the sliding safety shutter 1616 can include a tab inserted into a slot 1619 within the headset body, and a spring within the slot can also push the sliding safety shutter. In some additional alternative implementations, the spring may be omitted and the sliding safety shutter 1616 may operate by gravity. The spring-loaded rail 1619 can push the shutter when closing the sliding safety shutter 1616 and ensure that the safety shutter opens quickly when the safety system is activated. The system-operated clasp 1605 interacts with a clasp 1618 on the sliding safety shutter 1616 to close and secure the slide.

[0063] If the safety system is to be activated when the ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendage. The safety system can use, as described above, the determination of the user's body, appendage, ground or landmark, and their relative positions and orientations. When the safety system is activated, a signal to open the clasp can be sent to the system-operated clasp 1605. When the clasp opens, the spring 1619 pushes open the sliding safety shutter 1616, allowing the user to see through the display 1602 and avoid the risk of activating the safety system.

[0064] Figure 16C shows a headset with a louvered safety shutter according to an aspect of the present disclosure. In this implementation, the slats of the louvered safety shutter 1626 are longer in the first dimension than in the second dimension. In the closed position, the longitudinal dimension of the slats is substantially parallel to the optical system, and each slat overlaps either another slat or the headset body, blocking light passing through the display screen 1602. In the open position, the slats change position so that the short dimension is parallel to the optical system 1603, allowing light to pass through the slats and reach the display screen. The actuator rod 1628 can hinge each of the slats 1626. The safety system-controlled actuator 1629 can push or pull the actuator rod 1628 to open the louvered safety shutter when the safety system is activated. In some other implementations, the safety system-controlled actuator may be spring-loaded, and the clasp connected to the actuator rod 1628 and the system-controlled actuator 1629 may include a clasp that interlocks with the clasp of the actuator rod. When the safety system is activated, the clasp of the system-controlled actuator can open, allowing the actuator rod to move under spring pressure and the louvers to open.

[0065] When the ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendage, the safety system may be activated. As described above, the safety system can use the determination of the user's body, appendage, ground or landmark, and their relative positions and orientations. When activated, a signal to move the actuator rod can be sent to the system-operated actuator. When the actuator rod 1629 moves, it pushes open the slats of the louvered safety shutter 1626, enabling the user to see through the display 1602 and avoid the risk of operating the safety system.

[0066] Figure 16D shows a fabric safety shutter according to an aspect of the present disclosure. In this implementation, the safety shutter 1636 is composed of an opaque fabric, for example, but not limited to, a densely woven cotton fabric, polyester, vinyl, or a densely woven wool fabric. The fabric safety shutter 1636 may be coupled to a fabric roller 1639. The fabric roller 1639 may be configured to wind up the fabric shutter when the safety system is activated. The fabric roller may be, for example, but not limited to, spring-loaded using a clock spring, so that the clock spring is tensioned when the safety shutter closes, or alternatively, an electric motor may be used to wind up the fabric safety shutter. The system-operated clasp 1605 is interlocked with a clasp 1638 coupled to the fabric safety shutter 1636 to ensure that the fabric safety shutter does not open unintentionally.

[0067] During operation, the fabric safety shutter may be in the closed position. If a ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendage, the safety system may be activated. The safety system can use the determination of the user's body, appendage, ground or landmark, and their relative positions and orientations, as described above. When the safety system is activated, a signal to open the clasp can be sent to the system-operated clasp 1605. When the clasp opens the fabric roller 1619, the fabric safety shutter 1616 is wound up, allowing the user to view through the display 1602 and avoid the risk of activating the safety system.

[0068] FIG. 16E shows a headset with a liquid crystal safety shutter according to an aspect of the present disclosure. In this implementation, the liquid crystal screen 1646 is integrated into the headset 1601 and otherwise covers a hole in the headset. The safety system-controlled liquid crystal screen driver 1649 is communicatively coupled to the liquid crystal screen 1646.

[0069] The liquid crystal screen 1646 may be, for example, but not limited to, a liquid crystal shutter having a first polarizer and a second polarizer. The first polarizer has a polarization difference of 90 degrees from the second polarizer and has a cavity filled with a fluid. The cavity filled with the fluid can contain liquid crystals, and these liquid crystals are configured to have a first alignment when there is no electric field that changes the polarization to allow light to pass from the first polarizer to the second polarizer. The liquid crystals can be further configured to align to a second alignment under an electric field. Since the second alignment of the liquid crystals does not change the polarization of the light, the light passing through the first polarizer is blocked by the second polarizer. Electrodes may be disposed along the surface of the cavity filled with the liquid to enable control of the liquid crystals. If the safety system control type liquid crystal screen driver can be communicatively coupled to these electrodes, the electrodes can control the liquid crystals within the cavity filled with the fluid. As used herein, being communicatively coupled means that an electrical signal representing a message or instruction can be transmitted and / or received from one coupling element to the other coupling element, and these signals can pass through intermediate elements and their formats may change, but the messages contained therein remain unchanged.

[0070] During operation, the safety system control driver 1649 can send a signal to the liquid crystal safety shutter 1646 to make it opaque while the display screen 1602 is active. When the safety system becomes active, the driver 1649 can make the liquid crystal safety shutter 1646 transparent. For example, but not limited to, if the driver can reduce the voltage supplied to the liquid crystal safety shutter and return the liquid crystal to their first orientation, the polarization of light changes, allowing light to pass through the second polarizer. The safety system may be activated when a ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendages. The safety system can use the determination of the user's body, appendages, ground or landmark, and their relative positions and orientations as described above.

[0071] Tracking of finger position Aspects of the present disclosure may be applied to finger tracking. FIG. 17 shows finger tracking by a DVS and a controller according to an aspect of the present disclosure. Here, the controller 1701 includes two or more light sources and one or more buttons 1705. The two or more light sources include one or more light sources 1706 proximate to the one or more buttons 1705 and two or more other tracking light sources 1703. As described above, the two or more other tracking light sources 1703 can generate events in the DVS 1702 for determining the position and orientation of the controller 1701. The DVS having two DVSs 1702, or two arrays and three other tracking light sources 1703 shown here is used for determining the position and orientation of the controller.

[0072] One or more light sources 1706 proximate to one or more buttons 1705 can be used for finger tracking. For example, but not limited to, finger tracking can be achieved by the DVS 1702 using the occlusion of one or more light sources 1706 proximate to the button 1705. One or more light sources 1706 proximate to the button can be turned off and on at a predetermined interval. The DVS 1702 can generate an event for each blink. By analyzing the event, the occlusion of one or more light sources 1706 proximate to the button 1705 can be determined. By knowing the configuration of one or more light sources proximate to the button, for example, but not limited to, when occluded by a finger or palm, the pattern of light detected in the events generated by the DVS will be different from when the light source is not occluded. As described above with respect to determining the position and orientation of the controller, here the timing of the blinking is used to determine which light is occluded, and thus the position of the corresponding finger or palm can be determined. If the light source 1706 proximate to the button 1705 has a reduced intensity or no intensity during the interval, it can be determined that the light source proximate to the button is "on". It is determined that the light source is occluded. Similarly, when the detected intensity of the light known to be "on" changes, an event can be generated, and from that event, it may be determined that the user's finger or hand has moved and the button has been aborted. The occlusion of one or more of the light sources proximate to one or more buttons can be correlated to the position of the finger or palm based on their positions. For example, but not limited to, the position of the user's hand 1704 can be determined using a light source located near the user's palm when the controller 1701 is held.

[0073] The light source 1706 can be positioned around each button 1705 and can use the button configuration and design of the controller to determine the finger position. For example, but not limited to, the controller 1701 may be designed such that when held, each finger of the user is placed near the button 1705. Next, the occlusion pattern of the light source determined from the DVS event may be used to determine when the user's finger is hovering in the air over an inactive button, and moreover, may be used to determine when the user's finger has passed over the button. This can help provide additional interaction options to the user, such as a half - press of the button, or an insufficient press, or other button options. By having multiple light sources surround each button, it becomes possible to refine the determination of the position of the user's finger or palm. For example, but not limited to, in some implementations, more than 10 light sources may surround each button, and in other implementations, a single light source may shine light through a translucent diffuser around the button and use interruptions in the diffused light profile to determine the finger position.

[0074] In some implementations, the button 1705 itself can also be a light source. One or more buttons 1705 can be turned off and on at different intervals from one or more light sources 1706 or other light - tracking light sources 1703 proximate to the button. Alternatively, the light source of the button 1705 may have a different wavelength or polarization from one or more light sources 1706 or other light - tracking light sources 1703 proximate to it.

[0075] Also, the power saving mode of the light source can be enabled using buttons and tracking. For example, if a button is determined to be pressed, one or more light sources proximate to the button may be dimmed or turned off. Further, if the controller 1701 determines that it is outside the field of view of the DVS 1702, one or more lights proximate to the button may be dimmed or turned off. In some implementations, data from an IMU can be used to determine whether the controller is being held by the user. For example, but not limited to, if no changes in IMU data such as acceleration, angular velocity, etc. have been detected over a threshold period, the light source may be dimmed or turned off. When a change in the IMU data is detected, the light source can be turned back on.

[0076] In an alternative implementation, finger tracking can be performed without using one or more light sources. A machine learning model can be trained using a machine learning algorithm to detect the position of a finger from events generated from changes in ambient light due to the movement of the finger. The machine learning model may be a general machine learning model such as the CNN, RNN, or DNN described above. In some implementations, for example, but not limited to, a special machine learning model such as a spiking (or spark) neural network (SNN) can be trained with a special machine learning algorithm. The SNN mimics biological NNs by having activation thresholds and weights, and these weights are adjusted according to the relative spike times within intervals, also called spike-timing-dependent plasticity (STDP). When the activation threshold is reached, the SNN is said to spike and send its weights to the next layer. The SNN can be trained via STDP and supervised or unsupervised learning techniques. Further information on SNNs can be found in Tavanaei, Amirhossein et al. "Deep Learning in Spiking Neural Networks" Neural Networks (2018) arXiv:1804.08150, the contents of which are hereby incorporated by reference for all purposes.

[0077] Alternatively, HDR (high dynamic range) images may be constructed using events aggregated from ambient data. The machine learning model is trained to recognize the position or the position and orientation of the hand or the controller from the HDR image. The trained machine learning model can be applied to the HDR image generated from the events to determine the position or the position and orientation of the hand / finger or the controller. The machine learning model may be a general machine learning model trained with a supervised learning technique as described in the section on training of general neural networks.

[0078] Visual target tracking Aspects of the present disclosure may be applied to eye gaze tracking. Generally, eye gaze tracking image analysis utilizes characteristics specific to how light is reflected from the eye to determine the line of sight direction from an image. For example, an image may be analyzed to identify the position of the eye based on the corneal reflection in the image data, and further analyzed to determine the line of sight direction based on the relative position of the pupil in the image.

[0079] Two common eye gaze tracking techniques for determining the line of sight direction based on the position of the pupil are known as the bright pupil method and the dark pupil method. The bright pupil method involves illuminating the eye with a light source that is substantially aligned with the optical axis of the DVS, which reflects the emitted light off the retina and back through the pupil to the DVS. The pupil appears in the image as a distinguishable bright spot at the position of the pupil, similar to the red-eye effect that occurs in an image during conventional flash photography. In this method of eye gaze tracking, when the contrast between the pupil and the iris is not sufficient, the bright reflection from the pupil itself helps the system to locate the pupil.

[0080] The dark pupil method involves illuminating with a light source that is substantially offset from the optical axis of the DVS, which reflects the light directed through the pupil away from the optical axis of the DVS, resulting in a distinguishable dark spot at the position of the pupil for an event. In another dark pupil method system, an infrared light source and a camera directed at the eye can see the corneal reflection. Such a DVS-based system tracks the positions of the pupil and the corneal reflection, thereby obtaining a parallax due to the different depths of the reflections and improving the accuracy.

[0081] FIG. 18A shows an example of a dark pupil eye tracking system 1800 that may be used in the context of the present disclosure. The eye tracking system tracks the orientation of the user's eye E with respect to a display screen 1801 on which a visible image is presented. Although a display screen is utilized in the exemplary system of FIG. 18A, certain alternative embodiments may utilize an image projection system that can project an image directly onto the user's eye. In these embodiments, the user's eye E is tracked with respect to the image projected onto the user's eye. In the example of FIG. 18A, the eye E collects light from the screen 1801 through a variable iris I; the lens L projects an image onto the retina R. The opening of the iris is known as the pupil. Muscles control the rotation of the eye E in response to nerve impulses from the brain. The upper and lower eyelid muscles ULM, LLM control the upper and lower eyelids UL, LL, respectively, in response to other nerve impulses.

[0082] The light-sensitive cells on the retina R generate electrical impulses that are sent via the optic nerve ON to the user's brain (not shown). The visual field of the brain interprets the impulses. Not all parts of the retina R have equal light sensitivity. Specifically, the light-sensitive cells are concentrated in a region known as the fovea.

[0083] The illustrated image tracking system includes one or more infrared light sources 1802, for example, light-emitting diodes (LEDs) that direct non-visible light (e.g., infrared light) towards the eye E. A portion of the non-visible light is reflected by the cornea C of the eye and a portion is reflected by the iris. The reflected non-visible light is directed by a wavelength-selective mirror 1806 towards a DVS 1804 that is sensitive to infrared light. The mirror transmits visible light from the screen 1801 but reflects non-visible light reflected from the eye.

[0084] The DVS 1804 generates events of the eye E that can be analyzed to determine the line of sight direction GD from the relative position of the pupil. This event can be generated by a processor 1805. The DVS 1804 is advantageous in this implementation because the very high update rate of the events provides near real-time information regarding changes in the user's line of sight.

[0085] As can be seen in FIG. 18B, an event 1811 indicating the user's head H can be analyzed to determine the line-of-sight direction GD from the relative position of the pupils. For example, the analysis may determine the two-dimensional offset of the pupil P from the center of the eye E in the image. The position of the pupil relative to the center can be converted into the line-of-sight direction with respect to the screen 1801 by simple geometric calculations of a three-dimensional vector based on the known size and shape of the eyeball. The determined line-of-sight direction GD can indicate the rotation and acceleration of the eye E as the eye E moves relative to the screen 1801.

[0086] As can also be seen in FIG. 18B, the event may include reflections 1807 and 1808 of non-visible light from the cornea C and the lens L, respectively. Since the depths of the cornea and the lens are different, the parallax and the refractive index between the reflections can be used to improve the accuracy in determining the line-of-sight direction GD. An example of this type of eye-tracking system is a dual Purkinje tracker, where the corneal reflection is the first Purkinje image and the lens reflection is the fourth Purkinje image. When the user wears glasses, a reflection 1808 from the user's glasses 1809 may also be present.

[0087] The performance of the eye-tracking system depends on a number of factors including the arrangement of the light source (IR, visible light, etc.) and the DVS, whether the user wears glasses or contacts, the headset optics, the latency of the tracking system, the speed of eye movement, the shape of the eyes (which may change during the day or as a result of movement), the state of the eyes, such as amblyopia, line-of-sight stability, fixation on a moving object, the scene presented to the user, and the movement of the user's head. The DVS reduces the output of extra information to the processor and provides a very high update rate for events. This enables faster processing and faster determination of the state of eye tracking and error parameters.

[0088] Error parameters that can be determined from eye-tracking data can include, but are not limited to, rotational speed and prediction error, fixation error, confidence intervals regarding current and / or future eye positions, and smooth pursuit error. State information regarding the user's eye includes the individual state of the user's eye and / or gaze. Thus, exemplary state parameters that can be determined from eye-tracking data can include, but are not limited to, blink metrics, saccade metrics, depth of field response, color vision anomalies, gaze stability, and eye movements as precursors to head movements.

[0089] In certain implementations, the eye-tracking error parameter can include a confidence interval regarding the current eye position. The confidence interval can be determined by examining the rotational speed and acceleration of the user's eye for no change from the last position. In alternative embodiments, the eye-tracking error and / or state parameter can include a prediction of a future eye position. The future eye position can be determined by examining the rotational speed and acceleration of the eye and extrapolating the possible future positions of the user's eye. Generally speaking, the DVS update rate of the eye-tracking system can result in a smaller error between the determined future position and the actual future position for users with high rotational speed and acceleration values due to the very high DVS update rate, and this small error can be significantly smaller than that of existing camera-based systems.

[0090] In yet another alternative implementation, the gaze tracking error parameter can include a measurement of the eye velocity, such as the rotational velocity. In certain alternative embodiments, determining the gaze tracking state parameter includes measuring a measurement criterion of the user's blink. During a normal blink, typically a period of 150 milliseconds (ms) elapses and the user's vision is not focused on the presented image. Thus, depending on the frame rate of the display device, the user's vision may not be focused on the presented image for up to 20-30 frames. However, when the blink ends, the user's line of sight direction may not correspond to the last measured line of sight direction determined by the acquired gaze tracking data. Thus, the measurement criterion of the user's line of sight can be determined from the acquired gaze tracking data. These measurement criteria can include, but are not limited to, the measured start time and end time of the user's blink, as well as the predicted end time.

[0091] In yet an additional alternative implementation, determining the gaze tracking state parameter includes measuring a measurement criterion of the user's saccade. During a normal saccade, typically a period of 20-200 ms elapses and the user's vision is not focused on the presented image. Thus, depending on the frame rate of the display device, the user's vision may not be focused on the presented image for up to 40 frames. However, by the nature of the saccade, when the saccade ends, the user's line of sight direction moves to another region of interest. Thus, the gaze tracking data can be used to establish the measurement criterion of the user's saccade based on the actual time or predicted time elapsed during the saccade. These measurement criteria can include, but are not limited to, the measured start time and end time of the user's saccade, as well as the predicted end time.

[0092] In certain alternative implementations, determining the gaze tracking state parameter includes determining a transition in the user's line of sight direction between regions of interest as a result of a change in the depth of field between the presented images. This is because when a transition between regions of interest of the presented image is provided, the user will experience a saccade.

[0093] In further additional alternative implementations, the determined gaze tracking state parameters can be adapted to color vision deficiencies. For example, regions of interest may be present in an image presented to a user such that these regions are not recognized by a user having a particular form of color vision deficiency. The resulting gaze tracking data determines, for example, whether the user's gaze has identified or responded to a region of interest as a result of a change in the user's gaze direction. Thus, as a gaze tracking error parameter, it can be determined whether the user is color vision deficient with respect to a particular color or spectrum.

[0094] In a particular alternative implementation, the determined gaze tracking state parameters include a measure of the user's gaze stability. Determining gaze stability can be done by measuring the radius of the user's microsaccades, and the smaller the overshoot and undershoot of fixation, the more stable the user's gaze.

[0095] In further additional alternative implementations, the determined gaze tracking error and / or state parameters include the user's ability to fixate on a moving object. These parameters may include a measure of the user's eye's ability to perform smooth pursuit and the maximum object tracking speed of the eye. Typically, the jitter of eye movements experienced by a user with excellent smooth pursuit ability is reduced.

[0096] In a particular alternative implementation, the determined gaze tracking error and / or state parameters include determining eye movements as precursors to head movements. The offset between the orientation of the head and the eyes can affect certain error and / or state parameters, such as in smooth pursuit or fixation, as described above.

[0097] Further information regarding the determination of gaze tracking and error parameters can be found in U.S. Patent No. 10,192,528, the content of which is hereby incorporated by reference for all purposes.

[0098] System FIG. 19 is a block system diagram of a tracking system by DVS according to an aspect of the present disclosure. By way of example and not limitation, according to an aspect of the present disclosure, system 1900 may be an embedded system, a mobile phone, a personal computer, a tablet computer, a portable game device, a workstation, a game console, or the like.

[0099] System 1900 generally includes a central processing unit (CPU) 1903 and a memory 1904. System 1900 may also include well-known support functions 1906 that can communicate with other components of the system via, for example, a data bus 1905. Such support functions may include, but are not limited to, input / output (I / O) elements 1907, a power supply (P / S) 1911, a clock (CLK) 1912, and a cache 1913.

[0100] System 1900 may include a display device 1931 for presenting rendered graphics to a user. In an alternative implementation, the display device is a separate component that functions in cooperation with system 1900. The display device 1931 may be in the form of a flat panel display, a head-mounted display (HMD), a cathode ray tube (CRT) screen, a projector, or other device capable of displaying visible text, numbers, graphic symbols, or images.

[0101] Here, the display device 1931 is coupled to DVS 1901A, and the controller 1902 includes two or more light sources 1932A, which may be of any configuration described herein. In an alternative implementation, the DVS can be coupled to a game controller, and the display device can instead include two or more light sources. In yet another alternative implementation, the DVS is a separate unit decoupled from either the display device or the controller, in which case both the controller and the display device may include two or more light sources for tracking.

[0102] In some implementations where the display device is part of a head-mounted display (HMD), such an HMD can include an inertial measurement unit (IMU) such as an accelerometer or a gyroscope. Also, as discussed above in this specification, such an HMD can include light sources 1932B, and these light sources can be tracked using a DVS that is separate from the display device 1901 and coupled to the CPU 1903. As an example, a separate DVS 1901B can be attached to the controller 1902.

[0103] In some implementations, DVS 1901A or DVS 1901B can be part of a hybrid sensor, as described above with respect to FIGS. 11A, 11B, 11C, or 11D. Such a hybrid sensor can include a depth sensor, such as a DTOF sensor, and in this case, the hybrid sensor can include an illumination unit (not shown).

[0104] Furthermore, when the display device 1931 is part of an HMD, an optional safety shutter 1933 can be externally fitted to this device, which can be operably coupled to a processor such as the CPU 1903 and can operate as described above with respect to FIGS. 16A, 16B, 16C, 16D, or 16E. Alternatively, the safety shutter can be controlled by a separate processor attached to the HMD.

[0105] Also, the system 1900 includes a mass storage device 1915, such as a disk drive, a CD-ROM drive, a flash memory, a solid state drive (SSD), a tape drive, etc., to provide non-volatile storage of programs and / or data. The system 1900 may also optionally include a user interface unit 1916 to facilitate interaction between the system 1900 and the user. The user interface 1916 may include a keyboard, a mouse, a joystick, a light pen, or other devices that can be used with a graphical user interface (GUI). The system 1900 may also include a network interface 1914 to enable the device to communicate with other devices through the network 1920. The network 1920 can be, for example, a local area network (LAN), a wide area network such as the Internet, a personal area network such as a Bluetooth (registered trademark) network, or other types of networks. These components can be implemented in hardware, software, or firmware, or any combination of two or more of these.

[0106] Each of the CPUs 1903 may include one or more processor cores, for example, a single core, two cores, four cores, eight cores, or more. In some implementations, the CPU 1903 may include GPU cores or multiple cores of the same APU (Accelerated Processing Unit). The memory 1904 may be in the form of an integrated circuit that provides addressable memory, such as random access memory (RAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), and the like. The main memory 1904 may include application data 1923 used by the processor 1903 during processing. The main memory 1904 may also include event data 1909 received from the DVS 1901. As described in FIG. 9, a trained neural network (NN) 1910 can be loaded into the memory 1904 to determine position and orientation data. Further, the memory 1904 may include a machine learning algorithm 1921 for training or adjusting the NN 1910. A database 1922 can be included in the memory 1904. The database can include information regarding the light source configuration, the respective predetermined blinking intervals of one or more light sources, and the like. The memory may also include the output from an IMU coupled to the controller 1902 or the display device 1931. In some implementations, the memory 1904 can include the output from one or more light sources, such as a timestamp when each of the light sources is on or off.

[0107] According to aspects of the present disclosure, the processor 1903 can execute methods for determining the position and orientation of a controller or user, as described in FIGS. 8 and 9, and these methods can be loaded into the memory 1904 as the application 1923. As a result of the processor executing the methods described in FIGS. 8 and 9 and further described with respect to FIG. 15, the processor can generate the orientation and configuration of one or more of a controller, a headset, a user's body or appendage, the ground, or a landmark. These positions and orientations can be stored in the database 1922 and used for successive iterations of the methods of FIGS. 8 and 9. In some implementations, the processor 1903 can utilize such positions and / or orientations in a machine learning algorithm trained to perform SLAM (simultaneous localization and mapping).

[0108] The mass storage 1915 can include an application or program 1917 that is loaded into the main memory 1904 when processing is initiated in the application 1923. Further, the mass storage 1915 can include data 1918 used by the processor during the processing of the application 1923, the NN 1910, the machine learning algorithm 1921, and during the filling of the database 1922.

[0109] As used herein and as generally understood by those skilled in the art, an application specific integrated circuit (ASIC) is an integrated circuit that is customized for a particular use rather than for general use.

[0110] As used herein and as generally understood by those skilled in the art, a field programmable gate array (FPGA) is an integrated circuit designed to be configured by a customer or designer after manufacture - and thus "field programmable". FPGA configurations are generally specified using a hardware description language (HDL) similar to that used for ASICs.

[0111] As used herein and as generally understood by those skilled in the art, a system on a chip, i.e., a system-on-chip (SoC or SOC), is an integrated circuit (IC) that integrates all components of a computer or other electronic system onto a single chip. This can include digital, analog, mixed-signal, and often radio frequency functions - all on a single chip substrate. Typical applications are in the field of embedded systems.

[0112] Common SoCs include the following hardware components: One or more processor cores (e.g., a microcontroller, microprocessor, or digital signal processor (DSP) core). Memory blocks, such as read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), and flash memory. A timing source, such as an oscillator or a phase-locked loop. Peripherals, such as a counter timer, real-time timer, or power-on reset generator. External interfaces, such as industry standard specifications like the Universal Serial Bus (USB), FireWire (registered trademark), Ethernet, USART (universal asynchronous receiver / transmitter), Serial Peripheral Interface (SPI) bus, etc. Analog interfaces (including analog-to-digital converters (ADC) and digital-to-analog converters (DAC)). Voltage regulators and power management circuits.

[0113] These components are connected by either a proprietary bus or an industry standard bus. A direct memory access (DMA) controller improves the data throughput of the SoC by directly routing data between the external interface and memory, bypassing the processor core.

[0114] A typical SoC includes both the hardware components described above and executable instructions (such as software or firmware) that control the processor core(s), peripheral devices, and interfaces.

[0115] Aspects of the present disclosure provide image-based tracking that features a higher sample rate than is possible with conventional image-based tracking systems, resulting in improved tracking fidelity. Further advantages include cost reduction, weight reduction, reduction of irrelevant data generation, and reduction of processing requirements when using a DVS-based tracking system. These advantages enable, among other applications, the improvement of virtual reality (VR) and augmented reality (AR) systems.

[0116] The foregoing is a complete description of preferred embodiments of the invention, but various alternatives, modifications, and equivalents may be used. Accordingly, the scope of the invention should not be determined with reference to the foregoing, but instead should be determined with reference to the appended claims, which are to be accorded the full scope of equivalents. Whether or not preferred, any feature described herein may be combined with any other feature described herein, whether or not preferred. In the following claims, the indefinite articles "A" or "An" refer to one or more of the items following the article, unless expressly stated otherwise. The appended claims should not be construed to include means-plus-function limitations, unless such a limitation is expressly recited in a given claim using the phrase "means for".

Claims

1. A processor, a controller operably coupled to the processor, two or more light sources mounted in a configuration known relative to each other and relative to the controller, the two or more light sources being configured to turn on and off within a predetermined time sequence, a dynamic vision sensor (DVS) operably coupled to the processor, the DVS having an array of photosensitive elements in a known configuration, the dynamic vision sensor being configured to output signals corresponding to one or more events at one or more corresponding photosensitive elements within the array in response to changes in the light output from the two or more light sources, the output signals including information corresponding to the time of the one or more events and the position of the one or more corresponding photosensitive elements within the array, the DVS, comprising, the processor is configured to determine an association between each of the one or more events and one or more corresponding specific light sources of the two or more light sources, and to determine occlusion of one or more of the two or more light sources from the association, the processor is configured to estimate the position of one or more objects relative to the controller using the determined occlusion, the known configuration of the two or more light sources relative to each other and relative to the controller body, and the position of the one or more corresponding photosensitive elements within the array, a tracking system.

2. The system of claim 1, wherein the processor is further configured to determine the position or movement of a user's hand or finger relative to the controller using the determined occlusion, the known configuration of the two or more light sources relative to each other and relative to the controller body, and the position of the one or more corresponding photosensitive elements within the array.

3. The system of claim 1, wherein the two or more light sources include one or more light sources proximate to a control element of the controller, and the processor is configured to predict an operation of the control element by the user from the determined occlusion, the known configuration of the two or more light sources relative to each other and relative to the controller body, and the position of the one or more corresponding photosensitive elements within the array.

4. The above two or more light sources include one or more light sources proximate to each of the two or more control elements of the controller, and the processor is configured to predict and enhance a virtual reality (VR) representation of a user's finger from the determined position or movement of the user's hand or finger relative to the controller. The system according to claim 1.

5. Further comprising a headset having a display screen configured to be visible to the user when the user wears the headset, and the processor is configured to cause the VR representation of the user's hand or finger to be presented on the display screen. The system according to claim 4.

6. Further comprising a headset having a display screen configured to be visible to the user when the user wears the headset, the DVS is located on the headset, and the processor is configured to cause the VR representation of the user's hand or finger to be presented on the display screen. The system according to claim 4.

7. The processor is configured to selectively turn off a particular light source among the two or more light sources when the determined occlusion indicates that the particular light source is occluded. The system according to claim 1.

8. The two or more light sources include two or more light sources arranged around a control element. The system according to claim 1.

9. The two or more light sources arranged around the control element include ten or more light sources arranged around the control element. The system according to claim 8.

10. The control element includes another light source, and the processor is further configured to interpret occlusion of the another light source as a button press. The system according to claim 8.

11. A tracking method, Receiving signals corresponding to one or more events at one or more corresponding photosensitive elements in an array of dynamic vision sensors (DVSs) in response to changes in the light output from two or more light sources, the signals including information corresponding to the time of the one or more events and the position of the one or more corresponding photosensitive elements in the array, the two or more light sources being attached in a known configuration relative to each other and relative to a controller, the receiving, Determining an association between each of one or more events and the one or more corresponding specific light sources of the two or more light sources; Determining occlusion of one or more of the two or more light sources from the association; Estimating the position of one or more objects relative to the controller using the determined occlusion, the known configuration of the two or more light sources relative to each other and the controller body, and the position of the one or more corresponding photosensitive elements in the array; A method comprising. **Claim 12** The method of claim 11, further comprising determining the position or movement of a user's hand or finger relative to the controller using the determined occlusion, the known configuration of the two or more light sources relative to each other and the controller body, and the position of the one or more corresponding photosensitive elements in the array. **Claim 13** The method of claim 11, further comprising predicting an operation of a control element of the controller by the user from the determined occlusion, the known configuration of the two or more light sources relative to each other and the controller body, and the position of the one or more corresponding photosensitive elements in the array, wherein the two or more light sources include one or more light sources proximate to the control element of the controller. **Claim 14** The method of claim 11, further comprising predicting and enhancing a virtual reality (VR) representation of the user's finger from the determined position or movement of the user's hand or finger relative to the controller, wherein the two or more light sources include one or more light sources proximate to each control element of the two or more control elements of the controller. **Claim 15** The method of claim 14, further comprising causing the VR representation of the user's hand or finger to be presented on a display screen, wherein the display screen is coupled to a headset and is configured to be visible to the user when the user wears the headset. **Claim 16** The method of claim 14, further comprising causing the VR representation of the user's hand or finger to be presented on a display screen, wherein the display screen is coupled to a headset and is configured to be visible to the user when the user wears the headset, and the DVS is located on the headset. **Claim 17** The method according to claim 11, further comprising determining the shielding of a specific light source among the two or more light sources, and selectively turning off the specific light source among the two or more light sources.

18. The method according to claim 11, wherein the two or more light sources include two or more light sources arranged around a control element.

19. The method according to claim 18, wherein the two or more light sources arranged around the control element include 10 or more light sources arranged around the control element.

20. The method according to claim 18, further comprising interpreting the shielding of another light source as a button press, wherein the control element includes the another light source.

Citation Information

Patent Citations

  • Device and method for displaying composite sense of reality, recording medium and computer program

    JP2003256876A

  • Simulation system, program and controller

    JP2018126340A

  • Systems and methods for detecting objects within the boundary of a defined space while in artificial reality

    US20210319220A1