Asynchronous dynamic visual sensor LED AI tracking system and method
The use of a dynamic vision system with high update rates and IMUs, combined with machine learning, addresses the limitations of existing VR and AR tracking systems by providing efficient and accurate motion tracking for game controllers.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-04-12
- Publication Date
- 2026-03-10
AI Technical Summary
Modern VR and AR implementations face challenges in accurate and fast motion tracking due to the limitations of existing infrared camera frame rates and the need for expensive hardware to process large amounts of data, which is not suitable for modern inside-out detection methods.
Utilizing a dynamic vision system (DVS) with high update rates to detect changes in light intensity from multiple light sources, combined with inertial measurement units (IMUs) and machine learning algorithms, to track game controllers with reduced data processing requirements.
Provides accurate and efficient motion tracking with reduced data processing needs, enabling smooth motion feedback and supporting modern inside-out detection methods without the need for expensive hardware.
Smart Images

Figure 0007827887000001 
Figure 0007827887000002 
Figure 0007827887000003
Abstract
Description
[Technical Field]
[0001] Aspects of the present disclosure relate to tracking game controllers, and more particularly, aspects of the present disclosure relate to tracking game controllers using dynamic vision sensors. [Background technology]
[0002] Modern virtual reality (VR) and augmented reality (AR) implementations rely on accurate and fast motion tracking for user interaction with the device. AR and VR often rely on information about the position and orientation of the controller relative to other objects. Many VR and AR implementations rely on a combination of inertial measurements obtained by accelerometers or gyroscopes within the controller and visual detection of the controller by an external camera to determine the controller's position and orientation.
[0003] Some of the earliest implementations use infrared light detected by an infrared camera with a defined detection radius on a game controller pointed at the screen. The camera captures images at a moderately fast rate of 200 frames per second to determine the location of the infrared light. The distance between the infrared lights is predetermined, and from the relative position of the infrared light in the camera image, the position of the controller relative to the screen can be calculated. Accelerometers may also be used to provide information about relative three-dimensional changes in the controller's position or orientation. These traditional implementations rely on a fixed position of the screen and the controller pointed at the screen. Modern VR and AR implementations allow the screen to be placed near the user's face in a head-mounted display that moves with the user. Therefore, having an absolute light position (also known as a lighthouse) becomes undesirable because it requires extra setup time and forces the user to set up an independent lighthouse point, limiting the user's range of motion. Furthermore, even the moderately fast frame rate of 200 frames per second of an infrared camera is not fast enough to provide smooth motion feedback. Furthermore, this setup is too simplistic and unsuitable for more modern inside-out detection methods such as room mapping and hand detection.
[0004] Recent implementations use cameras and accelerometers in combination with trained machine learning algorithms that are trained to detect hands, controllers, and / or other body parts. Smoothness of motion detection requires the use of a high-frame-rate camera to generate image frames for body part / controller detection. This generates a large amount of data that must be processed quickly for smooth update rates. Therefore, expensive hardware must be used to process the frame data. Furthermore, much of the frame data within each frame is discarded as unnecessary because it is not relevant to motion tracking.
[0005] It is in this context that aspects of the present disclosure arise. Summary of the Invention
[0006] The teachings of the present invention can be readily understood by considering the following detailed description in conjunction with the accompanying drawings, in which: [Brief explanation of the drawings]
[0007] [Figure 1] 1 is a diagram illustrating an implementation of tracking a game controller using a DVS with a single sensor array, according to one aspect of the disclosure. [Figure 2] 1 is a diagram illustrating an implementation of tracking a game controller using a DVS with a dual sensor array, according to one aspect of the disclosure. [Figure 3] 1 is a diagram illustrating an implementation of game controller tracking using a DVS in combination with a single sensor array and a camera, according to one aspect of the present disclosure. [Figure 4] 1 is a diagram illustrating a DVS tracking the movement of a game controller with two or more light sources, according to one embodiment of the present disclosure. [Figure 5] 1 is a diagram illustrating an implementation of head tracking or other device tracking using a controller including a DVS with a single sensor array, according to one aspect of the disclosure. [Figure 6] 1 is a diagram illustrating an implementation of head tracking or other device tracking using a game controller including a DVS with a dual sensor array, according to one embodiment of the present disclosure. [Figure 7] 1 is a diagram illustrating an implementation of head tracking or other device tracking using a controller with a DVS in combination with a single sensor array and camera, according to one aspect of the present disclosure. [Figure 8] FIG. 1 is a flow diagram illustrating a method for motion tracking with DVS using one or more illuminants and a conformal model of an illuminant, according to one embodiment of the present disclosure. [Figure 9] FIG. 1 is a flow diagram illustrating a method for motion tracking with DVS using time-stamped light source position information, according to one aspect of the present disclosure. [Figure 10A]1 is a diagram illustrating a basic form of an RNN having layers of nodes, each of which is characterized by an activation function, one input weight, a regression hidden node transition weight, and an output transition weight, according to an embodiment of the present disclosure. [Figure 10B] FIG. 1 is a simplified diagram illustrating that an RNN can be viewed as a series of nodes with the same activation function moving over time, according to aspects of the present disclosure. [Figure 10C] 1 illustrates an example layout of a convolutional neural network, such as a CRNN, according to an embodiment of the present disclosure. [Figure 10D] 1 shows a flow diagram depicting a method for supervised training of a machine learning neural network, according to an aspect of the present disclosure. [Figure 11A] 1 is a diagram illustrating a hybrid DVS with multiple co-location sensor types, according to an embodiment of the present disclosure. [Figure 11B] 1 is an illustration of a hybrid DVS having multiple sensor types arranged in a checkerboard pattern in an array, according to an embodiment of the present disclosure. [Figure 11C] 1 is a schematic cross-sectional view of a hybrid DVS having multiple sensor types arranged in a pattern in an array, according to an embodiment of the present disclosure. FIG. [Figure 11D] 1 is a schematic cross-sectional view of a hybrid DVS having multiple filter types arranged in a pattern in an array, according to an embodiment of the present disclosure. [Figure 12] 1 is a diagram illustrating a hybrid DVS including multiple sensor types with inputs separated by optical separators, according to an embodiment of the present disclosure. [Figure 13] 1 is a diagram illustrating a hybrid DVS in which multiple sensor types have inputs separated by micro-electromechanical (MEMS) mirrors, according to an embodiment of the present disclosure. [Figure 14] 1 is a diagram illustrating a hybrid DVS in which multiple sensor types have temporally filtered inputs, according to an embodiment of the present disclosure. [Figure 15] 1 is an illustration showing body tracking with DVS, according to an aspect of the present disclosure. [Figure 16A]1 is an illustration showing a headset with a safety shutter door according to an aspect of the present disclosure. [Figure 16B] 1 is an illustration showing a headset with a sliding safety shutter according to an aspect of the present disclosure. [Figure 16C] 1 is an illustration showing a headset with a louvered safety shutter according to an aspect of the present disclosure. [Figure 16D] 1 is an illustration showing a headset with a fabric safety shutter according to an aspect of the present disclosure. [Figure 16E] 1 is an illustration showing a headset with a liquid crystal safety shutter according to an aspect of the present disclosure. [Figure 17] 1 is an illustration showing finger tracking by a DVS and a controller, according to an aspect of the present disclosure. [Figure 18A] FIG. 1 is a schematic diagram illustrating gaze tracking within the context of aspects of the present disclosure. [Figure 18B] FIG. 1 is a schematic diagram illustrating gaze tracking within the context of aspects of the present disclosure. [Figure 19] FIG. 1 is a block system diagram of a DVS-based tracking system according to an aspect of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0008] Although the following detailed description contains many specific details for purposes of illustration, those skilled in the art will recognize that many variations and modifications to the following details are within the scope of the invention. Accordingly, the exemplary embodiments of the invention described below are set forth without loss of generality to, and without imposing limitations on, the claimed invention.
[0009] preface A new type of vision system called a dynamic vision system (DVS) has recently been developed. This DVS uses only changes in light intensity of an array of light-sensitive pixels to resolve changes in a scene. A DVS has a very fast update rate, and instead of delivering a stream of image frames, a DVS provides a nearly continuous stream of the locations of changes in pixel intensity. Each change in pixel intensity is sometimes called an event. This has the added benefit of greatly reducing extraneous data output.
[0010] The two or more light sources can provide continuous updates on the DVS camera's position relative to the position indicator, with an update rate determined by the flashing rate of the light. In some implementations, the two or more light sources can be infrared light sources, and the DVS can use infrared-sensitive pixels. Alternatively, the DVS can be sensitive to the visible light spectrum, and the two or more light sources can be multiple visible light sources or a single visible light at a known wavelength. In implementations with a DVS sensitive to visible light, the DVS can also be sensitive to motion occurring within its field of view (FOV). The DVS can detect changes in light intensity caused by the reflection of light from a moving surface. In implementations using an infrared-sensitive DVS, the light from an infrared illuminator can be used to detect motion within the FOV by reflection.
[0011] implementation FIG. 1 illustrates an example implementation of tracking a game controller using a DVS 101 with a single sensor array, according to one embodiment of the present disclosure. In the illustrated implementation, the DVS is attached to a headset 102, which may be part of a head-mounted display. A controller 103, which includes two or more light sources, is within the field of view of the DVS 101. In the illustrated example, the controller 103 includes four light sources 104, 105, 106, and 107. These light sources have a known configuration relative to each other and relative to the controller 103. Here, we have one DVS with a single photosensitive array. Using these four light sources, the position and orientation of the controller 103 relative to the DVS 101 can be accurately determined. Known information about the light sources can include the distance between each of the other light sources relative to each other and the respective positions of the light sources on the controller 103. As shown, the three light sources 104, 105, and 106 can describe a plane, and the light source 107 can be out of plane relative to the plane described by the three light sources 104, 105, and 106. The light sources here have known configurations, such as, but not limited to, a first light source 104 located at the top left of the front, a second light source 105 located at the top right of the front, a third light source 106 located at the left side away from the top, and a fourth light source 107 located at the bottom center of the front of the controller. With four light sources, a DVS with a single photosensitive array may be able to determine the motion of the controller in the X, Y, and Z axes. Additionally, an inertial measurement unit (IMU) 108 may be coupled to the controller 103. By way of example, the IMU 108 may include an accelerometer configured to measure acceleration about one, two, or three axes. Alternatively, the IMU may include a gyroscope configured to sense changes in rotation about one, two, or three axes. In some implementations, the IMU may include both an accelerometer and a gyroscope. The IMU 108 may be used to refine the determination of motion, position, and orientation with information from the DVS 101 using a processor. The processor may be located within the headset 102, a game console, or another computing device (not shown).The DVS 101, headset 102, and IMU 108 may be operably coupled to a processor 110, which may be located on the headset 102, the controller 103, or a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor may implement tracking as described herein, for example, as described below with respect to Figures 4, 8, and 9. Additionally, the processor 110 may control the blinking of the light sources 104, 105, and 106.
[0012] During operation, the DVS 101 with a photosensitive array can detect movement of the light sources 104, 105, 106, 107 with the photosensitive array, and changes in light detected by the photosensitive array can be transmitted to a processor. In some implementations, the light sources can be configured to turn on and off in a predetermined pattern, for example, without limitation, using signals from the circuitry and / or processor. The processor can use the predetermined pattern to determine the identity of each light source. The identity of the light source can include a known position relative to the controller and relative to other light sources. In other implementations, each light source can be configured to turn off and on in a predetermined pattern, and the pattern can be used to determine the identity of that specific light source. In some implementations, the processor can match the known configuration of the light sources relative to the controller to events detected by the photosensitive array.
[0013] A DVS can have a nearly continuous update rate, which can be discretely approximated to about 1 million updates per second. A DVS with a high update rate may be able to resolve very fast blinking patterns of a light source. The blinking rate is primarily limited by the Nyquist frequency, which is half the sample rate of the DVS. The light source can blink with a duty cycle suitable for detection of blinking by the DVS. Generally speaking, the "on" time of the blink needs to be long enough that it can be consistently detected by the DVS. Furthermore, because the update rate is high, slight differences in the blinking rate may be detectable.
[0014] The light sources 104, 105, 106, and 107 may be broad visible spectrum light sources, such as incandescent lamps or white light-emitting diodes. Alternatively, the light sources 104, 105, 106, and 107 may be infrared light sources, or the light sources may have a specific light spectral profile detectable by the DVS 101. The DVS 101 may include a photosensitive array configured to detect the emission spectrum of the light sources 104, 105, 106, and 107. For example, but not by way of limitation, if the light sources are infrared light sources, the photosensitive array of the DVS may be sensitive to infrared light, or if the light sources have a specific emission spectrum, the photosensitive array may be configured to enhance sensitivity to the specific emission spectrum of the light source. Additionally, for example, but not by way of limitation, the photosensitive array of the DVS may be insensitive to or exclude light of other wavelengths not emitted by the light source; for example, if the light source is infrared light, the photosensitive array may be configured to detect only infrared light.
[0015] FIG. 2 illustrates an example implementation of tracking a game controller using a DVS with dual sensor arrays according to one embodiment of the present disclosure. In this implementation, a headset 203 includes a first DVS 201 and a second DVS 202. Alternatively, the headset 203 may include a DVS with a first photosensitive array 201 and a second photosensitive array 202. The general functionality of the light source and DVS is similar to that described above with reference to FIG. 1. Information from the second DVS or second array can be integrated with information from the first array to provide a better fit for controller orientation and some depth information. The two DVS or photosensitive arrays can have overlapping fields of view, allowing for the use of binocular parallax.
[0016] Two DVSs or two photosensitive arrays provide binocular vision for depth sensing, which can further reduce the number of light sources. The first and second DVSs or first and second photosensitive arrays may be separated by a known distance, for example, but not limited to, about 50-100 millimeters or more than 100 millimeters. More typically, this spacing is large enough to provide sufficient parallax for the desired depth sensitivity, but not so large that there is no overlap between the fields of view. As shown, the controller 207 may include a first light source 204, a second light source 205, and a third light source 206. A fourth light source coupled to the controller may not be necessary, as information from the two DVSs or two arrays provides sufficient information for determining the controller's position and orientation. The controller may include an IMU 208, which can provide additional inertial information used to refine the position and orientation determination.
[0017] The first DVS 201, the second DVS 202, the headset 203, and the IMU 208 may be operably coupled to a processor 210, which may be located on the headset 203, the controller 207, or a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor may implement tracking as described herein, for example, as described below with respect to Figures 4, 8, and 9. Additionally, the processor 210 may control the blinking of the light sources 204, 205, and 206.
[0018] While FIG. 2 shows two DVSs or two photosensitive arrays, aspects of the present disclosure are not so limited. A device can include any number of DVSs or photosensitive arrays. For example, without limitation, a three DVS or a DVS with three separate photosensitive arrays may allow the use of only two light sources coupled to a controller. The third DVS or photosensitive array may not be collinear with the other two DVSs or arrays. Similar to binocular parallax, each of the additional DVSs or photosensitive arrays can be separated by a known distance and may have overlapping fields of view, allowing for greater parallax effects to be used. Furthermore, some implementations may include multiple DVSs, each with multiple photosensitive arrays. For example, without limitation, there may be two DVSs, each with two separate photosensitive arrays.
[0019] FIG. 3 illustrates an example implementation of tracking a game controller using a DVS in combination with a single sensor array and a camera, according to one embodiment of the present disclosure. In this implementation, a camera 302 is added to a DVS 301. The DVS 301 and camera 302 may be coupled to a headset 303. The DVS 301 and camera 302 may have overlapping fields of view or may share the same field of view. The controller may include three or more light sources 304, 305, and 306 on the controller 307. The DVS 301 and camera 302 can be used together to determine the position and orientation of the controller. Frames from the camera 302 can be interpolated using events from the DVS 301. The image frames can also be used to simultaneously perform localization and mapping to improve the determination of the controller's orientation and position. Additionally, the image frames from the camera can be used to perform inside-out tracking of the user using machine learning algorithms, such as hand tracking or foot tracking. An IMU 308 can provide additional inertial information used to further refine the determination of the controller's position and orientation.
[0020] The first DVS 301, the second DVS 302, the headset 303, and the IMU 308 may be operably coupled to a processor 310, which may be located on the headset 303, the controller 307, or a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor may implement tracking as described herein, for example, as described below with respect to Figures 4, 8, and 9. Additionally, the processor 310 may control the blinking of the light sources 304, 305, and 306.
[0021] FIG. 5 illustrates an example implementation of head tracking or other device tracking using a controller including a DVS with a single sensor array, according to one embodiment of the present disclosure. Here, the DVS 507 is attached to a controller 508. A headset 501 within the field of view of the DVS 507 is tracked using light sources attached to the headset. The headset 501 may include two or more light sources, here four light sources 502, 503, 504, and 505. Four light sources provide accurate information for determining three-dimensional position with a single DVS. Using additional DVSs or cameras may reduce the number of light sources used. The position of each light source relative to the headset 501 is known to the system and may be somewhat important for providing relevant information to the DVS. Here, light sources 502, 503, and 504 represent a plane. Light source 505 is positioned outside the plane of the other light sources 502, 503, and 504. This facilitates three-dimensional detection of the headset's position, orientation, or movement. The headset 501 may include an IMU 506, which may use information from the DVS 507 to improve position and orientation estimation. Additionally, the controller 508 may include an IMU 509. Information from the controller IMU 509 may be used to further refine the determination of the headset's position and orientation, for example, but not by way of limitation, the IMU information may be used to determine whether the controller is moving relative to the headset and to determine the rate or acceleration of that movement.
[0022] The headset 501, IMU 506, DVS 507, and controller IMU 509 may be operably coupled to a processor 510, which may be located on the headset 501, the controller 508, or a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor may implement tracking as described herein, for example, as described below with respect to Figures 4, 8, and 9. Additionally, the processor 510 may control the blinking of the light sources 502, 503, 504, and 505.
[0023] FIG. 6 illustrates an example implementation of head tracking or other device tracking using a game controller including a DVS with dual sensor arrays, according to one embodiment of the present disclosure. Here, a controller 605 is coupled to two DVSs 606, 607, or a single DVS with two photosensitive arrays 606 and 607. As described above, the two DVSs or two photosensitive arrays can be separated by a suitable distance, e.g., between 500 and 1000 millimeters, and can have overlapping fields of view. This allows for the use of parallax effects in determining depth. Additionally, a headset 601 can include three light sources 602, 603, and 604 and an IMU 608. The use of two DVSs or two separated arrays 606, 607 may allow for the use of fewer than four light sources, such as, but not limited to, three light sources. Two light sources 602, 603 can form a line, and a third light source 604 can be outside the line formed by the other two light sources 602, 603. Information from an IMU 605 coupled to the headset 601 can be used to refine the position and orientation determination.
[0024] Headset 601, DVS 606, DVS 607, and IMU 608 may be operably coupled to processor 610, which may be located on headset 601, controller 605, or a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor may implement tracking as described herein, for example, as described below with respect to Figures 4, 8, and 9. Additionally, processor 610 may control the blinking of light sources 602, 603, 604.
[0025] 7 shows an example of an implementation of head tracking or other device tracking using a controller including a DVS in combination with a single sensor array and image camera, according to one embodiment of the present disclosure. Here, the DVS 707 and camera 708 are coupled to a controller 706. The camera 708 and DVS 707 may have partially overlapping or fully overlapping fields of view. In some implementations, the pixels of the camera and the photosensitive elements of the DVS may share the same photosensitive array, thereby functioning as an integrated DVS and camera.
[0026] Three or more light sources 702, 703, 704 can be coupled to the headset 701. For example, without limitation, the light sources may be integrated within the headset housing, and each light source may be an LED, incandescent, halogen, or fluorescent light source mounted on a circuit board within or on the headset housing. In some implementations, a single light source may create multiple light sources using a plastic or glass light pipe or optical fiber that splits the light from the single light source to two or more light sources on the headset housing.
[0027] The three or more light sources 702, 703, 704 may be configured to turn on and off in response to an electronic signal. In some implementations, the three or more light sources can turn on and off in a predetermined sequence. Each time a light source 702, 703, 704 moves or flashes, the DVS 707 can generate an event. The camera 708 generates image frames of its field of view at a set frame rate. The high update rate of the DVS may allow the image frames generated by the camera to be interpolated with DVS events.
[0028] In some implementations, the IMU 705 can also be coupled to the headset 701. The headset 701, IMU 705, DVS 707, camera 708, and IMU 608 can be operably coupled to a processor 710, which may be located on the headset 701, the controller 706, or a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor can implement tracking as described herein, for example, as described below with respect to FIGS. 4, 8, and 9. Additionally, the processor 710 can control the blinking of the light sources 702, 703, and 704.
[0029] Further, in some implementations, the camera 708 may be a depth camera, such as a depth time-of-flight (DTOF) sensor. A DToF camera acquires depth images by measuring the time it takes light to travel from a light source to an object in a scene and return to a pixel array. By way of example and not limitation, a DToF camera may operate using continuous wave (CW) modulation, which is an example of an indirect time-of-flight (ToF) sensing method. In a CW ToF camera, light from an amplitude-modulated light source is backscattered by objects within the camera's field of view (FOV), and the phase shift between the emitted and reflected waveforms is measured. By measuring the phase shift at multiple modulation frequencies, a depth value for each pixel can be calculated. The phase shift is obtained by measuring the correlation between the emitted and received waveforms at different relative delays using in-pixel photon mixing demodulation.
[0030] A DTOF system typically includes an illumination module and an imaging module. The illumination module consists of a light source, a driver that drives the light source at a high modulation frequency, and a diffuser that projects a light beam from the light source onto a designed field of illumination (FOI). The DToF illumination module can include one or more light sources, which may be amplitude-modulated emitters such as, but not limited to, vertical-cavity surface-emitting lasers (VCSELs) or edge-emitting lasers (EELs). The imaging module can include an imaging lens assembly, a band-pass filter (BPF), a microlens array, and an array of photosensitive elements that convert incident photon energy into electronic signals. The microlens array increases the amount of light reaching the photosensitive elements, while the BPF reduces the amount of ambient light reaching the photosensitive elements and the microlens array.
[0031] operation FIG. 4 illustrates a DVS tracking the movement of a game controller with two or more light sources, according to one embodiment of the present disclosure. The DVS 401 has a controller 402 within its field of view. As shown, the controller 402 includes multiple light sources coupled to the controller body. The multiple light sources can be configured to turn off and on again at a predetermined rate. Each flashing of a light source within the field of view of the DVS 401 can generate one or more events 403 in the DVS. In the illustrated event 403, the light sources all changed from an off state to an on state, so the event indicates that all light sources are in an on state. Alternatively, depending on the sensitivity of the DVS, a change in the brightness of the light sources may be sufficient to trigger the event 403. Note that in other implementations, each event may correspond to fewer than all light sources being illuminated, since the lights may turn off and on at different rates or for different times. The times of events generated by the flashing lights and their corresponding positions within the array can be recorded in memory (not shown). In some implementations, the position and orientation of the controller 402 can be reconstructed from one or more events from the DVS.
[0032] In some implementations, each light source may flash in a predetermined time sequence. The DVS may output the time each event occurred, for example, but not by way of limitation, as a timestamp along with each event. The predetermined time sequence may be used to determine the identity of each light within the event, e.g., which event corresponds to which light source position. In the example shown, one or more events 403 output by the DVS 401 indicate a first light 406, a second light 407, a third light 408, and a fourth light 409 detected by the photosensitive array. As described above, the identity of each light source may be determined from the information output by the DVS and the predetermined flash sequence of the light source. For example, but not by way of limitation, the photosensitive array of the DVS 401 may detect light event 403 at time T+1, and the predetermined sequence may provide that light source 406 is turned on at T+1, thereby determining that the light event corresponds to light source 406. The predetermined sequence may be stored in memory, for example, as a table listing the timing and position of each light source's sequence. In some implementations, the predetermined sequence may be encoded into the blinking of the light itself; for example, without limitation, each light may blink in a sequence that indicates its identity. For example, a light source labeled 1 may blink in a Morse code sequence that indicates the number 1. The identity of the light event may then be recovered through analysis of the light event. Alternatively, the sequence information may come from the light source itself or the driver of the light source and indicate when the light source is turned on or off. Alternatively, the light sources may be turned on and off simultaneously, and a machine learning algorithm may be applied to the detected light events 403 to match the pose of the controller 402 to the event and the known configuration of the light sources.
[0033] Additionally, information from the events, such as the size, intensity, and separation of the light, can be used to determine the orientation and position of the light sources. The system may have information defining the size, location on the controller body, and intensity of each light source. Differences between the detected sizes, intensities, and separations may determine the position and orientation with greater accuracy. Furthermore, if one or more additional DVSs or photosensitive arrays are present, disparity information can be used to further enhance the determination of position and orientation.
[0034] During operation, a user can change the position and orientation of the controller 410. The relatively high update rate of the DVS allows the DVS to capture an event sequence at 411 as the light events move at 412 during changes in position and orientation. The movement of the light events here is illustrated in FIG. 4 by the vector arrows 412. The determined position and orientation of the controller 410 may be updated as the detected light events move, or the determined position and orientation may be updated at regular intervals. Due to the flashing of the first light 415, the second light 416, the third light 417, and the fourth light 418, the position and time of the events detected by the photosensitive elements of the DVS may be adapted to the new position and orientation of the controller. Additionally, inertial information from the IMU can be used to refine the movement and position and orientation of the controller 410.
[0035] The flow diagram shown in FIG. 8 illustrates a method for tracking motion with a DVS 801 using one or more light sources and a light source configuration adaptation model, according to one embodiment of the present disclosure. In this implementation, the light sources (shown as LEDs) can be turned on and off simultaneously, independently, or in a predetermined sequence. Light source movement or the blinking of one or more of the light sources generates an event 802 from the DVS 801. Generally, an event includes an electronic signal that relays the following information: the time interval during which the event occurred, the location within the array of light-sensitive elements of the DVS 801 during which the event occurred, and binary data (e.g., 1 or 0) corresponding to a change in light intensity above some detection threshold. Each event can be processed in 803 to format the event into a usable format, such as, but not limited to, placing the event location within a data array, associating a timestamp with the event, and aggregating multiple events. As an example of aggregating events, all events occurring at individual elements within the array within a predetermined time interval can be combined into a single data structure for analysis.
[0036] The processed events can then be analyzed to associate detected DVS events with corresponding LED pulses 804. In some implementations, because each light source can turn off and on at unique, predetermined time intervals, the timestamps of the aggregated events can be used to determine the predetermined time interval from the event and associate a particular light source with a particular event. For example, without limitation, the spatial pattern of aggregated events occurring within a predetermined time interval or sequence of time intervals can be analyzed to determine whether this pattern is consistent with an LED pulse. Event patterns that are too large, too small, or too irregular in shape may be filtered out as LED events. Additionally, the timing of the events can be analyzed, and events that are too short or too long may be filtered out.
[0037] A trained machine learning model 805 may be applied to the processed events. The model 805 may include information about the configuration of the light sources, such as the size of the light sources and their relative position with respect to the controller body. The machine learning model may be trained with training event data having corresponding masked positions and orientations of the controller, as described in a later section. The trained machine learning model is applied to the processed event data to determine correspondences 806 between detected pulses 804 and poses 808. The trained machine learning model may match the poses 808, e.g., position and orientation, of the controller to one or more processed events, e.g., detected LED pulses 804. Alternatively, instead of the trained machine learning model 805, a fitting algorithm may be applied to the processed events. The fitting algorithm may use a manually developed light source model to fit the position and orientation of the controller to the processed events. Alternatively, the fitting algorithm may be a hypothesis-and-test type algorithm that tries all possible permutations of light correspondences and finds the best match using redundant light sources. A tracking / prediction algorithm may then be applied to continue tracking the light sources. Additionally, the predicted current pose may be used to predict its next pose at 809. At 810, inertial data from the IMU 807 can be fused with the predicted pose 808 to generate a final predicted position and orientation of the controller. Fusion may be performed by a trained machine learning algorithm that is trained to use the inertial data to refine the position and orientation of the controller. Alternatively, fusion may be performed by, for example, but not limited to, a Kalman filter or nonlinear optimization.
[0038] FIG. 9 illustrates a method for tracking motion using a DVS using time-stamped light source position information, according to one embodiment of the present disclosure. In this implementation, the times of events output by the DVS 901 are used to determine the position and orientation of the controller. Here, the light sources are depicted as LEDs, and each LED may turn on at a different time as shown in the graph. Each time an LED turns on or off, a DVS with the LED in its field of view can generate an event 902. Each event may include an electronic signal corresponding to information such as the time interval during which the event occurred, the location within the array of light-sensitive elements of the DVS 901 at which the event occurred, and binary data (e.g., 1 or 0) corresponding to a change in light intensity above some detection threshold. Each event can be processed in 903 to format it into a usable format, such as, but not limited to, placing the event location within a data array, as described above, and aggregating multiple events occurring at individual elements within the array within a given time interval into a single data structure. The processed events can then be analyzed to detect a corresponding LED pulse in 904. For example, but not by way of limitation, the shape and size of events may be analyzed for regularity and conformance to the light source, and events that are too large, too small, or too irregular in shape may be filtered out as LED events, and noise suppression such as event averaging may be performed to remove random events. Additionally, the timing of events may also be analyzed, and events that are too short or too long may be filtered out, or multiple events may be condensed into a single event by filtering out events that co-occur with and end in time but after an initial event.
[0039] Once the events are processed, individual LED positions may be determined at 905. Determining individual LED positions may be performed by using the time sequence in which the LEDs turn on and off. For example, without limitation, the time of one or more events may be compared to a known time sequence of LED blinking. The known time sequence may be, for example, a table with LED on and off times and the location on the controller body for each LED, or a timestamp from the LED driver when each LED is on or off. If timestamps are used, the timestamps may be correlated with the LED positions on the controller body. From the timing sequence, the LED position information, and the processed event information, a matching position and orientation of the controller may be determined. IMU data 907 may be integrated with previously determined LED positions and inertial data from the IMU by a Kalman filter 908. The Kalman filter may predict the position of the light source based on inertial information from the IMU, and this prediction may be integrated with position information determined from the LED time sequence to refine the motion data, generating a final pose at 909 and refining future estimates.
[0040] General neural network training According to aspects of the present disclosure, the tracking system may use machine learning with a neural network (NN). For example, the trained model 805 discussed above may use machine learning as discussed below. The machine learning algorithm may use a training dataset, which may include inputs from the DVS, such as events, or processed events with known controller positions and orientations as labeling. Additionally, the machine learning algorithm using the NN may perform fusion between the controller positions and orientations determined from the DVS information and inertial information from the IMU. The training set for fusion may be, for example, but not limited to, potential controller positions and orientations and inertial data with final positions and orientations. In some implementations, the machine learning algorithm may be trained to perform simultaneous localization and mapping (SLAM) using a training set including objects such as ground planes, landmarks, and body parts with hidden labels. The hidden labels may include the identities of the objects and their relative positions. As generally understood by those skilled in the art, SLAM techniques generally solve the problem of building or updating a map of an unknown environment while simultaneously keeping track of the position of an agent within it.
[0041] The NN may include one or more of several different types of neural networks and may have many different layers. By way of example and not limitation, the neural network may be composed of one or more convolutional neural networks (CNNs), recurrent neural networks (RNNs), and / or dynamic neural networks (DNNs). The motion detection neural network may be trained using the general training methods disclosed herein.
[0042] By way of example and not limitation, FIG. 10A illustrates a basic form of an RNN that may be used, for example, in the trained model 805. In the illustrated example, the RNN has a layer of nodes 1020, each characterized by an activation function S, one input weight U, a recurrent hidden node transition weight W, and an output transition weight V. The activation function S may be any nonlinear function known in the art and is not limited to the hyperbolic tangent (tanh) function. For example, the activation function S may be a sigmoid function or a ReLu function. Unlike other types of neural networks, an RNN has one set of activation functions and weights for the entire layer. As shown in FIG. 10B, an RNN may be thought of as a series of nodes 1020 with the same activation function across times T and T+1. Thus, the RNN maintains historical information by feeding results from the previous time T to the current time T+1.
[0043] In some implementations, a convolutional RNN may be used. Another type of RNN that can be used is a long-short-term memory (LSTM) neural network, which adds memory blocks to the RNN nodes with input gate activation functions, output gate activation functions, and forget gate activation functions, resulting in a gating memory that allows the network to retain some information for a longer period of time, as described by Hochreiter & Schmidhuber, "Long Short-term memory," Neural Computation 9(8):1735-1780 (1997), which is incorporated herein by reference.
[0044] FIG. 10C illustrates an exemplary layout of a convolutional neural network, such as a CRNN, that may be used in a trained model 805, according to embodiments of the present disclosure. In this representation, a convolutional neural network is generated for an input 1032 having a size of 4 units in height and 4 units in width, giving a total area of 16 units. The depicted convolutional neural network has a filter 1033 having a size of 2 units in height and 2 units in width, with a skip value of 1 and channels 136 of size 9. For clarity in FIG. 10C, only the connections 1034 between the first row of channels and their filter windows are depicted. However, embodiments of the present disclosure are not limited to such implementations. According to embodiments of the present disclosure, a convolutional neural network may have any number of additional neural network node layers 1031 and may include such layer types as additional convolutional layers, fully connected layers, pooling layers, max-pooling layers, local contrast normalization layers, etc., of any size.
[0045] As seen in Figure 10D, training a neural network (NN) begins with the initialization of the NN's weights at 1041. Generally, the initial weights should be randomly distributed. For example, a NN with a tanh activation function should have random values distributed between -1 / √n and 1 / √n, where n is the number of inputs to the node.
[0046] After initialization, activation functions and optimizers are defined. Then, at 1042, the NN is provided with feature vectors or input datasets. Each different feature vector may be generated by the NN from inputs with known labels. Similarly, the NN may be provided with feature vectors corresponding to inputs with known labeling or classification. The NN then predicts labels or classifications for the features or inputs at 1043. The predicted labels or classes are compared to the known labels or classes (also known as ground truth) at 1044, and a loss function measures the total error between the prediction and the ground truth across all training samples. By way of example and not limitation, the loss function may be a cross-entropy loss function cost, a triplet contrasitive function, an exponential cost, etc. Multiple different loss functions may be used depending on the purpose. By way of example and not limitation, to train a classifier, a cross-entropy loss function may be used, whereas to learn a pre-trained embedding, a triplet contrasitive function may be employed. The NN may then be optimized and trained using the results of the loss function and using known methods for training neural networks, such as adaptive gradient descent backpropagation, as shown at 1045. At each training epoch, the optimizer attempts to select model parameters (i.e., weights) that minimize the training loss function (i.e., total error). The data is partitioned into training samples, validation samples, and test samples.
[0047] During training, an optimizer minimizes a loss function on the training samples. After each training epoch, the model is evaluated on the validation samples by computing the validation loss and accuracy. If there is no significant change, training may be stopped and the resulting trained model may be used to predict labels for test data.
[0048] Thus, neural networks may be trained to identify and classify inputs with known labels or classifications from those inputs. Similarly, the described methods may be used to train NNs to generate feature vectors from inputs with known labels or classifications. While the above discussion pertains to RNNs and CRNNSs, these discussions can be applied to NNs that do not include recurrent or hidden layers.
[0049] Hybrid Sensor FIG. 11A is a diagram illustrating a hybrid DVS with multiple co-located sensor types, according to embodiments of the present disclosure. In some implementations, a hybrid DVS can be used to combine multiple sensor types into one device. The multiple sensor types can be, for example, but not limited to, DVS infrared light sensitive elements, DVS visible light sensitive elements, DVS wavelength-specific light sensitive elements, visible light camera pixels, infrared camera pixels, and DTOF camera pixels. A first sensor type 1102 can be interspersed with a second sensor type 1102 on the same array 1101. For example, but not limited to, DVS visible light sensitive elements 1103 can surround visible light camera pixels 1102, or DVS infrared light sensitive elements 1103 can surround DVS visible light sensitive elements 1102, or visible light camera pixels 1103 can surround DVS infrared light sensitive elements 1102, or any combination thereof. Although a single element 1102 is shown surrounded by other elements 1103, aspects of the present disclosure are not so limited. A single element may include multiple DVS photosensitive elements or clusters of camera pixels, for example, but not limited to, eight DVS photosensitive elements surrounding a cluster of four camera pixels, or one camera pixel surrounded by eight pairs, triplets, or quadlets of DVS photosensitive elements.
[0050] 11B, a hybrid DVS can have multiple sensor types arranged in a checkerboard pattern within an array, where blocks of a first sensor type 1102 are evenly distributed with blocks of a second sensor type 1103. Each of the first sensor type 1102 and second sensor type 1103 can be different. For example, without limitation, the first sensor type 1102 can be a DVS infrared light sensitive element, and the second sensor type 1103 can be a DVS visible light sensitive element.
[0051] Alternatively, sensor types may be distinguished by filtering. In these implementations, one or more filters selectively transmit light to photosensitive elements located behind the one or more filters. The one or more filters may, for example, without limitation, selectively transmit a specific wavelength or wavelengths of light or a specific polarization of light. Alternatively, the one or more filters may selectively block a specific wavelength or wavelengths of light or a specific polarization of light. The photosensitive elements behind the one or more filters may be configured for use as different sensor types. For example, without limitation, an infrared-pass filter 1102 that allows only infrared light to pass through may cover one or more sensor elements in the array 1101, while other sensor elements may be unfiltered or may be an infrared-cut filter 1103. In another alternative implementation, the one or more filters may, for example, without limitation, be optical notch filters that allow only specific wavelengths of light to pass through 1102, while other filters may block the specific wavelengths but allow other wavelengths to pass through 1103. This filtering may reduce the likelihood of false light source detection by allowing the use of illuminator light wavelengths for DTOF and specific wavelengths for DVS photosensitive elements, where the sensor elements may be any type of DVS photosensitive element or any type of camera pixel.
[0052] Patterned sensor or filter elements can be incorporated into hybrid imaging units, for example, as shown in Figures 11C and 11D. Figure 11C shows an example of a hybrid DVS imaging unit 1112C having one or more lens elements 1114, an optional microlens array 1116, a bandpass filter 1118, and a patterned hybrid sensor array 1120. The sensor array includes DVS sensor elements 1122 and conventional imaging sensor elements 1124, which may be arranged in a pattern, for example, as shown in Figures 11A or 11B. Imaging units such as these can be used with an illumination unit (not shown) in a DTOF sensor. The hybrid DVS imaging unit 1112D shown in Figure 11D includes a DVS sensor array 1126 and a patterned filter element 1118 including bandpass and bandcut regions 1118A and 1118B arranged in a pattern, for example, as shown in Figures 11A or 11B.
[0053] FIG. 12 is a diagram illustrating a hybrid DVS including multiple sensor types with inputs separated by a light separator, according to an embodiment of the present disclosure. In this implementation, the light separator 1204 can filter light based on wavelength, and the array 1201 can include multiple sensor element types physically separated based on the wavelength or polarization of the light desired to be detected. For example, without limitation, unpolarized white light 1205 (a mixture of at least all visible light wavelengths, and in most cases including some infrared wavelengths) can enter the light separator 1204, which can be a dispersive prism, a diffraction grating, a dichroic mirror, or the like. As shown, the white light 1205 entering the light separator 1204 can be separated by wavelength or polarization, with some light 1206 entering a first portion of the array 1202 and other light 1207 entering a second portion of the array 1203. Here, the light separator 1204 can be thought of as a filter that changes the diffraction angle based on wavelength or polarization. While the illustrated array depicts the array as a single unit with a thick separating line 1203, aspects of the present disclosure are not so limited. A first portion of array 1201 may have a separation of up to 1 millimeter between a second portion of array 1202, and while the arrays are shown as being separated vertically, other implementations may have horizontally, diagonally, or circumferentially separated portions. Additionally, the optical separators herein can be combined with different filtering or different sensor configurations, for example, as shown in Figures 11A and 11B, to provide additional optical wavelength separation for different sensor types.
[0054] FIG. 13 is a diagram illustrating a hybrid DVS in which multiple sensor types have inputs separated by a microelectromechanical (MEMS) mirror, according to an embodiment of the present disclosure. Here, a MEMS mirror 1304 can oscillate between different positions at set times to reflect light 1305 to a first portion 1301 or a second portion 1302 of the array depending on the time the light reaches the MEMS mirror 1304. The first portion 1301 and second portion 1302 of the array can be physically separated from each other at 1303 based on the angle of incidence of the light reflected by the MEMS mirror 1304. In this manner, light can be temporally filtered between different sensor types. Such temporal filtering can be timed according to a known blinking pattern of a light source on the controller or headset. For example, without limitation, a light source on the controller or headset can be turned on at regular intervals for a predetermined duration. For example, the light source can be turned on for 60 microseconds every 100 microseconds. In such a case, MEMS mirror 1304 can be synchronized to reflect light 1305 to the DVS portion of array 1301 for more than 60 microseconds every 100 microseconds or less to capture changes in the light source. Other reflected light 1307 can be detected by a second portion of array 1302 during times when the light source is off to capture ambient light for image tracking or DTOF.
[0055] Alternatively, the MEMS mirror 1304 can filter the light 1305 based on wavelength. The MEMS mirror in these implementations can be, for example, without limitation, a MEMS Fabry-Perot filter or a diffraction grating. The MEMS mirror can diffract light in a first wavelength range 1306 into at least a first portion 1301 of the array or light in a second wavelength range 1307 into a second portion 1302 of the array.
[0056] FIG. 14 is a diagram illustrating a hybrid DVS in which multiple sensor types have temporarily filtered inputs, according to embodiments of the present disclosure. Here, filters allow light to selectively pass to the array based on time. In some implementations, the filters may be optical waveguides. In a first time step, the filters may allow a first wavelength or polarization of light to pass to the array 1401 and block a second wavelength or polarization of light, or other wavelengths or polarizations. In a second time step, the filters may allow a second wavelength or polarization of light 1402 but block the first wavelength or polarization of light, or other wavelengths or polarizations. In this manner, light can be temporarily filtered, which may be useful for tracking with different sensor types. For example, without limitation, a light source coupled to a controller or headset may be infrared light or may have a specific wavelength. The light source can be configured to turn on and off at specific intervals. The specific intervals during which the light source is on and off may be a sequence or coded pattern. Temporal filtering may be activated during specific intervals to allow specific wavelengths to pass to the sensor and block other wavelengths. Furthermore, the filtering switching interval may be longer than the light source specific interval, taking into account the light travel time to the sensor.
[0057] Body Tracking FIG. 15 is a diagram illustrating body tracking with a DVS, according to an embodiment of the present disclosure. As shown, a user 1501 can wear a headset 1504 having a DVS 1503. Here, the DVS is shown with two arrays or DVS units, or a camera and DVS. The user 1501 holds two controllers 1502 with two or more light sources 1505. The DVS 1503 has the controllers 1502 with corresponding light sources 1505 within its field of view (FOV). Additionally, the DVS 1503 can have the user's appendages, such as the user's hands or arms 1507 or legs or feet 1508, within its FOV. The DVS can also have the ground or other landmarks 1509 within its FOV. For example, without limitation, photosensitive elements in the DVS can detect light reflections corresponding to the user's appendages or the ground or other landmarks when the user moves or the light changes. The camera detects light reflections from the field of view at the camera's frame rate. From the detected light reflections, 1506 can determine the user's appendages or the ground or landmarks. In some alternative implementations, an external DVS 1510 can be used to track the user's appendages. The external DVS 1510 can be at a distance from the user selected so that the user's appendages fit within the field of view of the external DVS 1510. For example, without limitation, the external DVS can be located on top of or below a television or computer monitor, or other freestanding or wall-mounted display.
[0058] The machine learning algorithm is trained to determine the user's body, appendages, ground, or landmarks, and their relative positions and orientations, from data such as events or frames. The machine learning algorithm may be a neural network, and training may be similar to the method described in the general neural network training section of FIGS. 10A-10D above. The neural network may be trained using a training set including labeled events and / or frames. The labeled events and / or frames may include, for example, but not limited to, labels for the user's body, appendages, ground, landmarks, and the relative positions and orientations of the user's body, appendages, ground, and landmarks. In addition to determining the position and orientation of the controller 1502, the determination of the user's body, appendages, ground, or landmarks, and their relative positions and orientations, may be performed, for example, using SLAM. Alternatively, the position and orientation of the controller may be determined using the same trained neural network that determines the labels for the user's body, appendages, ground, or landmarks, and their relative positions and orientations. As indicated by element 1506, a model can be fitted to the determined user's body and appendages to improve the determination of relative position and location.
[0059] Safety Shutter As described above, body tracking in conjunction with determining the position and orientation of the controller can be used to trigger a safety shutter in a VR or AR headset. FIG. 16A shows a headset with a safety shutter door according to an embodiment of the present disclosure. The headset 1601 may include a head strap 1604, an eyepiece 1603, and a display screen 1602. The eyepiece 1603 may include one or more lenses configured to focus on the display screen 1602. The one or more lenses may be, for example, Fresnel lenses or prescription lenses. The display screen 1602 may be transparent, and a hole 1607 in the body of the headset may allow vision through the display screen 1602 when the safety shutter is open.
[0060] In this implementation, the safety shutter 1606 is a door that swings away from the hole 1607 when the safety system is activated. Here, a system-operated clasp 1605 interacts with a clasp 1608 on the safety shutter door 1606 to secure the door closed over the hole 1607. A spring-loaded hinge 1609 can ensure that the safety shutter door 1606 opens quickly when the system-operated clasp 1605 opens. The spring-loaded hinge may, for example, but not be limited to, have a clock-type spring wound around the hinge and secured to the door, where the spring winds when the door closes and unwinds when the door opens. Alternatively, a flat spring may push against the safety shutter door 1606 when it is closed.
[0061] The system-operated clasp 1605 can be configured to open the clasp when the safety system is activated. The safety-system-operated clasp 1605 can include an electric motor or linear actuator that moves the clasp. The safety system may be activated when the ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendages. The safety system can use the determination of the user's body, appendage, ground, or landmarks and their relative positions and orientations, as described above. When the safety system is activated, a signal to open the clasp can be sent to the safety-system-operated clasp 1605. When the clasp opens, the spring-loaded hinge 1609 pushes open the safety shutter door 1606, allowing the user to view through the display 1602 and avoid the risk of activating the safety system.
[0062] Other safety shutter implementations may be used. For example, FIG. 16B shows an alternative headset with a sliding safety shutter according to an embodiment of the present disclosure. Here, when the safety system is activated, the safety shutter slide 1616 slides out of the way of the hole 1607. The safety shutter slide 1616 can move on a spring-loaded rail 1619. Alternatively, the sliding safety shutter 1616 can include a tab that inserts into a slot 1619 in the body of the headset, and a spring in the slot can also push against the sliding safety shutter. In some additional alternative implementations, the spring may be omitted, and the sliding safety shutter 1616 may operate under gravity. The spring-loaded rail 1619 can push against the sliding safety shutter 1616 when it is closed, ensuring that the safety shutter opens quickly when the safety system is activated. A system operated clasp 1605 interacts with a clasp 1618 on the sliding safety shutter 1616 to secure the slide closed.
[0063] If the ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendages, a safety system may be activated. The safety system may use the determination of the user's body, appendages, the ground, or landmarks and their relative positions and orientations, as described above. When the safety system is activated, a signal may be sent to the system-operated clasp 1605 to open the clasp. When the clasp opens, a spring 1619 pushes open a sliding safety shutter 1616, allowing the user to see through the display 1602 and avoid the risk of tripping the safety system.
[0064] FIG. 16C illustrates a headset with a louvered safety shutter according to an embodiment of the present disclosure. In this implementation, the slats of the louvered safety shutter 1626 have a first dimension longer than a second dimension. In the closed position, the long dimension of the slats is approximately parallel to the optics, and each slat overlaps either another slat or the headset body, blocking light from passing through the display screen 1602. In the open position, the slats change position so that their short dimension is parallel to the optics 1603, allowing light to pass through the slats to reach the display screen. An actuator rod 1628 can hinge each of the slats 1626. A safety-system-controlled actuator 1629 can push or pull the actuator rod 1628 to open the louvered safety shutter when the safety system is activated. In some other implementations, the safety system controlled actuator may be spring loaded, and the clasp coupled to the actuator rod 1628 and the system controlled actuator 1629 may include a clasp that interfaces with a clasp on the actuator rod. When the safety system is activated, the clasp on the system controlled actuator may open, allowing the actuator rod to move under spring pressure and open the louvers.
[0065] If the ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendages, the safety system may be activated. The safety system may use the determination of the user's body, appendages, the ground, or landmarks and their relative positions and orientations, as described above. When activated, a signal may be sent to a system-operated actuator to move an actuator rod. Movement of the actuator rod 1629 pushes open the slats of the louvered safety shutter 1626, allowing the user to see through the display 1602 and avoid the risk of tripping the safety system.
[0066] FIG. 16D shows a fabric safety shutter according to an embodiment of the present disclosure. In this implementation, the safety shutter 1636 is constructed from an opaque fabric, such as, but not limited to, a tightly woven cotton fabric, polyester, vinyl, or a tightly woven wool fabric. The fabric safety shutter 1636 may be coupled to a fabric roller 1639. The fabric roller 1639 may be configured to wind up the fabric shutter when the safety system is activated. The fabric roller may be spring-loaded, for example, but not limited to, using a clock spring, so that the clock spring is under tension as the safety shutter closes; alternatively, an electric motor may be used to wind up the fabric safety shutter. A system-operated clasp 1605 may interface with a clasp 1638 coupled to the fabric safety shutter 1636 to ensure that the fabric safety shutter does not open unintentionally.
[0067] During operation, the fabric safety shutter may be in a closed position. If the ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendages, the safety system may be activated. The safety system may use the determination of the user's body, appendages, the ground, or landmarks and their relative positions and orientations, as described above. When the safety system is activated, a signal may be sent to the system-operated clasp 1605 to open the clasp. The clasp opens the fabric roller 1619, which rolls up the fabric safety shutter 1616, allowing the user to see through the display 1602 and avoid the risk of activating the safety system.
[0068] 16E shows a headset with a liquid crystal safety shutter, according to an embodiment of the present disclosure. In this implementation, an LCD screen 1646 is integrated into the headset 1601 and covers an otherwise hole in the headset. A safety system-controlled LCD screen driver 1649 is communicatively coupled to the LCD screen 1646.
[0069] The liquid crystal screen 1646 may be, for example, but not limited to, a liquid crystal shutter having a first polarizer and a second polarizer, the first polarizer having a 90-degree polarization difference from the second polarizer, and a fluid-filled cavity. The fluid-filled cavity may contain liquid crystals configured to have a first orientation in the absence of an electric field that changes the polarization of light to allow it to pass from the first polarizer to the second polarizer. The liquid crystals may be further configured to align to a second orientation under an electric field. Because the second orientation of the liquid crystals does not change the polarization of the light, light passing through the first polarizer is blocked by the second polarizer. Electrodes may be disposed along the surface of the liquid crystals to enable control of the liquid crystals. A safety system-controlled liquid crystal screen driver may be communicatively coupled to these electrodes, allowing the electrodes to control the liquid crystals within the fluid-filled cavity. As used herein, communicatively coupled means that electrical signals representing messages or instructions can be sent and / or received from one coupled element to another coupled element; these signals may pass through intermediate elements and their format may change, but the messages contained therein remain unchanged.
[0070] In operation, the safety system controlled driver 1649 can send a signal to the liquid crystal safety shutter 1646 to make it opaque while the display screen 1602 is active. When the safety system is activated, the driver 1649 can make the liquid crystal safety shutter 1646 transparent. For example, without limitation, the driver can reduce the voltage supplied to the liquid crystal safety shutter, returning the liquid crystals to their first orientation, changing the polarization of the light and allowing the light to pass through the second polarizer. The safety system may be activated when the ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendages. The safety system can use the determination of the user's body, appendages, ground, or landmarks, and their relative positions and orientations, as described above.
[0071] Finger position tracking Aspects of the present disclosure may be applied to finger tracking. FIG. 17 illustrates finger tracking with a DVS and a controller according to aspects of the present disclosure. Here, the controller 1701 includes two or more light sources and one or more buttons 1705. The two or more light sources include one or more light sources 1706 proximate the one or more buttons 1705 and two or more other tracking light sources 1703. As described above, the two or more other tracking light sources 1703 can generate events in the DVS 1702 that are used to determine the position and orientation of the controller 1701. As shown here, two DVSs 1702, or a DVS with two arrays and three other tracking light sources 1703, are used to determine the position and orientation of the controller.
[0072] One or more light sources 1706 proximate one or more buttons 1705 can be used for finger tracking. For example, without limitation, finger tracking can be achieved with the DVS 1702 using occlusion of one or more light sources 1706 proximate the buttons 1705. The one or more light sources 1706 proximate the buttons can be turned off and on at predetermined intervals. The DVS 1702 can generate an event for each flash. The event can be analyzed to determine occlusion of one or more light sources 1706 proximate the buttons 1705. By knowing the configuration of one or more light sources proximate a button, for example, without limitation, when occluded by a finger or palm, the light pattern detected in the event generated by the DVS will be different than when the light source is not occluded. As described above with respect to determining the position and orientation of the controller, the timing of the flashes can now be used to determine which lights are occluded, and thereby the position of the corresponding finger or palm. If a light source 1706 proximate a button 1705 has reduced intensity or no intensity during the interval, the light source proximate the button is known to be “on.” That light source is determined to be obscured. Similarly, if the detected intensity of a light known to be “on” changes, an event may be generated from which it may be determined that the user's finger or hand has moved and the button has become unactivated. Obscuration of one or more of the light sources proximate one or more buttons can be correlated to the position of the finger or palm based on their location. For example, without limitation, a light source located near the user's palm when the controller 1701 is held can be used to determine the position of the user's hand 1704.
[0073] Light sources 1706 can be positioned around each button 1705, and the button configuration and design of the controller can be used to determine finger position. For example, without limitation, the controller 1701 can be designed so that each of the user's fingers rests near a button 1705 when held. The light source occlusion pattern determined from the DVS event can then be used to determine when the user's finger hovers over an inactive button and, further, when the user's finger passes over the button. This can be useful for providing the user with additional interaction options, such as a half-press, a full press, or other button options. Multiple light sources can surround each button, allowing for refined determination of the user's finger or palm position. For example, without limitation, in some implementations, 10 or more light sources can surround each button, while in other implementations, a single light source can shine light through a translucent diffuser around the button, and the interruptions in the diffuse light profile can be used to determine finger position.
[0074] In some implementations, the buttons 1705 themselves can also be light sources. One or more buttons 1705 can turn off and on at different intervals than one or more light sources 1706 or other light tracking light sources 1703 proximate the button. Alternatively, the light source of the button 1705 can have a different wavelength or polarization than one or more light sources 1706 or other light tracking light sources 1703 proximate the button.
[0075] Buttons and tracking can also be used to activate power-saving modes for light sources. For example, if a button is determined to be pressed, one or more light sources proximate the button may be dimmed or turned off. Additionally, if the controller 1701 is determined to be out of the field of view of the DVS 1702, one or more lights proximate the button may be dimmed or turned off. In some implementations, data from the IMU can be used to determine whether the controller is being held by a user, and the light source may be dimmed or turned off if no change in IMU data, such as, but not limited to, acceleration, angular velocity, etc., is detected for a threshold period. Once a change in the IMU data is detected, the light source can be turned back on.
[0076] In an alternative implementation, finger tracking can be performed without using one or more light sources. A machine learning model can be trained using a machine learning algorithm to detect the finger position from events generated from changes in ambient light due to finger movement. The machine learning model can be a general machine learning model such as the CNN, RNN, or DNN described above. In some implementations, a specialized machine learning model, such as, for example, but not limited to, a spiking (or sparking) neural network (SNN), can be trained with a specialized machine learning algorithm. SNNs mimic biological NNs by having activation thresholds and weights that are adjusted according to relative spike times within an interval, also known as STDP (Spike-timing-dependent-plasticity). When an activation threshold is reached, the SNN is said to spike and transmit its weights to the next layer. SNNs can be trained via STDP and supervised or unsupervised learning techniques. Further information on SNNs can be found in Tavanaei, Amirhossein et al., "Deep Learning in Spiking Neural Networks," Neural Networks (2018) arXiv:1804.08150, the contents of which are incorporated herein by reference for all purposes.
[0077] Alternatively, a high dynamic range (HDR) image may be constructed using events aggregated from ambient data. A machine learning model is trained to recognize hand positions or controller positions and orientations from the HDR images. The trained machine learning model can be applied to the HDR images generated from the events to determine hand / finger positions or controller positions and orientations. The machine learning model may be a general machine learning model trained with supervised learning techniques, as described in the section on training a general neural network.
[0078] Eye tracking Aspects of the present disclosure may be applied to eye tracking. Generally, eye tracking image analysis utilizes unique characteristics of how light is reflected from the eye to determine gaze direction from an image. For example, an image may be analyzed to identify the location of the eye based on corneal reflections in the image data, and the image may be further analyzed to determine gaze direction based on the relative position of the pupil in the image.
[0079] Two common eye-tracking techniques that determine gaze direction based on pupil position are known as bright pupil and dark pupil methods. Bright pupil methods involve illuminating the eye with a light source substantially aligned with the optical axis of the DVS, which reflects the emitted light off the retina and back through the pupil to the DVS. The pupil appears in the image as a distinct bright spot at the pupil's location, similar to the red-eye effect that occurs in images during traditional flash photography. In this method of eye-tracking, a bright reflection from the pupil itself helps the system locate the pupil when there is insufficient contrast between the pupil and the iris.
[0080] Dark pupil imaging involves illuminating the DVS with a light source substantially offset from its optical axis, which reflects light directed through the pupil away from the DVS's optical axis, resulting in a dark spot at the pupil where the event can be identified. In another dark pupil imaging system, an infrared light source and camera aimed at the eye can view the corneal reflection. Such DVS-based systems track the position of the pupil and corneal reflection, which provides parallax due to the different depths of the reflections, improving accuracy.
[0081] FIG. 18A illustrates an example of a dark pupil eye tracking system 1800 that may be used in the context of the present disclosure. The eye tracking system tracks the orientation of a user's eye E relative to a display screen 1801 on which a visible image is presented. While the exemplary system of FIG. 18A utilizes a display screen, certain alternative embodiments may utilize an image projection system that can project an image directly onto the user's eye. In these embodiments, the user's eye E is tracked relative to the image projected onto the user's eye. In the example of FIG. 18A, the eye E collects light from the screen 1801 through a variable iris I; a lens L projects the image onto the retina R. The opening in the iris is known as the pupil. Muscles control the rotation of the eye E in response to nerve impulses from the brain. The upper and lower eyelid muscles ULM and LLM control the upper and lower eyelids UL and LL, respectively, in response to other nerve impulses.
[0082] Light-sensitive cells on the retina R generate electrical impulses that are sent via the optic nerve ON to the user's brain (not shown). The brain's visual cortex interprets the impulses. Not all parts of the retina R are equally light-sensitive. Specifically, the light-sensitive cells are concentrated in an area known as the fovea.
[0083] The illustrated image tracking system includes one or more infrared light sources 1802, e.g., light emitting diodes (LEDs), that direct non-visible light (e.g., infrared light) toward the eye E. Some of the non-visible light is reflected off the cornea C of the eye, and some is reflected off the iris. The reflected non-visible light is directed by a wavelength-selective mirror 1806 to a DVS 1804 that is sensitive to infrared light. The mirror transmits visible light from the screen 1801 but reflects non-visible light reflected from the eye.
[0084] The DVS 1804 generates eye E events that can be analyzed to determine gaze direction GD from the relative positions of the pupils. The events can be generated by the processor 1805. The DVS 1804 is advantageous in this implementation because its very fast update rate of events provides near real-time information about changes in the user's gaze.
[0085] 18B, an event 1811 showing the user's head H can be analyzed to determine a gaze direction GD from the relative position of the pupil. For example, the analysis may determine the two-dimensional offset of the pupil P from the center of the eye E in the image. The position of the pupil relative to the center can be converted to a gaze direction relative to the screen 1801 by a simple geometric calculation of a three-dimensional vector based on the known size and shape of the eyeball. The determined gaze direction GD can indicate the rotation and acceleration of the eye E as it moves relative to the screen 1801.
[0086] As can also be seen in Figure 18B, events may include non-visible light reflections 1807 and 1808 from the cornea C and lens L, respectively. Because the cornea and lens are at different depths, the parallax and refractive index between the reflections may be used to improve accuracy in determining gaze direction GD. An example of this type of eye-tracking system is a dual Purkinje tracker, where the corneal reflection is the first Purkinje image and the lens reflection is the fourth Purkinje image. If the user is wearing glasses, there may also be a reflection 1808 from the user's glasses 1809.
[0087] The performance of an eye-tracking system depends on many factors, including the placement of the light source (IR, visible, etc.) and DVS, whether the user is wearing glasses or contacts, the headset optics, the latency of the tracking system, the speed of eye movements, the shape of the eye (which may change throughout the day or as a result of movement), the condition of the eye (e.g., amblyopia), gaze stability, gaze on moving objects, the scene being presented to the user, and the user's head movement. The DVS provides a very fast update rate of events, reducing the output of redundant information to the processor. This allows for rapid processing and fast determination of eye-tracking state and error parameters.
[0088] Error parameters that can be determined from the eye-tracking data can include, but are not limited to, rotational velocity and prediction error, fixation error, confidence intervals for current and / or future gaze positions, and smooth pursuit error. State information about the user's gaze includes the individual state of the user's eyes and / or gaze. Thus, exemplary state parameters that can be determined from the eye-tracking data can include, but are not limited to, blink metrics, saccade metrics, depth of field response, color blindness, gaze stability, and eye movements as precursors to head movement.
[0089] In certain implementations, the eye tracking error parameters may include a confidence interval for the current gaze position. The confidence interval may be determined by examining the rotational velocity and acceleration of the user's eyes for changes from the last position. In alternative embodiments, the eye tracking error and / or state parameters may include a prediction of a future gaze position. The future gaze position may be determined by examining the rotational velocity and acceleration of the eyes and extrapolating possible future positions of the user's eyes. Generally speaking, the DVS update rate of the eye tracking system may be so high that for users with high values of rotational velocity and acceleration, the error between the determined future position and the actual future position may be small, and this small error may be significantly smaller than existing camera-based systems.
[0090] In yet another alternative implementation, the eye tracking error parameter may include a measurement of eye velocity, e.g., rotational velocity. In a specific alternative embodiment, the determined eye tracking state parameter includes measuring metrics of the user's blink. During a typical blink, a period of 150 milliseconds (ms) typically elapses during which the user's vision is not focused on the presented image. Thus, depending on the frame rate of the display device, the user's vision may not be focused on the presented image for up to 20-30 frames. However, upon terminating a blink, the user's gaze direction may not correspond to the last measured gaze direction determined by the acquired eye tracking data. Therefore, metrics of the user's gaze may be determined from the acquired eye tracking data. These metrics may include, but are not limited to, the measured start and end times of the user's blink, as well as the predicted end times.
[0091] In yet another alternative implementation, the determined eye-tracking state parameters include measuring metrics of a user's saccade. During a typical saccade, a period of 20-200 ms typically elapses during which the user's vision is not focused on the presented image. Thus, depending on the frame rate of the display device, the user's vision may be unfocused on the presented image for up to 40 frames. However, due to the nature of saccades, once the saccade ends, the user's gaze direction shifts to another area of interest. Therefore, eye-tracking data may be used to establish metrics of a user's saccade based on actual or predicted time elapsed during the saccade. These metrics may include, but are not limited to, the measured start and end times of the user's saccade, as well as the predicted end times.
[0092] In certain alternative implementations, the determined eye-tracking state parameters include determining transitions in the user's gaze direction between regions of interest as a result of changes in depth of field between the presented images, since providing a transition between regions of interest in the presented images would cause the user to experience a saccade.
[0093] In yet another alternative implementation, the determined eye-tracking state parameters can be adapted for color blindness. For example, regions of interest may be present in an image presented to a user such that these regions are not recognized by a user with a particular form of color blindness. The obtained eye-tracking data can determine whether the user's gaze identified or responded to a region of interest, for example, as a result of a change in the user's gaze direction. Thus, the eye-tracking error parameter can determine whether the user is color blind to a particular color or spectrum.
[0094] In certain alternative implementations, the determined eye tracking state parameters include a measure of the user's gaze stability. Determining gaze stability can be done by measuring the microsaccade radius of the user's eyes; the smaller the fixation overshoot and undershoot, the more stable the user's gaze.
[0095] In yet a further alternative implementation, the determined eye tracking error and / or state parameters include the user's ability to fixate on a moving object. These parameters may include a measure of the user's eye's ability to perform smooth pursuit and a maximum object tracking velocity of the eyes. Typically, users with better smooth pursuit ability experience less jitter in their eye movements.
[0096] In certain alternative implementations, the determined gaze tracking error and / or state parameters include determining eye movements as precursors to head movements. An offset between head and eye orientation can affect certain errors and / or state parameters, for example, in smooth pursuit or fixation, as described above.
[0097] Further information regarding eye tracking and determining error parameters can be found in US Pat. No. 10,192,528, the contents of which are incorporated herein by reference for all purposes.
[0098] system 19 is a block system diagram of a DVS tracking system according to an embodiment of the present disclosure. By way of example and not limitation, according to an embodiment of the present disclosure, the system 1900 may be an embedded system, a mobile phone, a personal computer, a tablet computer, a portable gaming device, a workstation, a gaming console, etc.
[0099] System 1900 generally includes a central processing unit (CPU) 1903 and memory 1904. System 1900 may also include well-known support functions 1906 that may communicate with other components of the system, for example, via a data bus 1905. Such support functions may include, but are not limited to, input / output (I / O) elements 1907, a power supply (P / S) 1911, a clock (CLK) 1912, and a cache 1913.
[0100] System 1900 may include a display device 1931 for presenting rendered graphics to a user. In an alternative implementation, the display device is a separate component that works in conjunction with system 1900. Display device 1931 may be in the form of a flat panel display, a head-mounted display (HMD), a cathode ray tube (CRT) screen, a projector, or other device capable of displaying visible text, numbers, graphic symbols, or images.
[0101] Here, display device 1931 is coupled to DVS 1901A, and controller 1902 includes two or more light sources 1932A, which may be in any of the configurations described herein. In an alternative implementation, the DVS can be coupled to a game controller, and the display device can instead include two or more light sources. In yet another alternative implementation, the DVS is a separate unit decoupled from either the display device or the controller, in which case both the controller and display device may include two or more light sources for tracking.
[0102] In some implementations, for example, where the display device is part of a head-mounted display (HMD), such an HMD may include an inertial measurement unit (IMU) such as an accelerometer or gyroscope. Also, as discussed herein above, such an HMD may include light sources 1932B, which may be tracked using a DVS that is separate from the display device 1901 and coupled to the CPU 1903. By way of example, a separate DVS 1901B may be attached to the controller 1902.
[0103] In some implementations, DVS1901A or DVS1901B may be part of a hybrid sensor, for example, as described above with respect to Figure 11A, Figure 11B, Figure 11C, or Figure 11D. Such a hybrid sensor may include a depth sensor, for example, a DTOF sensor, in which case the hybrid sensor may include an illumination unit (not shown).
[0104] Additionally, if the display device 1931 is part of an HMD, the device may optionally be fitted with a safety shutter 1933, which may be operably coupled to a processor such as CPU 1903 and may operate as described above with respect to Figures 16A, 16B, 16C, 16D or 16E. Alternatively, the safety shutter may be controlled by a separate processor attached to the HMD.
[0105] The system 1900 also includes a mass storage device 1915, such as a disk drive, CD-ROM drive, flash memory, a solid-state drive (SSD), a tape drive, or the like, to provide non-volatile storage of programs and / or data. The system 1900 may also optionally include a user interface unit 1916 to facilitate interaction between the system 1900 and a user. The user interface 1916 may include a keyboard, a mouse, a joystick, a light pen, or other devices that can be used in conjunction with a graphical user interface (GUI). The system 1900 may also include a network interface 1914 to enable the device to communicate with other devices over a network 1920. The network 1920 may be, for example, a local area network (LAN), a wide area network such as the Internet, a personal area network such as a Bluetooth network, or other type of network. These components may be implemented in hardware, software, or firmware, or some combination of two or more of these.
[0106] The CPUs 1903 may each include one or more processor cores, e.g., a single core, two cores, four cores, eight cores, or more. In some implementations, the CPUs 1903 may include a GPU core or multiple cores of the same APU (Accelerated Processing Unit). The memory 1904 may be in the form of an integrated circuit providing addressable memory, e.g., random access memory (RAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), etc. The main memory 1904 may include application data 1923 used by the processors 1903 during processing. The main memory 1904 may also include event data 1909 received from the DVS 1901. As illustrated in FIG. 9 , a trained neural network (NN) 1910 may be loaded into the memory 1904 to determine position and orientation data. Additionally, the memory 1904 may include machine learning algorithms 1921 for training or tuning the NN 1910. A database 1922 can be included in memory 1904. The database can include information regarding the light source configuration, the predetermined blink interval for each of the one or more light sources, etc. The memory can also include outputs from an IMU coupled to the controller 1902 or the display device 1931. In some implementations, the memory 1904 can include outputs from one or more light sources, such as timestamps for when each of the light sources is on or off.
[0107] According to aspects of the disclosure, processor 1903 can execute methods for determining the position and orientation of a controller or a user, as described in Figures 8 and 9, which can be loaded into memory 1904 as application 1923. The processor can generate one or more orientations and configurations of the controller, headset, the user's body or appendages, the ground, or landmarks as a result of executing the methods described in Figures 8 and 9 and further described with respect to Figure 15. These positions and orientations can be persisted in database 1922 and can be used for successive iterations of the methods of Figures 8 and 9. In some implementations, processor 1903 can utilize such positions and / or orientations in machine learning algorithms trained to perform simultaneous localization and mapping (SLAM).
[0108] Mass storage 1915 can include applications or programs 1917 that are loaded into main memory 1904 when processing begins in application 1923. Additionally, mass storage 1915 can include data 1918 that is used by the processor during processing of application 1923, NN 1910, machine learning algorithm 1921, and during population of database 1922.
[0109] As used herein and generally understood by those skilled in the art, an application specific integrated circuit (ASIC) is an integrated circuit that is customized for a particular application, rather than for general use.
[0110] As used herein and generally understood by those skilled in the art, a field programmable gate array (FPGA) is an integrated circuit that is designed to be configured by a customer or designer after manufacture - hence "field programmable." FPGA configurations are generally specified using a hardware description language (HDL) similar to those used for ASICs.
[0111] As used herein and as commonly understood by those skilled in the art, a system on a chip or system-on-chip (SoC or SOC) is an integrated circuit (IC) that integrates all the components of a computer or other electronic system onto a single chip. This can include digital, analog, mixed-signal, and often radio frequency functionality—all on a single chip substrate. A typical application is in the field of embedded systems.
[0112] A typical SoC includes the following hardware components: One or more processor cores (e.g., a microcontroller, microprocessor, or digital signal processor (DSP) core). Memory blocks, such as read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM) and flash memory. A timing source, for example, an oscillator or a phase-locked loop. Peripherals, such as counter timers, real-time timers, or power-on reset generators. External interfaces, such as industry standards such as Universal Serial Bus (USB), FireWire®, Ethernet, universal asynchronous receiver / transmitter (USART), and Serial Peripheral Interface (SPI) bus. Analog interfaces (including analog-to-digital converters (ADCs) and digital-to-analog converters (DACs)). Voltage regulators and power management circuits.
[0113] These components are connected by either proprietary or industry-standard buses. Direct Memory Access (DMA) controllers increase the data throughput of the SoC by routing data directly between external interfaces and memory, bypassing the processor core.
[0114] A typical SoC includes both the hardware components described above and executable instructions (eg, software or firmware) that control the processor core(s), peripherals, and interfaces.
[0115] Aspects of the present disclosure provide image-based tracking featuring higher sample rates than are possible with conventional image-based tracking systems, resulting in improved tracking fidelity. Additional benefits include reduced cost, reduced weight, reduced extraneous data generation, and reduced processing requirements when using DVS-based tracking systems. Advantages such as these enable improvements in virtual reality (VR) and augmented reality (AR) systems, among other applications.
[0116] While the above is a complete description of the preferred embodiment of the present invention, it is possible to use various alternatives, modifications, and equivalents. Therefore, the scope of the invention should be determined not with reference to the above, but instead with reference to the appended claims, along with their full scope of equivalents. Any feature described in this specification, whether preferred or not, may be combined with any other feature described in this specification, whether preferred or not. In the following claims, the indefinite article "A" or "An" refers to a quantity of one or more of the item following the article, unless expressly stated otherwise. The appended claims should not be construed as including means-plus-function limitations unless such limitations are expressly recited in a given claim using the phrase "means for."
Claims
1. processor, a controller operably coupled to the processor; two or more light sources mounted in a known configuration relative to each other and relative to the controller, each light source of the two or more light sources configured to be turned on and off in a pattern indicative of the identity of each light source such that the two or more light sources are distinguishable from one another; a dynamic vision sensor (DVS) operably coupled to the processor, the DVS having an array of light sensitive elements in a known configuration, the dynamic vision sensor configured to, in response to changes in light output from the two or more light sources, output signals corresponding to two or more events at two or more corresponding light sensitive elements in the array, the output signals including information corresponding to times of the two or more events and positions of the two or more corresponding light sensitive elements in the array; Including, the processor is configured to determine an association between each of the two or more events and two or more corresponding particular light sources among the two or more light sources, and to adapt a position and orientation of the controller using the determined association, the known configuration of the two or more light sources relative to each other and relative to the controller body, and the positions of the two or more corresponding light sensitive elements within the array.
2. 2. The tracking system of claim 1, wherein adapting the position and orientation of the controller comprises using a machine learning model trained to adapt at least the two or more events to the position and orientation of the controller.
3. 3. The tracking system of claim 2, wherein the machine learning model is trained using artificial training data, the artificial training data including at least two or more events having corresponding masked controller positions and orientations.
4. The tracking system of claim 3 , wherein noise is added to the artificial training data.
5. The tracking system of claim 2 , wherein the machine learning model is trained from live controller tracking and masked controller position and orientation information.
6. The tracking system of claim 2 , wherein the machine learning model is further trained to predict an orientation and a position of a user's hand or arm from the two or more events.
7. 2. The tracking system of claim 1, wherein adapting the position and orientation of the controller comprises using an adaptation algorithm to adapt the known configurations of the two or more light sources relative to each other and relative to the controller to the two or more events.
8. The tracking system of claim 1 , wherein the two or more light sources are configured to emit infrared light, and the photosensitive element of the DVS is configured to detect infrared light.
9. a processor; a dynamic visual sensor (DVS) operably coupled to the processor, the DVS having an array of light sensitive elements in a known configuration, the dynamic visual sensor being mounted in a known configuration relative to each other and relative to a controller body, and configured to output signals corresponding to two or more events at two or more corresponding light sensitive elements in the array in response to changes in light output from two or more light sources, the output signals including information corresponding to times of the two or more events and positions of the two or more corresponding light sensitive elements in the array, each light source of the two or more light sources being configured to turn on and off in a pattern indicative of the identity of each light source such that the two or more light sources are distinguishable from one another; Including, the processor is configured to determine an association between each of the two or more events and two or more corresponding particular light sources among the two or more light sources, and to adapt at least the position and orientation of the controller using the determined association, the known configuration of the DVS relative to each other and relative to the controller body, and the positions of the two or more corresponding light sensitive elements within the array.
10. 10. The tracking system of claim 9, wherein adapting the position and orientation of the controller comprises using a machine learning model trained to adapt at least the two or more events to the positions and orientations of the two or more light sources relative to the controller.
11. 11. The tracking system of claim 10, wherein the machine learning model is trained using artificial training data, the artificial training data including at least two or more events having corresponding masked positions and orientations of the two or more light sources relative to the controller.
12. The tracking system of claim 11 , wherein noise is added to the artificial training data.
13. The tracking system of claim 10 , wherein the machine learning model is trained from live tracking data and masked positions and orientations of the two or more light sources.
14. The tracking system of claim 10 , wherein the machine learning model is further trained to predict an orientation and a position of a user's hand or arm from the two or more events.
15. 10. The tracking system of claim 9, wherein adapting at least the position and orientation of the controller comprises using an adaptation algorithm to adapt the known configuration of the two or more light sources relative to each other and relative to the controller to the two or more events.
16. The tracking system of claim 9 , wherein the DVS is attached to a controller, and the two or more light sources are attached to the controller body.
17. 1. A tracking method comprising: detecting two or more events corresponding to changes in light output from two or more light sources, the two or more light sources being configured to be turned on and off in a pattern indicative of the identity of each light source such that the two or more light sources are distinguishable from one another; determining an association between each of the two or more events and two or more corresponding light sources, the two or more light sources having a known configuration with respect to each other and with respect to the controller body; adapting at least a position and orientation of the controller using the determined associations, the known configuration of the two or more light sources relative to each other and relative to the controller body, and the positions of the two or more corresponding light sensitive elements within the array; A method comprising:
18. 20. The method of claim 17, wherein adapting the position and orientation of the controller comprises using a machine learning model trained to adapt at least the two or more events to the position and orientation of the controller.
19. 20. The method of claim 18, wherein the machine learning model is trained using artificial training data, the artificial training data including at least two or more events with corresponding masked controller positions and orientations.
20. 20. The method of claim 19, wherein noise is added to the artificial training data.
21. 20. The method of claim 17, wherein adapting the position and orientation of the controller comprises using an adaptation algorithm to adapt the known configuration of the two or more light sources relative to each other and relative to the controller to the two or more events.
Citation Information
Patent Citations
Signal generation system, detector system, and method for determining the position of a user's finger
JP2018500674A
Tracking object position and orientation in virtual reality systems
JP2020518808A
Infrared remocon and remocon receiving apparatus
KR1020130021247A
Utilizing a hybrid model to recognize fast and precise hand inputs in a virtual environment
US10956724B1
Virtual Reality
US20190088018A1