Methods for controller tracking and simultaneous body tracking, and the development of dynamic vision sensor hybrid elements in SLAM or safety shutters.
DVS with IMUs and machine learning algorithms address the limitations of infrared cameras and inertial sensors in VR and AR by enabling efficient and accurate game controller tracking without complex setups.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-04-12
- Publication Date
- 2026-03-25
AI Technical Summary
Existing VR and AR implementations face challenges in accurately and efficiently tracking game controllers due to the limitations of infrared cameras and inertial sensors, which require expensive hardware and complex setups, and do not provide smooth motion feedback at high frame rates.
Utilizing Dynamic Vision Sensors (DVS) with high update rates to detect changes in light intensity from multiple light sources on the controller, combined with inertial measurement units (IMUs) and machine learning algorithms, to determine precise position and orientation without the need for external lighthouses or complex setups.
Enables accurate and efficient tracking of game controllers with reduced data processing requirements, eliminating the need for expensive hardware and complex setups, while providing smooth motion feedback.
Smart Images

Figure 0007835896000001 
Figure 0007835896000002 
Figure 0007835896000003
Abstract
Description
Technical Field
[0001] Aspects of the present disclosure relate to the tracking of game controllers. Specifically, aspects of the present disclosure relate to the tracking of game controllers using dynamic vision sensors.
Background Art
[0002] Latest virtual reality (VR) and augmented reality (AR) implementations rely on accurate and fast movement tracking for user interaction with devices. AR and VR often rely on information regarding the position and orientation of a controller relative to other objects. Many VR and AR implementations rely on a combination of inertial measurements obtained by accelerometers or gyroscopes within the controller and visual detection of the controller by an external camera to determine the position and orientation of the controller.
[0003] Some of the earliest implementations use infrared light detected by an infrared camera with a defined detection radius on a game controller pointed at the screen. The camera captures images at a moderately fast rate of 200 frames per second to determine the position of the infrared light. The distance between infrared light sources is determined from the relative position of the infrared light sources in the camera image, which allows for the calculation of a given object and the position of the controller relative to the screen. Accelerometers may also be used to provide information about relative 3D changes in the controller's position or orientation. These earlier implementations rely on a fixed screen position and a controller pointed at the screen. Modern VR and AR implementations allow the screen to be positioned closer to the user's face in a head-mounted display that moves with the user. Therefore, having an absolute light position (also called a lighthouse) becomes undesirable because it requires extra setup time and the user must set up an independent lighthouse point that limits the user's range of motion. Furthermore, even the moderately fast frame rate of 200 frames per second for infrared cameras was not fast enough to provide smooth motion feedback. In addition, this setup was too simplified and unsuitable for more modern inside-out detection methods such as room mapping and hand detection.
[0004] Recent implementations use cameras and accelerometers in combination with trained machine learning algorithms that have been trained to detect both hands, controllers, and / or other body parts. For smooth motion detection, a high frame rate camera is required to generate image frames for body part / controller detection. This generates a large amount of data that needs to be processed quickly for smooth update rates. Therefore, expensive hardware is required to process the frame data. Furthermore, much of the frame data within each frame is discarded as unnecessary as it is not relevant to motion tracking.
[0005] It is in these circumstances that the nature of this disclosure arises. [Overview of the project]
[0006] The teachings of this invention can be easily understood by examining the following detailed description in conjunction with the accompanying drawings. [Brief explanation of the drawing]
[0007] [Figure 1] This is a diagram illustrating an implementation of game controller tracking using a DVS having a single sensor array, according to one aspect of the present disclosure. [Figure 2] This is a diagram illustrating an implementation of game controller tracking using a DVS having a dual sensor array, according to one aspect of the present disclosure. [Figure 3] This is a diagram illustrating an implementation of game controller tracking using a DVS in combination with a single sensor array and a camera, according to one aspect of the present disclosure. [Figure 4] This is a diagram illustrating a DVS that tracks the motion of a game controller having two or more light sources, according to one aspect of the present disclosure. [Figure 5] This is a diagram illustrating an implementation of head tracking or other device tracking using a controller including a DVS having a single sensor array, according to one aspect of the present disclosure. [Figure 6] This is a diagram illustrating an implementation of head tracking or other device tracking using a game controller including a DVS having a dual sensor array, according to one aspect of the present disclosure. [Figure 7] This is a diagram illustrating an implementation of head tracking or other device tracking using a controller having a DVS in combination with a single sensor array and a camera, according to one aspect of the present disclosure. [Figure 8] This is a flowchart illustrating a motion tracking method using DVS with one or more light sources and a configuration-adapted model of one light source, according to one aspect of the present disclosure. [Figure 9] This is a flowchart illustrating a motion tracking method using DVS with timestamped light source position information, according to one aspect of this disclosure. [Figure 10A]This diagram illustrates a basic form of an RNN having a node layer according to an aspect of the present disclosure, where each node is characterized by an activation function, one input weight, a regressive hidden node transition weight, and an output transition weight. [Figure 10B] This is a simplified diagram illustrating how a series of nodes having the same activation function that moves over time can be considered as an RNN according to aspects of this disclosure. [Figure 10C] An exemplary layout of a convolutional neural network, such as a CRNN, according to the aspects of this disclosure is shown. [Figure 10D] A flowchart illustrating a supervised training method for a machine learning neural network according to the aspects of this disclosure is shown. [Figure 11A] This is a diagram illustrating a hybrid DVS having multiple colocation sensor types according to an aspect of the present disclosure. [Figure 11B] This is an illustration of a hybrid DVS having multiple sensor types arranged in a grid pattern within an array, according to an aspect of the present disclosure. [Figure 11C] This is a schematic cross-sectional view of a hybrid DVS having multiple sensor types arranged in a pattern within an array, according to an aspect of the present disclosure. [Figure 11D] This is a schematic cross-sectional view of a hybrid DVS having multiple filter types arranged in a pattern within an array, according to an aspect of the present disclosure. [Figure 12] This diagram illustrates a hybrid DVS, according to an aspect of the present disclosure, which includes multiple sensor types having inputs separated by an optical separator. [Figure 13] This diagram illustrates a hybrid DVS in which multiple sensor types have inputs separated by micro-electromechanical (MEMS) mirrors, according to an aspect of the present disclosure. [Figure 14] This diagram illustrates a hybrid DVS in which multiple sensor types have temporarily filtered inputs, according to an aspect of the present disclosure. [Figure 15] This is an illustration illustrating body tracking using DVS in the manner of this disclosure. [Figure 16A]A diagram showing a headset with a safety shutter door according to an aspect of the present disclosure. [Figure 16B] A diagram showing a headset with a sliding safety shutter according to an aspect of the present disclosure. [Figure 16C] A diagram showing a headset with a louvered safety shutter according to an aspect of the present disclosure. [Figure 16D] A diagram showing a headset with a fabric safety shutter according to an aspect of the present disclosure. [Figure 16E] A diagram showing a headset with a liquid crystal safety shutter according to an aspect of the present disclosure. [Figure 17] A diagram showing finger tracking by DVS and a controller according to an aspect of the present disclosure. [Figure 18A] A schematic diagram showing gaze tracking within the context of an aspect of the present disclosure. [Figure 18B] A schematic diagram showing gaze tracking within the context of an aspect of the present disclosure. [Figure 19] A block system diagram of a tracking system by DVS according to an aspect of the present disclosure.
Mode for Carrying Out the Invention
[0008] The following detailed description includes many specific details for purposes of illustration, but those skilled in the art will recognize that many variations and modifications to the following details are within the scope of the invention. Thus, the exemplary embodiments of the invention described below are presented without loss of generality to the claimed invention and without imposing limitations on the claimed invention.
[0009] Preface A new type of visual system called Dynamic Visual Systems (DVS) has recently been developed. This DVS resolves changes in a scene by utilizing only changes in the light intensity of a light-sensitive pixel array. DVSs have a very fast update rate, and instead of delivering a stream of image frames, they provide a nearly continuous stream of the locations of changes in pixel intensity. Each change in pixel intensity is sometimes called an event. This has the added advantage of significantly reducing irrelevant data output.
[0010] Two or more light sources can provide continuous updates regarding the position of the DVS camera relative to a position indicator light at an update rate determined by the flashing speed of the light. In some implementations, the two or more light sources may be infrared light sources, and the DVS may use infrared-sensitive pixels. Alternatively, the DVS may be sensitive to the visible light spectrum, and the two or more light sources may be multiple visible light sources or a single visible light source at a known wavelength. In implementations with a visible-sensitive DVS, the DVS may also be sensitive to motion occurring within its field of view (FOV). The DVS may detect changes in light intensity caused by reflection of light from a moving surface. In implementations using an infrared-sensitive DVS, light from an infrared illuminator can be used to detect motion within the FOV by reflection.
[0011] implementation Figure 1 shows an example of an implementation of game controller tracking using a DVS 101 having a single sensor array, according to one aspect of the present disclosure. In the implementation shown, the DVS is mounted on a headset 102, which may be part of a head-mounted display. A controller 103, including two or more light sources, is within the field of view of the DVS 101. In the example shown, the controller 103 includes four light sources 104, 105, 106, and 107. These light sources have known configurations with respect to each other and with respect to the controller 103. Here we have one DVS with a single photosensitive array. Using such four light sources, the position and orientation of the controller 103 relative to the DVS 101 can be precisely determined. Known information about the light sources may include the distance between each of the other light sources relative to each of the light sources, and the position of each light source on the controller 103. As shown, the three light sources 104, 105, and 106 can draw a plane, and the light source 107 may be outside the plane with respect to the plane drawn by the three light sources 104, 105, and 106. The light sources here have known configurations, and for example, but are not limited to, a first light source 104 located on the upper left front, a second light source 105 located on the upper right front, a third light source 106 located on the left side away from the top, and a fourth light source 107 located on the lower center front of the controller. With four light sources, a DVS having a single photosensitive array may be able to determine the motion of the controller in the X, Y, and Z axes. Furthermore, an inertial measurement unit (IMU) 108 can be coupled to the controller 103. For example, the IMU 108 may include an accelerometer configured to measure acceleration in one, two, or three axes. Alternatively, the IMU may include a gyroscope configured to sense changes in rotation in one, two, or three axes. In some implementations, the IMU may include both an accelerometer and a gyroscope. Using the IMU 108, the determination of motion, position, and orientation can be refined with information from the DVS 101 using a processor. The processor may be located in a headset 102, a game console, or other computing device (not shown).The DVS101, headset102, and IMU108 may be operably coupled to a processor110, which may be located on the headset102, controller103, or on a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor may implement tracking as described herein, for example, as described later with respect to Figures 4, 8, and 9. Furthermore, the processor110 may control the blinking of the light sources 104, 105, and 106.
[0012] During operation, the DVS101, which has a photosensitive array, can detect the movement of light sources 104, 105, 106, and 107 using the photosensitive array, and when a change in light is detected by the photosensitive array, it can be transmitted to the processor. In some implementations, the light sources may be configured to turn on and off in a predetermined pattern using signals from the circuit and / or processor, for example, but not limited to these. The processor can use the predetermined pattern to determine the identity of each light source. The identity of a light source may include a known position relative to the controller and to other light sources. In other implementations, each light source may be configured to turn off and on in a predetermined pattern, and the identity of that specific light source can be determined using that pattern. In some implementations, the processor can adapt the known configuration of the light sources relative to the controller to the events detected by the photosensitive array.
[0013] DVS can have a nearly continuous update rate that can be discretely approximated to approximately 1 million updates per second. High-up-rate DVSs may be able to resolve very fast flickering patterns of light sources. The flickering rate is mainly limited by the Nyquist frequency, i.e., half the DVS's sample rate. The light source can flicker with a duty cycle suitable for detection by the DVS. Generally speaking, the "on" time of the flicker needs to be long enough to be consistently detected by the DVS. Furthermore, due to the high update rate, even slight differences in the flickering rate may be detectable.
[0014] Light sources 104, 105, 106, and 107 may be broad-spectrum light sources such as incandescent lamps or white light-emitting diodes. Alternatively, light sources 104, 105, 106, and 107 may be infrared light sources, or the light sources may have a specific optical spectral profile detectable by the DVS 101. The DVS 101 may include a photosensitive array configured to detect the emission spectra of light sources 104, 105, 106, and 107. For example, but not limited to, if the light sources are infrared light sources, the photosensitive array of the DVS may be sensitive to infrared light, or if the light sources have a specific emission spectrum, the photosensitive array may be configured to be highly sensitive to the specific emission spectrum of the light sources. In addition, for example, but not limited to, the photosensitive array of the DVS may not be sensitive to or may exclude light of other wavelengths not emitted by the light sources, for example, if the light sources are infrared light sources, the photosensitive array may be configured to detect only infrared light.
[0015] Figure 2 shows an example of an implementation of game controller tracking using a DVS having a dual sensor array according to one aspect of the present disclosure. In this implementation, the headset 203 includes a first DVS 201 and a second DVS 202. Alternatively, the headset 203 may include a DVS having a first photosensitive array 201 and a second photosensitive array 202. The general functions of the light source and DVS are the same as those described above with respect to Figure 1. Information from the second DVS or second array can be integrated with information from the first array to provide a better match for controller orientation and any depth information. The two DVSs or photosensitive arrays may have partially overlapping fields of view, enabling the use of binocular parallax.
[0016] Two DVSs or two photosensitive arrays provide binocular vision for depth sensing. This further reduces the number of light sources. The first and second DVSs or the first and second photosensitive arrays may be separated by a known distance, for example, not limited to, but about 50-100 millimeters or more than 100 millimeters. More generally, this distance is large enough to provide sufficient parallax for the desired depth sensitivity, but not so large so that there is no overlap between the fields of view. As shown, the controller 207 may include a first light source 204, a second light source 205, and a third light source 206. A fourth light source coupled to the controller may not be necessary because the information from the two DVSs or two arrays provides sufficient information for determining the controller's position and orientation. The controller may include an IMU 208 which can provide additional inertial information used to refine the determination of position and orientation.
[0017] The first DVS201, the second DVS202, the headset 203, and the IMU208 may be operably coupled to a processor 210, which may be located on the headset 203, the controller 207, or on a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor may implement tracking as described herein, for example, as described later with respect to Figures 4, 8, and 9. Furthermore, the processor 210 may control the flashing of the light sources 204, 205, and 206.
[0018] Figure 2 shows two DVSs or two photosensitive arrays, but the embodiments of this disclosure are not limited thereto. A device may include any number of DVSs or photosensitive arrays. For example, a DVS having three DVSs, or three separate photosensitive arrays, may allow the use of only two light sources coupled to the controller, although this is not limited. The third DVS or photosensitive array does not have to be collinear with the other two DVSs or arrays. Similar to binocular parallax, each of the additional DVSs or photosensitive arrays may be separated by a known distance, and by having partially overlapping fields of view, it is possible to increase the parallax effect used. Furthermore, some implementations may include multiple DVSs, each having multiple photosensitive arrays. For example, there may be two DVSs, each having two separate photosensitive arrays, although this is not limited thereto.
[0019] Figure 3 shows an example of an implementation of game controller tracking using a DVS in combination with a single sensor array and a camera, according to one aspect of the present disclosure. In this implementation, a camera 302 is added to the DVS 301. The DVS 301 and camera 302 may be coupled to a headset 303. The DVS 301 and camera 302 may have partially overlapping fields of view or may share the same field of view. The controller may include three or more light sources 304, 305, 306 on the controller 307. The DVS 301 and camera 302 can be used together to determine the position and orientation of the controller. Frames from camera 302 may be interpolated using events from the DVS 301. Image frames can also be used to perform localization and mapping simultaneously to improve the determination of the controller's orientation and position. Furthermore, image frames from the camera can be used to perform inside-out tracking of the user using machine learning algorithms, e.g., hand tracking or foot tracking. The IMU 308 can provide additional inertial information used to further refine the determination of the controller's position and orientation.
[0020] The first DVS301, the second DVS302, the headset303, and the IMU308 may be operably coupled to a processor310, which may be located on the headset303, the controller307, or on a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor may implement tracking as described herein, for example, as described later with respect to Figures 4, 8, and 9. Furthermore, the processor310 may control the flashing of the light sources 304, 305, and 306.
[0021] Figure 5 shows an example of an implementation of head tracking or tracking of other devices using a controller including a DVS having a single sensor array, according to one aspect of the present disclosure. Here, the DVS 507 is mounted on the controller 508. A headset 501 within the field of view of the DVS 507 is tracked using light sources mounted on the headset. The headset 501 may include two or more light sources, in this case four light sources 502, 503, 504, and 505. The four light sources provide precise information for determining the three-dimensional position by a single DVS. The number of light sources used may be reduced by using additional DVSs or cameras. The position of each light source relative to the headset 501 is known to the system and may be somewhat important for providing relevant information to the DVS. Here, light sources 502, 503, and 504 represent a plane. Light source 505 is located outside the plane of the other light sources 502, 503, and 504. This facilitates the three-dimensional detection of the headset's position, orientation, or motion. The headset 501 may include an IMU 506, which can improve position and orientation estimation using information from the DVS 507. Furthermore, the controller 508 may include an IMU 509. Information from the controller IMU 509 can be used to further refine the determination of the headset's position and orientation. For example, but not limited to, the IMU information can be used to determine whether the controller is moving relative to the headset and to determine the speed or acceleration of that movement.
[0022] The headset 501, IMU 506, DVS 507, and controller IMU 509 may be operably coupled to a processor 510, which may be located on the headset 501, controller 508, or on a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor may implement tracking as described herein, for example, as described later with respect to Figures 4, 8, and 9. Furthermore, the processor 510 may control the flashing of the light sources 502, 503, 504, and 505.
[0023] Figure 6 shows an example of an implementation of head tracking or other device tracking using a game controller including a DVS having a dual sensor array, according to one aspect of the present disclosure. Here, the controller 605 is coupled to two DVSs 606, 607, or a single DVS having two photosensitive arrays 606 and 607. As described above, the two DVSs or two photosensitive arrays may be separated by an appropriate distance, e.g., between 500 and 1000 millimeters, and may have partially overlapping fields of view. This makes it possible to use the parallax effect for depth determination. Furthermore, the headset 601 may include three light sources 602, 603, 604 and an IMU 608. The use of two DVSs or two separated arrays 606, 607 may allow the use of fewer than four light sources, e.g., three, but not limited to. Two light sources 602, 603 may draw a line, and a third light source 604 may be outside the line drawn by the other two light sources 602, 603. Information from the IMU 605 coupled to the headset 601 can be used to refine the determination of position and orientation.
[0024] The headset 601, DVS 606, DVS 607, and IMU 608 may be operably coupled to a processor 610, which may be located on the headset 601, controller 605, or on a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor may implement tracking as described herein, for example, as described later with respect to Figures 4, 8, and 9. Furthermore, the processor 610 may control the blinking of the light sources 602, 603, and 604.
[0025] Figure 7 shows an example of an implementation of head tracking or other device tracking using a controller including a DVS in combination with a single sensor array and an image camera, according to one aspect of the present disclosure. Here, the DVS 707 and camera 708 are coupled to controller 706. The camera 708 and DVS 707 may have partially overlapping fields of view or fully overlapping fields of view. In some implementations, the pixels of the camera and the photosensitive elements of the DVS may share the same photosensitive array, thereby functioning as an integrated DVS and camera.
[0026] Three or more light sources 702, 703, 704 can be combined with the headset 701. For example, but not limited to, the light sources may be integrated within the headset housing, and each light source may be an LED, incandescent, halogen, or fluorescent light source mounted within or on a circuit board on the headset housing. In some implementations, a single light source may create multiple light sources using a plastic or glass light pipe or optical fiber that splits the light from the single light source to two or more light sources on the headset housing.
[0027] Three or more light sources 702, 703, 704 may be configured to turn on and off in response to electronic signals. In some implementations, the three or more light sources can be turned on and off in a predetermined sequence. Each time a light source 702, 703, 704 moves or blinks, the DVS 707 can generate an event. The camera 708 generates image frames of its field of view at a set frame rate. Due to the high update rate of the DVS, it may be possible to interpolate the image frames generated by the camera with DVS events.
[0028] In some implementations, the IMU 705 can also be coupled to the headset 701. The headset 701, IMU 705, DVS 707, camera 708, and IMU 608 may be operably coupled to a processor 710, which may be located on the headset 701, controller 706, or on a separate device such as a personal computer, laptop computer, tablet computer, smartphone, or game console. The processor can implement tracking as described herein, for example, as described later with respect to Figures 4, 8, and 9. Furthermore, the processor 710 can control the blinking of the light sources 702, 703, and 704.
[0029] Furthermore, in some implementations, camera 708 may be a depth camera, such as a depth time-of-flight (DTOF) sensor. A DToF camera acquires a depth image by measuring the time it takes for light to travel from a light source to an object in the scene and back to the pixel array. As an example, but not an limitation, a DToF camera can be operated using continuous wave (CW) modulation, which is an example of an indirect time-of-flight (ToF) sensing method. In a CW ToF camera, light from an amplitude-modulated light source is backscattered by an object in the camera's field of view (FOV), and the phase shift between the emitted waveform and the reflected waveform is measured. By measuring the phase shift at multiple modulation frequencies, a depth value for each pixel can be calculated. The phase shift is obtained by measuring the correlation between the emitted and received waveforms at different relative delays using intra-pixel photon mixed demodulation.
[0030] A DToF system generally includes an illumination module and an imaging module. The illumination module consists of a light source, a driver that drives the light source at a high modulation frequency, and a diffuser that projects a light beam from the light source into a designed illumination field (FOI). A DToF illumination module may include one or more light sources, which may be amplitude-modulated light sources such as, for example, a vertical-cavity surface-emitting laser (VCSEL) or an end-face-emitting laser (EEL), for example. The imaging module may include an imaging lens assembly, a bandpass filter (BPF), a microlens array, and an array of photosensitive elements that convert incident photon energy into electronic signals. The microlens array increases the amount of light reaching the photosensitive elements, and the BPF reduces the amount of ambient light reaching the photosensitive elements and the microlens array.
[0031] operation Figure 4 shows a DVS that tracks the motion of a game controller having two or more light sources, according to one aspect of the present disclosure. The DVS 401 has a controller 402 within its field of view. As shown, the controller 402 includes a plurality of light sources coupled to the controller body. The plurality of light sources can be configured to turn off and on again at a predetermined rate. Each blink of a light source within the field of view of the DVS 401 can generate one or more events 403 in the DVS. In the shown event 403, the event indicates that all light sources are on, since all light sources have changed from an off state to an on state. Alternatively, depending on the sensitivity of the DVS, a change in the brightness of the light sources may be sufficient to trigger event 403. Note that in other implementations, the lights may turn off and on at different rates or at different times, so each event may correspond to fewer light sources than all of the light sources that are lit. The time of the events generated by the blinking lights and their corresponding positions in the array can be recorded in memory (not shown). In some implementations, the position and orientation of the controller 402 can be reconstructed from one or more events from the DVS.
[0032] In some implementations, each light source may blink in a predetermined time sequence. The DVS may output the time at which each event occurred, for example, as a timestamp along with each event. The predetermined time sequence may be used at the time the event occurred to determine the identity of each light within the event, for example, which event corresponds to which light source location. In the example shown, one or more events 403 output by the DVS 401 represent the first light 406, second light 407, third light 408, and fourth light 409 detected by the photosensitive array. As described above, the identity of each light source may be determined from the information output by the DVS and the predetermined blinking sequence of the light source. For example, the photosensitive array of the DVS 401 may detect a light event 403 at time T+1, and the predetermined sequence may provide that light source 406 turns on at T+1, so it is determined that the light event corresponds to light source 406. The predetermined sequence may be stored in memory, for example, as a table listing the timing and position of the sequence for each light source. In some implementations, a given sequence may be encoded in the flashing of the light itself, for example, each light may flash with a sequence indicating its identity. For example, a light source labeled 1 may flash with a Morse code sequence indicating the number 1. The identity of the light event can then be recovered through analysis of the light event. Alternatively, the sequence information may come from the light source itself or the light source driver and indicate when the light source is turned on or off. Alternatively, the light source may be turned on and off simultaneously, and a machine learning algorithm can be applied to the detected light event 403 to match the controller attitude 402 to the event and a known configuration of the light source.
[0033] Furthermore, information from events such as light size, intensity, and separation can be used to determine the orientation and position of the light sources. The system may have information defining the size, position on the controller body, and intensity of each light source. The position and orientation may be determined with greater accuracy from the differences between the detected size, intensity, and separation. In addition, if one or more additional DVSs or photosensitive arrays are present, parallax information can be used to further enhance the determination of position and orientation.
[0034] During operation, the user can change the position and orientation of the controller 410. Due to the relatively high update rate of the DVS, as the optical event moves in 412 during changes in position and orientation, the DVS at 411 is able to capture the event sequence. The motion of the optical event here is shown in Figure 4 by the vector arrow 412. As the detected optical event moves, the determined position and orientation of the controller 410 may be updated, or the determined position and orientation may be updated at regular intervals. Due to the flashing of the first light 415, second light 416, third light 417, and fourth light 418, the position and time of the event detected by the photosensitive element of the DVS can be adapted to the new position and orientation of the controller. Furthermore, inertial information from the IMU can be used to refine the motion, position, and orientation of the controller 410.
[0035] The flowchart shown in Figure 8 illustrates a method for tracking motion using DVS801 with one or more light sources and one light source configuration adaptation model, according to one aspect of the present disclosure. In this implementation, the light sources (shown as LEDs) can be turned on and off simultaneously, independently, or in a predetermined sequence. Movement of the light sources or flashing of one or more of the light sources generates an event 802 from DVS801. Generally, an event includes an electronic signal relaying the following information: the time interval in which the event occurred, the position of the DVS801 photosensitive element in the array where the event occurred, and binary data (e.g., 1 or 0) corresponding to a change in light intensity exceeding some detection threshold. Each event can be processed in 803 to format the event into a usable format, for example, by placing the event location in a data array, associating a timestamp with the event, and aggregating multiple events. As an example of aggregating events, all events occurring at individual elements in the array within a given time interval can be integrated into a single data structure for analysis.
[0036] Next, the processed events can be analyzed to associate the detected DVS events with the corresponding LED pulses 804. In some implementations, each light source can be turned on and off at unique predetermined time intervals, so the timestamps of the aggregated events can be used to determine the predetermined time interval from the events and associate specific light sources with specific events. For example, but not limited to, the spatial pattern of aggregated events occurring within a predetermined time interval or sequence of time intervals can be analyzed to determine whether this pattern is consistent with the LED pulses. Event patterns that are too large, too small, or too irregular in shape may be excluded as LED events. Furthermore, the timing of the events can also be analyzed, and events that are too short or too long may also be excluded.
[0037] A trained machine learning model 805 can be applied to the processed events. Model 805 may include information about the configuration of the light sources, such as the size of the light sources and their relative positions to the controller body. The machine learning model can be trained with training event data having corresponding masked positions and orientations of the controller, as described in a later section. The trained machine learning model is applied to the processed event data to determine the correspondence 806 between the detected pulse 804 and the attitude 808. The trained machine learning model can fit the attitude 808 of the controller, e.g., position and orientation, to one or more processed events, e.g., detected LED pulse 804. Alternatively, a fitting algorithm can be applied to the processed events instead of the trained machine learning model 805. The fitting algorithm can fit the position and orientation of the controller to the processed events using a manually developed light source model. Alternatively, the fitting algorithm may be a hypothesis-test type algorithm that tries all possible permutations of light correspondences and uses redundant light sources to find the best fit. A tracking / prediction algorithm can then be applied to continue tracking the light sources. Furthermore, 809 can use the predicted current attitude to predict the next attitude. In the 810, inertial data from IMU807 can be fused with the predicted attitude 808 to generate the controller's final predicted position and orientation. The fusion may be performed by a trained machine learning algorithm that has been trained to refine the controller's position and orientation using the inertial data. Alternatively, the fusion may be performed by, for example, a Kalman filter or nonlinear optimization, but is not limited to these.
[0038] Figure 9 shows a motion tracking method using a DVS with timestamped light source position information according to one aspect of the present disclosure. In this implementation, the time of events output by the DVS 901 is used to determine the position and orientation of the controller. Here, the light source is depicted as an LED, and each LED may be turned on at different times as shown in the graph. Each time an LED is turned on or off, a DVS having the LED in its field of view can generate an event 902. Each event may include an electronic signal corresponding to information such as the time interval in which the event occurred, the position of the photosensitive element in the DVS 901 where the event occurred, and binary data (e.g., 1 or 0) corresponding to a change in light intensity that exceeds some detection threshold. Each event can be processed by 903 to format the event into a usable format, for example, by arranging the event position in a data array, as described above, and by integrating multiple events occurring at individual elements in the array within a given time interval into a single data structure, thereby aggregating multiple events. The processed events can then be analyzed by 904 to detect the corresponding LED pulse. For example, though not limited to, the shape and size of events may be analyzed in relation to the regularity and fit of the light source, and events that are too large, too small, or too irregular in shape may be excluded as LED events, and noise suppression such as averaging of events may be performed to remove random events. Furthermore, the timing of events may also be analyzed, and events that are too short or too long may be excluded, or multiple events may be condensed into a single event by excluding events that co-occur with an initial event and are within the same timeframe but end after the initial event.
[0039] Once an event is processed, the 905 can determine the position of individual LEDs. Determining the position of individual LEDs can be done by using the time sequence in which the LEDs are turned on and off. For example, the time of one or more events can be compared to a known time sequence of LED blinking, for example. The known time sequence may be, for example, a table containing the on and off times of the LEDs and their positions on the controller body for each LED, or a timestamp from the LED driver when each LED is on or off. If timestamps are used, the timestamps can be correlated with the LED positions on the controller body. From the timing sequence, LED position information, and processed event information, the matching position and orientation of the controller can be determined. IMU data 907 can be integrated with the previously determined LED positions and inertial data from the IMU by the Kalman filter 908. The Kalman filter can predict the position of the light source based on the inertial information from the IMU, and this prediction can be integrated with the position information determined from the LED time sequence to refine the motion data, and the 909 can generate the final attitude and refine future projections.
[0040] Training of a typical neural network According to aspects of this disclosure, the tracking system may utilize machine learning with a neural network (NN). For example, the trained model 805 discussed above may utilize machine learning as discussed below. The machine learning algorithm may use a training dataset, which may include inputs from a DVS such as events, or processed events having known controller positions and orientations as labels. Furthermore, a machine learning algorithm using an NN may perform fusion between the controller positions and orientations determined from the DVS information and inertial information from an IMU. The training set for fusion may, for example, be inertial data with potential controller positions and orientations, as well as final positions and orientations, though not limited to these. In some implementations, the machine learning algorithm may be trained to perform SLAM (simultaneous localization and mapping) using a training set that includes objects such as the ground, landmarks, and body parts with hidden labels. Hidden labels may include the identity of objects and their relative positions. As is generally understood by those skilled in the art, SLAM techniques generally solve the problem of building or updating a map of an unknown environment while simultaneously tracking the position of an agent within it.
[0041] The NN may include one or more of several different types of neural networks and may have many different layers. For example, and not limited to, the neural network may consist of one or more convolutional neural networks (CNNs), recurrent neural networks (RNNs), and / or dynamic neural networks (DNNs). The motion detection neural network may be trained using the general training methods disclosed herein.
[0042] As an example rather than an limitation, Figure 10A shows a basic form of an RNN that may be used, for example, in the trained model 805. In the illustrated example, the RNN has a layer of 1020 nodes, each node characterized by an activation function S, one input weight U, a recursive hidden node transition weight W, and an output transition weight V. The activation function S may be any nonlinear function known in the art and is not limited to the hyperbolic tangent (tanh) function. For example, the activation function S may be a sigmoid function or a ReLU function. Unlike other types of neural networks, the RNN has one set of activation function and weights for the entire layer. As shown in Figure 10B, the RNN may be thought of as a series of nodes 1020 having the same activation function moving through time T and T+1. Thus, the RNN maintains historical information by feeding the results from the previous time T to the current time T+1.
[0043] In some implementations, convolutional RNNs may be used. Another type of RNN that can be used is the Long Short-Term Memory (LSTM) neural network, which adds memory blocks to the RNN nodes along with input gate activation functions, output gate activation functions and forget gate activation functions, which result in gating memory that allows the network to retain some information for longer periods, as described by Hochreiter & Schmidhuber, "Long Short-term memory," Neural Computation 9(8):1735-1780 (1997), incorporated herein by reference.
[0044] Figure 10C shows an exemplary layout of a convolutional neural network, such as a CRNN, which may be used for a trained model 805, etc., according to an aspect of the present disclosure. In this representation, the convolutional neural network is generated for an input 1032 having a size of 4 units in height and 4 units in width, giving a total area of 16 units. The represented convolutional neural network has a filter 1033 having a size of 2 units in height and 2 units in width, with channels 136 having a skip value of 1 and size 9. For clarity in Figure 10C, only the connections 1034 between the first row of channels and their filter window are represented. However, aspects of the present disclosure are not limited to such implementations. According to an aspect of the present disclosure, the convolutional neural network may have any number of additional neural network node layers 1031, which may include any type of additional convolutional layers, fully connected layers, pooling layers, maximum pooling layers, local contrast normalization layers, etc., of any size.
[0045] As seen in Figure 10D, training a neural network (NN) begins with initializing the NN's weights at 10⁴¹. Generally, the initial weights should be randomly distributed. For example, an NN with a tanh activation function should have random values distributed between -1 / √n and 1 / √n, where n is the number of inputs to the node.
[0046] After initialization, the activation function and optimizer are defined. Then, in step 1042, the NN is provided with feature vectors or an input dataset. Each of the different feature vectors may be generated by the NN from an input with known labels. Similarly, the NN may be provided with feature vectors corresponding to an input with known labeling or classification. The NN then, in step 1043, predicts the labels or classifications for the features or input. The predicted labels or classes are compared in step 1044 to the known labels or classes (also known as ground truth), and the loss function measures the total error between the predictions and ground truth throughout all training samples. The loss function may be a cross-entropy loss function cost, a triplet contrasive function, an exponential cost, etc., not as an example and not as an extension. Multiple different loss functions may be used depending on the purpose. For example, a cross-entropy loss function may be used to train a classifier, while a triplet contrasive function may be employed to learn a pre-trained embedding. The NN may then be optimized and trained using the results of the loss function and known methods for training neural networks, such as backpropagation by adaptive steepest descent, as shown in 1045. In each training epoch, the optimizer attempts to select model parameters (i.e., weights) that minimize the training loss function (i.e., total error). The data is divided into training samples, validation samples, and test samples.
[0047] During training, the optimizer minimizes the loss function for the training samples. After each training epoch, the model is evaluated on the validation samples by calculating the validation loss and accuracy. If there is no significant change, training may be stopped, and the resulting trained model may be used to predict the labels of the test data.
[0048] Therefore, a neural network may be trained from inputs to identify and classify inputs that have known labels or classifications. Similarly, a NN can be trained using the described methods to generate feature vectors from inputs that have known labels or classifications. The above discussion relates to RNNs and CRNNS, but these discussions can be applied to NNs that do not have recursive or hidden layers.
[0049] Hybrid Sensor Figure 11A illustrates a hybrid DVS having multiple colocation sensor types according to an aspect of the present disclosure. In some implementations, a hybrid DVS can be used to combine multiple sensor types into a single device. The multiple sensor types may include, for example, DVS infrared photosensitive elements, DVS visible light photosensitive elements, DVS wavelength-specific photosensitive elements, visible light camera pixels, infrared camera pixels, and DTOF camera pixels. Second sensor types 1102 may be scattered on the same array 1101 as a first sensor type 1102. For example, for example, a DVS visible light photosensitive element 1103 may surround a visible light camera pixel 1102, or a DVS infrared photosensitive element 1103 may surround a DVS visible light photosensitive element 1102, or a visible light camera pixel 1103 may surround a DVS infrared photosensitive element 1102, or any combination thereof. Although a single element 1102 is shown surrounded by other elements 1103, the embodiments of this disclosure are not limited thereto. A single element may include a plurality of DVS photosensitive elements or clusters of camera pixels, for example, a cluster of four camera pixels may be surrounded by eight DVS photosensitive elements, or a single camera pixel may be surrounded by eight pairs, triplets, or quadlets of DVS photosensitive elements.
[0050] Furthermore, as shown in Figure 11B, the hybrid DVS may have multiple sensor types arranged in a grid pattern within the array. Here, blocks of the first sensor type 1102 are evenly distributed with blocks of the second sensor type 1103. The first sensor type 1102 and the second sensor type 1103 may each be different. For example, but are not limited, the first sensor type 1102 may be a DVS infrared photosensitive element, and the second sensor type 1103 may be a DVS visible light photosensitive element.
[0051] Alternatively, the sensor types may be distinguished by filtering. In these implementations, one or more filters selectively transmit light to the photosensitive elements located behind the one or more filters. The one or more filters may, for example, selectively transmit one or more specific wavelengths or specific polarizations of light. The one or more filters may also selectively block one or more specific wavelengths or specific polarizations of light. The photosensitive elements behind the one or more filters may be configured to be used as different sensor types. For example, an infrared pass filter may cover one or more sensor elements in array 1101 that allows only infrared light to pass through 1102, while other sensor elements may not be filtered or may be an infrared cut filter 1103. In another alternative implementation, the one or more filters may, for example, be optical notch filters that allow only specific wavelengths of light to pass through 1102, while other filters may block those specific wavelengths but allow other wavelengths to pass through 1103. This filtering may reduce the likelihood of false light source detection by allowing the use of a specific wavelength for the illuminator light for the DTOF and a specific wavelength for the DVS photosensitive element. Here, the sensor element may be any type of DVS photosensitive element or any type of camera pixel.
[0052] Patterned sensor or filter elements can be incorporated into hybrid imaging units, for example, as shown in Figures 11C and 11D. Figure 11C shows an example of a hybrid DVS imaging unit 1112C having one or more lens elements 1114, an optional microlens array 1116, a bandpass filter 1118, and a patterned hybrid sensor array 1120. The sensor array includes DVS sensor elements 1122 and conventional imaging sensor elements 1124, which may be arranged in a pattern, for example, as shown in Figure 11A or Figure 11B. Such imaging units can be used with an illumination unit (not shown) in a DTOF sensor. The hybrid DVS imaging unit 1112D shown in Figure 11D has a DVS sensor array 1126, for example, as shown in Figure 11A or Figure 11B, and a patterned filter element 1118 including a patterned bandpass region 1118A and a bandcut region 1118B.
[0053] Figure 12 illustrates a hybrid DVS according to an aspect of the present disclosure, which includes multiple sensor types having inputs separated by an optical separator. In this implementation, the optical separator 1204 can filter light based on wavelength, and the array 1201 can include multiple sensor element types that are physically separated based on the wavelength or polarization of the light to be detected. For example, but not limited to, unpolarized white light 1205 (a mixture of at least all visible light wavelengths, most often including some infrared wavelengths) can enter the optical separator 1204, which may be a dispersion prism, diffraction grating, or dichroic mirror. As shown, the white light 1205 entering the optical separator 1204 can be separated by wavelength or polarization, with some light 1206 incident on the first part of the array 1202 and other light 1207 incident on the second part of the array 1203. Here, the optical separator 1204 can be thought of as a filter that changes the diffraction angle based on wavelength or polarization. The arrays shown are depicted as a single unit with a thick separation line 1203, but embodiments of the disclosure are not limited thereto. The first part of array 1201 may have a separation of up to 1 millimeter between it and the second part of array 1202, and the arrays are shown separated vertically, but other implementations may have parts separated horizontally, diagonally, or circumferentially. Furthermore, the optical separators present herein can be combined with different filtering or different sensor configurations, such as those shown in Figures 11A and 11B, to provide additional optical wavelength separation for different sensor types.
[0054] Figure 13 illustrates a hybrid DVS according to an aspect of the present disclosure, in which multiple sensor types have inputs separated by micro-electromechanical (MEMS) mirrors. Here, the MEMS mirror 1304 can oscillate between different positions at set intervals so as to reflect light 1305 to either a first part 1301 or a second part 1302 of the array, depending on the time it takes for light 1305 to reach the MEMS mirror 1304. The first part 1301 and the second part 1302 of the array can be physically separated from each other at 1303 based on the angle of incidence of the light reflected by the MEMS mirror 1304. In this way, light can be temporarily filtered between different sensor types. Such temporary filtering can be timed according to a known flashing pattern of a light source on a controller or headset. For example, but not limited to, a light source on a controller or headset can be turned on at regular intervals for a predetermined duration. For example, a light source can be turned on for 60 microseconds every 100 microseconds. In such cases, the MEMS mirror 1304 can be synchronized to reflect light 1305 to the DVS portion of array 1301 for more than 60 microseconds at intervals of less than 100 microseconds in order to capture changes in the light source. Other reflected light 1307 can be detected by the second portion of array 1302 during times when the light source is off and can capture ambient light for image tracking or DTOF.
[0055] Alternatively, the MEMS mirror 1304 can filter the light 1305 based on wavelength. The MEMS mirror in these implementations may be, for example, a MEMS Fabry-Perot filter or a diffraction grating, for example, but not limited to these. The MEMS mirror can diffract light in the first wavelength range 1306 to at least the first part 1301 of the array, or light in the second wavelength range 1307 to the second part 1302 of the array.
[0056] Figure 14 illustrates a hybrid DVS according to an aspect of the present disclosure, in which multiple sensor types have a transiently filtered input. Here, the filter allows light to selectively pass through the array based on time. In some implementations, the filter may be an optical waveguide. In the first time step, the filter may allow a first wavelength or polarization of light to pass through array 1401, while blocking a second wavelength or polarization of light, or other wavelengths or polarizations. In the second time step, the filter may allow the second wavelength or polarization 1402, but block the first wavelength or polarization of light, or other wavelengths or polarizations. In this way, light can be transiently filtered, which may be useful for tracking different sensor types. For example, but not limited to, a light source coupled to a controller or headset may be infrared light or have a specific wavelength. The light source may be configured to turn on and off at specific intervals. The specific intervals in which the light source turns on and off may be a sequence or coded pattern. The transient filtering may be activated during specific intervals to allow a specific wavelength to pass through the sensor and block other wavelengths. Furthermore, the filtering switching interval may be longer than a specific interval between light sources, taking into account the time it takes for light to travel to the sensor.
[0057] Body tracking Figure 15 is an illustration illustrating body tracking by DVS according to an aspect of the present disclosure. As shown, user 1501 can wear a headset 1504 having a DVS 1503, where the DVS is shown as two arrays or DVS units or together with a camera and DVS. User 1501 holds two controllers 1502, each having two or more light sources 1505. DVS 1503 has a controller 1502 with corresponding light sources 1505 within its field of view (FOV). Furthermore, DVS 1503 may have the user's hands or arms 1507, or the user's limbs such as legs or feet 1508 within its FOV. The DVS may also have the ground or other landmarks 1509 within its FOV. For example, but not limited to, the photosensitive elements of the DVS may detect reflections of light corresponding to the user's limbs or the ground or other landmarks when the user moves or when there is a change in light. The camera detects reflections of light from the field of view at the camera's frame rate. From the detected light reflection, the 1506 can determine the user's appendages, the ground, or a landmark. In some alternative implementations, an external DVS 1510 can be used to track the user's appendages. The external DVS 1510 may be located at a distance from the user, selected so that the user's appendages fit within the field of view of the external DVS 1510. For example, but not limited to, the external DVS may be located on or below a television or computer monitor, or other freestanding or wall-mounted display.
[0058] The machine learning algorithm is trained to determine the user's body, appendages, ground, or landmarks, as well as their relative positions and orientations, from data such as events or frames. The machine learning algorithm may be a neural network, and the training may be similar to the methods described in the general neural network training section of Figures 10A-10D above. The neural network can be trained using a training set that includes labeled events or frames, or both. Labelled events or frames, or both, may include, for example, labels for the user's body, appendages, ground, landmarks, and the relative positions and orientations of the user's body, appendages, ground, and landmarks. In addition to determining the position and orientation of the controller 1502, the determination of the user's body, appendages, ground, or landmarks, as well as their relative positions and orientations, can be performed using, for example, SLAM. Alternatively, the position and orientation of the controller can be determined using the same trained neural network that determines the labels for the user's body, appendages, ground, or landmarks, and their relative positions and orientations. As shown by element 1506, the model can be adapted to the determined user's body and appendages in order to improve the determination of relative position and location.
[0059] Safety shutter As described above, body tracking can be used in conjunction with determining the position and orientation of the controller to trigger a safety shutter in a VR or AR headset. Figure 16A shows a headset with a safety shutter door according to an aspect of the present disclosure. The headset 1601 may include a head strap 1604, an eyepiece 1603, and a display screen 1602. The eyepiece 1603 may include one or more lenses configured to focus on the display screen 1602. The one or more lenses may be, for example, Fresnel lenses or prescription lenses. The display screen 1602 may be transparent, and a hole 1607 in the body of the headset may allow vision through the display screen 1602 when the safety shutter is open.
[0060] In this implementation, the safety shutter 1606 is a door that swings away from the hole 1607 when the safety system is activated. Here, a system-operated clasp 1605 interacts with a clasp 1608 on the safety shutter door 1606 to close and secure the door over the hole 1607. A spring-loaded hinge 1609 can ensure that the safety shutter door 1606 opens quickly when the system-operated clasp 1605 opens. The spring-loaded hinge may have, for example, a clock spring wound around the hinge and fixed to the door, which is wound when the door closes and unwound when the door opens. Alternatively, a leaf spring may push the safety shutter door 1606 when it is closed.
[0061] The system-operated clasp 1605 can be configured to open the clasp when the safety system is activated. The safety system-operated clasp 1605 may include an electric motor or linear actuator to move the clasp. The safety system may be activated if the system, the user, the user's body, or the user's appendages detect the ground or one or more landmarks near them. The safety system may use the determination of the user's body, appendages, the ground or landmarks, and their relative positions and orientations, as described above. When the safety system is activated, a signal to open the clasp can be sent to the safety system-operated clasp 1605. When the clasp is opened, the spring-loaded hinge 1609 pushes open the safety shutter door 1606, allowing the user to see through the display 1602 and avoid the danger of activating the safety system.
[0062] Other safety shutter implementations may be used. For example, Figure 16B shows an alternative headset with a sliding safety shutter according to an aspect of the present disclosure, where, when the safety system is activated, the safety shutter slide 1616 slides so as not to obstruct the hole 1607. The safety shutter slide 1616 can move on a spring-loaded rail 1619. Alternatively, the sliding safety shutter 1616 may include a tab inserted into a slot 1619 in the body of the headset, and a spring in the slot may also push the sliding safety shutter. In some additional alternative implementations, the spring may be omitted, and the sliding safety shutter 1616 may operate by gravity. The spring-loaded rail 1619 can push the sliding safety shutter 1616 when closing it, and can ensure that the safety shutter opens quickly when the safety system is activated. The system-operated clasp 1605 interacts with the clasp 1618 on the sliding safety shutter 1616 to close and lock the slide.
[0063] The safety system may be activated if the system, the user, the user's body, or the user's appendages detect the ground or one or more landmarks nearby. The safety system may use the determination of the user's body, appendages, the ground or landmarks, and their relative positions and orientations, as described above. When the safety system is activated, a signal to open the clasp can be sent to the system-operated clasp 1605. When the clasp is opened, the spring 1619 pushes open the sliding safety shutter 1616, allowing the user to see through the display 1602 and avoid the danger of activating the safety system.
[0064] Figure 16C shows a headset with a louvered safety shutter according to an embodiment of the present disclosure. In this implementation, the slats of the louvered safety shutter 1626 have a first dimension that is longer than the second dimension. In the closed position, the longitudinal dimension of the slats is approximately parallel to the optical system, and each slat overlaps either another slat or the headset body, blocking light from passing through the display screen 1602. In the open position, the slats are positioned such that their short dimension is parallel to the optical system 1603, allowing light to pass through the slats and reach the display screen. An actuator rod 1628 can hinge each of the slats 1626. A safety system-controlled actuator 1629 can push or pull the actuator rod 1628 to open the louvered safety shutter when the safety system is activated. In some other implementations, the safety system-controlled actuator may be spring-loaded, and the clasps connected to the actuator rod 1628 and the system-controlled actuator 1629 may include clasps that interlock with the actuator rod clasps. When the safety system is activated, the clasps of the system-controlled actuator can open, allowing the actuator rod to move under spring pressure and the louvers to open.
[0065] The safety system may be activated if the system, the user, the user's body, or the user's appendages detect the ground or one or more landmarks nearby. The safety system may use the user's body, appendages, the ground or landmarks, and the determination of their relative positions and orientations, as described above. When activated, it may send a signal to the system-operated actuator to move the actuator rod. When the actuator rod 1629 moves, it pushes open the slats of the louvered safety shutter 1626, allowing the user to see through the display 1602 and avoid the danger of activating the safety system.
[0066] Figure 16D shows a fabric safety shutter according to an embodiment of the present disclosure. In this implementation, the safety shutter 1636 is made of an opaque fabric, for example, but not limited to, a tightly woven cotton fabric, polyester, vinyl, or a tightly woven woolen fabric. The fabric safety shutter 1636 may be coupled to a fabric roller 1639. The fabric roller 1639 may be configured to wind up the fabric shutter when the safety system is activated. The fabric roller may be spring-loaded, for example, using a clock spring, so that the clock spring is under tension when the safety shutter is closed, and an electric motor may be used to wind up the fabric safety shutter. A system-operated clasp 1605 works in conjunction with a clasp 1638 coupled to the fabric safety shutter 1636 to ensure that the fabric safety shutter does not open unintentionally.
[0067] During operation, the fabric safety shutter may be in the closed position. The safety system may be activated if the ground or one or more landmarks are detected near the system, the user, the user's body, or the user's appendages. The safety system may use the determination of the user's body, appendages, the ground or landmarks, and their relative positions and orientations, as described above. When the safety system is activated, a signal to open the clasp can be sent to the system-operated clasp 1605. When the clasp opens the fabric roller 1619, it rolls up the fabric safety shutter 1616, which the user can see through the display 1602, thus avoiding the risk of activating the safety system.
[0068] Figure 16E shows a headset with a liquid crystal safety shutter according to an embodiment of the present disclosure. In this implementation, the liquid crystal screen 1646 is integrated into the headset 1601 and otherwise covers a hole in the headset. A safety system-controlled liquid crystal screen driver 1649 is communicatively coupled to the liquid crystal screen 1646.
[0069] The liquid crystal screen 1646 may, for example, be a liquid crystal shutter having a first polarizer and a second polarizer, the first polarizer having a polarization difference of 90 degrees from the second polarizer, and having a fluid-filled cavity. The fluid-filled cavity may contain liquid crystals, which are configured to have a first orientation when there is no electric field that changes the polarization so that light can pass from the first polarizer to the second polarizer. The liquid crystals may be further configured to align to a second orientation under an electric field. Since the second orientation of the liquid crystals does not change the polarization of light, light passing through the first polarizer is blocked by the second polarizer. Electrodes may be positioned along the surface of the liquid-filled cavity to enable control of the liquid crystals. A safety system-controlled liquid crystal screen driver may be communicatively coupled to these electrodes, so that the electrodes can control the liquid crystals in the fluid-filled cavity. As used herein, communicative coupling means that electrical signals representing messages or commands can be transmitted from one coupling element to the other and / or received, and these signals can pass through intermediate elements, their format may change, but the message contained therein remains unchanged.
[0070] During operation, the safety system-controlled driver 1649 can send a signal to the liquid crystal safety shutter 1646 to make it opaque while the display screen 1602 is active. When the safety system is activated, the driver 1649 can make the liquid crystal safety shutter 1646 transparent. For example, but not limited to, the driver can reduce the voltage supplied to the liquid crystal safety shutter, returning the liquid crystals to their first orientation, which changes the polarization of light so that light can pass through the second polarizer. The safety system may be activated if the ground or one or more landmarks are detected near the system, user, user's body, or user's appendages. The safety system can use the determination of the user's body, appendages, ground or landmarks, as well as their relative positions and orientations, as described above.
[0071] Finger position tracking Aspects of this disclosure may be applied to finger tracking. Figure 17 shows finger tracking using a DVS and controller according to an aspect of this disclosure, where the controller 1701 includes two or more light sources and one or more buttons 1705. The two or more light sources include one or more light sources 1706 adjacent to one or more buttons 1705 and two or more other tracking light sources 1703. As described above, the two or more other tracking light sources 1703 can generate events in the DVS 1702 that are used to determine the position and orientation of the controller 1701. A DVS having two DVS 1702, or two arrays and three other tracking light sources 1703, as shown herein, is used to determine the position and orientation of the controller.
[0072] One or more light sources 1706 adjacent to one or more buttons 1705 can be used for finger tracking. For example, but not limited to, finger tracking can be achieved in DVS 1702 using the occlusion of one or more light sources 1706 adjacent to the buttons 1705. One or more light sources 1706 adjacent to the buttons can be turned on and off at predetermined intervals. DVS 1702 can generate an event with each blink. By analyzing the event, the occlusion of one or more light sources 1706 adjacent to the buttons 1705 can be determined. Knowing the configuration of one or more light sources adjacent to the buttons, for example, but not limited to, when occluded by a finger or palm, the light pattern detected in the event generated by DVS will be different from when the light source is not occluded. As described above regarding the determination of the controller's position and orientation, the timing of the blinks can be used here to determine which light is occluded, thereby determining the position of the corresponding finger or palm. If a light source 1706 adjacent to button 1705 has a reduced intensity or no intensity during the interval, the light source adjacent to the button is known to be "on". It is then determined that the light source is occluded. Similarly, if the detected intensity of a light known to be "on" changes, an event may be generated, from which it may be determined that the user's finger or hand has moved and the button is no longer pressed. The occlusion of one or more light sources adjacent to one or more buttons can be correlated to the position of the finger or palm based on their positions. For example, but not limited to, a light source located near the user's palm when the controller 1701 is held can be used to determine the position of the user's hand 1704.
[0073] Light sources 1706 can be positioned around each button 1705, and the button configuration and design of the controller can be used to determine the finger position. For example, but not limited to, the controller 1701 may be designed so that each of the user's fingers is positioned near a button 1705 when held. The occlusion pattern of the light sources determined from the DVS event may then be used to determine when the user's finger is airborne over an inactive button, and also when the user's finger has passed over the button. This can help provide the user with further interaction options, such as half-pressing or under-pressing a button, or having other button options. Multiple light sources surrounding each button allows for more refinement in determining the position of the user's finger or palm. For example, but not limited to, in some implementations, more than 10 light sources may surround each button, and in other implementations, a single light source may illuminate the button through a translucent diffuser, and interruptions in the diffuse light profile can be used to determine the finger position.
[0074] In some implementations, the button 1705 itself can also be a light source. One or more buttons 1705 can be turned on and off at different intervals than one or more adjacent light sources 1706 or other light-tracking light sources 1703. Alternatively, the light source of the button 1705 may have a different wavelength or polarization than the adjacent one or more light sources 1706 or other light-tracking light sources 1703.
[0075] Furthermore, power-saving modes for light sources can be enabled using buttons and tracking. For example, if it is determined that a button is pressed, one or more light sources adjacent to that button may be dimmed or turned off. Additionally, if it is determined that the controller 1701 is outside the field of view of the DVS 1702, one or more lights adjacent to the button may be dimmed or turned off. In some implementations, data from the IMU can be used to determine whether the controller is being held by the user, and for example, if no changes in IMU data such as acceleration or angular velocity have been detected for a threshold period (but are not limited to these), the light sources may be dimmed or turned off. When changes in IMU data are detected, the light sources can be turned back on.
[0076] In alternative implementations, finger tracking can be performed without using one or more light sources. A machine learning model can be trained using a machine learning algorithm to detect the finger's position from events generated from changes in ambient light due to finger movement. The machine learning model may be a general machine learning model such as the CNN, RNN, or DNN mentioned above. In some implementations, specialized machine learning models, such as, but not limited to, spiking (or spark) neural networks (SNNs), can be trained with specialized machine learning algorithms. SNNs mimic biological NNs by having an activation threshold and weights, which are adjusted according to the relative spike time within an interval, also known as STDP (Spike-timing-dependent-plasticity). When the activation threshold is reached, the SNN is said to spike and send its weights to the next layer. SNNs can be trained via STDP and supervised or unsupervised learning techniques. Further information regarding SNNs can be found in Tavanaei, Amirhossein et al., "Deep Learning in Spiking Neural Networks," Neural Networks (2018) arXiv:1804.08150, the contents of which are incorporated herein by reference for all purposes.
[0077] Alternatively, an HDR (high dynamic range) image may be constructed using events aggregated from ambient data. A machine learning model is trained to recognize the position and orientation of the hand or controller from the HDR image. The trained machine learning model can be applied to the HDR image generated from the events to determine the position and orientation of the hand / fingers or controller. The machine learning model may be a general machine learning model trained with supervised learning techniques, as described in the section on training general neural networks.
[0078] Target tracking Aspects of this disclosure may be applied to target tracking. Generally, target tracking image analysis determines the direction of gaze from an image by utilizing the specific properties of how light is reflected from the eye. For example, by analyzing an image, the position of the eye can be identified based on corneal reflection in the image data, and by further analyzing the image, the direction of gaze can be determined based on the relative position of the pupil in the image.
[0079] Two common eye-tracking techniques that determine gaze direction based on pupil position are known as the bright pupil method and the dark pupil method. The bright pupil method involves illuminating the eye with a light source that substantially coincides with the optical axis of the DVS, which reflects the emitted light off the retina and back through the pupil to the DVS. The pupil appears in the image as a identifiable bright spot at the pupil's position, similar to the red-eye effect that occurs in images during conventional flash photography. In this eye-tracking method, if the contrast between the pupil and iris is insufficient, the bright reflection from the pupil itself helps the system locate the pupil.
[0080] The dark pupil method involves illuminating the eye with a light source substantially offset from the optical axis of the DVS. This causes the light directed through the pupil to reflect away from the DVS's optical axis, resulting in a dark spot at the pupil's location that is identifiable to the event. In another dark pupil method system, an infrared light source and camera directed at the eye can observe corneal reflections. Such DVS-based systems track the position of the pupil and corneal reflections, thereby improving accuracy by obtaining parallax due to the different depths of reflection.
[0081] Figure 18A shows an example of a dark pupil eye-tracking system 1800 that may be used in the context of this disclosure. The eye-tracking system tracks the orientation of the user's eye E to a display screen 1801 on which a visible image is presented. While a display screen is used in the exemplary system of Figure 18A, certain alternative embodiments may utilize an image projection system that can project an image directly onto the user's eye. In these embodiments, the user's eye E is tracked relative to the image projected onto the user's eye. In the example of Figure 18A, the eye E collects light from the screen 1801 via a variable iris I; a lens L projects the image onto the retina R. The opening of the iris is known as the pupil. Muscles control the rotation of the eye E in response to nerve impulses from the brain. The upper and lower eyelid muscles ULM, LLM control the upper and lower eyelids UL, LL, respectively, in response to other nerve impulses.
[0082] Photosensitive cells on the retina (R) generate electrical impulses that are sent to the user's brain (not shown) via the optic nerve (ON). The visual cortex of the brain interprets these impulses. Not all parts of the retina (R) are equally photosensitive. Specifically, photosensitive cells are concentrated in a region known as the fovea.
[0083] The illustrated image tracking system includes one or more infrared light sources 1802, for example, light-emitting diodes (LEDs) that direct invisible light (e.g., infrared light) towards the eye E. Some of the invisible light is reflected by the cornea C of the eye, and some is reflected by the iris. The reflected invisible light is directed by a wavelength-selective mirror 1806 to an infrared-sensitive DVS 1804. The mirror transmits visible light from the screen 1801, but reflects invisible light reflected from the eye.
[0084] DVS1804 generates eye E events that can be analyzed to determine the gaze direction GD from the relative position of the pupil. These events can be generated by processor 1805. DVS1804 is advantageous in this implementation because its very fast event update rate provides near real-time information about changes in the user's gaze.
[0085] As shown in Figure 18B, by analyzing event 1811, which represents the user's head H, the gaze direction GD can be determined from the relative position of the pupil. For example, the analysis may determine the two-dimensional offset of the pupil P from the center of the eye E in the image. The position of the pupil relative to the center can be converted to the gaze direction relative to the screen 1801 by a simple geometric calculation of a three-dimensional vector based on the known size and shape of the eyeball. The determined gaze direction GD can indicate the rotation and acceleration of the eye E as it moves relative to the screen 1801.
[0086] As can be seen in Figure 18B, the event may include reflections of invisible light 1807 and 1808 from the cornea C and lens L, respectively. Due to the difference in depth between the cornea and lens, the parallax and refractive index between the reflections can be used to improve the accuracy of determining the line of sight direction GD. An example of this type of target tracking system is the dual Purkinje tracker, where the corneal reflection is the first Purkinje image and the lens reflection is the fourth Purkinje image. If the user is wearing glasses, there may also be a reflection 1808 from the user's glasses 1809.
[0087] The performance of an eye-tracking system depends on numerous factors, including the light source (IR, visible light, etc.) and DVS placement, whether the user is wearing glasses or contact lenses, the headset optics, the tracking system latency, the speed of eye movements, the shape of the eye (which can change throughout the day or as a result of movement), the state of the eye, such as amblyopia, gaze stability, fixation on moving objects, the scene presented to the user, and the user's head movements. The DVS reduces the output of extraneous information to the processor, providing a very fast update rate for events. This allows for faster processing and quicker determination of the eye-tracking state and error parameters.
[0088] Error parameters that can be determined from eye-tracking data may include, but are not limited to, rotational speed and prediction errors, fixation errors, confidence intervals for current and / or future gaze positions, and smoothness tracking errors. State information regarding the user's gaze includes the individual states of the user's eyes and / or gaze. Accordingly, exemplary state parameters that can be determined from eye-tracking data may include, but are not limited to, blink metrics, saccade metrics, depth-of-field response, color blindness, gaze stability, and eye movements as precursors to head movements.
[0089] In certain implementations, the eye-tracking error parameter may include a confidence interval for the current gaze position. The confidence interval can be determined by checking whether the user's eye rotation speed and acceleration have changed from the last position. In alternative embodiments, the eye-tracking error and / or state parameter may include a prediction of the future gaze position. The future gaze position can be determined by checking the eye rotation speed and acceleration and extrapolating the possible future positions of the user's eyes. Generally speaking, the DVS update rate of an eye-tracking system is very high, so for users with high rotation speed and acceleration values, the error between the determined future position and the actual future position can be small, and this small error can be significantly smaller than in existing camera-based systems.
[0090] In yet another alternative implementation, the eye-tracking error parameter may include a measurement of eye velocity, such as rotational speed. In a particular alternative embodiment, the determined eye-tracking state parameter includes measuring a metric for the user's blink. During a normal blink, a period of typically 150 milliseconds (ms) elapses, during which the user's vision is not focused on the presented image. Therefore, depending on the display device's frame rate, the user's vision may be out of focus on the presented image for up to 20-30 frames. However, once the blink ends, the user's gaze direction may not correspond to the last measured gaze direction, as determined by the acquired eye-tracking data. Therefore, the user's gaze metric may be determined from the acquired eye-tracking data. These metrics may include, but are not limited to, the predicted end time of the user's blink, as well as the measured start and end times.
[0091] In additional alternative implementations, the determined eye-tracking state parameters include measuring the user's saccade metrics. During a typical saccade, a period of 20–200 ms usually elapses, during which the user's vision is not focused on the presented image. Therefore, depending on the display device's frame rate, the user's vision may be out of focus on the presented image for up to 40 frames. However, due to the nature of saccades, once the saccade ends, the user's gaze direction shifts to another area of interest. Therefore, eye-tracking data can be used to establish metrics for the user's saccade based on the actual or predicted time elapsed during the saccade. These metrics may include, but are not limited to, the predicted end time of the user's saccade, as well as the measured start and end times.
[0092] In certain alternative implementations, the determined gaze-tracking state parameters include determining the transition in the user's gaze direction between regions of interest as a result of changes in depth of field between presented images. This is because providing transitions between regions of interest of presented images would cause the user to experience saccades.
[0093] In additional alternative implementations, the determined eye-tracking state parameters can be adapted to color blindness. For example, regions of interest may exist in the image presented to the user in such a way that these regions are not perceived by users with a particular form of color blindness. The resulting eye-tracking data can be used to determine, for example, whether the user's gaze identified or responded to a region of interest as a result of changes in the user's gaze direction. Thus, it can be determined whether the user is color-blind to a particular color or spectrum as an eye-tracking error parameter.
[0094] In certain alternative implementations, the determined eye-tracking state parameters include a measure of the user's gaze stability. Determining gaze stability can be done by measuring the microsaccadic radius of the user's eye; the smaller the fixation overshoot and undershoot, the more stable the user's gaze is.
[0095] In additional alternative implementations, the determined eye-tracking error and / or state parameters include the user's ability to fixate on moving objects. These parameters may include a measure of the user's ability to perform smooth tracking, and the maximum object tracking speed of the eyeball. Typically, users with superior smooth tracking ability experience less eye movement jitter.
[0096] In certain alternative implementations, the determined gaze tracking error and / or state parameters include the determination of eye movements as precursors to head movements. The offset between head and eye orientation can affect certain error and / or state parameters, for example, in smooth tracking or fixation, as described above.
[0097] Further information regarding eye tracking and error parameter determination can be found in U.S. Patent No. 10,192,528, which is incorporated herein by reference for all purposes.
[0098] system Figure 19 is a block system diagram of a DVS tracking system according to an aspect of the present disclosure. Without limitation, and as an example, according to an aspect of the present disclosure, system 1900 may be an embedded system, a mobile phone, a personal computer, a tablet computer, a portable game device, a workstation, a game console, and the like.
[0099] System 1900 generally includes a central processing unit (CPU) 1903 and memory 1904. System 1900 may also include well-known support functions 1906 that can communicate with other components of the system, for example, via a data bus 1905. Such support functions may include, but are not limited to, input / output (I / O) elements 1907, a power supply (P / S) 1911, a clock (CLK) 1912, and a cache 1913.
[0100] System 1900 may include a display device 1931 for presenting rendered graphics to the user. In alternative implementations, the display device is a separate component that functions in conjunction with System 1900. The display device 1931 may take the form of a flat panel display, a head-mounted display (HMD), a cathode ray tube (CRT) screen, a projector, or other device capable of displaying visible text, numbers, graphic symbols, or images.
[0101] Here, the display device 1931 is coupled with the DVS 1901A, and the controller 1902 includes two or more light sources 1932A, which may be in any of the configurations described herein. In an alternative implementation, the DVS may be coupled to the game controller, and the display device may instead include two or more light sources. In yet another alternative implementation, the DVS is a separate unit uncoupled from either the display device or the controller, in which case both the controller and the display device may include two or more light sources for tracking.
[0102] For example, in some implementations where the display device is part of a head-mounted display (HMD), such an HMD may include an inertial measurement unit (IMU) such as an accelerometer or gyroscope. Also, as discussed above in this specification, such an HMD may include light sources 1932B, which may be tracked using a DVS separate from the display device 1901 and coupled to the CPU 1903. For example, a separate DVS 1901B may be attached to the controller 1902.
[0103] In some implementations, the DVS1901A or DVS1901B may be part of a hybrid sensor, as described above with respect to, for example, Figures 11A, 11B, 11C, or 11D. Such a hybrid sensor may include a depth sensor, such as a DTOF sensor, in which case the hybrid sensor may include an illumination unit (not shown).
[0104] Furthermore, if the display device 1931 is part of the HMD, a safety shutter 1933 may optionally be fitted to the device, which can be operably coupled to a processor such as the CPU 1903 and can operate as described above with respect to Figures 16A, 16B, 16C, 16D, or 16E. Alternatively, the safety shutter may be controlled by a separate processor mounted on the HMD.
[0105] Furthermore, system 1900 includes a mass storage device 1915, such as a disk drive, CD-ROM drive, flash memory, solid-state drive (SSD), tape drive, etc., to provide non-volatile storage for programs and / or data. System 1900 may also optionally include a user interface unit 1916 to facilitate interaction between system 1900 and the user. The user interface 1916 may include a keyboard, mouse, joystick, light pen, or other devices that can be used with a graphical user interface (GUI). System 1900 may also include a network interface 1914 to enable the device to communicate with other devices through a network 1920. The network 1920 may be, for example, a local area network (LAN), a wide area network such as the Internet, a personal area network such as a Bluetooth® network, or other types of networks. These components may be implemented in hardware, software, or firmware, or any combination of two or more of these.
[0106] Each CPU 1903 may contain one or more processor cores, for example, a single core, two cores, four cores, eight cores, or more. In some implementations, CPU 1903 may contain a GPU core or multiple cores of the same APU (Accelerated Processing Unit). Memory 1904 may be in the form of an integrated circuit providing addressable memory, such as random access memory (RAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), etc. Main memory 1904 may contain application data 1923 used by the processor 1903 during processing. Main memory 1904 may also contain event data 1909 received from DVS 1901. As illustrated in Figure 9, a trained neural network (NN) 1910 can be loaded into memory 1904 to determine position and orientation data. Furthermore, memory 1904 may contain machine learning algorithms 1921 for training or tuning NN 1910. The database 1922 can be contained in memory 1904. The database may contain information such as the light source configuration and the predetermined blinking intervals of one or more light sources. The memory may also contain outputs from an IMU coupled to the controller 1902 or display device 1931. In some implementations, memory 1904 may contain outputs from one or more light sources, such as timestamps indicating when each light source is on or off.
[0107] According to aspects of this disclosure, the processor 1903 can perform methods for determining the position and orientation of a controller or user, as described in Figures 8 and 9, and these methods can be loaded into memory 1904 as applications 1923. As a result of performing the methods described in Figures 8 and 9, and further with respect to Figure 15, the processor can generate one or more orientations and configurations of a controller, a headset, a user's body or appendages, the ground, or a landmark. These positions and orientations can be stored in a database 1922 and used for sequential iterations of the methods in Figures 8 and 9. In some implementations, the processor 1903 may utilize such positions and / or orientations to machine learning algorithms trained to perform SLAM (simultaneous localization and mapping).
[0108] The mass storage 1915 may include an application or program 1917 that is loaded into main memory 1904 when processing is initiated by application 1923. Furthermore, the mass storage 1915 may include data 1918 used by the processor during the processing of application 1923, NN 1910, machine learning algorithm 1921, and while filling the database 1922.
[0109] As used herein and as generally understood by those skilled in the art, an application-specific integrated circuit (ASIC) is an integrated circuit customized for a specific application, rather than for general-purpose use.
[0110] As used herein and as generally understood by those skilled in the art, a field-programmable gate array (FPGA) is an integrated circuit designed to be configured by the customer or designer after manufacture—and thus "field-programmable." FPGA configurations are generally specified using a hardware description language (HDL) similar to that used for ASICs.
[0111] As used herein and as generally understood by those skilled in the art, a system on a chip, or system-on-a-chip (SoC or SOC), is an integrated circuit (IC) that integrates all the components of a computer or other electronic system onto a single chip. It can include digital, analog, mixed-signal, and often radio frequency functions—all on a single chip substrate. Typical applications are in the field of embedded systems.
[0112] A typical SoC includes the following hardware components: One or more processor cores (e.g., a microcontroller, microprocessor, or digital signal processor (DSP) core). Memory blocks, such as read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), and flash memory. Timing source, e.g., oscillator or phase-locked loop. Peripheral devices, such as counter timers, real-time timers, or power-on reset generators. External interfaces, such as industry standards including Universal Serial Bus (USB), FireWire®, Ethernet, USART (universal asynchronous receiver / transmitter), and Serial Peripheral Interface (SPI) bus. Analog interface (including analog-to-digital converters (ADCs) and digital-to-analog converters (DACs)). Voltage regulator and power management circuit.
[0113] These components are connected by either a proprietary bus or an industry-standard bus. The Direct Memory Access (DMA) controller improves the SoC's data throughput by routing data directly between the external interface and memory, bypassing the processor core.
[0114] A typical SoC includes both the hardware components mentioned above, as well as the processor core(s), peripherals, and executable instructions (e.g., software or firmware) that control the interfaces.
[0115] Aspects of this disclosure provide image-based tracking with a higher sample rate than is possible with conventional image-based tracking systems, thereby improving tracking fidelity. Further benefits include reduced costs, reduced weight, reduced irrelevant data generation, and reduced processing requirements when using DVS-based tracking systems. These benefits enable improvements in virtual reality (VR) and augmented reality (AR) systems, among other applications.
[0116] The above is a complete description of preferred embodiments of the present invention, but various alternatives, modifications, and equivalents are possible. Therefore, the scope of the present invention should not be determined by reference to the foregoing, but rather by reference to the appended claims, which encompass the full range of their equivalents. Any feature described herein, whether preferred or not, may be combined with any other feature described herein, whether preferred or not. In the following claims, the indefinite article “A” or “An” refers to one or more items following the article, unless explicitly stated otherwise. The appended claims should not be construed as including means-plus-function limitations unless such limitations are explicitly stated in a given claim using the phrase “means for”.
Claims
1. A tracking method, The determination of the position and orientation of the controller is made from known configurations of two or more light sources relating to each other and to the controller body, and from output signals from a dynamic visual sensor (DVS) generated in response to changes in the light output from the two or more light sources, wherein the output signals include information corresponding to the time of two or more events at two or more corresponding photosensitive elements in an array of two or more photosensitive elements in the DVS, and the position of the two or more corresponding photosensitive elements in the array. The position and orientation of one or more objects are determined from signals generated by two or more photosensitive elements due to other light reaching the two or more photosensitive elements, Includes, A tracking method in which the one or more objects include one or more objects other than the controller.
2. The tracking method according to claim 1, wherein the one or more objects include the controller.
3. The tracking method according to claim 1, wherein determining the position and orientation of one or more objects from signals generated by two or more photosensitive elements due to other light reaching the two or more photosensitive elements includes determining an environment map and the state of the one or more objects.
4. A tracking method, The determination of the position and orientation of the controller is made from known configurations of two or more light sources relating to each other and to the controller body, and from output signals from a dynamic visual sensor (DVS) generated in response to changes in the light output from the two or more light sources, wherein the output signals include information corresponding to the time of two or more events at two or more corresponding photosensitive elements in an array of two or more photosensitive elements in the DVS, and the position of the two or more corresponding photosensitive elements in the array. The position and orientation of one or more objects are determined from signals generated by two or more photosensitive elements due to other light reaching the two or more photosensitive elements, Includes, A tracking method further comprising triggering an optical safety shutter on a headset in response to the determined position of one or more objects.
5. The tracking method according to claim 4, wherein the optical safety shutter includes an opaque blind operably coupled to the headset, the opaque blind being configured to selectively allow light to pass into the eyes of the user wearing the headset.
6. The tracking method according to claim 5, wherein triggering the optical safety shutter includes triggering the optical safety shutter in response to the determined position of one or more objects to change from an opaque state to a light-transmitting state.
7. The tracking method according to claim 6, wherein the optical safety shutter includes a liquid crystal shutter.
8. The tracking method according to claim 1, wherein determining the position and orientation of one or more objects from signals generated by the two or more photosensitive elements as a result of the other light reaching the two or more photosensitive elements includes determining an environment map and the state of the one or more objects from signals generated by the two or more photosensitive elements in a DVS located on the controller.
9. The tracking method according to claim 1, wherein determining the position and orientation of one or more objects includes using a machine learning algorithm to match the position and orientation of the one or more objects to the signals generated by the two or more photosensitive elements.
10. A tracking method, The determination of the position and orientation of the controller is made from known configurations of two or more light sources relating to each other and to the controller body, and from output signals from a dynamic visual sensor (DVS) generated in response to changes in the light output from the two or more light sources, wherein the output signals include information corresponding to the time of two or more events at two or more corresponding photosensitive elements in an array of two or more photosensitive elements in the DVS, and the position of the two or more corresponding photosensitive elements in the array. The position and orientation of one or more objects are determined from signals generated by two or more photosensitive elements due to other light reaching the two or more photosensitive elements, Receiving information from an inertial measuring unit (IMU) that corresponds to changes in angular velocity, specific force, specific acceleration, or magnetic flux, Includes, A tracking method for determining the position and orientation of the controller, further comprising using information corresponding to changes in angular velocity or specific force or specific acceleration or magnetic flux when determining the position and orientation of the controller.
11. The tracking method according to claim 10, further comprising using a Kalman filter to merge information on changes in angular velocity or changes in specific force / acceleration or magnetic flux with the two or more events.
12. The tracking method according to claim 1, wherein determining the position and orientation of the controller includes determining the position of at least one of the two or more light sources in a known configuration using changes in the light output and known patterns of changes in the light output.
13. The tracking method according to claim 1, further comprising notifying the user of the location of the one or more objects based on the determined location and orientation of the one or more objects.
14. A tracking method, The determination of the position and orientation of the controller is made from known configurations of two or more light sources relating to each other and to the controller body, and from output signals from a dynamic visual sensor (DVS) generated in response to changes in the light output from the two or more light sources, wherein the output signals include information corresponding to the time of two or more events at two or more corresponding photosensitive elements in an array of two or more photosensitive elements in the DVS, and the position of the two or more corresponding photosensitive elements in the array. The position and orientation of one or more objects are determined from signals generated by two or more photosensitive elements due to other light reaching the two or more photosensitive elements, In response to the determined position of the controller, the optical safety shutter on the headset is triggered. Tracking methods, including those mentioned above.
15. The tracking method according to claim 14, wherein the optical safety shutter includes an opaque blind operably coupled to the headset, the opaque blind being configured to selectively allow light to pass into the eyes of the user wearing the headset.
16. The tracking method according to claim 14, wherein triggering the optical safety shutter includes triggering the optical safety shutter in response to the determined position of one or more objects to transition from an opaque state to a light-transmitting state.
17. The tracking method according to claim 16, wherein the optical safety shutter includes a liquid crystal shutter.
Citation Information
Patent Citations
Device and method for displaying composite sense of reality, recording medium and computer program
JP2003256876A
Simulation system, program and controller
JP2018126340A
Position detection device, position detection system, and position detection method
US20160328089A1
Tracking of position and orientation of objects in virtual reality systems
US20180314346A1
Systems and methods for detecting objects within the boundary of a defined space while in artificial reality
US20210319220A1