Tracking an apparatus for human-machine-interactions

The IMU-LED-feature tightly-coupled fusion algorithm addresses the instability of ringless CV-based controllers by fusing IMU, LED, and feature tracking measurements, improving tracking precision and robustness for enhanced user experience.

WO2025179955A1PCT designated stage Publication Date: 2025-09-04QUALCOMM INC +4
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/131709
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-27
Filing Date
2024-11-13
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Conventional ring-type controllers for human-machine interfaces are bulky, inconvenient, and prone to tracking failures due to LED tracking rings being blocked by hands, while ringless CV-based controllers face challenges with unstable feature relationships and occlusions, leading to tracking instability.

Method used

A novel inertial measurement unit (IMU)-LED-feature tightly-coupled fusion algorithm that simultaneously fuses IMU, LED tracking, and feature tracking measurements, using a sliding window and Mahalanobis distance test to maintain stable relationships and reject outliers, enhancing tracking robustness.

Benefits of technology

Improves tracking precision and robustness of ringless CV-based controllers, ensuring consistent performance even under occlusion and varying hand gestures, thereby enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024131709_04092025_PF_FP_ABST
    Figure CN2024131709_04092025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and techniques are described for tracking a human-machine-interface (HMI) device. For instance, a method may include determining light-emitting diode (LED) -tracking measurements based on images of the HMI device, wherein the HMI device comprises LEDs; determining feature-tracking measurements based on the images of the HMI device; determining movement data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; and determining a pose of the HMI device by fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data.
Need to check novelty before this filing date? Find Prior Art

Description

TRACKING AN APPARATUS FOR HUMAN-MACHINE-INTERACTIONS

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of PCT International Application No. PCT / CN2024 / 078685, filed February 27, 2024, titled “TRACKING APPARATUS FOR HUMAN-MACHINE INTERACTIONS, ” which is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0003] The present disclosure generally relates to a human-machine interactions. More particularly, it relates to tracking a controller apparatus for a human-machine-interface (HMI) device.BACKGROUND

[0004] A handheld controller is an example of a human-machine-interface (HMI) device. Such a handheld controller may allow a user to interact with a machine by pressing or activating buttons on the handheld controller and / or through a position and / or motion of the handheld controller. Handheld controllers can be used with extended-reality (XR) systems.SUMMARY

[0005] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.

[0006] Systems and techniques are described for tracking a human-machine-interface (HMI) device. According to at least one example, a method is provided for tracking an HMI device. The method includes: determining light-emitting diode (LED) -tracking measurements based on images of the HMI device, wherein the HMI device comprises LEDs; determining feature-tracking measurements based on the images of the HMI device; determining movement data of the HMI  device based on an inertial measurement unit (IMU) of the HMI device; and determining a pose of the HMI device by fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data.

[0007] In another example, an apparatus for tracking an HMI device is provided. The apparatus includes: at least one memory and at least one processor (e.g., configured in circuitry) coupled to the at least one memory. The at least one processor configured to: determine light-emitting diode (LED) -tracking measurements based on images of the HMI device, wherein the HMI device comprises LEDs; determine feature-tracking measurements based on the images of the HMI device; determine movement data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; and determine a pose of the HMI device by fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data.

[0008] In another example, a non-transitory computer-readable medium is provided that has stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: determine light-emitting diode (LED) -tracking measurements based on images of the HMI device, wherein the HMI device comprises LEDs; determine feature-tracking measurements based on the images of the HMI device; determine movement data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; and determine a pose of the HMI device by fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data.

[0009] In another example, an apparatus for tracking an HMI device. The apparatus includes: means for extracting image features from images of the HMI device captured by a camera of a head-mounted device (HMD) , the image features related to at least one of hands or arms of a user holding the HMI device; means for determining light-emitting diode (LED) -tracking measurements based on images of the HMI device, wherein the HMI device comprises LEDs; means for determining feature-tracking measurements based on the images of the HMI device; means for determining movement data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; and means for determining a pose of the HMI device by fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data.

[0010] In some aspects, one or more of the apparatuses described herein is, can be part of, or can include an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device) , a mobile device (e.g., a mobile telephone or so-called “smart phone” , a tablet computer, or other type of mobile device) , a smart or connected device (e.g., an Internet-of-Things (IoT) device) , a wearable device, a personal computer, a laptop computer, a video server, a television (e.g., a network-connected television) , a robotics device or system, a vehicle (or a computing device or system of a vehicle) , or other device. In some aspects, each apparatus can include an image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each apparatus can include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each apparatus can include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each apparatus can include one or more sensors. In some cases, the one or more sensors can be used for determining a location of the apparatuses, a state of the apparatuses (e.g., a tracking state, an operating state, a temperature, a humidity level, and / or other state) , and / or for other purposes.

[0011] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.

[0012] The foregoing, together with other features and aspects, will become more apparent upon referring to the following specification, claims, and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Illustrative examples of the present application are described in detail below with reference to the following figures:

[0014] FIG. 1 is a diagram illustrating a conventional hand-held controller including a number of light emitting diodes arranged on a tracking ring protruding from a handle of handheld controller;

[0015] FIG. 2A is a diagram illustrating an example handheld controller including a number of light sources arranged on a body of handheld controller, according to various aspects of the present disclosure;

[0016] FIG. 2B is a diagram illustrating handheld controller being held by a hand of a user, according to various aspects of the present disclosure;

[0017] FIG. 3A is a diagram illustrating an example system including a handheld controller and a tracking system, according to various aspects of the present disclosure;

[0018] FIG. 3B is a block diagram illustrating example elements of handheld controller and tracking system of system, according to various aspects of the present disclosure;

[0019] FIG. 4 is a hybrid block diagram  / flow chart illustrating operation of an example process 400 that may be used to track a controller, according to various aspects of the present disclosure;

[0020] FIG. 5A is a flowchart illustrating an example process for performing a motion-data-feature tightly-coupled fusion, according to various aspects of the present disclosure;

[0021] FIG. 5B is a flowchart illustrating an example process for performing a motion-data-feature tightly-coupled fusion, according to various aspects of the present disclosure;

[0022] FIG. 6A is a flow diagram illustrating another example process for tracking a human-machine-interface (HMI) device, in accordance with aspects of the present disclosure;

[0023] FIG. 6B is a flow diagram illustrating another example process for tracking a HMI device, in accordance with aspects of the present disclosure;

[0024] FIG. 7 is a block diagram illustrating an example of a deep learning neural network that can be used to perform various tasks, according to some aspects of the disclosed technology;

[0025] FIG. 8 is a block diagram illustrating an example of a convolutional neural network (CNN) , according to various aspects of the present disclosure; and

[0026] FIG. 9 is a block diagram illustrating an example computing-device architecture of an example computing device which can implement the various techniques described herein.DETAILED DESCRIPTION

[0027] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects of the application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and description are not intended to be restrictive.

[0028] The ensuing description provides example aspects only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the exemplary aspects will provide those skilled in the art with an enabling description for implementing an exemplary aspect. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the application as set forth in the appended claims.

[0029] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration. ” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as preferred or advantageous over other aspects. Likewise, the term “aspects of the disclosure” does not require that all aspects of the disclosure include the discussed feature, advantage, or mode of operation.

[0030] Object detection and tracking can be used to identify an object and track the object over time. For example, an image of an object can be obtained and object detection can be performed on the image to detect the object in the image. In some cases, the detected object can be classified into a category of object. Further, a bounding box can be generated to identify a position of the object in the image. Various types of systems can be used for object detection, including neural network-based object detectors. The position of the object in images can be tracked through a series of images. In some cases, tracking the object can include determining a pose of the object relative to a camera which captured the image of the object and / or relative to prior positions of the object. In the present disclosure, the term “pose” may refer to a position and orientation. Poses may be determined according to six degrees of freedom including three translational degrees of  freedom (e.g, . x, y, and z dimensions) and three rotational degrees of freedom (e.g., roll, pitch, and yaw) .

[0031] A handheld controller can be used as a human-machine-interface (HMI) device. For example, an object-tracking system may detect and track the handheld controller and a user can interface with a machine by moving and / or rotating the handheld controller. For example, the machine may receive inputs based on how the user moves and / or rotates the handheld controller. Additionally, the controller may include one or more buttons that the user may press or activate. Indications of the buttons being pressed or activated may be transmitted to the machine.

[0032] One example of a system that may be interacted with by a handheld controller is an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device) . An example of an XR device is a head-mounted display (HMD) device (e.g., in the form of a headset, glasses, etc. ) that can be worn by the user. The XR device may include one or more cameras that may capture images of a handheld controller. The user may move and / or rotate the handheld controller. The XR device may include a tracking system that can, using the cameras, track the handheld controller and interpret inputs based on how the user moves and / or rotates the handheld controller. Another example includes a gaming system with a camera mounted above a display.

[0033] Ring-type controllers are a class of computer vision (CV) -based controllers. Ringless CV-based controllers are becoming more and more common because ringless CV-based controllers are more compact and portable. But ringless CV-based controllers are more difficult to track (e.g., by a CV-based object detection and / or tracking algorithm) .

[0034] Systems, apparatuses, methods (also referred to as processes) , and computer-readable media (collectively referred to herein as “systems and techniques” ) are described herein for improving tracking of an HMI device. For example, the systems and techniques described herein may include a new multi-sensor fusion algorithm for ringless CV-based controllers. The new fusion algorithm can improve the tracking precision and robustness and improve the user experience as well.

[0035] Various aspects of the application will be described with respect to the figures below.

[0036] FIG. 1 is a diagram illustrating a conventional hand-held controller 100 including a number of light-emitting diodes 106 (LEDs 106) arranged on a tracking ring 104 protruding from a handle 102 of handheld controller 100. Handheld controller 100 includes a button 108 as an example of buttons that may be present on handle 102 of handheld controller 100. A conventional tracking system may rely on images of LEDs 106 on tracking ring 104 to localize and track handheld controller 100. Tracking ring 104 may be bulky, not easy to carry, and inconvenient to use.

[0037] To improve the performance of conventional tracking systems, tracking ring 104 protrudes from the outside of handheld controller 100. Because of the requirements of the tracking algorithm, light-emitting diode (LED) tracking rings (e.g., tracking ring 104) must be large enough for the tracking algorithm to accurately track. A large tracking ring affects portability, such as it may be more difficult to put the pocket in a pocket or back to be carried around.

[0038] Further, in order to avoid tracking failure caused by the LED tracking light ring being blocked by the hand (e.g., a hand holding handle 102) , the structural design of the LED tracking ring is usually protruding outside where a hand would hold the handle. But such a protruding structure is inconvenient to use, and LED tracking ring can easily hit other objects when in use. For example, when both hands of a person are holding respective controllers, the hands cannot touch each other due to the obstruction of the protruding part of the LED tracking ring.

[0039] Ring-type controllers (such as handheld controller 100) have following disadvantages: uncompact and unportable, may encounter obstacles during use and easily damaged. Additionally, there are certain restrictions on the cross-interaction between two controllers when two ring-type controllers are used at the same time.

[0040] To overcome the limitations of ring-type controllers, ringless CV-based controllers are becoming more and more popular. Ringless CV-based controllers do not include LED tracking rings. The tracking LEDs placed on the ring of ring-type controllers are placed on the controller’s panel and body. Ringless CV-based controllers are more compact and portable that ring-type controllers. Additionally, ringless CV-based controllers improve the experience of a user when the user uses two controllers at the same time.

[0041] FIG. 2A is a diagram illustrating an example handheld controller 200 including a number of light sources 202 arranged on a body 204 of handheld controller 200, according to various aspects of the present disclosure. In contrast to handheld controller 100 of FIG. 1, light sources 202 are not arranged on a tracking ring 104 protruding from a handle 102. Rather, light sources 202 are arranged on body 204 (e.g., not radially protruding from handle 206 of body 204) . Handheld controller 200 includes a button 208 as an example of buttons that may be present on body 204 of handheld controller 200. FIG. 2B is a diagram illustrating handheld controller 200 being held by a hand 210 of a user, according to various aspects of the present disclosure. Handheld controller 200 may be sized such that at least some of light sources 202 are visible when body 204 is held by hand 210.

[0042] Because the tracking LEDs are placed on controller’s panel and body, the LEDs are more likely to be blocked by fingers and palm when the controller is held in a hand. In extreme cases, all the LEDs may be blocked and tracking may fail.

[0043] Some tracking methods include hand-joint tracking, visual-features tracking on controllers, on hands, and even on arms. Hand-joint tracking, visual-features tracking (including tracking features of controllers, hands, and / or arms) may be referred to as “tracking methods, ” “aided visual feature tracking, ” or “feature tracking” to distinguish from “LED tracking. ” Where “LED tracking” may refer to tracking LEDs in images.

[0044] The challenges to tracking ringless controllers include: LEDs are more likely to blocked, and more easily to lose tracking, need to find other feature tracking methods to aid the tracking, and the relationship between the controller and the features of hands or arms is not stable, and may be changing with time and different holding gestures. That is to say, their connection is not rigid, which is very challenging for controller fusion algorithm.

[0045] This disclosure discloses a novel inertial measurement unit (IMU) -LED-feature tightly-coupled fusion algorithm to address at least some of the challenges facing ringless controllers. As previously mentioned, besides IMU and LED tracking measurements, feature tracking measurements may be fused in a ringless fusion algorithm. But the relationship between the controller and the features of hands or arms is not stable. So a novel IMU-LED-feature tightly-coupled fusion algorithm is disclosed to solve this problem. In the present disclosure, the term  “IMU” may refer to an inertial measurement unit, data from an inertial measurement unit (e.g., motion data) , and / or other motion data.

[0046] The fusion algorithm has at least the following aspects: a tightly-coupled fusion framework that can simultaneously fuse IMU, LED tracking measurements, and feature tracking measurements, a sliding window to the fusion state vector to estimate the relationships between the controller IMU and the tracking features at the same time, and using a distance threshold (e.g., the Mahalanobis distance test) method to detect the changes of relationships between the controller IMU and the tracking features, and then reject outliers.

[0047] FIG. 3A is a diagram illustrating an example system 300 including a handheld controller 302 and a tracking system 320, according to various aspects of the present disclosure. FIG. 3B is a block diagram illustrating example elements of handheld controller 302 and tracking system 320 of system 300, according to various aspects of the present disclosure. In some cases, the tracking system 320 can be part of or include an XR device (e.g., an HMD) .

[0048] Handheld controller 302 includes light-emitting diodes (LEDs) 306 on body 304 of handheld controller 302. LEDs 306 may emit visible light (of any color) , near infrared light, and / or infrared light. LEDs 306 may include groups of LEDs (e.g., each group including a red LED, a green LED and a blue LED such that each group of LEDs may vary wavelengths of light emitted) . LEDs 306 may be included on handheld controller 302 to enable tracking system 320 to track handheld controller 302. Additionally, handheld controller 302 includes at least one processor (e.g., processor 312) and a communication unit 314.

[0049] Tracking system 320 includes a camera 322 which may capture images of handheld controller 302 (including of LEDs 306 of handheld controller 302 and images of hands and / or arms of a user of controller 302) . Tracking system 320 includes at least one processor (e.g., processor 326) . Tracking system 320, using the processor 326, may track handheld controller 302 based on the images of LEDs 306 captured by camera 322. Tracking system 320 may further include a communication unit 324, with which tracking system 320 may communicate with handheld controller 302 (via a communication unit 314 of handheld controller 302) . For example, handheld controller 302 may communicate status message and / or movement data (e.g., based on measurements from an inertial measurement unit (IMU 308) to tracking system 320. Tracking  system 320 may use the motion data in tracking controller 302. Tracking system 320 may communicate control messages (e.g., instructing handheld controller 302 to illuminate LEDs 306) .

[0050] FIG. 4 is a hybrid block diagram  / flow chart illustrating operation of an example process 400 that may be used to track a controller, according to various aspects of the present disclosure. One or more operations of process 400 may be performed by a computing device (or apparatus) or a component (e.g., a chipset, codec, etc. ) of the computing device. The computing device may be a mobile device (e.g., a mobile phone) , a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capabilities to perform the process 400. The one or more operations of process 400 may be implemented as software components that are executed and run on one or more processors.

[0051] Process 400 may be implemented by tracking system 320 of FIG. 3A and FIG. 3B to track handheld controller 302 of FIG. 3A and FIG. 3B. For example, tracking system 320 may capture image 402 of handheld controller 302 (and / or of hands and arms of a user holding handheld controller 302) using camera (s) 322 of FIG. 3A and FIG. 3B. Additionally, IMU 308 of handheld controller 302 of FIG. 3B may capture motion data 428 and provide motion data 428 to tracking system 320 via communication unit 314 of FIG. 3B.

[0052] Process 400 (or a system or device implementing process 400) may receive as inputs: images 402 (which may include from multiple cameras, such as 2 or 4 cameras such as camera (s) 322 of FIG. 3A and FIG. 3B) , motion data 428 (e.g., from handheld controller 302, which may include IMU 308, which may include gyroscope data and / or accelerometer data) , and pose data (not illustrated in FIG. 4) for each camera (e.g., from a head-tracking algorithm running in tracking system 320) . Pose data for the various cameras may be referred to as The timestamps of image 402 and motion data 428 should be software or hardware synchronized. In some aspects, at block 430, process 400 may predict a pose of handheld controller 302 based on motion data 428. Additionally or alternatively, process 400 may determine motion-data preintegration 432 based on motion data 428.

[0053] Process 400 may generate (and / or output) pose 436 -a pose of a controller (e.g., handheld controller 302) . Pose data for the controller may be referred to as

[0054] Image 402 is provided as an example image input to process 400. Image 402 may include a number of images captured at a number of times (e.g., sequentially) as the controller is tracked over time. Image 402 is provided as an example of one of such images to illustrate process 400 occurring once on one image of the number of images.

[0055] Additionally or alternatively, image 402 may be, or may include, multiple images such as, for example, a short-exposure image and an auto-exposure image. For example, image 402 may include an image captured with a short exposure duration, which may allow LEDs (e.g., LEDs 306 of handheld controller 302) to stand out while the underexposed background and hands in the image are too dark. Further, image 402 may include an image captured according to an automatic exposure setting, which may include the hands and background exposed such that the hands and background are visible.

[0056] At extract LED blobs 404 and extract image features 408, process 400 may process image 402. For example, process 400 may extract LED blobs (LED features 406) and aided visual features (image features 410) in parallel, using traditional computer vision methods or deep learning. LEDs in image 402 are usually round or oval bright blobs. And image features 410 can be hand joints and visual feature points on hands, arms, and / or controllers.

[0057] At decision block 412, process 400 may check the tracking state. For example, process 400 may determine whether the controller is being tracked (based on having been detected and localized in a previous image) .

[0058] If the state is not in the tracking mode, that is to say not initialized or in the lost mode, process 400 may proceed to search 424 and localize 426. At search 424 and localize 426, process 400 may perform initialization when not initialized, or perform relocalization operations when in the lost mode to recovery the controller pose. After initialization or relocalization, the fusion module (e.g., fuse 434) may be initialized or reset.

[0059] If the state is in the tracking mode, that is to say the previous controller pose is known, process 400 may proceed to tracker 414, which includes Track LEDs 416 and track image features  420. At tracker 414, process 400 may predict the current controller pose using motion-data integration (e.g., based on motion data 428) based on a previous pose, and then predict the positions of LEDs and features in the current images. As a result, process 400 may search and match correspondent LED blobs and features near the predicted positions in the current images. This can effectively reduce the computation time and improve the tracking robustness.

[0060] After tracking LEDs and features, all the tracking measurements (e.g., LED tracking measurements 418 and image tracking measurements 422) , motion-data preintegration (motion-data preintegration 432) and camera poses will be passed to the tightly-coupled fusion module (e.g., fuse 434) . The output of the fusion module (e.g., fuse 434) is the controller pose (e.g., pose 436) at the current image timestamp.

[0061] FIG. 5A is a flowchart illustrating an example process 500 for performing an IMU-LED-feature tightly-coupled fusion, according to various aspects of the present disclosure. One or more operations of process 500 may be performed by a computing device (or apparatus) or a component (e.g., a chipset, codec, etc. ) of the computing device. The computing device may be a mobile device (e.g., a mobile phone) , a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capabilities to perform the process 500. The one or more operations of process 500 may be implemented as software components that are executed and run on one or more processors. For example, tracking system 320 of FIG. 3A and FIG. 3B may perform process 500. In some aspects, process 500 may be performed with relation to fuse 434 of FIG. 4.

[0062] Fusion methods can be divided into loosely-coupled and tightly-coupled according to the degree of information fusion. Loosely-coupled method fuse estimated poses from visual information (LED features and images features) and IMU information, and the measurements are the pose solved from LED and features. In general, to solve the pose of controller, at least 3 LED / feature points are needed. If less than 3 LED / feature points, the controller pose cannot be solved by the visual information, so the loosely-coupled method cannot fuse the visual information with less than 3 LED / feature points.

[0063] Tightly-coupled methods directly use the successfully-matched LED blobs / features on the image as the observation measurements, and the degree of information coupling is more tight. In contrast to loosely-coupled method, tightly-coupled methods can also fuse visual information at less than 3 LED / feature points, even if only one point can be fused. Therefore, tightly-coupled method requires less LED / feature point number, and the tracking robustness is higher, especially in the case of occlusion.

[0064] At a high-level, process 500 includes: propagating system state vector and covariance matrix with IMU preintegration, building a filter-measurement model (e.g., an extended Kalman filter (EKF) measurement model) with LED measurements and feature measurements, removing outliers using a threshold, (e.g., a Mahalanobis distance test) , augmenting and / or marginalizing the system state vector and the covariance matrix, performing a filter update (e.g., an EKF filter update) to get the optimal system error state estimations and post covariance matrix, correcting system states using the estimate system error state, maintaining feature / IMU relationships, marginalizing the oldest pose state from the sliding window.

[0065] Before describing each module of fusion algorithm in detail, coordination system and mathematical symbol definitions are introduced.

[0066] Coordinate definitions: W: the world coordinate system. C: the camera coordinate system. I: the controller IMU coordinate system.

[0067] R: a transformation (e.g., rotation) matrix. For example,  the rotation matrix from A coordinate system to B coordinate system. p the position of a coordinate system in another coordinate system. For example,  the position of A coordinate system in B coordinate system.

[0068] Error definitions:  where x is the ground truth,  is the estimated value of x, and δx is the error of x.

[0069] Process 500 may use a “state. ” For example, process 500 may propagate and update a state to track the state. The state may be referred to as a “system state. ” At each time step k, the fusion system may maintain the following error state vector:

[0070] where Xf is the error state vector of the IMU-feature relative position,

[0071] where and

[0072] for j=1, …, N, denotes the error of relative position of the feature fj in the IMU coordinate system {I} .

[0073] Further for i=k-M+1, …, k, represents controller IMU poses at time step i, where M is the sliding window size. Each controller IMU pose in sliding window is defined as:

[0074] where is the error state vector of rotation from IMU coordinate system {Ii} to World coordinate system {W}

[0075] where is the error state of the position of {Ii} in {W} .

[0076] XE is the extra error states for IMU biases and velocity:

[0077] where and are the bias errors of the gyroscope and accelemeter respectively, and δvk is velocity error at time step k.

[0078] As mentioned previously, process 500 may propagate the state, for example, at block 502. An IMU preintegration recursion formula may be:

[0079] where ΔRk, k-1, Δvk, k-1, Δpk, k-1 are IMU preintegration results;

[0080] ΔRk, k-1 is the rotation increment from k-1 time step to k time step;

[0081] Δvk, k-1 is the velocity increment from k-1 time step to k time step; and

[0082] Δpk, k-1 is the position increment from k-1 time step to k time step.

[0083] The IMU error state vector is defined as:

[0084] And the propagation of xIMU can be expressed as: xIMU, k=FIMUxIMU, k-1+GIMUnIMU

[0085] where nIMU= [σg σa σbg σba] T is the system noise.

[0086] The Covariance matrix of nIMU, QIMU, depends on the IMU noise characteristics and is calculated during sensor calibration.

[0087] So the propagation of the whole system state vector xk can be expressed as:

[0088] At block 504, process 500 may build a measurement model. The general linearized form of measurement model (which may be an EKF measurement model) may be: rk=Hxk+noise

[0089] where rk is the measurement residuals, H is the measurement jacobian matrix, and the noise term is zero-mean, Gaussian, and uncorrelated to the error state xk.

[0090] For IMU / LED / features fusion, the measurements include LED blobs measurements and feature measurements.

[0091] The j-th LED blob measurement residual:

[0092] where is the j-th LED blob measurement,  is the estimated j-th LED blob measurement.

[0093] is the j-th LED blob measurement Jacobian matrix of controller pose state at time step k.

[0094] And the image projection function π is defined as:

[0095] The j-th feature measurement residual

[0096] Where is the j-th feature measurement,  is the estimated j-th feature measurement.

[0097] is the j-th feature measurement jacobian matrix of feature-IMU relative position

[0098] is the j-th feature measurement Jacobian matrix of controller pose state at time step k.

[0099] Stacking all the LED blob and feature measurement residuals together, may result in:

[0100] At block 506, process 500 may remove outliers. The relationships between features and IMU are not stable, the relative positions are prone to change. So before employing feature measurement updates, an outlier rejection procedure may be applied. The fusion system employs, as an example, a Mahalanobis distance test. S=HPHT+σ2I

[0101] where γ represents the Mahalanobis distance,

[0102] is measurement residual, S is the residual covariance,

[0103] H is measurement Jacobian, P is covariance matrix, and

[0104] σ is measurement noise standard deviation.

[0105] All these values are all available from the filter (e.g., the EKF filter) . If the γ is larger than a certain threshold, the feature will be considered as an outlier.

[0106] At block 508, process 500 may augment and / or marginalize feature states. For example, if old features already in state vector becomes an outlier, the state vector and covariance matrix should be marginalized. If new features are added, the state vector and covariance matrix should be augmented.

[0107] At block 510, process 500 may update the filter (e.g., the EKF filter) . For example, by this point in process 500, all the states may be propagated and all the measurements are ready. Text the filter may be updated. For example, a Kalman gain may be:

[0108] where Rk is the measurement noise covariance matrix.

[0109] The estimated error state may be:

[0110] Finally, the state covariance matrix is updated according to:

[0111] At block 512, the state may be corrected. Further, at block 512, the feature-motion data relationship may be maintained. For example, after the filter is updated, the estimated error state may be obtained. The error may be used to correct states. To correct the relationships between the j-th feature and IMU:

[0112] To correct the i-th pose in the sliding window:

[0113] Correct the extra states:

[0114] At block 516, the controller pose states may be marginalized. For example, the oldest pose state should be marginalized out of the sliding window.

[0115] At block 514, the feature-motion data relationship may be maintained.

[0116] The inputs to process 500 may include:

[0117] · Prior error state estimate

[0118] · Prior covariance matrix

[0119] · IMU preintegration increments ΔRk, k-1, Δvk, k-1, Δpk, k-1

[0120] · Camera pose:

[0121] · LED blob measurements:

[0122] · Feature measurements:

[0123] Process 500 may include at least the following steps:

[0124] · Propagate states

[0125] · Build measurement model

[0126] · Remove outliers

[0127] · Augment / Marginalize states

[0128] · EKF update

[0129] · Correct states and Maintain feature / IMU relationship

[0130] · Marginalize states

[0131] The present disclosure discloses a fusion algorithm for ringless CV-based controllers, including a novel motion data-LED-feature tightly-coupled fusion algorithm to solve.

[0132] The present disclosure discloses a tightly-coupled fusion framework that can simultaneously fuse motion data, LED tracking measurements, and feature tracking measurements. Additionally, the present disclosure discloses a sliding window to the fusion state vector to estimate the relationships between the controller IMU and the tracking features at the same time. Additionally, the present disclosure discloses adopting a distance test (e.g., the Mahalanobis distance test) to detect the changes of relationships between the controller IMU data and the tracking features, and then reject outliers.

[0133] FIG. 5B is a flow diagram illustrating a process 550 for tracking a human-machine-interface (HMI) device, in accordance with aspects of the present disclosure. One or more operations of process 550 may be performed by a computing device (or apparatus) or a component (e.g., a chipset, codec, etc. ) of the computing device. The computing device may be a mobile device (e.g., a mobile phone) , a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capabilities to perform the process 550. The one or more operations of process 550 may be implemented as software components that are executed and run on one or more processors.

[0134] At block 552, a computing device (or one or more components thereof) may propagate a system state vector and a covariance matrix with IMU preintegration. For example, tracking system 320 of FIG. 3A and FIG. 3B may propagate a system state vector and a covariance matrix with IMU preintegration.

[0135] At block 554, the computing device (or one or more components thereof) may build a filter measurement model with LED measurements and feature measurements. For example, filter system 320 may build a filter measurement model with LED measurements and feature measurements.

[0136] At block 556, the computing device (or one or more components thereof) may remove outliers using a distance test. For example, filter system 320 may remove outliers using a distance test.

[0137] At block 558, the computing device (or one or more components thereof) may augment or marginalizing the system state vector and the covariance matrix. For example, filter system 320 may augment or marginalizing the system state vector and the covariance matrix.

[0138] At block 560, the computing device (or one or more components thereof) may update the filter measurement model to decrease system error state estimations and post covariance matrix. For example, filter system 320 may update the filter measurement model to decrease system error state estimations and post covariance matrix.

[0139] At block 562, the computing device (or one or more components thereof) may correct system states using the system error state estimations. For example, filter system 320 may correct system states using the system error state estimations.

[0140] At block 564, the computing device (or one or more components thereof) may maintain feature-motion data relationship. For example, filter system 320 may maintain feature-motion data relationship.

[0141] At block 566, the computing device (or one or more components thereof) may marginalize an oldest pose state from a sliding window. For example, filter system 320 may marginalize an oldest pose state from a sliding window.

[0142] FIG. 6A is a flow diagram illustrating a process 600 for tracking a human-machine-interface (HMI) device, in accordance with aspects of the present disclosure. One or more operations of process 600 may be performed by a computing device (or apparatus) or a component (e.g., a chipset, codec, etc. ) of the computing device. The computing device may be a mobile device (e.g., a mobile phone) , a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capabilities to perform the one or more operations of process 600. The one or more operations of process 600 may be implemented as software components that are executed and run on one or more processors.

[0143] At block 602, a computing device (or one or more components thereof) may extract image features from images of the HMI device captured by a camera of a head-mounted device (HMD) , the image features related to at least one of hands or arms of a user holding the HMI device. For example, tracking system 320 of FIG. 3A and FIG. 3B may obtain (e.g., capture using a camera of tracking system 320) images (e.g., image 402 of FIG. 4) of handheld controller 302, hand 310 and / or arm of a user holding handheld controller 302. Tracking system 320 may extract image features (e.g., image features 410 of FIG. 4) from the images.

[0144] At block 604, the computing device (or one or more components thereof) may extract light-emitting diode (LED) features from the images, the LED features related to LEDs on the HMI device. For example, tracking system 320 may extract LED features (e.g, . LED features 406 of FIG. 4) from the images.

[0145] At block 606, the computing device (or one or more components thereof) may track the image features and the LED features across multiple images. For example, tracking system 320 may track the image features (e.g., image features 410) and the LED features (e.g., LED features 406) across multiple images (e.g., of handheld controller 302, hand 310, and / or an arm of a holder of handheld controller 302, the images captured by, for example, tracking system 320) .

[0146] At block 608, the computing device (or one or more components thereof) may obtain motion data from the HMI device, the motion data measured by an inertial measurement unit (IMU) of the HMI device. For example, tracking system 320 may obtain motion data from handheld  controller 302. The motion data may be measured by IMU 308 of FIG. 3B of handheld controller 302.

[0147] At block 610, the computing device (or one or more components thereof) may determine motion-data preintegration based on the motion data. For example, tracking system 320 may determine motion-data preintegration (e.g., motion-data preintegration 432 of FIG. 4) based on the motion data (e.g., motion data 428) .

[0148] At block 612, the computing device (or one or more components thereof) may predict, based on the tracked image features, the tracked LED features, and the motion data, positions of the LEDs on the HMI device, and positions of the image features. For example, tracking system 320 may predict LED tracking measurements 418 and image tracking measurements 422 based on the tracking of LED features 406, image features 410, motion data 428, and / or motion-data preintegration 432.

[0149] At block 614, the computing device (or one or more components thereof) may fuse the motion-data preintegration, the positions of the LEDs, and the positions of the image features to determine a pose of the HMI device. For example, tracking system 320 may fuse motion-data preintegration 432, LED tracking measurements 418, and image tracking measurements 422 to generate pose 436.

[0150] In some aspects, to fuse the motion data, LED positions of the LEDs, and the positions of the image features, the computing device (or one or more components thereof) may track and update the pose of the HMI device over time using a filter. For example, tracking system 320 may use a filter to track and update the pose of handheld controller 302 over time. The filter may be, or may include, an algorithm used to estimate the state of a dynamic system by updating an estimate based on a sequence of observations. In some aspects, the filter may be, or may include, at least one of: an extended Kalman filter (EKF) an uncentered Kalman filter (UKF) , or a particle filter.

[0151] In some aspects, to fuse the motion data, the positions of the LEDs, and the positions of the image features, the computing device (or one or more components thereof) may determine a system state vector comprising errors of relative positions of image features in an IMU coordinate system. For example, tracking system 320 may determine a system state vector comprising errors  of relative positions of image features in an IMU coordinate system. In some aspects, the system state vector may be, or may include, a sliding window of the relative positions.

[0152] In some aspects, the computing device (or one or more components thereof) may filter values of the system state vector. In some aspects, the system state vector may be filtered according to a Mahalanobis distance.

[0153] FIG. 6B is a flow diagram illustrating a process 620 for tracking a human-machine-interface (HMI) device, in accordance with aspects of the present disclosure. One or more operations of process 620 may be performed by a computing device (or apparatus) or a component (e.g., a chipset, codec, etc. ) of the computing device. The computing device may be a mobile device (e.g., a mobile phone) , a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capabilities to perform the one or more operations of process 620. The one or more operations of process 620 may be implemented as software components that are executed and run on one or more processors.

[0154] At block 622, a computing device (or one or more components thereof) may determine light-emitting diode (LED) -tracking measurements based on images of the HMI device, wherein the HMI device comprises LEDs. For example, tracking system 320 of FIG. 3A and FIG. 3B may obtain (e.g., capture using a camera of tracking system 320) images (e.g., image 402 of FIG. 4) of handheld controller 302, hand 310 and / or arm of a user holding handheld controller 302. Tracking system 320 may extract LED features (e.g, . LED features 406 of FIG. 4) from the images. Tracking system 320 may determine LED-tracking measurements based on the images 402 and / or based on LED features 406.

[0155] In some aspects, , to determine the LED-tracking measurements, the computing device (or one or more components thereof) may identify LED features of the images of the HMI device; and track the LED features across the images of the HMI device. For example, tracking system 320 may may extract LED features (e.g, . LED features 406 of FIG. 4) from images 402 and track the LED features across instances of image 402.

[0156] In some aspects, to determine the LED-tracking measurements, the computing device (or one or more components thereof) may predict positions of the LED features. For example, tracking system 320 may predict positions of LED features 406 in an upcoming instance of image 402.

[0157] At block 624, the computing device (or one or more components thereof) may determine feature-tracking measurements based on the images of the HMI device. For example, tracking system 320 may obtain images 402 of handheld controller 302, hand 310 and / or arm of a user holding handheld controller 302. Tracking system 320 may extract image features (e.g., image features 410 of FIG. 4) from the images. Tracking system 320 may determine feature-tracking measurements based on the images 402 or based on image features 410.

[0158] In some aspects, to determine the feature-tracking measurements, the computing device (or one or more components thereof) may: identify image features of the images of the HMI device; and track the image features across the images of the HMI device. For example, tracking system 320 may extract image features (e.g., image features 410 of FIG. 4) from the images 402 and track the image features across instances of image 402.

[0159] In some aspects, to determine the feature-tracking measurements, the computing device (or one or more components thereof) may predict positions of the image features. For example, tracking system 320 may predict positions of image features 410 in an upcoming instance of image 402.

[0160] In some aspects, the image features may be related to at least one of: surface features of the HMI device, a hand of a user holding the HMI device, or an arm of the user. For example, image features 410 may be related to surface features of handheld controller 302, a hand of a user holding handheld controller 302, or an arm of the user.

[0161] At block 626, the computing device (or one or more components thereof) may determine movement data of the HMI device based on an inertial measurement unit (IMU) of the HMI device. For example, tracking system 320 may obtain motion data from handheld controller 302. The motion data may be measured by IMU 308 of FIG. 3B of handheld controller 302. Tracking system 320 may determine motion-data preintegration (e.g., motion-data preintegration 432 of FIG. 4) based on the motion data (e.g., motion data 428) .

[0162] In some aspects, to determine the movement data, the computing device (or one or more components thereof) may determine motion-data preintegration based on measurements of the IMU. For example, tracking system 320 may determine motion-data preintegration 432 based on motion data 428.

[0163] In some aspects, to determine the movement data, the computing device (or one or more components thereof) may predict a pose of the HMI device based on measurements of the IMU. For example, at block 430, tracking system 320 may predict a pose of handheld controller 302 based on motion data 428.

[0164] At block 628, the computing device (or one or more components thereof) may determine a pose of the HMI device by fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data. For example, tracking system 320 may fuse motion-data preintegration 432, LED tracking measurements 418, and image tracking measurements 422 to generate pose 436.

[0165] In some aspects, to fuse the LED-tracking measurements, the feature-tracking measurements, and the movement data, the computing device (or one or more components thereof) may track and update the pose of the HMI device over a period of time using a filter. For example, tracking system 320 may use a filter to track and update the pose of handheld controller 302 over time. The filter may be, or may include, an algorithm used to estimate the state of a dynamic system by updating an estimate based on a sequence of observations. In some aspects, the filter may be, or may include, at least one of: an extended Kalman filter (EKF) an uncentered Kalman filter (UKF) , or a particle filter.

[0166] In some aspects, to determine the pose of the HMI device, the computing device (or one or more components thereof) may update a state of a filter based on the LED-tracking measurements, the feature-tracking measurements, and the movement data. For example, tracking system 320 may, at block 510 of process 500 of FIG. 5A, update a state of a filter.

[0167] In some aspects, the filter may be, or may include, at least one of: an extended Kalman filter (EKF) , an uncentered Kalman filter (UKF) , or a particle filter. For example, the filter updated at block 510 may be, or may include, an EKF, UKF, or a particle filter.

[0168] In some aspects, to fuse the LED-tracking measurements, the feature-tracking measurements, and the movement data, the computing device (or one or more components thereof) may determine a system-state vector comprising errors of relative positions of image features in an IMU coordinate system. For example, tracking system 320 may determine a system-state vector including errors of positions of image features in a coordinate system of handheld controller 302.

[0169] In some aspects, the system-state vector may be, or may include, a sliding window of the relative positions. For example, tracking system 320 may maintain a sliding window of positions of handheld controller 302.

[0170] In some aspects, the system-state vector is filtered according to a distance. For example, tracking system 320 may filter the system-state vector based on distances.

[0171] In some aspects, the computing device (or one or more components thereof) may fuse the LED-tracking measurements, the feature-tracking measurements, and the movement data based on a camera pose of a camera used to capture the images of the HMI device. For example, tracking system 320 may fuse the LED-tracking measurements, the feature-tracking measurements, and the movement data based at least in part, on pose data for the various cameras may be referred to as

[0172] In some aspects, the computing device (or one or more components thereof) may be a computing device of a device. The device may include a head-mounted-display and the camera. For example, the computing device (or one or more components thereof) may be, or may include, tracking system 320. Tracking system 320 may include a display at camera 322.

[0173] In some examples, as noted previously, the methods described herein (e.g., process 400 of FIG. 4, process 500 of FIG. 5A, process 550 of FIG. 5B, process 600 of FIG. 6A, process 620 of FIG. 6B, and / or other methods described herein) can be performed, in whole or in part, by a computing device or apparatus. In one example, one or more of the methods can be performed by tracking system 320 of FIG. 3A and FIG. 3B, or by another system or device. In another example, one or more of the methods (e.g., process 400, process 500, process 550, process 600, process 620, and / or other methods described herein) can be performed, in whole or in part, by the computing-device architecture 900 shown in FIG. 9. For instance, a computing device with the computing- device architecture 900 shown in FIG. 9 can include, or be included in, the components of the tracking system 320 and can implement the operations of process 400, process 500, process 550, process 600, process 620, and / or other process described herein. In some cases, the computing device or apparatus can include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other component (s) that are configured to carry out the steps of processes described herein. In some examples, the computing device can include a display, a network interface configured to communicate and / or receive the data, any combination thereof, and / or other component (s) . The network interface can be configured to communicate and / or receive Internet Protocol (IP) based data or other type of data.

[0174] The components of the computing device can be implemented in circuitry. For example, the components can include and / or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs) , digital signal processors (DSPs) , central processing units (CPUs) , and / or other suitable electronic circuits) , and / or can include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein.

[0175] Process 400, process 500, process 550, process 600, process 620, and / or other process described herein are illustrated as logical flow diagrams, the operation of which represents a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the processes.

[0176] Additionally, process 400, process 500, process 550, process 600, process 620, and / or other process described herein can be performed under the control of one or more computer  systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code can be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.

[0177] As noted above, various aspects of the present disclosure can use machine-learning models or systems.

[0178] FIG. 7 is an illustrative example of a neural network 700 (e.g., a deep-learning neural network) that can be used to implement machine-learning based feature segmentation, implicit-neural-representation generation, rendering, classification, object detection, image recognition (e.g., face recognition, object recognition, scene recognition, etc. ) , feature extraction, authentication, gaze detection, gaze prediction, and / or automation. For example, neural network 700 may be an example of, or can perform feature extraction related to extract LED blobs 404 and / or extract image features 408 of FIG. 4

[0179] An input layer 702 includes input data. In one illustrative example, input layer 702 can include data representing image 402 of FIG. 4. Neural network 700 includes multiple hidden layers, for example, hidden layers 706a, 706b, through 706n. The hidden layers 706a, 706b, through hidden layer 706n include “n” number of hidden layers, where “n” is an integer greater than or equal to one. The number of hidden layers can be made to include as many layers as needed for the given application. Neural network 700 further includes an output layer 704 that provides an output resulting from the processing performed by the hidden layers 706a, 706b, through 706n. In one illustrative example, output layer 704 can provide LED features 406 and / or image features 410 of FIG. 4.

[0180] Neural network 700 may be, or may include, a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. Information associated with the nodes is shared among the different layers and each layer retains information as information is processed. In some cases, neural network 700 can include a feed-forward network, in which case there are no feedback connections where outputs of the network are fed back into itself. In some  cases, neural network 700 can include a recurrent neural network, which can have loops that allow information to be carried across nodes while reading in input.

[0181] Information can be exchanged between nodes through node-to-node interconnections between the various layers. Nodes of input layer 702 can activate a set of nodes in the first hidden layer 706a. For example, as shown, each of the input nodes of input layer 702 is connected to each of the nodes of the first hidden layer 706a. The nodes of first hidden layer 706a can transform the information of each input node by applying activation functions to the input node information. The information derived from the transformation can then be passed to and can activate the nodes of the next hidden layer 706b, which can perform their own designated functions. Example functions include convolutional, up-sampling, data transformation, and / or any other suitable functions. The output of the hidden layer 706b can then activate nodes of the next hidden layer, and so on. The output of the last hidden layer 706n can activate one or more nodes of the output layer 704, at which an output is provided. In some cases, while nodes (e.g., node 708) in neural network 700 are shown as having multiple output lines, a node has a single output and all lines shown as being output from a node represent the same output value.

[0182] In some cases, each node or interconnection between nodes can have a weight that is a set of parameters derived from the training of neural network 700. Once neural network 700 is trained, it can be referred to as a trained neural network, which can be used to perform one or more operations. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnection can have a tunable numeric weight that can be tuned (e.g., based on a training dataset) , allowing neural network 700 to be adaptive to inputs and able to learn as more and more data is processed.

[0183] Neural network 700 may be pre-trained to process the features from the data in the input layer 702 using the different hidden layers 706a, 706b, through 706n in order to provide the output through the output layer 704. In an example in which neural network 700 is used to identify features in images, neural network 700 can be trained using training data that includes both images and labels, as described above. For instance, training images can be input into the network, with each training image having a label indicating the features in the images (for the feature-segmentation machine-learning system) or a label indicating classes of an activity in each image.  In one example using object classification for illustrative purposes, a training image can include an image of a number 2, in which case the label for the image can be [0 0 1 0 0 0 0 0 0 0] .

[0184] In some cases, neural network 700 can adjust the weights of the nodes using a training process called backpropagation. As noted above, a backpropagation process can include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. The process can be repeated for a certain number of iterations for each set of training images until neural network 700 is trained well enough so that the weights of the layers are accurately tuned.

[0185] For the example of identifying objects in images, the forward pass can include passing a training image through neural network 700. The weights are initially randomized before neural network 700 is trained. As an illustrative example, an image can include an array of numbers representing the pixels of the image. Each number in the array can include a value from 0 to 255 describing the pixel intensity at that position in the array. In one example, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or luma and two chroma components, or the like) .

[0186] As noted above, for a first training iteration for neural network 700, the output will likely include values that do not give preference to any particular class due to the weights being randomly selected at initialization. For example, if the output is a vector with probabilities that the object includes different classes, the probability value for each of the different classes can be equal or at least very similar (e.g., for ten possible classes, each class can have a probability value of 0.1) . With the initial weights, neural network 700 is unable to determine low-level features and thus cannot make an accurate determination of what the classification of the object might be. A loss function can be used to analyze error in the output. Any suitable loss function definition can be used, such as a cross-entropy loss. Another example of a loss function includes the mean squared error (MSE) , defined as The loss can be set to be equal to the value of Etotal.

[0187] The loss (or error) will be high for the first training images since the actual values will be much different than the predicted output. The goal of training is to minimize the amount of loss so that the predicted output is the same as the training label. Neural network 700 can perform a  backward pass by determining which inputs (weights) most contributed to the loss of the network and can adjust the weights so that the loss decreases and is eventually minimized. A derivative of the loss with respect to the weights (denoted as dL / dW, where W are the weights at a particular layer) can be computed to determine the weights that contributed most to the loss of the network. After the derivative is computed, a weight update can be performed by updating all the weights of the filters. For example, the weights can be updated so that they change in the opposite direction of the gradient. The weight update can be denoted as where w denotes a weight, wi denotes the initial weight, and η denotes a learning rate. The learning rate can be set to any suitable value, with a high learning rate including larger weight updates and a lower value indicating smaller weight updates.

[0188] Neural network 700 can include any suitable deep network. One example includes a convolutional neural network (CNN) , which includes an input layer and an output layer, with multiple hidden layers between the input and out layers. The hidden layers of a CNN include a series of convolutional, nonlinear, pooling (for downsampling) , and fully connected layers. Neural network 700 can include any other deep network other than a CNN, such as an autoencoder, a deep belief nets (DBNs) , a Recurrent Neural Networks (RNNs) , among others.

[0189] FIG. 8 is an illustrative example of a convolutional neural network (CNN) 800. The input layer 802 of the CNN 800 includes data representing an image or frame. For example, the data can include an array of numbers representing the pixels of the image, with each number in the array including a value from 0 to 255 describing the pixel intensity at that position in the array. Using the previous example from above, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luma and two chroma components, or the like) . The image can be passed through a convolutional hidden layer 804, an optional non-linear activation layer, a pooling hidden layer 806, and fully connected layer 808 (which fully connected layer 808 can be hidden) to get an output at the output layer 810. While only one of each hidden layer is shown in FIG. 8, one of ordinary skill will appreciate that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers can be included in the CNN 800. As previously described, the output can indicate a single class of an object or can include a probability of classes that best describe the object in the image.

[0190] The first layer of the CNN 800 can be the convolutional hidden layer 804. The convolutional hidden layer 804 can analyze image data of the input layer 802. Each node of the convolutional hidden layer 804 is connected to a region of nodes (pixels) of the input image called a receptive field. The convolutional hidden layer 804 can be considered as one or more filters (each filter corresponding to a different activation or feature map) , with each convolutional iteration of a filter being a node or neuron of the convolutional hidden layer 804. For example, the region of the input image that a filter covers at each convolutional iteration would be the receptive field for the filter. In one illustrative example, if the input image includes a 28×28 array, and each filter (and corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in the convolutional hidden layer 804. Each connection between a node and a receptive field for that node learns a weight and, in some cases, an overall bias such that each node learns to analyze its particular local receptive field in the input image. Each node of the convolutional hidden layer 804 will have the same weights and bias (called a shared weight and a shared bias) . For example, the filter has an array of weights (numbers) and the same depth as the input. A filter will have a depth of 3 for an image frame example (according to three color components of the input image) . An illustrative example size of the filter array is 5 x 5 x 3, corresponding to a size of the receptive field of a node.

[0191] The convolutional nature of the convolutional hidden layer 804 is due to each node of the convolutional layer being applied to its corresponding receptive field. For example, a filter of the convolutional hidden layer 804 can begin in the top-left corner of the input image array and can convolve around the input image. As noted above, each convolutional iteration of the filter can be considered a node or neuron of the convolutional hidden layer 804. At each convolutional iteration, the values of the filter are multiplied with a corresponding number of the original pixel values of the image (e.g., the 5x5 filter array is multiplied by a 5x5 array of input pixel values at the top-left corner of the input image array) . The multiplications from each convolutional iteration can be summed together to obtain a total sum for that iteration or node. The process is next continued at a next location in the input image according to the receptive field of a next node in the convolutional hidden layer 804. For example, a filter can be moved by a step amount (referred to as a stride) to the next receptive field. The stride can be set to 1 or any other suitable amount. For example, if the stride is set to 1, the filter will be moved to the right by 1 pixel at each convolutional iteration. Processing the filter at each unique location of the input volume produces  a number representing the filter results for that location, resulting in a total sum value being determined for each node of the convolutional hidden layer 804.

[0192] The mapping from the input layer to the convolutional hidden layer 804 is referred to as an activation map (or feature map) . The activation map includes a value for each node representing the filter results at each location of the input volume. The activation map can include an array that includes the various total sum values resulting from each iteration of the filter on the input volume. For example, the activation map will include a 24 x 24 array if a 5 x 5 filter is applied to each pixel (a stride of 1) of a 28 x 28 input image. The convolutional hidden layer 804 can include several activation maps in order to identify multiple features in an image. The example shown in FIG. 8 includes three activation maps. Using three activation maps, the convolutional hidden layer 804 can detect three different kinds of features, with each feature being detectable across the entire image.

[0193] In some examples, a non-linear hidden layer can be applied after the convolutional hidden layer 804. The non-linear layer can be used to introduce non-linearity to a system that has been computing linear operations. One illustrative example of a non-linear layer is a rectified linear unit (ReLU) layer. A ReLU layer can apply the function f (x) = max (0, x) to all of the values in the input volume, which changes all the negative activations to 0. The ReLU can thus increase the non-linear properties of the CNN 800 without affecting the receptive fields of the convolutional hidden layer 804.

[0194] The pooling hidden layer 806 can be applied after the convolutional hidden layer 804 (and after the non-linear hidden layer when used) . The pooling hidden layer 806 is used to simplify the information in the output from the convolutional hidden layer 804. For example, the pooling hidden layer 806 can take each activation map output from the convolutional hidden layer 804 and generates a condensed activation map (or feature map) using a pooling function. Max-pooling is one example of a function performed by a pooling hidden layer. Other forms of pooling functions be used by the pooling hidden layer 806, such as average pooling, L2-norm pooling, or other suitable pooling functions. A pooling function (e.g., a max-pooling filter, an L2-norm filter, or other suitable pooling filter) is applied to each activation map included in the convolutional hidden  layer 804. In the example shown in FIG. 8, three pooling filters are used for the three activation maps in the convolutional hidden layer 804.

[0195] In some examples, max-pooling can be used by applying a max-pooling filter (e.g., having a size of 2x2) with a stride (e.g., equal to a dimension of the filter, such as a stride of 2) to an activation map output from the convolutional hidden layer 804. The output from a max-pooling filter includes the maximum number in every sub-region that the filter convolves around. Using a 2x2 filter as an example, each unit in the pooling layer can summarize a region of 2×2 nodes in the previous layer (with each node being a value in the activation map) . For example, four values (nodes) in an activation map will be analyzed by a 2x2 max-pooling filter at each iteration of the filter, with the maximum value from the four values being output as the “max” value. If such a max-pooling filter is applied to an activation filter from the convolutional hidden layer 804 having a dimension of 24x24 nodes, the output from the pooling hidden layer 806 will be an array of 12x12 nodes.

[0196] In some examples, an L2-norm pooling filter could also be used. The L2-norm pooling filter includes computing the square root of the sum of the squares of the values in the 2×2 region (or other suitable region) of an activation map (instead of computing the maximum values as is done in max-pooling) and using the computed values as an output.

[0197] The pooling function (e.g., max-pooling, L2-norm pooling, or other pooling function) determines whether a given feature is found anywhere in a region of the image and discards the exact positional information. This can be done without affecting results of the feature detection because, once a feature has been found, the exact location of the feature is not as important as its approximate location relative to other features. Max-pooling (as well as other pooling methods) offer the benefit that there are many fewer pooled features, thus reducing the number of parameters needed in later layers of the CNN 800.

[0198] The final layer of connections in the network is a fully-connected layer that connects every node from the pooling hidden layer 806 to every one of the output nodes in the output layer 810. Using the example above, the input layer includes 28 x 28 nodes encoding the pixel intensities of the input image, the convolutional hidden layer 804 includes 3×24×24 hidden feature nodes based on application of a 5×5 local receptive field (for the filters) to three activation maps, and the  pooling hidden layer 806 includes a layer of 3×12×12 hidden feature nodes based on application of max-pooling filter to 2×2 regions across each of the three feature maps. Extending this example, the output layer 810 can include ten output nodes. In such an example, every node of the 3x12x12 pooling hidden layer 806 is connected to every node of the output layer 810.

[0199] The fully connected layer 808 can obtain the output of the previous pooling hidden layer 806 (which should represent the activation maps of high-level features) and determines the features that most correlate to a particular class. For example, the fully connected layer 808 can determine the high-level features that most strongly correlate to a particular class and can include weights (nodes) for the high-level features. A product can be computed between the weights of the fully connected layer 808 and the pooling hidden layer 806 to obtain probabilities for the different classes. For example, if the CNN 800 is being used to predict that an object in an image is a person, high values will be present in the activation maps that represent high-level features of people (e.g., two legs are present, a face is present at the top of the object, two eyes are present at the top left and top right of the face, a nose is present in the middle of the face, a mouth is present at the bottom of the face, and / or other features common for a person) .

[0200] In some examples, the output from the output layer 810 can include an M-dimensional vector (in the prior example, M=10) . M indicates the number of classes that the CNN 800 has to choose from when classifying the object in the image. Other example outputs can also be provided. Each number in the M-dimensional vector can represent the probability the object is of a certain class. In one illustrative example, if a 10-dimensional output vector represents ten different classes of objects is [0 0 0.05 0.8 0 0.15 0 0 0 0] , the vector indicates that there is a 5%probability that the image is the third class of object (e.g., a dog) , an 80%probability that the image is the fourth class of object (e.g., a human) , and a 15%probability that the image is the sixth class of object (e.g., a kangaroo) . The probability for a class can be considered a confidence level that the object is part of that class.

[0201] FIG. 9 illustrates an example computing-device architecture 900 of an example computing device which can implement the various techniques described herein. In some examples, the computing device can include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR)  device) , a personal computer, a laptop computer, a video server, a vehicle (or computing device of a vehicle) , or other device. For example, the computing-device architecture 900 may include, implement, or be included in any or all of tracking system 320 of FIG. 3A and FIG. 3B. Additionally or alternatively, computing-device architecture 900 may be configured to perform process 400, process 500, process 550, process 600, process 620, and / or other process described herein.

[0202] The components of computing-device architecture 900 are shown in electrical communication with each other using connection 912, such as a bus. The example computing-device architecture 900 includes a processing unit (CPU or processor) 902 and computing device connection 912 that couples various computing device components including computing device memory 910, such as read only memory (ROM) 908 and random-access memory (RAM) 906, to processor 902.

[0203] Computing-device architecture 900 can include a cache of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 902. Computing-device architecture 900 can copy data from memory 910 and / or the storage device 914 to cache 904 for quick access by processor 902. In this way, the cache can provide a performance boost that avoids processor 902 delays while waiting for data. These and other modules can control or be configured to control processor 902 to perform various actions. Other computing device memory 910 may be available for use as well. Memory 910 can include multiple different types of memory with different performance characteristics. Processor 902 can include any general-purpose processor and a hardware or software service, such as service 1 916, service 2 918, and service 3 920 stored in storage device 914, configured to control processor 902 as well as a special-purpose processor where software instructions are incorporated into the processor design. Processor 902 may be a self-contained system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.

[0204] To enable user interaction with the computing-device architecture 900, input device 922 can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech and so forth. Output device 924 can also be one or more of a number of output mechanisms known to those of skill in  the art, such as a display, projector, television, speaker device, etc. In some instances, multimodal computing devices can enable a user to provide multiple types of input to communicate with computing-device architecture 900. Communication interface 926 can generally govern and manage the user input and computing device output. There is no restriction on operating on any particular hardware arrangement and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.

[0205] Storage device 914 is a non-volatile memory and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, random-access memories (RAMs) 906, read only memory (ROM) 908, and hybrids thereof. Storage device 914 can include services 916, 918, and 920 for controlling processor 902. Other hardware or software modules are contemplated. Storage device 914 can be connected to the computing device connection 912. In one aspect, a hardware module that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 902, connection 912, output device 924, and so forth, to carry out the function.

[0206] The term “substantially, ” in reference to a given parameter, property, or condition, may refer to a degree that one of ordinary skill in the art would understand that the given parameter, property, or condition is met with a small degree of variance, such as, for example, within acceptable manufacturing tolerances. By way of example, depending on the particular parameter, property, or condition that is substantially met, the parameter, property, or condition may be at least 90%met, at least 95%met, or even at least 99%met.

[0207] Aspects of the present disclosure are applicable to any suitable electronic device (such as security systems, smartphones, tablets, laptop computers, vehicles, drones, or other devices) including or coupled to one or more active depth sensing systems. While described below with respect to a device having or coupled to one light projector, aspects of the present disclosure are applicable to devices having any number of light projectors and are therefore not limited to specific devices.

[0208] The term “device” is not limited to one or a specific number of physical objects (such as one smartphone, one controller, one processing system and so on) . As used herein, a device may be any electronic device with one or more parts that may implement at least some portions of this disclosure. While the below description and examples use the term “device” to describe various aspects of this disclosure, the term “device” is not limited to a specific configuration, type, or number of objects. Additionally, the term “system” is not limited to multiple components or specific aspects. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. While the below description and examples use the term “system” to describe various aspects of this disclosure, the term “system” is not limited to a specific configuration, type, or number of objects.

[0209] Specific details are provided in the description above to provide a thorough understanding of the aspects and examples provided herein. However, it will be understood by one of ordinary skill in the art that the aspects may be practiced without these specific details. For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks including functional blocks including devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the aspects.

[0210] Individual aspects may be described above as a process or method which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.

[0211] Processes and methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions can include, for example, instructions and data which cause or otherwise configure a general-purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code, etc.

[0212] The term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction (s) and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD) , flash memory, magnetic or optical disks, USB devices provided with non-volatile memory, networked storage devices, any suitable combination thereof, among others. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.

[0213] In some aspects the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bit stream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.

[0214] Devices implementing processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description languages, or any  combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor (s) may perform the necessary tasks. Typical examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.

[0215] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure.

[0216] In the foregoing description, aspects of the application are described with reference to specific aspects thereof, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative aspects of the application have been described in detail herein, it is to be understood that the inventive concepts may be otherwise variously embodied and employed, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the above-described application may be used individually or jointly. Further, aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather than restrictive. For the purposes of illustration, methods were described in a particular order. It should be appreciated that in alternate aspects, the methods may be performed in a different order than that described.

[0217] One of ordinary skill will appreciate that the less than ( “<” ) and greater than ( “>” ) symbols or terminology used herein can be replaced with less than or equal to ( “≤” ) and greater than or equal to ( “≥” ) symbols, respectively, without departing from the scope of this description.

[0218] Where components are described as being “configured to” perform certain operations, such configuration can be accomplished, for example, by designing electronic circuits or other  hardware to perform the operation, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits) to perform the operation, or any combination thereof.

[0219] The phrase “coupled to” refers to any component that is physically connected to another component either directly or indirectly, and / or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and / or other suitable communication interface) either directly or indirectly.

[0220] Claim language or other language reciting “at least one of” a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any duplicate information or data (e.g., A and A, B and B, C and C, A and A and B, and so on) , or any other ordering, duplication, or combination of A, B, and C. The language “at least one of” a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases “at least one” and “one or more” are used interchangeably herein.

[0221] Claim language or other language reciting “at least one processor configured to, ” “at least one processor being configured to, ” “one or more processors configured to, ” “one or more processors being configured to, ” or the like indicates that one processor or multiple processors (in any combination) can perform the associated operation (s) . For example, claim language reciting “at least one processor configured to: X, Y, and Z” means a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each tasked with a certain subset of operations X, Y, and Z such that together the multiple processors perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language reciting “at least one processor configured to: X, Y, and Z” can mean that any single processor may only perform at least a subset of operations X, Y, and Z.

[0222] Where reference is made to one or more elements performing functions (e.g., steps of a method) , one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function) . Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions.

[0223] Where reference is made to an entity (e.g., any entity or device described herein) performing functions or being configured to perform functions (e.g., steps of a method) , the entity may be configured to cause one or more elements (individually or collectively) to perform the functions. The one or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of the functions, and / or any combination thereof. Where reference to the entity performing functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to collectively perform the functions. When the entity is configured to cause more than one component to collectively perform the functions, each function need not be performed by each of those components (e.g., different functions may be performed by different components) and / or each function need not be performed in whole by only one component (e.g., different components may perform different sub-functions of a function) .

[0224] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described  functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0225] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general-purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium including program code including instructions that, when executed, performs one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include memory or data storage media, such as random-access memory (RAM) such as synchronous dynamic random-access memory (SDRAM) , read-only memory (ROM) , non-volatile random-access memory (NVRAM) , electrically erasable programmable read-only memory (EEPROM) , flash memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as propagated signals or waves.

[0226] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs) , general-purpose microprocessors, an application specific integrated circuits (ASICs) , field programmable logic arrays (FPGAs) , or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor, ” as used herein may refer to any of the foregoing  structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein.

[0227] Illustrative aspects of the disclosure include:

[0228] Aspect 1. An apparatus for tracking a human-machine-interface (HMI) device, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: extract image features from images of the HMI device captured by a camera of a head-mounted device (HMD) , the image features related to at least one of hands or arms of a user holding the HMI device; extract light-emitting diode (LED) features from the images, the LED features related to LEDs on the HMI device; track the image features and the LED features across multiple images; obtain motion data from the HMI device, the motion data measured by an inertial measurement unit (IMU) of the HMI device; determine motion-data preintegration based on the motion data; predict, based on the tracked image features, the tracked LED features, and the motion data, positions of the LEDs on the HMI device, and positions of the image features; and fuse the motion-data preintegration, the positions of the LEDs, and the positions of the image features to determine a pose of the HMI device.

[0229] Aspect 2. The apparatus of aspect 1, wherein, to fuse the motion data, LED positions of the LEDs, and the positions of the image features, the at least one processor is configured to: track and update the pose of the HMI device over time using a filter.

[0230] Aspect 3. The apparatus of aspect 2, wherein the filter comprises at least one of: an extended Kalman filter (EKF) an uncentered Kalman filter (UKF) , or a particle filter.

[0231] Aspect 4. The apparatus of any one of aspects 1 to 3, wherein, to fuse the motion data, the positions of the LEDs, and the positions of the image features, the at least one processor is configured to: determine a system state vector comprising errors of relative positions of image features in an IMU coordinate system.

[0232] Aspect 5. The apparatus of aspect 4, wherein the system state vector comprises a sliding window of the relative positions.

[0233] Aspect 6. The apparatus of any one of aspects 4 or 5, wherein the at least one processor is further configured to filter values of the system state vector.

[0234] Aspect 7. The apparatus of aspect 6, wherein the system state vector is filtered according to a Mahalanobis distance.

[0235] Aspect 8. The apparatus of any one of aspects 1 to 7, wherein the at least one processor is further configured to: propagate a system state vector and a covariance matrix with IMU preintegration; build a filter measurement model with LED measurements and feature measurements; remove outliers using a distance test; augment or marginalize the system state vector and the covariance matrix; update the filter measurement model to obtain system error state estimations and post covariance matrix; correct system states using the system error state estimations; maintain feature-motion data relationship; and marginalize an oldest pose state from a sliding window.

[0236] Aspect 9. An apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: propagate a system state vector and a covariance matrix with IMU preintegration; build a filter measurement model with LED measurements and feature measurements; remove outliers using a distance test; augment or marginalizing the system state vector and the covariance matrix; update the filter measurement model to decrease system error state estimations and post covariance matrix; correct system states using the system error state estimations; maintain feature-motion data relationship; and marginalize an oldest pose state from a sliding window.

[0237] Aspect 10. A method for tracking a human-machine-interface (HMI) device, the method comprising: extracting image features from images of the HMI device captured by a camera of a head-mounted device (HMD) , the image features related to at least one of hands or arms of a user holding the HMI device; extracting light-emitting diode (LED) features from the images, the LED features related to LEDs on the HMI device; tracking the image features and the LED features across multiple images; obtaining motion data from the HMI device, the motion data measured by an inertial measurement unit (IMU) of the HMI device; determining motion-data preintegration based on the motion data; predicting, based on the tracked image features, the tracked LED features, and the motion data, positions of the LEDs on the HMI device, and positions of the image features;  and fusing the motion-data preintegration, the positions of the LEDs, and the positions of the image features to determine a pose of the HMI device.

[0238] Aspect 11. The method of aspect 10, wherein to fuse the motion data, LED positions of the LEDs, and the positions of the image features comprises: tracking and updating the pose of the HMI device over time using a filter.

[0239] Aspect 12. The method of aspect 11, wherein the filter comprises at least one of: an extended Kalman filter (EKF) an uncentered Kalman filter (UKF) , or a particle filter..

[0240] Aspect 13. The method of any one of aspects 10 to 12, wherein to fuse the motion data, the positions of the LEDs, and the positions of the image features comprises: determining a system state vector comprising errors of relative positions of image features in an IMU coordinate system.

[0241] Aspect 14. The method of aspect 13, wherein the system state vector comprises a sliding window of the relative positions.

[0242] Aspect 15. The method of any one of aspects 13 or 14, further comprising filtering values of the system state vector.

[0243] Aspect 16. The method of aspect 15, wherein the system state vector is filtered according to a Mahalanobis distance.

[0244] Aspect 17. The method of any one of aspects 10 to 16, further comprising: propagating a system state vector and a covariance matrix with IMU preintegration; building a filter measurement model with LED measurements and feature measurements; removing outliers using a distance test; augmenting or marginalizing the system state vector and the covariance matrix; updating the filter measurement model to obtain system error state estimations and post covariance matrix; correcting system states using the system error state estimations; maintaining feature-motion data relationship; and marginalizing an oldest pose state from a sliding window.

[0245] Aspect 18. A method comprising: propagating a system state vector and a covariance matrix with IMU preintegration; building a filter measurement model with LED measurements and feature measurements; removing outliers using a distance test; augmenting or marginalizing the system state vector and the covariance matrix; updating the filter measurement model to  decrease system error state estimations and post covariance matrix; correcting system states using the system error state estimations; maintaining feature-motion data relationship; and marginalizing an oldest pose state from a sliding window.

[0246] Aspect 19. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to: extract image features from images of the HMI device captured by a camera of a head-mounted device (HMD) , the image features related to at least one of hands or arms of a user holding the HMI device; extract light-emitting diode (LED) features from the images, the LED features related to LEDs on the HMI device; track the image features and the LED features across multiple images; obtain motion data from the HMI device, the motion data measured by an inertial measurement unit (IMU) of the HMI device; determine motion-data preintegration based on the motion data; predict, based on the tracked image features, the tracked LED features, and the motion data, positions of the LEDs on the HMI device, and positions of the image features; and fuse the motion-data preintegration, the positions of the LEDs, and the positions of the image features to determine a pose of the HMI device.

[0247] Aspect 20. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to: propagate a system state vector and a covariance matrix with IMU preintegration; build a filter measurement model with LED measurements and feature measurements; remove outliers using a distance test; augment or marginalizing the system state vector and the covariance matrix; update the filter measurement model to decrease system error state estimations and post covariance matrix; correct system states using the system error state estimations; maintain feature-motion data relationship; and marginalize an oldest pose state from a sliding window.

[0248] Aspect 21. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to perform operations according to any of aspects 10 to 18.

[0249] Aspect 22. An apparatus for providing virtual content for display, the apparatus comprising one or more means for perform operations according to any of aspects 10 to 18.

[0250] Aspect 23. An apparatus for tracking a human-machine-interface (HMI) device, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: determine light-emitting diode (LED) -tracking measurements based on images of the HMI device, wherein the HMI device comprises LEDs; determine feature-tracking measurements based on the images of the HMI device; determine movement data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; and determine a pose of the HMI device by fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data.

[0251] Aspect 24. The apparatus of aspect 23, wherein, to fuse the LED-tracking measurements, the feature-tracking measurements, and the movement data, the at least one processor is configured to track and update the pose of the HMI device over a period of time using a filter.

[0252] Aspect 25. The apparatus of any one of aspects 23 or 24, wherein, to determine the pose of the HMI device, the at least one processor is configured to update a state of a filter based on the LED-tracking measurements, the feature-tracking measurements, and the movement data.

[0253] Aspect 26. The apparatus of aspect 25, wherein the filter comprises at least one of: an extended Kalman filter (EKF) , an uncentered Kalman filter (UKF) , or a particle filter.

[0254] Aspect 27. The apparatus of any one of aspects 23 to 26, wherein, to fuse the LED-tracking measurements, the feature-tracking measurements, and the movement data, the at least one processor is configured to determine a system-state vector comprising errors of relative positions of image features in an IMU coordinate system.

[0255] Aspect 28. The apparatus of aspect 27, wherein the system-state vector comprises a sliding window of the relative positions.

[0256] Aspect 29. The apparatus of any one of aspects 27 or 28, wherein the system-state vector is filtered according to a distance.

[0257] Aspect 30. The apparatus of any one of aspects 23 to 29, wherein, to determine the LED-tracking measurements, the at least one processor is configured to: identify LED features of the images of the HMI device; and track the LED features across the images of the HMI device.

[0258] Aspect 31. The apparatus of aspect 30, wherein, to determine the LED-tracking measurements, the at least one processor is configured to predict positions of the LED features.

[0259] Aspect 32. The apparatus of any one of aspects 23 to 31, wherein, to determine the feature-tracking measurements, the at least one processor is configured to: identify image features of the images of the HMI device; and track the image features across the images of the HMI device.

[0260] Aspect 33. The apparatus of aspect 32, wherein, to determine the feature-tracking measurements, the at least one processor is configured to predict positions of the image features.

[0261] Aspect 34. The apparatus of any one of aspects 32 or 33, wherein the image features are related to at least one of: surface features of the HMI device, a hand of a user holding the HMI device, or an arm of the user.

[0262] Aspect 35. The apparatus of any one of aspects 23 to 34, wherein, to determine the movement data, the at least one processor is configured to determine motion-data preintegration based on measurements of the IMU.

[0263] Aspect 36. The apparatus of any one of aspects 23 to 35, wherein, to determine the movement data, the at least one processor is configured to predict a pose of the HMI device based on measurements of the IMU.

[0264] Aspect 37. The apparatus of any one of aspects 23 to 36, wherein the LED-tracking measurements, the feature-tracking measurements, and the movement data are fused based on a camera pose of a camera used to capture the images of the HMI device.

[0265] Aspect 38. The apparatus of claim 16, wherein the apparatus comprises a head-mounted-device comprising the camera.

[0266] Aspect 39. A method for tracking a human-machine-interface (HMI) device, the method comprising: determining light-emitting diode (LED) -tracking measurements based on images of the HMI device, wherein the HMI device comprises LEDs; determining feature-tracking measurements based on the images of the HMI device; determining movement data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; and determining a pose  of the HMI device by fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data.

[0267] Aspect 40. The method of aspect 39, wherein fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data, comprises tracking and updating the pose of the HMI device over a period of time using a filter.

[0268] Aspect 41. The method of any one of aspects 39 or 40, wherein determining the pose of the HMI device comprises updating a state of a filter based on the LED-tracking measurements, the feature-tracking measurements, and the movement data.

[0269] Aspect 42. The method of aspect 41, wherein the filter comprises at least one of: an extended Kalman filter (EKF) , an uncentered Kalman filter (UKF) , or a particle filter.

[0270] Aspect 43. The method of any one of aspects 39 to 42, wherein fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data comprises determining a system-state vector comprising errors of relative positions of image features in an IMU coordinate system.

[0271] Aspect 44. The method of aspect 43, wherein the system-state vector comprises a sliding window of the relative positions.

[0272] Aspect 45. The method of any one of aspects 43 or 44, wherein the system-state vector is filtered according to a distance.

[0273] Aspect 46. The method of any one of aspects 39 to 45, wherein determining the LED-tracking measurements comprises: identifying LED features of the images of the HMI device; and tracking the LED features across the images of the HMI device.

[0274] Aspect 47. The method of aspect 46, wherein determining the LED-tracking measurements comprises predicting positions of the LED features.

[0275] Aspect 48. The method of any one of aspects 39 to 47, wherein determining the feature-tracking measurements comprises: identifying image features of the images of the HMI device; and tracking the image features across the images of the HMI device.

[0276] Aspect 49. The method of aspect 48, wherein determining the feature-tracking measurements comprises predicting positions of the image features.

[0277] Aspect 50. The method of any one of aspects 48 or 49, wherein the image features are related to at least one of: surface features of the HMI device, a hand of a user holding the HMI device, or an arm of the user.

[0278] Aspect 51. The method of any one of aspects 39 to 50, wherein determining the movement data comprises determining motion-data preintegration based on measurements of the IMU.

[0279] Aspect 52. The method of any one of aspects 39 to 51, wherein determining the movement data comprises predicting a pose of the HMI device based on measurements of the IMU.

[0280] Aspect 53. The method of any one of aspects 39 to 52, wherein the LED-tracking measurements, the feature-tracking measurements, and the movement data are fused based on a camera pose of a camera used to capture the images of the HMI device.

[0281] Aspect 54. The method of aspect 53, wherein the method is implemented by a head-mounted-device comprising the camera.

[0282] Aspect 55. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to perform operations according to any of aspects 39 to 54.

[0283] Aspect 56. An apparatus for providing virtual content for display, the apparatus comprising one or more means for perform operations according to any of aspects 39 to 54.

Claims

1.An apparatus for tracking a human-machine-interface (HMI) device, the apparatus comprising:at least one memory; andat least one processor coupled to the at least one memory and configured to:determine light-emitting diode (LED) -tracking measurements based on images of the HMI device, wherein the HMI device comprises LEDs;determine feature-tracking measurements based on the images of the HMI device;determine movement data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; anddetermine a pose of the HMI device by fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data.2.The apparatus of claim 1, wherein, to fuse the LED-tracking measurements, the feature-tracking measurements, and the movement data, the at least one processor is configured to track and update the pose of the HMI device over a period of time using a filter.3.The apparatus of claim 1, wherein, to determine the pose of the HMI device, the at least one processor is configured to update a state of a filter based on the LED-tracking measurements, the feature-tracking measurements, and the movement data.4.The apparatus of claim 3, wherein the filter comprises at least one of: an extended Kalman filter (EKF) , an uncentered Kalman filter (UKF) , or a particle filter.5.The apparatus of claim 1, wherein, to fuse the LED-tracking measurements, the feature-tracking measurements, and the movement data, the at least one processor is configured to determine a system-state vector comprising errors of relative positions of image features in an IMU coordinate system.6.The apparatus of claim 5, wherein the system-state vector comprises a sliding window of the relative positions.7.The apparatus of claim 5, wherein the system-state vector is filtered according to a distance.8.The apparatus of claim 1, wherein, to determine the LED-tracking measurements, the at least one processor is configured to:identify LED features of the images of the HMI device; andtrack the LED features across the images of the HMI device.9.The apparatus of claim 8, wherein, to determine the LED-tracking measurements, the at least one processor is configured to predict positions of the LED features.10.The apparatus of claim 1, wherein, to determine the feature-tracking measurements, the at least one processor is configured to:identify image features of the images of the HMI device; andtrack the image features across the images of the HMI device.11.The apparatus of claim 10, wherein, to determine the feature-tracking measurements, the at least one processor is configured to predict positions of the image features.12.The apparatus of claim 10, wherein the image features are related to at least one of: surface features of the HMI device, a hand of a user holding the HMI device, or an arm of the user.13.The apparatus of claim 1, wherein, to determine the movement data, the at least one processor is configured to determine motion-data preintegration based on measurements of the IMU.14.The apparatus of claim 1, wherein, to determine the movement data, the at least one processor is configured to predict a pose of the HMI device based on measurements of the IMU.15.The apparatus of claim 1, wherein the LED-tracking measurements, the feature-tracking measurements, and the movement data are fused based on a camera pose of a camera used to capture the images of the HMI device.16.The apparatus of claim 15, wherein the apparatus comprises a head-mounted-device comprising the camera.17.A method for tracking a human-machine-interface (HMI) device, the method comprising:determining light-emitting diode (LED) -tracking measurements based on images of the HMI device, wherein the HMI device comprises LEDs;determining feature-tracking measurements based on the images of the HMI device;determining movement data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; anddetermining a pose of the HMI device by fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data.18.The method of claim 17, wherein fusing the LED-tracking measurements, the feature-tracking measurements, and the movement data, comprises tracking and updating the pose of the HMI device over a period of time using a filter.19.The method of claim 17, wherein determining the pose of the HMI device comprises updating a state of a filter based on the LED-tracking measurements, the feature-tracking measurements, and the movement data.20.The method of claim 19, wherein the filter comprises at least one of: an extended Kalman filter (EKF) , an uncentered Kalman filter (UKF) , or a particle filter.

Citation Information

Patent Citations

  • Pose recovery of an ultrasound transducer

    CN107238396A

  • Pose estimation method of mobile robot and computer readable storage medium

    CN112815939A

  • Posture estimation method and device

    CN115690201A

  • Joint camera and inertial measurement unit calibration

    WO2022066486A1