Tracking devices for human-computer interaction

CN122804205APending Publication Date: 2026-09-22QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480087578.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-27
Filing Date
2024-11-13
Publication Date
2026-09-22

Smart Images

  • Figure CN122804205A_ABST
    Figure CN122804205A_ABST
Patent Text Reader

Abstract

Systems and techniques for tracking a human-machine interface (HMI) device are described. For example, a method can include determining a light-emitting diode (LED) tracking measurement based on an image of the HMI device, where the HMI device includes an LED; determining a feature tracking measurement based on the image of the HMI device; determining movement data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; and determining a pose of the HMI device by fusing the LED tracking measurement, the feature tracking measurement, and the movement data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit of PCT International Application No. PCT / CN2024 / 078685, filed on February 27, 2024, entitled “TRACKING APPARATUS FOR HUMAN-MACHINE INTERACTIONS”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates in general to human-computer interaction. More specifically, this disclosure relates to controller devices for tracking human-machine interface (HMI) devices. Background Technology

[0004] A handheld controller is an example of a human-machine interface (HMI) device. This handheld controller allows a user to interact with a machine by pressing or activating buttons on the handheld controller and / or by positioning and / or moving the handheld controller. The handheld controller can be used with extended reality (XR) systems. Summary of the Invention

[0005] The following is a simplified summary of the invention relating to one or more aspects disclosed herein. Therefore, this summary should not be considered an exhaustive overview relating to all conceived aspects, nor should it be considered to identify key or decisive elements relating to all conceived aspects or to depict the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts in a simplified form relating to one or more aspects of the mechanisms disclosed herein, preceding the detailed description that follows.

[0006] Systems and techniques for tracking human-machine interface (HMI) devices are described. According to at least one example, a method for tracking an HMI device is provided. The method includes: determining light-emitting diode (LED) tracking measurements based on an image of the HMI device, wherein the HMI device includes LEDs; determining feature tracking measurements based on the image of the HMI device; determining motion data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; and determining the attitude of the HMI device by fusing the LED tracking measurements, feature tracking measurements, and motion data.

[0007] In another example, an apparatus for tracking an HMI device is provided. The apparatus includes: at least one memory; and at least one processor (e.g., configured in a circuit) coupled to the at least one memory. The at least one processor is configured to: determine light-emitting diode (LED) tracking measurements based on an image of the HMI device, wherein the HMI device includes LEDs; determine feature tracking measurements based on an image of the HMI device; determine motion data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; and determine the attitude of the HMI device by fusing the LED tracking measurements, feature tracking measurements, and motion data.

[0008] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon, which, when executed by one or more processors, cause one or more processors to: determine light-emitting diode (LED) tracking measurements based on images from an HMI device, wherein the HMI device includes LEDs; determine feature tracking measurements based on images from the HMI device; determine motion data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; and determine the attitude of the HMI device by fusing the LED tracking measurements, feature tracking measurements, and motion data.

[0009] In another example, an apparatus for tracking an HMI device. The apparatus includes: components for extracting image features from images of the HMI device captured by a camera of a head-mounted device (HMD), these image features being related to at least one of the hand or arm of a user holding the HMI device; components for determining light-emitting diode (LED) tracking measurements based on the images of the HMI device, wherein the HMI device includes LEDs; components for determining feature tracking measurements based on the images of the HMI device; components for determining motion data of the HMI device based on the inertial measurement unit (IMU) of the HMI device; and components for determining the attitude of the HMI device by fusing the LED tracking measurements, feature tracking measurements, and motion data.

[0010] In some aspects, one or more of the devices described herein are, may be part of, or may include: extended reality (XR) devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), mobile devices (e.g., mobile phones or so-called "smartphones," tablet computers, or other types of mobile devices), smart or connected devices (e.g., Internet of Things (IoT) devices), wearable devices, personal computers, laptop computers, video servers, televisions (e.g., network-connected televisions), robotic devices or systems, vehicles (or computing devices or systems of vehicles), or other devices. In some aspects, each device may include one image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each device may include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each device may include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each device may include one or more sensors. In some cases, the one or more sensors may be used to determine the location of the device, the state of the device (e.g., tracking state, operating state, temperature, humidity level and / or another state) and / or for other purposes.

[0011] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim.

[0012] The foregoing and other features and aspects will become more apparent from the following description, claims and accompanying drawings. Attached Figure Description

[0013] The following description, with reference to the accompanying drawings, details exemplary examples of this application: Figure 1 This is an illustration of a conventional handheld controller including multiple light-emitting diodes arranged on a tracking ring protruding from the handle of the handheld controller; Figure 2A This is a diagram illustrating an example handheld controller including multiple light sources arranged on the body of the handheld controller according to various aspects of this disclosure; Figure 2B This is an illustration of a handheld controller being held by a user's hand, according to various aspects of this disclosure; Figure 3A These are illustrations of example systems, including handheld controllers and tracking systems, according to various aspects of this disclosure; Figure 3B This is a block diagram illustrating example elements of a handheld controller and a tracking system according to various aspects of this disclosure; Figure 4 This is a mixed block diagram / flowchart illustrating the operation of an example process 400 that can be used to track a controller according to various aspects of this disclosure; Figure 5A This is a flowchart illustrating an example process for performing a tightly coupled fusion of motion data and features according to various aspects of this disclosure; Figure 5B This is a flowchart illustrating an example process for performing a tightly coupled fusion of motion data and features according to various aspects of this disclosure; Figure 6A This is a flowchart illustrating another example process for tracking a human-machine interface (HMI) device according to various aspects of this disclosure; Figure 6B This is a flowchart illustrating another example process for tracking HMI devices according to various aspects of this disclosure; Figure 7 This is a block diagram illustrating examples of deep learning neural networks that can be used to perform various tasks, based on some aspects of the disclosed techniques; Figure 8 This is a block diagram illustrating examples of convolutional neural networks (CNNs) according to various aspects of this disclosure; and Figure 9 This is a block diagram illustrating an example computing device architecture that can implement the various technologies described herein. Detailed Implementation

[0014] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently, and some may be applied in combination, as will be apparent to those skilled in the art. Specific details are set forth in the following description for purposes of explanation in order to provide a thorough understanding of the various aspects of this application. However, it will be apparent that various aspects may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0015] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary aspects will provide those skilled in the art with descriptions that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0016] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as superior to or better than other aspects. Similarly, the term “aspects of this disclosure” does not require that all aspects of this disclosure include the features, advantages, or modes of operation discussed.

[0017] Object detection and tracking can be used to identify and track objects over time. For example, an image of an object can be obtained, and object detection can be performed on the image to detect the object in the image. In some cases, the detected objects can be classified into a category of objects. Furthermore, bounding boxes can be generated to identify the object's location in the image. Various types of systems, including neural network-based object detectors, can be used for object detection. The location of an object in an image can be tracked through a series of images. In some cases, tracking an object may include determining the pose of the object relative to a camera that captured an image of the object and / or the pose relative to the object's previous location. In this disclosure, the term "pose" can specify position and orientation. The pose can be determined based on six degrees of freedom, including three translational degrees of freedom (e.g., x, y, and z dimensions) and three rotational degrees of freedom (e.g., roll, pitch, and yaw).

[0018] Handheld controllers can be used as human-machine interface (HMI) devices. For example, an object tracking system can detect and track a handheld controller, and a user can interact with the machine by moving and / or rotating the handheld controller. For example, the machine can receive input based on how the user moves and / or rotates the handheld controller. Additionally, the controller may include one or more buttons that the user can press or activate. Indications of button presses or activation can be sent to the machine.

[0019] An example of a system that can be interacted with via a handheld controller is an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device). An example of an XR device is a head-mounted display (HMD) device that can be worn by a user (e.g., in the form of a head-mounted device, glasses, etc.). An XR device may include one or more cameras that can capture images of the handheld controller. The user can move and / or rotate the handheld controller. An XR device may include a tracking system that uses the cameras to track the handheld controller and interprets input based on how the user moves and / or rotates the handheld controller. Another example includes a gaming system with a camera mounted above a display.

[0020] Acyclic controllers are a type of computer vision (CV) based controller. Because acyclic CV-based controllers are more compact and portable, they are becoming increasingly common. However, tracking acyclic CV-based controllers (e.g., through CV-based object detection and / or tracking algorithms) is more difficult.

[0021] This document describes systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively, “Systems and Technologies”) for improving tracking in HMI devices. For example, the systems and technologies described herein may include novel multi-sensor fusion algorithms for acyclic, CV-based controllers. These novel fusion algorithms improve tracking accuracy and robustness, and also enhance the user experience.

[0022] Various aspects of this application will be described below with reference to the accompanying drawings.

[0023] Figure 1 This is an illustration of a conventional handheld controller 100 including a plurality of light-emitting diodes 106 (LEDs 106) arranged on a tracking ring 104 protruding from the handle 102 of the handheld controller 100. The handheld controller 100 includes a button 108, as an example of a button that may be present on the handle 102 of the handheld controller 100. Conventional tracking systems may rely on the image of the LEDs 106 on the tracking ring 104 to locate and track the handheld controller 100. The tracking ring 104 may be bulky, inconvenient to carry, and difficult to use.

[0024] To improve the performance of conventional tracking systems, the tracking ring 104 protrudes from the outside of the handheld controller 100. Due to the requirements of the tracking algorithm, the light-emitting diode (LED) tracking ring (e.g., tracking ring 104) must be large enough for the tracking algorithm to track accurately. A large tracking ring affects portability, making it more difficult to carry, such as in a pocket or backpack.

[0025] Furthermore, to avoid tracking failure caused by the LED tracking ring being blocked by a hand (e.g., the hand holding the handle 102), the LED tracking ring is typically designed to protrude beyond the position where the hand will hold the handle. However, this protruding structure is inconvenient to use, and the LED tracking ring may easily collide with other objects during use. For example, when a person's hands are holding the corresponding controller, the hands cannot make contact with each other due to the obstruction of the protruding part of the LED tracking ring.

[0026] Ring controllers (such as the handheld controller 100) have the following disadvantages: they are not compact and are not portable, they may encounter obstacles during use, and they are easily damaged. Additionally, when two ring controllers are used simultaneously, there are certain limitations regarding the cross-interaction between the two controllers.

[0027] To overcome the limitations of ring controllers, acyclic-free CV-based controllers are becoming increasingly popular. Acyclic-free CV-based controllers do not include an LED tracking ring. The tracking LEDs, which are placed on the ring of a ring controller, are located on the controller's panel and body. Acyclic-free CV-based controllers are more compact and portable compared to ring controllers. Additionally, acyclic-free CV-based controllers improve the user experience when users operate two controllers simultaneously.

[0028] Figure 2A This is a diagram illustrating an example handheld controller 200, including a plurality of light sources 202 arranged on the body 204 of the handheld controller 200, according to various aspects of this disclosure. Figure 1 Compared to the handheld controller 100, the light source 202 is not arranged on the tracking ring 104 protruding from the handle 102. Instead, the light source 202 is arranged on the body 204 (e.g., not protruding radially from the handle 206 of the body 204). The handheld controller 200 includes a button 208, as an example of a button that may be present on the body 204 of the handheld controller 200. Figure 2B This is an illustration of a handheld controller 200 held by a user's hand 210 according to various aspects of this disclosure. The size of the handheld controller 200 may be designed such that at least some of the light sources 202 are visible when the hand 210 holds the main body 204.

[0029] Because the tracking LEDs are located on the controller's panel and body, they are more likely to be blocked by fingers and palms when the controller is held in the hand. In extreme cases, all LEDs may be blocked and tracking may fail.

[0030] Some tracking methods include hand joint tracking, visual feature tracking on the controller, hand, and even arm. Hand joint tracking and visual feature tracking (including tracking features of the controller, hand, and / or arm) can be referred to as "tracking methods," "assisted visual feature tracking," or "feature tracking" to distinguish them from "LED tracking," which can refer to tracking LEDs in an image.

[0031] The challenges of tracking acyclic controllers include: LEDs are more likely to be occluded and more prone to loss of tracking, requiring alternative feature tracking methods to assist tracking, and the relationship between the controller and features of the hand or arm is unstable and may change over time and with different grip postures. In other words, the connection between the controller and features of the hand or arm is not rigid, which is very challenging for controller fusion algorithms.

[0032] This disclosure discloses a novel inertial measurement unit (IMU)-LED-feature tightly coupled fusion algorithm for addressing at least some of the challenges faced by acyclic controllers. As previously mentioned, feature tracking measurements can be fused in acyclic fusion algorithms in addition to IMU and LED tracking measurements. However, the relationship between the controller and features of the hand or arm is unstable. Therefore, a novel IMU-LED-feature tightly coupled fusion algorithm is disclosed to address this problem. In this disclosure, the term "IMU" can refer to an inertial measurement unit, data from the inertial measurement unit (e.g., motion data), and / or other motion data.

[0033] The fusion algorithm has at least the following aspects: a tightly coupled fusion framework that can simultaneously fuse IMU, LED tracking measurements and feature tracking measurements; a sliding window applied to the fusion state vector for simultaneously estimating the relationship between the controller IMU and the tracking features; and the use of a distance threshold (e.g., Mahalanobis distance test) method to detect changes in the relationship between the controller IMU and the tracking features and then reject outliers.

[0034] Figure 3A This is a diagram illustrating an example system 300 including a handheld controller 302 and a tracking system 320 according to various aspects of this disclosure. Figure 3B This is a block diagram illustrating example elements of a handheld controller 302 and a tracking system 320 of a system 300 according to various aspects of this disclosure. In some cases, the tracking system 320 may be part of or include an XR device (e.g., an HMD).

[0035] The handheld controller 302 includes a light-emitting diode (LED) 306 located on the body 304 of the handheld controller 302. The LED 306 can emit visible light, near-infrared light, and / or infrared light (of any color). The LED 306 may include groups of LEDs (e.g., each group includes red LEDs, green LEDs, and blue LEDs, such that each group of LEDs can change the wavelength of the emitted light). The LED 306 may be included on the handheld controller 302 to enable the tracking system 320 to track the handheld controller 302. Additionally, the handheld controller 302 includes at least one processor (e.g., processor 312) and a communication unit 314.

[0036] The tracking system 320 includes a camera 322 that captures images of a handheld controller 302, including images of LED 306 on the handheld controller 302 and images of the user's hand and / or arm. The tracking system 320 includes at least one processor (e.g., processor 326). The tracking system 320 can use the processor 326 to track the handheld controller 302 based on images of LED 306 captured by the camera 322. The tracking system 320 may also include a communication unit 324 that the tracking system 320 can use to communicate with the handheld controller 302 (via communication unit 314 of the handheld controller 302). For example, the handheld controller 302 may transmit status messages and / or motion data (e.g., based on measurements from an inertial measurement unit (IMU 308)) to the tracking system 320. The tracking system 320 can use the motion data while tracking the controller 302. The tracking system 320 may transmit control messages (e.g., commanding the handheld controller 302 to illuminate LED 306).

[0037] Figure 4 This is a mixed block diagram / flowchart illustrating the operation of example process 400, which can be used for a tracking controller according to various aspects of this disclosure. One or more operations of process 400 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a vehicle or a component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capability to perform process 400. One or more operations of process 400 may be implemented as software components that execute and run on one or more processors.

[0038] Process 400 can be made by Figure 3A and Figure 3B The tracking system 320 achieves tracking Figure 3A and Figure 3B A handheld controller 302. For example, the tracking system 320 can use... Figure 3A and Figure 3B The camera 322 captures images 402 of the handheld controller 302 (and / or the hand and arm of the user holding the handheld controller 302). Additionally, Figure 3B The handheld controller 302's IMU 308 can capture motion data 428, and via... Figure 3B The communication unit 314 provides motion data 428 to the tracking system 320.

[0039] Process 400 (or the system or device implementing process 400) may receive the following as inputs: image 402 (which may include images from multiple cameras, such as 2 or 4 cameras, such as...) Figure 3A and Figure 3B The camera 322), motion data 428 (e.g., from a handheld controller 302, which may include an IMU 308, which may include gyroscope data and / or accelerometer data), and attitude data for each camera ( Figure 4 (Not illustrated) (e.g., from a head tracking algorithm running in tracking system 320). Pose data from various cameras can be referred to as... , The timestamps of image 402 and motion data 428 should be synchronized by software or hardware. In some respects, at box 430, process 400 may predict the attitude of handheld controller 302 based on motion data 428. Additionally or alternatively, process 400 may determine motion data pre-integration 432 based on motion data 428.

[0040] Process 400 may generate (and / or output) attitude 436, i.e., the attitude of the controller (e.g., handheld controller 302). The controller's attitude data may be referred to as... , .

[0041] Image 402 is provided as an example image input to process 400. Image 402 may include multiple images captured at multiple times (e.g., sequentially) while tracking the controller over time. Image 402 is provided as an example of one such image to illustrate that process 400 occurs once on one of the multiple images.

[0042] Additionally or alternatively, image 402 may be or may include multiple images, such as, for example, short-exposure images and automatic exposure images. For example, image 402 may include an image captured with a short exposure duration that allows an LED (e.g., LED 306 of the handheld controller 302) to stand out in an image with an underexposed background and a hand that is too dark. Furthermore, image 402 may include an image captured according to an automatic exposure setting that exposes the hand and background so that they are visible.

[0043] At the extraction of LED spots 404 and image features 408, process 400 can process image 402. For example, process 400 can use conventional computer vision methods or deep learning to extract LED spots (LED features 406) and auxiliary visual features (image features 410) in parallel. The LEDs in image 402 are typically circular or elliptical bright spots. And image features 410 can be hand joints and visual feature points on the hand, arm, and / or controller.

[0044] At decision box 412, process 400 can check the tracking status. For example, process 400 can determine whether the controller is being tracked (based on the controller being detected and located in a previous image).

[0045] If the state is not in tracking mode (i.e., uninitialized or in lost mode), process 400 may proceed to search 424 and localization 426. At search 424 and localization 426, process 400 may perform initialization if uninitialized, or perform a relocation operation to restore the controller attitude if in lost mode. After initialization or relocation, the fusion module (e.g., fusion 434) may be initialized or reset.

[0046] If the state is in tracking mode (i.e., the previous controller pose is known), process 400 can proceed to tracker 414, which includes tracking LED 416 and tracking image features 420. At tracker 414, process 400 can predict the current controller pose based on the previous pose using motion data integration (e.g., based on motion data 428), and then predict the localization of the LEDs and features in the current image. Therefore, process 400 can search for and match corresponding LED spots and features near the predicted localizations in the current image. This effectively reduces computation time and improves tracking robustness.

[0047] After tracking the LEDs and features, all tracking measurements (e.g., LED tracking measurement 418 and image tracking measurement 422), motion data pre-integration (motion data pre-integration 432), and camera pose ( , This will be passed to the tightly coupled fusion module (e.g., fusion 434). The output of the fusion module (e.g., fusion 434) is the controller pose at the current image timestamp (e.g., pose 436).

[0048] Figure 5A This is a flowchart illustrating an example process 500 for performing IMU-LED-feature tight coupling fusion according to various aspects of this disclosure. One or more operations of process 500 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a vehicle or a component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capability to perform process 500. One or more operations of process 500 may be implemented as software components that execute and run on one or more processors. For example, Figure 3A and Figure 3B The tracking system 320 can execute process 500. In some respects, process 500 can be combined with... Figure 4 The integration of 434 execution.

[0049] Based on the degree of information fusion, fusion methods can be classified into loosely coupled and tightly coupled methods. Loosely coupled methods fuse estimated pose from visual information (LED features and image features) and IMU information, and the measurement is the pose solved from the LEDs and features. Generally, at least three LEDs / feature points are required to solve for the controller's pose. If fewer than three LEDs / feature points are available, the controller pose cannot be solved from visual information; therefore, loosely coupled methods cannot fuse visual information with fewer than three LEDs / feature points.

[0050] Tightly coupled methods directly use successfully matched LED spots / features on the image as observation measurements, resulting in a tighter degree of information coupling. Compared to loosely coupled methods, tightly coupled methods can also fuse visual information with fewer than three LEDs / feature points, and even with just one point. Therefore, tightly coupled methods require fewer LEDs / feature points and offer higher tracking robustness, especially in the presence of occlusion.

[0051] At a higher level, process 500 includes: using IMU pre-integration to propagate the system state vector and covariance matrix; using LED measurements and feature measurements to construct a filter measurement model (e.g., an extended Kalman filter (EKF) measurement model); using a threshold (e.g., a Mahalanobis distance test) to remove outliers; augmenting and / or marginalizing the system state vector and covariance matrix; performing a filter update (e.g., an EKF filter update) to obtain the optimal system error state estimate and posterior covariance matrix; using the estimated system error state to correct the system state; maintaining the feature / IMU relationship; and marginalizing the oldest attitude state from a sliding window.

[0052] Before describing each module of the fusion algorithm in detail, we introduce the coordinate system and the definition of mathematical symbols.

[0053] Coordinate definitions: W: World coordinate system. C: Camera coordinate system. I: Controller IMU coordinate system.

[0054] R: Transformation (e.g., rotation) matrix. For example, : The rotation matrix from coordinate system A to coordinate system B. Positioning of one coordinate system within another. For example, Positioning of coordinate system A in coordinate system B.

[0055] Error definition: ,in It is the baseline truth value. yes The estimated value, and yes The error.

[0056] Process 500 can use a "state". For example, process 500 can propagate and update the state to track it. This state can be referred to as the "system state". At each time step k, the fused system can maintain the following error state vector:

[0057] in It is the error state vector of IMU-feature relative localization. in ,and for , Representation of features Error in relative positioning in the IMU coordinate system {I}.

[0058] In addition, for , Let M represent the controller IMU pose at time step i, where M is the sliding window size. Each controller IMU pose within the sliding window is defined as follows:

[0059] in From the IMU coordinate system { The error state vector of the rotation from {W} to the world coordinate system {W}. in yes{ The positioning error state in {W}.

[0060] This refers to additional error states related to IMU bias and speed:

[0061] in and These are the bias errors of the gyroscope and accelerometer, respectively. It is the velocity error at time step k.

[0062] As previously mentioned, process 500 can, for example, propagate the state at box 502. The IMU pre-integration recursive formula can be:

[0063] in , , This is the IMU pre-integration result; It is the rotation increment from time step k-1 to time step k; It is the velocity increment from time step k-1 to time step k; and It is the positioning increment from time step k-1 to time step k.

[0064] The IMU error state vector is defined as:

[0065] and The propagation can be expressed as:

[0066] in It is system noise.

[0067] covariance matrix It depends on the IMU noise characteristics and is calculated during sensor calibration.

[0068]

[0069] Therefore, the state vector of the entire system The propagation can be expressed as:

[0070] At box 504, process 500 can construct a measurement model. The general linearized form of the measurement model (which can be an EKF measurement model) can be:

[0071] in H is the measurement residual, and H is the measurement Jacobian matrix. Furthermore, the noise term is zero-mean, Gaussian distributed, and related to the error state. Irrelevant.

[0072] For IMU / LED / feature fusion, measurements include LED spot measurement and feature measurement.

[0073] Measurement residual of the j-th LED spot:

[0074] in This is the measurement of the j-th LED spot. It is the estimated measurement of the j-th LED spot.

[0075] It is the controller attitude state at time step k. The Jacobian matrix of the j-th LED spot is measured.

[0076] And image projection function Defined as:

[0077] The j-th feature measurement residual

[0078] in It is the j-th feature measurement. It is the estimated j-th feature measurement.

[0079] It is a feature-IMU relative positioning The j-th feature measurement Jacobian matrix, It is the controller attitude state at time step k. The j-th feature measurement Jacobian matrix.

[0080] By stacking all LED spots and feature measurement residuals together, we can obtain:

[0081] At box 506, procedure 500 removes outliers. The relationship between features and the IMU is unstable; relative positioning tends to change. Therefore, an outlier rejection procedure can be applied before feature measurement updates. The fusion system uses a Mahalanobis distance test as an example.

[0082]

[0083] in Represents Mahalanobis distance, It measures the residual. It is the residual covariance. H is the measurement Jacobian, P is the covariance matrix, and It measures the standard deviation of noise.

[0084] All these values ​​can be obtained from a filter (e.g., an EKF filter). If γ is greater than a certain threshold, the feature will be considered an outlier.

[0085] At box 508, process 500 may augment and / or marginalize feature states. For example, if an old feature already in the state vector becomes an outlier, the state vector and covariance matrix should be marginalized. If a new feature is added, the state vector and covariance matrix should be augmented.

[0086] At box 510, process 500 can update the filter (e.g., an EKF filter). For example, at this point in process 500, all states can be propagated and all measurements are ready. The filter can then be updated. For example, the Kalman gain could be:

[0087] in It is the measurement noise covariance matrix.

[0088] The estimated error state can be:

[0089] Finally, the state covariance matrix is ​​updated according to the following:

[0090] At box 512, this state can be corrected. Furthermore, at box 512, the feature-motion data relationship can be maintained. For example, after updating the filter, an estimated error state can be obtained. This error can be used to correct the state. To correct the relationship between the j-th feature and the IMU:

[0091] To correct the i-th pose in the sliding window:

[0092] Correcting additional states:

[0093] At box 516, the controller pose state can be edged out. For example, the oldest pose state should be edged out from the sliding window. .

[0094] At box 514, the feature-motion data relationship can be maintained.

[0095] Inputs to process 500 may include: Prior error state estimation Prior covariance matrix IMU pre-integral increment , , Camera pose: , LED spot measurement: Feature measurement:

[0096] Process 500 may include at least the following steps: Propagation status Build a measurement model Remove outliers Augmentation / Marginalization State EKF Update Correcting the state and maintaining the feature / IMU relationship Marginalized state This disclosure discloses a fusion algorithm for acyclic CV-based controllers, including a novel motion data-LED-feature tightly coupled fusion algorithm for solving.

[0097] This disclosure discloses a tightly coupled fusion framework that can simultaneously fuse motion data, LED tracking measurements, and feature tracking measurements. Additionally, this disclosure discloses a sliding window applied to the fused state vector for simultaneously estimating the relationship between the controller IMU and tracking features. Furthermore, this disclosure discloses employing a distance test (e.g., Mahalanobis distance test) to detect changes in the relationship between the controller IMU data and the tracking features, and then rejecting outliers.

[0098] Figure 5BThis is a flowchart illustrating a process 550 for tracking a human-machine interface (HMI) device according to various aspects of this disclosure. One or more operations of process 550 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable device such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a vehicle or a component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capability to perform process 550. One or more operations of process 550 may be implemented as software components that execute and run on one or more processors.

[0099] At box 552, the computing device (or one or more components thereof) may utilize IMU pre-integration to propagate the system state vector and covariance matrix. For example, Figure 3A and Figure 3B The tracking system 320 can utilize IMU pre-integration to propagate the system state vector and covariance matrix.

[0100] At box 554, a computing device (or one or more components thereof) may use LED measurements and characteristic measurements to construct a filter measurement model. For example, filter system 320 may use LED measurements and characteristic measurements to construct a filter measurement model.

[0101] At box 556, the computing device (or one or more components thereof) may use a distance check to remove outliers. For example, filter system 320 may use a distance check to remove outliers.

[0102] At box 558, a computing device (or one or more components thereof) may augment or marginalize the system state vector and covariance matrix. For example, filter system 320 may augment or marginalize the system state vector and covariance matrix.

[0103] At box 560, a computing device (or one or more components thereof) may update the filter measurement model to reduce the system error state estimate and posterior covariance matrix. For example, filter system 320 may update the filter measurement model to reduce the system error state estimate and posterior covariance matrix.

[0104] At box 562, a computing device (or one or more components thereof) may use a system error state estimate to correct the system state. For example, a filter system 320 may use a system error state estimate to correct the system state.

[0105] At box 564, a computing device (or one or more components thereof) may maintain feature-motion data relationships. For example, a filter system 320 may maintain feature-motion data relationships.

[0106] At box 566, the computing device (or one or more components thereof) may edge out the oldest pose state from a sliding window. For example, filter system 320 may edge out the oldest pose state from a sliding window.

[0107] Figure 6A This is a flowchart illustrating a process 600 for tracking a human-machine interface (HMI) device according to various aspects of this disclosure. One or more operations of process 600 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable device such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a vehicle or a component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device having the resource capability to perform one or more operations of process 600. One or more operations of process 600 may be implemented as software components that execute and run on one or more processors.

[0108] At box 602, the computing device (or one or more components thereof) can extract image features from images captured by a camera of the head-mounted device (HMD) of the HMI device, these image features being associated with at least one of the hand or arm of the user holding the HMI device. For example, Figure 3A and Figure 3B The tracking system 320 can acquire (e.g., capture using a camera of the tracking system 320) images of the handheld controller 302, the hand 310 of the user holding the handheld controller 302, and / or the arm (e.g., Figure 4 Image 402). The tracking system 320 can extract image features from the image (e.g., Figure 4 Image features 410).

[0109] At box 604, the computing device (or one or more components thereof) can extract light-emitting diode (LED) features from the image, which are associated with LEDs on the HMI device. For example, tracking system 320 can extract LED features from the image (e.g., Figure 4 LED feature 406).

[0110] At box 606, the computing device (or one or more components thereof) can track image features and LED features across multiple images. For example, tracking system 320 can track image features (e.g., image feature 410) and LED features (e.g., LED feature 406) across multiple images (e.g., multiple images of the handheld controller 302, the hand 310 of the holder of the handheld controller 302, and / or the arm of the holder, captured by, for example, tracking system 320).

[0111] At box 608, a computing device (or one or more components thereof) can obtain motion data from an HMI device, which is measured by the HMI device's inertial measurement unit (IMU). For example, tracking system 320 can obtain motion data from a handheld controller 302. The motion data can be obtained by... Figure 3B The handheld controller 302 measures the IMU 308.

[0112] At box 610, the computing device (or one or more components thereof) may determine motion data pre-integration based on motion data. For example, tracking system 320 may determine motion data pre-integration based on motion data (e.g., motion data 428). Figure 4 Motion data pre-integration 432).

[0113] At box 612, the computing device (or one or more components thereof) may predict the location of the LEDs on the HMI device and the location of the image features based on the tracked image features, the tracked LED features, and motion data. For example, the tracking system 320 may predict LED tracking measurement 418 and image tracking measurement 422 based on the tracking of LED features 406, the tracking of image features 410, motion data 428, and / or motion data pre-integration 432.

[0114] At box 614, the computing device (or one or more components thereof) may fuse motion data pre-integration, LED positioning, and image feature positioning to determine the attitude of the HMI device. For example, tracking system 320 may fuse motion data pre-integration 432, LED tracking measurement 418, and image tracking measurement 422 to generate attitude 436.

[0115] In some aspects, to fuse motion data, LED positioning of LEDs, and positioning of image features, a computing device (or one or more components thereof) may use filters to track and update the attitude of an HMI device over time. For example, tracking system 320 may use filters to track and update the attitude of handheld controller 302 over time. Filters may be or may include algorithms for estimating the state of a dynamic system through observation-based sequential update estimation. In some aspects, filters may be or may include at least one of the following: extended Kalman filter (EKF), unscented Kalman filter (UKF), or particle filter.

[0116] In some aspects, to fuse motion data, LED positioning, and image feature positioning, a computing device (or one or more components thereof) may determine a system state vector that includes the error in the relative positioning of the image features in the IMU coordinate system. For example, tracking system 320 may determine a system state vector that includes the error in the relative positioning of the image features in the IMU coordinate system. In some aspects, the system state vector may be or may include a sliding window of relative positioning.

[0117] In some respects, the computing device (or one or more components thereof) can filter the values ​​of the system state vector. In some respects, the system state vector can be filtered according to Mahalanobis distance.

[0118] Figure 6B This is a flowchart illustrating process 620 for tracking a human-machine interface (HMI) device according to various aspects of this disclosure. One or more operations of process 620 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable device such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a vehicle or a component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device having the resource capability to perform one or more operations of process 620. One or more operations of process 620 may be implemented as software components that execute and run on one or more processors.

[0119] At box 622, the computing device (or one or more components thereof) can determine light-emitting diode (LED) tracking measurements based on an image from an HMI device, wherein the HMI device includes LEDs. For example, Figure 3A and Figure 3B The tracking system 320 can acquire (e.g., capture using a camera of the tracking system 320) images of the handheld controller 302, the hand 310 of the user holding the handheld controller 302, and / or the arm (e.g., Figure 4 Image 402). Tracking system 320 can extract LED features from the image (e.g., Figure 4 (LED feature 406). The tracking system 320 can determine LED tracking measurements based on image 402 and / or based on LED feature 406.

[0120] In some respects, to determine LED tracking measurements, a computing device (or one or more components thereof) may: identify LED features in an image of the HMI device; and track LED features across the image of the HMI device. For example, tracking system 320 may extract LED features from image 402 (e.g., Figure 4 LED features 406), and LED features are tracked across instances of image 402.

[0121] In some respects, in order to determine LED tracking measurements, a computing device (or one or more components thereof) can predict the location of LED features. For example, tracking system 320 can predict the location of LED feature 406 in an upcoming instance of image 402.

[0122] At box 624, the computing device (or one or more components thereof) may determine feature tracking measurements based on an image of the HMI device. For example, tracking system 320 may acquire an image 402 of a handheld controller 302, a user's hand 310 holding the handheld controller 302, and / or an image of their arm. Tracking system 320 may extract image features from the image (e.g., Figure 4 Image features 410). The tracking system 320 can determine feature tracking measurements based on image 402 or image features 410.

[0123] In some aspects, to determine feature tracking measurements, a computing device (or one or more components thereof) may: identify image features of an image from an HMI device; and track image features across the HMI device. For example, tracking system 320 may extract image features from image 402 (e.g., Figure 4 Image features 410), and image features are tracked across instances of image 402.

[0124] In some respects, in order to determine feature tracking measurements, a computing device (or one or more components thereof) can predict the localization of image features. For example, tracking system 320 can predict the localization of image feature 410 in an upcoming instance of image 402.

[0125] In some respects, the image features may be associated with at least one of the following: surface features of the HMI device, the hand of the user holding the HMI device, or the user's arm. For example, image feature 410 may be associated with surface features of the handheld controller 302, the hand of the user holding the handheld controller 302, or the user's arm.

[0126] At box 626, the computing device (or one or more components thereof) can determine the motion data of the HMI device based on the inertial measurement unit (IMU) of the HMI device. For example, the tracking system 320 can obtain motion data from the handheld controller 302. The motion data can be obtained by... Figure 3BThe handheld controller 302 uses the IMU 308 for measurement. The tracking system 320 can determine motion data pre-integration (e.g., motion data 428) based on the motion data. Figure 4 Motion data pre-integration 432).

[0127] In some respects, to determine motion data, a computing device (or one or more components thereof) may determine motion data pre-integration based on IMU measurements. For example, tracking system 320 may determine motion data pre-integration 432 based on motion data 428.

[0128] In some respects, to determine motion data, a computing device (or one or more components thereof) can predict the attitude of an HMI device based on IMU measurements. For example, at block 430, tracking system 320 can predict the attitude of handheld controller 302 based on motion data 428.

[0129] At box 628, the computing device (or one or more components thereof) can determine the attitude of the HMI device by fusing LED tracking measurements, feature tracking measurements, and motion data. For example, tracking system 320 can fuse motion data pre-integration 432, LED tracking measurements 418, and image tracking measurements 422 to generate attitude 436.

[0130] In some aspects, to integrate LED tracking measurements, feature tracking measurements, and motion data, a computing device (or one or more components thereof) may use a filter to track and update the attitude of an HMI device over a time period. For example, tracking system 320 may use a filter to track and update the attitude of handheld controller 302 over time. The filter may be or may include an algorithm for estimating the state of a dynamic system through observation-based sequential update estimation. In some aspects, the filter may be or may include at least one of the following: an extended Kalman filter (EKF), an unscented Kalman filter (UKF), or a particle filter.

[0131] In some respects, to determine the attitude of an HMI device, a computing device (or one or more components thereof) can update the state of a filter based on LED tracking measurements, feature tracking measurements, and motion data. For example, the tracking system 320 can... Figure 5A The filter state is updated at frame 510 in process 500.

[0132] In some respects, the filter may be or may include at least one of the following: an extended Kalman filter (EKF), an unscented Kalman filter (UKF), or a particle filter. For example, the filter updated at block 510 may be or may include an EKF, a UKF, or a particle filter.

[0133] In some respects, to integrate LED tracking measurements, feature tracking measurements, and motion data, a computing device (or one or more components thereof) can determine a system state vector that includes the error in the relative positioning of image features in the IMU coordinate system. For example, tracking system 320 can determine a system state vector that includes the error in the positioning of image features in the coordinate system of handheld controller 302.

[0134] In some respects, the system state vector may be or may include a sliding window of relative positioning. For example, the tracking system 320 may maintain a sliding window of positioning for the handheld controller 302.

[0135] In some respects, the system state vector is filtered based on distance. For example, the tracking system 320 may filter the system state vector based on distance.

[0136] In some respects, the computing device (or one or more components thereof) can fuse LED tracking measurements, feature tracking measurements, and motion data based on the camera pose of the camera used to capture images of the HMI device. For example, the tracking system 320 can be based at least in part on pose data from various cameras (which may be referred to as...) , This integrates LED tracking measurements, feature tracking measurements, and motion data.

[0137] In some aspects, the computing device (or one or more components thereof) may be the computing device of the device. The device may include a head-mounted display and a camera. For example, the computing device (or one or more components thereof) may be or may include tracking system 320. Tracking system 320 may include a display at camera 322.

[0138] In some examples, as previously noted, the methods described herein (e.g., Figure 4 Process 400 Figure 5A Process 500 Figure 5B Process 550 Figure 6A Process 600 Figure 6B The process 620 and / or other methods described herein may be performed wholly or partially by a computing device or apparatus. In one example, one or more of these methods may be performed by... Figure 3A and Figure 3B The tracking system 320 may be used or performed by another system or device. In another example, one or more of these methods (e.g., process 400, process 500, process 550, process 600, process 620 and / or other methods described herein) may be used by Figure 9 The computing device architecture 900 shown is implemented wholly or partially. For example, it has Figure 9The computing device of the illustrated computing device architecture 900 may include or be included in the tracking system 320 and may implement the operation of processes 400, 500, 550, 600, 620 and / or other processes described herein. In some cases, the computing device or apparatus may include various components such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other types of data.

[0139] A component capable of implementing a computing device in a circuit. For example, the component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0140] Processes 400, 500, 550, 600, 620, and / or other processes described herein are illustrated as logic flowcharts, whose operations represent sequences of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement a process.

[0141] Additionally, processes 400, 500, 550, 600, 620, and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, by hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0142] As noted above, various aspects of this disclosure may utilize machine learning models or systems.

[0143] Figure 7 This is an exemplary example of a neural network 700 (e.g., a deep learning neural network) that can be used to implement machine learning-based feature segmentation, implicit neural representation generation, rendering, classification, object detection, image recognition (e.g., face recognition, object recognition, scene recognition, etc.), feature extraction, authentication, gaze detection, gaze prediction, and / or automation. For example, the neural network 700 could be... Figure 4 Examples of feature extraction related to LED spot extraction 404 and / or image feature extraction 408, or the feature extraction can be performed.

[0144] Input layer 702 includes input data. In an exemplary example, input layer 702 may include representations... Figure 4 The data for image 402. The neural network 700 includes multiple hidden layers, such as hidden layers 706a, 706b to 706n. Hidden layers 706a, 706b to 706n include "n" hidden layers, where "n" is an integer greater than or equal to one. Multiple hidden layers can be made to include as many layers as needed for a given application. The neural network 700 further includes an output layer 704 that provides the output produced by the processing performed by hidden layers 706a, 706b to 706n. In an exemplary example, output layer 704 provides... Figure 4 LED feature 406 and / or image feature 410.

[0145] Neural network 700 may be or may include a multi-layer neural network with interconnected nodes. Each node may represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains the information while processing it. In some cases, neural network 700 may include a feedforward network, in which case there are no feedback connections in which the network's output is fed back into itself. In some cases, neural network 700 may include a recurrent neural network, which may have loops that allow information to be carried across nodes when reading from the input.

[0146] Information can be exchanged between nodes through node-to-node interconnects between layers. Nodes in input layer 702 can activate the node set in the first hidden layer 706a. For example, as shown, each input node in input layer 702 is connected to each node in the first hidden layer 706a. Nodes in the first hidden layer 706a can transform the information of each input node by applying an activation function to the input node information. The information derived from this transformation can then be passed to nodes in the next hidden layer 706b, activating those nodes, which can then perform their own specified functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable function. The output of hidden layer 706b can then activate nodes in the next hidden layer, and so on. The output of the last hidden layer 706n can activate one or more nodes in output layer 704, at which the output is provided. In some cases, although a node in neural network 700 (e.g., node 708) is shown to have multiple output lines, the node has a single output, and all lines shown as outputs from the node represent the same output value.

[0147] In some cases, each node or the interconnection between nodes may have weights, which are a set of parameters derived from the training of the neural network 700. Once the neural network 700 is trained, it can be referred to as a trained neural network, which can be used to perform one or more operations. For example, the interconnection between nodes may represent a piece of information about what the interconnected nodes have learned. This interconnection may have tunable numerical weights that can be tuned (e.g., based on the training dataset), allowing the neural network 700 to adapt to the input and learn as more and more data is processed.

[0148] The neural network 700 can be pre-trained to process features from the data in the input layer 702 using different hidden layers 706a, 706b to 706n, so as to provide an output through the output layer 704. In an example where the neural network 700 is used to identify features in an image, the neural network 700 can be trained using training data that includes both images and labels, as described above. For example, training images can be input into the network, where each training image has a label indicating features in the image (for feature segmentation machine learning systems) or a label indicating the category of activity in each image. In an example where object classification is used for illustrative purposes, the training images may include images of the number 2, in which case the label of the image may be [0 0 1 0 0 0 0 0 0 0].

[0149] In some cases, the neural network 700 can use a training process called backpropagation to adjust the weights of its nodes. As noted above, the backpropagation process can include forward pass, loss function, back pass, and weight update. For each training iteration, forward pass, loss function, back pass, and parameter update are performed. For each set of training images, this process can be repeated up to a certain number of iterations until the neural network 700 is trained well enough to accurately tune the weights of each layer.

[0150] For an example of identifying objects in an image, the forward pass may include passing a training image through a neural network 700. The weights are initially randomized before training the neural network 700. As an illustrative example, the image may include a numerical array representing the pixels of the image. Each number in the array may include a value from 0 to 255 describing the intensity of the pixel at that location in the array. In one example, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or lightness and two chroma components, etc.).

[0151] As noted above, for the first training iteration of the neural network 700, the output will likely include values ​​due to the weights being randomly selected during initialization without prioritizing any particular class. For example, if the output is a vector with probabilities that an object includes different classes, the probability values ​​for each class may be equal or at least very similar (e.g., for ten possible classes, each class may have a probability value of 0.1). Using the initial weights, the neural network 700 cannot determine low-level features and therefore cannot make an accurate determination of what the object's classification might be. A loss function can be used to analyze the error in the output. Any suitable loss function can be defined, such as cross-entropy loss. Another example of a loss function is mean squared error (MSE), defined as... The loss can be set to equal E. 总计 The value of .

[0152] For the first training image, the loss (or error) will be high because the actual value will be significantly different from the predicted output. The goal of training is to minimize the loss so that the predicted output matches the training labels. The Neural Network 700 performs backpropagation by determining which inputs (weights) contribute most to the network's loss and can adjust the weights to reduce and eventually minimize the loss. The derivative of the loss with respect to the weights (denoted as...) can be calculated. dL / dW ,in W These are the weights at a specific layer, used to determine the weights that contribute the most to the network's loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, weights can be updated so that they change in the opposite direction of the gradient. A weight update can be represented as... ,in w Indicates weight, w i Let represent the initial weights, and η represent the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes larger weight updates, while a lower value indicates smaller weight updates.

[0153] Neural Network 700 can include any suitable deep network. An example includes a Convolutional Neural Network (CNN), which includes an input layer and an output layer, with multiple hidden layers between them. The hidden layers of a CNN include a series of convolutional layers, non-linear layers, pooling layers (for downsampling), and fully connected layers. Neural Network 700 can include any other deep network besides CNNs, such as autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), etc.

[0154] Figure 8 This is an exemplary example of a Convolutional Neural Network (CNN) 800. The input layer 802 of the CNN 800 includes data representing an image or frame. For example, the data may include a numerical array representing pixels of an image, where each number in the array includes a value from 0 to 255 describing the pixel intensity at that location in the array. Using the previous example from above, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or lightness and two chroma components, etc.). The image may be passed through a convolutional hidden layer 804, an optional non-linear activation layer, a pooling hidden layer 806, and a fully connected layer 808 (which may be hidden) to obtain an output at the output layer 810. Although Figure 8Only one hidden layer from each hidden layer is shown in the diagram, but those skilled in the art will understand that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers may be included in a CNN 800. As previously described, the output may indicate a single category of an object, or may include probabilities that best describe the category of an object in an image.

[0155] The first layer of a CNN 800 can be a convolutional hidden layer 804. The convolutional hidden layer 804 analyzes the image data from the input layer 802. Each node in the convolutional hidden layer 804 is connected to a region of the input image called a receptive field (pixel). The convolutional hidden layer 804 can be thought of as one or more filters (each filter corresponding to a different activation or feature map), where each convolutional iteration of the filter is a node or neuron in the convolutional hidden layer 804. For example, the region of the input image covered by the filter at each convolutional iteration will be the receptive field of the filter. In an exemplary example, if the input image consists of a 28×28 array and each filter (and its corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in the convolutional hidden layer 804. Each connection between a node and its receptive field learns weights and, in some cases, learns an overall bias, such that each node learns to analyze its specific local receptive field in the input image. Each node in the convolutional hidden layer 804 will have the same weights and biases (called shared weights and shared biases). For example, the filter has a weight (digital) array and the same depth as the input. For the image frame example, the filter would have a depth of 3 (based on the three color components of the input image). An exemplary example of the filter array size is 5×5×3, corresponding to the size of the receptive field of the node.

[0156] The convolutional property of the convolutional hidden layer 804 is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filters of the convolutional hidden layer 804 may begin at the top left corner of the input image array and may convolve around the input image. As noted above, each convolutional iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 804. In each convolutional iteration, the value of the filter is multiplied by the corresponding number of original pixel values ​​of the image (e.g., a 5×5 filter array is multiplied by a 5×5 array of input pixel values ​​at the top left corner of the input image array). The multiplications from each convolutional iteration can be summed to obtain the sum of that iteration or node. Next, the process continues at the next position in the input image based on the receptive field of the next node in the convolutional hidden layer 804. For example, the filter may move a step size (called stride) to the next receptive field. The stride may be set to 1 or any other suitable amount. For example, if the stride is set to 1, the filter will move 1 pixel to the right in each convolutional iteration. Processing the filter at each unique location in the input volume produces a number representing the filter result at that location, thus determining a sum value for each node of the convolutional hidden layer 804.

[0157] The mapping from the input layer to the convolutional hidden layer 804 is called an activation map (or feature map). An activation map includes values ​​for each node representing the filter results at each location in the input volume. Activation maps can include arrays containing various sums of values ​​produced by the filter on each iteration of the input volume. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map would consist of a 24×24 array. The convolutional hidden layer 804 can include several activation maps to identify multiple features in the image. Figure 8 The example shown includes three activation maps. Using these three activation maps, the convolutional hidden layer 804 can detect three different types of features, each of which is detectable across the entire image.

[0158] In some examples, nonlinear hidden layers can be applied after convolutional hidden layers 804. Nonlinear layers can be used to introduce nonlinearity into a system that has already computed linear operations. An exemplary example of a nonlinear layer is the Corrected Linear Unit (ReLU) layer. A ReLU layer applies the function f(x) = max(0, x) to all values ​​in the input volume, which changes all negative activations to 0. Therefore, ReLU can add nonlinearity to CNN 800 without affecting the receptive field of convolutional hidden layers 804.

[0159] A pooling hidden layer 806 can be applied after the convolutional hidden layer 804 (and, in use, after a non-linear hidden layer). The pooling hidden layer 806 is used to simplify the information in the output of the convolutional hidden layer 804. For example, the pooling hidden layer 806 takes each activation map output from the convolutional hidden layer 804 and uses a pooling function to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by the pooling hidden layer. The pooling hidden layer 806 uses other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. Pooling functions (e.g., max pooling filters, L2 norm filters, or other suitable pooling filters) are applied to each activation map included in the convolutional hidden layer 804. Figure 8 In the example shown, three pooling filters are used to convolve the three activation maps in the hidden layer 804.

[0160] In some examples, max pooling can be used by applying a max pooling filter (e.g., of size 2×2) with a stride (e.g., equal to the dimension of the filter, such as stride 2) to the activation map output from convolutional hidden layer 804. The output from the max pooling filter includes the maximum number in each sub-region of the filter convolution. Using a 2×2 filter as an example, each unit in the pooling layer summarizes a region of 2×2 nodes from the previous layer (each node is a value in the activation map). For example, four values ​​(nodes) in the activation map will be analyzed by the 2×2 max pooling filter at each iteration of the filter, with the maximum of the four values ​​being output as the "maximum" value. If such a max pooling filter is applied to an activation filter of 24×24 nodes from convolutional hidden layer 804, the output from pooling hidden layer 806 will be an array of 12×12 nodes.

[0161] In some examples, L2 norm pooling filters may also be used. L2 norm pooling filters involve calculating the square root of the sum of squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (instead of calculating the maximum value as done in max pooling), and using the calculated value as the output.

[0162] Pooling functions (e.g., max pooling, L2 norm pooling, or other pooling functions) determine whether a given feature is found anywhere in a region of an image and discard the exact localization information. This can be done without affecting the results of feature detection because once a feature has been found, its exact location is less important than its approximate location relative to other features. Max pooling (and other pooling methods) offers the benefit of having far fewer pooling features, thus reducing the number of parameters required in subsequent layers of a CNN 800.

[0163] The final connection in the network is a fully connected layer that connects each node from the pooling hidden layer 806 to each output node in the output layer 810. Using the example above, the input layer comprises 28×28 nodes encoding the pixel intensity of the input image, the convolutional hidden layer 804 comprises 3×24×24 hidden feature nodes based on applying a 5×5 local receptive field (for filtering) to three activation maps, and the pooling hidden layer 806 comprises a layer of 3×12×12 hidden feature nodes based on applying a max-pooling filter to a 2×2 region across each of the three feature maps. Extending this example, the output layer 810 may comprise ten output nodes. In this example, each node of the 3×12×12 pooling hidden layer 806 is connected to each node in the output layer 810.

[0164] The fully connected layer 808 takes the output of the preceding pooling hidden layer 806 (which should represent an activation map of high-level features) and determines the features most relevant to a particular class. For example, the fully connected layer 808 can determine the high-level features most relevant to a particular class and may include weights (nodes) for those high-level features. The product between the weights of the fully connected layer 808 and the pooling hidden layer 806 can be computed to obtain the probabilities for different classes. For example, if the CNN 800 is used to predict that the object in an image is a person, there will be high values ​​in the activation map representing the high-level features of a person (e.g., two legs, a face at the top of the object, two eyes at the top left and top right of the face, a nose in the middle of the face, a mouth at the bottom of the face, and / or other features common to people).

[0165] In some examples, the output from output layer 810 may include an M-dimensional vector (M=10 in the previous example). M indicates the number of classes from which CNN 800 must choose when classifying objects in an image. Other example outputs may also be provided. Each number in the M-dimensional vector represents the probability that an object belongs to a certain class. In an exemplary example, if the 10-dimensional output vector represents objects of ten different classes as [0 0 0.05 0.8 0 0.15 0 0 0 0], then the vector indicates a 5% probability that the image is an object of the third class (e.g., a dog), an 80% probability that the image is an object of the fourth class (e.g., a person), and a 15% probability that the image is an object of the sixth class (e.g., a kangaroo). The probability of a class can be considered as the confidence level that an object is part of that class.

[0166] Figure 9An example computing device architecture 900 is illustrated, illustrating example computing devices that can implement the various technologies described herein. In some examples, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device within a vehicle), or other devices. For example, computing device architecture 900 may include, implement Figure 3A and Figure 3B The tracking system 320 may be any part or all of it, or may be included therein. Additionally or alternatively, the computing device architecture 900 may be configured to execute process 400, process 500, process 550, process 600, process 620 and / or other processes described herein.

[0167] The components of the computing device architecture 900 are shown to communicate electrically with each other using a connection 912, such as a bus. The example computing device architecture 900 includes a processing unit (CPU or processor) 902 and a computing device connection 912 that couples various computing device components, including computing device memories 910 (such as read-only memory (ROM) 908 and random access memory (RAM) 906), to the processor 902.

[0168] The computing device architecture 900 may include a cache of high-speed memory directly connected to, very adjacent to, or integrated into the processor 902. The computing device architecture 900 may copy data from memory 910 and / or storage device 914 to cache 904 for fast access by the processor 902. In this way, the cache provides a performance improvement by avoiding latency for the processor 902 while waiting for data. These and other modules may control or be configured to control the processor 902 to perform various actions. Other computing device memory 910 may also be available. Memory 910 may include various different types of memory with different performance characteristics. The processor 902 may include any general-purpose processor and hardware or software services configured to control the processor 902 (such as services 1 916, 2 918, and 3 920 stored in storage device 914), as well as dedicated processors in which software instructions are incorporated into the processor design. The processor 902 may be a self-contained system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0169] To enable user interaction with the computing device architecture 900, input device 922 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. Output device 924 can also be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, projector, television, speaker equipment, etc. In some instances, multi-mode computing devices enable users to provide multiple types of input to communicate with the computing device architecture 900. Communication interface 926 typically controls and manages user input and computing device output. There are no limitations on operation on any particular hardware arrangement, and therefore the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0170] Storage device 914 is a non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as a magnetic tape cassette, flash memory card, solid-state memory device, digital multifunction disk, magnetic tape cartridge, random access memory (RAM) 906, read-only memory (ROM) 908, and hybrid forms thereof. Storage device 914 may include services 916, 918, and 920 for controlling processor 902. Other hardware or software modules are envisioned. Storage device 914 may be connected to computing device connection 912. In one aspect, a hardware module performing a specific function may include software components stored in a computer-readable medium connected to necessary hardware components, such as processor 902, connection 912, output device 924, etc., to perform that function.

[0171] With reference to a given parameter, property, or condition, the term "substantially" may mean that a person skilled in the art would understand that a given parameter, property, or condition is satisfied with a small degree of variance (such as, for example, within acceptable manufacturing tolerances). For example, depending on the specific parameter, property, or condition that is substantially satisfied, the parameter, property, or condition may be satisfied at least 90%, at least 95%, or even at least 99%.

[0172] Various aspects of this disclosure are applicable to any suitable electronic device (such as a security system, smartphone, tablet computer, laptop computer, vehicle, drone, or other device) that includes or is coupled to one or more active depth sensing systems. Although devices having or coupled to a light projector are described below, various aspects of this disclosure are applicable to devices having any number of light projectors and are therefore not limited to any particular device.

[0173] The term "device" is not limited to one or a specific number of physical objects (such as a smartphone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that implement at least some parts of this disclosure. Although the following description and examples use the term "device" to describe various aspects of this disclosure, the term "device" is not limited to a specific configuration, type, or number of objects. Additionally, the term "system" is not limited to multiple components or specific aspects. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. Although the following description and examples use the term "system" to describe various aspects of this disclosure, the term "system" is not limited to a specific configuration, type, or number of objects.

[0174] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects can be practiced without these specific details. For clarity, in some cases, the technology may be presented as comprising individual functional blocks, including functional blocks comprising devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring the aspects.

[0175] Various aspects described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. Although flowcharts can describe operations as sequential processes, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but a process may have additional steps not included in the accompanying diagrams. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.

[0176] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, cause or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion of the computer resources used may be accessible via a network. Computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc.

[0177] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, magnetic disks or optical disks, USB devices equipped with non-volatile memory, network storage devices, any suitable combinations thereof, etc. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, independent variables, parameters, data, etc., can be transmitted, forwarded, or sent through any suitable means, including memory sharing, message passing, token passing, network sending, etc.

[0178] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0179] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or interlocking cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.

[0180] Instructions, media for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0181] In the foregoing description, aspects of this application have been described with reference to their specific aspects, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that various inventive concepts may be embodied and employed in various other ways, and the appended claims are not intended to be construed as including these variations unless limited by prior art. The various features and aspects of the applications described above may be used individually or in combination. Furthermore, aspects may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.

[0182] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with less than or equal to ("≤") and greater than or equal to ("≥") symbols without departing from the scope of this description.

[0183] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0184] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0185] The claim language or other language that states "at least one of" and / or "one or more of" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, the claim language that states "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, the claim language that states "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any repetition is information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language "at least one of" and / or "one or more of" in a set does not limit the set to the items listed in the set. For example, the language of a claim stating "at least one of A and B" or "at least one of A or B" may refer to A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases "at least one" and "one or more" are used interchangeably herein.

[0186] Claims using phrases such as "at least one processor, at least one processor is configured to," "at least one processor is configured to," "one or more processors, one or more processors are configured to," or other languages ​​indicate that one or more processors (in any combination) are capable of performing associated operations. For example, a claim using the phrase "at least one processor, the at least one processor is configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each assigned a specific subset of tasks to perform operations X, Y, and Z, such that the multiple processors together perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, a claim using the phrase "at least one processor, the at least one processor is configured to: X, Y, and Z" could mean that any single processor can perform only a subset of operations X, Y, and Z.

[0187] When referring to one or more elements that perform functions (e.g., steps of a method), one element may perform all functions, or more than one element may jointly perform these functions. When more than one element jointly performs these functions, each function does not need to be performed by every single element (e.g., different functions may be performed by different elements), and / or each function does not need to be performed by only one element as a whole (e.g., different elements may perform different sub-functions of a function). Similarly, when referring to one or more elements configured to cause another element (e.g., a device) to perform functions, one element may be configured to cause another element to perform all functions, or more than one element may be jointly configured to cause another element to perform these functions.

[0188] When referring to an entity that performs or is configured to perform functions (e.g., steps of a method) (e.g., any entity or device described herein), the entity may be configured to cause one or more elements (individually or collectively) to perform those functions. One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more of those functions, and / or any combination thereof. When referring to an entity that performs functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to perform those functions collectively. When the entity is configured to cause more than one component to perform those functions collectively, each function does not need to be performed by every single component (e.g., different functions may be performed by different components), and / or each function does not need to be performed by only one component as a whole (e.g., different components may perform different sub-functions of a function).

[0189] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this application.

[0190] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.

[0191] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, a combination of one or more microprocessors with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.

[0192] The exemplary aspects of this disclosure include: Aspect 1. An apparatus for tracking a human-machine interface (HMI) device, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: extract image features from images of the HMI device captured by a camera of a head-mounted device (HMD), the image features being associated with at least one of a hand or arm of a user holding the HMI device; extract light-emitting diode (LED) features from the images, the LED features being associated with an LED on the HMI device; track the image features and the LED features across multiple images; obtain motion data from the HMI device, the motion data being measured by an inertial measurement unit (IMU) of the HMI device; determine motion data pre-integration based on the motion data; predict the localization of the LED on the HMI device and the localization of the image features based on the tracked image features, the tracked LED features and the motion data; and fuse the motion data pre-integration, the localization of the LED and the localization of the image features to determine the attitude of the HMI device.

[0193] Aspect 2. The apparatus according to aspect 1, wherein, in order to fuse the motion data, the LED positioning of the LED and the positioning of the image features, the at least one processor is configured to use a filter to track and update the pose of the HMI device over time.

[0194] Aspect 3. The apparatus according to aspect 2, wherein the filter comprises at least one of: an extended Kalman filter (EKF), an unscented Kalman filter (UKF), or a particle filter.

[0195] Aspect 4. The apparatus according to any one of Aspects 1 to 3, wherein, in order to fuse the motion data, the positioning of the LED, and the positioning of the image features, the at least one processor is configured to: determine a system state vector, the system state vector including the error of the relative positioning of the image features in the IMU coordinate system.

[0196] Aspect 5. The apparatus according to aspect 4, wherein the system state vector includes the relatively positioned sliding window.

[0197] Aspect 6. The apparatus according to any one of Aspects 4 or 5, wherein the at least one processor is further configured to filter the value of the system state vector.

[0198] Aspect 7. The apparatus according to aspect 6, wherein the system state vector is filtered according to Mahalanobis distance.

[0199] Aspect 8. The apparatus according to any one of Aspects 1 to 7, wherein the at least one processor is further configured to: propagate a system state vector and a covariance matrix using IMU pre-integration; construct a filter measurement model using LED measurements and feature measurements; remove outliers using distance tests; augment or marginalize the system state vector and the covariance matrix; update the filter measurement model to obtain a system error state estimate and a posterior covariance matrix; correct the system state using the system error state estimate; maintain feature-motion data relationships; and marginalize the oldest attitude state from a sliding window.

[0200] Aspect 9. An apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: propagate a system state vector and a covariance matrix using IMU pre-integration; construct a filter measurement model using LED measurements and feature measurements; remove outliers using distance tests; augment or marginalize the system state vector and the covariance matrix; update the filter measurement model to reduce system error state estimates and posterior covariance matrices; correct system states using the system error state estimates; maintain feature-motion data relationships; and marginalize the oldest attitude states from a sliding window.

[0201] Aspect 10. A method for tracking a human-machine interface (HMI) device, the method comprising: extracting image features from images of the HMI device captured by a camera of a head-mounted device (HMD), the image features being associated with at least one of a hand or arm of a user holding the HMI device; extracting light-emitting diode (LED) features from the images, the LED features being associated with an LED on the HMI device; tracking the image features and the LED features across multiple images; obtaining motion data from the HMI device, the motion data being measured by an inertial measurement unit (IMU) of the HMI device; determining motion data pre-integration based on the motion data; predicting the localization of the LED on the HMI device and the localization of the image features based on the tracked image features, the tracked LED features and the motion data; and fusing the motion data pre-integration, the localization of the LED and the localization of the image features to determine the attitude of the HMI device.

[0202] Aspect 11. The method according to aspect 10, wherein, in order to fuse the motion data, the LED positioning of the LED, and the positioning of the image features, the method includes: using a filter to track and update the pose of the HMI device over time.

[0203] Aspect 12. The method according to aspect 11, wherein the filter comprises at least one of: an extended Kalman filter (EKF), an unscented Kalman filter (UKF), or a particle filter.

[0204] Aspect 13. The method according to any one of Aspects 10 to 12, wherein, in order to fuse the motion data, the positioning of the LED, and the positioning of the image features, the method includes: determining a system state vector, the system state vector including an error in the relative positioning of the image features in the IMU coordinate system.

[0205] Aspect 14. The method according to aspect 13, wherein the system state vector includes the relative positioning sliding window.

[0206] Aspect 15. The method according to any one of Aspects 13 or 14, the method further comprising filtering the value of the system state vector.

[0207] Aspect 16. The method according to aspect 15, wherein the system state vector is filtered according to Mahalanobis distance.

[0208] Aspect 17. The method according to any one of Aspects 10 to 16, further comprising: propagating a system state vector and a covariance matrix using IMU pre-integration; constructing a filter measurement model using LED measurements and feature measurements; removing outliers using distance tests; augmenting or marginalizing the system state vector and the covariance matrix; updating the filter measurement model to obtain a system error state estimate and a posterior covariance matrix; correcting the system state using the system error state estimate; maintaining feature-motion data relationships; and marginalizing the oldest attitude state from a sliding window.

[0209] Aspect 18. A method comprising: propagating a system state vector and a covariance matrix using IMU pre-integration; constructing a filter measurement model using LED measurements and feature measurements; removing outliers using a distance test; augmenting or marginalizing the system state vector and the covariance matrix; updating the filter measurement model to reduce the system error state estimate and the posterior covariance matrix; correcting the system state using the system error state estimate; maintaining feature-motion data relationships; and marginalizing the oldest attitude state from a sliding window.

[0210] Aspect 19. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions, when executed by at least one processor, causing the at least one processor to: extract image features from images of an HMI device captured by a camera of a head-mounted device (HMD), the image features being associated with at least one of a user's hand or arm holding the HMI device; extract light-emitting diode (LED) features from the images, the LED features being associated with LEDs on the HMI device; track the image features and the LED features across multiple images; obtain motion data from the HMI device, the motion data being measured by an inertial measurement unit (IMU) of the HMI device; determine motion data pre-integration based on the motion data; predict the location of the LEDs on the HMI device and the location of the image features based on the tracked image features, the tracked LED features, and the motion data; and fuse the motion data pre-integration, the location of the LEDs, and the location of the image features to determine the attitude of the HMI device.

[0211] Aspect 20. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions, when executed by at least one processor, causing the at least one processor to: propagate a system state vector and a covariance matrix using IMU pre-integration; construct a filter measurement model using LED measurements and feature measurements; remove outliers using distance tests; augment or marginalize the system state vector and the covariance matrix; update the filter measurement model to reduce the system error state estimate and the posterior covariance matrix; correct the system state using the system error state estimate; maintain feature-motion data relationships; and marginalize the oldest attitude state from a sliding window.

[0212] Aspect 21. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing the at least one processor, when executed by the at least one processor, to perform any one of aspects 10 to 18.

[0213] Aspect 22. An apparatus for providing virtual content for display, the apparatus comprising one or more components for performing operations according to any one of aspects 10 to 18.

[0214] Aspect 23. An apparatus for tracking a human-machine interface (HMI) device, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: determine light-emitting diode (LED) tracking measurements based on an image of the HMI device, wherein the HMI device includes LEDs; determine feature tracking measurements based on the image of the HMI device; determine motion data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; and determine the attitude of the HMI device by fusing the LED tracking measurements, the feature tracking measurements, and the motion data.

[0215] Aspect 24. The apparatus according to aspect 23, wherein, in order to fuse the LED tracking measurement, the feature tracking measurement, and the motion data, the at least one processor is configured to use a filter to track and update the attitude of the HMI device over a time period.

[0216] Aspect 25. The apparatus according to any one of Aspects 23 or 24, wherein, in order to determine the attitude of the HMI device, the at least one processor is configured to update the state of the filter based on the LED tracking measurement, the feature tracking measurement, and the movement data.

[0217] Aspect 26. The apparatus according to aspect 25, wherein the filter comprises at least one of: an extended Kalman filter (EKF), an unscented Kalman filter (UKF), or a particle filter.

[0218] Aspect 27. The apparatus according to any one of Aspects 23 to 26, wherein, in order to fuse the LED tracking measurement, the feature tracking measurement, and the motion data, the at least one processor is configured to determine a system state vector, the system state vector including errors in the relative positioning of image features in the IMU coordinate system.

[0219] Aspect 28. The apparatus according to aspect 27, wherein the system state vector includes the relatively positioned sliding window.

[0220] Aspect 29. The apparatus according to any one of Aspects 27 or 28, wherein the system state vector is filtered according to distance.

[0221] Aspect 30. The apparatus according to any one of Aspects 23 to 29, wherein, in order to determine the LED tracking measurement, the at least one processor is configured to: identify LED features of the image of the HMI device; and track the LED features across the image of the HMI device.

[0222] Aspect 31. The apparatus according to aspect 30, wherein, in order to determine the LED tracking measurement, the at least one processor is configured to predict the localization of the LED features.

[0223] Aspect 32. The apparatus according to any one of aspects 23 to 31, wherein, in order to determine the feature tracking measurement, the at least one processor is configured to: identify image features of the image of the HMI device; and track the image features across the image of the HMI device.

[0224] Aspect 33. The apparatus according to aspect 32, wherein, in order to determine the feature tracking measurement, the at least one processor is configured to predict the localization of the image features.

[0225] Aspect 34. The apparatus according to any one of Aspects 32 or 33, wherein the image feature is associated with at least one of: surface features of the HMI device, the hand of a user holding the HMI device, or the arm of the user.

[0226] Aspect 35. The apparatus according to any one of Aspects 23 to 34, wherein, in order to determine the motion data, the at least one processor is configured to determine motion data pre-integration based on the measurements of the IMU.

[0227] Aspect 36. The apparatus according to any one of Aspects 23 to 35, wherein, in order to determine the movement data, the at least one processor is configured to predict the attitude of the HMI device based on the measurements of the IMU.

[0228] Aspect 37. The apparatus of any one of Aspects 23 to 36, wherein the LED tracking measurement, the feature tracking measurement, and the motion data are fused based on the camera pose of a camera used to capture the image of the HMI device. Aspect 38. The apparatus of claim 16, wherein the apparatus includes a head-mounted device including the camera.

[0229] Aspect 39. A method for tracking a human-machine interface (HMI) device, the method comprising: determining light-emitting diode (LED) tracking measurements based on an image of the HMI device, wherein the HMI device includes LEDs; determining feature tracking measurements based on the image of the HMI device; determining motion data of the HMI device based on an inertial measurement unit (IMU) of the HMI device; and determining the attitude of the HMI device by fusing the LED tracking measurements, the feature tracking measurements, and the motion data.

[0230] Aspect 40. The method according to aspect 39, wherein fusing the LED tracking measurement, the feature tracking measurement, and the motion data includes using a filter to track and update the attitude of the HMI device over a time period.

[0231] Aspect 41. The method according to any one of Aspects 39 or 40, wherein determining the attitude of the HMI device includes updating the state of the filter based on the LED tracking measurement, the feature tracking measurement, and the movement data.

[0232] Aspect 42. The method according to aspect 41, wherein the filter comprises at least one of: an extended Kalman filter (EKF), an unscented Kalman filter (UKF), or a particle filter.

[0233] Aspect 43. The method according to any one of Aspects 39 to 42, wherein fusing the LED tracking measurement, the feature tracking measurement, and the motion data includes determining a system state vector, the system state vector including errors in the relative positioning of image features in the IMU coordinate system.

[0234] Aspect 44. The method according to aspect 43, wherein the system state vector includes the relatively positioned sliding window.

[0235] Aspect 45. The method according to any one of Aspects 43 or 44, wherein the system state vector is filtered according to distance.

[0236] Aspect 46. The method according to any one of Aspects 39 to 45, wherein determining the LED tracking measurement comprises: identifying LED features of the image of the HMI device; and tracking the LED features across the image of the HMI device.

[0237] Aspect 47. The method according to aspect 46, wherein determining the LED tracking measurement includes predicting the location of the LED features.

[0238] Aspect 48. The method according to any one of Aspects 39 to 47, wherein determining the feature tracking measurement comprises: identifying image features of the image of the HMI device; and tracking the image features across the image of the HMI device.

[0239] Aspect 49. The method according to aspect 48, wherein determining the feature tracking measurement includes predicting the localization of the image features.

[0240] Aspect 50. The method according to any one of Aspects 48 or 49, wherein the image feature is associated with at least one of: surface features of the HMI device, the hand of a user holding the HMI device, or the arm of the user.

[0241] Aspect 51. The method according to any one of Aspects 39 to 50, wherein determining the motion data includes determining motion data pre-integration based on measurements of the IMU.

[0242] Aspect 52. The method according to any one of Aspects 39 to 51, wherein determining the motion data includes predicting the attitude of the HMI device based on measurements of the IMU.

[0243] Aspect 53. The method according to any one of Aspects 39 to 52, wherein the LED tracking measurement, the feature tracking measurement, and the motion data are fused based on the camera pose of the camera used to capture the image of the HMI device.

[0244] Aspect 54. The method according to aspect 53, wherein the method is implemented by a head-mounted device including the camera.

[0245] Aspect 55. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing the at least one processor, when executed, to perform any one of aspects 39 to 54.

[0246] Aspect 56. An apparatus for providing virtual content for display, the apparatus comprising one or more components for performing operations according to any one of aspects 39 to 54.

Claims

1. An apparatus for tracking a human-machine interface (HMI) device, the apparatus comprising: At least one memory; and At least one processor, the at least one processor being coupled to the at least one memory and being configured to: The light-emitting diode (LED) tracking measurement is determined based on the image from the HMI device, wherein the HMI device includes LEDs; Feature tracking measurements are determined based on the image from the HMI device; The movement data of the HMI device is determined based on the inertial measurement unit (IMU) of the HMI device; as well as The attitude of the HMI device is determined by fusing the LED tracking measurements, the feature tracking measurements, and the motion data.

2. The apparatus of claim 1, wherein, in order to fuse the LED tracking measurements, the feature tracking measurements, and the motion data, the at least one processor is configured to use a filter to track and update the attitude of the HMI device over a time period.

3. The apparatus of claim 1, wherein, in order to determine the attitude of the HMI device, the at least one processor is configured to update the state of the filter based on the LED tracking measurement, the feature tracking measurement, and the movement data.

4. The apparatus of claim 3, wherein the filter comprises at least one of: an extended Kalman filter (EKF), an unscented Kalman filter (UKF), or a particle filter.

5. The apparatus of claim 1, wherein, in order to fuse the LED tracking measurement, the feature tracking measurement, and the motion data, the at least one processor is configured to determine a system state vector, the system state vector including errors in the relative positioning of image features in the IMU coordinate system.

6. The apparatus of claim 5, wherein the system state vector includes the relatively positioned sliding window.

7. The apparatus of claim 5, wherein the system state vector is filtered based on distance.

8. The apparatus of claim 1, wherein, in order to determine the LED tracking measurement, the at least one processor is configured to: LED features that identify the image of the HMI device; and The image across the HMI device tracks the LED features.

9. The apparatus of claim 8, wherein, in order to determine the LED tracking measurement, the at least one processor is configured to predict the localization of the LED features.

10. The apparatus of claim 1, wherein, in order to determine the feature tracking measurement, the at least one processor is configured to: Image features that identify the image of the HMI device; and The image features are tracked across the HMI device.

11. The apparatus of claim 10, wherein, in order to determine the feature tracking measurement, the at least one processor is configured to predict the localization of the image features.

12. The apparatus of claim 10, wherein the image feature is associated with at least one of: surface features of the HMI device, the hand of a user holding the HMI device, or the arm of the user.

13. The apparatus of claim 1, wherein, in order to determine the motion data, the at least one processor is configured to determine motion data pre-integration based on measurements from the IMU.

14. The apparatus of claim 1, wherein, in order to determine the movement data, the at least one processor is configured to predict the attitude of the HMI device based on measurements from the IMU.

15. The apparatus of claim 1, wherein the LED tracking measurement, the feature tracking measurement, and the motion data are fused based on the camera pose of the camera used to capture the image of the HMI device.

16. The apparatus of claim 15, wherein the apparatus includes a head-mounted device, the head-mounted device including the camera.

17. A method for tracking a human-machine interface (HMI) device, the method comprising: The light-emitting diode (LED) tracking measurement is determined based on the image from the HMI device, wherein the HMI device includes LEDs; Feature tracking measurements are determined based on the image from the HMI device; The movement data of the HMI device is determined based on the inertial measurement unit (IMU) of the HMI device; as well as The attitude of the HMI device is determined by fusing the LED tracking measurements, the feature tracking measurements, and the motion data.

18. The method of claim 17, wherein fusing the LED tracking measurement, the feature tracking measurement, and the motion data includes using a filter to track and update the attitude of the HMI device over a time period.

19. The method of claim 17, wherein determining the attitude of the HMI device comprises updating the state of the filter based on the LED tracking measurement, the feature tracking measurement, and the motion data.

20. The method of claim 19, wherein the filter comprises at least one of: an extended Kalman filter (EKF), an unscented Kalman filter (UKF), or a particle filter.