Self-centered body posture tracking

By using wearable sensors and deep neural networks in augmented and virtual reality, the problem of self-centered, hands-free and unobstructed body tracking in the prior art is solved, and a higher user interaction experience is achieved.

CN120077344APending Publication Date: 2025-05-30SNAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380065966.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-15
Filing Date
2023-09-14
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to achieve self-centered, hands-free and unobstructed body tracking systems in augmented and virtual reality, resulting in limited immersion and naturalness of user interaction.

Method used

High-fidelity posture estimation is achieved through a human posture tracking system using wearable sensors, combined with EMF tracking sensors and IMU sensors, deep neural networks and reverse kinematic models are used to achieve high-fidelity posture estimation, and metal interference is detected and corrected by EMF-IMU fusion method.

Benefits of technology

The self-centered, hands-free and unobstructed body tracking is achieved, improving the user's interactive experience in virtual reality and augmented reality, and enhancing immersion and nature.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077344A_ABST
    Figure CN120077344A_ABST
Patent Text Reader

Abstract

A gesture tracking system is provided. The gesture tracking system includes an EMF tracking system having a head-worn EMF source worn by a user and one or more user-worn EMF tracking sensors attached to a wrist of the user. The EMF source is associated with a VIO tracking system such as AR glasses. The gesture tracking system determines a gesture of the user's head and a ground plane using a VIO tracking system, and determines a gesture of the user's hands using an EMF tracking system to determine a whole body gesture of the user. Metal interference to the EMF tracking system is minimized using an IMU mounted with the EMF tracking sensor. Using an EMF tracking system, long-term drift in IMU and VIO tracking systems is minimized.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Claim

[0002] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 375,811, filed on September 15, 2022, which is hereby incorporated by reference in its entirety. Technical Field

[0003] The present disclosure generally relates to augmented reality and virtual reality, and more particularly to body pose detection. Background Art

[0004] A head-mounted device can be implemented with a transparent or translucent display through which a user of the head-mounted device can view the surrounding environment. Such a device enables the user to look through the transparent or translucent display to view the surrounding environment and also to see objects (e.g., virtual objects such as renderings of 2D or 3D graphical models, images, videos, text, etc.) that are generated for display and appear as part of and / or overlaid on the surrounding environment. This is generally referred to as "augmented reality" or "AR". The head-mounted device can also completely occlude the user's field of view and display a virtual environment through which the user can move or be moved. This is generally referred to as "virtual reality" or "VR". In a hybrid form, a view of the surrounding environment is captured using a camera device and then displayed to the user together with augmentation on a display that occludes the user's eyes. As used herein, the terms extended reality "XR" and "AR" refer to augmented reality, virtual reality, and any hybrid of these technologies, unless the context indicates otherwise. Brief Description of the Drawings

[0005] To facilitate identification of discussion of any particular element or act, one or more of the most significant digits in the reference numerals refer to the figure number in which the element is first introduced.

[0006] Figure 1 is a perspective view of a head-mounted device according to some examples.

[0007] Figure 2 shows additional views of the head-mounted device in Figure 1 according to some examples.

[0008] Figure 3 is a graphical representation of a machine within which a set of instructions can be executed to cause the machine to perform any one or more of the methods discussed herein.

[0009] Figure 4A and Figure 4B are illustrations of components of a pose tracking system according to some examples.

[0010] Figure 4C is a pose tracking method of a pose tracking system according to some examples.

[0011] Figure 5A is an illustration of a 3D body model according to some examples, and Figure 5B is an illustration of 3D body model generation processing according to some examples.

[0012] Figure 6 is an illustration of an error introduced into a pose tracking system by metal interference according to some examples.

[0013] Figure 7 is a process flow diagram of a method for metal interference detection and mitigation method according to some examples.

[0014] Figure 8 illustrates an error correction method according to some examples.

[0015] Figure 9 illustrates a machine learning pipeline according to some examples.

[0016] Figure 10 illustrates the training and use of a machine learning program according to some examples.

[0017] Figure 11 is a block diagram showing a software architecture in which the present disclosure can be implemented according to some examples.

[0018] Figure 12 is a block diagram showing a networked system including details of a head-mounted XR or AR system according to some examples.

[0019] Figure 13 is a block diagram showing an example messaging system for exchanging data (e.g., messages and associated content) over a network according to some examples. DETAILED DESCRIPTION

[0020] Human pose tracking by wearable sensors has great potential in XR applications such as telecommunication using 3D avatars with full body expression. Some pose tracking systems use vision-based methods with hand-held controllers, limiting natural body-centered interactions such as hands-free movement, and vision-based systems may be robust to full or partial occlusion of the image sensor, while other pose tracking systems using body-worn inertial measurement unit (IMU) sensors fail due to insufficient accuracy.

[0021] Human motion tracking can be used in various human-computer interaction applications, especially in XR. Some devices use camera devices embedded in head-mounted displays to track the user's head pose and use two handheld controllers for spatial input in world coordinates. These inputs give sparse information about the body pose and may not be able to directly recover the full-body pose with more joints and degrees of freedom. This may detract from their utility in driving user avatars or designing full-body interactions in a virtual world. Additionally, due to the limited field of view of the camera device, the controllers may be prone to going out of view and losing tracking, restricting the user's interaction range. Additionally, the user needs to hold the controllers with both hands, which may impede their ability to interact with the virtual world using their fingers. These constraints, namely the lack of finger freedom and full-body tracking, have a negative impact on the immersion and naturalness of the overall experience in XR. Therefore, a self-centered, hands-free, and unobstructed body tracking system is desired.

[0022] In some examples, a pose tracking system includes a head-mounted XR system such as AR glasses and one or more wrist-mounted electromagnetic field (EMF) tracking sensors. The pose tracking system uses a trained deep neural network for inverse kinematics to achieve high-fidelity pose estimation.

[0023] In some examples, the pose tracking system uses IMU data combined with EMF sensor data to detect and correct metal interference in the EMF tracking sensors and improve the tracking improvement of the EMF-IMU fusion method to detect and correct the interfered EMF tracking.

[0024] In some examples, a full-body pose tracking system includes a head-mounted display (HMD) and magnetic tracking in the form of a wristband. In the pose tracking system, an electromagnetic field (EMF) source is combined with the visual inertial odometry (VIO) tracking of the HMD, and the pose tracking system is capable of tracking the 6 degrees of freedom (DoF) pose of three locations (the head and two wrists). The pose tracking system uses a neural network to reconstruct the body pose from these sparse signals, and the neural network is trained using human pose inverse kinematics (IK) to recognize human poses. The neural network is trained on a dataset to generate reasonable body poses.

[0025] In some examples, the pose tracking system addresses the problem of metal interference in magnetic tracking by leveraging an IMU sensor embedded with the EMF sensor. The pose tracking system detects metal interference in real time and, additionally, mitigates the impact by correcting the tracking with the EMF-IMU fusion method.

[0026] In an example, the performance of a pose tracking system is built for high performance in body reconstruction and robustness to metal interference. Inverse kinematics (IK) body models are trained on a lightweight upper body reconstruction model designed for resource-constrained XR devices and a full body reconstruction model that also implements more lower body dynamics (as described more fully with reference to Figure 9 and Figure 10 ). The pose tracking system is quantitatively superior to existing hand-held controller-based methods.

[0027] In some examples, the pose tracking system provides the following features:

[0028] · A hands-free, occlusion-robust body pose tracking system with a single head-mounted display (HMD) and two EMF sensors on the user's wrists.

[0029] · Two inverse kinematics models: a lightweight model for the upper body only and a complex model for the full body.

[0030] · To address problems in EMF tracking, namely metal interference, detect metal interference and mitigate tracking failures.

[0031] · EMF tracking data is used to correct the long-term drift of the VIO tracking system and / or the inertial measurement unit (IMU).

[0032] Other technical features will be readily apparent to those skilled in the art from the following drawings, description, and claims.

[0033] Figure 1 is a perspective view of a head-mounted XR or AR system (e.g., Figure 1 the glasses 100). The glasses 100 may include a frame 102 made of any suitable material such as plastic or metal, any suitable material including any suitable shape memory alloy. In one or more examples, the frame 102 includes a first optical element holder or left optical element holder 104 (e.g., a display or lens holder) and a second optical element holder or right optical element holder 106 connected by a bridge portion 112. A first optical element or left optical element 108 and a second optical element or right optical element 110 may be disposed within the left optical element holder 104 and the right optical element holder 106, respectively. The right optical element 110 and the left optical element 108 may be lenses, displays, display components, or a combination of the foregoing. Any suitable display component may be disposed in the glasses 100.

[0034] The frame 102 further includes a left arm piece or left temple piece 122 and a right arm piece or right temple piece 124. In some examples, the frame 102 may be formed from a single piece of material to have a unified or integral construction.

[0035] The glasses 100 may include a computing system such as the computer 120, which may be of any suitable type to be carried by the frame 102, and in one or more examples, the computing system may be of a suitable size and shape to be partially disposed in one of the temple pieces 122 or 124. The computer 120 may include multiple processors, memories, and various communication components that share a common power supply. As discussed below, various components of the computer 120 may include low-power circuitry, high-speed circuitry, and a display processor. Additional details of aspects of the computer 120 may be implemented as shown by the data processor 1202 discussed below.

[0036] The computer 120 further includes a battery 118 or other suitable portable power source. In some examples, the battery 118 is disposed in the left temple piece 122 and electrically coupled to the computer 120 disposed in the right temple piece 124. The glasses 100 may include a connector or port (not shown) adapted to charge the battery 118, a wireless receiver, a transmitter, or a transceiver (not shown), or a combination of such devices.

[0037] The glasses 100 include a first camera device or left camera device 114 and a second camera device or right camera device 116. Although two camera devices are depicted, other examples contemplate the use of a single or additional (i.e., more than two) camera devices. In one or more examples, in addition to the left camera device 114 and the right camera device 116, the glasses 100 further include any number of input sensors or other input / output devices. Such sensors or input / output devices may additionally include biometric sensors, position sensors, motion sensors, etc.

[0038] In some examples, the left camera device 114 and the right camera device 116 provide video frame data for the glasses 100 to extract 3D information from a real-world scene environment scene.

[0039] The glasses 100 may further include a touchpad 126 that is mounted to one or both of the left temple piece 122 and the right temple piece 124, or integrated with one or both of the left temple piece 120 and the right temple piece 122. The touchpad 126 is typically arranged vertically and, in some examples, is approximately parallel to the user's temple. As used herein, being typically vertically aligned means that the touchpad is more vertical compared to the horizontal, although potentially more vertical than this. Additional user input may be provided by one or more buttons 128, which in the illustrated example are disposed on the outer upper edges of the left optical element holder 104 and the right optical element holder 106. The one or more touchpads 126 and buttons 128 provide a means by which the glasses 100 can receive input from a user of the glasses 100.

[0040] Figure 2 The glasses 100 are shown from the user's perspective. For clarity, Figure 1 several of the elements shown have been omitted. As Figure 1 described, Figure 2 the glasses 100 shown include a left optical element 108 and a right optical element 110 that are respectively fixed within a left optical element holder 104 and a right optical element holder 106.

[0041] The glasses 100 include: a forward optical assembly 202 that includes a right projector 204 and a right near-eye display 206; and a forward optical assembly 210 that includes a left projector 212 and a left near-eye display 216.

[0042] In some examples, the near-eye display is a waveguide. The waveguide includes a reflective structure or a diffractive structure (e.g., optical elements such as mirrors, lenses, or prisms and / or gratings). The light 208 emitted by the projector 204 encounters the diffractive structure of the waveguide of the near-eye display 206, which directs the light towards the user's right eye to provide an image on or in the right optical element 110 that is superimposed on the view of the real-world scene environment seen by the user. Similarly, the light 214 emitted by the projector 212 encounters the diffractive structure of the waveguide of the near-eye display 216, which directs the light towards the user's left eye to provide an image on or in the left optical element 108 that is superimposed on the view of the real-world scene environment seen by the user. The combination of the GPU, the forward optical assembly 202, the left optical element 108, and the right optical element 110 provides the optical engine of the glasses 100. The glasses 100 use the optical engine to generate a superimposition of the view of the user's real-world scene environment, including displaying a user interface to the user of the glasses 100.

[0043] However, it will be understood that other display technologies or configurations may be utilized within the optical engine to display images to a user within the user's field of view. For example, instead of providing the projector 204 and waveguide, an LCD, LED, or other display panel or display surface may be provided.

[0044] In use, a user of the glasses 100 will be presented with information, content, and various user interfaces on the near-eye display. As described in more detail herein, the user may then interact with the glasses 100 using the touchpad 126 and / or buttons 128, touch input or voice input on an associated device (e.g., Figure 12 the client device 1226 shown in FIG. 1) and / or hand movements, positions, and orientations detected by the glasses 100.

[0045] In some examples, the glasses 100 include a stand-alone XR or AR system that provides an XR or AR experience to a user of the glasses 100. In some examples, the glasses 100 are a component of an XR or AR system that includes one or more other devices that provide additional computing resources and / or additional user input and output resources. The other devices may include a smart phone, a general-purpose computer, etc.

[0046] Figure 3 is a diagrammatic representation of a machine 300 in which instructions 310 (e.g., software, a program, an application, an applet, an app, or other executable code) may be executed to cause the machine 300 to perform any one or more of the methods discussed herein. The machine 300 may be used as, for example, Figure 1a computer 120 of an XR system in the form of an AR system such as glasses 100. For example, the instructions 310 may cause the machine 300 to perform any one or more of the methods described herein. The instructions 310 transform the general, unprogrammed machine 300 into a particular machine 300 programmed to perform the described and illustrated functions in the described manner. The machine 300 may operate as a stand-alone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine 300 may operate in a server-client network environment as a server machine or a client machine, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 300 in combination with other components of the XR system may be used as but not limited to: a server, a client, a computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a PDA, an entertainment media system, a cellular phone, a smartphone, a mobile device, a head-mounted device (e.g., a smartwatch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing the instructions 310 specifying the actions to be taken by the machine 300. Further, although a single machine 300 is shown, the term "machine" may also be regarded as including a collection of machines that individually or jointly execute the instructions 310 to perform any one or more of the methods discussed herein.

[0047] The machine 300 may include a processor 302, a memory 304, and an I / O device interface 306 that may be configured to communicate with each other via a bus 344. In an example, the processor 302 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), other processors, or any suitable combination thereof) may include, for example, a processor 308 and a processor 312 that execute the instructions 310. The term "processor" is intended to include multi-core processors that may include two or more independent processors (sometimes referred to as "cores") that may execute instructions simultaneously. Although Figure 3 a plurality of processors 302 are shown, the machine 300 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0048] The memory 304 includes a main memory 314, a static memory 316, and a storage unit 318, all of which are accessible by the processor 302 via the bus 344. The main memory 304, the static memory 316, and the storage unit 318 store instructions 310 that embody any one or more of the methods or functions described herein. The instructions 310 may also reside, completely or partially, within the main memory 314, within the static memory 316, within the storage unit 318, within the non-transitory machine-readable medium 320 within the storage unit 318, within one or more of the processors 302 (e.g., within a cache memory of the processor), or within any suitable combination thereof during execution by the machine 300.

[0049] The I / O device interface 306 couples the machine 300 to the I / O device 346. One or more of the I / O devices 346 may be components of the machine 300 or may be separate devices. The I / O device interface 306 may include various interfaces to the I / O device 346 that are used by the machine 300 to receive input, provide output, generate output, transmit information, exchange information, capture measurements, etc. The specific I / O device interface 306 included in a particular machine will depend on the type of the machine. It will be understood that the I / O device interface 306, the I / O device 346 may include Figure 3 many other components not shown. In various examples, the I / O device interface 306 may include an output component interface 328 and an input component interface 332. The output component interface 328 may include interfaces to visual components (e.g., displays such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. The input component interface 332 may include interfaces to alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, optical keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens that provide touch gestures or the location and / or force of a touch, or other haptic input components), audio input components (e.g., microphones), etc.

[0050] In another example, the I / O device interface 306 can include a biometric component interface 334, a motion component interface 336, an environmental component interface 338, or a location component interface 340, as well as various other component interfaces. For example, the biometric component interface 334 can include interfaces to components for detecting expressions (e.g., hand expressions, facial expressions, vocal expressions, body postures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retina recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), and so on.

[0051] The biometric component can include a brain-machine interface (BMI) system that allows communication between the brain and an external device or machine. This can be achieved by recording brain activity data, converting the data into a format that a computer can understand, and then using the resulting signals to control the device or machine.

[0052] Example types of BMI technologies include:

[0053] · Electroencephalogram (EEG)-based BMI, which uses electrodes placed on the scalp to record electrical activity in the brain.

[0054] · Invasive BMI, which uses electrodes surgically implanted into the brain.

[0055] · Optogenetic BMI, which uses light to control the activity of specific nerve cells in the brain.

[0056] Any biometric data collected by the biometric component is captured and stored in a temporary cache only with the user's approval and deleted upon the user's request. Additionally, such biometric data can be used for very limited purposes, such as identification verification. To ensure the restricted and authorized use of biometric information and other personally identifiable information (PII), access to this data is limited to authorized personnel (if any). Any use of biometric data can be strictly limited to identification verification purposes and the biometric data will not be shared or sold to any third party without the user's explicit consent. Additionally, appropriate technical and organizational measures are implemented to ensure the security and confidentiality of this sensitive information.

[0057] The motion component interface 336 may include interfaces to an inertial measurement unit (IMU), an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. The environmental component interface 338 may include interfaces to, for example, an illuminance sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect the ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas detection sensor that detects the concentration of a hazardous gas for safety or measures pollutants in the atmosphere), or other components that may provide indications, measurements, or signals associated with the surrounding physical environment. The position component interface 340 includes interfaces to a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure from which altitude can be obtained), an orientation sensor component (e.g., a magnetometer), etc.

[0058] Various techniques may be used to implement communication. The I / O device interface 306 also includes a communication component interface 342, which is operable to couple the machine 300 to the network 322 or the device 324 via the couplings 330 and 326, respectively. For example, the communication component interface 342 may include an interface to a network interface component that interfaces with the network 322 or another suitable device. In another example, the communication component interface 342 may include interfaces to a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, components (e.g., low power), components, and other communication components that provide communication via other modalities. The device 324 may be another machine or any one of a variety of peripheral devices (e.g., a peripheral device coupled via USB).

[0059] In addition, the communication component interface 342 may include an interface to a component operable to detect an identifier. For example, the communication component interface 342 may include an interface to a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as universal product code (UPC) barcodes, and multi-dimensional barcodes such as quick response (QR) codes, Aztec codes, data matrix, data glyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying tagged audio signals). Additionally, various information may be derived via the communication component interface 342, such as a location via Internet protocol (IP) geolocation, a location via signal triangulation, a location via detecting an NFC beacon signal that may indicate a particular location, etc.

[0060] Various memories (e.g., memory 304, main memory 314, static memory 316, and / or the memory of the processor 302) and / or storage unit 318 may store one or more sets of instructions and data structures (e.g., software) that embody any one or more of the methods or functions described herein or are used by the methods or functions. These instructions (e.g., instructions 310), when executed by the processor 302, cause the various operations to implement the disclosed examples.

[0061] Any one of many well-known transport protocols (e.g., hypertext transfer protocol (HTTP)) may be used and, via a network interface device (e.g., the network interface component included in the communication component interface 342), instructions 310 may be sent or received over the network 322 using a transmission medium. Similarly, instructions 310 may be sent or received using a transmission medium via a coupling 326 (e.g., a peer-to-peer coupling) to the device 324.

[0062] Figure 4A and Figure 4BFIG. is a diagram of components of a pose tracking system according to some examples of the present disclosure. User 442 wears one or more EMF tracking sensors, such as EMF tracking sensor 444, on one or more wrists of user 442, and one or more wrist-mountable EMF tracking sensors, such as EMF tracking sensor 446 and EMF tracking sensor 444, and wears a head-mounted EMF source 448. The EMF source 448 and the one or more EMF tracking sensors include components of an EMF tracking system 414 of the pose tracking system 412. The EMF tracking system 414 generates EMF tracking data 430, which is transmitted to the head and hand pose determination component 418 of the pose tracking system 412.

[0063] The visual inertial odometry (VIO) tracking system 416 of the AR glasses 450 is in a fixed relative pose with respect to the EMF source 448, such that the positions of the EMF tracking sensors 444 and 446 are fixed in a known relative position with respect to the VIO tracking system 416 coordinates. The VIO tracking system 416 uses the visual inertial odometry (VIO)-tracking data 432 generated by the AR glasses 450 to track the head of the user 442 in world coordinates.

[0064] The head and hand pose determination component 418 of the pose tracking system 412 receives the EMF tracking data 430 and the VIO tracking data 432, and calculates the 6DoF head and wrist pose data in the world coordinate system 420. The calculated head and wrist pose data includes the EMF pose in world coordinates based on the head pose determined according to the VIO tracking data 432 received from the VIO tracking system 416 of the AR glasses and the EMF tracking data 430 received from the EMF tracking system 414. One or more wrist poses are tracked by EMF tracking, which provides the relative pose from the wrists of the user 442 to the EMF source 448. By applying the transformation between the EMF source 448 in world coordinates and the relative wrist EMF source pose, the wrist pose in world coordinates can be obtained. Thus, the head and hand pose determination component 418 has the absolute world transformation of the head and two wrists as the measurement result.

[0065] The upper body inverse kinematics component 422 of the pose tracking system 412 receives 6DoF head and wrist pose data in the world coordinate system 420 and uses the inverse kinematics model 452, which is an upper body inverse kinematics model, to generate 3D body model data 426 of the reconstructed full body pose 454 of the user 442 based on these sparse measurements. In some examples, the full body inverse kinematics component 424 of the pose tracking system 412 receives 6DoF head and wrist pose data in the world coordinate system 420 and uses the inverse kinematics model 452, which is a full body inverse kinematics model, to generate 3D body model data 426 of the reconstructed full body pose 454 of the user 442 based on these sparse measurements.

[0066] The pose tracking system 412 transmits the 3D body model data 426 to the XR application 434. The XR application 434 receives the 3D body model data 426 and uses the graphics engine 440 to generate 3D body mesh rendering data 438 based on the 3D body model data 426. The XR application 434 generates or updates the XR user interface 428 for the XR experience of one or more users based on the 3D body mesh rendering data 438.

[0067] In some examples, the EMF source 448 and one or more EMF tracking sensors communicate with the head-mounted display (HMD) of the AR glasses 450 via Bluetooth Low Energy using the ESB (Enhanced Shock Burst) protocol for minimizing latency.

[0068] In some examples, the EMF tracking system 414 of one or more EMF tracking sensors and the EMF source 448 has 3D coils. One or more EMF tracking sensors measure the EMF B-field signals from three orthogonal transmission coils from the source and calculate the relative position and rotation with respect to the source. The basic physics for EMF tracking is Faraday's law: when the sensor moves within the alternating AC magnetic field generated from the source coil, a voltage is generated according to the following formula:

[0069]

[0070] where B is the magnetic field, is the vector of the cross-section of the winding area, and n x is the noise. To track all three axes, for both the source and the sensor, three coils are mounted orthogonally to each other, and each source axis generates magnetic fields of different frequencies for multiplexing. In some examples, the EMF tracking system has a range of 1.5 meters (typical reach of an arm), with a position RMS error of 0.9 mm and an angular RMS error of 0.5 degrees at a 1-meter range.

[0071] In some examples, the IMU sensor is integrated in the EMF tracking sensor. The IMU sensor generates IMU tracking data 436 for the EMF tracking sensor and is used in a sensor fusion method to solve the metal interference problem for the EMF tracking system 414. The head and hand pose determination component 418 receives the IMU tracking data 436 and detects the metal interference for the EMF tracking system 414, and replaces the EMF tracking data 430 with the IMU tracking data 436 when calculating the 6DoF head and wrist pose data in the world coordinate system 420, as further described with reference to Figure 6 and Figure 7 described further.

[0072] In some examples, the EMF tracking system 414 operates at 442 frames per second (fps) with a latency of approximately 15 ms. In some examples, each EMF tracking sensor includes a computing component that hosts the executable components of the corresponding EMF tracking system, and the EMF tracking algorithm runs locally on each EMF tracking sensor. In some examples, the EMF tracking data 430 and the IMU tracking data 436 data streams are synchronized. In some examples, the EMF tracking data 430 and the IMU tracking data 436 can be accessed from an external computing system via a Bluetooth connection.

[0073] In some examples, two tracking systems such as the EMF tracking system 414 and the VIO tracking system 416 provide information for body pose reconstruction. Multiple coordinate systems are defined as follows:

[0074] · World global coordinate: The AR glasses 450 of the VIO tracking system 416 provide the orientation in axis-angle representation and the absolute position of the coordinates, where g represents the glasses and W represents the world coordinates. · EMF local coordinate: The coordinates of the wrist-worn EMF tracking sensor tracked relative to the EMF source 448 on the head of the user 442. The position and rotation are represented as and where s represents the sensor and M represents the EMF coordinates.

[0075] · Body local coordinate: The human body pose can be represented by all joint positions and orientation in the body local coordinate, where j represents the body joint and B represents the body coordinates. A 3D body model is used to represent the human body pose and animate it. The 3D body model takes as input the relative rotations of all joints determined using inverse kinematics and outputs a 3D body mesh.

[0076] Figure 4Cis a pose tracking method 456 of a pose tracking system 412 according to some examples of the present disclosure. In operation 402, the pose tracking system 412 uses the EMF tracking system 414 to determine EMF tracking data 430 of one or more wrists of the user 442.

[0077] In operation 404, the pose tracking system 412 uses the VIO tracking system 416 to determine VIO tracking data 432 of the user's head.

[0078] In operation 406, the pose tracking system 412 determines 6DoF head and wrist pose data of the user's 442 head and one or more wrists of the user 442 in the world coordinate system 420 based on the EMF tracking data 430 and the VIO tracking data 432.

[0079] In operation 408, the pose tracking system 412 generates 3D body model data 426 of the user 442 based on the 6DoF head and wrist pose data in the world coordinate system 420.

[0080] In operation 410, the pose tracking system 412 transmits the 3D body model data 426 to the XR application 434 for use in the XR experience provided to the user 442 by the XR application 434.

[0081] Figure 5A is an illustration of a 3D body model 502, and Figure 5B is an illustration of a 3D body model generation process according to some examples of the present disclosure. In some examples, an SMPL 3D body model is used. The 3D body model 502 is a learned assembly template mesh with 6890 vertices for representing a 3D body shape mesh. The body shape S is defined by identity-related shape parameters β, pose parameters θ, and soft deformation φ as the following equation: S(β, θ, φ) = G(T(β, β, φ), J(β), θ, w), where T(β, θ, φ) represents the body vertices in the rest pose, J ∈ R 3k represents joint positions, θ ∈ R 3k represents body pose (i.e., the relative rotation angles between adjacent body joints), and w ∈ R 4x3N is the blending weight. In some examples, the parameter beta is a constant, meaning a standard body shape is used in IK model training and visualization.

[0082] In some examples, calibration includes scaling up / down the sensor readings to make the input conform to the model training.

[0083] In some examples, all joint positions J and rotations θ are provided to generate the full body shape, which is forward kinematics.

[0084] In some examples, user 442 wears AR glasses 450 and wears EMF tracking sensors 444 and 446 on their wrists. EMF source 448 is located at the back of the head and has a fixed relative pose with respect to AR glasses 450. Through coordinate transformation, body tracking data 510 is generated from EMF tracking data and / or IMU tracking data 436. Body tracking data 510 includes the absolute poses of AR glasses 450 and the two EMF tracking sensors in world coordinates. AR glasses 450 are mapped to joint 15 504 of 3D body model 502, one EMF tracking sensor is mapped to joint 20 506, and the other EMF tracking sensor is mapped to joint 21 508. Thus, the IK problem can be formulated as T = f({pj,θj}), j ∈ {15,20,21}, where pj and θj are the positions and orientations of three known body joints, and the function f represented in the IK model is used to map the three joints to all 22 body joints.

[0085] The IK model includes a linear embedding component 512, one or more transformation encoder components 514, a world transformation decoder 516, a joint rotation decoder 518, and a forward kinematics component 520 operatively connected to joint rotation decoder 518. The pose tracking system uses the IK model to generate 3D body model 502 based on body tracking data 510.

[0086] In some examples, multiple machine learning-based IK models are used according to different use cases and computing resources. Figure 4A The upper body inverse kinematics component 422 in () uses a lightweight per-frame IK model for upper body tracking, and Figure 4A the full body inverse kinematics component 424 in () uses a multi-frame full body IK model. In some use cases, the user uses an upper body representation in XR and may prefer to run the model in real time on a local device; thus, a lightweight machine learning model suitable for real-time operation on a mobile device is desired. In this model, emphasis is placed on model inference efficiency and upper body IK accuracy. Each frame is processed and an IK result is produced, rather than taking a sequence of frames as input, which results in heavier preprocessing and model inference computational overhead. In this model, the global transformation of the body is not considered, and upper body shape reconstruction is the focus. In some examples, the global transformation is ignored by setting the origin at the head position and taking the relative pose of the EMF tracking sensor with respect to the head as input. Since the ground plane data of the ground plane can be determined from the VIO tracking system, the height information that helps predict leg bending is retained. Assuming the head position is p = {x, y, z} and the initial height of the AR glasses is y0, the x and z transformations can be ignored, and the height y is normalized with respect to the initial height y0, then the head position can be represented as p = {0, y / y0, 0}.

[0087] Figure 6 It is an illustration of error graph 604 of errors introduced into the pose tracking system by metal interference according to some examples of the present disclosure. In some examples, the IMU sensor is embedded together with the EMF receiver, thus allowing both interference detection and interference mitigation. Depending on how the sensor moves near metal objects, there are multiple types of metal interference. On the one hand, if the sensor passes by a static metal object, spike-like short-duration errors 602 occur in the EMF tracking data. On the other hand, if the metal object and the EMF tracking sensor move around together, the error persists as long as they are in close proximity. Both of these situations can occur in end-user scenarios: the user may move their arm near a metal object such as a laptop or a door, or hold a smartphone or a metal container when interacting with XR content. Therefore, it is useful to break the problem into two parts, such as an interference detection part and an interference mitigation part. In some examples, the method of interference mitigation (i.e., correcting the tracking error under interference) targets short-duration errors, where the metal object is statically placed in the environment and the user occasionally encounters them while moving. This is because it is difficult for the EMF and IMU sensors to track the pose for a long time under metal interference.

[0088] Figure 6 Sample error graphs under three different conditions over a period of time are shown. From these graphs, the trends of two different types of errors are observed. Under the open space 608 and normal 610 conditions, the errors are mostly spike-like, which means that the error part occurs in a short period of time, corresponding to the moment when the moving wrist is closer to the metal. On the other hand, when the user holds or touches a metal object such as a metal container or a laptop, under the intentional condition 612, large errors persist for several seconds. Both of these errors occur in the end-user environment and will be considered.

[0089] Figure 7 It is a process flow chart of a method of metal interference detection and mitigation method 700 according to some examples of the present disclosure. The pose tracking system uses the interference detection and mitigation method 700 to detect interference caused by metal objects in the operating environment of the pose tracking system. In operation 702, the pose tracking system determines the position, rotation, and linear acceleration of the EMF tracking sensor within the local frame for one or more EMF tracking sensors such as the EMF tracking sensors 444 and 446 in ( Figure 4A ). For example, the pose tracking system determines the EMF quaternion based on the EMF tracking data 430 in ( Figure 4A ), and determines the IMU quaternion based on the IMU tracking data 436 in ( Figure 4A ).

[0090] In operation 704, the pose tracking system determines whether there is metal interference with the measurements of the EMF tracking system 414 in ( Figure 4A ). For example, when moving near the operating environment, the EMF tracking sensor streams two values regarding its orientation based on different principles: the angular momentum from the gyroscope sensor of the IMU and the pose angle from the EMF tracking sensor. Assume that at a given time t without metal interference, the orientation I(t) in axis-angle representation is used as a binary index to indicate whether there is interference. At time t + Δt, the orientation information from the EMF sensor is represented as The approximation of this value is where is the angular momentum in axis-angle representation with the same coordinates as . Then, an error threshold can be introduced to estimate the interference state I(t + Δt) as:

[0091] If then If then where d represents the intrinsic geodesic distance between two given angles.

[0092] In operation 706, if the pose tracking system does not detect metal interference in operation 704, the pose tracking system uses the EMF tracking data without correction. For example, the pose tracking system sets the acceleration value of the nodes in the 3D body model to the measured acceleration of the EMF quaternion of the corresponding EMF tracking sensor, and sets the position of the nodes of the 3D body model to the position of the corresponding EMF tracking sensor.

[0093] In operation 708, if the pose tracking system detects metal interference in operation 704, the pose tracking system corrects the metal interference based on the IMU tracking data 436. For example, the pose tracking system identifies the moment when the interference occurs and corrects the tracking within that moment. The error is corrected in real time in order to develop a real-time body tracking system. For example, if then the past tracking and sensor data up to time t are used to generate the correct value of the current position . Then, if the interference still exists at t + 2Δt, the data up to t + Δt in the past are used to correct the position. This may lead to long-term drift in the corrected values.

[0094] In some examples, IMU odometry data are used to correct the measured data. This method is a physics-based method: given a time series of accelerations from the IMU sensor and an initial velocity, the position is obtained by double integration:

[0095]

[0096] Where x(t), v(t), and a(t) represent position, velocity, and acceleration at time t, and t0 represents the initial time.

[0097] Figure 8 An EMF tracking data correction method 800 according to some examples of the present disclosure is shown. In some examples, ([ Figure 4A in) the pose tracking system 412 uses a previous history 810 of EMF position tracking data to predict future EMF position tracking data 812 of a future trajectory over a short time period, which can be used to correct EMF position tracking data under metal interference. For example, an AI method and one or more trained models are used to determine the corrected EMF position tracking data. The EMF tracking data correction method 800 utilizes historical EMF position tracking data to generate a predicted trajectory and correct errors caused by metal interference. The EMF tracking data correction method 800 uses a deep machine learning model for time series prediction. In some examples, the EMF tracking data correction method 800 utilizes a prediction model 802 with an N-BEATS architecture. The prediction model 802 includes one or more stacks such as stack 808, the stack includes one or more blocks such as block 806, and the block includes one or more fully connected layers 804. The prediction model 802 is trained on representations of different components of the time series of EMF position tracking data (as more fully described with reference to Figure 9 and Figure 10 ), the different components including a trend component, a seasonal component, and a residual component, as illustrated by the following equation:

[0098] {x(t0),..., x(t0 + toutput)} = Prediction_Model({x(t0 - tinput)),..., x(t0)}),

[0099] where toutput and tinput respectively correspond to how much future data the prediction model 802 outputs and how much previous data the prediction model 802 takes as input.

[0100] In some examples, the architecture of the prediction model 802 includes a backward residual link and a forward residual link. The backward residual link models the residual between the previous history 810 of EMF position tracking data and the predicted future EMF position tracking data 812, while the forward link models the prediction itself. The residual link is particularly helpful in accounting for errors introduced in the EMF signal. By decomposing the time series into these different components, the EMF tracking data correction method 800 is able to make accurate predictions by correcting errors in the prediction in real time.

[0101] In some examples,

[0102] In some examples, the prediction model 802 is trained by minimizing a loss function that compares the prediction to the actual values in a future time window, as described more fully with reference to Figure 9 and Figure 10 more fully described.

[0103] In some implementations, the prediction model can work well when there is no acceleration component in the previous EMF position tracking data history 810, as shown in FIG. 814 of the predicted EMF tracking data and the actual ground truth EMF tracking data without acceleration, but the prediction model may have difficulty predicting accurate predicted EMF tracking data, as shown in FIG. 816 of the predicted EMF tracking data and the actual ground truth EMF tracking data with acceleration. In some examples, a fusion model is used to correct the EMF tracking data. For example, to determine the future acceleration while avoiding errors due to noise v(t0), the trajectory is approximated as follows: x(t0 + Δt) = PredictionModel({x(t0 - tinput),..., x(t0)})t0 + Δt + a(t0)×Δt, for example, using the prediction model 802 model in an iterative manner. The model outputs an estimate of touput seconds, using a single predicted value corresponding to time t0 + Δt. After adding the acceleration component, if there is still metal interference, the pose tracking system uses the value predicted by the model 802 as the input for the next step. In this way, the pose tracking system can correct the prediction of the prediction model 802 by adding the acceleration component, further affecting the subsequent trajectory prediction.

[0104] In some examples, a large human motion database, AMASS (as a motion capture archive of surface shapes), is used to train the IK model. This database contains a collection of a series of existing high-precision MoCap datasets based on optical tracking. Specifically, the model training system uses a combination of CMU, Eyes_Japan, KIT, MPI_HDM05, and TotalCapture datasets as the training set, uses MPI_Limits as the validation set, and uses ACCAD and MPI_mosh as the test set. The model training system uses a total of 88,519 training samples, 1182 validation samples, and 2244 test samples. For full-body model training, the model training system downsamples the MoCap dataset from 120Hz to 60Hz and generates window segments of 40 frames (i.e., 2 / 3 second window) with a stride length of 0.1 second. To train the IK model, the model training system uses the Adam solver with a batch size of 32, a starting learning rate of 0.001, and decays 0.8 times every 20 epochs. The model training system trains the model using PyTorch on the Google Cloud Platform with NVIDIA Tesla V100 GPUs.

[0105] In some examples, the pose tracking system detects interference and notifies the user of lower body tracking performance. For example, if the user holds a smartphone that interferes with the pose tracking system, the pose tracking system notifies the user via the AR glasses that the tracking performance is low due to a metal object being close to the sensor. Such a remedy helps for a better user experience.

[0106] Machine learning pipeline 1000

[0107] Figure 10 is a flowchart depicting a machine learning pipeline 1000 according to some examples. The machine learning pipeline 200 can be used to generate a trained model such as the trained machine learning program 1002 in, for example, Figure 10 to perform operations associated with search and query response. ( Figure 4A in) Example models used by the pose tracking system 412 include, but are not limited to, inverse kinematics models, upper body reconstruction models, prediction models, etc.

[0108] Overview

[0109] Broadly, machine learning can involve using computer algorithms to automatically learn patterns and relationships in data, potentially without explicit programming. Machine learning algorithms can be classified into three main categories: supervised learning, unsupervised learning, and reinforcement learning.

[0110] · Supervised learning involves using labeled data to train a model to predict an output for new, unseen inputs. Examples of supervised learning algorithms include linear regression, decision trees, and neural networks.

[0111] · Unsupervised learning involves training a model on unlabeled data to find hidden patterns and relationships in the data. Examples of unsupervised learning algorithms include clustering, principal component analysis, and generative models such as autoencoders.

[0112] · Reinforcement learning involves training a model to make decisions in a dynamic environment by receiving feedback in the form of rewards or punishments. Examples of reinforcement learning algorithms include Q-learning and policy gradient methods.

[0113] According to some examples, examples of specific machine learning algorithms that can be deployed include logistic regression, which is a type of supervised learning algorithm for binary classification tasks. Logistic regression models the probability of a binary response variable based on one or more predictor variables. Another example type of machine learning algorithm is naive Bayes, which is another supervised learning algorithm for classification tasks. Naive Bayes is based on Bayes' theorem and assumes that the predictor variables are independent of each other. Random forest is another type of supervised learning algorithm for classification, regression, and other tasks. Random forest constructs a collection of decision trees and combines their outputs for prediction. Additional examples include neural networks, which include interconnected layers of nodes (or neurons) that process information based on input data and make predictions. Matrix factorization is another type of machine learning algorithm for recommendation systems and other tasks. Matrix factorization decomposes a matrix into two or more matrices to uncover hidden patterns or relationships in the data. Support vector machine (SVM) is a type of supervised learning algorithm for classification, regression, and other tasks. SVM finds a hyperplane that separates different classes in the data. Other types of machine learning algorithms include decision trees, k-nearest neighbors, clustering algorithms, and deep learning algorithms such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and transformer models. The choice of algorithm depends on the nature of the data, the complexity of the problem, and the performance requirements of the application.

[0114] The performance of a machine learning model is typically evaluated on a separate test dataset that is not used during training to ensure that the model can generalize to new, unseen data.

[0115] Although several specific examples of machine learning algorithms are discussed in this article, the principles discussed herein can also be applied to other machine learning algorithms. Deep learning algorithms such as convolutional neural networks, recurrent neural networks, and transformers, as well as more traditional machine learning algorithms such as decision trees, random forests, and gradient boosting, can be used in a variety of machine learning applications.

[0116] Problems of two example types in machine learning are classification problems and regression problems. Classification problems, also known as categorization problems, aim to classify items into one of several categorical values (e.g., is the object an apple or an orange?). Regression algorithms aim to quantify some items (e.g., by providing a value as a real number).

[0117] Training phase 1004

[0118] Generating a trained machine learning program 1002 can include multiple stages that form part of a machine learning pipeline 1000, including, for example Figure 9 the following stages shown in

[0119] · Data collection and preprocessing 902: This stage can include acquiring and cleaning data to ensure its suitability for use in a machine learning model. This stage can also include removing duplicates, handling missing values, and transforming the data into a suitable format.

[0120] · Feature engineering 904: This stage can include selecting and transforming training data 1006 to create features useful for predicting the target variable. Feature engineering can include (1) receiving features 1008 (e.g., as structured data or labeled data in supervised learning) and / or (2) identifying features 1008 in training data 1006 (e.g., for unstructured data or unlabeled data in unsupervised learning). · Model selection and training 906: This stage can include selecting an appropriate machine learning algorithm and training it on the preprocessed data. This stage can also involve splitting the data into a training set and a test set, using cross-validation to evaluate the model, and tuning hyperparameters to improve performance.

[0121] · Model evaluation 908: This stage can include evaluating the performance of the trained model (e.g., the trained machine learning program 1002) on a separate test data set. This stage can help determine whether the model is overfitting or underfitting and whether the model is suitable for deployment.

[0122] · Prediction 910: This stage involves using the trained model (e.g., the trained machine learning program 1002) to generate predictions for new unseen data.

[0123] · Validation, refinement, or retraining 912: This stage can include updating the model based on feedback such as new data or user feedback generated from the prediction stage.

[0124] · Deployment 914: This phase can include integrating the trained model (e.g., the trained machine learning program 1002) into a broader system or application such as a web service, a mobile app, or an IoT device. This phase can involve setting up APIs, building user interfaces, and ensuring that the model is scalable and can handle large amounts of data.

[0125] Figure 10 Other details of two example phases are shown, namely the training phase 1004 (e.g., the model selection and training 906 part) and the prediction phase 1010 (the prediction 910 part). Before the training phase 1004, feature engineering 904 is used to identify features 1008. This can include identifying informative, discriminative, and independent features for effectively operating the trained machine learning program 1002 in pattern recognition, classification, and regression. In some examples, the training data 1006 includes labeled data that is known for the pre-identified features 1008 and one or more outcomes. Each of the features 1008 can be a variable or an attribute, such as a separate measurable characteristic of an item, a system, a phenomenon, or a process represented by a data set (e.g., the training data 1006). By way of example only, the features 1008 can also have different types such as numerical features, strings, and graphs, and can include one or more of content 1012, concepts 1014, attributes 1016, historical data 1018, and / or user data 1020.

[0126] In the training phase 1004, the machine learning pipeline 1000 uses the training data 1006 to find correlations between the features 1008 that affect the prediction outcome or the prediction / inference data 1022.

[0127] Using the training data 1006 and the identified features 1008, the trained machine learning program 1002 is trained during the training phase 1004 during machine learning program training 1024. Machine learning program training 1024 evaluates the values of the features 1008 when the features 1008 are relevant to the training data 1006. The result of the training is the trained machine learning program 1002 (e.g., the trained or learned model).

[0128] In addition, the training phase 1004 can involve machine learning where the training data 1006 is structured (e.g., labeled during preprocessing operations). The trained machine learning program 1002 implements a neural network 1026 capable of performing operations such as classification and clustering. In other examples, the training phase 1004 can involve deep learning where the training data 1006 is unstructured, and the trained machine learning program 1002 implements a deep neural network 1026 capable of performing both feature extraction and classification / clustering operations.

[0129] In some examples, the neural network 226 can be generated during the training phase 1004 and implemented within the trained machine learning program 1002. The neural network 1026 includes a hierarchical (e.g., layered) organization of neurons, with each layer including multiple neurons or nodes. Neurons in the input layer receive input data, while neurons in the output layer produce the final output of the network. Between the input layer and the output layer, there can be one or more hidden layers, each including multiple neurons.

[0130] Each neuron in the neural network 1026 operationally computes a function, such as an activation function, that takes as input the weighted sum of the outputs of the neurons in the previous layer and a bias term. The output of this function is then passed as input to the neurons in the next layer. If the output of the activation function exceeds a certain threshold, then that output is transmitted from the neuron (e.g., the sending neuron) to the connected neurons (e.g., the receiving neurons) in the subsequent layer. The connections between neurons have associated weights that define the influence of the input from the sending neuron on the receiving neuron. During the training phase, these weights are adjusted by a learning algorithm to optimize the performance of the network. Different types of neural networks can use different activation functions and learning algorithms, affecting their performance on different tasks. The hierarchical organization of neurons and the use of activation functions and weights enable the neural network to model complex relationships between inputs and outputs and generalize to new inputs not seen during training.

[0131] In some examples, the neural network 1026 can also be one of several different types of neural networks, by way of example only, such as a single-layer feedforward network, a multi-layer perceptron (MLP), an artificial neural network (ANN), a recurrent neural network (RNN), a long short-term memory network (LSTM), a bidirectional neural network, a symmetrically connected neural network, a deep belief network (DBN), a convolutional neural network (CNN), a generative adversarial network (GAN), an autoencoder neural network (AE), a restricted Boltzmann machine (RBM), a Hopfield network, a self-organizing map (SOM), a radial basis function network (RBFN), a spiking neural network (SNN), a liquid state machine (LSM), an echo state network (ESN), a neural Turing machine (NTM), or a transformer network.

[0132] In addition to the training phase 1004, a validation phase can also be performed on a separate dataset called the validation dataset. The validation dataset is used to tune the hyperparameters of the model, such as the learning rate and regularization parameter. The hyperparameters are adjusted to improve the performance of the model on the validation dataset.

[0133] Once the model is fully trained and validated, in the testing phase, the model can be tested on a new dataset. The test dataset is used to evaluate the performance of the model and ensure that the model has not overfit to the training data.

[0134] In the prediction phase 1010, the trained machine learning program 1002 uses the features 1008 to analyze the query data 1028 to generate inferences, results, or predictions, as examples of prediction / inference data 1022. For example, during the prediction phase 1010, the trained machine learning program 1002 generates an output. The query data 1028 is provided as input to the trained machine learning program 1002, and in response to the receipt of the query data 1028, the trained machine learning program 1002 generates the prediction / inference data 1022 as output.

[0135] In some examples, the trained machine learning program 1002 can be a generative AI model. Generative AI is a term that can refer to any type of artificial intelligence that is capable of creating new content from the training data 1006. For example, generative AI can produce text, images, videos, audio, code, or synthetic data that is similar to but not identical to the original data.

[0136] Some techniques that can be used in generative AI are:

[0137] · Convolutional Neural Network (CNN): CNN can be used for image recognition and computer vision tasks. For example, a CNN can be designed to extract features from an image by using filters or kernels that scan the input image and highlight important patterns.

[0138] · Recurrent Neural Network (RNN): An RNN can be used, for example, to process sequential data such as speech, text, and time series data. An RNN employs a feedback loop that allows it to capture temporal dependencies and remember past inputs.

[0139] · Generative Adversarial Network (GAN): A GNN can include two neural networks: a generator and a discriminator. The generator network attempts to create realistic content that can "fool" the discriminator network, while the discriminator network attempts to distinguish between realistic and fake content. The generator network and the discriminator network compete with each other and improve over time.

[0140] · Variational Autoencoder (VAE): A VAE can encode input data into a latent space (e.g., a compressed representation) and then decode it back to output data. The latent space can be manipulated to generate new variations of the output data. VAEs can use self-attention mechanisms to process input data, allowing them to handle long text sequences and capture complex dependencies.

[0141] · Transformer Model: The Transformer model can use the attention mechanism to learn the relationships between different parts of the input data (such as words or pixels) and generate output data based on these relationships. The Transformer model can process sequential data such as text or speech, as well as non-sequential data such as images or code.

[0142] In the generative AI example, the output prediction / inference data 222 includes predictions, transformations, summaries, or media content.

[0143] Figure 11 FIG. 1100 is a block diagram showing a software architecture 1104 that can be installed on any one or more of the devices described herein. The software architecture 1104 is supported by hardware such as a machine 1102 including a processor 1120, a memory 1126, and an I / O component interface 1138. In this example, the software architecture 1104 can be conceptually represented as a stack of layers, where each layer provides a specific function. The software architecture 1104 includes layers such as an operating system 1112, libraries 1108, frameworks 1110, and applications 1106. In operation, the application 1106 activates an API call 750 through the software stack and receives a message 1152 in response to the API call 1150.

[0144] The operating system 1112 manages hardware resources and provides common services. The operating system 1112 includes, for example, a kernel 1114, services 1116, and drivers 1122. The kernel 1114 serves as an abstraction layer between the hardware and other software layers. For example, the kernel 1114 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security settings. The services 1116 can provide other common services for other software layers. The drivers 1122 are responsible for controlling or interfacing with the underlying hardware. For example, the drivers 1122 can include a display driver, a camera device driver, or a low-power driver, a flash driver, a serial communication driver (e.g., a Universal Serial Bus (USB) driver), a driver, an audio driver, a power management driver, etc.

[0145] Library 1108 provides low-level common infrastructure used by application 1106. Library 1108 may include system libraries 1118 (e.g., C standard library), which provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, library 1108 may include API libraries 1124, such as media libraries (e.g., libraries for supporting the presentation and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), High Efficiency Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), graphics libraries (e.g., OpenGL framework for rendering two-dimensional (2D) and three-dimensional (3D) graphical content on a display, GLMotif for implementing user interfaces), image feature extraction libraries (e.g., OpenIMAJ), database libraries (e.g., SQLite that provides various relational database functions), web libraries (e.g., WebKit that provides web browsing functions), etc. Library 1108 may also include various other libraries 1128 to provide many other APIs to application 1106.

[0146] Framework 1110 provides high-level common infrastructure used by application 1106. For example, framework 1110 provides various graphical user interface (GUI) functions, advanced resource management, and advanced location services. Framework 1110 may provide a wide range of other APIs that can be used by application 1106, some of which may be specific to a particular operating system or platform.

[0147] In an example, application 1106 may include a home application 1136, a contacts application 1130, a browser application 1132, a book reader application 1134, a location application 1142, a media application 1144, a messaging application 1146, a gaming application 1148, and a wide variety of other applications, such as third-party application 1140. Application 1106 is a program that executes functions defined in the program. One or more of the applications 1106 can be created using various programming languages, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In a particular example, a third-party application 1140 (e.g., an application developed using an ANDROID TM or IOS TM software development kit (SDK)) can be an application on platforms such as IOS TM 、ANDROID TM 、 Mobile software running on a mobile operating system such as Phone or another mobile operating system. In this example, the third-party application 1140 can activate the API call 1150 provided by the operating system 1112 to facilitate the functions described herein.

[0148] Figure 12 is a block diagram showing a networked system 1200 including details of the glasses 100 according to some examples. The networked system 1200 includes the glasses 100, a client device 1226, and a server system 1232. The client device 1226 can be a smart phone, a tablet, a phablet, a laptop computer, an access point, or any other such device capable of connecting to the glasses 100 using a low-power wireless connection 1236 and / or a high-speed wireless connection 1234. The client device 1226 is connected to the server system 1232 via a network 1230. The network 1230 can include any combination of wired and wireless connections. The server system 1232 can be one or more computing devices that are part of a service or network computing system. Any element of the client device 1226, as well as the server system 1232 and the network 830, can be implemented using Figure 11 and Figure 3 the software architecture 1104 or the details of the machine 300 described therein.

[0149] The glasses 100 include a data processor 1202, a display 1210, one or more imaging devices 1208, and additional input / output elements 1216. The input / output elements 1216 can include a microphone, an audio speaker, a biometric sensor, additional sensors, or additional display elements integrated with the data processor 1202. Examples of the input / output elements 1216 are referenced in Figure 11 and Figure 3 and discussed further. For example, the input / output elements 1216 can include any of the I / O device interfaces 306 including an output component interface 328, a motion component interface 336, etc. Examples of the display 1210 are discussed in Figure 2 In the specific example described herein, the display 1210 includes displays for the user's left and right eyes.

[0150] The data processor 1202 includes an image processor 1206 (e.g., a video processor), a GPU & display driver 1238, a tracking component 1240, an interface 1212, a low-power circuitry 1204, and a high-speed circuitry 1220. The components of the data processor 1202 are interconnected by a bus 1242.

[0151] Interface 1212 refers to any source of user commands provided to data processor 1202. In one or more examples, interface 1212 is a physical button that, when pressed, sends a user input signal from interface 1212 to low-power processor 1214. Low-power processor 1214 may process the pressing and then immediate release of such a button as a request to capture a single image, and vice versa. Low-power processor 1214 may process the pressing of such a button for a first time period as a request to capture video data when the button is pressed and stop video capture when the button is released, where the video captured when the button is pressed is stored as a single video file. Alternatively, a button press for an extended period of time may capture a still image. In some examples, interface 1212 may be any mechanical switch or physical interface capable of accepting user input associated with a data request from imaging device 1208. In other examples, interface 1212 may have a software component or may be associated with commands received wirelessly from another source such as client device 1226.

[0152] Image processor 1206 includes circuitry for receiving signals from imaging device 1208 and processing those signals from imaging device 1208 into a format suitable for storage in memory 1224 or suitable for transmission to client device 1226. In one or more examples, image processor 1206 (e.g., a video processor) includes a microprocessor integrated circuit (IC) customized for processing sensor data from imaging device 1208 and volatile memory used by the microprocessor in operation.

[0153] Low-power circuitry 1204 includes low-power processor 1214 and low-power radio circuitry 1218. These elements of low-power circuitry 1204 may be implemented as separate elements or may be implemented on a single IC as part of a single system-on-chip. Low-power processor 1214 includes logic for managing other elements of glasses 100. As described above, for example, low-power processor 1214 may accept user input signals from interface 1212. Low-power processor 1214 may also be configured to receive input signals or instruction communications from client device 1226 via low-power wireless connection 1236. Low-power radio circuitry 1218 includes circuit elements for implementing a low-power wireless communication system. Also known as Bluetooth TM Low Energy Bluetooth TM Smart (Bluetooth TM Smart) is a standard implementation of a low-power wireless communication system that may be used to implement low-power radio circuitry 1218. In other examples, other low-power communication systems may be used.

[0154] The high-speed circuit system 1220 includes a high-speed processor 1222, a memory 1224, and a high-speed wireless circuit system 1228. The high-speed processor 1222 can be any processor capable of managing the operations of any general computing system for the data processor 1202 and high-speed communications. The high-speed processor 1222 includes processing resources for managing high-speed data transmission on the high-speed wireless connection 1234 using the high-speed wireless circuit system 1228. In some examples, the high-speed processor 1222 runs an operating system such as the LINUX operating system or other such operating systems such as Figure 11 the operating system 1112. In addition to any other duties, the high-speed processor 1222 that runs the software architecture for the data processor 1202 is also used to manage data transmission with the high-speed wireless circuit system 1228. In some examples, the high-speed wireless circuit system 1228 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, which is also known as Wi-Fi herein. In other examples, the high-speed wireless circuit system 1228 may implement other high-speed communication standards.

[0155] The memory 1224 includes any storage device capable of storing camera device data generated by the camera device 1208 and the image processor 1206. Although the memory 1224 is shown as integrated with the high-speed circuit system 1220, in other examples, the memory 1224 may be a separate stand-alone component of the data processor 1202. In some such examples, electrical wiring may provide a connection from the image processor 1206 or the low-power processor 1214 to the memory 1224 through a chip including the high-speed processor 1222. In other examples, the high-speed processor 1222 may manage the addressing of the memory 1224 such that the low-power processor 1214 will initiate the high-speed processor 1222 at any time when a read operation or a write operation involving the memory 1224 is desired.

[0156] The tracking component 1240 estimates the pose of the glasses 100. For example, the tracking component 1240 uses image data and associated inertial data and GPS data from the camera device 1208 and the position component interface 340 to track the position and determine the pose of the glasses 100 relative to a reference frame (e.g., a real-world scene environment). The tracking component 1240 continuously collects and uses updated sensor data describing the movement of the glasses 100 to determine an updated three-dimensional pose of the glasses 100, which indicates changes in the relative position and orientation with respect to physical objects in the real-world scene environment. The tracking component 1240 allows the glasses 100 to visually place virtual objects relative to physical objects within the user's field of view via the display 1210.

[0157] The GPU & Display Driver 1238 can use the pose of the glasses 100 to generate frames of virtual content or other content to be presented on the display 1210 when the glasses 100 operate in a traditional augmented reality mode. In this mode, the GPU & Display Driver 1238 generates updated frames of virtual content based on the updated three-dimensional pose of the glasses 100, which reflects changes in the position and orientation of the user relative to physical objects in the user's real-world scene environment.

[0158] One or more of the functions or operations described herein may also be performed in an application residing on the glasses 100 or the client device 1226 or a remote server. For example, one or more of the functions or operations described herein may be performed by one of the applications 1106 such as the messaging application 1146.

[0159] Figure 13 FIG. is a block diagram showing an example messaging system 1300 for exchanging data (e.g., messages and associated content) over a network. The messaging system 1300 includes multiple instances of the client device 1226 that host multiple applications including the messaging client 1302 and other applications 1304. The messaging client 1302 is communicatively coupled via a network 1230 (e.g., the Internet) to the messaging server system 1306, a third-party server 1308, and other instances of the messaging client 1302 (e.g., hosted on corresponding other client devices 1226). The messaging client 1302 can also communicate with the locally hosted applications 1304 using an application programming interface (API).

[0160] The messaging client 1302 is capable of communicating and exchanging data with other messaging clients 1302 and with the messaging server system 1306 via the network 1230. The data exchanged between the messaging clients 1302 and between the messaging client 1302 and the messaging server system 1306 includes functions (e.g., commands for activating functions) and payload data (e.g., text, audio, video, or other multimedia data).

[0161] The messaging server system 1306 provides server-side functionality to a particular messaging client 1302 via a network 1230. Although some functions of the messaging system 1300 are described herein as being performed by the messaging client 1302 or by the messaging server system 1306, it can be a design choice whether to locate some functions within the messaging client 1302 or within the messaging server system 1306. For example, it may be technically preferable to initially deploy some technologies and functions within the messaging server system 1306, but then migrate the technologies and functions to the messaging client 1302 when the client device 1226 has sufficient processing power.

[0162] The messaging server system 1306 supports various services and operations provided to the messaging client 1302. Such operations include sending data to the messaging client 1302, receiving data from the messaging client 1302, and processing data generated by the messaging client 1302. As an example, the data can include message content, client device information, geographical location information, media enhancements and overlays, message content permanence conditions, social network information, and live event information. Data exchange within the messaging system 1300 is activated and controlled through functions available via the user interface (UI) of the messaging client 1302.

[0163] Turning now specifically to the messaging server system 1306, an application programming interface (API) server 1310 is coupled to an application server 1314 and provides a programming interface to the application server 1314. The application server 1314 is communicatively coupled to a database server 1316 that facilitates access to a database 1320, which stores data associated with messages processed by the application server 1314. Similarly, a web server 1324 is coupled to the application server 1314 and provides a web-based interface to the application server 1314. To this end, the web server 1324 processes incoming network requests via the Hypertext Transfer Protocol (HTTP) and several other related protocols.

[0164] The Application Programming Interface (API) Server 1310 receives and sends message data (e.g., commands and message payloads) between the client device 1226 and the application server 1314. Specifically, the Application Programming Interface (API) Server 1310 provides a set of interfaces (e.g., routines and protocols) that can be invoked or queried by the messaging client 1302 to activate the functions of the application server 1314. The Application Programming Interface (API) Server 1310 discloses various functions supported by the application server 1314, including: account registration, login function, sending a message from a specific messaging client 1302 to another messaging client 1302 via the application server 1314, sending a media file (e.g., an image or a video) from the messaging client 1302 to the messaging server 1312, and possible access for another messaging client 1302, setting of a media data collection (e.g., a story), retrieval of a friend list of a user of the client device 1226, retrieval of such collections, retrieval of messages and content, addition and deletion of entities (e.g., friends) to and from an entity graph (e.g., a social graph), the position of a friend in the social graph, and opening of an application event (e.g., related to the messaging client 1302).

[0165] The application server 1314 hosts multiple server applications and subsystems, including, for example, the messaging server 1312, the image processing server 1318, and the social network server 1322. The messaging server 1312 implements multiple message processing techniques and functions, particularly those related to the aggregation and other processing of the content (e.g., text and multimedia content) included in messages received from multiple instances of the messaging client 1302. As will be described in further detail, text and media content from multiple sources can be aggregated into collections of content (e.g., referred to as stories or galleries). These collections are then made available to the messaging client 1302. Given the hardware requirements for other processor- and memory-intensive processing, the processing of the data can also be performed on the server side by the messaging server 1312.

[0166] The application server 1314 also includes an image processing server 1318 that is dedicated to performing various image processing operations typically on images or videos within the payload of messages sent from or received at the messaging server 1312.

[0167] The social network server 1322 supports various social networking functions and services and makes these functions and services available to the messaging server 1312. To this end, the social network server 1322 maintains and accesses an entity graph within the database 1320. Examples of functions and services supported by the social network server 1322 include identifying other users of the messaging system 1300 with whom a particular user has a relationship or "follows", and also include identifying the interests of a particular user and other entities.

[0168] The messaging client 1302 can notify a user of the client device 1226 or other users associated with the user (e.g., "friends") of activities occurring in a shared or sharable session. For example, the messaging client 1302 can provide notifications related to the current or recent use of a game by one or more members of a user group to participants in a conversation (e.g., a chat session) in the messaging client 1302. One or more users can be invited to join an active session or initiate a new session. In some examples, a shared session can provide a shared augmented reality experience in which multiple people can collaborate or participate.

[0169] Additional examples include:

[0170] Example 1 is a computer-implemented method, including: determining, by one or more processors, EMF tracking data of one or more wrists of a user using an electromagnetic field (EMF) tracking system; determining, by the one or more processors, VIO tracking data of the head of the user using a visual inertial odometry (VIO) tracking system; determining, by the one or more processors, head pose data of the head of the user and wrist pose data of the one or more wrists of the user based on the EMF tracking data and the VIO tracking data; generating, by the one or more processors, 3D body model data of the user based on the head pose data and the wrist pose data; and transmitting, by the one or more processors, the 3D body model data to an AR application for use in an augmented reality (AR) user interface of the user.

[0171] In example 2, the subject matter of example 1 includes, wherein determining the head pose data and the wrist pose data further includes: determining inertial measurement unit (IMU) tracking data of one or more EMF tracking sensors of the EMF tracking system; detecting interference in the EMF tracking data based on the EMF tracking data; and correcting the EMF tracking data based on the IMU tracking data.

[0172] In Example 3, the subject matter of any one of Examples 1 to 2 includes: determining IMU tracking data of one or more EMF tracking sensors of the EMF tracking system; and using the EMF tracking data to correct long-term drift in the IMU tracking data.

[0173] In Example 4, the subject matter of any one of Examples 1 to 3 includes, wherein the EMF tracking system includes one or more wrist-mountable EMF tracking sensors, and the EMF tracking data is determined from the one or more wrist-mountable EMF tracking sensors.

[0174] In Example 5, the subject matter of any one of Examples 1 to 4 includes, wherein the EMF tracking system further includes a head-mounted EMF source having a fixed relationship with the VIO tracking system, and wherein the wrist pose data is determined based on the pose of the VIO tracking system and the relative pose of the one wrist-mountable EMF tracking sensor.

[0175] In Example 6, the subject matter of any one of Examples 1 to 5 includes: determining ground plane data by the one or more processors based on the VIO tracking data; and further generating the 3D body model data based on the ground plane data.

[0176] In Example 7, the subject matter of any one of Examples 1 to 6 includes: correcting the EMF tracking data by the one or more processors using a history of previous EMF position tracking data.

[0177] Example 8 is at least one machine-readable medium including instructions that, when executed by a processing circuitry, cause the processing circuitry to perform operations to implement any one of Examples 1 to 7.

[0178] Example 9 is an apparatus including means for implementing any one of Examples 1 to 7.

[0179] Example 10 is a system for implementing any one of Examples 1 to 7.

[0180] "Carrier signal" means any non-tangible medium capable of storing, encoding, or carrying instructions executable by a machine, and includes digital or analog communication signals or other non-tangible media facilitating the communication of such instructions. The instructions may be sent or received via a network interface device over a network using a transmission medium.

[0181] "Client device" refers to any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. A client device can be, but is not limited to, a mobile phone, desktop computer, laptop computer, portable digital assistant (PDA), smartphone, tablet, ultrabook, netbook, notebook, multiprocessor system, microprocessor-based or programmable consumer electronics, game console, set-top box, or any other communication device that a user can use to access a network.

[0182] "Communication network" refers to one or more portions of a network, which can be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the public switched telephone network (PSTN), plain old telephone service (POTS) network, cellular telephone network, wireless network, a network, another type of network, or a combination of two or more such networks. For example, a network or a portion of a network can include a wireless network or a cellular network, and the coupling can be a code division multiple access (CDMA) connection, global system for mobile communications (GSM) connection, or other type of cellular or wireless coupling. In this example, the coupling can implement any one of various types of data transmission technologies, such as single carrier radio transmission technology (1xRTT), evolved data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rates for GSM evolution (EDGE) technology, third generation partnership project (3GPP) including 3G, fourth generation wireless (4G) network, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), worldwide interoperability for microwave access (WiMAX), long term evolution (LTE) standard, other data transmission technologies defined by various standards-setting organizations, other long-range protocols, or other data transmission technologies.

[0183] "Machine-readable medium" refers to both machine storage media and transmission media. Thus, these terms include both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium", "machine-readable medium", and "device-readable medium" mean the same thing and can be used interchangeably in this disclosure.

[0184] "Machine storage medium" refers to a single or multiple storage devices and / or media (e.g., centralized or distributed databases, and / or associated caches and servers) that store executable instructions, routines, and / or data. The term includes, but is not limited to, solid-state memories as well as optical and magnetic media, including memories internal or external to a processor. Specific examples of machine storage media, computer storage media, and / or device storage media include, by way of example, non-volatile memories including semiconductor storage devices such as erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), FPGAs, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium", "device storage medium", and "computer storage medium" mean the same thing and may be used interchangeably throughout this disclosure. The terms "machine storage medium", "computer storage medium", and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are subsumed within the term "signal medium".

[0185] "Processor" refers to any circuit or virtual circuit (a physical circuit simulated by logic running on an actual processor) that manipulates data values according to control signals (e.g., "commands", "opcodes", "machine codes", etc.) and produces associated output signals that are applied to operate a machine. For example, a processor can be a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), or any combination thereof. A processor can also be a multi-core processor having two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously.

[0186] "Signal medium" refers to any intangible medium that is capable of storing, encoding, or carrying instructions executable by a machine and includes digital or analog communication signals or other intangible media that facilitate the communication of software or data. The term "signal medium" shall be deemed to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably throughout this disclosure.

[0187] Changes and modifications may be made to the disclosed examples without departing from the scope of the disclosure. These and other changes or modifications are intended to be included within the scope of the disclosure as expressed in the appended claims.

Claims

1. A computer-implemented method, comprising: determining, by one or more processors, EMF tracking data of one or more wrists of a user using an electromagnetic field (EMF) tracking system; determining, by the one or more processors, VIO tracking data of the head of the user using a visual inertial odometry (VIO) tracking system; determining, by the one or more processors, head pose data of the head of the user and wrist pose data of the one or more wrists of the user based on the EMF tracking data and the VIO tracking data; generating, by the one or more processors, 3D body model data of the user based on the head pose data and the wrist pose data; and transmitting, by the one or more processors, the 3D body model data to an XR application for use in an extended reality (XR) user interface of the user.

2. The computer-implemented method according to claim 1, wherein determining the head pose data and the wrist pose data further comprises: determining inertial measurement unit (IMU) tracking data of one or more EMF tracking sensors of the EMF tracking system; detecting interference in the EMF tracking data based on the EMF tracking data; and correcting the EMF tracking data based on the IMU tracking data.

3. The computer-implemented method according to claim 2, further comprising: determining IMU tracking data of one or more EMF tracking sensors of the EMF tracking system; and correcting long-term drift in the IMU tracking data using the EMF tracking data.

4. The computer-implemented method according to claim 1, wherein the EMF tracking system includes one or more wrist-mountable EMF tracking sensors, and the EMF tracking data is determined from the one or more wrist-mountable EMF tracking sensors.

5. The computer-implemented method according to claim 4, wherein the EMF tracking system further includes a head-mounted EMF source having a fixed relationship with the VIO tracking system, and wherein the wrist pose data is determined based on the pose of the VIO tracking system and the relative pose of one wrist-mountable EMF tracking sensor.

6. The computer-implemented method according to claim 1, further comprising: determining, by the one or more processors, ground plane data based on the VIO tracking data; and further generating the 3D body model data based on the ground plane data.

7. The computer-implemented method according to claim 1, further comprising: correcting, by the one or more processors, the EMF tracking data using a history of previous EMF position tracking data.

8. A machine, comprising: one or more processors; and one or more memories storing instructions that, when executed by the one or more processors, cause the machine to perform operations, the operations including: determining EMF tracking data of one or more wrists of a user using an electromagnetic field (EMF) tracking system; Determine VIO tracking data of the user's head using a visual inertial odometry (VIO) tracking system; Determine head pose data of the user's head and wrist pose data of the user's one or more wrists based on the EMF tracking data and the VIO tracking data; Generate 3D body model data of the user based on the head pose data and the wrist pose data; and Transmit the 3D body model data to an XR application for use in an extended reality (XR) user interface of the user.

9. The machine according to claim 8, wherein, Determining the head pose data and the wrist pose data further includes: Determining inertial measurement unit (IMU) tracking data of one or more EMF tracking sensors of the EMF tracking system; Detecting interference in the EMF tracking data based on the EMF tracking data; and Correcting the EMF tracking data based on the IMU tracking data.

10. The machine according to claim 9, wherein, The operation further includes: Determining IMU tracking data of one or more EMF tracking sensors of the EMF tracking system; and Using the EMF tracking data to correct long-term drift in the IMU tracking data.

11. The machine according to claim 8, wherein, The EMF tracking system includes one or more wrist-mountable EMF tracking sensors, and the EMF tracking data is determined from the one or more wrist-mountable EMF tracking sensors.

12. The machine according to claim 11, wherein, The EMF tracking system further includes a head-mounted EMF source in a fixed relationship with the VIO tracking system, and wherein the wrist pose data is determined based on the pose of the VIO tracking system and the relative pose of one wrist-mountable EMF tracking sensor.

13. The machine according to claim 8, wherein, The operation further includes: Determining ground plane data based on the VIO tracking data by the one or more processors; and Generating the 3D body model data further based on the ground plane data.

14. The machine according to claim 8, wherein, The operation further includes: Correcting the EMF tracking data by the one or more processors using a history of previous EMF position tracking data.

15. A machine storage medium storing instructions that, when executed by one or more processors of a machine, cause the machine to perform operations, the operations include: Determine EMF tracking data of one or more wrists of a user using an electromagnetic field (EMF) tracking system; Determine VIO tracking data of the user's head using a visual inertial odometry (VIO) tracking system; Determine head pose data of the user's head and wrist pose data of the user's one or more wrists based on the EMF tracking data and the VIO tracking data; Generate 3D body model data of the user based on the head pose data and the wrist pose data; and Transmit the 3D body model data to an XR application for use in the user's extended reality (XR) user interface.

16. The machine-readable storage medium according to claim 15, wherein, determining the head pose data and the wrist pose data further comprises: determining inertial measurement unit (IMU) tracking data of one or more EMF tracking sensors of the EMF tracking system; detecting interference in the EMF tracking data based on the EMF tracking data; and correcting the EMF tracking data based on the IMU tracking data.

17. The machine-readable storage medium according to claim 16, wherein, the operation further comprises: determining IMU tracking data of one or more EMF tracking sensors of the EMF tracking system; and correcting long-term drift in the IMU tracking data using the EMF tracking data.

18. The machine-readable storage medium according to claim 15, wherein, the EMF tracking system includes one or more wrist-mountable EMF tracking sensors, and the EMF tracking data is determined from the one or more wrist-mountable EMF tracking sensors.

19. The machine-readable storage medium according to claim 18, wherein, the EMF tracking system further includes a head-mounted EMF source in a fixed relationship with the VIO tracking system, and wherein the wrist pose data is determined based on the pose of the VIO tracking system and the relative pose of the one wrist-mountable EMF tracking sensor.

20. The machine-readable storage medium according to claim 15, wherein, the operation further comprises: determining ground plane data by the one or more processors based on the VIO tracking data; and generating the 3D body model data further based on the ground plane data.