Estimating body pose based on image and head pose information

By utilizing head pose information to determine the image search region and performing calculations only on the search region, the problem of computationally intensive large images is solved, achieving efficient body pose estimation while reducing power consumption and time costs.

CN121816602APending Publication Date: 2026-04-07QUALCOMM INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing body pose estimation techniques are computationally intensive when processing large images, resulting in excessive power consumption and time costs, and they cannot effectively save computational operations for users who are interested in them.

Method used

By utilizing head pose information to determine the search area of ​​the image, calculations are performed only on the search area to estimate body pose, reducing the scope of computational operations.

Benefits of technology

By reducing the computational area, computational resources are saved, computational efficiency is improved, and the focus is placed on the pose estimation of people that users are interested in, thereby reducing power consumption and time costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121816602A_ABST
    Figure CN121816602A_ABST
Patent Text Reader

Abstract

Systems and techniques for pose estimation of a person are described herein. For example, a method for pose estimation of a person is provided. The method may include: obtaining an image of a person; obtaining a head posture of a person; determining a search area of the image based on the head pose; and determining a gesture of the person based on the search area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates in general to estimating body pose. For example, aspects of this disclosure include systems and techniques for estimating head pose based on images and head pose information. Background Technology

[0002] Body pose estimation techniques seek to estimate the posture of a human body. Some body pose estimation techniques use a trained machine learning model to generate body poses based on images of the body. As an example, a machine learning model can be trained by providing images. The machine learning model can generate body poses based on the images. The body poses generated by the machine learning model can be compared with ground truth body poses corresponding to the images. The parameters (e.g., weights) of the machine learning model can be adjusted based on the differences (e.g., errors) between the body poses generated by the machine learning model and the ground truth body poses. After being trained (e.g., using many images and corresponding ground truth body poses), the machine learning model can be used to infer body poses based on images (e.g., newly generated images in real time). Summary of the Invention

[0003] The following is a simplified summary of the invention relating to one or more aspects disclosed herein. Therefore, this summary should not be considered an exhaustive overview relating to all conceived aspects, nor should it be considered to identify key or decisive elements relating to all conceived aspects or to depict the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts in a simplified form relating to one or more aspects of the mechanisms disclosed herein, preceding the detailed description that follows.

[0004] Systems and techniques for human pose estimation are described. Based on at least one example, a method for human pose estimation is provided. The method includes: acquiring an image of a person; acquiring the person's head pose; determining a search region of the image based on the head pose; and determining the person's pose based on the search region.

[0005] In another example, an apparatus for human pose estimation is provided, the apparatus comprising: at least one memory; and at least one processor (e.g., configured in a circuit) coupled to the at least one memory. The at least one processor is configured to: acquire an image of a person; acquire a head pose of the person; determine a search region of the image based on the head pose; and determine the human pose based on the search region.

[0006] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to: acquire an image of a person; acquire a head pose of the person; determine a search region of the image based on the head pose; and determine the pose of the person based on the search region.

[0007] In another example, an apparatus for human pose estimation is provided. The apparatus includes: components for acquiring an image of a person; components for acquiring the person's head pose; components for determining a search region of the image based on the head pose; and components for determining the person's pose based on the search region.

[0008] In some aspects, one or more of the devices described herein are, may be part of, or may include: mobile devices (e.g., mobile phones or so-called "smartphones," tablet computers, or other types of mobile devices), extended reality devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), vehicles (or computing devices or systems of vehicles), smart or connected devices (e.g., Internet of Things (IoT) devices), wearable devices, personal computers, laptop computers, video servers, televisions (e.g., network-connected televisions), robotic devices or systems, or other devices. In some aspects, each device may include one image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each device may include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each device may include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each device may include one or more sensors. In some cases, one or more sensors may be used to determine the location of the device, the state of the device (e.g., tracking state, operating state, temperature, humidity level, and / or another state), and / or for other purposes.

[0009] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim.

[0010] The foregoing and other features and aspects will become more apparent from the following description, claims and accompanying drawings. Attached Figure Description

[0011] The following description, with reference to the accompanying drawings, details exemplary examples of this application:

[0012] Figure 1 This is a block diagram illustrating the architecture of an example extended reality (XR) system according to some aspects of this disclosure;

[0013] Figure 2A These are illustrations of scenarios in which a system according to various aspects of this disclosure can be used to determine the posture of a human body;

[0014] Figure 2B This is a block diagram illustrating various aspects of a system for determining the posture of a human body according to the present disclosure;

[0015] Figure 3 It is a block diagram of an apparatus for determining the posture of a human body according to various aspects of this disclosure;

[0016] Figure 4 This illustrates various aspects of this disclosure. Figure 2A An illustration of the scene, to illustrate... Figure 2A and Figure 2B Systems and / or Figure 3 The various concepts described in the equipment description;

[0017] Figure 5 This is an illustration of the search area of ​​an image according to various aspects of this disclosure, to illustrate regarding... Figure 2A and Figure 2B Systems and / or Figure 3 The various concepts described in the equipment description;

[0018] Figure 6 It is a block diagram of an apparatus for determining the posture of a human body according to various aspects of this disclosure;

[0019] Figure 7 This is a flowchart illustrating another example process for estimating body posture according to various aspects of this disclosure;

[0020] Figure 8 This is a block diagram illustrating examples of deep learning neural networks that can be used to implement a perception module and / or one or more verification modules, based on some aspects of the disclosed techniques.

[0021] Figure 9 This is a block diagram illustrating examples of convolutional neural networks (CNNs) according to various aspects of this disclosure; and

[0022] Figure 10 This is a block diagram illustrating an example computing device architecture that can implement the various technologies described herein. Detailed Implementation

[0023] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently, and some may be applied in combination, as will be apparent to those skilled in the art. Specific details are set forth in the following description for purposes of explanation in order to provide a thorough understanding of the various aspects of this application. However, it will be apparent that various aspects may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0024] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0025] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as superior to or better than other aspects. Similarly, the term “aspects of this disclosure” does not require that all aspects of this disclosure include the features, advantages, or modes of operation discussed.

[0026] As previously mentioned, machine learning models can be trained to generate body poses based on images of the body. In this disclosure, the term "pose" can refer to the position and / or orientation of an object. Position can be defined according to three degrees of freedom, for example, along three orthogonal axes, such as the x-axis, y-axis, and z-axis, relative to a coordinate system and / or a reference point. Orientation can be defined according to three additional degrees of freedom, such as roll, pitch, and yaw. In this disclosure, the term "pose" (e.g., "body pose," "body posture," or "person's pose") referring to the body or a person can refer to the pose of one or more corresponding parts of a person's body. For example, body pose can include the pose of a person's torso, shoulders, hips, and corresponding limbs. In this disclosure, the term "pose" (e.g., "head pose" or "head posture") referring to the head can refer to the pose of a person's head (including position and orientation). In this disclosure, the term "pose" (e.g., "hand pose" or "hand posture") referring to one or more hands can refer to the pose of one or both hands of a person (including position and orientation).

[0027] Body pose estimation techniques can be computationally intensive. For example, many computational operations (which may consume power and / or take time) may be required to perform the body pose estimation technique. Furthermore, the computational intensity can be exacerbated when such body pose estimation techniques operate on large images (e.g., images with a large field of view and / or images comprising millions of pixels). For example, the neural network of a body pose estimation technique may include input neurons for each pixel of the input image and may store the activation of each input neuron. Therefore, larger input images require larger neural networks and / or more operations (e.g., associated with activations between neurons) to process. As an example of a computationally intensive image, a megapixel image with a large field of view may include several people. These people may be relatively far from the camera and may each occupy only tens of thousands of pixels (e.g., a portion of the camera's field of view). In some cases, the user of the camera capturing the image may only be interested in one or a few of these people. The body pose estimation technique may operate on the entire image. Furthermore, the body pose estimation technique may seek to estimate the corresponding body pose of each of these people. This may waste computational operations, for example, by processing pixels that do not include people and by estimating the poses of people that the user is not interested in.

[0028] This document describes systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively, “Systems and Techniques”) for estimating body pose based on head pose. The systems and techniques described herein can acquire images of a person and obtain the person’s head pose (e.g., using a camera). For example, the person may be wearing a head-mounted device (e.g., a head-mounted display, such as an extended reality (XR) headset). The head-mounted device can determine the person’s head pose (e.g., using an inertial measurement unit (IMU) and / or simultaneous localization and mapping (SLAM) techniques). The person’s head-mounted device can send the person’s head pose to the systems and techniques. The systems and techniques can determine a search region in an image based on the head pose. For example, the systems and techniques can determine a search region within an image that can represent the person based on the person’s head pose. The systems and techniques can determine the person’s pose based on the search region. For example, the systems and techniques can feed the search region to a trained pose estimation model that can generate the person’s pose.

[0029] By providing a search region (rather than the entire image) to a trained pose estimation model, this system and technique save computational operations. For example, the trained pose estimation model can operate only on the search region of the image (which may represent a person) instead of the entire image, thus performing fewer activation computations between the input layer and other layers of the trained pose estimation model. Furthermore, by providing the search region of the image to the trained pose estimation model, this system and technique save computational operations by limiting the number of people whose poses are determined by the trained pose estimation model to only those that the user of the system and technique is interested in.

[0030] Various aspects of this application will be described below with reference to the accompanying drawings.

[0031] Figure 1 This is a diagram illustrating the architecture of an example extended reality (XR) system 100 according to some aspects of this disclosure. The XR system 100 can execute XR applications and implement XR operations.

[0032] In this exemplary example, the XR system 100 includes one or more image sensors 102, accelerometers 104, gyroscopes 106, storage devices 108, input devices 110, displays 112, computing components 114, XR engines 124, image processing engines 126, rendering engines 128, and communication engines 130. It should be noted that... Figure 1 The components 102-130 shown are non-limiting examples provided for illustrative and explanatory purposes, and other examples may include those with... Figure 1 The components shown may be more numerous, fewer, or different than those shown. For example, in some cases, the XR system 100 may include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radar, light detection and ranging (LIDAR) sensors, radio detection and ranging (RADAR) sensors, sound detection and ranging (SODAR) sensors, sound navigation and ranging (SONAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or Figure 1 One or more other software and / or hardware components not shown. While various components of the XR system 100 (such as image sensor 102) may be referred to herein in the singular, it should be understood that the XR system 100 may include multiple components discussed herein (e.g., multiple image sensors 102).

[0033] Display 112 may be or may include glass, screen, lens, projector and / or other display mechanism that allows users to see a real-world environment and also allows XR content to be overlaid on, overlapped with, mixed with or otherwise displayed on the real-world environment.

[0034] XR system 100 may include input device 110 or be able to communicate with such input device (wired or wireless). Input device 110 may include any suitable input device, such as a touchscreen, pen or other pointing device, keyboard, mouse, buttons or keys, microphone for receiving voice commands, gesture input device for receiving gesture commands, video game controller, steering wheel, joystick, set of buttons, trackball, remote control, any other input device discussed herein, or any combination thereof. In some cases, image sensor 102 may capture images that can be processed to interpret gesture commands.

[0035] The XR system 100 can also communicate with one or more other electronic devices (wired or wireless). For example, the communication engine 130 can be configured to manage connections and communicate with one or more electronic devices. In some cases, the communication engine 130 may correspond to... Figure 10 The communication interface is 1026.

[0036] In some embodiments, the image sensor 102, accelerometer 104, gyroscope 106, storage device 108, display 112, computing component 114, XR engine 124, image processing engine 126, and rendering engine 128 may be part of the same computing device. For example, in some cases, the image sensor 102, accelerometer 104, gyroscope 106, storage device 108, display 112, computing component 114, XR engine 124, image processing engine 126, and rendering engine 128 may be integrated into an HMD, extended reality glasses, a smartphone, a laptop computer, a tablet computer, a gaming system, and / or any other computing device. However, in some embodiments, the image sensor 102, accelerometer 104, gyroscope 106, storage device 108, display 112, computing component 114, XR engine 124, image processing engine 126, and rendering engine 128 may be part of two or more independent computing devices. For example, in some cases, some of the components 102-130 may be part of or implemented by a computing device, and the remaining components may be part of or implemented by one or more other computing devices. For instance, such as in a discrete-sensory XR system, XR system 100 may include a first device (e.g., an HMD) that includes a display 112, an image sensor 102, an accelerometer 104, a gyroscope 106, and / or one or more computing components 114. XR system 100 may also include a second device that includes additional computing components 114 (e.g., implementing an XR engine 124, an image processing engine 126, a rendering engine 128, and / or a communication engine 130). In such examples, the second device may generate virtual content based on information or data (e.g., images, sensor data, such as measurements from accelerometer 104 and gyroscope 106) and may provide the virtual content to the first device for display at the first device. The second device may be or may include a smartphone, laptop computer, tablet computer, personal computer, gaming system, server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device, or a mobile device acting as a server device), any other computing device and / or combinations thereof.

[0037] Storage device 108 can be any storage device used for storing data. Furthermore, storage device 108 can store data from any component of the XR system 100. For example, storage device 108 can store data from image sensor 102 (e.g., image or video data), data from accelerometer 104 (e.g., measurements), data from gyroscope 106 (e.g., measurements), data from computing component 114 (e.g., processing parameters, preferences, virtual content, rendered content, scene maps, tracking and positioning data, object detection data, privacy data, XR application data, facial recognition data, occlusion data, etc.), data from XR engine 124, data from image processing engine 126, and / or data from rendering engine 128 (e.g., output frames). In some examples, storage device 108 may include a buffer for storing frames processed by computing component 114.

[0038] Computing component 114 may be or may include a central processing unit (CPU) 116, a graphics processing unit (GPU) 118, a digital signal processor (DSP) 120, an image signal processor (ISP) 122, and / or other processors (e.g., a neural processing unit (NPU) implementing one or more trained neural networks). Computing component 114 may perform various operations such as image enhancement, computer vision, graphics rendering, extended reality operations (e.g., tracking, localization, pose estimation, map building, content anchoring, content rendering, prediction, etc.), image and / or video processing, sensor processing, recognition (e.g., text recognition, face recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), trained machine learning operations, filtering, and / or any of the various operations described herein. In some examples, computing component 114 may implement (e.g., control, operate, etc.) an XR engine 124, an image processing engine 126, and a rendering engine 128. In other examples, computing component 114 may also implement one or more other processing engines.

[0039] Image sensor 102 may include any image and / or video sensor or capture device. In some examples, image sensor 102 may be part of a multi-camera assembly, such as a dual-camera assembly. Image sensor 102 may capture image and / or video content (e.g., raw image and / or video data), which may then be processed by computing component 114, XR engine 124, image processing engine 126, and / or rendering engine 128, as described herein.

[0040] In some examples, image sensor 102 may capture image data and may generate an image (also referred to as a frame) based on that image data and / or may provide the image data or frame to XR engine 124, image processing engine 126, and / or rendering engine 128 for processing. The image or frame may include a video frame in a video sequence or a still image. The image or frame may include an array of pixels representing a scene. For example, the image may be: a red-green-blue (RGB) image with red, green, and blue color components per pixel; a lightness, redness, and blueness (YCbCr) image with a lightness component and two chromaticity (redness and blueness) components per pixel; or any other suitable type of color or monochrome image.

[0041] In some cases, image sensor 102 (and / or other cameras of XR system 100) may be configured to also capture depth information. For example, in some implementations, image sensor 102 (and / or other cameras) may include an RGB depth (RGB-D) camera. In some cases, XR system 100 may include one or more depth sensors (not shown) that are separate from image sensor 102 (and / or other cameras) and can capture depth information. For example, such depth sensors may acquire depth information independently of image sensor 102. In some examples, depth sensors may be physically mounted in the same general location or orientation as image sensor 102, but may operate at a different frequency or frame rate than image sensor 102. In some examples, depth sensors may take the form of a light source that projects a structured or textured light pattern (which may include one or more narrowband lights) onto one or more objects in a scene. Depth information can then be obtained by utilizing the geometric deformation of the projected pattern caused by the surface shape of the objects. In one example, depth information may be obtained from a stereo sensor, such as a combination of an infrared structured light projector and an infrared camera registered to a camera (e.g., an RGB camera).

[0042] The XR system 100 may also include other sensors among its one or more sensors. The one or more sensors may include one or more accelerometers (e.g., accelerometer 104), one or more gyroscopes (e.g., gyroscope 106), and / or other sensors. The one or more sensors may provide velocity, orientation, and / or other position-related information to the computing component 114. For example, accelerometer 104 may detect the acceleration of the XR system 100 and may generate an acceleration measurement based on the detected acceleration. In some cases, accelerometer 104 may provide one or more translation vectors (e.g., up / down, left / right, forward / backward) that can be used to determine the position or attitude of the XR system 100. Gyroscope 106 may detect and measure the orientation and angular velocity of the XR system 100. For example, gyroscope 106 may be used to measure the pitch, roll, and yaw of the XR system 100. In some cases, gyroscope 106 may provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, the image sensor 102 and / or the XR engine 124 may use measurements obtained by the accelerometer 104 (e.g., one or more translation vectors) and / or measurements obtained by the gyroscope 106 (e.g., one or more rotation vectors) to calculate the attitude of the XR system 100. As previously noted, in other examples, the XR system 100 may also include other sensors such as an inertial measurement unit (IMU), a magnetometer, a gaze and / or eye-tracking sensor, a machine vision sensor, a smart scene sensor, a voice recognition sensor, a shock sensor, a vibration sensor, a position sensor, a tilt sensor, etc.

[0043] As noted above, in some cases, one or more sensors may include at least one IMU. An IMU is an electronic device that uses a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers to measure specific forces, angular velocities, and / or orientations of the XR system 100. In some examples, one or more sensors may output measurement information associated with the capture of images by the image sensor 102 (and / or other cameras of the XR system 100) and / or depth information obtained using one or more depth sensors of the XR system 100.

[0044] The XR engine 124 can use the output of one or more sensors (e.g., accelerometer 104, gyroscope 106, one or more IMUs and / or other sensors) to determine the attitude of the XR system 100 (also referred to as head attitude) and / or the attitude of the image sensor 102 (or other cameras of the XR system 100). In some cases, the attitude of the XR system 100 and the attitude of the image sensor 102 (or other cameras) can be the same. The attitude of the image sensor 102 refers to the position and orientation of the image sensor 102 relative to (e.g., relative to the field of view) a reference frame. In some specific implementations, the camera attitude can be determined for 6 degrees of freedom (6DoF), which refers to three translational components (e.g., which can be given by X (horizontal), Y (vertical), and Z (depth) coordinates relative to a reference frame such as the image plane) and three angular components (e.g., roll, pitch, and yaw relative to the same reference frame). In some implementations, camera attitude can be determined for 3 degrees of freedom (3DoF), which refers to three angular components (e.g., roll, pitch, and yaw).

[0045] In some cases, a device tracker (not shown) may use measurements from one or more sensors and image data from image sensor 102 to track the pose (e.g., 6DoF pose) of the XR system 100. For example, the device tracker may fuse visual data from the image data (e.g., using a visual tracking solution) with inertial data from the measurements to determine the position and motion of the XR system 100 relative to the physical world (e.g., a scene) and a map of the physical world. As described below, in some examples, when tracking the pose of the XR system 100, the device tracker may generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate updates to the 3D map for that scene. 3D map updates may include, for example, but not limited to, new or updated features and / or features or landmarks associated with the scene and / or the 3D map of that scene, localization updates identifying or updating the position of the XR system 100 within the scene and the 3D map of that scene, etc. The 3D map provides a digital representation of the scene in the real / physical world. In some examples, 3D maps can anchor location-based objects and / or content to real-world coordinates and / or objects. XR system 200 can use mapped scenes (e.g., scenes in the physical world represented by a 3D map and / or scenes associated with that 3D map) to merge the physical and virtual worlds and / or merge virtual content or objects with the physical environment.

[0046] In some aspects, computing component 114 may use a visual tracking solution to determine and / or track the pose of image sensor 102 and / or the XR system 100 as a whole, based on images captured by image sensor 102 (and / or other cameras of XR system 100). For example, in some examples, computing component 114 may use computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques to perform tracking. For example, computing component 114 may perform SLAM or may communicate (wired or wirelessly) with a SLAM engine (not shown). SLAM refers to a class of techniques that create a map of an environment (e.g., a map of the environment modeled by XR system 100) while tracking the pose of the camera (e.g., image sensor 102) and / or XR system 100 relative to that map. This map may be called a SLAM map and may be three-dimensional (3D). SLAM technology can be performed using color or grayscale image data captured by image sensor 102 (and / or other cameras of XR system 100) and can be used to generate an estimate of the 6DoF attitude measurement of image sensor 102 and / or XR system 100. Such SLAM technology configured to perform 6DoF tracking can be referred to as 6DoF SLAM. In some cases, the output of one or more sensors (e.g., accelerometer 104, gyroscope 106, one or more IMUs and / or other sensors) can be used to estimate, correct, and / or otherwise adjust the estimated attitude.

[0047] Figure 2A This is an illustration of a scenario 202 in which the system 200 according to various aspects of this disclosure can be used to determine the posture of a person 234's body 238. Figure 2B This is a block diagram illustrating a system 200 for determining the posture of a person 234's body 238 according to various aspects of the present disclosure. According to various aspects of the present disclosure, system 200 includes a head-mounted device 240 wearable on the head 236 of the person 234. System 200 also includes a head-mounted device 208, provided as an example for the purpose of describing systems and techniques implementing the present disclosure. Typically, head-mounted device 240 can determine the posture of the person 234's head 236. Head-mounted device 240 can send posture information 244 indicating the posture of the head 236 to head-mounted device 208. Head-mounted device 208 can capture an image 252 of the person 234. Head-mounted device 208 can determine the posture of the person 234's body 238 based on the posture of the head 236 and the image 252.

[0048] Each of head-mounted devices 208, 216, 228, and 240 may be Figure 1An example of an XR system 100 is provided. For the purpose of describing the systems and techniques implementing this disclosure, a head-mounted device 208 for user 204 is provided as an example. The system and techniques may be implemented by any or all of head-mounted devices 216, 228, and / or 240. Furthermore, the system and techniques may be implemented by another device (e.g., another head-mounted device or another device, such as a gaming system including a camera or a computing system including a camera).

[0049] Each of head-mounted devices 208, 216, 228, and 240 may include a corresponding posture determiner for determining the posture of the respective head-mounted device. For example, head-mounted device 208 may include a position determiner 210 for determining the posture of head-mounted device 208 (and / or the head 206 of the user 204 wearing head-mounted device 208), head-mounted device 216 may include a position determiner 218 for determining the posture of head-mounted device 216 (and / or the head 214 of the person 212 wearing head-mounted device 216), head-mounted device 228 may include a position determiner 230 for determining the posture of head-mounted device 228 (and / or the head 224 of the person 222 wearing head-mounted device 228), and head-mounted device 240 may include a position determiner 242 for determining the posture of head-mounted device 240 (and / or the head 236 of the person wearing head-mounted device 240).

[0050] In some aspects, each of head-mounted devices 208, 216, 228, and 240 may include a corresponding accelerometer 104, a corresponding gyroscope 106, and / or a corresponding inertial measurement unit (IMU). Each of head-mounted devices 208, 216, 228, and 240 may determine its corresponding attitude (and / or the attitude of the head wearing it) based on data from its corresponding accelerometer 104, its corresponding gyroscope 106, and / or its corresponding inertial measurement unit (IMU). For example, head-mounted device 208 (e.g., using position determiner 210) can determine the attitude of head 206 based on data from its accelerometer 104, gyroscope 106 and / or IMU, head-mounted device 216 (e.g., using position determiner 218) can determine the attitude of head 214 based on data from its accelerometer 104, gyroscope 106 and / or IMU, head-mounted device 228 (e.g., using position determiner 230) can determine the attitude of head 224 based on data from its accelerometer 104, gyroscope 106 and / or IMU, and head-mounted device 240 (e.g., using position determiner 242) can determine the attitude of head 236 based on data from its accelerometer 104, gyroscope 106 and / or IMU.

[0051] In some aspects, each of head-mounted devices 208, 216, 228, and 240 may include a corresponding image sensor 102 capable of capturing images. Additionally, each of head-mounted devices 208, 216, 228, and 240 may implement computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques based on images captured by its corresponding image sensor 102. For example, head-mounted device 208 (e.g., using position determiner 210) can determine the pose of head 206 based on images from its image sensor 102, head-mounted device 216 (e.g., using position determiner 218) can determine the pose of head 214 based on images from its image sensor 102, head-mounted device 228 (e.g., using position determiner 230) can determine the pose of head 224 based on images from its image sensor 102, and head-mounted device 240 (e.g., using position determiner 242) can determine the pose of head 236 based on images from its image sensor 102.

[0052] Additionally, each of head-mounted devices 208, 216, 228, and 240 may include a corresponding communication engine 130, which can be used to communicate with other head-mounted devices among head-mounted devices 208, 216, 228, and 240. Each of head-mounted devices 208, 216, 228, and 240 may provide its determined head posture to each of the other head-mounted devices among head-mounted devices 208, 216, 228, and 240. For example, head-mounted device 216 can send posture information 220 indicating the posture of head 214 to head-mounted device 208, head-mounted device 228 can send posture information 232 indicating the posture of head 224 to head-mounted device 208, and head-mounted device 240 can send posture information 244 indicating the posture of head 236 to head-mounted device 208. Although in Figure 2B Not illustrated, but head-mounted device 216 may send posture information 220 to head-mounted device 228 and / or head-mounted device 240, head-mounted device 228 may send posture information 232 to head-mounted device 216 and / or head-mounted device 240, head-mounted device 240 may send posture information 244 to head-mounted device 216 and / or head-mounted device 228, and / or head-mounted device 208 may send posture information indicating the posture of head 206 to head-mounted device 216, head-mounted device 228 and / or head-mounted device 240.

[0053] In some respects, head-mounted devices 208, 216, 228, and 240 can establish a common reference system. For example, all head-mounted devices 208, 216, 228, and 240 can establish a common coordinate system for describing attitude. Therefore, when head-mounted devices 208, 216, 228, and 240 transmit attitude information, that attitude information can be relative to the common reference system.

[0054] As an example of the operation of the systems and techniques disclosed herein, head-mounted device 208 may determine the posture of the body 238 of person 234 based on the posture of head 236 (e.g., as received by head-mounted device 208 in posture information 244) and image 252. Regarding Figures 3 to 6 Additional details are described regarding how the head-mounted device 208 can determine the pose of the boundary volume generator 308. In some aspects, the system and technique can provide the pose of the head 236 as a prior to a pose estimation machine learning model, and enable the pose estimation machine learning model to use the pose of the head 236 as a prior to determine the pose of the body 238 based on image 252.

[0055] In some aspects, one or more of head-mounted devices 208, 216, 228, and 240 may be associated with a corresponding handheld device. Handheld device 248 is provided as an example. Head-mounted device 240 is associated with handheld device 248. A person 234 may hold handheld device 248 in their hand 246. Head-mounted device 240 may determine the orientation of handheld device 248. In some aspects, handheld device 248 may include an accelerometer, gyroscope, and / or IMU. In such aspects, handheld device 248 may generate motion data 254 based on the accelerometer, gyroscope, and / or IMU, and transmit the motion data 254 to head-mounted device 240. Head-mounted device 240 may determine the orientation of handheld device 248 based on motion data 254. Additionally or alternatively, in some aspects, handheld device 248 may determine its pose (e.g., based on motion data 254) and may transmit the pose of handheld device 248 to head-mounted device 240. Additionally or alternatively, in some aspects, head-mounted device 240 may capture an image 256 of handheld device 248 and may (e.g., using a tracking algorithm) determine the pose of handheld device 248. In some aspects, head-mounted device 208 may determine the pose of person 234 at least in part based on the pose of handheld device 248. For example, head-mounted device 240 may transmit the pose of handheld device 248 to head-mounted device 208 (e.g., as part of pose information 244 or separately), and head-mounted device 208 may use the pose of handheld device 248 when determining the pose of person 234. In some aspects, the system and technique may provide the pose of hand 246 as a prior to a pose estimation machine learning model, and enable the pose estimation machine learning model to use the pose of hand 246 as a prior to determine the pose of body 238 based on image 252. Additionally or alternatively, the system and technique may provide the pose of hand 246 and head 236 as priors to a pose estimation machine learning model, and enable the pose estimation machine learning model to use the pose of hand 246 and head 236 as priors to determine the pose of body 238 based on image 252.

[0056] Figure 3 This is a block diagram of a device 300 for determining the posture 324 of a person 234's body 238, according to various aspects of this disclosure. Provided Figure 3Device 300 is provided as an example of a system and technology for implementing this disclosure. Device 300 may be an example of head-mounted device 208, head-mounted device 216, or head-mounted device 228. Additionally or alternatively, device 300 may be an example of another device or system (e.g., a computing device including a camera or a gaming system including a camera). Head-mounted device 208, head-mounted device 216, head-mounted device 228 and / or head-mounted device 228 and / or the aforementioned other devices or systems may include [related to device 300 and] Figure 3 The described components are substantially similar or identical. Furthermore, head-mounted devices 208, 216, 228, and 240, and / or other aforementioned devices or systems, can perform operations related to device 300 and... Figure 3 The described operation. For example, a gaming system that includes a camera or a computing system that includes a camera may include information about... Figure 3 The device 300 describes the elements and is capable of performing the following: Figure 3 The described operation.

[0057] Figure 4 This illustrates various aspects of this disclosure. Figure 2A The illustration of image 304 in scene 202 is used to illustrate the relevant information. Figure 2A and Figure 2B System 200 and / or Figure 3 The device 300 describes various concepts. Posture information 306 can be an example of an image captured and used according to the system and technology. For example (regardless of the viewpoint), image 304 can be an example of image 252 described with respect to head-mounted device 208 and person 234. For example, image 304 can be an example of an image captured by head-mounted device 208 of person 234. Alternatively, image 304 can be captured by another device. Figure 2A Example of an image for scene 202.

[0058] Figure 5 This is an illustration of the search area 320 of an image (e.g., image 304) according to various aspects of this disclosure, to illustrate regarding Figure 2A and Figure 2B System 200 and / or Figure 3 The device 300 describes various concepts. Search area 320 can be an example of a search area for an image captured and used according to the system and technology. For example (regardless of the viewing angle), search area 320 can be an example of a search area for an image 252 described with respect to head-mounted device 208 and person 234. For example, search area 320 can be an example of a search area for an image captured by head-mounted device 208 of person 234. Alternatively, search area 320 can be captured by another device. Figure 2A Example of the search area for the image in scene 202.

[0059] Go to Figure 3 The camera 302 of device 300 (which can be used with XR system 100) Figure 1 An image sensor 102 (which is the same as or substantially similar to the image sensor 102) can capture an image 304. Image 304 may include a person 234. In other words, image 304 may include pixels representing a person 234.

[0060] Device 300 can receive posture information 306, which can indicate the head posture 402 of person 234's head 236 (e.g., Figure 4 (As illustrated). For example, person 234 may wear head-mounted device 240 on head 236. Head-mounted device 240 may determine head pose 402 of person 234's head 236 (e.g., based on data from accelerometers, gyroscopes, and / or IMUs, and / or using SLAM technology based on images captured by head-mounted device 240). Head-mounted device 240 may send pose information 306 to device 300. As an example, device 300 may be an example of head-mounted device 208, and head-mounted device 240 may send pose information 244 to head-mounted device 208 (e.g., as shown in the example). Figure 2B (as described).

[0061] The boundary volume generator 308 of device 300 can generate a boundary volume 310 based on the head pose 402 received in pose information 244. The boundary volume 310 can be a three-dimensional volume in the three-dimensional coordinate system of device 300. In some cases, the boundary volume 310 can be in a common reference frame shared by device 300 and head-mounted device 240, and / or other head-mounted devices in head-mounted device 208, head-mounted device 216, and head-mounted device 228.

[0062] Boundary volume generator 308 can generate boundary volume 310 based on head pose 402. For example, boundary volume generator 308 can generate boundary volume 310 below head pose 402. Furthermore, in some aspects, boundary volume generator 308 can generate boundary volume 310 based on SLAM technology (e.g., based on ground and / or walls defined by SLAM technology). Additionally, boundary volume generator 308 can generate boundary volume 310 based on predetermined parameters, which are based on the human body. For example, boundary volume generator 308 can generate boundary volume 310 with a lateral extension of one meter based on an arm length parameter. Boundary volume 310 in... Figure 4 The boundary volume 310 is exemplified as a cylinder. The boundary volume 310 can have any shape. For example, the boundary volume 310 can be a cylinder, a sphere, a box, a spot, a mesh, etc.

[0063] The boundary volume projector 312 of device 300 can generate a mask 316 based on the boundary volume 310. For example, the boundary volume projector 312 can project the three-dimensional boundary volume 310 onto an image plane (e.g., the image plane of image 304). Because Figure 4 It is a two-dimensional image, so Figure 4 The outer boundary of the boundary volume 310 appearing in the image can represent the outer boundary of the boundary volume 310 projected onto the image plane. The boundary volume projector 312 can generate a mask 316 based on the boundary volume 310 projected onto the image plane. For example, the boundary volume projector 312 can generate a mask 316 to correspond to the outer boundary of the boundary volume 310 projected onto the image 304. Therefore, the mask 316 can be two-dimensional and corresponds to the outer boundary of the boundary volume 310 projected onto the image 304. Furthermore, the mask 316 can correspond to the position of the person 234 in the image 304. For example, the mask 316 can define the pixels in the image 304 that can represent the person 234. For example, since the boundary volume 310 is defined based on parameters of the human body and head pose 402, and since the boundary volume 310 is the projection of the boundary volume 310 onto the image 304, the mask 316 can indicate the pixels in the image 304 that represent the body 238 of the person 234.

[0064] The mask applicator 318 of device 300 can apply mask 316 to image 304 to generate search region 320. Search region 320 can be a portion of image 304 defined by mask 316 (e.g., as shown in the image). Figure 4 and Figure 5 (As illustrated by the comparison between them). The search area 320 defined by mask 316 may include pixels representing person 234.

[0065] The pose estimator 322 of device 300 can estimate the pose 324 of the body 238 of person 234 based on search region 320. For example, the pose estimator 322 can be a trained pose estimation model. More specifically, the pose estimator 322 can be trained to receive images of the body and determine the body pose based on those images. Pose 324 can include the pose of various parts of the body 238 of person 234 in six degrees of freedom. Pose 324 in Figure 5 The middle is illustrated as including multiple lines corresponding to the shoulders, torso, hips, arms, and legs of person 234. Because Figure 5 It is a two-dimensional image, so pose 324 is... Figure 5 While it appears to be two-dimensional, pose 324 can be three-dimensional, depending on the three-dimensional coordinate system of device 300 (which can be shared by head-mounted device 240).

[0066] Device 300 may output a pose 324. Device 300 (which may be one of head-mounted device 208, head-mounted device 216, or head-mounted device 228) or another device may use the pose 324 for various purposes, including, for example, as input and / or for extended reality (XR) purposes. For example, device 300 may track the pose 324 of the body 238 of person 234 and use the pose 324 as input (e.g., to control a character in a game). Additionally or alternatively, an XR device (e.g., head-mounted device 208) may track the pose 324 and use the generated content to cover the body 238 of person 234 within the user 204's field of view.

[0067] In some aspects, device 300 may include a verifier 326 that can use pose information 306 to verify pose 324. For example, verifier 326 may receive pose 324 from pose estimator 322 and (e.g., from head-mounted device 240) pose information 306, and verify pose 324 based on pose information 306. For example, verifier 326 may determine the accuracy of pose 324 based on pose information 306. For example, verifier 326 may use pose information 306 to understand head pose 402 of head 236 in the three-dimensional coordinate system of device 300. Verifier 326 may use head pose 402 to verify pose 324 of body 238 of person 234 in the three-dimensional coordinate system of device 300. In some aspects, verifier 326 may generate an indicative confidence value 328 based on the correctness of pose 324. Device 300 may output (e.g., confidence value 328 of pose 324).

[0068] As previously mentioned, according to various aspects of this disclosure, a device (e.g., device 300) can determine the posture 324 of a person 234 based on the posture of a handheld device 248. For example, as previously described, a head-mounted device 240 can determine the posture of the handheld device 248 (e.g., based on motion data 254 generated by sensors of the handheld device 248 and / or based on images of the handheld device 248 captured by the head-mounted device 240). The head-mounted device 240 can provide the posture of the handheld device 248 to the device 300, and the device 300 can determine the posture 324 at least in part based on the posture of the handheld device 248. For example, a boundary volume generator 308 can determine a boundary volume 310 at least in part based on the posture of the handheld device 248. For example, the boundary volume generator 308 can generate a boundary volume 310 that includes the handheld device 248, and, in the case that the boundary volume 310 has posture information of two handheld devices, the boundary volume 310 can generate a boundary volume 310 that does not include spaces where neither of the two handheld devices exists. Figure 3 The remainder of the operation described for device 300 may be performed as described above based on the boundary volume 310 generated according to the orientation of handheld device 248.

[0069] By generating pose 324 based on search region 320 (instead of image 304), pose estimator 322 can generate pose 324 faster (e.g., consuming less time and / or power) than pose estimator 322 generating pose 324 based on image 304. Therefore, device 300 can save time and / or power in generating pose 324. Furthermore, by generating pose 324 based on search region 320 (instead of image 304), pose estimator 322 can save time and / or power if the pose of person 212 is uncertain. For example, the system capturing image 304 may determine that user 204 and / or person 212 is not of interest (e.g., based on the fact that user 204 and / or person 212 is not wearing a head-mounted device, which is consistent with...). Figure 2A and Figure 4 Conversely, as illustrated in the example, or based on some other factors, such as user input.

[0070] Figure 6 This is a block diagram of a device 600 for determining the posture 624 of a person's body (e.g., body 238) according to various aspects of this disclosure. Provided Figure 6 Device 600 is provided as an example of a system and technology for implementing this disclosure. Device 600 may be an example of head-mounted device 208, head-mounted device 216, or head-mounted device 228. Additionally or alternatively, device 600 may be an example of another device or system (e.g., a computing device including a camera or a gaming system including a camera). Head-mounted device 208, head-mounted device 216, head-mounted device 228 and / or head-mounted device 228 and / or the aforementioned other devices or systems may include [related to device 600 and] Figure 6 The described components are substantially similar or identical. Furthermore, head-mounted devices 208, 216, 228, and 240, and / or other aforementioned devices or systems, can perform operations related to device 600 and... Figure 6 The described operation. For example, a gaming system that includes a camera or a computing system that includes a camera may include information about... Figure 6 The device 600 describes the element and is capable of performing the following: Figure 6 The described operation.

[0071] The camera 602 of the device 600 (which is compatible with the XR system 100) Figure 1 An image sensor 102 (which is the same as or substantially similar to) can capture an image 604. Image 604 may include a person (e.g., person 234). In other words, image 604 may include pixels representing a person.

[0072] Device 600 can receive attitude information 606, which can indicate head posture (e.g., Figure 4Head pose 402). For example, a human-wearable head-mounted device (e.g., head-mounted device 240). The head-mounted device can determine the head pose (e.g., based on data from an accelerometer, gyroscope, and / or IMU, and / or using SLAM technology based on images captured by the head-mounted device). The head-mounted device can send pose information 606 to device 600. As an example, device 600 can be an example of head-mounted device 208, and head-mounted device 240 can send pose information 244 to head-mounted device 208 (e.g., as per the head pose 402). Figure 2B (as described).

[0073] In some aspects, the posture information 606 may also include the posture of a person's hand. For example, a head-mounted device may be associated with a handheld device, and the posture of the handheld device may be determined (e.g., based on the handheld device's IMU and / or based on an image of the handheld device captured by the head-mounted device). The head-mounted device may transmit the handheld device posture to device 600 (e.g., in the posture information 606).

[0074] The pose estimator 630 of device 600 can determine a person's pose 624 based on image 604 and pose information 606. Pose 624 can be compared with... Figure 3 and Figure 5 The pose 624 is the same as or substantially similar to the pose 624. The pose estimator 630 can be a machine learning model trained to determine the body pose based on an image of the body and a head pose. As an example, the pose estimator 630 can be trained by providing an image and a head pose. The pose estimator 630 can generate a body pose based on the image and head pose. The body pose generated by the pose estimator 630 can be compared to a ground truth body pose corresponding to the image and head pose. The parameters (e.g., weights) of the pose estimator 630 can be adjusted based on the difference (e.g., error) between the body pose generated by the pose estimator 630 and the ground truth body pose. After being trained (e.g., using many images, head poses, and corresponding ground truth body poses), the pose estimator 630 can be deployed in device 600 and used to infer pose 624 based on image 604 and pose information 606.

[0075] Figure 7This is a flowchart illustrating a process 700 for estimating body posture according to various aspects of this disclosure. One or more operations of process 700 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable device such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a vehicle or a component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capability to perform process 700. One or more operations of process 700 may be implemented as software components that execute and run on one or more processors.

[0076] At box 702, a computing device (or one or more components thereof) can acquire an image of a person. For example, a head-mounted device 208 of system 200 can capture an image 252 of a person 234.

[0077] At box 704, a computing device (or one or more components thereof) can obtain the head pose of a person. For example, a head-mounted device 208 can determine the head pose of the head 236 of a person 234.

[0078] In some respects, a person's head posture can be determined by a head-mounted device on the person's head. For example, a head-mounted device 240 on the head 236 of person 234 can determine the head posture of person 234.

[0079] In some aspects, a person's head pose can be determined based on at least one of the following: motion data from one or more inertial measurement units (IMUs) of a head-mounted device, or based on simultaneous localization and mapping (SLAM) techniques using images captured at the head-mounted device. For example, a head-mounted device 240 on the head 236 of person 234 may include an IMU and may determine the head pose of the head 236 of person 234. Additionally or alternatively, the head-mounted device 240 may perform SLAM techniques and use SLAM techniques to determine the head pose of the head 236.

[0080] In some respects, in order to obtain head pose, a computing device (or one or more components thereof) may receive head pose from a head-mounted device on a person's head. For example, head-mounted device 208 may receive the head pose of head 236 from head-mounted device 240.

[0081] In some aspects, a computing device (or one or more components thereof) can determine and transmit a user's head pose. For example, head-mounted device 208 can determine the head pose of user 204's head 206 and transmit that head pose to other HMDs in, for example, system 200. By transmitting the head pose of user 204's head 206, head-mounted device 208 can enable other HMDs in system 200 to determine user 204's pose, for example, using process 700.

[0082] In some aspects, the user's head pose may be determined based on at least one of the following: motion data from one or more inertial measurement units (IMUs) of a head-mounted device on the user's head, or based on simultaneous localization and mapping (SLAM) techniques using images captured at the head-mounted device. For example, a head-mounted device 208 on the head 206 of user 204 may include an IMU and may determine the head pose of the head 206 of user 204. Additionally or alternatively, the head-mounted device 208 may perform SLAM techniques and use SLAM techniques to determine the head pose of the head 206.

[0083] At box 706, a computing device (or one or more components thereof) may determine the search region of an image based on head pose. For example, head-mounted device 208 may determine the search region 320 of image 304 based on the head pose of person 234's head 236.

[0084] At box 708, a computing device (or one or more components thereof) may determine a person's pose based on a search region. For example, a head-mounted device 208 may determine a person 234's pose 324 based on a search region 320 of image 304.

[0085] In some respects, a computing device (or one or more components thereof) may determine the boundary volume of a person's body based on the person's head pose. The person's pose may be further determined based on this boundary volume. For example, a head-mounted device 208 may determine the boundary volume 310 based on the head pose of the head 236 of the person 234. The head-mounted device 208 may determine the pose 324 of the person 234 based on the boundary volume 310.

[0086] In some aspects, a computing device (or one or more components thereof) may determine a mask for an image based on a boundary volume. This mask may be associated with a person's position in the image. The person's pose may be further determined based on this mask. For example, a head-mounted device 208 may determine a mask 316 for an image 304 based on a boundary volume 310. The boundary volume 310 may be associated with the position of a person 234 in the image 304 (e.g., the head pose of the person 234's head 236 has been determined based on the boundary volume 310). The head-mounted device 208 may determine the pose 324 of the person 234 based on the mask 316.

[0087] In some respects, to determine a mask, a computing device (or one or more components thereof) may project a boundary volume onto an image and define the mask based on the two-dimensional projection of the boundary volume onto the image. For example, a head-mounted device 208 may project a boundary volume 310 onto an image 304 and define a mask 316 as the two-dimensional projection of the boundary volume 310 onto the image 304.

[0088] In some respects, to determine a person's pose, a computing device (or one or more components thereof) may determine the person's pose based on a search region of an image. The search region is defined by a mask. For example, search region 320 may be defined by mask 316. Head-mounted device 208 may determine the pose 324 of person 234 based on search region 320.

[0089] In some respects, a person's pose can be determined using a pose estimation machine learning model that is trained to determine pose based on an image. For example, device 300 can use pose estimator 322 to determine pose 324, which can be trained to determine pose based on an image.

[0090] In some respects, a person's pose can be determined using a pose estimation machine learning model trained to determine pose based on an image. A computing device (or one or more components thereof) can provide the head pose as a priori information to the pose estimation machine learning model. For example, device 300 can use pose estimator 322 to determine pose 324, which can be trained to determine pose based on an image. Furthermore, head-mounted device 208 can provide the head pose of person 234's head 236 as a priori information to pose estimator 322.

[0091] In some aspects, a computing device (or one or more components thereof) may determine a confidence value for a person's posture based on a comparison between the person's pose and the person's head pose, wherein the confidence value is associated with the confidence of using the person's posture. For example, device 300 may include a verifier 326 that may determine a confidence value 328 based on the head pose of head 236 and pose 324. For example, verifier 326 may determine a confidence value 328 based on a comparison between the torso and head poses of pose 324.

[0092] In some respects, a computing device (or one or more components thereof) can acquire a person's hand gesture. The person's posture can then be further determined based on that hand gesture. For example, a head-mounted device 208 can acquire the hand gesture of a person 234's hand 246. The head-mounted device 208 can determine the posture 324 of person 234 based on the hand gesture of hand 246.

[0093] In some aspects, a person's hand posture can be determined by a head-mounted device on the person's head. For example, the hand posture of person 234's head 236 can be determined by a head-mounted device 240 wearable on person 234's head 236. In some aspects, a person's hand posture can be determined based on at least one of the following: hand tracking techniques based on one or more images captured at the head-mounted device, or motion data from an inertial measurement unit (IMU) in a handheld device in the person's hand. For example, head-mounted device 240 can capture an image 252 of person 234's hand 246 and determine the hand posture of hand 246 based on image 252. Additionally or alternatively, hand 246 may hold handheld device 248. Handheld device 248 may include an IMU. Handheld device 248 may provide motion data 254 to head-mounted device 240, and head-mounted device 240 may determine the hand posture of hand 246 based on motion data 254.

[0094] In some aspects, a computing device (or one or more components thereof) may determine the boundary volume of a person's body based on the person's head pose and hand pose. The person's pose may be further determined based on this boundary volume. For example, head-mounted device 208 may determine the boundary volume 310 based on the head pose of the person 234's head 236 and / or based on the hand pose of the person 234's hand 246. Head-mounted device 208 may also determine the pose 324 of the person 234 based on the boundary volume 310.

[0095] In some respects, a person's posture can be determined using a posture estimation machine learning model trained to determine posture based on images. A computing device (or one or more components thereof) can provide hand posture as a priori information to the posture estimation machine learning model. For example, device 300 can use posture estimator 322 to determine posture 324, which can be trained to determine posture based on images. Furthermore, head-mounted device 208 can provide the hand posture of person 234's hand 246 as a priori information to posture estimator 322.

[0096] In some aspects, a person's posture can be determined using a posture estimation machine learning model trained to determine posture based on images, and at least one processor is further configured to determine the upper body posture based on hand posture using inverse kinematics, and to provide this upper body posture as a prior to the posture estimation machine learning model. For example, device 300 can use a posture estimator 322 trained to determine posture based on images to determine posture 324. Device 300 can also use inverse kinematics to determine the upper body posture of person 234's body 238 based on hand posture. Furthermore, head-mounted device 208 can provide the hand posture of person 234's hand 246 as a prior to the posture estimator 322.

[0097] In some aspects, the computing device (or one or more components thereof) may determine a confidence value for a person's posture based on a comparison between the person's hand posture and the person's hand posture. For example, device 300 may include a verifier 326 that may determine a confidence value 328 based on the hand posture 246 and posture 324. For example, verifier 326 may determine a confidence value 328 based on a comparison between the hand or arm posture 324 and the hand posture.

[0098] In some respects, a person's pose can be determined using a pose estimation machine learning model that is trained to determine the pose based on an image and head pose. For example, device 600 can determine the pose 624 of pose 324 based on an image 604 of person 234 and pose information 606 (which may include the head pose of person 234's head 236).

[0099] In some respects, a computing device (or one or more components thereof) can acquire a person's hand pose. A pose estimation machine learning model can be trained to determine the pose based on an image, head pose, and hand pose. The person's pose can be further determined based on that hand pose. For example, device 600 can determine pose 624 of pose 324 based on an image 604 of person 234 and pose information 606 (which may include the head pose of person 234's head 236 and the hand pose of person 234's hand 246).

[0100] In some examples, as previously noted, the methods described herein (e.g., Figure 7 The process 700 and / or other methods described herein may be performed wholly or partially by a computing device or apparatus. In one example, one or more of these methods may be performed by... Figure 1 XR system 100, Figure 2A System 200 Figure 2A and Figure 2B Head-mounted devices 208 Figure 2A and Figure 2B Head-mounted devices 216 Figure 2A and Figure 2B Head-mounted devices 228 Figure 2A and Figure 2B Head-mounted devices 240 Figure 3 The device 300 or another system or device may execute these methods. In another example, these methods (e.g., Figure 7 One or more of the processes 700 and / or other methods described herein may be used by Figure 10 The computing device architecture 1000 shown is implemented wholly or partially. For example, it has Figure 10 The computing device of the computing device architecture 1000 shown may include or be included in Figure 1 XR system 100, Figure 2A System 200 Figure 2A and Figure 2B Head-mounted devices 208 Figure 2A and Figure 2B Head-mounted devices 216 Figure 2A and Figure 2B Head-mounted devices 228 Figure 2A and Figure 2B Head-mounted devices 240 Figure 3 The computing device 300 comprises components that enable the operation of process 700 and / or other processes described herein. In some cases, the computing device or apparatus may include various components such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other types of data.

[0101] A component capable of implementing a computing device in a circuit. For example, the component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0102] Process 700 and / or other processes described herein are illustrated as logic flowcharts, whose operations represent sequences of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0103] Additionally, process 700 and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0104] As noted above, various aspects of this disclosure may utilize machine learning models or systems.

[0105] Figure 8 This is an exemplary example of a neural network 800 (e.g., a deep learning neural network) that can be used to implement machine learning-based feature segmentation, implicit neural representation generation, rendering, classification, object detection, image recognition (e.g., face recognition, object recognition, scene recognition, etc.), feature extraction, authentication, gaze detection, gaze prediction, and / or automation. For example, a neural network 800 could be... Figure 3 pose estimator 322 and / or Figure 6 Examples of pose estimators 630, or implementations thereof.

[0106] Input layer 802 includes input data. In an exemplary example, input layer 802 may include data representing search region 320 and / or image 604 and pose information 606. Neural network 800 includes multiple hidden layers 806a, 806b to 806n. Hidden layers 806a, 806b to 806n include "n" hidden layers, where "n" is an integer greater than or equal to one. Multiple hidden layers can be made to include as many layers as needed for a given application. Neural network 800 also includes an output layer 804 that provides the output produced by the processing performed by hidden layers 806a, 806b to 806n. In an exemplary example, output layer 804 may provide pose 324 and / or pose 624.

[0107] The neural network 800 may be a multi-layer neural network with interconnected nodes or may include interconnected nodes. Each node may represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains the information while processing it. In some cases, the neural network 800 may include a feedforward network, in which case there are no feedback connections in which the network's output is fed back into itself. In some cases, the neural network 800 may include a recurrent neural network, which may have loops that allow information to be carried across nodes when reading input.

[0108] Information can be exchanged between nodes through node-to-node interconnects between layers. Nodes in input layer 802 can activate the node set in the first hidden layer 806a. For example, as shown, each input node in input layer 802 is connected to each node in the first hidden layer 806a. Nodes in the first hidden layer 806a can transform the information of each input node by applying an activation function to the input node information. The information derived from this transformation can then be passed to nodes in the next hidden layer 806b, activating those nodes, which can then perform their own specified functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable function. The output of hidden layer 806b can then activate nodes in the next hidden layer, and so on. Finally, the output of hidden layer 806n can activate one or more nodes in output layer 804, providing the output at those nodes. In some cases, although a node in neural network 800 (e.g., node 808) is shown as having multiple output lines, the node has a single output and all lines shown as outputs from the node represent the same output value.

[0109] In some cases, each node or the interconnection between nodes may have weights, which are a set of parameters derived from the training of the neural network 800. Once the neural network 800 is trained, it can be called a trained neural network, which can be used to perform one or more operations. For example, the interconnection between nodes may represent a piece of information about what the interconnected nodes have learned. The interconnection may have tunable numerical weights that can be tuned (e.g., based on the training dataset), allowing the neural network 800 to adapt to the input and learn as more and more data is processed.

[0110] The neural network 800 can be pre-trained to process features from the data in the input layer 802 using different hidden layers 806a, 806b to 806n, so as to provide an output through the output layer 804. In an example where the neural network 800 is used to identify features in an image, the neural network 800 can be trained using training data that includes both images and labels, as described above. For example, training images can be input into the network, where each training image has a label indicating features in the image (for feature segmentation machine learning systems) or a label indicating the category of activity in each image. In an example where object classification is used for illustrative purposes, the training images may include images of the number 2, in which case the label of the image may be [0 0 1 0 0 0 0 0 0 0].

[0111] In some cases, the neural network 800 can use a training process called backpropagation to adjust the weights of its nodes. As noted above, the backpropagation process can include forward pass, loss function, back pass, and weight update. For each training iteration, forward pass, loss function, back pass, and parameter update are performed. For each set of training images, this process can be repeated up to a certain number of iterations until the neural network 800 is trained well enough to accurately tune the weights of each layer.

[0112] For an example of identifying objects in an image, the forward pass may include passing a training image through a neural network 800. The weights are initially randomized before training the neural network 800. As an illustrative example, the image may include a numerical array representing the pixels of the image. Each number in the array may include a value from 0 to 255 describing the intensity of the pixel at that location in the array. In one example, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or lightness and two chroma components, etc.).

[0113] As noted above, for the first training iteration of a neural network 800, the output may include values ​​due to the weights being randomly selected during initialization without prioritizing any particular class. For example, if the output is a vector with probabilities that an object includes different classes, the probability values ​​for each class may be equal or at least very similar (e.g., for ten possible classes, each class may have a probability value of 0.1). With the initial weights, the neural network 800 cannot determine low-level features and therefore cannot make an accurate determination of what the object's classification might be. A loss function can be used to analyze the error in the output. Any suitable loss function can be defined, such as cross-entropy loss. Another example of a loss function includes mean squared error (MSE), which is defined as... The loss can be set to equal E. 总计 The value of .

[0114] For the first training image, the loss (or error) will be high because the actual value will be significantly different from the predicted output. The goal of training is to minimize the loss so that the predicted output matches the training labels. The Neural Network 800 performs backpropagation by determining which inputs (weights) contribute most to the network's loss and can adjust the weights to reduce and eventually minimize the loss. The derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a specific layer) can be calculated to determine the weights that contribute most to the network's loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. A weight update can be represented as... Where w represents the weight, wi Let represent the initial weights, and η represent the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes larger weight updates, while a lower value indicates smaller weight updates.

[0115] Neural Network 800 can include any suitable deep network. An example includes a Convolutional Neural Network (CNN), which includes an input layer and an output layer, with multiple hidden layers between them. The hidden layers of a CNN include a series of convolutional layers, non-linear layers, pooling layers (for downsampling), and fully connected layers. Neural Network 800 can include any other deep network besides CNNs, such as autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), etc.

[0116] Figure 9 This is an exemplary example of a Convolutional Neural Network (CNN) 900. The input layer 902 of the CNN 900 includes data representing an image or frame. For example, the data could include a numerical array representing pixels of an image, where each number in the array includes a value from 0 to 255 describing the pixel intensity at that location in the array. Using the previous example from above, the array could include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or lightness and two chroma components, etc.). The image can be passed through a convolutional hidden layer 904, an optional non-linear activation layer, a pooling hidden layer 906, and a fully connected layer 908 (which can be hidden) to obtain the output at the output layer 910. Although... Figure 9 Only one hidden layer from each hidden layer is shown in the diagram, but those skilled in the art will understand that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers may be included in a CNN 900. As previously described, the output may indicate a single category of an object, or may include probabilities that best describe the category of an object in an image.

[0117] The first layer of a CNN 900 can be a convolutional hidden layer 904. The convolutional hidden layer 904 analyzes the image data from the input layer 902. Each node in the convolutional hidden layer 904 is connected to a region of the input image called a receptive field (pixel). The convolutional hidden layer 904 can be thought of as one or more filters (each filter corresponding to a different activation or feature map), where each convolutional iteration of the filter is a node or neuron in the convolutional hidden layer 904. For example, the region of the input image covered by the filter at each convolutional iteration will be the filter's receptive field. In an exemplary example, if the input image consists of a 28×28 array and each filter (and its corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in the convolutional hidden layer 904. Each connection between a node and its receptive field learns weights and, in some cases, learns an overall bias, allowing each node to learn to analyze its specific local receptive field in the input image. Each node in the convolutional hidden layer 904 will have the same weights and biases (called shared weights and shared biases). For example, the filter has a weight (digital) array and the same depth as the input. For the image frame example, the filter would have a depth of 3 (based on the three color components of the input image). An exemplary example of the filter array size is 5×5×3, corresponding to the size of the receptive field of a node.

[0118] The convolutional nature of the convolutional hidden layer 904 is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filters of the convolutional hidden layer 904 can start at the top left corner of the input image array and can convolve around the input image. As noted above, each convolutional iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 904. In each convolutional iteration, the value of the filter is multiplied by the corresponding number of original pixel values ​​of the image (e.g., a 5×5 filter array is multiplied by a 5×5 array of input pixel values ​​at the top left corner of the input image array). The multiplications from each convolutional iteration can be summed to obtain the sum of that iteration or node. Next, the process continues at the next position in the input image based on the receptive field of the next node in the convolutional hidden layer 904. For example, the filter can move a step size (called stride) to the next receptive field. The stride can be set to 1 or any other suitable amount. For example, if the stride is set to 1, the filter will move 1 pixel to the right in each convolutional iteration. Processing the filter at each unique location in the input volume produces a number representing the filter result at that location, thus determining a sum value for each node of the convolutional hidden layer 904.

[0119] The mapping from the input layer to the convolutional hidden layer 904 is called an activation map (or feature map). An activation map includes node-specific values ​​representing the filter results at each location within the input volume. Activation maps can include arrays containing various sums of values ​​produced by the filter on each iteration of the input volume. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map would consist of a 24×24 array. The convolutional hidden layer 904 can include several activation maps to identify multiple features in the image. Figure 9 The example shown includes three activation maps. Using these three activation maps, the convolutional hidden layer 904 can detect three different types of features, each of which is detectable across the entire image.

[0120] In some examples, a nonlinear hidden layer can be applied after the convolutional hidden layer 904. Nonlinear layers can be used to introduce nonlinearity into a system that has already computed linear operations. An exemplary example of a nonlinear layer is the Corrected Linear Unit (ReLU) layer. A ReLU layer applies the function f(x) = max(0, x) to all values ​​in the input volume, which changes all negative activations to 0. Therefore, ReLU can add nonlinearity to the CNN 900 without affecting the receptive field of the convolutional hidden layer 904.

[0121] A pooling hidden layer 906 can be applied after the convolutional hidden layer 904 (and, in use, after the non-linear hidden layer). The pooling hidden layer 906 is used to simplify the information in the output of the convolutional hidden layer 904. For example, the pooling hidden layer 906 takes each activation map output from the convolutional hidden layer 904 and uses a pooling function to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by the pooling hidden layer. The pooling hidden layer 906 uses other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. Pooling functions (e.g., max pooling filters, L2 norm filters, or other suitable pooling filters) are applied to each activation map included in the convolutional hidden layer 904. Figure 9 In the example shown, three pooling filters are used to convolve the three activation maps in the hidden layer 904.

[0122] In some examples, max pooling can be used by applying a max pooling filter (e.g., of 2×2 size) with a stride (e.g., equal to the dimension of the filter, such as stride 2) to the activation map output from convolutional hidden layer 904. The output from the max pooling filter includes the maximum number in each sub-region of the filter convolution. Using a 2×2 filter as an example, each unit in the pooling layer summarizes a region of 2×2 nodes from the previous layer (each node is a value in the activation map). For example, four values ​​(nodes) in the activation map will be analyzed by the 2×2 max pooling filter at each iteration of the filter, with the maximum of the four values ​​being output as the "maximum" value. If such a max pooling filter is applied to an activation filter of 24×24 nodes from convolutional hidden layer 904, the output from pooling hidden layer 906 will be an array of 12×12 nodes.

[0123] In some examples, L2 norm pooling filters may also be used. L2 norm pooling filters involve calculating the square root of the sum of squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (instead of calculating the maximum value as done in max pooling), and using the calculated value as the output.

[0124] Pooling functions (e.g., max pooling, L2 norm pooling, or other pooling functions) determine whether a given feature is found anywhere within a region of the image, discarding the exact location information. This can be done without affecting the results of feature detection, because once a feature has been found, its exact location is less important than its approximate location relative to other features. Max pooling (and other pooling methods) offers the benefit of having far fewer pooling features, thus reducing the number of parameters required in subsequent layers of a CNN 900.

[0125] The final connection in the network is a fully connected layer, which connects each node from the pooling hidden layer 906 to each output node in the output layer 910. Using the example above, the input layer comprises 28×28 nodes encoding the pixel intensity of the input image, the convolutional hidden layer 904 comprises 3×24×24 hidden feature nodes based on applying a 5×5 local receptive field (for filtering) to three activation maps, and the pooling hidden layer 906 comprises a layer of 3×12×12 hidden feature nodes based on applying a max-pooling filter to a 2×2 region in each of the three feature maps. Extending this example, the output layer 910 may comprise ten output nodes. In such an example, each node of the 3×12×12 pooling hidden layer 906 is connected to each node of the output layer 910.

[0126] The fully connected layer 908 takes the output of the previous pooling hidden layer 906 (which should represent an activation map of high-level features) and determines the features most relevant to a particular class. For example, the fully connected layer 908 can determine the high-level features most relevant to a particular class and may include weights (nodes) for those high-level features. The product between the weights of the fully connected layer 908 and the pooling hidden layer 906 can be computed to obtain the probabilities for different classes. For example, if the CNN 900 is used to predict that the object in an image is a person, there will be high values ​​in the activation map representing the high-level features of a person (e.g., two legs, a face at the top of the object, two eyes at the top left and top right of the face, a nose in the middle of the face, a mouth at the bottom of the face, and / or other features common to people).

[0127] In some examples, the output from output layer 910 may include an M-dimensional vector (M=10 in the previous example). M indicates the number of classes the CNN 900 must choose from when classifying objects in an image. Other example outputs may also be provided. Each number in the M-dimensional vector represents the probability that an object belongs to a certain class. In an exemplary example, if the 10-dimensional output vector representing objects of ten different classes is [0 0 0.05 0.8 0 0.15 0 0 0 0], then the vector indicates a 5% probability that the image is an object of the third class (e.g., a dog), an 80% probability that the image is an object of the fourth class (e.g., a person), and a 15% probability that the image is an object of the sixth class (e.g., a kangaroo). The probability of a class can be considered as the confidence level that an object is part of that class.

[0128] Figure 10 An example computing device architecture 1000 is illustrated, illustrating example computing devices capable of implementing the various technologies described herein. In some examples, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device within a vehicle), or other devices. For example, computing device architecture 1000 may include, implement, or be included in... Figure 1 XR system 100, Figure 2A System 200 Figure 2A and Figure 2B Head-mounted devices 208 Figure 2A and Figure 2B Head-mounted devices 216 Figure 2A and Figure 2B Head-mounted devices 228 Figure 2A and Figure 2B Head-mounted devices 240 Figure 3 Equipment 300 Figure 6The computing device architecture 1000 may be configured to execute process 700 and / or other processes described herein.

[0129] The components of the computing device architecture 1000 are shown to communicate electrically with each other using a connection 1012, such as a bus. The example computing device architecture 1000 includes a processing unit (CPU or processor) 1002 and a computing device connection 1012 that couples various computing device components, including computing device memories 1010 (such as read-only memory (ROM) 1008 and random access memory (RAM) 1006), to the processor 1002.

[0130] The computing device architecture 1000 may include a cache of high-speed memory that is directly connected to, very close to, or integrated into the processor 1002. The computing device architecture 1000 may copy data from memory 1010 and / or storage device 1014 to cache 1004 for fast access by the processor 1002. In this way, the cache can provide performance improvements by avoiding latency for the processor 1002 while waiting for data. These and other modules may control or be configured to control the processor 1002 to perform various actions. Other computing device memory 1010 may also be used. Memory 1010 may include various different types of memory with different performance characteristics. The processor 1002 may include any general-purpose processor and hardware or software services configured to control the processor 1002 (such as services 11016, 2 1018, and 3 1020 stored in storage device 1014), as well as dedicated processors in which software instructions are incorporated into the processor design. The processor 1002 may be a self-contained system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors can be symmetric or asymmetric.

[0131] To enable user interaction with the computing device architecture 1000, input device 1022 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. Output device 1024 can also be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, projector, television, speaker equipment, etc. In some instances, a multi-mode computing device allows the user to provide multiple types of input to communicate with the computing device architecture 1000. Communication interface 1026 typically controls and manages user input and computing device output. There are no limitations on operation on any particular hardware arrangement, and therefore the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0132] Storage device 1014 is a non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as a magnetic tape cassette, flash memory card, solid-state memory device, digital multifunction disk, magnetic tape cartridge, random access memory (RAM) 1006, read-only memory (ROM) 1008, and hybrid forms thereof. Storage device 1014 may include services 1016, 1018, and 1020 for controlling processor 1002. Other hardware or software modules are envisioned. Storage device 1014 may be connected to computing device connection 1012. In one aspect, a hardware module performing a specific function may include software components stored in a computer-readable medium connected to necessary hardware components, such as processor 1002, connection 1012, output device 1024, etc., to perform that function.

[0133] In box 1102, routine 1100 acquires an image of a person. In box 1104, routine 1100 acquires the person's head pose. In box 1106, routine 1100 determines a search region of the image based on the head pose. In box 1108, routine 1100 determines the person's pose based on the search region.

[0134] In some respects, body pose estimation can be used to determine good priors for visual positioning system (VPS) localization. For example, head-mounted device 208 can acquire an image of person 234. Head-mounted device 208 can use pose estimation techniques to determine the pose 324 of person 234 based on the image. Furthermore, head-mounted device 208 can provide the pose 324 of person 234 to head-mounted device 240 of person 234. Head-mounted device 240 can use the pose 324 of person 234 as a prior to determine the person's position. Determining the position of person 234 may include determining the position and / or orientation of person 234 relative to scene 202. Head-mounted device 240 can further determine the position of person 234 based on VPS localization, for example, based on motion data from one or more inertial measurement units (IMUs) of head-mounted device 240 and / or based on simultaneous localization and mapping (SLAM) techniques on images captured at head-mounted device 240.

[0135] With reference to a given parameter, property, or condition, the term "substantially" may mean that a person skilled in the art would understand that a given parameter, property, or condition is satisfied with a small degree of variance (such as, for example, within acceptable manufacturing tolerances). For example, depending on the specific parameter, property, or condition that is substantially satisfied, the parameter, property, or condition may be satisfied at least 90%, at least 95%, or even at least 99%.

[0136] Various aspects of this disclosure are applicable to any suitable electronic device (such as a security system, smartphone, tablet, laptop, vehicle, drone, or other device) that includes or is coupled to one or more active depth sensing systems. Although devices having or coupled to a light projector are described below, various aspects of this disclosure are applicable to devices having any number of light projectors and are therefore not limited to any particular device.

[0137] The term "device" is not limited to one or a specific number of physical objects (such as a smartphone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that implement at least some parts of this disclosure. Although the following description and examples use the term "device" to describe various aspects of this disclosure, the term "device" is not limited to a specific configuration, type, or number of objects. Additionally, the term "system" is not limited to multiple components or specific aspects. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. Although the following description and examples use the term "system" to describe various aspects of this disclosure, the term "system" is not limited to a specific configuration, type, or number of objects.

[0138] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects can be practiced without these specific details. For clarity, in some cases, the technology may be presented as comprising individual functional blocks, including functional blocks comprising devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring aspects.

[0139] Various aspects described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. While flowcharts may describe operations as sequential processes, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but it may have additional steps not included in the diagrams. Processes can correspond to methods, functions, procedures, subroutines, subroutines, etc. When a process corresponds to a function, its termination may correspond to the function returning to its calling function or the main function.

[0140] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, cause or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion of the computer resources used may be accessible via a network. Computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc.

[0141] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media can include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, magnetic disks or optical disks, USB devices equipped with non-volatile memory, network storage devices, any suitable combinations thereof, etc. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments can be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, independent variables, parameters, data, etc., can be transmitted, forwarded, or sent through any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0142] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0143] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor performs the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or interlocking cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.

[0144] Instructions, media for delivering such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0145] In the foregoing description, aspects of this application have been described with reference to their specific aspects, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that the inventive concepts can be implemented and employed in various other ways, and the appended claims are not intended to be construed as including these variations unless limited by prior art. The various features and aspects of the applications described above can be used individually or in combination. Furthermore, aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.

[0146] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with less than or equal to ("≤") and greater than or equal to ("≥") symbols without departing from the scope of this description.

[0147] When a component is described as being “configured” to perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0148] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0149] Claim language or other languages ​​that state "at least one of" and / or "one or more of" in a set indicate that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language stating "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language stating "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any repetition is information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language "at least one of" and / or "one or more of" in a set does not limit the set to the items listed in the set. For example, the language of a claim stating "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases "at least one" and "one or more" are used interchangeably herein.

[0150] Claims using phrases such as "at least one processor, the at least one processor being configured to," "at least one processor being configured to," "one or more processors, the one or more processors being configured to," or "one or more processors being configured to," or other languages, indicate that one or more processors (in any combination) are capable of performing associated operations. For example, a claim using the phrase "at least one processor, the at least one processor being configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each assigned a specific subset of tasks to perform operations X, Y, and Z, such that the multiple processors together perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, a claim using the phrase "at least one processor, the at least one processor being configured to: X, Y, and Z" could mean that any single processor can perform only at least one subset of operations X, Y, and Z.

[0151] When referring to one or more elements that perform functions (e.g., steps of a method), one element may perform all functions, or more than one element may jointly perform these functions. When more than one element jointly performs these functions, each function does not need to be performed by every single element (e.g., different functions may be performed by different elements), and / or each function does not need to be performed by only one element as a whole (e.g., different elements may perform different sub-functions of a function). Similarly, when referring to one or more elements configured to cause another element (e.g., a device) to perform functions, one element may be configured to cause another element to perform all functions, or more than one element may be jointly configured to cause another element to perform these functions.

[0152] When referring to an entity that performs or is configured to perform functions (e.g., steps of a method) (e.g., any entity or device described herein), the entity may be configured to cause one or more elements (individually or collectively) to perform those functions. One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more of those functions, and / or any combination thereof. When referring to an entity that performs functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to perform those functions collectively. When the entity is configured to cause more than one component to perform those functions collectively, each function does not need to be performed by every single component (e.g., different functions may be performed by different components), and / or each function does not need to be performed by only one component as a whole (e.g., different components may perform different sub-functions of a function).

[0153] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0154] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.

[0155] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.

[0156] The exemplary aspects of this disclosure include:

[0157] Aspect 1. An apparatus for human pose estimation, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: acquire an image of the human; acquire a head pose of the human; determine a search region of the image based on the head pose; and determine the human pose based on the search region.

[0158] Aspect 2. The apparatus according to aspect 1, wherein the at least one processor is further configured to determine the boundary volume of the person's body based on the person's head pose, and wherein the person's pose is further determined based on the boundary volume.

[0159] Aspect 3. The apparatus according to aspect 2, wherein the at least one processor is further configured to determine a mask of the image based on the boundary volume, wherein the mask is associated with the position of the person in the image, and wherein the pose of the person is further determined based on the mask.

[0160] Aspect 4. The apparatus according to aspect 3, wherein, in order to determine the mask, the at least one processor is configured to project the boundary volume onto the image and define the mask based on the two-dimensional projection of the boundary volume onto the image.

[0161] Aspect 5. The apparatus according to any one of Aspects 3 or 4, wherein, in order to determine the pose of the person, the at least one processor is configured to determine the pose of the person based on the search region of the image, and wherein the search region is defined by the mask.

[0162] Aspect 6. The apparatus according to aspect 5, wherein the pose of the person is determined using a pose estimation machine learning model trained to determine the pose based on an image.

[0163] Aspect 7. The apparatus according to any one of Aspects 1 to 6, wherein the pose of the person is determined using a pose estimation machine learning model trained to determine the pose based on an image, and wherein the at least one processor is further configured to provide the head pose as a priori to the pose estimation machine learning model.

[0164] Aspect 8. The apparatus according to any one of Aspects 1 to 7, wherein the at least one processor is further configured to obtain the hand posture of the person, wherein the posture of the person is further determined based on the hand posture.

[0165] Aspect 9. The apparatus according to aspect 8, wherein the hand gesture of the person is determined by a head-mounted device on the person's head.

[0166] Aspect 10. The apparatus according to aspect 9, wherein the person’s hand posture is determined based on at least one of the following: hand tracking technology based on one or more images captured at the head-mounted device, or motion data from an inertial measurement unit (IMU) in a handheld device in the person’s hand.

[0167] Aspect 11. The apparatus according to any one of Aspects 8 to 10, wherein the at least one processor is further configured to determine the boundary volume of the person's body based on the person's head posture and the person's hand posture, wherein the person's posture is further determined based on the boundary volume.

[0168] Aspect 12. The apparatus according to aspect 11, wherein the at least one processor is further configured to determine a mask of the image based on the boundary volume, wherein the mask is associated with the position of the person in the image, wherein the pose of the person is further determined based on the mask.

[0169] Aspect 13. The apparatus according to aspect 12, wherein, in order to determine the mask, the at least one processor is configured to project the boundary volume onto the image and define the mask based on the two-dimensional projection of the boundary volume onto the image.

[0170] Aspect 14. The apparatus according to any one of Aspects 12 or 13, wherein, in order to determine the pose of the person, the at least one processor is configured to determine the pose of the person based on the search region of the image, and wherein the search region is defined by the mask.

[0171] Aspect 15. The apparatus according to aspect 14, wherein the pose of the person is determined using a pose estimation machine learning model trained to determine the pose based on an image.

[0172] Aspect 16. The apparatus according to any one of Aspects 8 to 15, wherein the person’s pose is determined using a pose estimation machine learning model trained to determine the pose based on an image, and wherein the at least one processor is further configured to provide the hand pose as a priori to the pose estimation machine learning model.

[0173] Aspect 17. The apparatus according to any one of Aspects 8 to 16, wherein the person’s posture is determined using a posture estimation machine learning model trained to determine the posture based on an image, and wherein the at least one processor is further configured to determine the upper body posture based on the hand posture using inverse kinematics, and to provide the upper body posture as a priori to the posture estimation machine learning model.

[0174] Aspect 18. The apparatus according to any one of Aspects 8 to 17, wherein the at least one processor is further configured to determine a confidence value of the person's posture based on a comparison between the person's posture and the person's hand posture.

[0175] Aspect 19. The apparatus according to any one of Aspects 1 to 18, wherein the person’s pose is determined using a pose estimation machine learning model trained to determine the pose based on an image and head pose.

[0176] Aspect 20. The apparatus according to aspect 19, wherein the at least one processor is further configured to obtain the person's hand pose, wherein the pose estimation machine learning model is trained to determine the pose based on an image, head pose, and hand pose, and wherein the person's pose is further determined based on the hand pose.

[0177] Aspect 21. The apparatus according to any one of aspects 1 to 20, wherein the head posture of the person is determined by a head-mounted device on the person's head.

[0178] Aspect 22. The apparatus according to aspect 21, wherein the head pose of the person is determined based on at least one of the following: motion data from one or more inertial measurement units (IMUs) of the head-mounted device, or based on simultaneous localization and mapping (SLAM) technology of images captured at the head-mounted device.

[0179] Aspect 23. The apparatus according to any one of aspects 1 to 22, wherein, in order to obtain the head posture, the at least one processor is configured to receive the head posture from a head-mounted device on the head of the person.

[0180] Aspect 24. The apparatus according to any one of aspects 1 to 23, wherein the at least one processor is further configured to determine the user's head pose and send the user's head pose.

[0181] Aspect 25. The apparatus according to aspect 24, wherein the user’s head pose is determined based on at least one of the following: motion data from one or more inertial measurement units (IMUs) of a head-mounted device on the user’s head, or based on simultaneous localization and mapping (SLAM) technology of images captured at the head-mounted device.

[0182] Aspect 26. The apparatus according to any one of aspects 1 to 25, wherein the at least one processor is further configured to determine a confidence value of the person's posture based on a comparison between the person's posture and the person's head posture, wherein the confidence value is associated with a confidence level using the person's posture.

[0183] Aspect 27. A method for human pose estimation, the method comprising: obtaining an image of the person; obtaining a head pose of the person; determining a search region of the image based on the head pose; and determining the pose of the person based on the search region.

[0184] Aspect 28. The method according to aspect 27, the method further comprising determining a boundary volume of the person's body based on the person's head pose, wherein the person's pose is further determined based on the boundary volume.

[0185] Aspect 29. The method according to aspect 28, the method further comprising determining a mask of the image based on the boundary volume, wherein the mask is associated with the position of the person in the image, and wherein the pose of the person is further determined based on the mask.

[0186] Aspect 30. The method according to aspect 29, wherein determining the mask includes projecting the boundary volume onto the image and defining the mask based on the two-dimensional projection of the boundary volume onto the image.

[0187] Aspect 31. The method according to any one of Aspects 29 or 30, wherein determining the pose of the person comprises determining the pose of the person based on the search region of the image, and wherein the search region is defined by the mask.

[0188] Aspect 32. The method according to aspect 31, wherein the pose of the person is determined using a pose estimation machine learning model trained to determine the pose based on an image.

[0189] Aspect 33. The method according to any one of Aspects 27 to 32, wherein the pose of the person is determined using a pose estimation machine learning model trained to determine the pose based on an image, and the method further includes providing the head pose as a priori to the pose estimation machine learning model.

[0190] Aspect 34. The method according to any one of Aspects 27 to 33, the method further comprising obtaining the person’s hand posture, wherein the person’s posture is further determined based on the hand posture.

[0191] Aspect 35. The method according to aspect 34, wherein the hand gesture of the person is determined by a head-mounted device on the person's head.

[0192] Aspect 36. The method according to aspect 35, wherein the person’s hand posture is determined based on at least one of the following: hand tracking technology based on one or more images captured at the head-mounted device, or motion data from an inertial measurement unit (IMU) in a handheld device in the person’s hand.

[0193] Aspect 37. The method according to any one of Aspects 34 to 36, the method further comprising determining a boundary volume of the person's body based on the person's head posture and the person's hand posture, wherein the person's posture is further determined based on the boundary volume.

[0194] Aspect 38. The method according to aspect 37, the method further comprising determining a mask of the image based on the boundary volume, wherein the mask is associated with the position of the person in the image, wherein the pose of the person is further determined based on the mask.

[0195] Aspect 39. The method according to aspect 38, wherein determining the mask includes projecting the boundary volume onto the image, and defining the mask based on the two-dimensional projection of the boundary volume onto the image.

[0196] Aspect 40. The method according to any one of Aspects 38 or 39, wherein determining the pose of the person comprises determining the pose of the person based on the search region of the image, and wherein the search region is defined by the mask.

[0197] Aspect 41. The method according to aspect 40, wherein the pose of the person is determined using a pose estimation machine learning model trained to determine the pose based on an image.

[0198] Aspect 42. The method according to any one of Aspects 34 to 41, wherein the person’s pose is determined using a pose estimation machine learning model trained to determine the pose based on an image, and wherein the method further comprises providing the hand pose as a priori to the pose estimation machine learning model.

[0199] Aspect 43. The method according to any one of Aspects 34 to 42, wherein the person’s posture is determined using a posture estimation machine learning model trained to determine the posture based on an image, and wherein the method further comprises determining an upper body posture based on the hand posture using inverse kinematics, and providing the upper body posture as a priori to the posture estimation machine learning model.

[0200] Aspect 44. The method according to any one of aspects 34 to 43, the method further comprising determining a confidence value of the person's posture based on a comparison between the person's posture and the person's hand posture.

[0201] Aspect 45. The method according to any one of Aspects 27 to 44, wherein the pose of the person is determined using a pose estimation machine learning model trained to determine the pose based on an image and head pose.

[0202] Aspect 46. The method according to aspect 45, the method further comprising obtaining the person's hand pose, wherein the pose estimation machine learning model is trained to determine the pose based on an image, head pose, and hand pose, and wherein the person's pose is further determined based on the hand pose.

[0203] Aspect 47. The method according to any one of Aspects 27 to 46, wherein the head posture of the person is determined by a head-mounted device on the person's head.

[0204] Aspect 48. The method according to aspect 47, wherein the head pose of the person is determined based on at least one of the following: motion data from one or more inertial measurement units (IMUs) of the head-mounted device, or based on simultaneous localization and mapping (SLAM) technology of images captured at the head-mounted device.

[0205] Aspect 49. The method according to any one of Aspects 27 to 48, wherein obtaining the head posture includes receiving the head posture from a head-mounted device on the person's head.

[0206] Aspect 50. The method according to any one of Aspects 27 to 49, the method further comprising determining a user’s head pose and sending the user’s head pose.

[0207] Aspect 51. The method according to aspect 50, wherein the user’s head pose is determined based on at least one of the following: motion data from one or more inertial measurement units (IMUs) of a head-mounted device on the user’s head, or based on simultaneous localization and mapping (SLAM) techniques of images captured at the head-mounted device.

[0208] Aspect 52. The method according to any one of Aspects 27 to 51, the method further comprising determining a confidence value of the person's posture based on a comparison between the person's posture and the person's head posture, wherein the confidence value is associated with a confidence level using the person's posture.

[0209] Aspect 53. A method comprising: acquiring an image of a person; determining the person's pose based on the image using a pose estimation technique; and providing the person's pose to a device of the person, wherein the device of the person uses the person's pose as a priori to determine the person's position, wherein the person's position is further determined based on at least one of: motion data from one or more inertial measurement units (IMUs) of the device, or simultaneous localization and mapping (SLAM) techniques based on images captured at the device.

[0210] Aspect 54. A system for training a pose estimation model, the system comprising: a head-mounted device configured to determine the pose of a person's head; an image capture device configured to capture an image of the person; and at least one processor configured to train a pose estimation machine learning model using the image as input and the pose as a ground truth.

[0211] Aspect 55. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing the at least one processor, when executed, to perform any one of aspects 27 to 53.

[0212] Aspect 56. An apparatus for providing virtual content for display, the apparatus comprising one or more components for performing operations according to any one of aspects 27 to 53.

Claims

1. An apparatus for human posture estimation, the apparatus comprising: At least one memory; and At least one processor, the at least one processor being coupled to the at least one memory and being configured to: Obtain an image of the person; Obtain the head posture of the person; The search area of ​​the image is determined based on the head pose; as well as The person's posture is determined based on the search area.

2. The apparatus of claim 1, wherein the at least one processor is further configured to determine the boundary volume of the person's body based on the person's head pose, and wherein the person's pose is further determined based on the boundary volume.

3. The apparatus of claim 2, wherein the at least one processor is further configured to determine a mask of the image based on the boundary volume, wherein the mask is associated with the position of the person in the image, and wherein the pose of the person is further determined based on the mask.

4. The apparatus of claim 3, wherein, in order to determine the mask, the at least one processor is configured to project the boundary volume onto the image and define the mask based on the two-dimensional projection of the boundary volume onto the image.

5. The apparatus of claim 3, wherein, in order to determine the pose of the person, the at least one processor is configured to determine the pose of the person based on the search region of the image, and wherein the search region is defined by the mask.

6. The apparatus of claim 5, wherein the person’s pose is determined using a pose estimation machine learning model trained to determine the pose based on an image.

7. The apparatus of claim 1, wherein the person’s pose is determined using a pose estimation machine learning model trained to determine the pose based on an image, and wherein the at least one processor is further configured to provide the head pose as a priori to the pose estimation machine learning model.

8. The apparatus of claim 1, wherein the at least one processor is further configured to obtain the person's hand gesture, wherein the person's gesture is further determined based on the hand gesture.

9. The apparatus of claim 8, wherein the hand posture of the person is determined by a head-mounted device on the person's head.

10. The apparatus of claim 9, wherein the person’s hand posture is determined based on at least one of the following: hand tracking technology based on one or more images captured at the head-mounted device, or motion data from an inertial measurement unit (IMU) in a handheld device in the person’s hand.

11. The apparatus of claim 8, wherein the at least one processor is further configured to determine a boundary volume of the person's body based on the person's head pose and the person's hand pose, wherein the person's pose is further determined based on the boundary volume.

12. The apparatus of claim 11, wherein the at least one processor is further configured to determine a mask of the image based on the boundary volume, wherein the mask is associated with the position of the person in the image, wherein the pose of the person is further determined based on the mask.

13. The apparatus of claim 12, wherein, in order to determine the mask, the at least one processor is configured to project the boundary volume onto the image and define the mask based on the two-dimensional projection of the boundary volume onto the image.

14. The apparatus of claim 12, wherein, in order to determine the pose of the person, the at least one processor is configured to determine the pose of the person based on the search region of the image, and wherein the search region is defined by the mask.

15. The apparatus of claim 14, wherein the person’s pose is determined using a pose estimation machine learning model trained to determine the pose based on an image.

16. The apparatus of claim 8, wherein the person’s pose is determined using a pose estimation machine learning model trained to determine the pose based on an image, and wherein the at least one processor is further configured to provide the hand pose as a priori to the pose estimation machine learning model.

17. The apparatus of claim 8, wherein the person’s posture is determined using a posture estimation machine learning model trained to determine the posture based on an image, and wherein the at least one processor is further configured to determine the upper body posture based on the hand posture using inverse kinematics, and to provide the upper body posture as a priori to the posture estimation machine learning model.

18. The apparatus of claim 8, wherein the at least one processor is further configured to determine a confidence value of the person's posture based on a comparison between the person's posture and the person's hand posture.

19. The apparatus of claim 1, wherein the person’s pose is determined using a pose estimation machine learning model trained to determine the pose based on an image and head pose.

20. The apparatus of claim 19, wherein the at least one processor is further configured to obtain the person's hand pose, wherein the pose estimation machine learning model is trained to determine the pose based on an image, head pose, and hand pose, and wherein the person's pose is further determined based on the hand pose.

21. The apparatus of claim 1, wherein the head posture of the person is determined by a head-mounted device on the person's head.

22. The apparatus of claim 21, wherein the head posture of the person is determined based on at least one of the following: motion data from one or more inertial measurement units (IMUs) of the head-mounted device, or based on simultaneous localization and mapping (SLAM) technology of images captured at the head-mounted device.

23. The apparatus of claim 1, wherein, in order to obtain the head posture, the at least one processor is configured to receive the head posture from a head-mounted device on the person's head.

24. The apparatus of claim 1, wherein the at least one processor is further configured to determine the user's head pose and send the user's head pose.

25. The apparatus of claim 24, wherein the user’s head pose is determined based on at least one of the following: motion data from one or more inertial measurement units (IMUs) of a head-mounted device on the user’s head, or based on simultaneous localization and mapping (SLAM) technology of images captured at the head-mounted device.

26. The apparatus of claim 1, wherein the at least one processor is further configured to determine a confidence value of the person's posture based on a comparison between the person's posture and the person's head posture, wherein the confidence value is associated with a confidence level using the person's posture.

27. A method for human pose estimation, the method comprising: Obtain an image of the person; Obtain the head posture of the person; The search area of ​​the image is determined based on the head pose; as well as The person's posture is determined based on the search area.

28. The method of claim 27, further comprising determining a boundary volume of the person's body based on the person's head pose, wherein the person's pose is further determined based on the boundary volume.

29. The method of claim 28, further comprising determining a mask of the image based on the boundary volume, wherein the mask is associated with the position of the person in the image, and wherein the pose of the person is further determined based on the mask.

30. The method of claim 29, wherein determining the mask comprises projecting the boundary volume onto the image, and defining the mask based on a two-dimensional projection of the boundary volume onto the image.

31. A method, the method comprising: Obtain an image of a person; The person's pose is determined based on the image using pose estimation techniques. as well as Provide the person's posture to the person's device. The device of the person uses the person's posture as a priori to determine the person's position, wherein the person's position is further determined based on at least one of the following: motion data from one or more inertial measurement units (IMUs) of the device, or simultaneous localization and mapping (SLAM) technology based on images captured at the device.

Citation Information

Patent Citations

  • Method and device for determining posture of whole human body

    CN115050097A

  • Human body posture determination method and device, equipment and storage medium

    CN115560750A

  • Egocentric pose estimation from human vision span

    US20220319041A1

  • Hand Pose Estimation for Machine Learning Based Gesture Recognition

    US20230214458A1