Keypoint-based sampling for pose estimation

CN116997941BActive Publication Date: 2026-08-11QUALCOMM TECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2026-08-11

Smart Images

  • Figure CN116997941B_ABST
    Figure CN116997941B_ABST
Patent Text Reader

Abstract

Systems and techniques are provided for determining one or more poses of one or more objects. For example, the process may include determining multiple keypoints from an image using a machine learning system. The multiple keypoints are associated with at least one object in the image. The process may include determining multiple features from the machine learning system based on the multiple keypoints. The process may include classifying the multiple features into multiple joint types. The process may include determining pose parameters for at least one object based on the multiple joint types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] In summary, this disclosure relates to determining or estimating the pose of an object in an image or frame. For example, aspects of this disclosure relate to performing keypoint-based sampling for determining or estimating the pose of an object (e.g., multiple hands, a single hand, and physical objects, etc.) in an image or frame. Background Technology

[0002] Determining the objects present in an image and their properties is useful for many applications. For example, a system can determine the pose of objects in an image (e.g., a person, a part of a person (such as a hand or face), a vehicle, a building, etc.). In some cases, the pose can be used to determine or generate a model (e.g., a three-dimensional (3D) model) representing the object. For example, a model can be generated using the pose determined for the object.

[0003] The pose determined for objects in an image (and in some cases, the model generated with that pose) can be used to facilitate the efficient operation of a variety of systems and / or applications. Examples of such applications and systems include extended reality (XR) systems (e.g., augmented reality (AR), virtual reality (VR), and / or mixed reality (MR) systems), robotics, automotive and aerospace, 3D scene understanding, object grasping, and object tracking, among many other applications and systems. In various illustrative examples, a 3D model with a determined pose can be displayed (e.g., by a mobile device, by an XR system, and / or by other systems or devices) to determine the position of objects represented by the 3D model (e.g., for scene understanding and / or navigation, for object grasping, for autonomous vehicle operation, and / or for other uses), and other applications.

[0004] Determining the accurate pose of an object allows the system to generate a precise localized and oriented representation of the object (e.g., a model). Summary of the Invention

[0005] In some examples, systems and techniques for performing keypoint-based sampling to determine or estimate the pose of objects in an image are described. For example, these systems and techniques can be used to determine the pose of a person's two hands in an image, the pose of one hand and the pose of a physical object positioned relative to that hand (e.g., a cup or other object held or near that hand), the pose of a person's two hands and the pose of a physical object positioned relative to one or both hands and / or other objects.

[0006] According to at least one example, a method is provided for determining one or more poses of one or more objects. The method includes: determining multiple keypoints from an image using a machine learning system, the multiple keypoints being associated with at least one object in the image; determining multiple features from the machine learning system based on the multiple keypoints; classifying the multiple features into multiple joint types; and determining pose parameters for the at least one object based on the multiple joint types.

[0007] In another example, an apparatus is provided for determining one or more poses of one or more objects. The apparatus includes at least one memory and a processor (e.g., implemented in a circuit) coupled to the at least one memory. The at least one processor is configured and capable of performing the following operations: determining a plurality of keypoints from an image using a machine learning system, the plurality of keypoints being associated with at least one object in the image; determining a plurality of features from the machine learning system based on the plurality of keypoints; classifying the plurality of features into a plurality of joint types; and determining pose parameters for the at least one object based on the plurality of joint types.

[0008] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon, which, when executed by one or more processors, cause one or more processors to: determine multiple keypoints from an image using a machine learning system, the multiple keypoints being associated with at least one object in the image; determine multiple features from the machine learning system based on the multiple keypoints; classify the multiple features into multiple joint types; and determine pose parameters for at least one object based on the multiple joint types.

[0009] In another example, an apparatus is provided for determining one or more poses of one or more objects. The apparatus includes: a unit for determining multiple keypoints from an image using a machine learning system, the multiple keypoints being associated with at least one object in the image; a unit for determining multiple features from the machine learning system based on the multiple keypoints; a unit for classifying the multiple features into multiple joint types; and a unit for determining pose parameters for at least one object based on the multiple joint types.

[0010] In some aspects, at least one object comprises two objects. In such aspects, multiple keypoints comprise keypoints for both objects. In such aspects, pose parameters may include pose parameters for both objects.

[0011] In some aspects, at least one object includes at least one hand. In some cases, at least one hand includes two hands. In this case, multiple keypoints may include keypoints for both hands. In this case, pose parameters may include pose parameters for both hands.

[0012] In some aspects, at least one object includes a single hand. In such aspects, the methods, apparatus, and computer-readable media described above may further include: using a machine learning system to determine a plurality of object keypoints from an image, the plurality of object keypoints being associated with an object associated with a single hand; and determining pose parameters for the object based on the plurality of object keypoints.

[0013] In some aspects, each of the multiple keypoints corresponds to a joint of at least one object.

[0014] In some aspects, in order to determine multiple features from a machine learning system based on multiple key points, the above-described methods, apparatus, and computer-readable media may further include: determining a first set of features corresponding to the multiple key points from a first feature map of the machine learning system, the first feature map including a first resolution; and determining a second set of features corresponding to the multiple key points from a second feature map of the machine learning system, the second feature map including a second resolution.

[0015] In some aspects, the methods, apparatus, and computer-readable media described above may further include: generating a feature representation for each of a plurality of keypoints, wherein the plurality of features are classified into a plurality of joint types using the feature representation for each keypoint. In some cases, the feature representation for each keypoint includes an encoded vector.

[0016] In some respects, machine learning systems include neural networks that use images as input.

[0017] In some aspects, multiple features are classified into multiple joint types by the encoder of the transformer neural network, and pose parameters for at least one object are determined by the decoder of the transformer neural network based on multiple joint types.

[0018] In some aspects, pose parameters are determined for at least one object based on multiple joint types and one or more learned joint queries. In some cases, the at least one object includes a first object and a second object. In this case, one or more learned joint queries can be used to predict at least one of the following: relative translation between the first and second objects, a set of object shape parameters, and camera model parameters.

[0019] In some aspects, the pose parameters for at least one object include a three-dimensional vector for each of a plurality of joint types. In some cases, the three-dimensional vector for each of the plurality of joint types includes horizontal, vertical, and depth components. In some cases, the three-dimensional vector for each of the plurality of joint types includes a vector between each joint and the parent joint associated with each joint.

[0020] In some aspects, the pose parameters for at least one object include the position of each joint and the difference between the depth of each joint and the depth of the parent joint associated with each joint.

[0021] In some aspects, the pose parameters for at least one object include the translation of at least one object relative to another object in the image.

[0022] In some aspects, the pose parameters for at least one object include the shape of at least one hand.

[0023] In some aspects, the methods, apparatus, and computer-readable media described above may further include determining user input based on attitude parameters.

[0024] In some aspects, the methods, apparatus, and computer-readable media described above may further include rendering virtual content based on pose parameters.

[0025] In some aspects, the apparatus may include mobile devices (e.g., mobile phones or so-called "smartphones"), wearable devices, extended reality devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices, such as head-mounted displays (HMDs), XR glasses, etc.), personal computers, laptop computers, vehicles (or computing devices or components of vehicles), server computers, televisions, video game consoles, or other devices, or part of the aforementioned devices. In some aspects, the apparatus also includes at least one camera for capturing one or more images or video frames. For example, the apparatus may include one or more cameras (e.g., an RGB camera) for capturing one or more images and / or one or more videos including video frames. In some aspects, the apparatus includes a display for displaying one or more images, one or more 3D models, one or more videos, one or more notifications, any combination thereof, and / or other displayable data. In some aspects, the apparatus includes a transmitter configured to transmit data (e.g., data representing images, videos, 3D models, etc.) to at least one device via a transmission medium. In some respects, a processor includes a neural processing unit (NPU), a central processing unit (CPU), a graphics processing unit (GPU), or other processing devices or components.

[0026] This invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. The subject matter should be understood with reference to the appropriate portions of the entire specification, any or all of the drawings, and each claim.

[0027] The foregoing and other features and embodiments will become more apparent upon reference to the following description, claims and drawings. Attached Figure Description

[0028] The illustrative embodiments of this application are described in detail below with reference to the accompanying drawings:

[0029] Figures 1A-1D These are images illustrating examples of hand interactions based on some sample images;

[0030] Figure 2A This is a diagram illustrating examples of pose estimation systems based on some examples;

[0031] Figure 2B This is a diagram illustrating an example of the process of determining the pose of one or more objects from an input image, based on some examples;

[0032] Figure 3 This is another example of a process for determining the pose of one or more objects from an input image, based on some examples;

[0033] Figure 4 This is a diagram illustrating examples of hands with various joints and associated joint labels or identifiers, based on some examples.

[0034] Figure 5 Includes example images of poses from the H2O-3D dataset and annotations, based on some examples;

[0035] Figure 6 This is a diagram illustrating the cross-focus of queries for the three joints of the right hand, based on some examples;

[0036] Figure 7 This is a flowchart illustrating an example of a process for determining one or more poses of one or more objects based on some examples;

[0037] Figure 8 This is a block diagram illustrating examples of deep learning neural networks based on some examples;

[0038] Figure 9 This is a block diagram illustrating examples of convolutional neural networks (CNNs) based on some examples;

[0039] Figure 10 Examples of computing systems that can implement one or more of the techniques described herein are shown, based on some examples. Detailed Implementation

[0040] Certain aspects and embodiments of this disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments can be applied independently, and some can be applied in combination. Specific details are set forth in the following description for purposes of explanation in order to provide a thorough understanding of embodiments of this application. However, it will be apparent that various embodiments can be practiced without these specific details. The drawings and description are not intended to be limiting.

[0041] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the subsequent description of the exemplary embodiments will provide those skilled in the art with an implementation description for carrying out the exemplary embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0042] As described above, the system can determine the pose of an object in an image or frame. The terms image and frame are used interchangeably herein. For example, an image or frame can refer to a standalone image (e.g., a still image), an image or frame from a sequence of frames (e.g., from a video), a depth image including depth information, and / or other types of images or frames. Pose can include position (e.g., 3D translation) and orientation (e.g., pitch, roll, and yaw). In some cases, the system can use a pose generation model (e.g., a 3D model) to represent the object. The pose determined for an object in an image (and in some cases, the model generated using the determined pose) can be used to facilitate efficient operation of a variety of systems and / or applications. Examples of systems and applications that can utilize pose information include extended reality (XR) systems (e.g., augmented reality (AR) systems, virtual reality (VR) systems, and / or mixed reality (MR) systems), robotics, automotive and aerospace, 3D scene understanding, object grasping, object tracking, and / or other systems and applications.

[0043] In one example, based on determining the pose of an object in an image, the system can use the determined pose to generate a three-dimensional (3D) model of the object. The 3D model can be displayed (e.g., by a mobile device, by an XR system, and / or by other systems or devices), used to determine the position of the object represented by the 3D model (e.g., for scene understanding and / or navigation, for object grasping, for autonomous vehicle operation, and / or for other uses), and for other purposes.

[0044] For example, in some AR systems, users can view images that combine artificial or virtual graphics with their natural environment. In some cases, users can view the real world through the display of an AR system (e.g., the lenses of AR glasses), where virtual objects or graphics can also be displayed. In others, users view images of the real-world environment along with virtual objects or graphics. Such AR applications allow for the manipulation of real-world images to add virtual objects to the images and to align virtual objects with the images in multiple dimensions. For example, real-world objects that exist in reality can be represented using models that are similar to or exactly match the real-world objects. In one example, as a user continues to view his or her natural environment in an AR environment, a model of a virtual airplane can be presented in the view of an AR device (e.g., glasses, goggles, or other devices) representing a real airplane sitting on a runway. Viewers may be able to manipulate the model while viewing the real-world scene. In another example, an actual object sitting on a table can be identified and rendered in an AR environment using models with different colors or different physical properties. In some cases, computer-generated copies of artificial virtual objects that do not exist in reality or actual objects or structures in the user's natural environment can also be added to the AR environment.

[0045] Determining the accurate pose of an object in an image can help generate a precisely localized and oriented representation of the object (e.g., a 3D model). For example, 3D hand pose estimation has the potential to enhance various systems and applications (e.g., XR systems such as VR, AR, MR, etc.) and interactions with computers and robots by making them more efficient and / or more intuitive. In one case, using an AR system as an illustrative example, improved 3D hand pose estimation can allow the AR system to more accurately interpret gesture-based input from a user. In another example, improved 3D hand estimation can significantly improve the accuracy of remote control operations (e.g., remote operation of autonomous vehicles, remote surgery using robotic systems, etc.) by allowing the system to accurately determine the correct position and orientation of the hand in one or more images (e.g., images or frames of video).

[0046] In some cases, accurately determining the pose of certain objects can be difficult. For example, determining or estimating the pose of objects interacting with each other in an image can be challenging. In an illustrative example, one of the user's hands may be interacting closely with the user's other hand in an image. Figures 1A-1D These are images illustrating examples of close interaction between hands. For example, such as... Figure 1A As shown, the user's right hand 102 and left hand 104 are clasped together in an overlapping manner, causing obscuring and blurring of which fingers and joints are under which hand. Figure 1BAs shown, the user's right hand 102 and left hand 104 are joined together, with the fingertips of the right hand 102 touching the fingertips of the left hand 104. Figure 1C As shown, the user's right hand 102 is grasping their left hand 104. Figure 1D In the image, the right hand 102 and the left hand 104 are joined together, with the fingers of both hands 102 and 104 alternating in the vertical direction (relative to the image plane).

[0047] As described above, identify objects in the image that interact with each other (e.g., Figures 1B-1D The pose of the hand (as shown) can be a challenging problem. Using the hand as an illustrative example, the joints of the hand can be identified from the image and used to determine the hand's pose. However, there may be occlusion between the joints of the hand, as well as uncertainty about which joints visible in the image belong to which hand. References Figures 1A-1D As an illustrative example, the system may not be able to easily determine which joints belong to the right hand 102 and which belong to the left hand 104. Furthermore, some joints may not be visible due to occlusion (e.g., Figure 1A The thumb joint in the image Figure 1D (e.g., some joints in the left hand 104). Similar problems arise when processing images with other objects being interacted with, leading to occlusion and / or other blurring. For example, an image might contain one or more hands of a user holding a coffee cup. However, due to the interaction between (one or more) hands and (one or more) hands with the coffee cup, a portion of one or both hands and / or a portion of the coffee cup may be occluded.

[0048] Significant progress has been made in single-hand pose estimation based on depth maps and monochrome images (e.g., images with red (R), blue (B), and green (G) components per pixel (called RGB images), YUV or YCbCr images including luma or luminance components Y and chroma or chrominance components U and V or Cb and Cr per pixel, or other types of images). The ability to determine the pose of an object in an RGB image can be attractive because it does not require power-intensive (e.g., continuously active sensors) motion sensors (e.g., depth sensors). Numerous methods have been proposed for performing hand pose estimation (e.g., using different convolutional network architectures to directly predict 3D joint positions or angles, as well as rendering-dependent fine-grained pose estimation and tracking).

[0049] Compared to single-hand pose estimation, progress in bi-hand pose estimation has been limited. The problem of bi-hand pose estimation is arguably more challenging than single-hand pose estimation. For example, the joints of the two hands in an image may appear similar, making accurate identification of the hand joints difficult (e.g., as shown in the image). Figures 1A-1D (As shown). Furthermore, some joints in one hand are likely to be obscured by the other hand or the same hand (e.g., as shown). Figures 1A-1D (As shown). In this scenario, existing solutions fail to accurately estimate hand poses. For example, detecting the left and right hands before independently predicting their 3D poses performs poorly in scenarios where both hands interact closely. Figures 1A-1D As shown. In another example, when attempting to identify joints in an image, a bottom-up approach that first estimates the 2D joint positions and their depths can struggle to handle joint similarities and occlusions.

[0050] This document describes systems, methods (also referred to as processes), apparatuses, and computer-readable media (collectively, the “Systems and Techniques”) for performing keypoint-based sampling to determine or estimate the pose of objects (e.g., hands, a hand, and physical objects) in an image. The systems and techniques described herein can be used to determine the pose of any type of object or combination of objects (e.g., objects of the same type, objects of different types, etc.) in one or more images. For example, these systems and techniques can be used to estimate or determine the pose of a person’s two hands in an image, the pose of a hand and the pose of a physical object positioned relative to that hand (e.g., a cup or other object held or near that hand, such as being obscured by that hand), the pose of a person’s two hands and the pose of a physical object positioned relative to one or both hands, and / or the pose of other types of objects in the image. While this document uses hands and objects interacting with or near hands as illustrative examples of objects, it will be understood that these systems and techniques can be used to determine the pose of any type of object.

[0051] Figure 2A This is a diagram illustrating an example of a pose estimation system 200, configured to perform keypoint-based sampling for determining or estimating the pose of an object in an image. The pose estimation system 200 includes one or more image sensors 224, a memory 226, and one or more depth sensors 222 (e.g., which are optional, such as those via...). Figure 2AThe dashed outlines shown indicate the processing system 230, keypoint determination engine 250, feature determination engine 254, and pose estimation engine 256. In some examples, the keypoint determination engine 250 includes a machine learning system 252, which may include one or more neural networks and / or other machine learning systems. In one illustrative example, the machine learning system 252 of the keypoint determination engine 250 may include a U-shaped neural network. In some examples, the pose estimation engine 256 includes a machine learning system 257, which may include one or more neural networks and / or other machine learning systems. In one illustrative example, the machine learning system 257 of the pose estimation engine 256 may include a transformer neural network (e.g., including a transformer encoder and a transformer decoder). The following discusses... Figure 8 and Figure 9 An illustrative example of a neural network is described.

[0052] Processing system 230 may include components, such as, but not limited to, a central processing unit (CPU) 232, a graphics processing unit (GPU) 234, a digital signal processor (DSP) 236, an image signal processor (ISP) 238, a cache memory 251, and / or a memory 253, which processing system 230 may use to perform one or more of the operations described herein. For example, CPU 232, GPU 234, DSP 236, and / or ISP 238 may include electronic circuitry or other electronic hardware, such as one or more programmable electronic circuits. CPU 232, GPU 234, DSP 236, and / or ISP 238 may implement or execute computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of processing system 230. In some cases, one or more of the CPU 232, GPU 234, DSP 236, and / or ISP 238 can implement the keypoint determination engine 250, the feature determination engine 254, and / or the pose estimation engine 256. It should be noted that in some examples, the processing system 230 can implement features not present in the CPU 232, GPU 234, DSP 236, and / or ISP 238. Figure 2A One or more computation engines are shown in the figure. For illustrative and explanatory purposes, this document provides a key point determination engine 250, a feature determination engine 254, and a pose estimation engine 256, and for simplicity, other possible computation engines are not shown.

[0053] The pose estimation system 200 may be part of or implemented by a single computing device or multiple computing devices. In some examples, the pose estimation system 200 may be part of one or more electronic devices, such as extended reality (XR) devices (e.g., head-mounted displays (HMDs), XR glasses, AR glasses, or other extended reality devices for rendering virtual reality (VR), augmented reality (AR), and / or mixed reality (MR), head-up displays (HUDs), mobile devices (e.g., smartphones, cellular phones, or other mobile devices), vehicles or computing components or systems in vehicles (e.g., autonomous vehicles or human-driven vehicles), camera systems or devices (e.g., digital cameras, camera phones, video phones, IP cameras, video cameras, security cameras, or other camera systems or devices), laptops or notebook computers, tablets, set-top boxes, televisions, display devices, digital media players, game consoles, video streaming devices, drones, Internet of Things (IoT) devices, smart wearable devices, or any other suitable electronic devices.

[0054] In some implementations, one or more depth sensors 222, image sensors 224, storage devices 226, processing systems 230, keypoint determination engines 250, feature determination engines 254, and pose estimation engines 256 may be part of the same computing device. For example, in some cases, one or more depth sensors 222, image sensors 224, storage devices 226, processing systems 230, keypoint determination engines 250, feature determination engines 254, and pose estimation engines 256 may be integrated into XR devices, vehicle computing systems, smartphones, cameras, laptops, tablets, smart wearable devices, HMDs, IoT devices, gaming systems, and / or any other computing devices. However, in some implementations, one or more of the depth sensors 222, image sensors 224, storage devices 226, processing systems 230, keypoint determination engines 250, feature determination engines 254, and / or pose estimation engines 256 may be part of, or implemented by, two or more independent computing devices.

[0055] Figure 2B This is a diagram illustrating an illustrative example of a process 260 that can be performed by a pose estimation system 200 to determine the pose of one or more objects from an input image. The pose estimation system 200 may obtain the input image from an image source (not shown), such as... Figure 2BThe input image 261 is shown. Image sources may include one or more image sensors (e.g., cameras) (included in or communicating with the attitude estimation system 200), image and / or video storage devices (e.g., storage device 226 or other storage devices of system 200 or another system or device), image and / or video archives containing stored images, image and / or video servers or content providers that provide image and / or video data, image and / or video feed interfaces that receive images from video servers or content providers, computer graphics systems for generating computer graphics images and / or video data, combinations of such sources, or other sources of image frame content. In some cases, multiple image sources (e.g., multiple image sensors, multiple storage devices, etc.) may provide images to the attitude estimation system 200. The image can be a red-green-blue (RGB) image with red, green, and blue components per pixel; a luminance, chrominance-red, chrominance-blue (YCbCr) image with a luminance component per pixel and two chrominance (color) components (chrominance-red and chrominance-blue); or any other suitable type of color or monochrome image. In some cases, the image may include depth information, such as an RGB-depth (RGB-D) image containing RGB color components and depth information per pixel.

[0056] exist Figure 2B At box 262 of process 260, the keypoint determination engine 250 can process the input image 261 to determine keypoints associated with one or more objects (e.g., a hand, multiple hands, or other objects) in image 261. In an illustrative example, keypoints can be determined from image 261 using machine learning system 252. For example, image 261 can be input to machine learning system 252. Machine learning system 252 can process image 261 to output keypoints detected from image 261. Machine learning system 252 may include keypoints trained to determine keypoints from images (e.g., using information about...). Figure 8 and Figure 9 The neural network described uses the backpropagation technique.

[0057] Keypoints correspond to specific portions of one or more objects in image 261. In an illustrative example, image 261 may include one or more hands. In some cases, image 261 may include a physical object held or near one or more hands (e.g., occluded by one or more hands). In such an example, each keypoint may correspond to a joint of one or more hands in image 261 (and in some cases, to a point of a physical object held or near a hand). In some cases, machine learning system 252 may process image 261 to generate a graph (which may be called a heatmap) or array of points associated with the input image, such as... Figure 2BFigure 263 is shown. The graph or array can be a two-dimensional (2D) graph or a point array. For example, a graph can be an image with a 2D array of numbers, where each number includes values ​​within a specific range (e.g., a range of values ​​between 0 and 1).

[0058] In some examples, the keypoint determination engine 250 (e.g., machine learning system 252) may determine keypoints associated with one or more hands (and in some cases, physical objects) in image 261, at least in part, by extracting a set of candidate 2D locations of joints of one or more hands (and / or physical objects) from Figure 263. In one example, the keypoint determination engine 250 (e.g., machine learning system 252) selects extreme values ​​in Figure 263 as the set of candidate 2D locations of joints. For example, at locations in Figure 263 corresponding to features of interest (e.g., joints of one or more hands) in image 261, Figure 263 may have values ​​close to (e.g., within a threshold difference, such as 0.1, 0.15, 0.2, etc.) the maximum value of a range of possible values ​​in Figure 263 (e.g., a value of 1 in this example when the range includes values ​​between 0 and 1). The extreme values ​​selected by the keypoint determination engine 250 may include locations in Figure 263 that are close to the maximum value (e.g., locations where the value is close to 1), such as within a threshold difference. In some examples, the key point determination engine 250 can use non-maximum suppression techniques to ensure that two extreme values ​​do not get too close to each other.

[0059] While this paper uses joints as a representation feature of the hand, other parts of the hand can be used in other examples. In some implementations, when one or more physical objects are in an image (e.g., being held or near one or more hands), the keypoint determination engine 250 can generate a map for the hand (e.g., Figure 263) and can generate separate maps for one or more objects. In some examples, when the image includes one or more physical objects other than the hand, the keypoint determination engine 250 (e.g., using machine learning system 252) can perform object segmentation to separate one or more physical objects from the background and / or other objects in the image (e.g., from one or more hands in the image). For example, the keypoint determination engine 250 can generate a segmentation map or image that includes values ​​corresponding to one or more objects. The keypoint determination engine 250 can select (e.g., by random selection) points within the segmentation map and use the selected points as keypoints for one or more objects. In an example where the image includes one or more objects and one or more hands, the key point determination engine 250 can also generate a map (e.g., Figure 263, such as a heatmap) for one or more hands and determine key points for one or more hands as described above.

[0060] In some cases, it is not required that all keypoint locations correctly correspond to specific parts of one or more objects (e.g., joints of a hand) in image 261. In some cases, all specific parts of one or more objects in image 261 (e.g., all joints of one or more hands in image 261) may not be detectable from image 261, for example, due to occlusion and / or other factors.

[0061] exist Figure 2B At frame 264 of process 260, the feature determination engine 254 of the pose estimation system 200 can determine features based on keypoints determined from image 261. For example, as described above, the machine learning system 252 may include a neural network for determining keypoints (e.g., based on 2D image 263). The neural network of the machine learning system 252 may include multiple layers (e.g., convolutional layers, which may be followed by other layers such as pooling layers, activation functions, etc.) for processing image data from image 261. (See also: ...) Figure 8 and Figure 9 As described, each layer may include filters or nodes for processing data associated with image 261, resulting in one or more feature maps (or feature arrays) generated by each layer. The feature maps may be stored in a storage device such as storage device 226. For example, each filter or node may include an array of parameter values ​​(e.g., weights, biases, etc.) that are applied to the output of the previous layer in the neural network of machine learning system 252. The parameter values ​​can be tuned during training, such as regarding... Figure 8 and Figure 9 The example described is as follows. In an illustrative example, the neural network of the machine learning system 252 includes a U-network neural network, which includes an encoder part and a decoder part.

[0062] Using keypoints, the feature determination engine 254 can determine or extract features from one or more feature maps of certain layers (e.g., the last M layers, where M is a positive integer value) of the neural network of the machine learning system 252. For example, features in the feature maps that share the same spatial location (in 2D) as keypoints in 2D graph 263 can be extracted to represent keypoints. In some examples, the feature determination engine 254 can determine linear interpolation features from multiple layers of the decoder portion of the neural network of the machine learning system 252, resulting in multi-scale features because each layer has a different scale or resolution. In some cases, when the feature map has a different resolution or dimension compared to the 2D graph output by the neural network of the machine learning system 252, the feature map or 2D graph can be scaled up or down so that the feature map and the 2D graph have the same resolution. In an illustrative example, graph 263 may have a resolution of 128x128 pixels (in both the horizontal and vertical directions), and the feature maps from the various layers may have different resolutions. Figure 263 can be enlarged (by increasing resolution) or reduced (by decreasing resolution) to match the resolution of each feature map from which features are extracted. This scaling allows the feature determination engine 254 to extract features from the correct locations within the feature maps (e.g., locations that precisely correspond to keypoints in Figure 263).

[0063] Feature determination engine 254 can generate feature representations for the determined features. An example of a feature representation is a feature vector. In some cases, feature determination engine 254 can determine a feature representation (e.g., a feature vector) for each keypoint determined by keypoint determination engine 250. For example, feature determination engine 254 can combine features from multiple layers of a neural network from machine learning system 252 corresponding to a particular keypoint (e.g., by concatenating features) to form a single feature vector for that particular keypoint. In one example, for the keypoint at location (3, 7) in Figure 263 (corresponding to the third column and seventh row in the 2D array of Figure 263), feature determination engine 254 can combine (e.g., concatenate) all features from various feature maps of the neural network corresponding to location (3, 7) into a single feature vector. In some examples, the feature representation can be generated as an appearance and spatial encoding of the location from Figure 263. For example, as described in more detail below, feature determination engine 254 can combine (e.g., concatenate) the location encoding with features extracted for a particular keypoint to form a feature representation (e.g., a feature vector) for that particular keypoint.

[0064] The pose estimation engine 256 can use feature representations as input to determine the correct configuration of specific parts of one or more objects (e.g., the joints of one or more hands) and the 3D pose of one or more objects (e.g., the 3D pose of two hands in an image, one or more hands, and the 3D pose of a physical object grasping or near a hand, etc.). For example, in Figure 2B At box 266 of process 260, the encoder of machine learning system 257 can process the feature representations output by feature determination engine 254 to determine the classification of a specific part of one or more objects. For example, the encoder of the neural network of machine learning system 257 can process feature representations of one or more hands in image 261 to determine the joint category or joint type for each feature representation in the feature representations. Joint category (or classification) is as follows: Figure 2B Image 267 is shown. Figure 2B At box 268 of process 260, the decoder of the neural network of machine learning system 257 can process classification (and in some cases, learned queries, such as learned joint queries described below) to determine the 3D pose of one or more objects in image 261 (shown in image 269). For example, using classification and learned queries, the decoder of the neural network of machine learning system 257 can determine the 3D pose of one or more hands and / or the pose of a physical object held or near one or more hands (e.g., occluded by one or more hands) in image 261. Machine learning system 257 may include a neural network independent of the neural network of machine learning system 252. In an illustrative example, the neural network of machine learning system 257 includes a transformer neural network, which includes a transformer encoder and a transformer decoder. As described in detail below, the 3D pose of one or more hands (and / or physical objects) can be determined using various types of pose representations (e.g., parent-relative joint vectors, parent-relative 2.5D poses, or joint angles (or other point angles)).

[0065] As mentioned above, hand gesture estimation can be difficult, and existing solutions are inadequate for various reasons. Figure 2AThe pose estimation system 200 can be used to identify the joints of both hands (or the joints of one or both hands and the location of points on another physical object in an image that is grasped or near one or more hands) and jointly predict their 3D positions and / or angles using a neural network (e.g., a transformer neural network) of the machine learning system 257, as described herein. For example, as described above, the keypoint determination engine 250 and the feature determination engine 254 can use the determined keypoints (e.g., determined as local maxima in a 2D graph or heatmap) to locate joints in 2D, which can lead to accurate 3D pose. At this stage, keypoints may not yet be associated with a specific joint. In some cases, one or more keypoints may not correspond to a joint at all, and some joint points may not be detected as keypoints (e.g., due to occlusion or other factors). However, keypoints can provide a useful starting point for predicting accurate 3D pose for both hands (and / or one or both hands and physical objects in an image that are grasped or near one or more hands). The pose estimation engine 256 can perform joint association (or classification) and pose estimation using a neural network of an end-to-end trained machine learning system 257 (e.g., a transformer encoder-decoder architecture) and a neural network of a machine learning system 252 for keypoint detection. This system architecture collaboratively analyzes the hand joint positions in the input image, resulting in more reliable pose estimations compared to other existing methods, such as during close interaction between the hand and / or other physical objects. The neural network architecture of the machine learning system 257 (e.g., a transformer neural network architecture) can also accept varying numbers of inputs, allowing the system to handle the fact that different numbers of keypoints can be detected in different input images. The self-attention of the transformer neural network and the variability in the number of inputs allow the machine learning system 257 to accurately determine the 3D pose of the hand and / or other physical objects in the image.

[0066] As previously mentioned, in some implementations, the machine learning system 257 of the pose estimation engine 256 may include a transformer neural network. The transformer architecture can be designed to estimate single-hand pose, two-hand pose, and hand-object pose (where one or more hands in the image are grasping or near a physical object) from an input image (e.g., an RGB image). The transformer neural network can model the relationships between features at different locations (e.g., each location) in the image, which in some cases can increase computational complexity as the resolution of the feature maps increases. Generally, due to this constraint of the transformer neural network, the transformer typically operates on lower-resolution feature maps that cannot capture finer image details, such as closely spaced hand joints. As indicated by the experimental results provided below, lower-resolution feature maps may be insufficient for accurately estimating hand pose. One solution to this problem is to allow features at each spatial location to focus on a small set of features from sampling locations across different scales, thereby more accurately detecting small objects in the image. The pose estimation system 200 can model the relationships between sampled features from high-resolution and low-resolution feature maps, where the sampling locations are keypoints provided by a convolutional neural network (CNN), which is efficient in detecting finer image details. For the pose estimation task, sparse sampled features are efficient in accurately estimating the 3D pose of hands and / or physical objects when they are interacting closely with each other.

[0067] Figure 3 This is a diagram illustrating another example of a process 300 that can be performed by the pose estimation system 200. In example process 300, a U-shaped neural network (including an encoder part and a decoder part) is used as an example implementation of a machine learning system 252 for a keypoint determination engine 250, and a transformer neural network (including a transformer encoder 366 and a transformer decoder 368) is used as an example implementation of a machine learning system 257 for a pose estimation engine 256. Process 300 will be described as estimating two hands in an image (e.g., ...). Figure 1A The pose of the two hands in the image. However, it will be understood that process 300 can be used to estimate the pose of a single hand and the pose of a physical object (e.g., a bottle) that is being held or near the hand (e.g., being occluded or occluding the hand), the pose of two hands and the pose of one or more physical objects that are being held or near the two hands, and / or the pose of two or more physical objects that are interacting with each other (e.g., touching each other, occluding each other, etc.).

[0068] At operation 362, the machine learning system 252 of the pose estimation system 200 can process the input image 361 to generate a keypoint heatmap 363. As described above, the machine learning system 252 in Figure 3The image is shown as a U-shaped neural network. Machine learning system 252 can perform keypoint detection to detect keypoints (from keypoint heatmap 363) that may correspond to the 2D hand position in input image 361. At operation 364, feature determination engine 254 of pose estimation system 200 can encode (e.g., encode as feature vectors) features extracted from machine learning system 252 based on keypoints from heatmap 363. For example, feature determination engine 254 can use the keypoints to determine which features to extract from one or more feature maps generated by the decoder portion of the U-shaped neural network.

[0069] At operation 365, the pose estimation system 200 can use encoded features as input to a transformer encoder 366 (which may be part of a machine learning system 257 of the pose estimation engine 256). The transformer encoder 366 can determine a joint type category or type for each feature representation of one or more hands in image 361 (corresponding to each visible joint in image 361). Using a learned joint query 367 (described below) and the joint categories output by the transformer encoder 366, the transformer decoder 368 can predict pose parameters relative to each joint of both hands. In some cases, the transformer decoder 368 can also predict additional parameters, such as translation between the hands and / or hand shape parameters. In some examples, the pose estimation system 200 may consider an auxiliary loss on the transformer encoder 366 to identify keypoints. The auxiliary loss may not directly affect the pose estimation, but it can guide the transformer decoder 368 to select more appropriate features, thereby significantly improving accuracy.

[0070] Further details regarding operations 362 and 364 (performing keypoint detection and encoding) of process 300 will now be described. At operation 362, given an input image 361, a keypoint determination engine 250 (e.g., using a machine learning system 252) can extract keypoints that may correspond to 2D hand joint positions. In an illustrative example, the machine learning system 252 can predict a keypoint heatmap 363 (which can be represented as a heatmap H) from the input image 361 using a U-network architecture. The predicted heatmap 363 can have a single channel (e.g., a single value, such as a value between 0 and 1 as described above). For example, heatmap 363 can have a dimension of 128x128x1 (corresponding to a 128x128 2D array, where each position in the 2D array has a single value). In some cases, the keypoint determination engine 250 may select local maxima of heatmap 363 as keypoints, as described above regarding... Figure 2AAs described above. In an illustrative example, the keypoint determination engine 250 may determine the maximum value of N keypoints (e.g., having N = 100 or other integer values) by determining the N local maxima of the heatmap 363, for example. At this stage, the pose estimation system 200 may not attempt to identify which keypoint corresponds to which joint.

[0071] The keypoint identification engine 250 can learn to predict heatmaps by applying a 2D Gaussian kernel to each of several ground-based joint locations and using L2 norm loss. To calculate the ground-based heat map

[0072]

[0073] Ground-based joint locations may include annotated joint locations (labels for training), provided along with a dataset used for training (and, in some cases, testing). The annotated joint location labels can serve as training data for supervised training of machine learning systems 252 and / or 257. For example, the loss function provided by equation (1) can be used based on heatmaps output by machine learning system 252. The loss is determined by the predicted joint positions and the ground-based joint positions provided by the labels. Based on the loss at each training iteration or epoch, the parameters of the machine learning system 252 (e.g., weights, biases, etc.) can be tuned to minimize the loss.

[0074] In some cases, the U-Net architecture uses a ResNet architecture up to layer C5 as the backbone, followed by upsampling layers and convolutional layers with jump connections. In an illustrative example, input image 361 can have a resolution of 256×256 pixels, and heatmap 363 can have a resolution of 128×128 pixels.

[0075] At operation 364, the feature determination engine 254 can then extract features around each keypoint from one or more feature maps of the machine learning system 252 (e.g., all or some feature maps from the decoder portion of the U-network architecture). For example, in some examples, such as Figure 3As shown in points 369a, 369b, 369c, 369d, 369e, 369f, 369g, 369h, and 369i, the feature determination engine 254 can determine linear interpolated features from multiple layers of the U-Net decoder. These linear interpolated features can be extracted from the feature map. The feature determination engine 254 can use the extracted features, along with spatial encoding in some cases, to represent keypoints as feature representations (e.g., feature vectors). The feature determination engine 254 can combine features (e.g., by concatenating features from different feature maps) to form feature representations, such as feature vectors. For example, the feature determination engine 254 can concatenate features to generate a 3968-dimensional (3968-D) feature vector. In some cases, the feature determination engine 254 can use a three-layer multilayer perceptron (MLP) to downscale the 3968-D feature vector to a 224-D encoded vector. In some examples, the feature determination engine 254 can further combine (e.g., concatenate) spatial or positional encodings (e.g., 32-D sinusoidal positional encoding). Using the 224-D encoded vector from the example above, the 32-D sinusoidal position encoded vector can be combined with the 224-D encoded vector to form a 256-D vector representation for each keypoint. This 256-D vector representation for each keypoint can be provided as input to the transformer encoder 366. In some cases, keypoint detection via non-maximum suppression is indiscriminate, and gradients do not flow through the peak detection operation during training.

[0076] The feature determination engine 254 can output the feature representation to the storage device 226 (in which case, the transformer encoder 366 can obtain the feature representation from the storage device 226), and / or can provide the feature representation as input to the transformer encoder 366. The transformer encoder 366 and the transformer decoder 368 can use the feature representation (e.g., each of the 256-D vector representations described above) to predict the 3D pose of the two hands in the input image 361. Once operations 362 and 364 are performed, for each keypoint K... i Generate encoding vector For example, the encoding vector This can be represented by the 256-D vector above. The converter encoder 366 can use the encoded vector. As input, the transformer encoder 366 may include a self-focused module or engine, which can process the encoded vector. The relationships between associated keypoints are modeled, and global context-aware features can be generated. These global context-aware features help associate each keypoint with a hand joint (e.g., by using encoded vectors). (Classified by joint category or type). In some implementations, to help the transformer encoder 366 model this relationship, an auxiliary joint association loss can be used to train the transformer encoder 366. Using the learned joint query 367 as input, the transformer decoder 368 processes the joint-aware (based on joint category or type) features from the transformer encoder 366 to predict the 3D pose of the hand in the input image 361.

[0077] The learned joint query 367 used by the transformer decoder 368 can correspond to a joint type embedding. The joint query 367 can be transformed by a series of self-attention and cross-attention modules in the decoder. For example, for each joint query, the cross-attention module or engine in the decoder 368 can soft-select features from the transformer encoder 366 that best represent the queried hand joint, and can transform the selected features (e.g., by passing the features through one or more layers, such as one or more MLP layers). For example, the cross-attention module or engine can soft-select features by determining a linear combination of features. The transformed features can be fed into a feedforward network (FFN) within the transformer decoder 368 to predict joint-related pose parameters. The FFN can include two MLP layers: a linear projection layer and a layer with standard cross-entropy loss. The softmax layer is used. In some examples, multiple decoder layers can be used. In some cases, the transformer decoder 368 can use an FFN with shared weights to predict the pose after each decoder layer. For example, cross-entropy loss can be applied after each layer. The FFN used for joint type prediction can share weights across layers. In some examples, together with the joint query 367, the transformer decoder 368 can use additional learned queries to predict the relative translation between hands. MANO hand shape parameters (For example, it could be 10D or 10-D) and / or weak perspective camera model parameters (scale). and 2D translation ).

[0078] In some examples, (x) can be obtained by performing a proximity test. i ,y iThe joint association is defined as the ground truth joint type for the detected keypoint. For example, if the distance from a keypoint to the nearest joint is less than a threshold distance γ, the keypoint can be assigned the joint type of the nearest joint in the 2D image plane. When multiple joints are within the threshold distance γ from the keypoint, the joint with the smallest depth can be selected as the joint type for that keypoint. If there are no joints within the threshold distance γ, the keypoint can be assigned to the background class. In some cases, this joint association may result in multiple keypoints being assigned to a single joint type. However, as mentioned above, joint association is not directly used in pose estimation but serves as a guide for the transformer decoder to select appropriate features for pose estimation.

[0079] The pose estimation system 200 can output a 3D pose 258 for one or more objects in an input image. In some cases, the 3D pose may include 3D pose parameters that define the orientation and / or translation of objects in the input image (e.g., a hand in an image). As described below, the pose parameters can be defined using various pose representations. The pose estimation system 200 can use the 3D pose 258 for various purposes, such as generating a 3D model of an object in the determined pose, performing operations based on the pose (e.g., navigating a vehicle, robotic device, or other device or system around an object), and other uses. In the implementation where a 3D model is generated, the 3D model can be displayed (e.g., by a mobile device, by an XR system, and / or by other systems or devices), used to determine the position of objects represented by the 3D model (e.g., for scene understanding and / or navigation, for object grasping, for autonomous vehicle operation, and / or for other uses), and other uses.

[0080] Various pose representations and losses can be used for the estimated 3D pose 258. In some cases, direct regression of 3D joint positions can be more accurate (in terms of joint error) than regression of model parameters (e.g., MANO joint angles from a CNN architecture). However, regression of MANO joint angles provides access to the complete hand mesh needed to model contact and interpenetration during interaction or to learn in a weakly supervised setting. Pose estimation system 200 can be configured to output multiple types of pose representations (e.g., 3D joint positions and joint angles). Using the techniques described herein, pose estimation system 200 can provide joint angle representations while achieving performance competitive with joint position representations.

[0081] Examples of pose representations that can be used include the parent relative joint vector. Parent relative 2.5D pose and MANO joint angle The parent relative joint vector and the MANO joint angle can be the root relative pose. For the parent relative 2.5D pose, the complete pose is reconstructed from the estimated parameters using the absolute root depth, thus obtaining the absolute pose.

[0082] In some cases, a weak perspective camera model is used to determine the root relative 3D joint position of each hand. Projected onto the image plane

[0083]

[0084] Where Π represents orthographic projection. The scale parameters of a weak perspective camera model. The translation parameters of a weak perspective camera model.

[0085] For the parent relative joint vector This means that each joint j can be coupled with V. j =J 3D (j)-J 3D (p(j))J 3D p(j) gives the 3D "joint vector" V j Related, where J 3D This refers to the 3D joint position, where p(j) is the index of the parent joint of point j. Parent relative joint vector. The advantage of this representation is that it defines the hand pose relative to its root without requiring knowledge of the camera's intrinsic parameters. This solution can be useful when the camera's intrinsic parameters are unavailable or require extensive computation to determine.

[0086] Figure 4 This is a diagram showing an example of a hand with various joints. Figure 4 Each joint in the model is labeled according to its joint type, including thumb tip (F0), index finger tip (F1), middle finger tip (F2), ring finger tip (F3), little finger tip (F4), distal interphalangeal (DIP) joint, proximal interphalangeal (PIP) joint, metacarpophalangeal (MCP) joint, interphalangeal (IP) joint, and Carpo-Metacarpal (CMC) joint. Each joint has a parent joint. In one example, the parent joint of the DIP joint is the PIP joint, the parent joint of the PIP joint is the MCP joint, and so on. Parent relative joint vectors are used. This means that 3D vectors can be determined from each joint to its parent joint (or alternatively from the parent joint to the child joint). In some cases, 20 joint vectors can be determined for each hand, for example, for... Figure 4The diagram shows a joint vector from each joint to each corresponding parent joint (or alternatively from a parent joint to each corresponding child joint). Based on the joint vectors, the pose estimation system 200 can compute the root relative 3D position of each joint by performing an accumulation function (e.g., by summing the parent relative joint vectors).

[0087] When using parent relative joint vector In this representation, the neural network of the machine learning system 257 of the pose estimation engine 256 can use 3D joint loss (represented as...). ) and joint vector loss (represented as in and V * The neural network of the machine learning system 257 can be trained using reprojection loss (representing ground real-time values). In some cases, the neural network can use reprojection loss (representing ground real-time values). Training was conducted, among which It is the ground-based 2D joint position, and This is a weak perspective camera projection (e.g., the weak perspective camera projection defined in equation (2) above). The loss can be calculated using the L1 distance between the estimated value and the ground reality value. The attitude loss can be defined as:

[0088] For the parent relative to the 2.5D pose representation, each joint is represented by its 2D position J. 2D And the difference Z between its depth and the depth of its parent joint. p Parameterization. Then, the camera intrinsic parameter matrix K, and the absolute depth Z of the root (wrist) joint. root The proportions of the hand are used to reconstruct the 3D pose of the hand in the camera coordinate system, as shown below:

[0089]

[0090] Among them, Z r From Z p The relative depth of the joint root was obtained, and It is J 2D The x and y coordinates. Each joint query is used as input to the transformer decoder 368 to predict the x and y coordinates for each of the 20 joints and the point on the wrist. Figure 4 J (marked as "wrist") 2D and Z p In some cases, 43 joint queries can be used when estimating 2.5D pose (including one query for relative hand translation, such as in...). Figure 4 (at the wrist position).

[0091] The pose estimation engine 256 can use a machine learning system 257 to predict the root depth Z.root For example, the neural architecture of machine learning system 257 can be used at 2D joint locations. and relative depth The L1 loss is used for training, where and These are the estimated 2.5D attitude and the actual 2.5D attitude on the ground. Attitude loss can be calculated by... Provided.

[0092] In the MANO joint angle representation, each 3D hand pose can be represented by 16 3D joint angles in the hand kinematics tree. In this example, to train the neural network architecture of the machine learning system 257 of the pose estimation engine 256, the 3D joint positions can be... Joint angle and joint reprojection Apply L1 loss. The total pose loss can be calculated using... The joint angle loss can be represented as regularization and can avoid unrealistic poses. In some cases, the pose estimation engine 256 can only estimate the root relative pose in this representation, and therefore may not require camera intrinsic parameters to estimate the pose. Therefore, this representation may be advantageous when camera intrinsic parameters are unavailable.

[0093] In some examples, pose estimation system 200 (e.g., machine learning system 252 and machine learning system 257) can be trained end-to-end by performing end-to-end training, where all stages are connected to the final loss, as follows:

[0094]

[0095] During the first few epochs of training (in which case the estimated keypoint heatmap may not be very accurate), the ground truth heatmap can be fed to the multi-scale feature sampler, and then switched to the estimated heatmap.

[0096] Table 1 below shows the results of the system and techniques described in this paper with three different pose representations on the InterHandV0.0 dataset (the last three rows of the table). Figure 3 As shown, the pose estimation system 200 achieves 14% higher accuracy compared to InterNet (described in Gyeongsik Moon et al., “Interhand 2.6m: A dataset and baseline for 3D interacting hand poseestimation from a single RGB image”). In ECCV, it uses a CNN architecture.

[0097]

[0098] Table 1

[0099] The asterisk (*) for the first MANO joint angle entry in Table 1 indicates the ground-based 3D joint obtained according to the fitted MANO model. MPJPE measures the Euclidean distance (in mm) between the predicted 3D joint position and the ground-based 3D joint position after root joint alignment and indicates the accuracy of the root-relative 3D pose. Alignment was performed separately for the right and left hands. MRRPE measures the left hand's position relative to the right hand (in mm) using Euclidean distance. The InterNet system uses a full CNN architecture to predict the 2.5D pose of the interacting hands. Even when using joint vector representations that do not require camera intrinsic parameters to reconstruct the pose, the pose estimation system 200 achieves better accuracy than InterNet, particularly with respect to the interacting hands. The pose estimation system 200 achieves a 3 mm (or 14%) improvement on the interacting hands when predicting similar 2.5D poses. These results suggest that explicitly modeling the relationships between CNN features belonging to hand joints using a transformer is more accurate than directly using CNNs to estimate pose.

[0100] The parent relative joint vector representation (which does not require camera intrinsic parameters to reconstruct the root relative pose) also outperforms InterNet (which requires camera intrinsic parameters and is slightly less accurate than the 2.5D pose representation). This decrease in accuracy can be attributed to the fact that the parent relative joint vector representation fully predicts the 3D pose, unlike the 2.5D representation, which partially predicts the 2D pose and relies on known camera intrinsic parameters to project the pose into 3D. Table 1 also shows that, using the MANO joint angle representation, the pose estimation system 200 performs similarly to InterNet, which outputs direct 3D joint positions. This is important because previous work estimating joint angle representations or their PCA components has reported results that are generally inferior to methods that directly estimate 3D joint positions, suggesting that regressing model parameters is more difficult than estimating joint positions. With the aid of a transformer architecture, the example neural network architecture of the machine learning system 257 described above, which softly selects multi-scale CNN features specific to each joint position in the input image, achieves accurate estimation of any joint-related parameters, regardless of their representation.

[0101] Another dataset on which the pose estimation system 200 can be evaluated is the HO-3D dataset. The HO-3D dataset is helpful for evaluating the pose estimation system 200 when estimating the pose of one or more hands and the pose of objects held or near those hands. The HO-3D dataset includes hand-object interaction sequences for only the right hand and 10 objects from the YCB object dataset. Annotations for training using the HO-3D dataset can be automatically obtained and can include 66,000 training images and 11,000 test images. In some cases, 20 hand pose joint queries, one weak perspective camera model parameter query, and two object pose queries can be used. The mean joint error and area under the curve (AUC) metric after scaling and translation alignment of the root joints can be used to evaluate the hand pose results. The object pose can be calculated relative to the hand's frame of reference. Mean 3D angular error. It can be used to evaluate the accuracy of an object's pose. In some cases, 3D angular error can be used. The variation is used to consider the symmetry of the object, where and P * These refer to the estimated attitude matrix and the ground-based actual attitude matrix, respectively. The object attitude error metric can be defined as:

[0102]

[0103]

[0104] Among them B i Let represent the i-th corner of the object's bounding box, and S represent the 3D rotation set of objects that do not change the object's appearance. The HO-3D test set contains three visible objects (mustard bottle, bleach, and canned meat) and one object not seen in the training data. Only the visible object is used for the evaluation below. Hand pose can be estimated using joint vector representations, and this technique can be trained and tested on loosely cropped 256x256 image patches.

[0105] Table 2 below shows the accuracy of the pose estimation system 200 evaluated on the HO-3D dataset relative to other methods, including the technique described by Shreyas Hampali et al. in “Honnotate: A method for 3D notation of hand and object poses” (hereinafter referred to as “Honnotate”) at CVPR 2020; the technique described by Yana Hasson et al. in “Learning joint reconstruction of hands and manipulated objects” (hereinafter referred to as “Learning Joint Reconstruction”) at CVPR 2019; and the technique described by Yana Hasson et al. in “Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction” (hereinafter referred to as “Leveraging Photometric Consistency”) at CVPR 2020.

[0106] Table 3 below compares the accuracy of object pose estimation using the pose estimation system and "using photometric consistency" which uses the mean object angle error. Using photometric consistency uses a CNN backbone followed by fully connected layers to estimate object pose, which regress object rotation (axis-angle representation) and object translation in the camera coordinate system. Since using photometric consistency does not consider object symmetry, results are shown with and without symmetry consideration during training and testing. Pose estimation system 200 obtains a more accurate hand-relative object pose. As shown in Tables 2 and 3, pose estimation system 200 performs significantly better than existing methods. All errors are in centimeters (cm).

[0107]

[0108] Table 2

[0109]

[0110]

[0111] Table 3

[0112] Introducing the H2O-3D dataset, this dataset contains 3D pose annotations for both hands and objects, and is automatically annotated. The dataset was captured using six subjects manipulating ten different YCB objects with functional intent using both hands. It was captured in a multi-view setup with five RGB depth (RGBD) cameras and includes 50,000 training images and 12,000 test images. The H2O-3D dataset is potentially more challenging than previous hand interaction datasets due to the significant occlusion between the hands and objects. Figure 5 Example images and annotated poses from the H2O-3D dataset are included. The pose estimation system 200 can estimate the parent-relative joint vector representations of both hand poses (40 joint queries), the left-hand relative translation to the right-hand (1 query), and the right-hand relative object pose (2 queries), for a total of 43 queries at the transformer decoder. Training data from the HO-3D dataset can be used, and images can be randomly flipped during training to obtain only right-hand and left-hand images. This data can then be combined with the H2O-3D training dataset. The H2O-3D test set includes three objects (water jug ​​base, bleach, and drill) visible in the training data. Object pose estimation accuracy can only be evaluated on these objects.

[0113] After root joint alignment, the MPJPE metric was used to evaluate the accuracy of hand pose estimation, and the MRRPE metric was used to evaluate the relative translation between the hands. The symmetry-perceived mean 3D angular distance metric defined in Equation (5) was used to evaluate object pose. Table 5 below shows the accuracy of hand pose estimation using the pose estimation system 200 on the H2O-3D dataset. Estimating the translation between the hands was observed to be more challenging due to significant mutual occlusion. Table 4 shows the accuracy of object pose estimation by the pose estimation system 200.

[0114]

[0115] Table 4

[0116]

[0117] Table 5

[0118] Figure 6This is a graph showing the cross-focus of three joint queries for the right hand, namely the index finger's fingertip joint (red), the middle finger's PIP joint (blue), and the little finger's MCP joint (yellow). The first column contains the input image, and the second column contains a keypoint heatmap generated by the keypoint determination engine 250 for the input image. For each joint query, the corresponding colored circle in the third column (labeled "Joint Focus") indicates the location of the keypoint focused by that query. The radius of the circle in the Joint Focus column is proportional to the focus weight. For each joint query, Figure 3 The transformer decoder 368 (as an example of a neural network that can be used in the machine learning system 257 of the pose estimation engine 256) can select image features only from the joint positions. Joint-specific features enable the estimation of various joint-related pose parameters, such as joint angles and joint vectors. The output pose for the input image is... Figure 6 It is shown in the fourth column.

[0119] Figure 7 An example of a process 700 for determining one or more poses of one or more objects using the techniques described herein is shown. This process is used to perform... Figure 7 The functional units of one or more boxes shown may include hardware and / or software components of a computer system, such as those with... Figure 2A One or more components of the attitude estimation system 200 and / or Figure 10 The computer system shown is a computing device architecture 1000.

[0120] At box 702, process 700 includes determining multiple keypoints from an image using a machine learning system. The multiple keypoints are associated with at least one object in the image. In some cases, the at least one object includes at least one hand. In some examples, each of the multiple keypoints corresponds to a joint of at least one object (e.g., a joint of at least one hand). In some aspects, the machine learning system includes a neural network that uses the image as input. In an illustrative example, the neural network is... Figure 2A The keypoint determination engine 250 is part of the machine learning system 252. Units used to perform the functionality of box 702 may include one or more software and / or hardware components of a computer system, such as the keypoint determination engine 250 (e.g., utilizing the machine learning system 252), and one or more components of the processing system 230 (e.g., CPU 232, GPU 234, DSP 236, and / or ISP 238). Figure 10 The processor 1010 and / or other software and / or hardware components of the computer system.

[0121] At box 704, process 700 includes determining multiple features from a machine learning system based on multiple keypoints. In some examples, to determine multiple features from a machine learning system based on multiple keypoints, process 700 may include determining a first set of features corresponding to the multiple keypoints from a first feature map of the machine learning system. The first feature map includes a first resolution. Process 700 may further include determining a second set of features corresponding to the multiple keypoints from a second feature map of the machine learning system. The second feature map includes a second resolution different from the first resolution.

[0122] The unit used to perform the functionality of box 704 may include one or more software and / or hardware components of a computer system, such as feature determination engine 254 (e.g., utilizing machine learning system 252), and one or more components of processing system 230 (e.g., CPU 232, GPU 234, DSP 236, and / or ISP 238). Figure 10 The processor 1010 and / or other software and / or hardware components of the computer system.

[0123] At box 706, process 700 includes classifying multiple features into multiple joint types. In some aspects, process 700 may include generating a feature representation for each of the multiple keypoints. In such aspects, process 700 may use the feature representation for each keypoint to classify the multiple features into multiple joint types. In some cases, the feature representation for each keypoint includes an encoded vector.

[0124] The unit used to perform the functionality of box 706 may include one or more software and / or hardware components of a computer system, such as pose estimation engine 256 (e.g., utilizing machine learning system 257), one or more components of processing system 230 (e.g., CPU 232, GPU 234, DSP 236, and / or ISP 238). Figure 10 The processor 1010 and / or other software and / or hardware components of the computer system.

[0125] At box 708, process 700 includes determining pose parameters for at least one object (e.g., at least one hand or other object) based on multiple joint types. In some examples, a neural network can be used to classify multiple features into multiple joint types and determine the pose parameters. For example, process 700 can use an encoder of a transformer neural network to classify multiple features into multiple joint types, and can use a decoder of a transformer neural network to determine pose parameters for at least one object (e.g., at least one hand or other object) based on multiple joint types.

[0126] In some examples, at least one object includes two objects. For example, multiple keypoints include keypoints for both objects, and pose parameters include pose parameters for both objects. In one illustrative example, using a hand as an example of an object, at least one object includes two hands. In such examples, multiple keypoints may include keypoints for both hands, and pose parameters may include pose parameters for both hands. In another illustrative example, at least one object includes a single hand. In some examples, the image includes at least one hand (e.g., a single hand or two hands as examples of at least one object) and at least one physical object that is held or near the at least one hand (e.g., occluded by the at least one hand). In such examples, process 700 may include determining multiple object keypoints from the image using a machine learning system. The multiple object keypoints are associated with objects associated with the at least one hand (e.g., objects held or near the at least one hand). Process 700 may include determining pose parameters for the object based on the multiple object keypoints using the techniques described herein.

[0127] In some examples, pose parameters are determined based on multiple joint types and one or more learned joint queries for at least one object (e.g., at least one hand or other object). In some cases, at least one hand includes a first hand and a second hand. In this case, one or more learned joint queries can be used to predict at least one of the following: the relative translation between the first and second hands, the set of object (e.g., hand or other object) shape parameters, and camera model parameters.

[0128] As described above, various types of pose representations can be used for pose parameters. In some aspects, pose parameters for at least one object (e.g., at least one hand or other object) include a three-dimensional vector for each joint in a plurality of joint types. For example, the three-dimensional vector for each joint in a plurality of joint types may include horizontal, vertical, and depth components. In another example, the three-dimensional vector for each joint in a plurality of joint types may include a vector between each joint and a parent joint associated with each joint. In some aspects, pose parameters for at least one object (e.g., at least one hand or other object) include the position of each joint and the difference between the depth of each joint and the depth of the parent joint associated with each joint. In some aspects, pose parameters for at least one object (e.g., at least one hand or other object) include the translation of at least one object (e.g., at least one hand or other object) relative to another object (e.g., another hand or physical object) in the image. In some cases, pose parameters for at least one object (e.g., at least one hand or other object) include the shape of at least one object (e.g., at least one hand or other object).

[0129] The unit used to perform the functionality of box 708 may include one or more software and / or hardware components of a computer system, such as pose estimation engine 256 (e.g., utilizing machine learning system 257), one or more components of processing system 230 (e.g., CPU 232, GPU 234, DSP 236, and / or ISP 238). Figure 10 The processor 1010 and / or other software and / or hardware components of the computer system.

[0130] 20. The device of claim 1, wherein the at least one processor is configured to determine user input based on the attitude parameters.

[0131] 21. The device of claim 1, wherein the at least one processor is configured to render virtual content based on the pose parameters.

[0132] (For example, above or below the hand). The hand is very important for occlusion. We might want to render content "below" the hand, and there might also be use cases for rendering "above" the hand (e.g., to allow user input without disrupting / occluding the virtual screen).

[0133] In some examples, process 700 may be performed by a computing device or apparatus (e.g., having...) Figure 10 The computing device architecture 1000 shown is used for execution. In one example, process 700 can be performed by a computing device with an implementation... Figure 2AThe computational device architecture 1000 of the pose estimation system 200 shown is executed by the computational device. In some cases, the computational device or apparatus may include an input device, a key point determination engine (e.g., key point determination engine 250), a feature determination engine (e.g., feature determination engine 254), a pose estimation engine (e.g., pose estimation engine 256), an output device, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components configured to perform the steps of process 700. In some examples, the computational device or apparatus may include a camera configured to capture images. For example, the computational device may include a camera device. As another example, the computational device may include a mobile device or part of a mobile device, which may include one or more cameras (e.g., a mobile phone or tablet computer including one or more cameras), an XR device (e.g., a head-mounted display, XR glasses, or other XR device), a vehicle, a robotic device, or other device that may include one or more cameras. In some cases, the computational device may include a communication transceiver and / or a video codec. In some cases, the computational device may include a display for displaying images. In some examples, the camera or other capturing device that captures video data is separate from the computational device, in which case the computational device receives the captured video data. The computing device may further include a network interface configured to transmit video data. The network interface may be configured to transmit Internet Protocol (IP) based data or any other suitable data.

[0134] Components of a computing device (e.g., one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) can be implemented in circuitry. For example, a component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or may include computer software, firmware, or a combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or a combination thereof for performing the various operations described herein.

[0135] Process 700 is shown as a logic flowchart, where operations represent a sequence of operations that can be implemented by hardware, computer instructions, or a combination thereof. In the context of computer instructions, operations represent computer-executable instructions stored on one or more computer-readable storage media, which, when executed by one or more processors, perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as limiting, and any number of described operations can be combined in any order and / or in parallel to implement the process.

[0136] Furthermore, process 700 can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or a combination thereof. As described above, the code can be stored, for example, in the form of a computer program comprising multiple instructions executable by one or more processors on a computer-readable or machine-readable storage medium. The computer-readable or machine-readable storage medium can be non-transitory.

[0137] As described above, various aspects of this disclosure can utilize machine learning systems, such as machine learning system 252 for keypoint determination engine and machine learning system 257 for pose estimation engine 256. Figure 8 This is an illustrative example of a deep learning neural network 800 that can be used to implement the overall video understanding system described above. Input layer 820 includes input data. In one illustrative example, input layer 820 may include data representing pixels of an input video frame. Neural network 800 includes multiple hidden layers 822a, 822b through 822n. Hidden layers 822a, 822b through 822n include "n" hidden layers, where "n" is an integer greater than or equal to one. The number of hidden layers can be as many as required for a given application. Neural network 800 also includes an output layer 821, which provides the output produced by the processing performed by hidden layers 822a, 822b through 822n. In one illustrative example, output layer 821 can provide a classification of objects in the input video frame. The classification may include categories that identify the type of activity (e.g., playing football, playing the piano, listening to the piano, playing the guitar, etc.).

[0138] Neural network 800 is a multi-layered neural network with interconnected nodes. Each node can represent a piece of information. The information associated with a node is shared between different layers, and each layer retains the information as it is processed. In some cases, neural network 800 may include a feedforward network, in which there is no feedback connection that feeds the network's output back to itself. In some cases, neural network 800 may include a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read.

[0139] Information can be exchanged between nodes through node-to-node interconnections between layers. Nodes in input layer 820 can activate the node set in the first hidden layer 822a. For example, as shown, each input node in input layer 820 is connected to each node in the first hidden layer 822a. Nodes in the first hidden layer 822a can transform the information of each input node by applying an activation function to the input node information. The information derived from this transformation can then be passed to nodes in the next hidden layer 822b, and the nodes in the next hidden layer 822b can be activated, allowing them to perform their own specified functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable functions. The output of hidden layer 822b can then activate nodes in the next hidden layer, and so on. The output of the final hidden layer 822n can activate one or more nodes in output layer 821, providing the output at those nodes. In some cases, although a node in neural network 800 (e.g., node 826) is shown as having multiple output lines, the node has a single output and all lines shown as outputs from the node represent the same output value.

[0140] In some cases, each node or the interconnection between nodes can have weights, which are a set of parameters derived from the training of the neural network 800. Once the neural network 800 is trained, it can be called a trained neural network, which can be used to classify one or more activities. For example, the interconnection between nodes can represent a piece of information about what the interconnected nodes have learned. The interconnection can have tunable numerical weights that can be tuned (e.g., based on the training dataset), allowing the neural network 800 to adapt to the input and learn as more and more data is processed.

[0141] The neural network 800 is pre-trained to process features from the data in the input layer 820 using different hidden layers 822a, 822b to 822n, in order to provide an output through the output layer 821. In an example where the neural network 800 is used to identify an activity being performed by a driver in a frame, the neural network 800 can be trained using training data that includes both frames and labels, as described above. For example, training frames can be input into the network, where each training frame has a label indicating either the features in the frame (for a feature extraction machine learning system) or a label indicating the category of activity in each frame. In an example using object classification for illustrative purposes, the training frame could include an image of the number 2, in which case the image label could be [0010000000].

[0142] In some cases, the neural network 800 can use a training process called backpropagation to adjust the weights of its nodes. As mentioned above, the backpropagation process can include forward pass, loss function, back pass, and weight update. For each training iteration, forward pass, loss function, back pass, and parameter update are performed. For each set of training images, this process can be repeated a certain number of iterations until the neural network 800 is trained well enough that the weights of each layer are accurately tuned.

[0143] For an example of recognizing objects in a frame, the forward pass may include passing a training frame through a neural network 800. The weights are initially randomized before training the neural network 800. As an illustrative example, a frame may include a numerical array representing pixels of an image. Each number in the array may include a value from 0 to 255, describing the pixel intensity at that location in the array. In one example, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luminance and two chromaticity components, etc.).

[0144] As mentioned above, for the first training iteration of the neural network 800, the output will likely include values ​​due to the weights, which are randomly selected during initialization, without prioritizing any particular class. For example, if the output is a vector with probabilities that an object includes different classes, the probability values ​​for each distinct class can be equal or at least very similar (e.g., for ten possible classes, each class could have a probability value of 0.1). Using the initial weights, the neural network 800 cannot determine low-level features and therefore cannot accurately determine what the object's classification might be. A loss function can be used to analyze the error in the output. Any suitable loss function can be defined, such as cross-entropy loss. Another example of a loss function includes mean squared error (MSE), defined as... The loss can be set to equal E.total The value of .

[0145] For the first training image, the loss (or error) will be high because the actual value will be significantly different from the predicted output. The goal of training is to minimize the loss so that the predicted output matches the training label. A neural network 800 can perform backpropagation by determining which inputs (weights) contribute most to the network's loss, and the weights can be adjusted to reduce and ultimately minimize the loss. The derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a specific layer) can be calculated to determine the weights that contribute most to the network's loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. A weight update can be expressed as... Where w represents the weight, indicating that w i The initial weights are given, and η represents the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes larger weight updates, and a lower value indicates smaller weight updates.

[0146] Neural Network 800 can include any suitable deep network. An example includes a Convolutional Neural Network (CNN), which consists of an input layer and an output layer, with multiple hidden layers between them. The hidden layers of a CNN include a series of convolutional layers, non-linear layers, pooling (for downsampling) layers, and fully connected layers. Neural Network 800 can also include any other deep network besides CNNs, such as autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), etc.

[0147] Figure 9 This is an illustrative example of a Convolutional Neural Network (CNN) 900. The input layer 920 of the CNN 900 includes data representing an image or frame. For example, the data could include a numerical array representing pixels of an image, where each number in the array includes a value from 0 to 255 describing the pixel intensity at that location in the array. Using the previous example from above, the array could include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luminance and two chroma components, etc.). The image can be passed through a convolutional hidden layer 922a, an optional non-linear activation layer, a pooling hidden layer 922b, and a fully connected hidden layer 922c to obtain an output at the output layer 924. Although Figure 9 Only one hidden layer from each hidden layer is shown in the diagram, but those skilled in the art will understand that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers can be included in a CNN 900. As previously described, the output can indicate a single category of an object, or it can include probabilities that best describe the category of an object in an image.

[0148] The first layer of CNN 900 is a convolutional hidden layer 922a. Convolutional hidden layer 922a analyzes the image data from the input layer 920. Each node in convolutional hidden layer 922a is connected to a region of the input image (pixels) called the receptive field. Convolutional hidden layer 922a can be thought of as one or more filters (each filter corresponding to a different activation or feature map), where each convolutional iteration of the filter is a node or neuron in convolutional hidden layer 922a. For example, the region of the input image covered by the filter at each convolutional iteration will be the filter's receptive field. In an illustrative example, if the input image consists of a 28×28 array and each filter (and its corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in convolutional hidden layer 922a. Weights are learned for each connection between a node and its receptive field, and in some cases, an overall bias is learned, allowing each node to learn to analyze its specific local receptive field in the input image. Each node in hidden layer 922a will have the same weights and biases (referred to as shared weights and shared biases). For example, the filter has a weight (digital) array and the same depth as the input. For the video frame example, the filter would have a depth of 3 (based on the three color components of the input image). An illustrative example of the filter array size is 5×5×3, which corresponds to the size of the receptive field of a node.

[0149] The convolutional property of the convolutional hidden layer 922a is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filters of the convolutional hidden layer 922a can start at the top left corner of the input image array and can be convolved around the input image. As described above, each convolutional iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 922a. In each convolutional iteration, the value of the filter is multiplied by the corresponding number of original pixel values ​​of the image (e.g., a 5x5 filter array multiplied by a 5x5 array of input pixel values ​​at the top left corner of the input image array). The multiplications from each convolutional iteration can be summed to obtain the sum of that iteration or node. Next, the process continues at the next position in the input image based on the receptive field of the next node in the convolutional hidden layer 922a. For example, the filter can move a step (called stride) to the next receptive field. The stride can be set to 1 or other suitable amounts. For example, if the stride is set to 1, the filter will move 1 pixel to the right at each convolutional iteration. Processing the filter at each unique location in the input volume produces a number representing the filter result at that location, which leads to determining a sum value for each node of the convolutional hidden layer 922a.

[0150] The mapping from the input layer to the convolutional hidden layer 922a is called an activation map (or feature map). An activation map includes values ​​representing the filter results at each location in the input volume for each node. Activation maps can include arrays containing various sums of values ​​produced by the filter for each iteration of the input volume. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map would consist of a 24×24 array. The convolutional hidden layer 922a can include several activation maps to identify multiple features in the image. Figure 9 The example shown includes three activation maps. Using these three activation maps, the convolutional hidden layer 922a can detect three different types of features, each of which is detectable across the entire image.

[0151] In some examples, nonlinear hidden layers can be applied after convolutional hidden layers 922a. Nonlinear layers can be used to introduce nonlinearity into a system that has already computed linear operations. An illustrative example of a nonlinear layer is the Corrected Linear Unit (ReLU) layer. A ReLU layer applies the function f(x) = max(0,x) to all values ​​in the input volume, which changes all negative activations to 0. Therefore, ReLU can add nonlinearity to CNN 900 without affecting the receptive field of convolutional hidden layers 922a.

[0152] A pooling hidden layer 922b can be applied after the convolutional hidden layer 922a (and, in use, after the non-linear hidden layer). The pooling hidden layer 922b is used to simplify the information in the output of the convolutional hidden layer 922a. For example, the pooling hidden layer 922b can take each activation map from the output of the convolutional hidden layer 922a and use a pooling function to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by the pooling hidden layer. The pooling hidden layer 922a uses other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. Pooling functions (e.g., max pooling filters, L2 norm filters, or other suitable pooling filters) are applied to each activation map included in the convolutional hidden layer 922a. Figure 9 In the example shown, three pooling filters are used for the three activation maps in the convolutional hidden layer 922a.

[0153] In some examples, max pooling can be used by applying a max pooling filter (e.g., of size 2x2) with a stride (e.g., equal to the filter dimension, such as a stride of 2) to the activation map output from the convolutional hidden layer 922a. The output from the max pooling filter includes the maximum number in each sub-region of the filter convolution. Using a 2x2 filter as an example, each unit in the pooling layer can summarize a region of 2×2 nodes from the previous layer (each node being a value in the activation map). For example, four values ​​(nodes) in the activation map will be analyzed by the 2x2 max pooling filter at each iteration of the filter, and the maximum of the four values ​​will be output as the "maximum" value. If such a max pooling filter is applied to the activation filter from a convolutional hidden layer 922a with a dimension of 24x24 nodes, the output from the pooling hidden layer 922b will be an array of 12x12 nodes.

[0154] In some examples, an L2 norm pooling filter can also be used. An L2 norm pooling filter involves calculating the square root of the sum of the squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (instead of calculating the maximum value as done in max pooling), and using the calculated value as the output.

[0155] Intuitively, pooling functions (e.g., max pooling, L2-norm pooling, or other pooling functions) determine whether a given feature is found anywhere within a region of an image, discarding the exact location information. This can be done without affecting the results of feature detection, because once a feature has been found, its exact location is less important than its approximate location relative to other features. Max pooling (and other pooling methods) offers the benefit of pooling far fewer features, thus reducing the number of parameters required in subsequent layers of a CNN 900.

[0156] The final connection in the network is a fully connected layer, which connects each node from the pooling hidden layer 922b to each output node in the output layer 924. Using the example above, the input layer comprises 28x28 nodes for encoding pixel intensities of the input image; the convolutional hidden layer 922a comprises 3x24x24 hidden feature nodes based on applying a 5x5 local receptive field (for filtering) to three activation maps; and the pooling hidden layer 922b comprises a layer of 3x12x12 hidden feature nodes based on applying a max-pooling filter to a 2x2 region across each of the three feature maps. Extending this example, the output layer 924 could comprise ten output nodes. In such an example, each node of the 3x12x12 pooling hidden layer 922b is connected to each node of the output layer 924.

[0157] The fully connected layer 922c takes the output of the previous pooling hidden layer 922b (which should represent the activation map of high-level features) and determines the features most relevant to a particular class. For example, the fully connected layer 922c can determine the high-level features most relevant to a particular class and can include weights (nodes) for those high-level features. The product between the weights of the fully connected layer 922c and the pooling hidden layer 922b can be computed to obtain the probabilities for different classes. For example, if CNN 900 is used to predict that an object in a video frame is a person, there will be high values ​​in the activation map representing the high-level features of a person (e.g., two legs, a face at the top of the object, two eyes at the top left and top right of the face, a nose in the middle of the face, a mouth at the bottom of the face, and / or other features common to people).

[0158] In some examples, the output from output layer 924 may include an M-dimensional vector (M = 10 in the previous example). M indicates the number of classes the CNN 900 must choose from when classifying objects in an image. Other example outputs may also be provided. Each number in the M-dimensional vector can represent the probability that an object belongs to a certain class. In an illustrative example, if the 10-dimensional output vector represents objects in ten different classes as [000.050.800.150000], then the vector indicates a 5% probability that the image is an object in the third class (e.g., a dog), an 80% probability that the image is an object in the fourth class (e.g., a person), and a 15% probability that the image is an object in the sixth class (e.g., a kangaroo). The probability for a class can be thought of as the confidence level that an object is part of that class.

[0159] Figure 10 An example computing device with a computing device architecture 1000 is shown, which incorporates components of a computing device that can be used to perform one or more of the technologies described herein. Figure 10 The computing devices illustrated may be incorporated as part of any computerized system herein. For example, computing device architecture 1000 may represent some components of a mobile device or a computing device that performs a 3D model retrieval system or tool. Examples of computing device architecture 1000 include, but are not limited to, desktop computers, workstations, personal computers, supercomputers, video game consoles, tablet computers, smartphones, laptop computers, netbooks, or other portable devices. Figure 10 A schematic diagram of one embodiment of a computing device having architecture 1000 is provided, which can perform methods provided by various other embodiments as described herein, and / or can be used as a host computing device, a remote self-service terminal / terminal, a point-of-sale device, a mobile multifunction device, a set-top box, and / or a computing device. Figure 10This is intended only to provide a general overview of the various components; any or all of these components may be used at your discretion. Therefore, Figure 10 It roughly illustrates how individual system components can be implemented in a relatively separate or relatively more integrated manner.

[0160] A computing device architecture 1000 is illustrated, comprising hardware components that can be electrically coupled (or otherwise communicated as appropriate) via a bus 1005. The hardware components may include one or more processors 1010, including but not limited to one or more general-purpose processors and / or one or more dedicated processors (e.g., digital signal processing chips, graphics accelerators, etc.); one or more input devices 1015, which may include, but are not limited to, cameras, sensors 1050, mice, keyboards, etc.; and one or more output devices 1020, which may include, but are not limited to, display units, printers, etc.

[0161] The computing device architecture 1000 may further include (and / or communicate with) one or more non-transitory storage devices 1025, which may include, but are not limited to, local and / or network-accessible storage, and / or may include, but are not limited to, disk drives, drive arrays, optical storage devices, solid-state storage devices (such as random access memory (“RAM”) and / or read-only memory (“ROM”)), which may be programmable, flash-updatable, etc. Such storage devices can be configured to implement any suitable data storage, including but not limited to various file systems, database structures, etc.

[0162] The computing device architecture 1000 may also include a communication subsystem 1030. The communication subsystem 1030 may include a transceiver or wired and / or wireless media for receiving and transmitting data. The communication subsystem 1030 may also include, but is not limited to, a modem, a network interface card (NIC) (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (e.g., Bluetooth). TM Devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication facilities, etc. The communication subsystem 1030 may allow the exchange of data with networks (such as those described below, to name just one example), other computing devices, and / or any other devices described herein. In many embodiments, the computing device architecture 1000 will further include a non-transitory working memory 1035, which may include RAM or ROM devices as described above.

[0163] The computing device architecture 1000 may include software elements shown as currently residing within working memory 1035, including an operating system 1040, device drivers, executable libraries, and / or other code (e.g., one or more application code 1045). This code may include computer programs provided by various embodiments and / or may be designed to implement methods and / or configure systems provided by other embodiments, as described herein. By way of example only, one or more processes described with respect to the methods discussed above may be implemented as code and / or instructions executable by a computer (and / or a processor within a computer); in one aspect, this code and / or instructions may be used to configure and / or tune a general-purpose computer (or other device) to perform one or more operations according to the described methods.

[0164] These sets of instructions and / or code may be stored on a computer-readable storage medium, such as the storage device 1025 described above. In some cases, the storage medium may be incorporated within a computing device (e.g., a computing device having computing device architecture 1000). In other embodiments, the storage medium may be separate from the computing device (e.g., a removable medium, such as an optical disc) and / or provided as an installation package, such that the storage medium can be used to program, configure, and / or adapt a general-purpose computer having instructions / code stored thereon. These instructions may take the form of executable code (which is executable by computing device architecture 1000) and / or may take the form of source code and / or installable code (which then takes the form of executable code when compiled and / or installed on computing device architecture 1000 (e.g., using any of a variety of generally available compilers, installers, compression / decompression utilities, etc.)).

[0165] Substantial modifications can be made to suit specific requirements. For example, custom hardware can be used, and / or specific elements can be implemented in hardware, software (including portable software such as applets), or both. Furthermore, connectivity to other computing devices, such as network input / output devices, can be employed.

[0166] Some embodiments may employ a computing device (e.g., a computing device having computing device architecture 1000) to perform the methods according to this disclosure. For example, some or all of the processes of the described methods may be executed by a computing device having computing device architecture 1000 in response to processor 1010 executing one or more sequences of instructions contained in working memory 1035 (which may be incorporated into operating system 1040 and / or other code such as application 1045). These instructions may be read into working memory 1035 from another computer-readable medium (e.g., one or more storage devices 1025 in storage device 1025). By way of example only, executing a sequence of instructions contained in working memory 1035 may cause processor 1010 to perform one or more processes of the methods described herein.

[0167] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media on which data can be stored but does not include carrier waves and / or transient electronic signals propagated wirelessly or via a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as CDs or DVDs, flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means, including memory sharing, messaging, token passing, network transmission, etc.

[0168] In some embodiments, computer-readable storage devices, media, and memories may include cable or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0169] Specific details are provided in the foregoing description to provide a thorough understanding of the embodiments and examples provided herein. However, those skilled in the art will understand that the embodiments can be practiced without these specific details. For clarity, in some instances, the technology may be presented as comprising individual functional blocks, including functional blocks containing devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid obscuring the embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.

[0170] Individual embodiments may be described above as processes or methods, which are depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flowchart can describe operations as a sequential process, many operations within an operation can be performed in parallel or simultaneously. Furthermore, the order of operations can be rearranged. When its operations are completed, the process terminates, but may have additional steps not included in the diagram. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.

[0171] The processes and methods described in the examples above can be implemented using computer-executable instructions that are stored or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. Parts of the computer resources used may be accessible via a network. For example, computer-executable instructions may be binary files, intermediate format instructions (e.g., assembly language), firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information generated during the method according to the described examples include hard disks or optical disks, flash memory, USB devices provided with non-volatile memory, network storage devices, etc.

[0172] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. One (or more) processors may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or intercalation cards. By further example, such functionality may also be implemented on a circuit board between different chips or different processes performed in a single device.

[0173] Instructions, media for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are example units for providing the functionality described in this disclosure.

[0174] In the foregoing description, various aspects of this application have been described with reference to specific embodiments thereof; however, those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative embodiments of this application have been described in detail herein, it should be understood that these inventive concepts may be embodied and employed differently in other ways, and the appended claims are intended to be construed as including such variations beyond those limited by the prior art. Various features and aspects of the applications described above may be used individually or in combination. Furthermore, without departing from the broader spirit and scope of this specification, the embodiments may be used in any number of environments and applications beyond those described herein. Therefore, the specification and drawings are to be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that in alternative embodiments, these methods may be performed in a different order than that described.

[0175] Those skilled in the art will understand that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively.

[0176] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0177] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to said other component via a wired or wireless connection and / or other suitable communication interface).

[0178] The claim language, or other language, that states "at least one of" and / or "one or more" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, the claim language that states "one or more of A or B" means A, B, or A and B. In another example, the claim language that states "one or more of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language that states "at least one of" and / or "one or more" in a set does not limit the set to the items listed in the set. For example, the claim language that states "at least one of A or B" can mean A, B, or A and B, and can additionally include items not listed in the set of A and B.

[0179] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above in general terms of their functionality. Whether this functionality is implemented in hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0180] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication handheld devices, or integrated circuit devices with multiple uses (including applications in wireless communication handheld devices and other devices). Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product that may include packaging material. The computer-readable medium can include memory or data storage media, such as random access memory (RAM) (e.g., synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, etc. Alternatively or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or transmits program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.

[0181] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuit systems. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessor and DSP cores, or any other such configuration. Therefore, as used herein, the term "processor" can refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.

[0182] The illustrative aspects of this disclosure include:

[0183] Aspect 1: A method for determining one or more poses of one or more objects, the method comprising: determining a plurality of keypoints from an image using a machine learning system, the plurality of keypoints being associated with at least one object in the image; determining a plurality of features from the machine learning system based on the plurality of keypoints; classifying the plurality of features into a plurality of joint types; and determining pose parameters for the at least one object based on the plurality of joint types.

[0184] Aspect 2: According to the method of aspect 1, wherein the at least one object includes at least one hand.

[0185] Aspect 3: The method according to either aspect 1 or 2, wherein the at least one object comprises two objects, wherein the plurality of keypoints comprises keypoints for the two objects, and wherein the pose parameters comprise pose parameters for the two objects.

[0186] Aspect 4: The method according to either aspect 1 or 2, wherein the at least one object includes two hands, wherein the plurality of keypoints includes keypoints for the two hands, and wherein the pose parameters include pose parameters for the two hands.

[0187] Aspect 5: The method according to either aspect 1 or 2, wherein the at least one object includes a hand, and further includes: using the machine learning system to determine a plurality of object keypoints from the image, the plurality of object keypoints being associated with an object associated with the hand; and determining pose parameters for the object based on the plurality of object keypoints.

[0188] Aspect 6: The method according to any one of aspects 1 to 5, wherein each of the plurality of key points corresponds to a joint of the at least one object.

[0189] Aspect 7: The method according to any one of Aspects 1 to 6, wherein determining the plurality of features from the machine learning system based on the plurality of key points comprises: determining a first feature set corresponding to the plurality of key points from a first feature map of the machine learning system, the first feature map including a first resolution; and determining a second feature set corresponding to the plurality of key points from a second feature map of the machine learning system, the second feature map including a second resolution.

[0190] Aspect 8: The method according to any one of aspects 1 to 7 further includes: generating a feature representation for each of the plurality of key points, wherein the plurality of features are classified into the plurality of joint types using the feature representation for each key point.

[0191] Aspect 9: According to the method of aspect 8, wherein the feature representation for each keypoint includes an encoding vector.

[0192] Aspect 10: The method according to any one of aspects 1 to 9, wherein the machine learning system includes a neural network that uses the image as input.

[0193] Aspect 11: The method according to any one of aspects 1 to 10, wherein the plurality of features are classified into the plurality of joint types by an encoder of a transformer neural network, and wherein the pose parameters determined for the at least one object are determined by a decoder of the transformer neural network based on the plurality of joint types.

[0194] Aspect 12: The method according to any one of aspects 1 to 11, wherein the posture parameters are determined for the at least one object based on the plurality of joint types and based on one or more learned joint queries.

[0195] Aspect 13: The method according to aspect 12, wherein the at least one object includes a first object and a second object, and wherein the one or more learned joint queries are used to predict at least one of the relative translation between the first object and the second object, a set of object shape parameters, and camera model parameters.

[0196] Aspect 14: The method according to any one of aspects 1 to 13, wherein the pose parameters for the at least one object include a three-dimensional vector for each of the plurality of joint types.

[0197] Aspect 15: According to the method of aspect 14, wherein the three-dimensional vector for each of the plurality of joint types includes a horizontal component, a vertical component, and a depth component.

[0198] Aspect 16: The method according to aspect 14, wherein the three-dimensional vector for each of the plurality of joint types includes a vector between each joint and a parent joint associated with each joint.

[0199] Aspect 17: The method according to any one of aspects 1 to 13, wherein the pose parameters for the at least one object include the position of each joint and the difference between the depth of each joint and the depth of the parent joint associated with each joint.

[0200] Aspect 18: The method according to any one of aspects 1 to 17, wherein the pose parameters for the at least one object include a translation of the at least one object relative to another object in the image.

[0201] Aspect 19: The method according to any one of aspects 1 to 18, wherein the pose parameter for the at least one object includes the shape of the at least one object.

[0202] Aspect 20: The method according to any one of aspects 1 to 19 further includes: determining user input based on the attitude parameters.

[0203] Aspect 21: The method according to any one of aspects 1 to 20 further includes: rendering virtual content based on the pose parameters.

[0204] Aspect 22: The method according to any one of aspects 1 to 21, wherein the apparatus is an augmented reality device (e.g., a head-mounted display, augmented reality glasses or other augmented reality device).

[0205] Aspect 23: An apparatus for determining one or more poses of one or more objects, comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: determine a plurality of keypoints from an image using a machine learning system, the plurality of keypoints being associated with at least one object in the image; determine a plurality of features from the machine learning system based on the plurality of keypoints; classify the plurality of features into a plurality of joint types; and determine pose parameters for the at least one object based on the plurality of joint types.

[0206] Aspect 24: The apparatus according to aspect 23, wherein the at least one object comprises at least one hand.

[0207] Aspect 25: The apparatus according to any one of aspects 23 or 24, wherein the at least one object comprises two objects, wherein the plurality of key points comprises key points for the two objects, and wherein the attitude parameters comprise attitude parameters for the two objects.

[0208] Aspect 26: The device according to any one of aspects 23 or 24, wherein the at least one object includes two hands, wherein the plurality of keypoints includes keypoints for the two hands, and wherein the posture parameters include posture parameters for the two hands.

[0209] Aspect 27: The apparatus according to any one of aspects 23 or 24, wherein the at least one object comprises a single hand, and wherein the at least one processor is configured to: use the machine learning system to determine a plurality of object keypoints from the image, the plurality of object keypoints being associated with an object associated with the single hand; and determine pose parameters for the object based on the plurality of object keypoints.

[0210] Aspect 28: The apparatus according to any one of aspects 23 to 27, wherein each of the plurality of key points corresponds to a joint of the at least one object.

[0211] Aspect 29: The apparatus according to any one of aspects 23 to 28, wherein, in order to determine the plurality of features from the machine learning system based on the plurality of key points, the at least one processor is configured to: determine a first feature set corresponding to the plurality of key points from a first feature map of the machine learning system, the first feature map including a first resolution; and determine a second feature set corresponding to the plurality of key points from a second feature map of the machine learning system, the second feature map including a second resolution.

[0212] Aspect 30: The apparatus according to any one of aspects 23 to 29, wherein the at least one processor is configured to: generate a feature representation for each of the plurality of keypoints, wherein the plurality of features are classified into the plurality of joint types using the feature representation for each keypoint.

[0213] Aspect 31: The apparatus according to aspect 23, wherein the feature representation for each key point includes an encoding vector.

[0214] Aspect 32: The apparatus according to any one of aspects 23 to 31, wherein the machine learning system includes a neural network that uses the image as input.

[0215] Aspect 33: The apparatus according to any one of aspects 23 to 32, wherein the plurality of features are classified into the plurality of joint types by an encoder of a transformer neural network, and wherein the pose parameters determined for the at least one object are determined by a decoder of the transformer neural network based on the plurality of joint types.

[0216] Aspect 34: The apparatus according to any one of aspects 23 to 33, wherein the attitude parameters are determined for the at least one object based on the plurality of joint types and based on one or more learned joint queries.

[0217] Aspect 35: The apparatus according to aspect 34, wherein the at least one object comprises a first object and a second object, and wherein the one or more learned joint queries are used to predict at least one of the relative translation between the first object and the second object, a set of object shape parameters, and camera model parameters.

[0218] Aspect 36: The apparatus according to any one of aspects 23 to 35, wherein the posture parameters for the at least one object include a three-dimensional vector for each of the plurality of joint types.

[0219] Aspect 37: The apparatus according to aspect 36, wherein the three-dimensional vector for each of the plurality of joint types includes a horizontal component, a vertical component, and a depth component.

[0220] Aspect 38: The apparatus according to aspect 36, wherein the three-dimensional vector for each of the plurality of joint types includes a vector between each joint and a parent joint associated with each joint.

[0221] Aspect 39: The apparatus according to any one of aspects 23 to 35, wherein the attitude parameters for the at least one object include the position of each joint and the difference between the depth of each joint and the depth of the parent joint associated with each joint.

[0222] Aspect 40: The apparatus according to any one of aspects 23 to 39, wherein the pose parameters for the at least one object include a translation of the at least one object relative to another object in the image.

[0223] Aspect 41: The apparatus according to any one of aspects 23 to 40, wherein the attitude parameters for the at least one object include the shape of the at least one object.

[0224] Aspect 42: The apparatus according to any one of aspects 23 to 41, wherein the at least one processor is configured to determine user input based on the attitude parameters.

[0225] Aspect 43: The apparatus according to any one of aspects 23 to 42, wherein the at least one processor is configured to render virtual content based on the pose parameters.

[0226] Aspect 44: The apparatus according to any one of aspects 23 to 43, wherein the apparatus includes a mobile device.

[0227] Aspect 45: The apparatus according to any one of aspects 23 to 44, wherein the apparatus includes an augmented reality device (e.g., a head-mounted display, augmented reality glasses or other augmented reality device).

[0228] Aspect 46: The apparatus according to any one of aspects 23 to 45, wherein the at least one processor includes a neural processing unit (NPU).

[0229] Aspect 47: The apparatus according to any one of aspects 23 to 46 further includes: a display configured to display one or more images.

[0230] Aspect 48: The apparatus according to any one of aspects 23 to 47 further includes: an image sensor configured to capture one or more images.

[0231] Aspect 49: A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the operations described in any one of aspects 1 to 48.

[0232] Aspect 50: An apparatus for determining one or more poses of one or more objects, the apparatus comprising one or more units for performing operations according to any one of aspects 1 to 48.

Claims

1. An apparatus for determining one or more orientations of one or more objects, comprising: At least one memory; as well as At least one processor coupled to the at least one memory, the at least one processor being configured to: A machine learning system is used to identify multiple key points from an image, the multiple key points being associated with at least one object in the image; Based on the aforementioned key points, multiple features are determined from the machine learning system; The multiple features are classified into multiple joint types; as well as Pose parameters for the at least one object are determined based on the plurality of joint types, wherein the pose parameters are determined for the at least one object based on the plurality of joint types and based on one or more learned joint queries, and wherein the pose parameters are defined using a plurality of pose representations, and the learned joint queries correspond to joint type embeddings.

2. The apparatus according to claim 1, wherein, The at least one object includes at least one hand.

3. The apparatus according to claim 1, wherein, The at least one object includes two objects, wherein the plurality of keypoints includes keypoints for the two objects, and wherein the pose parameters include pose parameters for the two objects.

4. The apparatus according to claim 1, wherein, The at least one object includes two hands, wherein the plurality of keypoints includes keypoints for the two hands, and wherein the pose parameters include pose parameters for the two hands.

5. The apparatus according to claim 1, wherein, The at least one object includes a single hand, and wherein the at least one processor is configured to: The machine learning system is used to determine multiple object keypoints from the image, the multiple object keypoints being associated with objects related to the hand; and The pose parameters for the object are determined based on the multiple object key points.

6. The apparatus according to claim 1, wherein, Each of the plurality of key points corresponds to a joint of the at least one object.

7. The apparatus according to claim 1, wherein, In order to determine the multiple features from the machine learning system based on the multiple key points, the at least one processor is configured to: Determine a first feature set corresponding to the plurality of key points from a first feature map of the machine learning system, wherein the first feature map includes a first resolution; and A second feature set corresponding to the plurality of key points is determined from the second feature map of the machine learning system, the second feature map including a second resolution.

8. The apparatus according to claim 1, wherein, The at least one processor is configured to: A feature representation is generated for each of the plurality of key points, wherein the plurality of features are classified into the plurality of joint types using the feature representation for each key point.

9. The apparatus according to claim 8, wherein, The feature representation for each keypoint includes an encoding vector.

10. The apparatus according to claim 1, wherein, The machine learning system includes a neural network that uses the image as input.

11. The apparatus according to claim 1, wherein, The plurality of features are classified into the plurality of joint types by the encoder of the transformer neural network, and wherein the pose parameters determined for the at least one object are determined by the decoder of the transformer neural network based on the plurality of joint types.

12. The apparatus according to claim 1, wherein, The at least one object includes a first object and a second object, and wherein the one or more learned joint queries are used to predict at least one of the following: relative translation between the first object and the second object, a set of object shape parameters, and camera model parameters.

13. The apparatus according to claim 1, wherein, The pose parameters for the at least one object include a three-dimensional vector for each of the plurality of joint types.

14. The apparatus according to claim 13, wherein, The three-dimensional vector for each of the plurality of joint types includes a horizontal component, a vertical component, and a depth component.

15. The apparatus according to claim 13, wherein, The three-dimensional vector for each of the plurality of joint types includes a vector between each joint and the parent joint associated with each joint.

16. The apparatus according to claim 1, wherein, The pose parameters for the at least one object include the position of each joint and the difference between the depth of each joint and the depth of the parent joint associated with each joint.

17. The apparatus according to claim 1, wherein, The pose parameters for the at least one object include the translation of the at least one object relative to another object in the image.

18. The apparatus according to claim 1, wherein, The pose parameters for the at least one object include the shape of the at least one object.

19. The apparatus according to claim 1, wherein, The at least one processor is configured to determine user input based on the attitude parameters.

20. The apparatus according to claim 1, wherein, The at least one processor is configured to render virtual content based on the pose parameters.

21. The apparatus according to claim 1, wherein, The device is an augmented reality device.

22. A method for determining one or more poses of one or more objects, the method comprising: A machine learning system is used to identify multiple key points from an image, the multiple key points being associated with at least one object in the image; Multiple features are determined from the machine learning system based on the aforementioned key points; The multiple features are classified into multiple joint types; as well as Pose parameters for the at least one object are determined based on the plurality of joint types, wherein the pose parameters are determined for the at least one object based on the plurality of joint types and based on one or more learned joint queries, and wherein the pose parameters are defined using a plurality of pose representations, and the learned joint queries correspond to joint type embeddings.

23. The method according to claim 22, wherein, The at least one object includes at least one hand.

24. The method according to claim 22, wherein, The at least one object includes two hands, wherein the plurality of keypoints includes keypoints for the two hands, and wherein the pose parameters include pose parameters for the two hands.

25. The method according to claim 22, wherein, The at least one object includes a single hand, and the method further includes: The machine learning system is used to determine multiple object keypoints from the image, the multiple object keypoints being associated with objects related to the hand; and The pose parameters for the object are determined based on the multiple object key points.

26. The method according to claim 22, wherein, Determining the multiple features from the machine learning system based on the multiple key points includes: Determine a first feature set corresponding to the plurality of key points from a first feature map of the machine learning system, wherein the first feature map includes a first resolution; and A second feature set corresponding to the plurality of key points is determined from the second feature map of the machine learning system, the second feature map including a second resolution.

27. The method of claim 22, further comprising: A feature representation is generated for each of the plurality of key points, wherein the plurality of features are classified into the plurality of joint types using the feature representation for each key point.

28. The method according to claim 22, wherein, The plurality of features are classified into the plurality of joint types by the encoder of the transformer neural network, and wherein the pose parameters determined for the at least one object are determined by the decoder of the transformer neural network based on the plurality of joint types.