Annotation network for three-dimensional pose estimation
An annotation network addresses the challenge of generating high-quality 3D pose labels by automatically estimating and refining 3D poses, enhancing the robustness and efficiency of 3D motion tracking applications.
Patent Information
- Application Number
- PCT/CN2023/138071
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-19
AI Technical Summary
Current deep learning-based methods for 3D pose estimation face challenges in generating high-quality ground-truth 3D pose labels, which are time-consuming and difficult to annotate, especially from single input images.
An annotation network is employed to automatically estimate and refine 3D poses with high accuracy and robustness, utilizing a regression module to output latent representations, orientation, and camera parameters, and a decoder to convert these into pose parameters, with a kinematic layer estimating the 3D pose and a projection layer optimizing the graphical representation.
The proposed solution enables efficient labeling of 3D poses from images, improving robustness and practicability for 3D motion tracking applications by automatically producing accurate 3D pose estimations that can be refined through user edits and retraining of the annotation network.
Smart Images

Figure CN2023138071_19062025_PF_FP_ABST
Abstract
Description
ANNOTATION NETWORK FOR THREE-DIMENSIONAL POSE ESTIMATIONTechnical Field
[0001] This disclosure relates generally to neural networks (also referred to as “deep neural networks” or “DNNs” ) , and more specifically, to annotation network for three-dimensional (3D) pose estimation.Background
[0002] The last decade has witnessed a rapid rise in artificial intelligence (AI) based data processing, particularly based on DNNs. DNNs are widely used in the domains of image recognition, video understanding, image or video generation, machine translation, mathematical reasoning, and so on. For instance, deep learning based generative models can produce content of various kinds, including images, sounds, and texts. Variational autoencoders (VAEs) are a type of deep generative models. A VAE typically includes two neural networks that are referred to as the encoder and decoder, respectively.Brief Description of the Drawings
[0003] Embodiments will be readily understood by the following detailed description in conjunction with the accompanying drawings. To facilitate this description, like reference numerals designate like structural elements. Embodiments are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings.
[0004] FIG. 1 is a block diagram of a pose estimation system, in accordance with various embodiments.
[0005] FIG. 2 is a block diagram of an AI annotation module, in accordance with various embodiments.
[0006] FIG. 3 illustrates an example pose estimation process, in accordance with various embodiments.
[0007] FIG. 4 illustrates an example annotation network, in accordance with various embodiments.
[0008] FIG. 5 illustrates a 2D annotation file 510 generated from an image, in accordance with various embodiments.
[0009] FIG. 6 illustrates an example 3D pose graphical representation, in accordance with various embodiments.
[0010] FIG. 7 illustrates an optimized 3D pose graphical representation, in accordance with various embodiments.
[0011] FIG. 8 illustrates an editable 3D skeleton, in accordance with various embodiments.
[0012] FIG. 9 illustrates an example DNN, in accordance with various embodiments.
[0013] FIG. 10 illustrates an AI-based 3D pose estimation environment, in accordance with various embodiments.
[0014] FIG. 11 is a flowchart showing a method of pose estimation, in accordance with various embodiments.
[0015] FIG. 12 is a block diagram of an example computing device, in accordance with various embodiments.Detailed Description
[0016] Overview
[0017] A DNN typically includes a sequence of layers. A DNN layer may include one or more deep learning operations (also referred to as “neural network operations” ) , such as convolution, pooling, elementwise operation, linear operation, nonlinear operation, and so on. A DL operation in a DNN may be performed on one or more internal parameters of the DNNs (e.g., weights) , which are determined during the training phase, and one or more activations. An activation may be a data point (also referred to as “data elements” or “elements” ) . Activations or weights of a DNN layer may be elements of a tensor of the DNN layer. A tensor is a data structure having multiple elements across one or more dimensions. Example tensors include a vector, which is a one-dimensional tensor, and a matrix, which is a two-dimensional tensor. There can also be three-dimensional tensors and even higher dimensional tensors. A DNN layer may have an input tensor (also referred to as “input feature map (IFM) ” ) including one or more input activations (also referred to as “input elements” ) and a weight tensor including one or more weights. A weight is an element in the weight tensor. A weight tensor of a convolution may be a kernel, a filter, or a group of filters. The output data of the DNN layer may be an output tensor (also referred to as “output feature map (OFM) ” ) that includes one or more output activations (also referred to as “output elements” ) .
[0018] 3D pose estimation aims to regress 3D pose parameters from images. 3D pose estimation can be widely used in many applications, including augmented reality, virtual reality, sports analysis, telepresence, film and game production, action recognition, and so on. DNN models are used to estimate 3D pose from a single RGB (red, green, blue) image. However, when it comes to the technology landing on specific scenarios, the performance of many DNN models is tied to the quality and quantity of annotated 3D poses. Sometimes building a proper training dataset can be more important than the model selection. Given a collection of real images, high-quality ground-truth 3D pose labels (e.g., 3D joint positions, rotations, etc. ) for each image need to be generated for training an estimation network model for monocular 3D pose tracking applications. Compared with other tasks including detection, segmentation, and 2D pose estimation, for which it is relatively easy to achieve ground-truth labels by manual annotation for the given image, 3D pose annotation can be more difficult and time-consuming, especially in the case of a single input image. Some learning based algorithms can automatically generate pseudo-ground-truth 3D pose data for real-world images. However, the generated poses usually lack the quality for being used as ground-truth labels.
[0019] Embodiments of the present disclosure may improve on at least some of the challenges and issues described above by using an annotation network to automatically estimate and refine 3D poses with high accuracy and robustness. A 3D graphic interface may also be provided to facilitate users to view and edit the estimated 3D poses.
[0020] In an example embodiment of the present disclosure, an annotation network may be used to estimate 3D pose of an object using an image of the object. The image may be cropped to generate an input image. The input image is input into the annotation network. A regression module in the annotation network may output a latent representation of the input image in a latent space, an orientation of an object, and one or more camera parameters indicating one or more settings of a camera capturing the input image. A decoder in the annotation network may convert the latent representation of the input image into one or more pose parameters indicating a pose of the object in an image space. The image space may be two dimensional. The latent space may be one dimensional. A kinematic layer in the neural network may use a forward kinematic function to estimate a 3D pose of the object using the orientation of the object and the one or more pose parameters. A graphical representation of the estimated 3D pose of the object may be generated. A projection layer of the annotation network may project points in the graphical representation of the estimated 3D pose from the camera space into the image space. The camera space may be three dimensional. The graphical representation of the estimated 3D pose may be optimized based on the projected points. For instance, key points of the object may be identified from the image. A distance between the projected points and the key points may be measured. The graphical representation of the estimated 3D pose may be optimized using the measured distance. Also, the latent representation of the input image to generate a modified latent representation of the input image. A similarity between the latent representation of the input image and the modified latent representation of the input image may be measured and used to optimize the graphical representation of the estimated 3D pose. The optimized graphical representation of the estimated 3D pose may be provided for display in an interactive visualization interface where a user can view and modify one or more points in the graphical representation of the estimated 3D pose. The interactive visualization interface interface may present the graphical representation to the user in various views. The user’s modification can be used to further train the annotation neural network. The user’s modification can be real-time responses.
[0021] The present disclosure provides a more efficient approach for labeling 3D poses from images, including in-the-wild images. This approach can improve robustness and practicability for various 3D motion tracking applications. Given the input dataset consisting of the images and corresponding 2D pose annotations, 3D pose estimation can be automatically produced. The annotation network may be weakly supervised with 2D ground-truth labels without 3D supervision and may be tested on the same input dataset to predict 3D poses. To reach a fast and better training performance, the annotation network can be initialized by a pretrained pose estimation network. This data-driven learning design can enable the annotation network to better learn the coherent characteristics of the input dataset and generate far more accurate 3D poses for input images. The estimated 3D poses can be refined, and the refinement can be guided by a pose priori model to further reduce misalignment errors while preventing anatomically implausible articulations. The automatically estimated and refined 3D poses can be sufficiently accurate for many applications. In some cases, 3D poses may be further optimized based on user edits. User edits can be used to retrain the annotation network and improve the accuracy of the annotation network and therefore, the likelihood that user edits will be needed for future pose estimation tasks can be reduced.
[0022] For purposes of explanation, specific numbers, materials and configurations are set forth in order to provide a thorough understanding of the illustrative implementations. However, it will be apparent to one skilled in the art that the present disclosure may be practiced without the specific details or / and that the present disclosure may be practiced with only some of the described aspects. In other instances, well known features are omitted or simplified in order not to obscure the illustrative implementations.
[0023] Further, references are made to the accompanying drawings that form a part hereof, and in which is shown, by way of illustration, embodiments that may be practiced. It is to be understood that other embodiments may be utilized, and structural or logical changes may be made without departing from the scope of the present disclosure. Therefore, the following detailed description is not to be taken in a limiting sense.
[0024] Various operations may be described as multiple discrete actions or operations in turn, in a manner that is most helpful in understanding the claimed subject matter. However, the order of description should not be construed as to imply that these operations are necessarily order dependent. In particular, these operations may not be performed in the order of presentation. Operations described may be performed in a different order from the described embodiment. Various additional operations may be performed or described operations may be omitted in additional embodiments.
[0025] For the purposes of the present disclosure, the phrase “A or B” or the phrase "Aand / or B" means (A) , (B) , or (A and B) . For the purposes of the present disclosure, the phrase “A, B, or C” or the phrase "A, B, and / or C" means (A) , (B) , (C) , (A and B) , (A and C) , (B and C) , or (A, B, and C) . The term "between, " when used with reference to measurement ranges, is inclusive of the ends of the measurement ranges.
[0026] The description uses the phrases "in an embodiment" or "in embodiments, " which may each refer to one or more of the same or different embodiments. The terms "comprising, " "including, " "having, " and the like, as used with respect to embodiments of the present disclosure, are synonymous. The disclosure may use perspective-based descriptions such as "above, " "below, " "top, " "bottom, " and "side" to explain various features of the drawings, but these terms are simply for ease of discussion, and do not imply a desired or required orientation. The accompanying drawings are not necessarily drawn to scale. Unless otherwise specified, the use of the ordinal adjectives “first, ” “second, ” and “third, ” etc., to describe a common object, merely indicates that different instances of like objects are being referred to and are not intended to imply that the objects so described must be in a given sequence, either temporally, spatially, in ranking or in any other manner.
[0027] In the following detailed description, various aspects of the illustrative implementations will be described using terms commonly employed by those skilled in the art to convey the substance of their work to others skilled in the art.
[0028] The terms “substantially, ” “close, ” “approximately, ” “near, ” and “about, ” generally refer to being within + / -20%of a target value as described herein or as known in the art. Similarly, terms indicating orientation of various elements, e.g., “coplanar, ” “perpendicular, ” “orthogonal, ” “parallel, ” or any other angle between the elements, generally refer to being within + / -5-20%of a target value as described herein or as known in the art.
[0029] In addition, the terms “comprise, ” “comprising, ” “include, ” “including, ” “have, ” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a method, process, device, or DNN accelerator that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such method, process, device, or DNN accelerators. Also, the term “or” refers to an inclusive “or” and not to an exclusive “or. ”
[0030] The systems, methods and devices of this disclosure each have several innovative aspects, no single one of which is solely responsible for all desirable attributes disclosed herein. Details of one or more implementations of the subject matter described in this specification are set forth in the description below and the accompanying drawings.
[0031] Example Pose Estimation System
[0032] FIG. 1 is a block diagram of a pose estimation system 100, in accordance with various embodiments. The pose estimation system 100 uses a deep learning approach to estimate 3D poses of objects from images of the objects. The pose estimation system 100 includes an interface module 110, a preprocessing module 120, a point labeling module 130, an AI annotation module 140, an optimization module 150, a user editing module 160, an output module 170, and a datastore 180. In other embodiments, alternative configurations, different or additional components may be included in the pose estimation system 100. Further, functionality attributed to a component of the pose estimation system 100 may be accomplished by a different component included in the pose estimation system 100 or by a different module.
[0033] The interface module 110 facilitates communications of the pose estimation system 100 with other systems, devices, or modules. For example, the interface module 110 may receive images from an online system (e.g., a social media system, an online image gallery, an online search tool, etc. ) , a camera, and so on. As another example, the interface module 110 may receive one or more data for training or testing the annotation network. Yet another example, the interface module 110 may transmit data generated by the pose estimation system 100 to other systems, devices, or modules. For instance, the interface module 110 may transmit estimated 3D poses of objects, animations of objects, motion analysis results, or other types of motion tracking information to other systems, devices, or modules.
[0034] The preprocessing module 120 generates images that will be input into the annotation network. In some embodiments, the preprocessing module 120 may receive an image captured by a camera. The image may be a 2D image in a 2D image space. The 2D image space is also referred to as an image plane. The preprocessing module 120 may extract an input image from the received image. The input image may include an object, the 3D pose of which is to be estimated. The preprocessing module 120 may crop the received image to generate the input image. For instance, the preprocessing module 120 may remove one or more other objects or a background from the receive image.
[0035] In some embodiments, the preprocessing module 120 may generate the input image based on detection of the object. To detect a target object, the preprocessing module 120 may determine or receive a class of the target object. Examples of the class of the target object may include person, car, truck, robot, dog, cat, and so on. The preprocessing module 120 may also classify one or more objects in the received image and determine whether any of the objects fall into the class of the target object. The preprocessing module 120 may use a trained model to classify objects. For instance, the preprocessing module 120 may input the received image into the trained model, and the trained model may output the classes of the objects in the received image. The trained model may be a DNN, such as a CNN.
[0036] After the preprocessing module 120 detects the object, the preprocessing module 120 may generate a bounding box based on the detection of the object. The bounding box may be a 2D bounding box that surrounds the object in the received image. The preprocessing module 120 may use the bounding box to extract the input image from the received image.
[0037] The point labeling module 130 performs 2D annotation on 2D images. For instance, the point labeling module 130 may use a 2D image of an object to label one or more key points of the object. The key points may be 2D key points in the image space. In an example, the point labeling module 130 may input the image into a trained model, and the trained model may output labeled key points of the object. The trained model may be a trained DNN. In another example, the point labeling module 130 may identify and label the key points from the image. The point labeling module 130 may also determine connections between the labeled key points. For example, the point labeling module 130 may link two or more key points that are physically connected in the object. As another example, the point labeling module 130 may link two or more key points that are located closely to each other in the object. As yet another example, the point labeling module 130 may link two or more key points that are functionally related, e.g., the key points are associated with the same function of the object.
[0038] In some embodiments, the point labeling module 130 may generate a 2D annotation file. For instance, the point labeling module 130 may save information of the labeled key points in the 2D annotation file. The 2D annotation file may include a graph showing the key points. Additionally or alternatively, the 2D annotation file may include text describing the key points, such as locations, sizes, numbers, or other attributes of the key points. The point labeling module 130 may facilitate various file formats, such as JavaScript Object Notation (JSON) , TXT, and so on. An example of the 2D annotation file is shown in FIG. 5. The image and the 2D annotation file may be used by the AI annotation module 140 for estimating 3D poses. In some embodiments, the point labeling module 130 may combine the original image of the object with the 2D annotation file to generate a single file that includes the original image and labels of the key points. The original image or the combined image may be used as an input image by the AI annotation module 140.
[0039] The AI annotation module 140 uses the annotation network to estimate 3D poses of objects. In some embodiments, the AI annotation module 140 may generate an annotation network. The AI annotation module 140 may define the architecture of the annotation network. For instance, the AI annotation module 140 may determine what types of layers are included in the annotation network, positions of layers in the annotation network, data flows between layers, and so on. The AI annotation module 140 may also train and test the annotation network. After the annotation network is trained, the AI annotation module 140 may facilitate inference of the annotation network for estimating 3D poses. In some embodiments, to estimate a 3D pose of an object, the AI annotation module 140 may provide the input image to the annotation network. The input image may be processed in the layers of the annotation network to generate the output of the annotation network.
[0040] The AI annotation module 140 may obtain the output of the annotation network, which includes information indicating the estimated 3D pose. The output of the annotation network may include a 3D pose graphical representation of the object that shows the estimated 3D pose of the object. Examples of the 3D pose graphical representation may include 3D skeleton of the object, 3D projection, 3D mesh, other types of 3D graphical representations, or some combination thereof. The output of the annotation network may also include an orientation of the object. The orientation may be referred to as a global orientation, as it may be the orientation of the object as a whole. The output of the annotation network may further include one or more camera parameters. The camera parameters may include camera intrinsic parameters, such as parameters indicating estimated optical center, focal length, or other attributes or settings of the camera at the time the camera captures the input image. The AI annotation module 140 may provide the global orientation and camera parameters to the optimization module 150 for optimizing the estimated 3D pose. Certain aspects of the AI annotation moule 140 are described below in conjunction with FIG. 2.
[0041] The optimization module 150 optimizes 3D pose estimation results (e.g., 3D pose graphical representations) obtained by the AI annotation module 140. In some embodiments, the optimization module 150 may use an objection function to refine estimated 3D poses to achieve a better alignment between the 3D pose graphical representation obtained by the AI annotation module 140 and the 2D input image. In an example, the optimization module may select a root point from the key points determined by the preprocessing module 120 and determine the motion of a root point. For instance, the optimization module 150 may compute the root point’s motion dinit in a 3D camera space by minimizing the objective function defined as: E (dinit) =‖Π (P) -K‖2,
[0042] where P=Preg+dinit is the 3D point positions in the 3D camera space, K represents the 2D key points determined by the preprocessing module 120, and Preg represents a regressed 3D skeleton, which may be generated by the annotation network. Π is a pinhole projection function from the 3D camera space to the 2D image space ( “image plane” ) using the camera intrinsic parameters (fx, fy, cx, cy) . The camera intrinsic parameters may be camera parameters determined by the annotation network using the input image. In an example, fx, fy are the focal lengths in x, y directions, cx, cy are the x and y coordinates of the optical center in the image space. In an embodiment where the image of the object has a resolution (w, h) , fx=fy=max (w, h) .
[0043] In some embodiments, the optimization module 150 may optimize the initialization dinit, θinit and determine the optimal d, θ by minimizing another objective function. For instance, the optimization module 150 may employ pose priori in the optimization, rather than directly optimize θ. The optimization module 150 may optimize low-dimensional latent code z and transform this back into point angles θ in an axis-angle representation. The optimization module 150 may compute the minimal reconstruction loss based on the following equations: E (z, d) =Eproj (z, d) +Ereg (z, d) , Eproj (z, d) =λ1‖Π (P) -K‖2, and Ereg (z, d) =λ2 (‖z-zinit‖2+‖d-dinit‖2) ,
[0044] where Eproj measures the distance between the projection of 3D point positions and the labeled 2D key points; and Ereg measures the similarity between the optimized parameters and the initial ones.
[0045] In some embodiments, the optimization module 150 may output optimized 3D pose graphical representations. An optimized 3D pose graphical representation represents the optimized 3D pose estimation. Compared with the 3D pose estimated by the AI annotation module 140, the optimized 3D pose may have a smaller difference from the real pose of the object. The optimized 3D pose estimations may be used as ground-truth labels for the AI annotation module 140 to further train the annotation network.
[0046] The user editing module 160 provides a graphical interface for users to view and edit optimized 3D pose graphical representations. The graphical interface may present optimized 3D pose graphical representations in various views. For instance, the graphical interface may support front view, side view, back view, and so on. The graphical interface may allow the user to change the view. The user may be able to use a keyboard shortcut key or the mouse to change the view. For instance, the user can rotate, move, zoom in, or zoom out the optimized 3D pose graphical representation. In embodiments where the user is not satisfied with the optimized 3D pose graphical representation, the user can select the optimized 3D pose graphical representation (e.g., the user can select the 3D mesh) and start an editing mode of the graphical interface.
[0047] In some embodiments, before the user editing module 160 provides a 3D pose graphical representation for display to the user, the user editing module 160 may evaluate whether the user’s editing would be needed. For instance, the user editing module 160 may evaluate the accuracy of the optimized 3D pose graphical representation. The user editing module 160 may determine an accuracy score that indicates how accurate the optimized 3D pose graphical representation is. The user editing module 160 may determine the accuracy score by measuring a difference between the 3D pose shown in the 3D pose graphical representation and the pose shown in the input image. When the accuracy score is below a target accuracy, the user editing module 160 may provide the 3D pose graphical representation to the user for viewing and editing. When the difference meets a target accuracy, the user editing module 160 may choose not to provide the 3D pose graphical representation to the user for viewing and editing.
[0048] The graphical interface may be interactive in the editing mode. In some embodiments, the graphical interface may include one or more interactive elements, such as buttons, dropdown lists, editable elements, and so on. In some embodiments, the user editing module 160 may place editable points on a 3D pose graphical representation. The graphical interface may allow the user to use a client device to select and edit one or more of the editable points to adjust the 3D pose. For instance, the user may move one or more editable points or change the connection between editable points through one or more mouse or keyboard operations with the client device. The user editing module 160 may receive the edited 3D pose graphical representation from the client device. In some embodiments, after users’ editing, the user editing module 160 may evaluate the edited 3D pose graphical representation to make sure it is natural and realistic.
[0049] In some embodiments, the user editing module 160 may provide the edited 3D pose graphical representation to the optimization module 150 for further refining the edited 3D pose graphical representation. This can ensure that even when a user makes a rough edit, the optimization module 150 can help to produce the final accurate result. This can significantly improve the efficiency of manual annotation by users. In some embodiments, the user editing module 160 may provide the edited 3D pose graphical representation (or the optimization module 150 may provide the further refined 3D pose graphical representation) to the AI annotation module 140. The AI annotation module 140 may use the edited 3D pose graphical representation or the further refined 3D pose graphical representation to retrain the annotation network used by the AI annotation module 140 or update the object functions used by the optimization module 150. It can help produce more accurate predictions by the updated models for future pose estimation tasks. Such an approach can bring the benefit that the more images are annotated, the less manual intervention is needed.
[0050] The output module 170 generates output files of the pose estimation system 100. The output module 170 may receive 3D pose estimation results from the optimization module 150 or the user editing module 160. The output module 170 may store the 3D pose estimation result in a file. An example 3D pose estimation result may include a 3D pose graphical representation, 3D points positions, 3D point motion parameters, a 3D skeleton template, and so on. The output module 170 may use the 3D skeleton template to produce 3D point positions, e.g., through a forward kinematics function. The 3D point positions and the 3D point motion parameters may be used as ground-truth labels for training the annotation network. The output module 170 may facilitate various file formats, such as JSON, TXT, NPZ, and so on.
[0051] The datastore 180 stores data associated with the pose estimation system 100, such as data received, generated, or used by components of the pose estimation system 100. For instance, the datastore 180 may store parameters (e.g., internal parameters, hyperparameters, etc. ) of the annotation network. The datastore 180 may also store training data and validation data used to train and validate the annotation network. The datastore 180 may further store images received by the interface module 110, 3D pose graphical representations, outputs of the annotation network, outputs of the optimization module 150, files generated by the preprocessing module 120, files generated by the output module 170, and so on. In some embodiments, the pose estimation system 100 may include or be associated with more than one datastore. The datastore 180 may be implemented as a random-access memory (RAM) , such as a static RAM (SRAM) , disk storage, nearline storage, online storage, offline storage, and so on.
[0052] FIG. 2 is a block diagram of an AI annotation module 200, in accordance with various embodiments. The AI annotation module 200 uses an annotation module to process images and predict 3D poses illustrated in the images. The AI annotation module 200 may be an example of the AI annotation module 140 in FIG. 1. In the embodiments of FIG. 2, the AI annotation module 200 includes a training module 210, a validating module 220, an annotation network 230, and a deployment module 240. In other embodiments, alternative configurations, different or additional components may be included in the AI annotation module 200. Further, functionality attributed to a component of the AI annotation module 200 may be accomplished by a different component included in the AI annotation module 200 or by a different module.
[0053] The training module 210 trains the annotation network 230 to perform pose estimation tasks. The annotation network 230 can receive images as inputs and estimate 3D poses of objects in the input images. In some embodiments, the annotation network 230 may include a regression module, a decoder, a kinematic layer, and a projection layer. The regression module processes input images to predict latent representation of the input images, global orientations of objects, and camera parameters. The decoder may map the latent representation of the input images from a latent space to an image space and generate pose parameters indicating a pose in the image space. The image space may have more dimensions than the latent space. In an example, the image space may be two dimensions, and the latent space may be one dimensional. The latent representation of an input image may be a vector. The latent representation may be latent code of a VAE pose prior model. The kinematic layer may apply a forward kinematic function on the pose parameters and the global orientation to predict a 3D pose of the object. The projection layer may use the camera parameters and the 3D pose predicted by the kinematic layer to predict a 2D pose of the object. The 2D pose of the object may be used, e.g., by the optimization module 150, to optimize the 3D pose of the object.
[0054] In some embodiments, the training module 210 trains the annotation network 230 by using one or more training datasets. In some embodiments, the training module 210 forms the one or more training dataset. The training dataset includes training samples and ground-truth labels of the training samples. A training sample may be an image, e.g., a cropped frame of a video. A ground-truth label may include verified or known 2D key points, 3D pose parameters, or masks of the corresponding training sample. In some embodiments, a part of the training dataset may be used to initially train the annotation network 230, and the rest of the training dataset may be held back as a validation subset used by the validating module 220 to validate performance of the annotation network 230 after being trained. The portion of the training dataset not including the validation subset may be used to train the annotation network 230.
[0055] The training module 210 may input a training dataset into the annotation network 230. The training module 210 may modify the parameters inside the annotation network 230 ( “internal parameters of the annotation network 230” ) to minimize the error between labels of the training samples that are generated by the annotation network 230 and the ground-truth labels of the training samples. In some embodiments, the training module 210 may fix the decoder in the annotation network during training, meaning internal parameters of the decoder are not changed during training. The training module 210 may change other internal parameters of the annotation network, such as internal parameters of the regression module, the kinematic layer, or the projection layer. In some embodiments, the training module 210 may pretrain the annotation network 230 by available datasets with the 3D ground-truth labels. The training module 210 may retrain annotation network 230 using a dataset consisting of the input images with 2D ground-truth labels and a small collection of images with 3D ground-truth poses. The training efficiency can be improved.
[0056] During training (either pretraining or retraining) , the training module 210 may update internal parameters of the annotation network based on a loss. In some embodiments, the training module 210 may use a total loss defined as:
[0057] where is an L2 regularization applied to the estimated latent code to enforce the prediction to be in the latent space of the pose priori model’ is an L2 loss to measure the difference between the predicted global rotation and the ground-truth ones; is an L2 loss applied to the pose parameters decoded from the latent code by the decoder; and is an L2 loss applied to the estimated 3D points position P3D obtained by a forward kinematics function. With the inferred weak-perspective camera, the 2D projection of the 3D points may be computed as P2D=sP3D+t. may be a 2D consistency loss that minimizes the reprojection error between the predicted 2D joint locations and the 2D projections of the corresponding joints arising from the pose latent code.
[0058] In some embodiments, the training module 210 may also determine hyperparameters for training the annotation network 230, e.g., before the training process is started. Hyperparameters may be variables specifying the training process. Hyperparameters may be different from parameters inside the annotation network 230 (e.g., weights) . In some embodiments, hyperparameters include variables determining the architecture of the annotation network 230, such as number of layers in backbone, types of layers in backbone, number of layers in branches, types of layers in branches, connections between backbone and branches, connections between branches, and so on.
[0059] Hyperparameters also include variables which determine how the annotation network 230 is trained, such as batch size, number of epochs, etc. A batch size defines the number of training samples to work through before updating the parameters of the annotation network 230. The batch size is the same as or smaller than the number of samples in the training dataset. The training dataset can be divided into one or more batches. The number of epochs defines how many times the entire training dataset is passed forward and backwards through the entire network. The number of epochs defines the number of times that the deep learning algorithm works through the entire training dataset. One epoch means that each training sample in the training dataset has had an opportunity to update the parameters inside the DNN. An epoch may include one or more batches. The number of epochs may be 1, 5, 10, 50, 100, 500, 1000, or even larger.
[0060] The training module 210 may define the architecture of the annotation network 230, e.g., based on some of the hyperparameters. The architecture of the annotation network 230 may include a backbone and a plurality of branches associated with the backbone. The backbone may be a network that includes an input layer, an output layer, and a plurality of hidden layers. The input layer may include tensors (e.g., a multidimensional array) specifying attributes of the input image, such as the height of the input image, the width of the input image, and the depth of the input image (e.g., the number of bits specifying the color of a pixel in the input image) . The output layer includes labels of objects in the input layer. The hidden layers are layers between the input layer and output layer. The hidden layers include one or more convolutional layers and one or more other types of layers, such as pooling layers, fully-connected layers, normalization layers, SoftMax or logistic layers, and so on. The output layer may include an OFM representing features extracted by the backbone. In the process of defining the architecture of the annotation network 230, the training module 210 may also add an activation function to a hidden layer or the output layer. An activation function of a layer transforms the weighted sum of the input of the layer to an output of the layer. The activation function may be, for example, a rectified linear unit (ReLU) activation function, a tangent activation function, or other types of activation functions. The training module 210 may define one or more attributes of tensors computed in the backbone, such as spatial size, datatype, and so on.
[0061] The training module 210 may train the annotation network 230 for a predetermined number of epochs. The number of epochs is a hyperparameter that defines the number of times that the deep learning algorithm will work through the entire training dataset. One epoch means that each sample in the training dataset has had an opportunity to update internal parameters of the annotation network 230. After the training module 210 finishes the predetermined number of epochs, the training module 210 may stop updating the parameters in the annotation network 230.
[0062] After the annotation network is trained, the training module 210 may further train the annotation network. The training module 210 may further train the annotation network during a deployment stage. In an example, the training module 210 may use an input image as a new training sample and use a user-edited 3D pose graphical representation as a ground-truth label of the training sample. The user-edited 3D pose graphical representation may be generated by a user manually editing a 3D pose graphical representation predicted by the annotation network 230. Alternatively, the user-edited 3D pose graphical representation may be generated by a user manually editing a 3D pose graphical representation generated by the optimization module 150 by optimizing a 3D pose graphical representation predicted by the annotation network 230. The training module 210 may update one or more internal parameters of the annotation network 230 based on a loss determined using the 3D pose graphical representation predicted by the annotation network 230 and the user-edited 3D pose graphical representation. In some embodiments, during the retraining process, the training module 210 may update one or more internal parameters of the regression module, the kinematic layer, or the projection layer. The training module 210 may keep the internal parameters of the decoder fixed.
[0063] The validating module 220 verifies accuracy of the annotation network 230 after it is trained by the training module 210. In some embodiments, the validating module 220 inputs samples in a validation dataset into the annotation network 230 and uses the outputs of the annotation network 230 to determine the model accuracy. In some embodiments, a validation dataset may be formed of some or all the samples in the training dataset. Additionally or alternatively, the validation dataset includes additional samples, other than those in the training sets. In some embodiments, the validating module 220 may determine an accuracy score measuring the precision, recall, or a combination of precision and recall of the DNN. The validating module 220 may use the following metrics to determine the accuracy score: Precision = TP / (TP + FP) and Recall = TP / (TP + FN) , where precision may be how many the annotation network 230 correctly predicted (TP or true positives) out of the total it predicted (TP + FP or false positives) , and recall may be how many the annotation network 230 correctly predicted (TP) out of the total number of objects that did have the property in question (TP + FN or false negatives) . The F-score (F-score = 2 *PR / (P + R) ) unifies precision and recall into a single measure.
[0064] The validating module 220 may compare the accuracy score with a threshold score. In an example where the validating module 220 determines that the accuracy score of the annotation network 230 is less than the threshold score, the validating module 220 instructs the training module 210 to retrain the annotation network 230. In one embodiment, the training module 210 may iteratively retrain the annotation network 230 until the occurrence of a stopping condition, such as the accuracy measurement indication that the annotation network 230 may be sufficiently accurate, or a number of training rounds having taken place.
[0065] The annotation network 230 is DNN that. An example of the annotation network 230 is the annotation network 400 in FIG. 4.
[0066] The deployment module 240 deploys the annotation network 230 to perform 3D pose estimation tasks. The deployment module 240 may input images into the annotation network 230. An image may be a cropped frame from a video. The deployment module 240 may input the images into the annotation network 230 one by one. For each input image, the deployment module 240 may obtain information of a 3D pose estimated by the annotation network 230, e.g., a graphical representation that illustrates the estimated 3D pose. In some embodiments, the deployment module 240 may transmit the outputs of the annotation network 230 to optimization module 150 for optimizing the estimated 3D pose.
[0067] Example 3D Pose Estimation
[0068] FIG. 3 illustrates an example pose estimation process 300, in accordance with various embodiments. The pose estimation process 300 may be performed by the pose estimation system 100. The pose estimation process 300 includes six steps: 310, 320, 330, 340, 350, and 360. Although the pose estimation process 300 is described with reference to the flowchart illustrated in FIG. 3, many other processes for pose estimation may alternatively be used. For example, the order of execution of the steps 310, 320, 330, 340, 350, and 360 may be changed. As another example, some of the steps may be changed, eliminated, or combined.
[0069] In Step 310, input is received. The input includes images and 2D pose annotation files are obtained. In some embodiments, an image may be paired with a 2D pose annotation file. The image and the 2D pose annotation file may constitute an example. The image may be generated by cropping another image, e.g., an image captured by a camera. The 2D pose annotation file may be generated by the point labeling module 130 using the image.
[0070] In Step 320, the annotation network is trained and tested using the input images and 2D pose annotation files obtained in Step 310. The annotation network may be the annotation network 230 in FIG. 2. The annotation network may process the images and 2D pose annotation files and predict 3D poses shown in the images. In some embodiments, for each example, the annotation network may generate a 3D graphical representation (e.g., a 3D skeleton) that illustrates the predicted 3D pose.
[0071] In Step 330, 3D pose refinement is performed on each example. The 3D pose refinement may include optimization of the 3D pose predicted by the annotation network by using the corresponding 2D pose annotation file. The optimization may be done by using one or more objective functions.
[0072] In Step 340, it is determined whether the 3D pose is correct. In some embodiments, the user editing module 160 may automatically determine whether the 3D pose is correct, e.g., by determining whether an accuracy score of the 3D pose meets a target accuracy. In other embodiments, a user may determine whether the 3D pose is correct. For instance, a graphical representation of the 3D pose may be presented to the user in a 3D graphical interface. The user may provide, through the 3D graphical interface, feedback indicating whether the 3D pose is correct or not.
[0073] In embodiments where the 3D pose is correct, Step 350 is carried out, in which manual editing on 3D skeletons is performed, e.g., by users in a 3D graphical interface. The 3D graphical interface may present the 3D skeletons to users through client devices associated with the users. A 3D skeleton may include 3D points that a user can change. For instance, the user can add, remove, or move one or more 3D points in the 3D skeleton. The user may also be able to add, remove, or change connections between 3D points in the 3D skeleton. After the manual editing is done, the edited 3D skeletons will be refined in Step 330. Step 330 and the subsequent steps will be performed.
[0074] In embodiments where the 3D pose is correct, Step 360 is carried out, in which files saving 3D pose parameters are output. The files may have various formats, such as JSON, TXT, NPZ, and so on. In some embodiments, the same set of 3D pose parameters may be stored in multiple files having different formats. The 3D pose parameters may be used to update the annotation network used in Step 320 or the object functions used in Step 330.
[0075] FIG. 4 illustrates an example annotation network 400, in accordance with various embodiments. The annotation network 400 may be an example of the annotation network 230 in FIG. 2. In the embodiments of FIG. 4, the annotation network 400 includes a regression module 410, a decoder 420, a kinematic layer 430, and a projection layer 440.
[0076] The annotation network 400 receives an input image 401. The input image 401 captures an object, such as a person, animal, machine, vehicle, etc. The input image 401 may be a cropped image. For instance, the input image 401 may be generated by removing a background or one or more other objects from an image generated by a camera. The input image 401 may be associated with labeled 2D key points.
[0077] The regression module 410 receives and processes the input image 401. In some embodiments, the regression module 410 may be a neural network that includes a plurality of layers, e.g., an input layer, one or more hidden layers, and one or more output layers corresponding to one or more tasks that the regression module 410 can perform. For instance, the regression module 410 may perform an encoding task to generate latent code 402. In an example, the regression module 410 may include an encoder. The encoder may be a network including one or more encoding layers. The encoder may map the input image 401 to a latent space that corresponds to the parameters of a variational distribution. The latent code 402 may be a latent representation of the input image 401 in the latent space. The latent code 402 is input into the decoder 420. The decoder 420 may be a network including one or more decoding layers. The decoder 420 generates pose parameters 403 using the latent code 402. The decoder 420 may map the latent code 402 from the latent space to the image space to produce the pose parameters 403. The image space may be a 2D image plane. The pose parameters 403 may indicate a pose of the object in the 2D image plane. In some embodiments, the decoder 420 may have an opposite function from the encoder. The encoder in the regression module 410 and the decoder 420 may constitute a pose prior model, such as a VAE-based pose prior model, which can learn pose prior (e.g., 3D human pose prior) . In some embodiments, the regression module 410 and the decoder 420 may be trained separately. For instance, the decoder 420 may be pretrained. During the training of the annotation network 400, the internal parameters of the decoder 420 may be fixed while the regression module 410 is being trained.
[0078] The regression module 410 performs another regression task to predict a global orientation 404. The global orientation 404 may indicate an orientation of the object. The orientation may be position, direction, or a combination of both. The global orientation 404 is input into the kinematic layer 430. In the kinematic layer 430, a forward kinematic function may be applied on the global orientation 404 to determine a 3D pose 405. In some embodiments, the kinematic layer 430 may output one or more graphical representations of the 3D pose 405.
[0079] The regression module 410 performs yet another regression task to determine camera parameters 406. The camera parameters 406 may include camera intrinsic parameters, such as the camera intrinsic parameters described above. The 3D pose 405 and the camera parameters 406 are input into a projection layer 440. In the projection layer, a pinhole projection function may be applied on the 3D pose 405 and the camera parameters 406 to predict 2D pose 407.
[0080] The latent code 402, pose parameters 403, global orientation 404, 3D pose 405, and 2D pose 407 may be used to compute a loss of the annotation network 400 during training. Internal parameters of the annotation network 400, such as internal parameters of the regression module 410, the kinematic layer 430, or the projection layer 440, may be adjusted to minimize the loss. After the annotation network 400 is trained, the annotation network 400 may be retrained during deployment. For instance, the annotation network 400 may receive and process another input image after the training is done and predicts latent code, pose parameters, global orientation, 3D pose, and 2D pose based on the other input image. The 3D pose may be viewed and modified by a user. The annotation network 400 may be further trained using the other input image as a new training sample and using the modified 3D pose a ground-truth label of the new training sample. Internal parameters of the regression module 410, the kinematic layer 430, or the projection layer 440, may be further adjusted during retraining.
[0081] FIG. 5 illustrates a 2D annotation file 510 generated from an image 520, in accordance with various embodiments. The image 520 may be extracted from a camera image. The image 520 shows a person. The image 520 may be in a 2D image space, which may be defined by the X and Z axes shown in FIG. 5. The 2D annotation file 510 includes a plurality of key points 515, individually referred to as “key point 515. ” The key points 515 may be identified from the image 520, e.g., by the point labeling module 130. In some embodiments, each key point may correspond to a bone joint of the person. The key points 515 are connected using lines. The connections between the key points 515 may be determined based on the connections of the corresponding bone joints. The 2D annotation file 510 represents a 2D pose of the person. For the purpose of illustration, the 2D annotation file 510 includes 21 key points 515 in FIG. 5. In other embodiments, the 2D annotation file 510 may include different, fewer, or more key points 515. Also, the connections between the key points may be different from the connections shown in FIG. 5.
[0082] FIG. 6 illustrates an example 3D pose graphical representation 610, in accordance with various embodiments. The 3D pose graphical representation 610 represents a 3D pose of a person in a camera space. The camera space may be a 3D space defined by the X, Y, and Z axes in FIG. 6. The 3D pose graphical representation 610 may be generated by an annotation network, such as the annotation network 230, from an image 620. The image 620 may be a 2D image, e.g., in a 2D plane defined by the X-Z axes.
[0083] FIG. 7 illustrates an optimized 3D pose graphical representation 710, in accordance with various embodiments. The optimized 3D pose graphical representation 710 may be generated by optimizing the 3D pose graphical representation 610 in FIG. 6, e.g., by the optimization module 150. Compared with the 3D pose graphical representation 610, the 3D pose shown in the optimized 3D pose graphical representation 710 matches the person’s pose in the image 620 better. For instance, the pose of the left hand of the person shown in the optimized 3D pose graphical representation 710 is more accurate than the 3D pose graphical representation 610.
[0084] FIG. 8 illustrates an editable 3D skeleton 810, in accordance with various embodiments. The editable 3D skeleton 810 is part of a 3D pose graphical representation 820. The editable 3D skeleton 810 shows a pose of a person. The 3D pose graphical representation 820 also includes a graphical representation illustrating the body of the person. The 3D pose graphical representation 820 may be generated by the optimization module 150, e.g., by optimizing a 3D pose graphical representation generated by the annotation network 230. The 3D pose graphical representation 820 may be in a 3D camera space defined by the X, Y, and Z axes in FIG. 8. The editable 3D skeleton 810 includes a plurality of editable points 815, individually referred to as “editable point 815. ” As shown in FIG. 8, the editable points 815 are 3D points with three dimensions along the X, Y, and Z axes, respectively.
[0085] The editable 3D skeleton 810 (or the entire 3D pose graphical representation 820) may be presented to a user in a user interface. The user interface may be an interactive graphical interface. The user can edit the editable points 815. For instance, the user can remove an editable point 815, add a new editable point, move the position of an editable point 815, change the connection of an editable point 815 with another editable point 815, and so on. The user interface may also allow the user to get multiple, different views of the editable 3D skeleton 810 (or the entire 3D pose graphical representation 820) . For instance, the user interface may allow the user to rotate or move the editable 3D skeleton 810 (or the entire 3D pose graphical representation 820) . In an example, the user may select the editable point 815 that representing the joint of left ankle. The user may also change the local rotation value along the Y-axis to switch views between front view and side view to see 3D results. This process may be repeated until the satisfactory results are achieved. In another example, the user interface may allow the user to modify the rotations on the editable points 815 representing joints of right ankle, right shoulder, right elbow, and right wrist. After the user edits the editable 3D skeleton 810, the 3D pose graphical representation 820 may be updated. The updated 3D pose graphical representation may be further optimized by the optimization module 150. The annotation network 230 may be further trained using the updated 3D pose graphical representation or the further optimized 3D pose graphical representation.
[0086] Example DNN
[0087] FIG. 9 illustrates an example DNN 900, in accordance with various embodiments. The DNN 900 (or part of the DNN 900) may be an example of the regression module 410 in FIG. 4. In the embodiments of FIG. 9, the DNN 900 includes a sequence of layers comprising a plurality of convolutional layers 910 (individually referred to as “convolutional layer 910” ) , a plurality of pooling layers 920 (individually referred to as “pooling layer 920” ) , and a plurality of fully-connected layers 930 (individually referred to as “fully-connected layer 930” ) . In other embodiments, the DNN 900 may include fewer, more, or different layers. In an inference of the DNN 900, the layers of the DNN 900 execute tensor computation that includes many tensor operations, such as convolution (e.g., multiply-accumulate (MAC) operations, etc. ) , pooling operations, elementwise operations (e.g., elementwise addition, elementwise multiplication, etc. ) , other types of tensor operations, or some combination thereof.
[0088] The convolutional layers 910 summarize the presence of features in the input to the DNN 900. The convolutional layers 910 function as feature extractors. The first layer of the DNN 900 is a convolutional layer 910. In an example, a convolutional layer 910 performs a convolution on an input tensor 940 (also referred to as IFM 940) and a filter 950. As shown in FIG. 9, the IFM 940 is represented by a 7×7×3 three-dimensional (3D) matrix. The IFM 940 includes 3 input channels, each of which is represented by a 7×7 two-dimensional (2D) matrix. The 7×7 2D matrix includes 7 input elements (also referred to as input points) in each row and seven input elements in each column. The filter 950 is represented by a 3×3×3 3D matrix. The filter 950 includes 3 kernels, each of which may correspond to a different input channel of the IFM 940. A kernel is a 2D matrix of weights, where the weights are arranged in columns and rows. A kernel can be smaller than the IFM. In the embodiments of FIG. 9, each kernel is represented by a 3×3 2D matrix. The 3×3 kernel includes 3 weights in each row and three weights in each column. Weights can be initialized and updated by backpropagation using gradient descent. The magnitudes of the weights can indicate importance of the filter 950 in extracting features from the IFM 940.
[0089] The convolution includes MAC operations with the input elements in the IFM 940 and the weights in the filter 950. The convolution may be a standard convolution 963 or a depthwise convolution 983. In the standard convolution 963, the whole filter 950 slides across the IFM 940. All the input channels are combined to produce an output tensor 960 (also referred to as OFM 960) . The OFM 960 is represented by a 5×5 2D matrix. The 5×5 2D matrix includes 5 output elements (also referred to as output points) in each row and five output elements in each column. For the purpose of illustration, the standard convolution includes one filter in the embodiments of FIG. 9. In embodiments where there are multiple filters, the standard convolution may produce multiple output channels in the OFM 960.
[0090] The multiplication applied between a kernel-sized patch of the IFM 940 and a kernel may be a dot product. A dot product is the elementwise multiplication between the kernel-sized patch of the IFM 940 and the corresponding kernel, which is then summed, always resulting in a single value. Because it results in a single value, the operation is often referred to as the “scalar product. ” Using a kernel smaller than the IFM 940 is intentional as it allows the same kernel (set of weights) to be multiplied by the IFM 940 multiple times at different points on the IFM 940. Specifically, the kernel is applied systematically to each overlapping part or kernel-sized patch of the IFM 940, left to right, top to bottom. The result from multiplying the kernel with the IFM 940 one time is a single value. As the kernel is applied multiple times to the IFM 940, the multiplication result is a 2D matrix of output elements. As such, the 2D output matrix (i.e., the OFM 960) from the standard convolution 963 is referred to as an OFM.
[0091] In the depthwise convolution 983, the input channels are not combined. Rather, MAC operations are performed on an individual input channel and an individual kernel and produce an output channel. As shown in FIG. 9, the depthwise convolution 983 produces a depthwise output tensor 980. The depthwise output tensor 980 is represented by a 5×5×3 3D matrix. The depthwise output tensor 980 includes 3 output channels, each of which is represented by a 5×5 2D matrix. The 5×5 2D matrix includes 5 output elements in each row and five output elements in each column. Each output channel is a result of MAC operations of an input channel of the IFM 940 and a kernel of the filter 950. For instance, the first output channel (patterned with dots) is a result of MAC operations of the first input channel (patterned with dots) and the first kernel (patterned with dots) , the second output channel (patterned with horizontal strips) is a result of MAC operations of the second input channel (patterned with horizontal strips) and the second kernel (patterned with horizontal strips) , and the third output channel (patterned with diagonal stripes) is a result of MAC operations of the third input channel (patterned with diagonal stripes) and the third kernel (patterned with diagonal stripes) . In such a depthwise convolution, the number of input channels equals the number of output channels, and each output channel corresponds to a different input channel. The input channels and output channels are referred to collectively as depthwise channels. After the depthwise convolution, a pointwise convolution 993 is then performed on the depthwise output tensor 980 and a 9×1×3 tensor 990 to produce the OFM 960.
[0092] The OFM 960 is then passed to the next layer in the sequence. In some embodiments, the OFM 960 is passed through an activation function. An example activation function is ReLU. ReLU is a calculation that returns the value provided as input directly, or the value zero if the input is zero or less. The convolutional layer 910 may receive several images as input and calculate the convolution of each of them with each of the kernels. This process can be repeated several times. For instance, the OFM 960 is passed to the subsequent convolutional layer 910 (i.e., the convolutional layer 910 following the convolutional layer 910 generating the OFM 960 in the sequence) . The subsequent convolutional layers 910 perform a convolution on the OFM 960 with new kernels and generate a new feature map. The new feature map may also be normalized and resized. The new feature map can be kernelled again by a further subsequent convolutional layer 910, and so on.
[0093] In some embodiments, a convolutional layer 910 has four hyperparameters: the number of kernels, the size F kernels (e.g., a kernel is of dimensions F×F×D pixels) , the S step with which the window corresponding to the kernel is dragged on the image (e.g., a step of one means moving the window one pixel at a time) , and the zero-padding P (e.g., adding a black contour of P pixels thickness to the input image of the convolutional layer 910) . The convolutional layers 910 may perform various types of convolutions, such as 2-dimensional convolution, dilated or atrous convolution, spatial separable convolution, depthwise separable convolution, transposed convolution, and so on. The DNN 900 includes 96 convolutional layers 910. In other embodiments, the DNN 900 may include a different number of convolutional layers.
[0094] The pooling layers 920 down-sample feature maps generated by the convolutional layers, e.g., by summarizing the presence of features in the patches of the feature maps. A pooling layer 920 is placed between two convolution layers 910: a preceding convolutional layer 910 (the convolution layer 910 preceding the pooling layer 920 in the sequence of layers) and a subsequent convolutional layer 910 (the convolution layer 910 subsequent to the pooling layer 920 in the sequence of layers) . In some embodiments, a pooling layer 920 is added after a convolutional layer 910, e.g., after an activation function (e.g., ReLU, etc. ) has been applied to the OFM 960.
[0095] A pooling layer 920 receives feature maps generated by the preceding convolution layer 910 and applies a pooling operation to the feature maps. The pooling operation reduces the size of the feature maps while preserving their important characteristics. Accordingly, the pooling operation improves the efficiency of the CNN and avoids over-learning. The pooling layers 920 may perform the pooling operation through average pooling (calculating the average value for each patch on the feature map) , max pooling (calculating the maximum value for each patch of the feature map) , or a combination of both. The size of the pooling operation is smaller than the size of the feature maps. In various embodiments, the pooling operation is 2×2 pixels applied with a stride of two pixels, so that the pooling operation reduces the size of a feature map by a factor of 2, e.g., the number of pixels or values in the feature map is reduced to one quarter the size. In an example, a pooling layer 920 applied to a feature map of 6×6 results in an output pooled feature map of 3×3. The output of the pooling layer 920 is inputted into the subsequent convolution layer 910 for further feature extraction. In some embodiments, the pooling layer 920 operates upon each feature map separately to create a new set of the same number of pooled feature maps.
[0096] The fully-connected layers 930 are the last layers of the CNN. The fully-connected layers 930 may be convolutional or not. The fully-connected layers 930 may also be referred to as linear layers. In some embodiments, a fully-connected layer 930 (e.g., the first fully-connected layer in the DNN 900) may receive an input operand. The input operand may define the output of the convolutional layers 910 and pooling layers 920 and includes the values of the last feature map generated by the last pooling layer 920 in the sequence. The fully-connected layer 930 may apply a linear transformation to the input operand through a weight matrix. The weight matrix may be a kernel of the fully-connected layer 930. The linear transformation may include a tensor multiplication between the input operand and the weight matrix. The result of the linear transformation may be an output operand. In some embodiments, the fully-connected layer may further apply a nonlinear transformation (e.g., by using a nonlinear activation function) on the result of the linear transformation to generate an output operand. The output operand may contain as many elements as there are classes: element i represents the probability that the image belongs to class i. Each element is therefore between 0 and 9, and the sum of all is worth one. These probabilities are calculated by the last fully-connected layer 930 by using a logistic function (binary classification) or a SoftMax function (multi-class classification) as an activation function.
[0097] Example AI-based Pose Estimation Environment
[0098] FIG. 10 illustrates an AI-based pose estimation environment 1000, in accordance with various embodiments. The AI-based pose estimation environment 1000 includes a pose estimation system 1010, client devices 1020 (individually referred to as client device 1020) , and a third-party system 1030. In other embodiments, the AI-based pose estimation environment 1000 may include fewer, more, or different components. For instance, the AI-based pose estimation environment 1000 may include a different number of client devices 1020 or more than one third-party system 1030.
[0099] The pose estimation system 1010 estimates 3D poses of objects from 2D images of the objects. In some embodiments, the pose estimation system 1010 uses an annotation network to estimate 3D poses. The pose estimation system 1010 may receive images from one or more client devices 1020 or from the third-party system 1030. Also, the pose estimation system 1010 may transmit files storing estimated 3D pose parameters to one or more client devices 1020 or from the third-party system 1030. In some embodiments, the pose estimation system 1010 also provides an interactive graphical interface that allows users to edit 3D poses estimated by the pose estimation system 1010. The interactive graphical interface may be available online, e.g., in the Cloud, or can be downloaded and installed on the client devices 1020. An example of the pose estimation system 1010 is the pose estimation system 100 in FIG. 1.
[0100] A client device 1020 allows its user (s) to interact with the pose estimation system 1010, e.g., through one or more user interfaces supported by the pose estimation system 1010. For example, the client device 1020 may receive 3D pose estimation results from the pose estimation system 1010 and display the 3D pose graphical representations to one or more users associated with the client device 1020 in an interactive graphical interface executed by the client device. The interactive graphical interface may be provided by the pose estimation system 1010 to present. The client devices 1020 may include components (e.g., keyboards, mouses, etc. ) that can facilitate the users to edit the 3D pose graphical representations in the interactive graphical interface.
[0101] In some embodiments, a client device 1020 may execute one or more applications allowing one or more users of the client device 1020 to interact with the pose estimation system 1010. For example, a client device 1020 executes a browser application to enable interaction between the client device 1020 and the pose estimation system 1010. In another embodiment, a client device 1020 interacts with the pose estimation system 1010 through an application programming interface (API) running on a native operating system of the client device 1020, such as or ANDROIDTM.
[0102] A client device 1020 may be one or more computing devices capable of receiving user input as well as transmitting and / or receiving data via the network 1040. In one embodiment, a client device 1020 is a conventional computer system, such as a desktop or a laptop computer. Alternatively, a client device 1020 may be a device having computer functionality, such as a personal digital assistant (PDA) , a mobile telephone, a smartphone, an autonomous vehicle, or another suitable device. A client device 1020 is configured to communicate via the network 1040. In an embodiment, a client device 1020 is an integrated computing device that operates as a standalone network-enabled device. For example, the client device 1020 includes display, speakers, microphone, camera, and input device. In another embodiment, a client device 1020 is a computing device for coupling to an external media device such as a television or other external display and / or audio output system. In this embodiment, the client device 1020 may couple to the external media device via a wireless interface or wired interface (e.g., an HDMI cable) and may utilize various functions of the external media device such as its display, speakers, microphone, camera, and input devices. Here, the client device 1020 may be configured to be compatible with a generic external media device that does not have specialized software, firmware, or hardware specifically for interacting with the client device 1020.
[0103] The pose estimation system 1010, client devices 1020, and third-party system 1030 are connected through a network 1040. The network 1040 may comprise any combination of local area and / or wide area networks, using both wired and / or wireless communication systems. In one embodiment, the network 1040 may use standard communications technologies and / or protocols. For example, the network 1040 may include communication links using technologies such as Ethernet, 8010.11, worldwide interoperability for microwave access (WiMAX) , 3G, 4G, code division multiple access (CDMA) , digital subscriber line (DSL) , etc. Examples of networking protocols used for communicating via the network 1040 may include multiprotocol label switching (MPLS) , transmission control protocol / Internet protocol (TCP / IP) , hypertext transport protocol (HTTP) , simple mail transfer protocol (SMTP) , and file transfer protocol (FTP) . Data exchanged over the network 1040 may be represented using any suitable format, such as hypertext markup language (HTML) or extensible markup language (XML) . In some embodiments, all or some of the communication links of the network 1040 may be encrypted using any suitable technique or techniques.
[0104] A third-party system 1030 is an online system that may communicate with the pose estimation system 1010 or at least one of the client devices 1020. In some embodiments, the third-party system 1030 may provide images to the pose estimation system 1010 for 3D pose estimation. The third-party system 1030 may be a social media system, an online image gallery, an online searching system, and so on. Additionally or alternatively, the third-party system 1030 may use results of 3D pose estimation in various applications. For instance, the third-party system 1030 may use files storing 3D pose parameters estimated by the pose estimation system 1010 for motion tracking, action recognition, sport analysis, virtual reality, augmented reality, film and game production, telepresence, and so on.
[0105] Example Method of Pose Estimation
[0106] FIG. 11 is a flowchart showing a method 1100 of pose estimation, in accordance with various embodiments. The method 1100 may be used for 3D human pose estimation. The method 1100 may be performed by the pose estimation system 100 in FIG. 1. Although the method 1100 is described with reference to the flowchart illustrated in FIG. 11, many other methods for pose estimation may alternatively be used. For example, the order of execution of the steps in FIG. 11 may be changed. As another example, some of the steps may be changed, eliminated, or combined.
[0107] The pose estimation system 100 determines 1110, by a regression module in a neural network using an input image, a latent representation of the input image in a latent space, an orientation of an object, and one or more camera parameters. The one or more camera parameters indicate one or more settings of a camera capturing the input image. In some embodiments, the one or more settings of the camera include position of an optical center of the camera, a focal length of the camera, or other settings of the camera. In some embodiments, the regression module includes an encoder that maps the input image to the latent representation of the input image in the latent space. The latent space corresponds to a variational distribution. In some embodiments, the input image is generated by cropping an image captured by the camera.
[0108] The pose estimation system 100 coverts 1120, by a decoder in the neural network, the latent representation of the input image into one or more pose parameters indicating a pose of the object in an image space. In some embodiments, the image space has more dimensions than the latent space. In some embodiments, the image space is 2D.
[0109] The pose estimation system 100 estimates 1130, by a kinematic layer in the neural network using the orientation of the object and the one or more pose parameters, a 3D pose of the object. In some embodiments, the kinematic layer has a forward kinematic function that is applied on the orientation of the object and the one or more pose parameters and the one or more pose parameters to estimate the 3D pose of the object.
[0110] The pose estimation system 100 generates 1140 a graphical representation of the estimated 3D pose of the object. In some embodiments, the graphical representation of the estimated 3D pose of the object includes a 3D skeleton, a 3D mesh, or other types of 3D graphical representations.
[0111] The pose estimation system 100 optimizes 1150 the graphical representation of the estimated 3D pose based on the one or more camera parameters. In some embodiments, the pose estimation system 100 estimates a motion of a root point of the object based on the graphical representation of the estimated 3D pose and key points of the object. The key points are identified using the input image. The key points include the root point. The pose estimation system 100 adjusts the graphical representation of the estimated 3D pose to minimize an object function of the estimated motion of the root point.
[0112] In some embodiments, the pose estimation system 100 projects, by a projection layer in the neural network using the one or more camera parameters, points in the graphical representation of the estimated 3D pose from the camera space into the image space. The pose estimation system 100 measures a distance between the projected points and key points of the object, the key points identified using the input image. The pose estimation system 100 adjusts the graphical representation of the estimated 3D pose based on the measured distance. In some embodiments, the camera space is a three-dimensional space.
[0113] In some embodiments, the pose estimation system 100 optimizes the latent representation of the input image to generate a modified latent representation of the input image. The pose estimation system 100 measures a similarity between the latent representation of the input image and the modified latent representation of the input image. The pose estimation system 100 adjusts the graphical representation of the estimated 3D pose based on the measured similarity.
[0114] In some embodiments, the pose estimation system 100 provides the optimized graphical representation of the estimated 3D pose for display in a graphical interface. The graphical interface allows a user to view and modify the optimized graphical representation of the estimated 3D pose. In some embodiments, the pose estimation system 100 receives a modified graphical representation of the estimated 3D pose from a client device associated with the user. The pose estimation system 100 further trains the neural network based on the input image and the modified graphical representation of the estimated 3D pose. In some embodiments, the pose estimation system 100 further trains the neural network by modifying one or more internal parameters of the regression module while keeping one or more internal parameters of the decoder fixed.
[0115] Example Computing Device
[0116] FIG. 12 is a block diagram of an example computing device 1200, in accordance with various embodiments. In some embodiments, the computing device 1200 can be used as at least part of the pose estimation system 100. A number of components are illustrated in FIG. 12 as included in the computing device 1200, but any one or more of these components may be omitted or duplicated, as suitable for the application. In some embodiments, some or all of the components included in the computing device 1200 may be attached to one or more motherboards. In some embodiments, some or all of these components are fabricated onto a single system on a chip (SoC) die. Additionally, in various embodiments, the computing device 1200 may not include one or more of the components illustrated in FIG. 12, but the computing device 1200 may include interface circuitry for coupling to the one or more components. For example, the computing device 1200 may not include a display device 1206, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which a display device 1206 may be coupled. In another set of examples, the computing device 1200 may not include an audio input device 1218 or an audio output device 1208, but may include audio input or output device interface circuitry (e.g., connectors and supporting circuitry) to which an audio input device 1218 or audio output device 1208 may be coupled.
[0117] The computing device 1200 may include a processing device 1202 (e.g., one or more processing devices) . The processing device 1202 processes electronic data from registers and / or memory to transform that electronic data into other electronic data that may be stored in registers and / or memory. The computing device 1200 may include a memory 1204, which may itself include one or more memory devices such as volatile memory (e.g., DRAM) , nonvolatile memory (e.g., read-only memory (ROM) ) , high bandwidth memory (HBM) , flash memory, solid state memory, and / or a hard drive. In some embodiments, the memory 1204 may include memory that shares a die with the processing device 1202. In some embodiments, the memory 1204 includes one or more non-transitory computer-readable media storing instructions executable to perform operations for pose estimation, e.g., the method 1100 described above in conjunction with FIG. 11 or some operations performed by the pose estimation system 100. The instructions stored in the one or more non-transitory computer-readable media may be executed by the processing device 1202.
[0118] In some embodiments, the computing device 1200 may include a communication chip 1212 (e.g., one or more communication chips) . For example, the communication chip 1212 may be configured for managing wireless communications for the transfer of data to and from the computing device 1200. The term "wireless" and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data through the use of modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not.
[0119] The communication chip 1212 may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.10 family) , IEEE 802.16 standards (e.g., IEEE 802.16-2005 Amendment) , Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as "3GPP2" ) , etc. ) . IEEE 802.16 compatible Broadband Wireless Access (BWA) networks are generally referred to as WiMAX networks, an acronym that stands for worldwide interoperability for microwave access, which is a certification mark for products that pass conformity and interoperability tests for the IEEE 802.16 standards. The communication chip 1212 may operate in accordance with a Global System for Mobile Communication (GSM) , General Packet Radio Service (GPRS) , Universal Mobile Telecommunications System (UMTS) , High Speed Packet Access (HSPA) , Evolved HSPA (E-HSPA) , or LTE network. The communication chip 1212 may operate in accordance with Enhanced Data for GSM Evolution (EDGE) , GSM EDGE Radio Access Network (GERAN) , Universal Terrestrial Radio Access Network (UTRAN) , or Evolved UTRAN (E-UTRAN) . The communication chip 1212 may operate in accordance with CDMA, Time Division Multiple Access (TDMA) , Digital Enhanced Cordless Telecommunications (DECT) , Evolution-Data Optimized (EV-DO) , and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. The communication chip 1212 may operate in accordance with other wireless protocols in other embodiments. The computing device 1200 may include an antenna 1222 to facilitate wireless communications and / or to receive other wireless communications (such as AM or FM radio transmissions) .
[0120] In some embodiments, the communication chip 1212 may manage wired communications, such as electrical, optical, or any other suitable communication protocols (e.g., the Ethernet) . As noted above, the communication chip 1212 may include multiple communication chips. For instance, a first communication chip 1212 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second communication chip 1212 may be dedicated to longer-range wireless communications such as global positioning system (GPS) , EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some embodiments, a first communication chip 1212 may be dedicated to wireless communications, and a second communication chip 1212 may be dedicated to wired communications.
[0121] The computing device 1200 may include battery / power circuitry 1214. The battery / power circuitry 1214 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 1200 to an energy source separate from the computing device 1200 (e.g., AC line power) .
[0122] The computing device 1200 may include a display device 1206 (or corresponding interface circuitry, as discussed above) . The display device 1206 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD) , a light-emitting diode display, or a flat panel display, for example.
[0123] The computing device 1200 may include an audio output device 1208 (or corresponding interface circuitry, as discussed above) . The audio output device 1208 may include any device that generates an audible indicator, such as speakers, headsets, or earbuds, for example.
[0124] The computing device 1200 may include an audio input device 1218 (or corresponding interface circuitry, as discussed above) . The audio input device 1218 may include any device that generates a signal representative of a sound, such as microphones, microphone arrays, or digital instruments (e.g., instruments having a musical instrument digital interface (MIDI) output) .
[0125] The computing device 1200 may include a GPS device 1216 (or corresponding interface circuitry, as discussed above) . The GPS device 1216 may be in communication with a satellite-based system and may receive a location of the computing device 1200, as known in the art.
[0126] The computing device 1200 may include another output device 1210 (or corresponding interface circuitry, as discussed above) . Examples of the other output device 1210 may include an audio codec, a video codec, a printer, a wired or wireless transmitter for providing information to other devices, or an additional storage device.
[0127] The computing device 1200 may include another input device 1220 (or corresponding interface circuitry, as discussed above) . Examples of the other input device 1220 may include an accelerometer, a gyroscope, a compass, an image capture device, a keyboard, a cursor control device such as a mouse, a stylus, a touchpad, a bar code reader, a Quick Response (QR) code reader, any sensor, or a radio frequency identification (RFID) reader.
[0128] The computing device 1200 may have any desired form factor, such as a handheld or mobile computer system (e.g., a cell phone, a smart phone, a mobile internet device, a music player, a tablet computer, a laptop computer, a netbook computer, an ultrabook computer, a PDA, an ultramobile personal computer, etc. ) , a desktop computer system, a server or other networked computing component, a printer, a scanner, a monitor, a set-top box, an entertainment control unit, a vehicle control unit, a digital camera, a digital video recorder, or a wearable computer system. In some embodiments, the computing device 1200 may be any other electronic device that processes data.
[0129] Select Examples
[0130] The following paragraphs provide various examples of the embodiments disclosed herein.
[0131] Example 1 provides a method, including determining, by a regression module in a neural network using an input image, a latent representation of the input image in a latent space, an orientation of an object, and one or more camera parameters indicating one or more settings of a camera capturing the input image; converting, by a decoder in the neural network, the latent representation of the input image into one or more pose parameters indicating a pose of the object in an image space; estimating, by a kinematic layer in the neural network using the orientation of the object and the one or more pose parameters, a three-dimensional (3D) pose of the object in a camera space; generating a graphical representation of the estimated 3D pose of the object; and optimizing the graphical representation of the estimated 3D pose based on the one or more camera parameters.
[0132] Example 2 provides the method of example 1, in which the regression module includes an encoder that maps the input image to the latent representation of the input image in the latent space, and the latent space corresponds to a variational distribution.
[0133] Example 3 provides the method of example 1 or 2, in which optimizing the graphical representation of the estimated 3D pose of the object includes estimating a motion of a root point of the object based on the graphical representation of the estimated 3D pose and key points of the object, the key points identified using the input image, the key points including the root point; and adjusting the graphical representation of the estimated 3D pose to minimize an object function of the estimated motion of the root point.
[0134] Example 4 provides the method of any one of examples 1-3, in which optimizing the estimated 3D pose of the object includes projecting, by a projection layer in the neural network using the one or more camera parameters, points in the graphical representation of the estimated 3D pose from the camera space into the image space; measuring a distance between the projected points and key points of the object, the key points identified using the input image; and adjusting the graphical representation of the estimated 3D pose based on the measured distance.
[0135] Example 5 provides the method of any one of examples 1-4, in which optimizing the estimated 3D pose of the object includes optimizing the latent representation of the input image to generate a modified latent representation of the input image; measuring a similarity between the latent representation of the input image and the modified latent representation of the input image; and adjusting the graphical representation of the estimated 3D pose based on the measured similarity.
[0136] Example 6 provides the method of any one of examples 1-5, further including providing the optimized graphical representation of the estimated 3D pose for display in a graphical interface, the graphical interface allowing a user to view and modify the optimized graphical representation of the estimated 3D pose.
[0137] Example 7 provides the method of example 6, further including receiving a modified graphical representation of the estimated 3D pose from a client device associated with the user; and further training the neural network based on the input image and the modified graphical representation of the estimated 3D pose.
[0138] Example 8 provides the method of example 7, in which further training the neural network includes modifying one or more internal parameters of the regression module; and keeping one or more internal parameters of the decoder fixed.
[0139] Example 9 provides the method of any one of examples 1-8, in which the image space has more dimensions than the latent space.
[0140] Example 10 provides the method of any one of examples 1-9, in which the image space is a two-dimensional space, and the camera space is a three-dimensional space.
[0141] Example 11 provides one or more non-transitory computer-readable media storing instructions executable to perform operations, the operations including determining, by a regression module in a neural network using an input image, a latent representation of the input image in a latent space, an orientation of an object, and one or more camera parameters indicating one or more settings of a camera capturing the input image; converting, by a decoder in the neural network, the latent representation of the input image into one or more pose parameters indicating a pose of the object in an image space; estimating, by a kinematic layer in the neural network using the orientation of the object and the one or more pose parameters, a three-dimensional (3D) pose of the object in a camera space; generating a graphical representation of the estimated 3D pose of the object; and optimizing the graphical representation of the estimated 3D pose based on the one or more camera parameters.
[0142] Example 12 provides the one or more non-transitory computer-readable media of example 11, in which the regression module includes an encoder that maps the input image to the latent representation of the input image in the latent space, and the latent space corresponds to a variational distribution.
[0143] Example 13 provides the one or more non-transitory computer-readable media of example 11 or 12, in which optimizing the graphical representation of the estimated 3D pose of the object includes estimating a motion of a root point of the object based on the graphical representation of the estimated 3D pose and key points of the object, the key points identified using the input image, the key points including the root point; and adjusting the graphical representation of the estimated 3D pose to minimize an object function of the estimated motion of the root point.
[0144] Example 14 provides the one or more non-transitory computer-readable media of any one of examples 11-13, in which optimizing the estimated 3D pose of the object includes projecting, by a projection layer in the neural network using the one or more camera parameters, points in the graphical representation of the estimated 3D pose from the camera space into the image space; measuring a distance between the projected points and key points of the object, the key points identified using the input image; and adjusting the graphical representation of the estimated 3D pose based on the measured distance.
[0145] Example 15 provides the one or more non-transitory computer-readable media of any one of examples 11-14, in which optimizing the estimated 3D pose of the object includes optimizing the latent representation of the input image to generate a modified latent representation of the input image; measuring a similarity between the latent representation of the input image and the modified latent representation of the input image; and adjusting the graphical representation of the estimated 3D pose based on the measured similarity.
[0146] Example 16 provides the one or more non-transitory computer-readable media of any one of examples 11-15, in which the operations further include providing the optimized graphical representation of the estimated 3D pose for display in a graphical interface, the graphical interface allowing a user to view and modify the optimized graphical representation of the estimated 3D pose; receiving a modified graphical representation of the estimated 3D pose from a client device associated with the user; and further training the neural network based on the input image and the modified graphical representation of the estimated 3D pose.
[0147] Example 17 provides the one or more non-transitory computer-readable media of any one of examples 11-16, in which the image space is a two-dimensional space, and the camera space is a three-dimensional space.
[0148] Example 18 provides an apparatus, including a computer processor for executing computer program instructions; and a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations including determining, by a regression module in a neural network using an input image, a latent representation of the input image in a latent space, an orientation of an object, and one or more camera parameters indicating one or more settings of a camera capturing the input image, converting, by a decoder in the neural network, the latent representation of the input image into one or more pose parameters indicating a pose of the object in an image space, estimating, by a kinematic layer in the neural network using the orientation of the object and the one or more pose parameters, a three-dimensional (3D) pose of the object in a camera space, generating a graphical representation of the estimated 3D pose of the object, and optimizing the graphical representation of the estimated 3D pose based on the one or more camera parameters.
[0149] Example 19 provides the apparatus of example 18, in which the regression module includes an encoder that maps the input image to the latent representation of the input image in the latent space, and the latent space corresponds to a variational distribution.
[0150] Example 20 provides the apparatus of example 19, in which the operations further include providing the optimized graphical representation of the estimated 3D pose for display in a graphical interface, the graphical interface allowing a user to view and modify the optimized graphical representation of the estimated 3D pose; receiving a modified graphical representation of the estimated 3D pose from a client device associated with the user; and further training the neural network based on the input image and the modified graphical representation of the estimated 3D pose.
[0151] The above description of illustrated implementations of the disclosure, including what is described in the Abstract, is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. While specific implementations of, and examples for, the disclosure are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the disclosure, as those skilled in the relevant art will recognize. These modifications may be made to the disclosure in light of the above detailed description.
Claims
1.A method, comprising:determining, by a regression module in a neural network using an input image, a latent representation of the input image in a latent space, an orientation of an object, and one or more camera parameters indicating one or more settings of a camera capturing the input image;converting, by a decoder in the neural network, the latent representation of the input image into one or more pose parameters indicating a pose of the object in an image space;estimating, by a kinematic layer in the neural network using the orientation of the object and the one or more pose parameters, a three-dimensional (3D) pose of the object in a camera space;generating a graphical representation of the estimated 3D pose of the object; andoptimizing the graphical representation of the estimated 3D pose based on the one or more camera parameters.2.The method of claim 1, wherein the regression module includes an encoder that maps the input image to the latent representation of the input image in the latent space, and the latent space corresponds to a variational distribution.3.The method of claim 1, wherein optimizing the graphical representation of the estimated 3D pose of the object comprises:estimating a motion of a root point of the object based on the graphical representation of the estimated 3D pose and key points of the object, the key points identified using the input image, the key points including the root point; andadjusting the graphical representation of the estimated 3D pose to minimize an object function of the estimated motion of the root point.4.The method of claim 1, wherein optimizing the estimated 3D pose of the object comprises:projecting, by a projection layer in the neural network using the one or more camera parameters, points in the graphical representation of the estimated 3D pose from the camera space into the image space;measuring a distance between the projected points and key points of the object, the key points identified using the input image; andadjusting the graphical representation of the estimated 3D pose based on the measured distance.5.The method of claim 1, wherein optimizing the estimated 3D pose of the object comprises:optimizing the latent representation of the input image to generate a modified latent representation of the input image;measuring a similarity between the latent representation of the input image and the modified latent representation of the input image; andadjusting the graphical representation of the estimated 3D pose based on the measured similarity.6.The method of claim 1, further comprising:providing the optimized graphical representation of the estimated 3D pose for display in a graphical interface, the graphical interface allowing a user to view and modify the optimized graphical representation of the estimated 3D pose.7.The method of claim 6, further comprising:receiving a modified graphical representation of the estimated 3D pose from a client device associated with the user; andfurther training the neural network based on the input image and the modified graphical representation of the estimated 3D pose.8.The method of claim 7, wherein further training the neural network comprises:modifying one or more internal parameters of the regression module; andkeeping one or more internal parameters of the decoder fixed.9.The method of claim 1, wherein the image space has more dimensions than the latent space.10.The method of claim 1, wherein the image space is a two-dimensional space, and the camera space is a three-dimensional space.11.One or more non-transitory computer-readable media storing instructions executable to perform operations, the operations comprising:determining, by a regression module in a neural network using an input image, a latent representation of the input image in a latent space, an orientation of an object, and one or more camera parameters indicating one or more settings of a camera capturing the input image;converting, by a decoder in the neural network, the latent representation of the input image into one or more pose parameters indicating a pose of the object in an image space;estimating, by a kinematic layer in the neural network using the orientation of the object and the one or more pose parameters, a three-dimensional (3D) pose of the object in a camera space;generating a graphical representation of the estimated 3D pose of the object; andoptimizing the graphical representation of the estimated 3D pose based on the one or more camera parameters.12.The one or more non-transitory computer-readable media of claim 11, wherein the regression module includes an encoder that maps the input image to the latent representation of the input image in the latent space, and the latent space corresponds to a variational distribution.13.The one or more non-transitory computer-readable media of claim 11, wherein optimizing the graphical representation of the estimated 3D pose of the object comprises:estimating a motion of a root point of the object based on the graphical representation of the estimated 3D pose and key points of the object, the key points identified using the input image, the key points including the root point; andadjusting the graphical representation of the estimated 3D pose to minimize an object function of the estimated motion of the root point.14.The one or more non-transitory computer-readable media of claim 11, wherein optimizing the estimated 3D pose of the object comprises:projecting, by a projection layer in the neural network using the one or more camera parameters, points in the graphical representation of the estimated 3D pose from the camera space into the image space;measuring a distance between the projected points and key points of the object, the key points identified using the input image; andadjusting the graphical representation of the estimated 3D pose based on the measured distance.15.The one or more non-transitory computer-readable media of claim 11, wherein optimizing the estimated 3D pose of the object comprises:optimizing the latent representation of the input image to generate a modified latent representation of the input image;measuring a similarity between the latent representation of the input image and the modified latent representation of the input image; andadjusting the graphical representation of the estimated 3D pose based on the measured similarity.16.The one or more non-transitory computer-readable media of claim 11, wherein the operations further comprise:providing the optimized graphical representation of the estimated 3D pose for display in a graphical interface, the graphical interface allowing a user to view and modify the optimized graphical representation of the estimated 3D pose;receiving a modified graphical representation of the estimated 3D pose from a client device associated with the user; andfurther training the neural network based on the input image and the modified graphical representation of the estimated 3D pose.17.The one or more non-transitory computer-readable media of claim 11, wherein the image space is a two-dimensional space, and the camera space is a three-dimensional space.18.An apparatus, comprising:a computer processor for executing computer program instructions; anda non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations comprising:determining, by a regression module in a neural network using an input image, a latent representation of the input image in a latent space, an orientation of an object, and one or more camera parameters indicating one or more settings of a camera capturing the input image,converting, by a decoder in the neural network, the latent representation of the input image into one or more pose parameters indicating a pose of the object in an image space,estimating, by a kinematic layer in the neural network using the orientation of the object and the one or more pose parameters, a three-dimensional (3D) pose of the object in a camera space,generating a graphical representation of the estimated 3D pose of the object, andoptimizing the graphical representation of the estimated 3D pose based on the one or more camera parameters.19.The apparatus of claim 18, wherein the regression module includes an encoder that maps the input image to the latent representation of the input image in the latent space, and the latent space corresponds to a variational distribution.20.The apparatus of claim 19, wherein the operations further comprise:providing the optimized graphical representation of the estimated 3D pose for display in a graphical interface, the graphical interface allowing a user to view and modify the optimized graphical representation of the estimated 3D pose;receiving a modified graphical representation of the estimated 3D pose from a client device associated with the user; andfurther training the neural network based on the input image and the modified graphical representation of the estimated 3D pose.
Citation Information
Patent Citations
Forecasting multiple poses based on a graphical image
CN108694369A
Scene depth and camera position and posture solving method based on deep learning
CN110264526A
Two-dimensional human body posture estimation method, system and device and storage medium
CN116524541A
Three-dimensional (3D) pose estimation from a monocular camera
US20190278983A1
Neural network system for 3D pose estimation
WO2022245281A1