Using an attention model to elevate a 2D representation to 3D
By using an encoder architecture and a Transformer model to process datasets, the problem of small models in existing technologies struggling to accurately and in real-time upscale two-dimensional data to three-dimensional positions is solved, achieving high efficiency and accuracy in human joint estimation, suitable for video games and virtual reality applications.
Patent Information
- Application Number
- CN202080102235.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-31
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2040-08-31
AI Technical Summary
Existing technologies struggle to effectively elevate two-dimensional datasets to the position of three-dimensional objects in real-time using small, efficient models, particularly in the three-dimensional position estimation of human joints, where there is a trade-off between model size and accuracy.
An encoder architecture with multiple encoder layers, including a self-attention mechanism and a feedforward network, is employed. The dataset is processed by a Transformer model to improve the estimation accuracy of the 3D state, and the dimensionality of the dataset is adjusted by one-dimensional convolution to match the model input and output.
It achieves efficient and accurate upscaling of 2D human keypoints to 3D positions using a small model in real time, outperforming existing methods. The model is smaller but performs better, providing accurate human joint position estimation in video games and virtual reality applications.
Smart Images

Figure CN115917597B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to estimating the three-dimensional position of an object from a two-dimensional data set at a processing device. Background Art
[0002] In recent years, estimating the three-dimensional (3D) position of human joints has become a widely studied topic. Particular attention has been devoted to defining methods for extrapolating two-dimensional data (in the form of x, y keypoints) to 3D to predict the root-relative coordinates of joints associated with the human skeleton. The human skeleton is described by 17 keypoints, including the head, shoulder, elbow, wrist, pelvis, knee, and ankle.
[0003] Initially, the work was mainly based on pre-designed models with a large number of constraints to account for the high dependencies between different human joints.
[0004] With the development of Convolutional Neural Networks (CNNs), pose estimators have been developed that reconstruct 3D pose end-to-end directly from RGB images or intermediate two-dimensional (2D) predictions. This approach has rapidly surpassed the accuracy of previous hand-crafted estimators.
[0005] Current state-of-the-art methods generally fall into two main approaches. Some methods predict 3D keypoints directly from monocular images end-to-end. This often produces good results but requires very large models. Other methods perform lifting, where a 2D predictor is used to predict the human pose from the image, and then the 2D keypoints (relative to the image) are expanded to 3D. This two-step approach typically uses a temporal convolution layer to aggregate information from the pose derived from the video.
[0006] There is a need to develop a method to lift the 2D projection of an object's joints to 3D, which overcomes the limitations of existing methods by estimating depth by extrapolating input data at real-time using a small and efficient model. Summary of the Invention
[0007] According to one aspect, a processing device is provided for forming a system for estimating a three-dimensional state of one or more jointed objects represented by multiple data sets, each data set indicating a projection of a joint of the object onto a two-dimensional space, the processing device comprising one or more processors configured to: receive a data set; process the data set using an encoder architecture having multiple encoder layers, each encoder layer comprising a respective self-attention mechanism and a respective feedforward network, each self-attention mechanism implementing multiple attention heads; and train the encoder layers based on the data set to improve the accuracy of the encoder architecture in estimating the three-dimensional state of one or more objects represented by the data set.
[0008] This can allow a processing device to estimate depth by extrapolating input data while running in real time using a small, efficient model, thereby forming a system for lifting a 2D dataset comprising, for example, 2D human body key points (x, y coordinates) to 3D (x, y, z coordinates) relative to a root (such as a pelvic joint).
[0009] The object or each object may be a person and each data set may indicate a human pose. Each data set may include a plurality of key points of the human body (each having 2D coordinates). Typically, there are 17 key points describing the human skeleton, including the head, shoulders, elbows, wrists, pelvis, knees, and ankles. Each data set indicating a human pose may define 2D coordinates for each of the plurality of key points. This may allow estimation of the 3D position of human joints, which may be useful in video games or virtual reality applications.
[0010] The encoder architecture can implement the Transformer architecture. In the task of lifting 2D keypoints to 3D, using the Transformer encoder model produces good accuracy.
[0011] The operation of each self-attention mechanism can be defined by a set of parameters, and the device can be configured to share one or more such parameters between multiple self-attention mechanisms. Weight sharing can make the model more efficient while maintaining (or in some cases improving) the accuracy of predicting the 3D state of the object.
[0012] The device can be configured to adjust one or more parameters in a self-attention mechanism in response to a continuous data set, and after adjusting the parameters, propagate the parameters to one or more other self-attention mechanisms. Thus, the parameters of one or some attention layers can be shared with other attention layers. This can further improve the efficiency of the model.
[0013] The operation of each encoder layer can be defined by a set of parameters, and the device can be configured to adjust one or more of these parameters in an encoder layer in response to successive data sets, and the device can be configured not to propagate such parameters to any other encoder layer. Thus, in some embodiments, the parameters of an encoder layer may not be shared with other encoder layers. This can further improve the model.
[0014] The one or more processors can be configured to perform a one-dimensional convolution on the dataset to form convolution data and process the convolution data using an encoder architecture. This can allow the dimensions of the dataset to be adjusted to match the dimensions of the model input and output.
[0015] One or more processors can be configured to perform a one-dimensional convolution on a series of consecutive data sets. This can allow the 3D state of an object to be estimated for a 2D input sequence. For example, a 3D motion sequence of a human body can be estimated.
[0016] The series of datasets can be an odd number of datasets. This allows datasets on either side of a central dataset to be considered during training to achieve a three-dimensional state.
[0017] The device can be configured to estimate the 3D state of an intermediate dataset in the series of datasets based on the series of datasets. The intermediate dataset in the series of datasets can correspond to the center of an original receptive field (the total number of poses used to predict a 3D pose). During training, half of the receptive field corresponds to past poses and the other half corresponds to future poses. The intermediate pose within the receptive field can be the pose currently being lifted from 2D to 3D.
[0018] The one or more processors can be configured to train an encoder architecture to lift the dataset into three dimensions. Training the encoder can allow the model to more accurately predict the 3D state of the object.
[0019] Each dataset can represent the position of a human joint relative to a predetermined joint or structure in the body. This structure can be the pelvis. The position of the pelvis can serve as the root. A separate model can then be used to determine the distance from the camera to the pelvis. This allows the position of each joint to be determined relative to the camera.
[0020] The processing device may be configured to receive a plurality of images, each image representing an articulated object; and for each image, detect joint positions of the object in the image, thereby forming one of the datasets. A 2D pose estimator may be used to obtain an accurate 2D pose of a human body in the plurality of images. This may allow 2D poses to be predicted from the plurality of images, which may then be used as input to the dataset of the apparatus described above to enhance the 2D dataset to 3D.
[0021] Once trained, the system can be used during the inference phase to estimate the three-dimensional state of one or more articulated objects represented by multiple such datasets.
[0022] According to another aspect, there is provided a system for estimating a three-dimensional state of one or more articulated objects, the system being formed by the processing device described above.
[0023] According to another aspect, a method is provided for estimating a three-dimensional state of one or more joint objects represented by multiple data sets, each data set indicating a projection of a joint of the object onto a two-dimensional space, the method comprising: receiving the data sets; processing the data sets using an encoder architecture having multiple encoder layers, each encoder layer comprising a respective self-attention mechanism and a respective feedforward network, each self-attention mechanism implementing multiple attention heads; and training the encoder layers based on the data sets to improve the accuracy of the encoder architecture in estimating the three-dimensional state of one or more objects represented by the data sets.
[0024] The method may also include estimating a three-dimensional state of one or more joint objects represented by a plurality of such data sets.
[0025] Use of this approach can allow estimating depth from a 2D dataset consisting of, for example, 2D human keypoints (x, y coordinates) to 3D (x, y, z) relative to a root (such as a pelvic joint) by extrapolating the input data at runtime using a small, efficient model. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The present invention will now be described by way of examples with reference to the accompanying drawings.
[0027] In the attached figure:
[0028] Figure 1 Shows an overview of the model architecture, which takes as input a dataset sequence of 2D keypoints of human poses and produces 3D pose estimates using self-attention on long-term information.
[0029] Figure 2 A processing device is shown for forming a system for estimating a three-dimensional state of one or more articulated objects represented by a plurality of data sets.
[0030] Figure 3 An exemplary flow chart detailing method steps performed by a processing device is shown.
[0031] Figure 4Qualitative results are shown for several actions using human poses from the Humans3.6M dataset: (a) original RGB image with 2D keypoint prediction using CPN; (b) 3D reconstruction using the method described in this paper (n=243, where n is the receptive field of the pose in the input sequence); (c) ground truth 3D keypoints. DETAILED DESCRIPTION
[0032] The method described herein is exemplified by processing datasets, each of which indicates human poses. However, it should be understood that the method can also be applied to other datasets and objects where data needs to be converted from 2D to 3D.
[0033] Typically, there are 17 key points that describe the human skeleton, including the head, shoulders, elbows, wrists, pelvis, knees, and ankles. In the examples described herein, each data set indicating a human posture defines 2D coordinates for each of these key points. Preferably, each data set indicating a human posture represents the position of a human joint relative to a predetermined joint or structure of the human body. In a preferred embodiment, the structure is the human pelvis.
[0034] In the examples described herein, individual datasets (eg, individual human poses) that are input to the model can be derived from a larger motion capture dataset that includes images depicting different human poses. Such larger datasets may include motion capture datasets such as Human3.6M (see C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu, “Human3.6M: Large scale datasets and predictive methods for 3d human sensing in natural environments,” IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), 7 (2013), pp. 1325-1339) or HumanEva (see L. Sigal, A.O. Balan, and M.J. Black, “HumanEva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion,” International Journal of Computer Vision (IJCV) (2010), 87(1-2): 4). Human3.6M includes 3.6 million frames from 11 different subjects, but only seven of them are annotated. The subjects performed up to 15 different types of actions, which were recorded from four different angles. In contrast, HumanEva is a smaller motion capture dataset with only three subjects and recorded from three angles.
[0035] The method described herein may include a two-step pose estimation approach. The data set indicating human poses that is input to the model may first be obtained by using a 2D pose estimator to obtain accurate 2D poses of the human body from images. This may be done in a top-down manner. These joints may then be lifted by predicting their depth relative to the root (e.g., the pelvis).
[0036] Thus, the processing device may be configured to receive a plurality of images, each image representing a jointed object. For each image, the processing device may then detect the locations of the joints of the human body (or other object) in the image, thereby forming one of the data sets indicating a single posture.
[0037] For example, the 2D pose can be obtained by using a ground truth human bounding box and then using a 2D pose estimator. Some common 2D estimation models that can be used to obtain a 2D pose sequence are: Stacked Hourglass (SH), as described by A. Newell, K. Yang, and J. Deng in "Stacked hourglass networks for human pose estimation" (Proceedings of the IEEE European Conference on Computer Vision (ECCV) (2016), pp. 483–499); Mask RCNN (Mask-RCNN), as described by K. He, G. Gkioxari, P. Dollar, and R. Girshick in "Mask RCNN (Mask-RCNN)" (Proceedings of the IEEE International Conference on Computer Vision (ICCV) (2017), pp. 2961–2969); or Cascaded Pyramid Networks (Cascaded Pyramid Networks). The Cascaded Pyramid Network (CPN) is a CNN-based neural network that uses a CNN to perform multi-person pose estimation. The Cascaded Pyramid Network (CPN) is described in Y. Chen, Z. Wang, Y. Peng, Z. Zhang, G. Yu, and J. Sun, “Cascaded Pyramid Network for Multi-person Pose Estimation” (Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018), pp. 7103–7112).
[0038] Alternatively, a 2D detector that does not rely on ground truth can be used. For example, SH and CPN can be used as detectors for the Human3.6M motion capture dataset, and Mask-RCNN can be used as a detector for the HumanEva motion capture dataset. However, it is also possible to use ground truth 2D poses for training. In one specific example, SH can be pre-trained on the MPII motion capture dataset (L. Pishchulin, E. Insafutdinov, S. Tang, B. Andres, M. Andriluka, P. Gehler, and B. Schiele, “DeepCut: Joint Subset Partition and Labeling for MultiPerson Pose Estimation”), following the approach described in D. Pavllo, C. Feichtenhofer, D. Grangier, and M. Auli, “3D human pose estimation in video with temporal convolutions and semi-supervised training” (Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2019), pp. 7753–7762). Estimation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016), pp. 4929–4937), and fine-tuned on the Human3.6M motion capture dataset. Both Mask-RCNN and CPN can be pre-trained on COCO (T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and CL Zitnick, “Microsoft COCO: Common objects in context,” in Proceedings of the IEEE European Conference on Computer Vision (ECCV) (2014), pp. 740–755), and then fine-tuned on the 2D poses of Human3.6M, since keypoints are defined differently in each motion capture dataset. More specifically, Mask-RCNN uses ResNet-101 with FPN. Since CPN requires bounding boxes, Mask-RCNN can be used first to detect people. It can then determine keypoints from the image using ResNet-50 with an input resolution of 384×384.
[0039] In the example described in this article, an open-source Transformer model is used to improve a keypoint dataset. In this example, a 2D keypoint sequence is passed through several Transformer encoder blocks to generate a 3D pose prediction (corresponding to the pose at the center of the input sequence / receptive field). Further 3D estimates can be calculated in a sliding window manner.
[0040] Figure 1 Shows an overview of the model architecture, which takes as input a dataset sequence of 2D keypoints of human poses and produces 3D pose estimates using self-attention on long-term information.
[0041] As shown at 101, the device takes as input a sequence of datasets including 2D keypoints (the input sequence is collectively referred to as a receptive field). The method can be used for different numbers of datasets. The series of datasets is preferably a series of an odd number of datasets. For example, receptive fields including datasets with 27, 81, and 243 keypoints can be used. Preferably, the middle dataset of the sequence is selected for boosting because it corresponds to the center of the original receptive field.
[0042] As mentioned above, the dataset corresponding to the input sequence can come from a 2D predictor that estimates the 2D pose from image frames, or directly from the 2D ground truth. Therefore, these input datasets are in image space.
[0043] like Figure 1 As shown, in some embodiments, certain modifications may be performed to match the dimensions of the model's input and output. In this example, the sequence of input datasets is passed through the convolutional layer 102 to change the dimensions. In this example, the input to the Transformer is reprojected from the input dimensions [B, N, 34] to [B, N, 512], where B is the batch size, N is the receptive field (i.e., the number of human poses input to the model in each processing step), and 34 corresponds to 17 joints times 2 (the number of coordinates, i.e., x and y).
[0044] Then, a temporal encoding is added to embed the order of the input sequence of the dataset (poses). As shown at 103, the temporal encoding is aggregated by adding a vector to the input embedding. This allows the model to exploit the order of the pose sequence. This can also be called a positional encoding, which is used to inject information about the relative or absolute position of a token in the sequence. These temporal embeddings can be created using sine and cosine functions and then added as the input to the reprojection. In this embodiment, the injected temporal embedding has the same dimension as the input.
[0045] The temporally embedded dataset is then fed into a Transformer encoder model that processes it. Figure 1 As shown in 104.
[0046] The self-attention model is used in Natural Language Processing (NLP) to embed gestures over time, rather than using a fully convolutional approach. The encoder architecture has multiple encoder layers, each of which includes its own self-attention mechanism 107 and its own feedforward network 109. Each self-attention mechanism implements multiple attention heads. The outputs from blocks 107 and 109 can be summed and normalized, as shown at 108 and 110.
[0047] A.Vaswani, N.Shazeer, N.Parmar, J.Uszkoreit, L.Jones, ANGomez, The basic Transformer encoder described by I. Kaiser and I. Polosukhin in “Attention is all you need” (IEEE Advances in Neural Information Processing Systems (NeurIPS) (2017), pp. 5998–6008) can be used as a baseline. In this example, there are 512 hidden layers, 8 multi-attention heads, and 6 encoder blocks.
[0048] Since the decoder portion of the Transformer is not used in this implementation, and due to the residual connections within the self-attention, the dimensions of the output are the same as the input, i.e. [B, N, 512] in this example.
[0049] The 1D convolutional layer 105 changes the dimensionality so that the output of the model (as shown at 106) is the x, y, and z coordinates of all joints relative to the pelvis for the pose currently being lifted.
[0050] The output embedding is reprojected to the required dimension using a 1D convolutional layer 105, from [B, 1, 512] to [B, 1, 51], where 51 corresponds to 17 joints multiplied by 3 (for x, y, and z coordinates). The loss is then calculated, for example using the mean-per-joint position error (MPJPE) against the dataset's 3D ground truth. The error is then backpropagated and the next training iteration begins.
[0051] Preferably, the model output labels (the current pose lifted from 2D to 3D within the pose input sequence) correspond to an intermediate dataset (intermediate poses) of the receptive field. This is because during training, half of the receptive field corresponds to past poses and the other half to future poses. Therefore, during training, the model utilizes temporal data from both past and future frames to be able to create temporally consistent predictions.
[0052] During inference, the model architecture remains the same as during training, but only past pose frames are used in the receptive field, rather than the frames on either side of the current pose being lifted as used during training. The model works in a sliding window manner so that it can eventually obtain 3D reconstructions for all 2D representations in the input sequence.
[0053] One benefit of this architecture is that the receptive field and the number of attention heads can be modified without affecting the model size. In addition, the hyperparameters of the Transformer can also be modified.
[0054] In some embodiments, weight sharing can be used to maintain or improve final accuracy while significantly reducing the total number of parameters, thereby building a more efficient model.
[0055] Optionally, parameters can be shared across each encoder block. Specifically, attention layer parameters can be shared.
[0056] In one embodiment of shared attention layer parameters, the operation of each self-attention mechanism is defined by a set of parameters, and one or more of these parameters are shared between multiple self-attention mechanisms. One or more of the parameters in a self-attention mechanism can be adjusted in response to successive data sets, and after adjusting a parameter, the parameter can be propagated to one or more other self-attention mechanisms.
[0057] Alternatively or additionally, the operation of each encoder layer can be defined by a set of parameters. The device can be configured to adjust one or more of these parameters in an encoder layer in response to successive data sets. In one embodiment, the device can be configured not to propagate such parameters to any other encoder layer.
[0058] In particular, sharing only the attention layer parameters (and not the encoder block parameters) can improve the final accuracy while significantly reducing the total number of parameters.
[0059] In some embodiments, additional data augmentation can be applied to the dataset during training and testing. For example, each pose can be flipped horizontally.
[0060] During training of the attention layer of the Transformer model, an optimizer such as the Adam optimizer (as described in S.J. Reddi, S. Kale, and S. Kumar, “On the convergence of Adam and beyond” (Proceedings of the International Conference on Learning Representation (ICLR) (2018)) can be used. For example, a training run can last for 80 epochs on the Human3.6M motion capture dataset and 1000 epochs on the HumanEva motion capture dataset.
[0061] The learning rate can be increased linearly for the first training steps (e.g., 1000 iterations with a learning rate factor of 12), and then decreased proportionally to the inverse square root of the number of steps. This is often called NoamOpt.
[0062] The batch size can be proportional to the receptive field value. For example, when n=27, 81, and 243, the values are b=5120, 3072, and 1536, respectively.
[0063] In terms of hardware, the system can be trained and evaluated using eight NVIDIA V100 GPUs, with parallel optimization. Typically, considering the batch size, the training time per receptive field can be approximately 8, 14, and 40 hours, respectively.
[0064] Figure 2 2 is a schematic representation of a system 200 configured to perform the methods described herein. The system 200 may be implemented on a device such as a laptop, a tablet, a smartphone, or a television (TV).
[0065] The system 200 includes a processor 201 configured to process a data set in the manner described herein. For example, the processor 201 may be implemented as a computer program running on a programmable device such as a graphics processing unit (GPU) or a central processing unit (CPU). The system 200 includes a memory 202 arranged to communicate with the processor 201. The memory 202 may be a non-volatile memory. The processor 201 may also include a cache ( Figure 2(not shown) which can be used to temporarily store data from memory 202. The system may include more than one processor and more than one memory. The memory may store data executable by the processor. The processor may be configured to operate according to a computer program stored in a non-transitory form on a machine-readable storage medium. The computer program may store instructions for causing the processor to perform its method in the manner described herein.
[0066] Figure 3 A flowchart summarizing an example of a method for estimating a three-dimensional state of one or more articulated objects represented by multiple data sets (e.g., multiple human poses), each data set indicating a projection of a joint of the object (e.g., a human body) onto a two-dimensional space is shown.
[0067] At step 301, the method includes receiving a dataset. At step 302, the method includes processing the dataset using an encoder architecture having a plurality of encoder layers, each encoder layer including a respective self-attention mechanism and a respective feed-forward network, each self-attention mechanism implementing a plurality of attention heads. At step 303, the method includes training the encoder layers based on the dataset to improve the accuracy of the encoder architecture in estimating the three-dimensional state of one or more objects represented by the dataset.
[0068] This paper describes the use of a self-attention Transformer model to estimate depth from 2D keypoints. The encoder's self-attention architecture allows the model to produce temporally consistent poses by leveraging long-range temporal information across frames / poses.
[0069] In some embodiments, it has been found that the methods described herein can provide better results and allow for smaller model sizes than previous methods.
[0070] For input 2D predictions (Mask-RCNN and CPN), the method described in this paper was found to outperform previous improvement methods and perform comparable to methods using keypoints and features extracted from raw RGB images. For ground truth input, the model was found to outperform previous models, achieving results comparable to Skinned Multi-Person Linear Model (SMPL) or multi-view methods that simultaneously predict body shape and pose. The number of parameters in the model is easy to adjust and can be smaller (e.g., 9.5 million) than current methods (which can have around 11-17 million parameters) while still achieving better performance. Therefore, compared to the state of the art, this method can achieve better results with a smaller model size.
[0071] Figure 4Qualitative results are shown for several actions using human poses from the Humans3.6M dataset. Column (a) shows the original RGB image with 2D keypoint prediction using the CPN. Column (b) shows the 3D reconstruction (with a receptive field of n=243) using the method described in this paper. Column (c) shows the ground truth 3D keypoints. As can be seen, the obtained 3D reconstruction closely matches the ground truth 3D keypoints.
[0072] Therefore, the method described in this paper allows estimating depth by extrapolating input data at runtime using a small and efficient model, lifting 2D human keypoints (x, y coordinates) to 3D (x, y, z) relative to the root (such as the pelvic joint).
[0073] Applicants hereby disclose separately each individual feature described herein and any combination of two or more such features, provided that such feature or combination can be implemented based on the present specification as a whole according to the common general knowledge of a person skilled in the art, regardless of whether such feature or combination of features solves any problem disclosed herein, and without limiting the scope of the claims. Applicants indicate that aspects of the present invention may include any such individual feature or combination of features. In view of the above description, it will be apparent to those skilled in the art that various modifications can be made within the scope of the present invention.
Claims
1. A processing device (200), characterized in that For forming a system for estimating a three-dimensional state of one or more jointed objects represented by a plurality of data sets (101), each data set indicating a projection of a joint of the object onto a two-dimensional space, the processing device comprising one or more processors (201) configured to: receiving (301) said data set; Performing a one-dimensional convolution (102) on the data set to adjust the dimension of the data set; Adding time coding to the dimensionally adjusted dataset to embed the order of the input sequence of the dataset; processing (302) the temporally encoded dataset using an encoder architecture (104) having a plurality of encoder layers, each encoder layer including a respective self-attention mechanism (107) and a respective feed-forward network (109), each self-attention mechanism implementing a plurality of attention heads; as well as training (303) the encoder layers based on the dataset to improve the accuracy of the encoder architecture in estimating the three-dimensional state (106) of one or more objects represented by the dataset; wherein the operation of each self-attention mechanism (107) is defined by a set of parameters, and the device is configured to share one or more of the parameters between multiple self-attention mechanisms.
2. The processing device (200) according to claim 1, characterized in that Each object is a person and each dataset indicates a human pose.
3. The processing device (200) according to claim 1, characterized in that The encoder architecture (104) implements a Transformer architecture.
4. The processing device (200) according to claim 1, characterized in that The apparatus is configured to adjust one or more of the parameters in a self-attention mechanism (107) in response to a continuous data set, and after adjusting the parameters, propagate the parameters to one or more other self-attention mechanisms.
5. The processing device (200) according to claim 4, characterized in that The operation of each encoder layer is defined by a set of parameters, the apparatus is configured to adjust one or more of the parameters in an encoder layer in response to successive data sets, and the apparatus is configured not to propagate the parameters to any other encoder layer.
6. The processing device (200) according to claim 1, characterized in that The one or more processors (201) are configured to perform the one-dimensional convolution (102) on a series of consecutive data sets (101).
7. The processing device (200) according to claim 6, characterized in that The series of data sets (101) is a series of an odd number of data sets.
8. The processing device (200) according to claim 7, characterized in that The device is configured to estimate, based on the series of data sets (101), a three-dimensional state of an intermediate data set in the series of data sets.
9. The processing device (200) according to any one of claims 1 to 5, characterized in that The one or more processors (201) are configured to train the encoder architecture (104) to lift the dataset to three dimensions.
10. The processing device (200) according to claim 9, characterized in that Each dataset represents the position of a human joint relative to a predetermined joint or structure of the human body.
11. The processing device (200) according to claim 10, characterized in that The structure is the pelvis.
12. The processing device (200) according to any one of claims 1 to 5, characterized in that The processing device is configured to: receiving a plurality of images, each image representing an articulated object; and For each image, joint positions of the object in the image are detected, thereby forming one of the data sets (101).
13. A system for estimating the three-dimensional state of one or more joint objects, characterized in that The system is formed by a processing device (200) as claimed in any one of the preceding claims.
14. A method (300) for estimating a three-dimensional state of one or more articulated objects represented by a plurality of data sets (101), characterized in that Each data set indicates a projection of a joint of the object in a two-dimensional space, and the method comprises: receiving (301) said data set; Performing a one-dimensional convolution (102) on the data set to adjust the dimension of the data set; Adding time coding to the dimensionally adjusted dataset to embed the order of the input sequence of the dataset; processing (302) the temporally encoded dataset using an encoder architecture (104) having a plurality of encoder layers, each encoder layer including a respective self-attention mechanism (107) and a respective feed-forward network (109), each self-attention mechanism implementing a plurality of attention heads; and training (303) the encoder layers based on the dataset to improve the accuracy of the encoder architecture in estimating the three-dimensional state (106) of one or more objects represented by the dataset; Wherein the operation of each self-attention mechanism (107) is defined by a set of parameters, and one or more of said parameters are shared among multiple self-attention mechanisms.
Citation Information
Patent Citations
Voice recognition method and device, medium and equipment
CN110797018A