Modular pipeline for high fidelity hand-arm motion synthesis and multi-view rendering

By generating diverse hand and arm pose datasets through a biconditional variational autoencoder and combining arm mesh templates with hand mesh models, the problem of lack of motion modularity and semantic meaning in existing databases is solved, thereby improving high-fidelity hand pose estimation and pose recognition.

CN120853249APending Publication Date: 2025-10-28SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510529947.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-04-10
Filing Date
2025-04-25
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing synthetic hand pose databases lack motion modularity, semantic meaning, and data variability, making it impossible to effectively train and test 3D hand pose estimation and hand pose recognition systems, and they also lack complete hand-arm dynamic coordination.

Method used

A dual-conditional variational autoencoder architecture is used to generate finger poses and wrist movements separately. Diverse hand and arm pose datasets are generated by combining Cartesian products. The arm mesh template is combined with the hand mesh model using a cut-and-stitch method to generate a high-fidelity hand-arm mesh model, which is then captured using a real-world camera configuration.

Benefits of technology

It generates a diverse and flexible hand pose database, improves the training effect of hand-related models, enhances the accuracy of hand pose estimation and pose recognition, provides complete hand-arm dynamics coordination, and reduces data capture and annotation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853249A_ABST
    Figure CN120853249A_ABST
Patent Text Reader

Abstract

A computer-implemented method of generating a synthetic dataset of hand-arm gestures includes: generating a set of finger gestures from a first conditional variational auto-encoder including a first potential space and a first transducer decoder; generating a set of wrist motions from a second conditional variational auto-encoder including a second potential space and a second transducer decoder; and combining the set of finger gestures and the set of wrist motions to generate a synthetic dataset of hand-arm gestures.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority and benefit to U.S. Provisional Application No. 63 / 639339, filed April 26, 2024, and U.S. Non-Provisional Application No. 19 / 175941, filed April 10, 2025, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to hand gesture recognition and hand-arm mesh models. Background Technology

[0004] Hand pose databases are crucial for addressing R&D needs in extended reality (XR), human-computer interaction (HCI), and other fields requiring data to train and evaluate hand-related models. Synthetic hand pose databases are less costly than 3D capture and annotation of real-world data. However, dynamic hand pose datasets in related technologies often constrain poses to fixed combinations of global wrist movements and specific finger postures (i.e., rigid definitions of hand poses and a lack of motion modularity). Some synthetic hand pipelines can focus on constrained 3D hands with random poses at limited viewpoints. Consequently, some synthetic hand datasets may lack semantically meaningful poses, motion dynamics, and data variability. For example, in some systems, simple wrist movements (such as moving a fist left, right, up, or down) can be treated as distinct, unrelated poses. This rigid definition may fail to capture the semantic meaning of hand movements and their potential variability and flexibility. Therefore, some synthetic hand datasets may lack sufficient variation in hand shape, pose, dynamics, and viewpoint to robustly train and test 3D hand pose estimation (HPE) and hand gesture recognition (HGR) systems.

[0005] Additionally, some hand pose databases may lack complete hand-arm dynamics (i.e., the dataset may lack realistic coordination between the forearm, wrist, and fingers). For example, unless specifically designed for a very constrained 3D model, the forearm may not be dynamically aligned with the hand.

[0006] The information disclosed in the background section is only for enhancing the understanding of the background art of the present invention, and therefore may contain information that does not constitute prior art. Summary of the Invention

[0007] This disclosure relates to various embodiments of a computer-implemented method for generating a synthetic dataset of hand and arm poses. In one embodiment, the method includes: generating a set of finger poses according to a first conditional variational autoencoder including a first latent space and a first transformer decoder; generating a set of wrist movements according to a second conditional variational autoencoder including a second latent space and a second transformer decoder; and combining the set of finger poses and the set of wrist movements to generate a synthetic dataset of hand and arm poses.

[0008] This combination can include the Cartesian product of the set of finger gestures and the set of wrist movements.

[0009] The first conditional variational autoencoder can be different from the second conditional variational autoencoder.

[0010] The first transformer decoder can have eight layers, and the second transformer decoder can have two layers.

[0011] The method may also include generating a hand-arm mesh model of the hand-arm pose.

[0012] This set of finger gestures may include at least one digital gesture, at least one trigger gesture, and at least one special gesture.

[0013] This disclosure also relates to various embodiments of a computer-based method for generating hand-arm mesh models. In one embodiment, the method includes: generating a hand mesh model; and combining an arm mesh model with a hand mesh model. Combining the arm mesh model with the hand mesh model includes: identifying wrist boundary vertices of the hand mesh model and the arm mesh model; ensuring that the number of wrist boundary vertices of the hand mesh model is equal to the number of wrist boundary vertices of the arm mesh model; and applying a wrist rotation matrix to the hand mesh model.

[0014] The method may also include removing the overlapping surfaces between the hand mesh model and the arm mesh model at the wrist of the hand-arm mesh model.

[0015] The method may also include interpolation at the wrist between the hand mesh model and the arm mesh model to prevent visual seams between the hand mesh model and the arm mesh model.

[0016] The method may also include: applying skin texture to the hand mesh model; and propagating the skin texture of the hand mesh model to the arm mesh model.

[0017] The hand mesh model can be a NIMBLE model.

[0018] The arm mesh model can be an SMPL-X model.

[0019] This method may include applying a global transformation to the hand-arm mesh model.

[0020] Generating a hand mesh model can include converting a MANO hand model into a NIMBLE hand model.

[0021] The hand mesh model can be a Handy model.

[0022] This disclosure also relates to various embodiments simulating real-world camera configurations. The method may include: arranging a camera in a hemispherical configuration around a hand-arm mesh model; and capturing hand movements of the hand-arm mesh model from different perspectives using the camera.

[0023] Cameras can include still cameras.

[0024] Cameras can include dynamic cameras.

[0025] The dynamic camera may include a first camera with a close-up lens facing the palm side of the hand-arm mesh model, and a pair of stereo cameras facing the back of the hand-arm mesh model.

[0026] The method may further include generating a hand-arm mesh model, which may include: generating a hand mesh model by a processor; and combining an arm mesh model with a hand mesh model by a processor to generate a hand-arm mesh model. Combining the arm mesh model with the hand mesh model may include: identifying wrist boundary vertices of the hand mesh model and the arm mesh model by a processor; controlling the number of wrist boundary vertices of the hand mesh model to be equal to the number of wrist boundary vertices of the arm mesh model by a processor; and applying a wrist rotation matrix to the hand mesh model by a processor.

[0027] This summary is provided to introduce a series of concepts further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. One or more described features may be combined with one or more other described features to provide a feasible method or apparatus. Attached Figure Description

[0028] The features and advantages of embodiments of this disclosure will be better understood by referring to the following detailed description when considered in conjunction with the accompanying drawings. In the drawings, the same reference numerals are used throughout to refer to the same features and components. The drawings are not necessarily drawn to scale.

[0029] Figure 1 This is a flowchart illustrating a task of generating a high-fidelity synthetic dataset of hand and arm poses according to an embodiment of the present disclosure;

[0030] Figure 2 A conditional variational autoencoder (CVAE) for synthesizing various three-dimensional finger pose sequences and wrist motion sequences is described according to an embodiment of the present disclosure;

[0031] Figure 3 The present disclosure describes finger gestures, wrist movements, and the Cartesian product of these finger gestures and wrist movements according to one embodiment.

[0032] Figure 4 This is a flowchart illustrating a task of a "cut-and-stitch" method for generating a hand-arm mesh model according to an embodiment of the present disclosure;

[0033] Figures 5A to 5E A cut-and-stitch method for generating a hand-arm mesh model according to an embodiment of the present disclosure is described;

[0034] Figures 6A to 6D Simulations of a near-static camera, a dynamic camera, a distant static camera, and a combination of a near-static camera, a dynamic camera, and a distant static camera, according to embodiments of the present disclosure, are described respectively.

[0035] Figures 7A to 7B These are perspective views and schematic block diagrams of a virtual reality and / or augmented reality (VR / AR) system according to an embodiment of the present disclosure; and

[0036] Figure 8 This is an overview of a hand-arm assembly line according to one embodiment of the present disclosure. Detailed Implementation

[0037] In the following, exemplary embodiments will be described in more detail with reference to the accompanying drawings, wherein the same reference numerals throughout refer to the same elements. However, the invention may be embodied in a variety of different forms and should not be construed as limited to the embodiments shown herein. Rather, these embodiments are provided by way of example so that this disclosure will be comprehensive and complete, and will fully convey aspects and features of the invention to those skilled in the art. Accordingly, processes, elements, and techniques that are not essential for a full understanding of the aspects and features of the invention by those skilled in the art may not be described. Unless otherwise stated, the same reference numerals denote the same elements throughout all drawings and written description, and therefore their description will not be repeated.

[0038] In the accompanying drawings, for clarity, the relative dimensions of elements, layers, and regions may be exaggerated and / or simplified. For ease of interpretation, spatial relative terms such as “below,” “below,” “under,” “below,” “above,” and “above” are used herein to describe the relationship of an element or feature to other elements or features shown in the drawings. It will be understood that, in addition to the orientations depicted in the drawings, spatial relative terms are also intended to cover different orientations of the device in use or operation. For example, if the device in the drawings is flipped, an element described as “below” or “below” to other elements or features will be oriented “above” to those other elements or features. Thus, the example terms “below” and “below” can cover both above and below orientations. The device may be oriented in other ways (e.g., rotated 90 degrees or in other orientations), and the spatial relative descriptors used herein should be interpreted accordingly.

[0039] It will be understood that although the terms “first,” “second,” “third,” etc., may be used herein to describe various elements, components, regions, layers, and / or segments, these elements, components, regions, layers, and / or segments should not be limited by these terms. These terms are used to distinguish one element, component, region, layer, or segment from another element, component, region, layer, or segment. Therefore, without departing from the spirit and scope of the invention, the first element, component, region, layer, or segment described below may be referred to as the second element, component, region, layer, or segment.

[0040] It will be understood that when a component or layer is referred to as "on another component or layer," "connected to another component or layer," or "coupled to another component or layer," it can be directly on, directly connected to, or directly coupled to another component or layer, or there can be one or more intermediate components or layers. Furthermore, it will be understood that when a component or layer is referred to as "between" two components or layers, it can be the only component or layer between those two components or layers, or there can be one or more intermediate components or layers.

[0041] The terminology used herein is for describing particular embodiments and is not intended to limit the invention. As used herein, the singular forms “a” and “an” are also intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that the terms “comprising,” “including,” “covering,” and “containing” as used in this specification designate the presence of the stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or combinations thereof. As used herein, the term “and / or” includes all combinations of any one and one or more of the associated listed items. Expressions such as “at least one of…” modify the entire list of elements when placed after the list of elements, rather than the individual elements in the list.

[0042] As used herein, the terms “substantially,” “approximately,” and similar terms are used as approximate terms rather than terms of degree, and are intended to account for inherent variations in measured or calculated values ​​that would be recognized by those skilled in the art. Furthermore, the use of “may” in describing embodiments of the invention refers to “one or more embodiments of the invention.” As used herein, the terms “use,” “in use,” and “being used” can be considered synonymous with the terms “utilizing,” “being utilized,” and “being exploited,” respectively. Additionally, the term “exemplary” is intended to refer to an example or illustration.

[0043] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It will be further understood that terms such as those defined in common dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the relevant technology and / or this specification, and shall not be interpreted in an idealized or overly formal sense, unless expressly defined herein.

[0044] This disclosure relates to various embodiments of a method for generating high-fidelity synthetic datasets of hand and arm poses using a stage-aware conditional variational autoencoder (CVAE) framework. The hand and arm pose datasets can be used for applications such as hand pose inspection (HPE), hand pose recognition (HGR), extended reality (XR), human-computer interaction (HCI), or other domains that require real synthetic data to train and evaluate hand-related models. Configuring the generation and utilization of synthetic hand pose datasets, including semantically meaningful poses, motion dynamics, and data variations, to improve the training of hand-related models (e.g., hand-related models for HPE, HGR, XR, or HCI applications) is less costly than 3D capture and annotation of real-world data (e.g., capturing real-world hand poses using a camera array and then annotating these poses before using the captured images to train hand-related models). In one or more embodiments, the generation of the synthetic dataset is split into two CVAE streams (i.e., a dual CVAE architecture), one CVAE stream for finger poses (i.e., local poses) and the other CVAE stream for wrist movements (i.e., global movements). The two CVAE streams are combined via a Cartesian product to create a diverse and flexible hand pose database that includes pose categories beyond those of existing related technology datasets.

[0045] This disclosure also relates to various embodiments of a cut-and-stitch method for generating a hand-arm mesh model by stitching an arm mesh template (e.g., an SMPL-X arm mesh model) to a NIMBLE hand mesh model. In one or more embodiments, the cut-and-stitch method is configured to achieve dynamic hand articulation while maintaining attachment stability of the arm. Additionally, in one or more embodiments, the cut-and-stitch method is configured to minimize (or at least reduce) the visual seam between the hand mesh model and the arm model, and to propagate the skin texture / tone of the hand model to the arm to ensure a uniform (or substantially uniform) skin tone between the hand and arm, thereby maintaining realism.

[0046] This disclosure also relates to various embodiments simulating real-world camera setups configured in a hemispherical shape around a hand-arm mesh model. The camera is configured to capture diverse perspectives of hand poses clearly expressed by the hand-arm mesh model. The camera may include a static camera and / or a dynamic camera.

[0047] Figure 1 This is a flowchart illustrating aspects of a method 100 for generating a high-fidelity synthetic dataset of hand and arm poses according to an embodiment of the present disclosure. Although Figure 1Various operations in a method for generating a high-fidelity synthetic dataset of hand and arm poses according to some embodiments are illustrated, but embodiments of this disclosure are not limited thereto. For example, according to various embodiments, the method may include additional or fewer operations, or the order of operations may be varied, without departing from the spirit and scope of embodiments of this disclosure, unless otherwise stated or implied.

[0048] As shown and described below, embodiments of the present disclosure can utilize a biconditional variational autoencoder (CVAE) that separately models global wrist movements and local finger poses, where global wrist movements and local finger poses can be combined to create relatively diverse and flexible hand poses. For example, in the illustrated embodiment, method 100 includes a task 110 of generating a set of finger poses (e.g., finger poses representing semantic meaning, such as extending the index finger to represent the number "1" or making a circle with the thumb and index finger to represent "OK"). In one or more embodiments, the task 110 of generating finger poses utilizes a first conditional variational autoencoder (CVAE). A CVAE is an unsupervised generative model configured to generate samples from input by encoding input data into a latent representation and then reconstructing the input from the latent space. CVAEs extend the variational autoencoder (VAE) by incorporating conditional information (such as class labels) during training and inference, which enables control over the generation of data based on specific attributes or labels.

[0049] Additionally, in the illustrated embodiment, method 100 further includes a task 120 for generating a set of wrist movements. These wrist movements indicate or represent global hand movements. In one or more embodiments, task 120 for generating wrist movements utilizes a second CVAE. The second CVAE may differ from the first CVAE used in task 110 for generating a set of finger poses.

[0050] Therefore, compared to some systems or datasets that can constrain poses to a fixed combination of global wrist movements and specific finger postures, embodiments of this disclosure enable separate modeling of global wrist movements and local finger postures, as illustrated in tasks 110 and 120 and discussed in more detail below.

[0051] Figure 2 This is a schematic representation of the CVAE 200 used in missions 110 and 120. Figure 2 The left side of the middle image depicts CVAE 200 during training, and... Figure 2 The right side of the middle image depicts CVAE 200 during inference (e.g., generation of finger poses or wrist movements). Figure 2As shown, CVAE 200 includes a transformer encoder 201 and a transformer decoder 202. During training of CVAE 200, pose labels 203, 3D joints 204, pose parameters 205, and stage labels 206 are linearized and symbolized, and then fed into the transformer encoder 201. The transformer encoder 201 encodes the input data and outputs parameters of the probability distribution (e.g., the mean and variance of a Gaussian distribution) into a latent space 207, which is a low-dimensional continuous space that encodes the input data. Pose parameters 205 refer to finger configurations (e.g., index finger pointing; thumb raised; OK shape), and stage labels 206 refer to the degree to which the finger configuration has transformed into a final finger pose (e.g., initial position, transitioning, or final position). The transformer decoder 202 is configured to create new data similar to the input data by sampling from the distribution in the latent space 207 (e.g., using a reparameterized gradient estimator 208, also known as a reparameterization trick)). In one or more embodiments, the CVAE 200 used in Mission 110 and / or Mission 120 may be the same as or similar to the CVAE described in U.S. Provisional Application No. 63 / 707422, the entire contents of which are incorporated herein by reference.

[0052] In one or more embodiments, the transformer decoder 202 for generating the first CVAE of finger pose in task 110 has more layers than the transformer decoder 202 for generating the second CVAE 200 of wrist motion in task 120. For example, in one or more embodiments, the transformer decoder 202 for the first CVAE of task 110 has eight (8) layers, and the transformer decoder 202 for the second CVAE of task 120 has two (2) layers.

[0053] During Task 110, inference of the first CVAE 200 is performed by inputting text-based finger gesture labels 203 (e.g., “pinching fingers,” “sliding fingers,” or “snap fingers”) into the first CVAE 200. The latent space 207 is then randomly sampled, and this sample is input into the transformer decoder 202. The transformer decoder 202 generates a projection 209 based on the sample from the latent space 207, and the projection 209 outputs 3D joints 210, pose parameters 211, and stage labels 212. In one or more embodiments, the output is a configured joint skeleton corresponding to the input finger gesture label (e.g., in response to the input finger gesture label being “pinching fingers,” the output is a finger joint skeleton with the tips of the thumb and forefinger touching each other). In this way, the inference process in Task 110 is configured to synthesize diverse sequences of 3D finger gestures (i.e., by sampling from the latent space, generating various different 3D joints / meshes corresponding to the input text-based finger gesture labels).

[0054] During Task 120, inference in the second CVAE 200 is performed by inputting text-based global (wrist) pose labels (e.g., “circle,” “cross,” or “move up”) into the second CVAE 200. The latent space 207 is then randomly sampled, and this sample is fed into the transformer decoder 202. The transformer decoder 202 generates a projection based on the sample from the latent space 207, and this projection outputs 3D joints / meshes 210, pose parameters 211, and stage labels 212. In this way, the inference process in Task 120 is configured to synthesize diverse sequences of wrist poses (i.e., by sampling from the latent space, generating various different 3D joints / meshes corresponding to the input text-based wrist pose labels).

[0055] Reference again Figure 1 In the illustrated embodiment, method 100 further includes task 130, which combines a set of finger poses generated in task 110 with a set of wrist movements generated in task 120 to generate a synthetic dataset of hand and arm poses. Accordingly, finger poses and global wrist movements are synthesized separately in tasks 110 and 120 and then combined. The dual process streams (i.e., finger poses generated from a first CVAE and wrist movements generated from a second CVAE) together generate diverse sequences of three-dimensional hand poses (i.e., integrating diverse finger-wrist combinations to form a diverse dataset of hand poses). In this way, combining the two streams generates a wide variety of meaningful hand poses that transcend alternative datasets that may have rigidly defined or constrained combinations of hand poses, postures, or movement trajectories. In other words, the diverse datasets of hand poses generated according to embodiments of this disclosure provide an improvement over related technical datasets with finite hand poses, and these enriched datasets represent a wide variety of pose classes.

[0056] In one or more embodiments, method 100 may include task 140, which utilizes the hand and arm pose database generated in task 130 to train a hand pose estimation model (e.g., Mobile-StereoHPE) and / or a hand pose recognition model (e.g., Fast-DNN). In one or more embodiments, the trained models (e.g., trained hand pose estimation models and / or trained hand pose recognition models) may be incorporated into extended reality (XR) devices (such as augmented reality (AR) devices, virtual reality (VR) devices, or mixed reality devices). In one or more embodiments, the hand and arm pose database generated in task 130 may be used for a synthetic saliency bokeh video dataset for a film video project.

[0057] Figure 3 A set of local gestures (i.e., finger gestures) and global wrist movements are described. Local gestures include finger gestures that represent numbers (e.g., extending the thumb or index finger to represent the number 1, extending two fingers to represent the number 2, etc.), finger-triggered gestures (e.g., fingertip contact between the thumb and index finger, fingertip contact between the thumb and middle finger, fingertip contact between the thumb and ring finger, or fingertip contact between the thumb and little finger), and special finger postures (e.g., snapping fingers, heart shape, phone call gesture, OK gesture, etc.). Global wrist movements include moving to the left, moving to the right, circling, cross-shaped, moving forward, moving backward, moving upward, moving downward, etc. Figure 3 The right side depicts the Cartesian product of these finger gestures and wrist movements, such as the index finger pointing and the wrist making a circular motion, the fingers coming together to form a fist and sliding to the right (translation), the palm opening and rotating to the left, the palm opening and moving in a circular motion, and so on.

[0058] Figure 4 This is a flowchart illustrating the task of a "cut-and-stitch" method 300 for generating a hand-arm mesh model according to an embodiment of the present disclosure. Although Figure 4 Various operations in a method for generating a hand-arm mesh model according to some embodiments are illustrated, but the embodiments of this disclosure are not limited thereto. For example, according to various embodiments, the method may include additional or fewer operations, or the order of operations may be varied, without departing from the spirit and scope of the embodiments of this disclosure, unless otherwise stated or implied.

[0059] In the illustrated embodiment, method 300 includes a task 310 for generating a hand mesh model. In one or more embodiments, in task 310, the output from the first CVAE (i.e., jointed skeletons in various finger poses) can be transformed into a three-dimensional hand mesh using MANO (hand Model with Articulated and Non-rigiddefOrmations), where MANO is a parametric hand model. The MANO hand model, described in Romero et al.'s "Embodied Hands: Modeling and Capturing Hands and Bodies Together" (arXiv:2201.02610), comprises a kinematic tree with 16 joints, and the rotation of each joint is represented in an axis-angle format aligned with an orthogonal reference to the wrist. In one or more embodiments, the method can utilize A-MANO, an anatomically constrained version of the MANO hand model, to achieve an anatomically accurate hand mesh. The A-MANO hand model redefines the canonical hand pose to compute the torsional, extension, and flexion axes in an anatomically aligned orthogonal reference. The A-MANO hand model is described in Yang et al.'s "CPF: Learning a ContactPotential Field to Model the Hand-Object Interaction" (International Conference on Computer Vision, ICCV, 2021), the entire contents of which are incorporated herein by reference. The formulation of MANO can be described according to Equation 1 as follows:

[0060]

[0061] in This represents a composite function that aligns joint angles with anatomical axes. Let θ represent the flattened hand layer in MANO, and (θ) m c,β m These are the joint angles and shape parameters in the anatomically aligned space, respectively.

[0062] In one or more embodiments, the MANO hand mesh is then transformed into a NIMBLE hand mesh, which is a high-resolution, non-rigid parametric hand model that includes bone, muscle, and skin texture. The NIMBLE hand model, as described in Li et al.'s "NIMBLE: An On-rigid Hand Model with Bones and Muscles" (arXiv:2202.04533), extends the MANO hand mesh by combining both geometric and appearance modeling, as shown in Equation 2 below:

[0063]

[0064] in Model the geometry of the hand, and Capture the texture appearance of the hand. Parameter (θ) n ,β n α) correspond to posture, shape, and appearance, respectively.

[0065] In one or more embodiments, the task of converting a MANO hand mesh to a NIMBLE hand mesh involves using a NIMBLE model with functions ranging from 5990 to 778 vertices. To deterministically subsample the MANO mesh, as shown in Equation 3 below:

[0066]

[0067] Where M v This represents the MANO vertex derived from the NIMBLE vertex. Additionally, in one or more embodiments, the task of converting the MANO hand mesh into a NIMBLE hand mesh includes utilizing a gradient-based optimization method to make the generated MANO mesh M(θ)... m ,β m ) and subsampling NIMBLE grid (M v To minimize the deviation between θ and MANO attitude parameters, we fit the MANO attitude parameters θ. m and shape parameter β m Additionally, in one or more embodiments, to ensure reliable results, the method incorporates additional regularization terms for the attitude and shape parameters, as shown in Equation 4 below:

[0068]

[0069] The objective function is formulated as:

[0070] E = ||M v -M(θ m ,β m )||2 (Equation 5)

[0071] Where E is the reconstruction error based on the L2 norm, which measures the vertex-by-vertex distance between the subsampled MANO vertices and the optimized MANO mesh.

[0072] In the opposite direction, the MANO mesh is transformed and aligned to NIMBLE by optimizing the NIMBLE parameters and wrist translation to ensure consistency.

[0073] Although the method utilizes the MANO and NIMBLE hand models in one or more embodiments, it can achieve similar high-fidelity hand reconstruction using any other suitable parametric hand model, such as Handy, in one or more embodiments. The Handy hand model is described in “Handy: Towards a highfidelity 3D hand shape and appearance model” by Potamias et al. (Proceedings of the IEEE / CVF conference on Computer Vision and Pattern Recognition (June 2023), pp. 4670-4680), the entire contents of which are incorporated herein by reference.

[0074] The NIMBLE hand mesh ends at the wrist, which negatively impacts the realism of the generated synthetic hand model. Therefore, in one or more embodiments, method 300 includes a task 320 of combining a hand mesh model (e.g., a NIMBLE hand mesh model) with a forearm template model (e.g., an arm template model) to generate a complete hand-arm mesh model.

[0075] The forearm template model can be any suitable arm model, such as an SMPL-X hand-arm model that has already been stitched together at the wrist (i.e., cut out at the wrist). The SMPL-X hand-arm model is described in “Expressive Body Capture: 3D Hands, Face, and Body from a Single Image” by Pavlakos et al., the entire contents of which are incorporated herein by reference. In one or more embodiments, the forearm mesh can be an upsampled (e.g., high-resolution) version of the SMPL-X forearm model.

[0076] In one or more embodiments, the task 320 of attaching the forearm mesh model to the hand mesh model includes: identifying the wrist boundary vertices of the NIMBLE hand mesh. And the wrist boundary vertex of the forearm model Set to the wrist boundary vertex equal to the NIMBLE hand mesh. As shown in equation 6 below:

[0077]

[0078] In one or more embodiments, task 320, which attaches the forearm mesh model to the hand mesh model, is configured to maintain flexible wrist rotation of the hand-arm model while ensuring proper alignment between the hand mesh model and the forearm mesh model. In one or more embodiments, task 320 includes applying a wrist rotation matrix R to the hand mesh model. w Wrist rotation matrix R w This is a matrix (e.g., a 3×3 matrix) representing the orientation of the wrist relative to a reference frame. The wrist rotation matrix R can be calculated using Euler angles or other angular rotation representations. w In one or more embodiments, the vertex V at the seam between the hand mesh model and the forearm mesh model... stitch The wrist rotation matrix R w The function implements dynamic hand articulation of the hand-arm mesh model and stable attachment between the hand mesh model and the forearm mesh model, as shown in Equation 7 below:

[0079]

[0080] Additionally, in one or more embodiments, the task 320 of attaching the forearm mesh model to the hand mesh model includes applying a global transformation (R) to the entire hand and arm mesh model. g ,t g As shown in Equation 8 below:

[0081] V final =R g ·V stitch +t g (Equation 8)

[0082] In one or more embodiments, method 300 further includes task 330 of removing the overlapping surface between the hand mesh model and the forearm mesh model at the wrist.

[0083] In one or more embodiments, the "cut-and-stitch" method 300 further includes a task 340 of propagating the skin texture of the hand mesh model to the forearm to ensure a uniform skin tone and maintain realism. In one or more embodiments, task 340 includes UV mapping and blending textures for the hand and arm to achieve consistent (or substantially consistent) skin texture. UV mapping is the process of opening or unfolding a 3D model into 2D space to allow textures to be applied to the surface of the model (i.e., a UV map is a vertex map that stores horizontal (U) positions and vertical (V) positions on a 2D texture map). In one or more embodiments, in task 340, UV mapping at the wrist is adjusted by interpolating between the hand mesh model (e.g., a NIMBLE hand model) and the arm mesh model (e.g., an SMPL-X arm model or an upsampled SMPL-X arm model) to prevent visual seams between the hand mesh model and the forearm mesh model.

[0084] Figures 5A to 5E Further details are provided regarding aspects of a cut-and-stitch method 300 for generating a hand-arm mesh model according to an embodiment of the present disclosure. As shown and described, the method for generating a hand-arm mesh model can cut or divide portions of the hand and arm into different segments to allow for flexible movement and consistent skin tone between the different segments.

[0085] Figure 5A A forearm mesh model 400 (e.g., SMPL-X forearm mesh model) that has been “cut” (sectioned or segmented) at the wrist is depicted (i.e., a parametric arm model 400 that has been sectioned or divided between the hand portion 401 and the forearm portion 402). Figure 5B A hand mesh model 403 (e.g., the NIMBLE hand mesh model) and Figure 5A The forearm mesh model 402 shown is an upsampled (e.g., higher resolution) version 404. In one or more embodiments, Figure 5B The task of generating a hand mesh model is described 310. Figure 5C A hand mesh model 403 and a forearm mesh model 404 (e.g., an upsampled SMPL-X forearm model) are depicted attached ("stitched") together at the wrist 405. In one or more embodiments, Figure 5C A task 320 is described, which combines a hand mesh model with a forearm template model to generate a complete hand-arm mesh model. In one or more embodiments, Figure 5CThe task 330 of removing the overlapping surface between the hand mesh model and the forearm mesh model at the wrist is also described. As described above, the process of combining the hand mesh model 403 with the forearm mesh model 404 includes: identifying the wrist boundary vertex 406 of the hand mesh model 403; and setting the wrist boundary vertex 407 of the forearm mesh model 404 to be equal to the wrist boundary vertex 406 of the hand mesh model 403 according to Equation 6 above. In addition, as described above, the process of combining the hand mesh model 403 with the forearm mesh model 404 also includes: applying the wrist rotation matrix R, which realizes the dynamic hand articulation of the hand-arm mesh model and the stable attachment between the hand mesh model 403 and the forearm mesh model 404, according to Equation 7 above. w ; and apply the global transformation according to Equation 8 above.

[0086] Figure 5D The front view 404 of the forearm mesh model 404 is depicted from the hand mesh texture 409. f And the back is 404 b Texture 408 (e.g., skin texture and / or skin tone) (e.g., skin texture and / or skin tone selected from the UV skin texture map 409 of the NIMBLE hand mesh model 403 for the forearm mesh model 404). Figure 5E Skin texture and / or skin tone 408 are depicted and applied to both the hand mesh model 403 and the forearm mesh model 404. In one or more embodiments, Figure 5D and Figure 5E Task 340: Depicting the skin texture of the hand mesh model 403 from the forearm mesh model 404 to ensure a uniform skin tone and maintain the realism of the hand-arm mesh model.

[0087] Figures 6A to 6D Simulations of real-world camera setups arranged in a hemispherical configuration or arrangement are depicted, where the real-world camera setup is configured to capture hand-arm mesh models (e.g., Figures 5A-5E The description and according to Figure 4 The hand motion (including complex finger movements and global wrist movements) is depicted in the hand-arm mesh model formed by the cut-and-stitch method 300. These camera configurations support rendering from various camera viewpoints (e.g., monocular / stereo, egocentric / holographic, static / dynamic), which enhances the robustness of real-world scenes.

[0088] Figure 6AA still camera 501 with a standard lens and an ultra-wide-angle lens is depicted, wherein the standard lens and the ultra-wide-angle lens are positioned in front, behind, to the side, and diagonally around the hand-arm mesh model, and relatively close to the hand-arm mesh model. In one or more embodiments, the still camera 501 may include four cameras (positioned in front, behind, to the left, and to the right of the hand-arm mesh model), each camera having a focal length of 35mm, a sensor width (W) and height (H) of 36.0 × 24.0 mm, an image size of 1200 × 800, and an aperture of f / 1.8. In one or more embodiments, the still camera 501 may also include two cameras (positioned at the right front diagonal and left front diagonal relative to the hand-arm mesh model), each camera having a focal length of 18mm, a sensor width (W) and height (H) of 22.3 × 14.9 mm, an image size of 1200 × 800, and an aperture of f / 2.8. In one or more embodiments, the still camera 501 may also include two ultra-wide-angle cameras (positioned at the right front diagonal and left front diagonal relative to the hand-arm mesh model), each ultra-wide-angle camera having a focal length of 13mm, a sensor width (W) and height (H) of 6.17 × 4.55 mm, an image size of 3200 × 2400, and an aperture of f / 2.2. Figure 6A The static camera 501 depicted is positioned at a key angle, which is configured to capture a full and distortion-free coverage of the motion of the hand-arm mesh model.

[0089] Figure 6BA dynamic camera 502 is depicted, comprising a macro lens and a pair of stereo lenses, wherein the macro lens is positioned on the palm side of the hand-arm mesh model and configured to capture complex finger movements of the hand-arm mesh model, and the pair of stereo lenses are positioned on the back of the hand-arm mesh model and configured to track movement and three-dimensional (3D) depth behind the hand-arm mesh model. In one or more embodiments, the dynamic camera 502 may include a camera (positioned on the palm side of the hand-arm mesh model) and a pair of stereo cameras (positioned on the back of the hand-arm mesh model), the camera having a focal length of 85mm, a sensor width (W) and height (H) of 36.0 × 24.0 mm, an image size of 1200 × 800, and an aperture of f / 1.4, and the pair of stereo cameras having a focal length of 18mm, a sensor width (W) and height (H) of 27.36 × 24.0 mm, an image size of 1200 × 800, and an aperture of f / 2. In one or more embodiments, the dynamic camera 502 may include a camera (positioned on the palm side of the hand-arm mesh model) and a pair of stereo cameras (positioned on the back of the hand-arm mesh model), the camera having a focal length ranging from 70 mm to 100 mm, a sensor width (W) and height (H) in millimeters (mm) ranging from 24 × 16 to 48 × 32, an image size ranging from 800 × 600 to 1600 × 1000, and an aperture ranging from f / 1.2 to f / 1.6, and the pair of stereo cameras having a focal length ranging from 12 mm to 24 mm, a sensor width (W) and height (H) in millimeters (mm) ranging from 18 × 16 to 36 × 32, and an image size ranging from 800 × 600 to 1600 × 1000.

[0090] Figure 6C A still camera 503 is depicted, featuring a standard lens and an ultra-wide-angle lens, wherein the standard lens and ultra-wide-angle lens are positioned around a hand-arm mesh model at forward, rear, side, and diagonal positions, and are... Figure 6A The camera arrangement in the model is relatively far from the hand-arm mesh. Figure 6D Depicting Figures 6A to 6C A combined view of all camera settings depicted in the image.

[0091] Now refer to Figures 7A to 7BA virtual reality and / or augmented reality (VR / AR) system 600 according to one embodiment of the present disclosure includes a digital display 601 (e.g., a digital microdisplay, such as an organic light-emitting diode (OLED) display) and a lens system 602 (i.e., viewing optics) in front of the digital display 601. When a user wears the VR / AR system 600, the lens system 602 is positioned between the digital display 601 and the user's eyes.

[0092] In one or more embodiments, the VR / AR system 600 further includes a processor 603 coupled to a digital display 601, a non-volatile memory device 604 coupled to the processor 603 (e.g., flash memory, ferroelectric random access memory (F-RAM), magnetostrictive RAM (MRAM), FeFET memory, and / or resistive RAM (ReRAM) memory), and a power supply 605 coupled to the processor 603 (e.g., one or more secondary batteries). The non-volatile memory device 604 includes executable instructions (i.e., computer-readable code) that, when executed by the processor 603, cause the processor 603 to control the display 601 to display various images. In one or more embodiments, the VR / AR system 600 may include an input device 606 (e.g., a handheld controller) configured to perform various operations, such as modifying images displayed by the display 601. In one or more embodiments, the VR / AR system 600 may include a communication module (e.g., a network adapter) 607 configured to support the establishment of a direct (e.g., wired) or wireless communication channel between the VR / AR system 600 and an external electronic device (e.g., another electronic device or a server), and to perform communication via the established communication channel. In one or more embodiments, the VR / AR system 600 may also include a headband or strap 608 (e.g., an adjustable strap) configured to secure the VR / AR system 600 to a user's head.

[0093] The term "processor" is used herein to include any combination of hardware, firmware, and / or software for processing data or digital signals. The hardware of a processor may include, for example, application-specific integrated circuits (ASICs), general-purpose or special-purpose central processing units (CPUs), digital signal processors (DSPs), graphics processing units (GPUs), and programmable logic devices (such as field-programmable gate arrays (FPGAs)). As used herein, in a processor, each function is either performed by hardware configured (i.e., hardwired) to perform that function, or by more general-purpose hardware (such as a CPU) configured to execute instructions stored in a non-transitory storage medium. A processor may be manufactured on a single printed wiring board (PWB) or distributed across several interconnected PWBs. A processor may contain other processors; for example, a processor may include two processors (an FPGA and a CPU) interconnected on PWBs. Processor 603 may include a main processor (e.g., a central processing unit (CPU) or application processor (AP)) and an auxiliary processor (e.g., a graphics processing unit (GPU), image signal processor (ISP), sensor hub processor, or communication processor (CP)), wherein the auxiliary processor may operate independently of the main processor or in conjunction with the main processor. Additionally or alternatively, the auxiliary processor may be adapted to consume less power than the main processor or to perform specific functions. The auxiliary processor may be implemented separately from or as part of the main processor. The auxiliary processor may, when the main processor is inactive (e.g., in a sleep state), control at least some functions or states associated with at least one component of the electronic device, in place of the main processor, or, when the main processor is active (e.g., executing an application), control at least some functions or states associated with at least one component of the electronic device together with the main processor. The auxiliary processor (e.g., an image signal processor or a communication processor) may be implemented as part of another component functionally associated with the auxiliary processor.

[0094] Communication module 607 may include one or more communication processors, which may operate independently of processor 603 (e.g., AP) and support direct (e.g., wired) or wireless communication with another device. Communication module 607 may include a wireless communication module (e.g., a cellular communication module, a short-range wireless communication module, or a Global Navigation Satellite System (GNSS) communication module) or a wired communication module (e.g., a local area network (LAN) communication module or a power line communication (PLC) module). One of these communication modules may communicate via a short-range communication network (e.g., Bluetooth). TMCommunication modules can communicate with external electronic devices or servers via Wi-Fi Direct or Infrared Data Association (IrDA) standards or remote communication networks (e.g., cellular networks, the Internet, or computer networks such as LANs or WANs). These various types of communication modules can be implemented as a single component (e.g., a single IC) or as multiple separate components (e.g., multiple ICs). Communication module 607 can use subscriber information (e.g., International Mobile Subscriber Identity (IMSI)) stored in the subscriber identification module to identify and authenticate communication networks (such as short-range communications). A VR / AR system 600 is located in a network or remote communication network. In one or more embodiments, a communication module 507 may include an antenna configured to transmit signals and / or power to or from the outside of the VR / AR system 600 (e.g., an external electronic device). In one or more embodiments, the communication module 607 may include one or more antennas, whereby at least one antenna suitable for a communication scheme used in a communication network (such as a long-range or short-range communication network) can be selected by the communication module 607, for example. Signals or power can then be transmitted or received between the communication module 607 and the external electronic device via the selected at least one antenna.

[0095] Commands or data can be sent or received between the VR / AR system 600 and external electronic devices via a server coupled to a remote communication network. Each electronic device can be of the same or different type as the VR / AR system 600. All or some operations to be performed at the VR / AR system 600 can be performed at one or more external electronic devices. For example, if the VR / AR system 600 is required to automatically perform a function or service, or to perform a function or service in response to a request from a user or another device, the VR / AR system 600 may request one or more external electronic devices to perform at least a portion of that function or service instead of performing it. The one or more external electronic devices receiving the request may perform at least a portion of the requested function or service, or additional functions or services related to the request, and transmit the result of the performance to the VR / AR system 600. The VR / AR system 600 may provide the result, or not, with further processing on the result as at least part of a response to the request. For this purpose, cloud computing, distributed computing, or client-server computing technologies may be used, for example.

[0096] In one or more embodiments, the memory device 604 of the VR / AR system 600 may include a hand pose estimation model or module (e.g., a moving stereo HPE) and / or a hand pose recognition model or module (e.g., a fast DNN), such that the VR / AR system 600 is configured to recognize a user's hand poses as input commands. In one or more embodiments, the hand pose estimation model or module and / or the hand pose recognition model or module may have been trained on a hand and arm pose database generated according to embodiments of the present disclosure.

[0097] Figure 8 This is an overview of a hand-arm assembly line 700 according to an embodiment of the present disclosure. Figure 8 The left side depicts a database including multiple different finger gestures and hand movements 701 (e.g., pinching, sliding, grasping, pointing gestures, etc.), a hand and arm mesh model 702, various different skin textures (appearances) 703 applied to the hand and arm mesh model 702, and an environment map 704. Figure 8 The middle part depicts a multi-camera setup 705 (e.g., camera lens, AR / VR headset, smartphone, etc.) surrounding a hand and arm mesh model 706 that forms a semantic gesture (e.g., extending the index and middle fingers to represent the number 2). Figure 8 The right side depicts the rendered RGBD sequence 707, which combines color (red, green, blue (RGB)) and depth (D) information of real hand poses (e.g., "grasping" hand poses) selected from a database of different finger poses and hand movements 701, dynamic lighting, and a variety of environments selected from environment map 704.

[0098] Electronic or electrical devices and / or any other related devices or components according to embodiments of the invention described herein can be implemented using any suitable hardware, firmware (e.g., application-specific integrated circuits), software, or a combination of software, firmware, and hardware. For example, various components of these devices can be formed on an integrated circuit (IC) chip or a discrete IC chip. Furthermore, various components of these devices can be implemented on flexible printed circuit films, tape-on-a-carrier packages (TCPs), printed circuit boards (PCBs), or formed on a substrate. Additionally, various components of these devices can be processes or threads running on one or more processors in one or more computing devices, executing computer program instructions and interacting with other system components to perform the various functions described herein. The computer program instructions are stored in memory, which in the computing device can be implemented using standard memory devices such as, for example, random access memory (RAM). The computer program instructions can also be stored in other non-transitory computer-readable media, such as, for example, CD-ROMs, flash drives, etc. Furthermore, those skilled in the art will recognize that, without departing from the spirit and scope of exemplary embodiments of the invention, the functions of various computing devices can be combined or integrated into a single computing device, or the functions of a particular computing device can be distributed among one or more other computing devices.

[0099] While aspects of some embodiments of this disclosure have been described in detail with reference to some examples thereof, the disclosed embodiments described herein are not intended to be exhaustive or to limit the scope of the invention to the exact forms disclosed. Those skilled in the art will understand that modifications and alterations to the described structures and methods of assembly and operation can be practiced without intentionally departing from the principles, spirit, and scope of the invention as set forth in the following claims and their equivalents.

Claims

1. A method for generating synthetic datasets, comprising: A set of finger poses is generated by the processor from a first conditional variational autoencoder, which includes a first potential space and a first transformer decoder; The processor generates a set of wrist movements from a second conditional variational autoencoder, which includes a second latent space and a second transformer decoder; as well as The processor combines the set of finger poses and the set of wrist movements to generate a synthetic dataset of hand and arm poses.

2. The method of claim 1, wherein the combination comprises the Cartesian product of the set of finger gestures and the set of wrist movements executed by the processor.

3. The method according to claim 1, wherein the first conditional variational autoencoder is different from the second conditional variational autoencoder.

4. The method of claim 3, wherein the first converter decoder has eight layers and the second converter decoder has two layers.

5. The method according to claim 1, further comprising generating a hand-arm mesh model of the hand-arm pose by the processor.

6. The method of claim 1, wherein the set of finger gestures includes at least one digital gesture, at least one trigger gesture, and at least one special gesture.

7. A method for generating models, comprising: The processor generates a hand mesh model; as well as The processor combines the arm mesh model with the hand mesh model to generate a hand-arm mesh model, wherein combining the arm mesh model with the hand mesh model includes: The processor identifies the wrist boundary vertices of the hand mesh model and the arm mesh model; The processor controls the number of wrist boundary vertices in the hand mesh model to be equal to the number of wrist boundary vertices in the arm mesh model; and The processor applies a wrist rotation matrix to the hand mesh model.

8. The method of claim 7, further comprising removing, by the processor, the overlapping surface between the hand mesh model and the arm mesh model at the wrist of the hand-arm mesh model.

9. The method of claim 8, further comprising the processor interpolating at the wrist between the hand mesh model and the arm mesh model to prevent visual seams between the hand mesh model and the arm mesh model.

10. The method of claim 7, further comprising: The processor applies skin texture to the hand mesh model; as well as The processor transmits the skin texture of the hand mesh model to the arm mesh model.

11. The method of claim 7, wherein the hand mesh model comprises a NIMBLE model.

12. The method of claim 11, wherein the arm mesh model comprises the SMPL-X model.

13. The method of claim 7, further comprising applying a global transformation to the hand-arm mesh model by the processor.

14. The method of claim 7, wherein generating the hand mesh model comprises the processor converting the MANO hand model into a NIMBLE hand model.

15. The method of claim 7, wherein the hand mesh model comprises a Handy model.

16. A method for simulating a real-world camera configuration, the method comprising: Multiple cameras are arranged in a hemispherical configuration around the hand-arm mesh model; as well as The hand movements of the hand-arm mesh model are captured from different perspectives using the multiple cameras.

17. The method of claim 16, wherein the plurality of cameras comprises a plurality of still cameras.

18. The method of claim 16, wherein the plurality of cameras comprises a plurality of dynamic cameras.

19. The method of claim 18, wherein the plurality of dynamic cameras includes a first camera having a macro lens facing the palm side of the hand-arm mesh model, and a pair of stereo cameras facing the back side of the hand-arm mesh model.

20. The method of claim 16, further comprising generating the hand-arm mesh model, including: The processor generates a hand mesh model; as well as The processor combines the arm mesh model with the hand mesh model to generate the hand-arm mesh model, wherein combining the arm mesh model with the hand mesh model includes: The processor identifies the wrist boundary vertices of the hand mesh model and the wrist boundary vertices of the arm mesh model; The processor controls the number of wrist boundary vertices in the hand mesh model to be equal to the number of wrist boundary vertices in the arm mesh model; and The processor applies a wrist rotation matrix to the hand mesh model.