Reversible nerve skin
By constructing a reversible neural skin (INS) pipeline, combining the attitude conditional reversible network (PIN) and the differentiable linear hybrid skin (LBS) module, the problems of frequent grid extraction and nonlinear deformation capture in the prior art are solved, and efficient pose change animation production and detail capture are achieved.
Patent Information
- Application Number
- CN202380090010.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-29
- Filing Date
- 2023-12-27
- Publication Date
- 2025-08-05
AI Technical Summary
The prior art requires frequent extraction of grids and difficult to effectively capture nonlinear surface deformations when animating characters in clothing, resulting in volume loss and artifacts, and the existing methods are inefficient in pose changes.
Reversible neural network is used to construct a reversible neural skin (INS) pipeline. By combining the attitude condition reversible network (PIN) and the differentiable linear hybrid skin (LBS) module, nonlinear surface deformation is learned and volume loss is reduced. The pose can be redirected by just one mesh extraction.
It realizes efficient posture change animation production, reduces the number of grid extraction times, improves the accuracy and efficiency of posture correspondence, and can better capture the nonlinear deformation details of clothes.
Smart Images

Figure CN120435728A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. application serial number 18 / 090,724, filed on December 29, 2022, the contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] Examples described herein generally relate to animation of three-dimensional (3D) objects, and more particularly to methods and systems for animating 3D meshes of deformable objects by extending linear blend skinning using reversible neural networks. Background Art
[0004] Being able to generate animatable representations of clothed people out of a skinned mesh is extremely useful for building realistic augmented or virtual reality experiences and improving simulators. The state of the art in this area has moved from building parametric models of the human body to the latest state-of-the-art techniques that learn implicit 3D neural representations from data in a canonical space. The canonical representations of the state-of-the-art animate the character's skin (represented as a deformable mesh model) by following the motion of an underlying abstract skeleton, learning a skin weight field around it, and applying Linear Blend Skinning (LBS) or skeletal animation to deform the character's skin (represented as a deformable mesh model) to a new pose, where the pose is defined by the skeleton underlying the character's skin's 3D surface.
[0005] Parametric models typically define correspondences between poses represented as a set of skeletons and mesh vertices using LBS weights. These weights provide a soft assignment of vertices to human skeletons. Therefore, for animation, these models transform vertices using linear combinations of the skeleton transformations. When parametric models are unavailable, these weights need to be discovered. To this end, recent prior art has adopted learning-based solutions to discover LBS weights. They typically assume a shared canonical space and learn a canonical LBS weight field, which is used to deform the body to novel poses during inference. However, during training, the character needs to be backward warped from the deformed space to the canonical space (i.e., given a deformed point, the corresponding canonical point needs to be obtained). Therefore, some prior art methods learn LBS weights separately in the deformed and canonical spaces, which can be used to establish correspondences. These methods often require a cycle consistency loss for adjustment. Recently, differentiable forward skinning has been used to animate non-rigid neural implicit shapes. Differentiable forward skinning for animation of non-rigid neural implicit shapes (SNARF) computes these correspondences by finding solutions to the LBS equations using an iterative solver.
[0006] Numerous existing techniques have been developed for constructing parametric representations of the human body or specific parts such as hands and faces. In addition to the human body, recent techniques have developed parametric animal models. Some existing techniques have explored constructing implicit representations of the human body with and without clothing. However, representing characters as implicit functions requires time-consuming mesh extraction via marching cubes. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In the accompanying drawings, which are not necessarily drawn to scale, like numerals may describe similar components in different views. Some non-limiting examples are shown in the images of the accompanying drawings, in which:
[0008] Figure 1 is a schematic diagram depicting the bidirectional correspondence of a canonical representation of a clothed person in a deformation space using fast and reversible poses.
[0009] Figure 2 is a schematic diagram depicting the pose-conditioned 2D coupling layer of an invertible neural network (INN) in an example configuration.
[0010] Figure 3 is a diagram depicting the spatial and pose-aware conditions whereby, in an example configuration, body pose is encoded using a per-bone MLP network operating on individual bone transformations.
[0011] Figure 4A is a schematic diagram depicting, from right to left, the training of an Invertible Neural Skinning (INS) pipeline for points in deformable space that are processed to find correspondences in canonical space.
[0012] Figure 4B Figure 1 is a diagram depicting, from left to right, how a character is animated using the trained INS pipeline to handle a given pose using a common set of skeletons.
[0013] Figure 5 is a graph showing a time comparison of redirection attitudes between INS and SNARF.
[0014] Figure 6 is a block diagram of a machine within which instructions (eg, software, program, application, applet, application, or other executable code) may be executed for causing the machine to perform any one or more of the methodologies discussed herein.
[0015] Figure 7 is a block diagram illustrating a software architecture in which the examples described herein may be implemented. DETAILED DESCRIPTION
[0016] The Pose-conditioned Invertible Network (PIN) architecture extends the pure LBS process by using a pose-free canonical representation to learn additional pose-varying deformations. PIN is also combined with a differentiable LBS module to build an expressive and end-to-end learnable pose-redirected Invertible Neural Skinning (INS) pipeline, which addresses the shortcomings of the prior art by allowing the animation of implicit surfaces (e.g., 3D meshes using skeletons) with complex pose-varying effects without the need to extract a mesh for each pose, while also preserving the correspondence between poses. The described INS network is shown to correct the artifacts introduced by LBS.
[0017] The subject matter described in this paper uses Invertible Neural Networks (INNs), which are bijective functions that can maintain an exact correspondence between their input and output spaces while learning complex nonlinear transformations between them. This capability of INNs makes them suitable candidates for pose retargeting. This paper utilizes INNs to construct an Invertible Neural Skinning (INS) pipeline. To this end, a pose-conditional Invertible Network (PIN) is constructed to learn pose-conditional deformations. To produce an end-to-end Invertible Neural Skinning (INS) pipeline, two PINs are placed around a differentiable LBS module using a pose-free canonical representation. These PINs help capture the nonlinear surface deformations of the garment at different poses and mitigate the volume loss caused by the LBS operation. Since the canonical representation remains pose-free, the expensive mesh extraction only needs to be performed once, and the mesh is warped with the learned LBS during the reverse pass through the INS pipeline to retarget the pose.
[0018] The present disclosure provides methods and instructions on a computer-readable medium for implementing a method for animating a three-dimensional (3D) mesh of a deformable object using an invertible neural skinning (INS) pipeline. The method includes receiving as input a given pose of the deformable object defined by a set of universal skeletons by a first pose-conditional invertible network (PIN), and mapping, by the first PIN, canonical points of the deformable object in a pose-independent canonical space to points in a pose-dependent canonical space. The points in the pose-dependent canonical space are transformed by a differentiable linear blend skinning (LBS) network into deformed points in a novel pose of the deformable object. A second PIN performs error correction on the deformed points in the deformed space. The method also includes extracting a mesh of the deformable object from a pose-free canonical occupancy network or neural representation to obtain a mesh of the deformable object in a pose-independent canonical space. Mesh vertices of the extracted mesh of the deformable object can be repositioned using the universal skeleton set by passing through the first PIN, a differential LBS network, and a second PIN. The mesh of the deformable object is extracted once for different poses, and the repositioning of the extracted mesh includes warping it using the differential LBS network.
[0019] The method also includes linking together one-dimensional (1D) and two-dimensional (2D) pose-conditional coupling layers of a reversible neural network (INN) to form a first PIN and a second PIN. In an example configuration, the first PIN and the second PIN are reversible to maintain an exact correspondence between input and output.
[0020] The method may also include encoding each bone transformation in a given pose of the deformable object using an operation graph that takes a six-dimensional (6D) input that concatenates three-dimensional (3D) translations and rotations; and obtaining a pose embedding by concatenating the outputs of each bone.
[0021] The INS pipeline can be trained in the following way: the second PIN receives input scans of a deformable object at different poses in the deformation space, the second PIN provides the pose corresponding to the input scan, the differentiable LBS network obtains the pose correspondences of the canonical points in the pose-independent canonical space from the given pose, and the first PIN maps the points in the pose-dependent canonical space to the canonical points in the pose-independent canonical space, and passes the canonical points in the pose-independent canonical space to the pose-free occupancy network.
[0022] In an example configuration, a reversible neural skinning (INS) pipeline for animating a three-dimensional (3D) mesh of a deformable object includes a first pose-conditional reversible network (PIN) that receives as input a given pose of the deformable object defined by a set of general bones and maps canonical points of the deformable object in a pose-independent canonical space to points in a pose-dependent canonical space; a differentiable linear blend skinning (LBS) network that receives points in the pose-dependent canonical space and transforms these points to obtain deformed points in a novel pose of the deformable object; and a second PIN that receives the novel pose of the deformable object and performs error correction on the deformed points in the deformation space.
[0023] The INS pipeline can also include a pose-free canonical occupancy network or neural representation from which the mesh of the deformable object is extracted to obtain a mesh in a pose-independent canonical space. A set of common skeletons can be used to reposition the mesh vertices of the mesh extracted from the canonical occupancy network by passing it through a first PIN, a differential LBS network, and a second PIN. The mesh is extracted once for different poses and repositioned by warping the extracted mesh with a differential LBS network.
[0024] In an example configuration, the first and second PINs are reversible to maintain an exact correspondence between input and output, and each of the first and second PINs includes one-dimensional (1D) and two-dimensional (2D) pose-conditional coupling layers of a reversible neural network (INN) linked together.
[0025] Now refer to Figure 1-7 Detailed description of a method for animating a 3D mesh of a deformable object is described. Although this specification provides detailed descriptions of possible implementations, it should be noted that these details are intended to be exemplary and not to limit the scope of the inventive subject matter in any way.
[0026] Reversible neural networks (INNs) were originally designed for tractable density estimation in high dimensions and generative modeling, i.e., normalized flows. Typically, INNs are constructed by chaining together multiple conditionally coupled layers, where a single coupled layer defines a reversible transformation between its input and output. The main idea behind the coupling layer is that if the input is split into two parts and only the first part is modified while the second part is modified conditionally, the input should generally be reversible. Another popular type of reversible transformation is the reversible residual layer with a small condition number. They use fixed-point iteration to find the inverse. However, the present disclosure relies primarily on coupling layers because they are faster. In the context of 3D vision, INNs have been used to learn primitives for 3D representations, perform 3D shape completion tasks, and reconstruct dynamic scenes. The present disclosure extends the use of INNs to animated 3D characters.
[0027] The systems and methods described herein learn a 3D representation of the human body that allows the generation of novel poses beyond the original training data (i.e., redirected poses). For each subject, assume that there are N pairs of available data consisting of a skeletal pose and a 3D mesh, denoted as (θ t ,M t ) N t=1 Such data can be obtained from human scans, and pose can be estimated by fitting a parameterized SMPL-like body model to these scans. Given this data, we learn an implicit neural representation of a specific subject in a canonical space and a method to animate this representation.
[0028] Figure 1 The bidirectional correspondence between the canonical representation of a clothed person in the canonical space 100 and the canonical representation of the clothed person in the deformed (bidirectional correspondence) space 110 obtained using fast and reversible poses is shown. The input point in the deformed space 110 is represented as And a point in the canonical space 100 is represented as Since the input consists of a series of deformed (pose-bearing) meshes, the superscript t is used to indicate the time step of the capture. Since the canonical space 100 is pose-independent, it is shared across all time steps. Therefore, p c There is no time index.
[0029] To identify deformed and canonical poses, we follow the skinned multi-person linear (SMPL) model, which represents the body pose as a set of bones in a kinematic tree. When retargeting the pose, since only the relative pose between the canonical space and the deformed space 110 is required at any given time t, the retargeted pose is given by θ t =[B1,...,B nb ] indicates that B i =[R i |t i ] represents the transformation of the i-th bone in 3D space, that is, B i ∈SE(3), the corresponding rotation With translation The total number of bones is n b express.
[0030] To represent a specific subject, an occupancy network O is used, which only takes as input a point p c is conditioned to provide the pose-free canonical occupancy. Then, the canonical surface S c is implicitly represented as the level set of the occupancy network (σ = 0.5):
[0031]
[0032] To extract this canonical isosurface into a mesh, the MISE algorithm is used. This differs from prior art methods that use additional pose conditions in the canonical occupancy network.
[0033] For training and evaluation of the INS, 3D points are sampled in the deformation space 110 and their ground truth occupancy values of zero or one are obtained based on whether they lie outside the grid (scan) or not.
[0034] Differentiable forward blending skinning
[0035] To animate the subject from its canonical pose to a deformed pose, linear blend skinning (LBS) is used. LBS involves deforming the canonical surface according to a convex combination of rigid bone transformations. Specifically, a differentiable LBS formulation from SNARF is used, as described below.
[0036] The learnable weight field in the canonical space 100 is defined and parameterized by a neural network: w lbs : For a given point in the canonical space 100, this weight field predicts the blending weights corresponding to each bone:
[0037]
[0038] In order to make the weight (w i) For LBS to be convex, use softmax to constrain them to always be non-negative and sum to 1.
[0039] Given the above weight field and bone transformation, the relative body pose θ t =[B1,...,B nb ], LBS can be used to normalize any point p in the space 100 c Warp forward to deformation space 110 as follows:
[0040]
[0041] in Indicates LBS after p c The corresponding points fall in the deformation space 100.
[0042] The canonical correspondences can be searched. When training on the original scans, only the points in the deformed space 110 are provided. To find their possible correspondences in the canonical space 100, an iterative solver is used to solve the roots of equation (3) while maintaining w lbs Specifically, Broyden's method is used to initialize the root finding algorithm at K different points in the canonical space 100, for each deformation point Find a group The point correspondence is as follows:
[0043]
[0044] The above formula is end-to-end differentiable because the weight field w can be computed via implicit differentiation as shown in SNARF lbs Relative to the input point The following also extends these derivations to calculate the corresponding relationship The gradient with respect to the input point.
[0045] The above differentiable formulations share the same limitations as traditional LBS, such as the inability to represent clothed surfaces and the introduction of volume loss. For example, SNARF struggles to represent finer details such as cloth wrinkles, while SNARF-NC struggles to handle LBS artifacts such as volume loss and the candy wrapper effect. This is particularly problematic when learning from real-world data of people wearing clothing in various poses.
[0046] Posture-Conditioned Reversible Network (PIN)
[0047] As mentioned above, a reversible neural network is a bijective function composed of modular components called coupling layers, which maintain a one-to-one correspondence between input and output. The following describes the construction of the proposed pose-conditional coupling layers, which are chained together to construct the PIN.
[0048] Figure 2 FIG. 1 is a diagram illustrating a pose-conditioned 2D coupled layer of an invertible neural network (INN) 200 in an example configuration. The spatial pose condition is used to generate a 2D coupled layer using a multilayer perceptron (MLP) r and m t The two operation maps in the form of are used to predict the operation parameters, and MLP is used to rotate (R xy ) and translation ([t x ,t y ]) Input split [x,y]. In this case, [z] remains unchanged.
[0049] The coupling layer operates by splitting its input into two parts using a fixed fracture pattern. Figure 2 In , the numbers "1" and "2" represent the dimensions of the vector. Figure 2 As shown, after segmentation, the first part of the input (e.g., [x, y]) is transformed by applying a series of reversible operations (such as translation and rotation). The parameters of these operations can be generated by arbitrary functions that are jointly conditioned on the second part of the input (e.g., z) and external conditions (such as pose).
[0050] Formally, when the system operates in 3D space, the input point can be defined as [x, y, z], and the input is split into [x, y] and [z]. Then, the 2D coupling layer G xy ([x,y,z],θ t ) defines a reversible transformation as follows:
[0051] [x′, y′] = R xy [x, y] T + [t x , t y ] and z′ = z, (5)
[0052] in and is a rotation matrix and a translation vector generated by an arbitrary function that simply transforms the skeleton pose θ t and coordinate z as input. The inverse G of the coupling layer xy -1 ([x,y,z],θ t ) can be calculated as:
[0053] [x, y] = Rxy -1 ([x′, y′] - [t x , t y ]) and z = z′. (6)
[0054] The operating parameter R will now be described. xy and [t x ,t y ] calculation.
[0055] Posture θ t Each bone in is transformed using MLP m b is encoded, and MLP m b Takes a 6D input of concatenated 3D translation and rotation (as Euler angles). To obtain the pose embedding, each skeleton e θ The output is concatenated as follows:
[0056]
[0057] A learned periodic position encoding (e.g., a simple neural network architecture using a sine as an implicit neural representation of a periodic activation function (SIREN)) can be used to map spatial coordinates to:
[0058]
[0059] Such encoding helps to better represent high-frequency surface details, such as cloth wrinkles.
[0060] When the relative pose θ between the deformed space and the canonical space t When B is zero (i.e. i = [I|0], all bone transformations have identity rotation and zero translation), the coupling layer should not introduce any spatial variation (i.e., z-condition) changes. To enforce this, the Hadamard product of the spatial and pose-aware embeddings can be performed on the spatial and pose-aware conditions and then concatenated to obtain:
[0061]
[0062] Figure 3 is the description of the spatial and posture perception conditions e sp Schematic diagram of a per-bone MLP network m operating on a single bone transformation in the example configuration. b To adjust body posture θ t Then the posture is embedded into e θ is fused with spatial embeddings (e.g., by SIREN) to generate a pose-aware conditional vector e for the PIN sp .
[0063] To generate the parameters of the coupling operation, including the translation map m t and rotation graph m r The two MLPs transform the above conditional vector e sp As input:
[0064]
[0065] It should be noted that m r The output of only predicts the rotation angle γ in radians xy (Single value). The axis of rotation passes through the origin of the split input space, which is XY space in this example. γ xy is converted into a rotation matrix R xy .
[0066] Regarding the above Figure 2 Unlike the 2D coupled layers described above, rotation operators cannot be used in 1D. In this case, only translation is used. For a layer G with partitioning patterns [x] and [y, z] x ([x,y,z],θ t ), the coupling operation becomes:
[0067] x′ = x + t x and [y′, z′] = [y, z], (12)
[0068] in Use pan map t Produced in a similar manner to 2D coupled layers: where a single scalar is output instead of a 2D translation. In the 1D case, the spatial embedding of Eq. (8) takes two coordinates as input e xy =Φ([x,y]):
[0069] The pose-conditioned reversible network (PIN) can be composed by chaining together multiple 1D and 2D pose-conditioned coupling layers as follows:
[0070]
[0071] Where p represents a point in 3D space, G i represents the coupling layer, and θ t Represents the posture. Inverse PIN is equivalent to inverting each coupling layer in reverse order:
[0072]
[0073] Since the PIN is reversible by construction, it maintains the exact correspondence between its input and output spaces:
[0074]
[0075] The single coupling layer of PIN Figure 2 is visualized.
[0076] Reversible neural skin
[0077] Figure 4A and 4B An attitude reversible neural skin (INS) pipeline 400 is shown, comprising the three previously described components linked together:
[0078] ·H c : Pose-Conditioned Invertible Network (PIN) H operating after the canonical space 100 and before the LBS network 420 c 410.
[0079] ·H d : A pose-conditioned invertible network (PIN) 430 operates before the deformation space 110 and after the LBS network 420.
[0080] ·In PIN H c 410 and H d 430 operates between differentiable LBS networks 420 .
[0081] These PINs (H c and H d ) 410 and 430 capture the nonlinear surface deformation of the clothes and attenuate LBS artifacts. As mentioned above, the canonical representation is not conditioned on the target pose and only needs to extract the mesh once.
[0082] As described below, an invertible mapping can be formulated that preserves the correspondence between the deformation space 110 and the canonical space 100.
[0083] Transform to Standard (Training)
[0084] Figure 4A is a schematic diagram depicting, from right to left, the training of an Invertible Neural Skin (INS) pipeline 400 for points in the deformation space 110 that are processed to find correspondences in the canonical space 100. Figure 4A As shown from right to left, for a point in the deformation space 110 Use PIN H d 430 This point is processed to obtain a novel pose in the post-LBS space Next, Broyden's algorithm is used to obtain the normalized space 100 The corresponding relationship, for example, Finally, use the second PIN H c410 Mapping the correspondences to points in the pose-independent canonical space 100 In particular:
[0085]
[0086] To obtain the best-fitting canonical correspondence, arg max is used for all predicted canonical occupancies:
[0087]
[0088] During training, the argmax is approximated with a softmax function in order to softly backpropagate gradients through all correspondences after SNARF.
[0089] In the training dataset, points in the deformation space 110 and the corresponding ground truth occupancy values of zero or one are provided. These deformation points are mapped to the canonical space 100 and a binary cross entropy loss is applied to jointly train all components of the pose network according to the following formula:
[0090]
[0091] H d ,H c ,w lbs ,O
[0092] Furthermore, after SNARF, a prior on the canonical pose can be enforced by using two additional losses during the first cycle. First, additional points are sampled on the skeleton of the canonical pose and their occupancy is encouraged to be 1. Second, the skin weights of the skeleton joints are encouraged to be equal. However, no ground truth skin weights are required during these steps.
[0093] Therefore, the INS pipeline 400 is trained to provide PIN H d 430 provides input scans of the human body in different postures in the deformation space 110 to obtain novel postures The differentiable LBS network 420 uses the Broyden algorithm to obtain the The corresponding relationship Second PIN H c 410 maps these correspondences to points in the pose-independent canonical space 100 To train the occupancy network 440.
[0094] Normalization to deformation (inference)
[0095] Figure 4BSchematic diagram depicting, from left to right, how the trained INS pipeline 400 is used to animate a character (e.g., via skeletal joints) using a generic skeleton set to handle a given pose. Figure 4B As shown, once the INS pipeline 400 is trained, it can be used to train the INS pipeline 400 at any given pose θ n The character is animated in two steps. First, a mesh extraction is run on the canonical occupancy network O 440 to animate the character from the canonical point p c Get the pose q in the canonical space 100 c Next, via the reverse pass of the pose INS pipeline 400, the mesh vertices are reoriented using bones to poses:
[0096]
[0097] Since the canonical occupancy network O 440 is related to θ n is irrelevant, so the mesh can be extracted exactly once. Re-positioning the mesh for a series of poses is equivalent to performing multiple inferences as described in equation (19).
[0098] In an alternative configuration, the occupancy network 440 can be replaced with a neural representation that can process texture and lightning, thereby learning directly from 2D images and videos rather than from raw scans.
[0099] Thus, during inference by the INS pipeline 400, given the pose θ n and a common set of skeletons m b As input. PIN H c 410 sets the normative point in the pose-independent normative space 100 Mapping to the former LBS space to obtain the point p in the normative space 100 c The posture correspondence Posture correspondence is applied to the differentiable LBS network 420 to correct the error, thereby obtaining a novel pose in the post-LBS space PIN H d 430 will be a novel posture Transformed into a point in deformation space 110 Training Evaluation
[0100] The INS method is trained using sample points in the deformation space 110 and the corresponding occupancy and pose. The INS pipeline 400 is benchmarked on two datasets: CAPE, which features scans of people wearing loose clothing, and DFAUST, which contains scans of people wearing only minimal clothing.
[0101] CAPE contains scans of 11 subjects (8 males and 3 females) wearing 8 different types of clothing while performing a large number of actions. The actions were recorded using a high-resolution body scanner (3dMD LLC, Atlanta, GA) and the scans were registered using the SMPL model. Similar to SNARF, the INS pipeline 400 trains a new model for each subject-cloth pair. It should be understood that exhaustive training for every combination would quickly become expensive. To manage computational costs, a subset of 15 sequences was used. This subset covers all clothing types and most subjects at least once, thereby capturing variations in both body shape and clothing.
[0102] DFAUST is a subset of the AMASS dataset, consisting of 10 minimally clothed subjects. Each subject was scanned similarly to CAPE while performing 10 different actions. Because the subjects wore minimal clothing, most of their motions can be accurately represented by rigid body transformations. Due to the subjects' minimal clothing, DFAUST was observed to contain significantly fewer pose-specific deformations. Therefore, the DFAUST dataset is not suitable for testing the true capabilities of the INS pipeline 400.
[0103] For a given subject in DFAUST or subject-clothing pair in CAPE, multiple time series are provided, each containing a different action. These series are split into training and test sets in a 9:1 ratio. This split is similar to SNARF. According to SNARF, we report the average intersection-over-union ratio of points sampled near the grid surface (IoU surface) and points sampled uniformly in space (IoU bbox).
[0104] In the canonical occupancy network, SNARF is used without pose conditioning as the first and main baseline. To this end, the pose conditioning used by SNARF is removed so that the canonical space 100 becomes pose-independent, i.e., O(p c No other changes were made. This setup is comparable to INS in that it allows for fast poses and maintains correspondence between different poses.
[0105] The INS method is also compared with the original SNARF, which uses a pose-conditioned occupancy network, i.e., O(p c ,θ t ). However, the above pose-conditional occupancy comes at the expense of fast pose posing, as expensive mesh extraction is required for each new pose while not preserving the correspondence between them. These drawbacks make a direct comparison between INS and SNARF based solely on their performance somewhat unbalanced.
[0106] In addition to the strong learning baselines described above, we also provide results for two simpler baselines that use SMPL-fitted LBS weights to un-pose the mesh (scan) by forward skinning. For the Average LBS baseline, the average of all normalized training meshes is taken to produce a final canonical mesh, which is deformed to an unseen given pose using forward LBS and SMPL weights. The First LBS baseline is similar to the Average LBS baseline described and uses SMPL-fitted weights for re-orienting the pose. Instead of taking the average of all training meshes, only the first mesh is used, thus containing less pose-conditioning details.
[0107] result
[0108] Results show that INS outperforms the state-of-the-art pose retargeting method, SNARF. On clothed human data, INS provides approximately 1% absolute gain compared to SNARF with pose conditioning and approximately 6% absolute gain compared to SNARF without pose conditioning. Experiments on simpler minimally clothed human data yield competitive results. INS is found to be an order of magnitude faster at retargeting long pose sequences. The following results also clearly demonstrate that INS can effectively correct for LBS artifacts.
[0109] Table 1 shows the results of INS on clothed human data (CAPE). Given the challenges of modeling the cloth deformations contained in this dataset, INS is found to outperform SNARF-NC (without pose conditioning) by an average of +6.24% and +6.41% absolute percentage points in surface IoU and bounding box IoU, respectively. Furthermore, INS also outperforms the pose-conditioned vanilla SNARF by +0.89% and +1.02% absolute percentage points in surface IoU and bounding box IoU, while also benefiting from fast pose generation and correspondence matching across a wide range of poses. It can be observed that a simple aggregate baseline of average LBS (AVG-LBS) closely matches the performance of SNARF-NC, with performance drops of only 1.88% and 1.66% percentage points. However, AVG-LBS benefits from the strong priors and correspondence fitting weights of the parameterized SMPL model.
[0110]
[0111] Table 1 Quantitative results for clothed human body
[0112] Table 2 shows the results of INS on a simpler, minimally clothed person from the DFAUST dataset. As shown, INS outperforms SNARF-NC (no pose conditioning) by an average of +3.37% and +0.63% absolute percentage points in the surface IoU and bounding box IoU metrics, respectively. Compared to SNARF with pose conditioning, INS lags behind by -1.42% and -0.86% absolute percentage points in the surface IoU and bounding box IoU metrics, respectively. Given the minimal clothing and minimal pose conditioning nonlinearities in DFAUST, this performance degradation can be attributed to SNARF's tendency to overfit to this baseline. This result also reflects the importance of testing on many real-world datasets, such as CAPE.
[0113]
[0114]
[0115] Table 2 Quantitative results for minimally clothed human bodies
[0116] Figure 5 is a graph 500 illustrating a comparison of redirection attitude times between the INS 510 and the SNARF 520 . Figure 5 The SNARF 520 and INS 510 are shown redirecting the attitude at 125 different target attitudes to 128 3 The time taken to extract the mesh of a clothed character at a resolution of 128 is shown in the figure. 3 Running a single mesh extraction process on a cube of 500 mm takes nearly 1.5 seconds. When repositioning, SNARF 520 performs this operation for each given pose, while INS 510 performs a single mesh extraction. Inferring the repositioned pose on the extracted mesh INS 510 takes 0.13 seconds, which is an order of magnitude faster than SNARF.
[0117] resection
[0118] We also perform multiple cutouts on the INS setting. Table 3 summarizes the INS cutout results for clothed subject 03375 (Table 1, row 1) in the CAPE dataset.
[0119]
[0120] Table 3 Resection table
[0121] The results show that adding pose and spatial embedding is important. The pose condition can be reformulated by a simple connection, i.e. [e z ,e θ ] instead of multiplying first and then concatenating, i.e. [e z ,ez ⊙e θ ] results in a significant performance drop of about 11% for both metrics (row 2 of Table 3).
[0122] The removal also shows that the reversible networks do not contribute equally to the performance. d Compared to removing the canonical space PIN H c This results in a sharp drop of 4.94% in IoU (rows 5 and 6 in Table 3). c The contribution of H is much greater than d This is partly due to the fact that editing the LBS deformed mesh introduces additional complexity in solving part correspondences, such as locating the new positions of joints and limbs.
[0123] Replacing the SIREN position embedding with an MLP can hurt the results. If the learned sinusoidal embedding is replaced with a simple MLP layer, the surface IoU drops by 3.16%. When using an MLP, fine surface details such as cloth wrinkles become blurred, while SIREN prevents this from happening.
[0124] Furthermore, PINs without rotations may perform slightly worse. Removing 2D operations from PINs results in a 0.92% drop in IoU. This is because the distortion in the surface must be represented solely by displacement, which has previously been shown to be difficult to learn.
[0125] Simply using PIN without LBS performs even worse. Completely removing the differential LBS module and relying solely on PIN to capture the full joint motion results in a substantial drop of 32% in both metrics (row 7 of Table 3).
[0126] Qualitative analysis
[0127] Compared to SNARF, INS can represent finer details. It was found that SNARF fails to capture fine details in cloth wrinkles and also lacks fingers. It was found that SNARF-NC combats LBS artifacts, such as volume loss, by shrinking the arched back and exhibiting a candy wrapper effect. It was found that INS can capture clearer local details around body joints, such as around the waist and neck.
[0128] PIN can well represent pose-dependent deformations. We found that PIN learns to introduce pose-dependent deformations, such as raising the cloth contours around the neck and shoulder joints, introducing clothing wrinkles near the limbs, and even adjusting the limbs, such as adjusting the orientation of the feet.
[0129] in conclusion
[0130] We present a reversible, end-to-end differentiable, and trainable pipeline for repositioning humans, called Reversible Neural Skinning, which includes a pose-conditioned invertible network (PIN) that robustly handles nonlinear surface deformations of clothing and skin while also preserving correspondence between different poses. We create INS by placing two PINs around a differentiable LBS network and using a pose-free occupancy network. We find that INS outperforms previous methods for clothed humans while remaining competitive for simple and minimally clothed humans. Since pose repositioning using INS requires only a single, expensive mesh extraction, INS provides an order of magnitude speedup when animating long pose sequences compared to previous methods.
[0131] Processing Platform
[0132] Figure 6 6 is an illustration of a machine 600 within which instructions 610 (e.g., software, programs, applications, applet, application, or other executable code) may be executed to cause the machine 600 to perform one or more of the methods discussed herein. For example, the instructions 610 may cause the machine 600 to perform any one or more of the methods described herein. The instructions 610 transform a general, unprogrammed machine 600 into a specialized machine 600 that is programmed to perform the functions described and illustrated in the manner described. The machine 600 may operate as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine 600 may operate as a server or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 600 may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook computer, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smartphone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a network device, a network router, a network switch, a network bridge, or any machine capable of executing instructions 610, sequentially or otherwise, that specify actions to be taken by the machine 600. Furthermore, while only a single machine 600 is shown, the term "machine" should also be construed to include a collection of machines that individually or jointly execute instructions 610 to perform one or more of the methodologies discussed herein. For example, the machine 600 may include: Figure 4A and 4B In some examples, the machine 600 may also include a client and server system, where some operations of a particular method or algorithm are performed on the server side and some operations of a particular method or algorithm are performed on the client side.
[0133] The machine 600 may include a processor 604, a memory 606, and an input / output I / O component 602, which may be configured to communicate with each other via a bus 640. In an example, the processor 604 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 608 and a processor 612 that execute instructions 610. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously. Although Figure 6 Multiple processors 604 are shown, but the machine 600 may include a single processor with a single core, a single processor with multiple cores (eg, a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0134] The memory 606 includes a main memory 614, a static memory 616, and a storage unit 618, all of which are accessible by the processor 604 via a bus 640. The main memory 606, the static memory 616, and the storage unit 618 store instructions 610 for one or more of the methods or functions described herein. During execution of the instructions by the machine 600, the instructions 610 may also reside, completely or partially, within the main memory 614, within the static memory 616, within the machine-readable medium 620 within the storage unit 618, within at least one of the processors 604 (e.g., within a cache memory of the processor), or within any suitable combination thereof.
[0135] The I / O components 602 may include a variety of components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurements, etc. The specific I / O components 602 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine may not include such a touch input device. It should be understood that the I / O components 602 may include Figure 6Many other components are not shown in the drawings. In various examples, the I / O components 602 may include user output components 626 and user input components 628. The user output components 626 may include visual components (e.g., displays such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), tactile components (e.g., vibration motors, resistive mechanisms), other signal generators, etc. The user input components 628 may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), a point-based input component (e.g., a mouse, a touch pad, a trackball, a joystick, a motion sensor, or other directional instrument), a tactile input component (e.g., a physical button, a touch screen or other tactile input component that provides the location and force of a touch or touch gesture), an audio input component (e.g., a microphone), and the like.
[0136] In further examples, the I / O component 602 may include a biometric component 630, a motion component 632, an environmental component 634, or a position component 636, as well as various other components. For example, the biometric component 630 includes components for detecting expressions (e.g., hand expressions, facial expressions, vocal expressions, body gestures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), identifying individuals (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), and the like. The motion component 632 includes an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, and a rotation sensor component (e.g., a gyroscope).
[0137] Environmental components 634 include, for example, one or more cameras (with still image / photo and video capabilities), lighting sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers that detect ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., infrared sensors that detect nearby objects), gas sensors (e.g., gas detection sensors for detecting hazardous gas concentrations to ensure safety or to measure pollutants in the atmosphere), or other components that can provide identification, measurements, or signals corresponding to the surrounding physical environment.
[0138] Position component 636 includes a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure from which altitude can be derived), an orientation sensor component (e.g., a magnetometer), and the like.
[0139] Communication can be implemented using a variety of technologies. The I / O components 602 also include a communication component 638 that is operable to couple the machine 600 to the network 622 or device 624 via corresponding couplings or connections. For example, the communication component 638 may include a network interface component or another suitable device that interfaces with the network 622. In further examples, the communication component 638 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, Components (e.g. Low power consumption), Components and other communication components to provide communication via other modes. Device 624 can be another machine or any of a variety of peripheral devices (e.g., a peripheral device coupled via USB).
[0140] In addition, the communication component 638 can detect the identifier or include components that can be used to detect the identifier. For example, the communication component 638 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional bar codes (such as universal product code (UPC) bar codes), multi-dimensional bar codes (such as Quick Response (QR) codes, Aztec codes, Data Matrix, Data Glyphs, MaxiCode, PDF417, Hypercode, UCC RSS-2D bar codes and other optical codes) or an acoustic detection component (e.g., a microphone for identifying tagged audio signals). In addition, various information can be derived via the communication component 638, such as location via Internet Protocol (IP) geolocation, location via Internet Protocol (IP), location information ... Location of signal triangulation, location of NFC beacon signals that can indicate a specific location via detection, etc.
[0141] Various memories (e.g., main memory 614, static memory 616, and memory of processor 604) and storage unit 618 may store one or more sets of instructions and data structures (e.g., software) that embody or are used by one or more methods or functions described herein. When executed by processor 604, these instructions (e.g., instructions 610) cause various operations to implement the disclosed examples.
[0142] Instructions 610 may be sent or received over network 622 using a transmission medium via a network interface device (e.g., a network interface component included in communications component 638) and using any of several well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 610 may be sent or received using a transmission medium via a coupling (e.g., a peer-to-peer coupling) to device 624.
[0143] Figure 7700 is a block diagram illustrating a software architecture 704 that may be installed on one or more devices described herein. The software architecture 704 is comprised of hardware, such as a machine 702 (see Figure 6 )) support, the machine 702 includes a processor 720, a memory 726, and an I / O component 738. In this example, the software architecture 704 can be conceptualized as a stack of layers, where each layer provides specific functionality. The software architecture 704 includes layers such as an operating system 712, a library 710, a framework 708, and an application 706. In operation, the application 706 invokes an API call 750 through the software stack and receives a message 752 in response to the API call 750.
[0144] The operating system 712 manages hardware resources and provides common services. The operating system 712 includes, for example, a kernel 714, services 716, and drivers 722. The kernel 714 acts as an abstraction layer between the hardware and other software layers. For example, the kernel 714 provides memory management, processor management (e.g., scheduling), component management, network and security settings, and other functions. Services 716 can provide other common services to other software layers. Drivers 722 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 722 may include display drivers, camera drivers, or Low power drivers, Flash memory drivers, serial communication drivers (such as USB drivers), drivers, audio drivers, power management drivers, etc.
[0145] The libraries 710 provide a common low-level infrastructure used by the applications 706. The libraries 710 may include system libraries 718 (e.g., C standard libraries) that provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. In addition, the libraries 710 may include API libraries 724, such as media libraries (e.g., libraries for supporting the rendering and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), graphics libraries (e.g., OpenGL frameworks for rendering graphical content on a display in two dimensions (2D) and three dimensions (3D), database libraries (e.g., SQLite for providing various relational database functions), network libraries (e.g., WebKit for providing web browsing functions), etc. The libraries 710 may also include a variety of other libraries 76 to provide many other APIs to the applications 706.
[0146] The framework 708 provides a common high-level infrastructure used by the applications 706. For example, the framework 708 provides various graphical user interface (GUI) functions, advanced resource management, and advanced location services. The framework 708 can provide a wide range of other APIs that can be used by the applications 706, some of which may be specific to a particular operating system or platform.
[0147] In an example, applications 706 may include a home application 736, a contacts application 730, a browser application 732, a book reader application 734, a location application 742, a media application 744, a messaging application 746, a game application 748, and various other applications (such as third-party applications 740). Applications 706 are programs that perform functions defined in the program. Various programming languages can be used to generate one or more applications 706, which are structured in various ways, such as object-oriented programming languages (such as Objective-C, Java, or C++) or procedural programming languages (such as C or assembly language). In a specific example, third-party applications 740 (for example, those used by entities other than the vendor of a particular platform using ANDROID) TM or IOS TM Applications developed with a software development kit (SDK) can be developed on mobile operating systems such as IOS TM ANDROID TM 、 In this example, third-party applications 740 can call API calls 750 provided by the operating system 712 to facilitate the functions described herein.
[0148] "Carrier signal" means any intangible medium that is capable of storing, encoding, or carrying instructions for execution by a machine, and includes digital or analog communications signals or other intangible media to facilitate communication of such instructions. Instructions may be sent or received over a network using a transmission medium via a network interface device.
[0149] "Client Device" means any machine that connects to a communications network to obtain resources from one or more server systems or other client devices. A Client Device may be, but is not limited to, a mobile phone, desktop computer, laptop computer, portable digital assistant (PDA), smartphone, tablet computer, ultrabook, netbook, notebook computer, multiprocessor system, microprocessor-based or programmable consumer electronic device, game console, set-top box, or any other communications device that a user can use to access a network.
[0150] "Communications network" means one or more parts of a network, which may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a part of the Internet, a part of the Public Switched Telephone Network (PSTN), a Plain Old Telephone Service (POTS) network, a cellular telephone network, a wireless network,
[0014] The present invention relates to a network, another type of network, or a combination of two or more such networks. For example, the network or a portion of the network may include a wireless or cellular network, and the coupling may be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or other type of cellular or wireless coupling. In this example, the coupling may implement any of various types of data transmission technologies, such as single carrier radio transmission technology (1xRTT), evolution data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rates for GSM evolution (EDGE) technology, the third generation partnership project (3GPP) including 3G, fourth generation wireless (4G) networks, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), worldwide interoperability for microwave access (WiMAX), long term evolution (LTE) standards, other standards defined by various standards development organizations, other long range protocols, or other data transmission technologies.
[0151] "Component" refers to a device, physical entity, or logic with boundaries defined by function or subroutine calls, branch points, APIs, or other techniques that provide for the partitioning or modularization of specific processing or control functions. A component can be combined with other components via its interface to perform a machine process. A component can be a packaged functional hardware unit designed for use with other components, and is also part of a program that generally performs a specific function among related functions. A component can constitute a software component (e.g., code contained on a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit capable of performing an operation and can be configured or arranged in some physical manner. In various examples, one or more computer systems (e.g., a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or a group of processors) can be configured by software (e.g., an application or application portion) to operate as a hardware component to perform certain operations described herein. A hardware component can also be implemented mechanically, electronically, or any suitable combination thereof. For example, a hardware component can include dedicated circuits or logic that are permanently configured to perform certain operations. A hardware component can be a dedicated processor, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). The hardware components may also include programmable logic or circuits that are temporarily configured by software to perform certain operations. For example, the hardware components may include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware components become specific machines (or specific components of machines) that are specifically customized to perform the configured functions, and are no longer general-purpose processors. It should be understood that the decision to mechanically implement the hardware components in a dedicated and permanently configured circuit or in a temporarily configured circuit (for example, configured by software) may be driven by cost and time factors. Therefore, the phrase "hardware components" (or "hardware-implemented components") should be understood to include tangible entities, that is, entities that are physically constructed, permanently configured (for example, hardwired) or temporarily configured (for example, programmed) to operate or perform certain operations described herein in some way. Considering the example that the hardware components are temporarily configured (for example, programmed), each of the hardware components does not need to be configured or instantiated at any one point in time. For example, in the case where the hardware components include a general-purpose processor that is configured as a special-purpose processor by software, the general-purpose processor can be configured as correspondingly different special-purpose processors (for example, including different hardware components) at different times. Software accordingly configures one or more specific processors, for example, to constitute a specific hardware component at one point in time and to constitute a different hardware component at a different point in time. Hardware components can provide information to other hardware components and can also receive information from other hardware components. Thus, the described hardware components can be considered to be communicatively coupled.When there are multiple hardware components at the same time, communication can be achieved by signal transmission (for example, through appropriate circuits and buses) between or among two or more hardware components. In the example where multiple hardware components are configured or instantiated at different times, for example, communication between such hardware components can be achieved by storing and retrieving information in a memory structure that multiple hardware components can access. For example, a hardware component can perform an operation and store the output of the operation in a memory device coupled to its communication ground. Then, later, another hardware component can access the memory device to retrieve and process the stored output. The hardware component can also initiate communication with an input or output device and can operate on resources (for example, information collection). The various operations of the example methods described herein can be performed at least in part by one or more processors, which are temporarily configured or permanently configured to perform related operations (for example, by software). Whether it is a temporary configuration or a permanent configuration, such a processor can constitute a processor-implemented component for performing one or more operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors.
[0152] Similarly, the method described herein can be implemented at least in part by a processor, wherein specific one or more processors are examples of hardware. For example, at least some operations of the method can be performed by one or more processors or the components implemented by the processor. In addition, one or more processors can also be used as the performance of supporting related operations in a "cloud computing" environment or as "software as a service" (SaaS). For example, at least some operations can be performed by a group of computers (as an example of a machine including a processor), and these operations can be accessed via a network (for example, the Internet) and via one or more appropriate interfaces (for example, API). The performance of some operations can be distributed between processors, not only resides in a single machine, but also is deployed on multiple machines. In some examples, the processor or the components implemented by the processor can be located in a single geographical location (for example, in a home environment, an office environment, or a server farm). In other examples, the processor or the components implemented by the processor can be distributed in multiple geographical locations.
[0153] "Computer-readable storage media" refers to both machine storage media and transmission media. Thus, these terms encompass both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" refer to the same thing and are used interchangeably in this disclosure.
[0154] “Machine storage media” refers to a single or multiple storage devices and media (e.g., centralized or distributed databases and associated caches and servers) that store executable instructions, routines, and data. Thus, the term shall be taken to include, but not be limited to, solid-state memory, optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and device storage media include non-volatile memory, including, by way of example, semiconductor memory devices (e.g., Electrically Erasable Programmable Read-Only Memory (EPROM)), Electrically Erasable Programmable Read-Only Memory (EEPROM), FPGAs, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms “machine storage media,” “computer storage media,” and “device storage media” expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are encompassed by the term “signal media.”
[0155] “Non-transitory computer-readable storage medium” refers to a tangible medium capable of storing, encoding, or carrying instructions for execution by a machine.
[0156] "Signal medium" means any intangible medium that is capable of storing, encoding, or carrying instructions for execution by a machine, and includes digital or analog communication signals or other intangible media to facilitate the communication of software or data. The term "signal medium" shall be taken to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. The terms "transmission medium" and "signal medium" have the same meaning and are used interchangeably in this disclosure.
Claims
1. A reversible neural skinning (INS) pipeline for animating a three-dimensional (3D) mesh of a deformable object, comprising: a first pose-conditional invertible network (PIN) that receives as input a given pose of the deformable object defined by a set of generic skeletons and maps canonical points of the deformable object in a pose-independent canonical space to points in a pose-dependent canonical space; a differentiable linear blend skinning (LBS) network that receives points in the pose-dependent canonical space and transforms the points to obtain deformed points in a novel pose of the deformable object; as well as A second PIN receives the novel pose of the deformable object and performs error correction on the deformation points in deformation space.
2. The INS pipeline of claim 1 , further comprising a pose-free canonical occupancy network, the mesh of the deformable object being extracted from the pose-free canonical occupancy network to obtain a mesh in the pose-independent canonical space.
3. The INS pipeline according to claim 2, wherein: Mesh vertices of a mesh extracted from the canonical occupancy network are re-posed using the set of universal skeletons by passing through the first PIN, the differential LBS network, and the second PIN.
4. The INS pipeline according to claim 3, wherein: The mesh is extracted once for different poses, and the extracted mesh is reoriented to the pose by warping it with the differential LBS network.
5. The INS pipeline of claim 1 , further comprising a neural representation from which the mesh is extracted to obtain a mesh in the pose-independent canonical space.
6. The INS pipeline according to claim 1, wherein: The first PIN and the second PIN are reversible to maintain the exact correspondence between input and output, and each of the first PIN and the second PIN includes a one-dimensional (1D) posture conditional coupling layer and a two-dimensional (2D) posture conditional coupling layer of a reversible neural network (INN) linked together.
7. The INS pipeline according to claim 1, wherein: During training of the INS pipeline, the second PIN receives input scans of the deformable object in different poses in the deformed space and receives poses corresponding to the input scans, the differentiable LBS network obtains pose correspondences with canonical points in the pose-independent canonical space from the given poses using the Broyden algorithm, and the first PIN maps points in the pose-dependent canonical space to canonical points in the pose-independent canonical space and passes the canonical points in the pose-independent canonical space to a pose-free occupancy network.
8. A method for animating a three-dimensional (3D) mesh of a deformable object using an invertible neural skinning (INS) pipeline, comprising: A first pose-conditional invertible network (PIN) receives as input a given pose of the deformable object defined by a set of general skeletons; Mapping, by the first PIN, canonical points of the deformable object in the pose-independent canonical space to points in the pose-dependent canonical space; transforming points in the pose-dependent canonical space to deformation points in a novel pose of the deformable object by a differentiable linear blend skinning (LBS) network; as well as Error correction is performed on the deformation point in the deformation space by the second PIN.
9. The method of claim 8, further comprising extracting a mesh of the deformable object from a pose-free canonical occupancy network to obtain a mesh of the deformable object in the pose-independent canonical space.
10. The method of claim 9, further comprising re-posing mesh vertices of the extracted mesh of the deformable object using the set of common skeletons by passing through the first PIN, the differential LBS network, and the second PIN.
11. The method of claim 10, further comprising extracting a mesh of the deformable object once for different poses, and repositioning the extracted mesh by warping the extracted mesh using the differential LBS network.
12. The method of claim 8, further comprising extracting a mesh of the deformable object from a neural representation to obtain a mesh from the pose-independent space.
13. The method according to claim 8 further includes linking a one-dimensional (1D) posture conditional coupling layer and a two-dimensional (2D) posture conditional coupling layer of a reversible neural network (INN) to form the first PIN and the second PIN, wherein the first PIN and the second PIN are reversible to maintain an exact correspondence between input and output.
14. The method according to claim 8 also includes training the INS pipeline by the following operations: receiving, by the second PIN, input scans of the deformable object at different poses in the deformed space, providing, by the second PIN, poses corresponding to the input scans, obtaining, by the differentiable LBS network, pose correspondences with canonical points in the pose-independent canonical space from the given poses, and mapping, by the first PIN, points in the pose-dependent canonical space to canonical points in the pose-independent canonical space, and passing the canonical points in the pose-independent canonical space to a pose-free occupancy network.
15. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processor, cause the processor to animate a three-dimensional (3D) mesh of a deformable object using an invertible neural skinning (INS) pipeline comprising a first pose-conditioned invertible network (PIN), a differentiable linear blend skinning (LBS) network, and a second PIN by performing the following operations: receiving as input a given pose of the deformable object defined by a set of general bones; Mapping canonical points of the deformable object in a pose-independent canonical space to points in a pose-dependent canonical space; transforming a point in the pose-dependent canonical space into a deformation point in a novel pose of the deformable object; as well as Error correction of the deformation point in the deformation space is performed.
16. The medium of claim 15, further comprising instructions that, when executed by the processor, cause the processor to perform operations comprising: extracting a mesh of the deformable object from a pose-free canonical occupancy network to obtain a mesh of the deformable object in the pose-independent canonical space.
17. The medium of claim 16, further comprising instructions that, when executed by the processor, cause the processor to perform operations comprising: re-posing mesh vertices of the extracted mesh of the deformable object using the generic skeleton set via passage through the INS pipeline.
18. The medium of claim 17, further comprising instructions that, when executed by the processor, cause the processor to perform operations comprising: extracting a mesh of the deformable object once for different poses, and repositioning the extracted mesh by warping the extracted mesh with the differential LBS network.
19. The medium of claim 15 further comprising instructions that, when executed by the processor, cause the processor to perform operations comprising: encoding each bone transformation in a given pose of the deformable object using an operation graph, the operation graph taking a six-dimensional (6D) input of concatenated three-dimensional (3D) translations and rotations; and obtaining a pose embedding by concatenating the outputs of each bone.
20. The medium of claim 15, further comprising instructions that, when executed by the processor, cause the processor to perform operations comprising training the INS pipeline by receiving input scans of the deformable object at different poses in the deformed space, providing poses corresponding to the input scans, obtaining pose correspondences for the canonical points in the pose-independent canonical space from the novel poses, and mapping points in the pose-dependent canonical space to canonical points in the pose-independent canonical space, and passing the canonical points in the pose-independent canonical space to a pose-free occupancy network.