Machine Learning Based on Volume Capture and Mesh Tracking
Through volume capture and grid tracking technology, combined with machine learning and MoCAP animation workflow, time-consuming and expensive digital mannequin creation problems in the existing technology are solved, efficient and automated natural animation generation is achieved, and model fidelity and animation efficiency are improved.
Patent Information
- Application Number
- CN202180005113.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-31
- Filing Date
- 2021-03-31
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-03-31
AI Technical Summary
The prior art has the challenge of time-consuming, expensive and ineffective capturing facial expressions and body movements when creating realistic digital mannequins. In particular, video-based 4D scanners cannot generate novel facial expressions or body movements, and motion capture animations cannot automatically generate natural surface deformations.
High-quality 4D scanning is performed using volume capture system, combining grid tracking technology to establish time correspondence on face and full-body grid sequences, and establish spatial correspondence with 3D CG physics simulator through grid registration, surface deformation training is used using machine learning, and standard MoCAP animation workflow prediction and synthesis of deformation of natural animation.
It realizes efficient and automated generation of dynamic facial and full-body modeling with natural facial expressions and body language, reducing the workload of handmade production, and improving the fidelity of the model and the efficiency of animation.
Smart Images

Figure CN114730480B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATION(S)
[0002] This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Patent Application Serial No. 63 / 003,097, filed on March 31, 2020, entitled “VOLUMETRICCAPTURE AND MESH-TRACKING BASED MACHINE LEARNING 4D FACE / BODY DEFORMATION TRAINING,” which is incorporated herein by reference in its entirety for all purposes. Technical Field
[0003] The present invention relates to three-dimensional computer vision and graphics for the entertainment industry. More particularly, the present invention relates to acquiring and processing three-dimensional computer vision and graphics for use in film, TV, music, and gaming content creation. Background Art
[0004] Creating virtual humans is highly manual, time-consuming, and expensive. The recent trend is to efficiently create realistic digital human models using multi-view camera 3D / 4D scanners, rather than manually crafting CG artwork from scratch. Various 3D scanner studios (3Lateral, Avatta, TEN24, Pixel Light Effect, Eisko) and 4D scanner studios (4DViews, Microsoft, 8i, DGene) exist around the world for camera-based human digitization.
[0005] Photo-based 3D scanner studios consist of multiple arrays of high-resolution photographic cameras. Existing 3D scanning techniques are typically used to create assembly models and require hand-crafted animation because they don't capture deformations. Video-based 4D scanner studios (4D = 3D + time) consist of multiple arrays of high-frame-rate machine vision cameras. They capture natural surface dynamics, but due to the fixed video and motion, they cannot create novel facial expressions or body movements. Virtual actors need to perform many action sequences, which means a huge workload for the actors. Summary of the Invention
[0006] Mesh tracking-based dynamic 4D modeling for machine learning deformation training includes: using a volumetric capture system for high-quality 4D scanning, using mesh tracking to establish temporal correspondence on a sequence of 4D scanned face and full body meshes, using mesh registration to establish spatial correspondence between the 4D scanned face and full body meshes and a 3D CG physics simulator, and using machine learning to train the surface deformations as increments to the physics simulator. Deformations can be predicted and synthesized for natural animation using a standard MoCAP animation workflow. Animation and machine learning-based deformation synthesis using a standard MoCAP animation workflow includes single or multi-view of a MoCAP actor. Figure 2 3D video as input, solve the 3D model parameters for animation (excluding deformation) (3D solution), and predict 4D surface deformation from ML training based on the 3D model parameters obtained by 3D solution.
[0007] In one aspect, a method for programming in a non-transient state of a device includes: using mesh tracking to establish temporal correspondence on a 4D scanned face and full body mesh sequence, using mesh registration to establish spatial correspondence between the 4D scanned face and full body mesh sequence and a 3D computer graphics physics simulator, and using machine learning to train surface deformations as increments of the 3D computer graphics physics simulator. The method also includes using a volumetric capture system for high-quality 4D scanning. The volumetric capture system is configured to capture high-quality photos and videos simultaneously. The method also includes acquiring multiple separate 3D scans. The method also includes using standard motion capture animation to predict and synthesize deformations for natural animation. The method also includes using a single view or multiple views of the motion capture actor Figure 2 3D video as input, solve the 3D model parameters for animation, and predict 4D surface deformation from machine learning training based on the 3D model parameters obtained by 3D solution.
[0008] On the other hand, a device includes a non-volatile memory for storing an application and a processor coupled to the memory, the application being configured to: use mesh tracking to establish temporal correspondences on a 4D scanned face and full body mesh sequence, use mesh registration to establish spatial correspondences between the 4D scanned face and full body mesh sequence and a 3D computer graphics physics simulator, and use machine learning to train surface deformations as increments of the 3D computer graphics physics simulator, the processor being configured to process the application. The application is further configured to use a volumetric capture system for high-quality 4D scanning. The volumetric capture system is configured to capture high-quality photos and videos simultaneously. The application is further configured to acquire multiple separate 3D scans. The application is further configured to use standard motion capture animation to predict and synthesize deformations for natural animation. The application is further configured to: use a single view or multiple views of a motion captured actor Figure 23D video as input, solve the 3D model parameters for animation, and predict 4D surface deformation from machine learning training based on the 3D model parameters obtained by 3D solution.
[0009] In another aspect, a system includes: a volumetric capture system for high-quality 4D scanning and a computing device configured to: use mesh tracking to establish temporal correspondence on a 4D scanned face and full-body mesh sequence, use mesh registration to establish spatial correspondence between the 4D scanned face and full-body mesh sequence and a 3D computer graphics physics simulator, and use machine learning to train the 3D computer graphics physics simulator using surface deformations as increments. The computing device is further configured to predict and synthesize deformations for natural animation using standard motion capture animation. The computing device is further configured to: use a single view or multiple views of the motion captured actor Figure 2 The system uses 3D video as input, solves 3D model parameters for animation, and predicts 4D surface deformation from machine learning training based on the 3D model parameters obtained through 3D solving. The volumetric capture system is configured to capture high-quality photos and videos simultaneously.
[0010] In another aspect, a method for programming in a non-transient state of a device includes using a single view or multiple views of a motion capture actor. Figure 2 3D video as input, solve the 3D model parameters for animation, and predict 4D surface deformation from machine learning training based on the 3D model parameters obtained by 3D solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figures 1A-1B Illustrated is a flow chart of a method for implementing volumetric capture and mesh tracking based machine learning 4D facial / body deformation training for time / cost efficient natural animation of game characters, in accordance with some embodiments.
[0012] Figure 2 A block diagram of an exemplary computing device configured to implement a deformation training method according to some embodiments is illustrated. DETAILED DESCRIPTION
[0013] Unlike prior art 3D scanning techniques, the deformation training implementation described in this article is able to generate dynamic facial and full-body models with implicit deformations through machine learning (ML), i.e., the synthesis of arbitrary novel movements of facial expressions or body language with natural deformations.
[0014] The method described herein is based on photo-video capture from a "photo-video volumetric capture system." Photo-video based capture is described in PCT patent application PCT / US2019 / 068151, filed on December 20, 2019, entitled "PHOTO-VIDEO BASED SPATIAL-TEMPORALVOLUMETRIC CAPTURE SYSTEM FOR DYNAMIC 4D HUMAN FACE AND BODY DIGITIZATION," which is incorporated herein by reference in its entirety for all purposes. As described, the photo-video capture system is capable of capturing high-fidelity textures in sparse time, and between photo captures, a video is captured, and the video can be used to establish correspondences (e.g., transitions) between the sparse photos. The correspondence information can be used to implement mesh tracking.
[0015] Game studios use motion capture (MoCAP) in their animation workflows (ideally with face / body unified for natural motion capture), but it does not automatically generate animations with natural surface deformations (e.g., flesh dynamics). Typically, game computer graphics (CG) designers add handcrafted deformations (4D) on top of 3D rigged models, which is time-consuming.
[0016] Other systems generate natural animations with deformations, but still require a high level of manual (handcrafted) work. Such systems cannot be automated through machine learning (ML) training. Other systems can synthesize deformations for facial animation, but require new workflows (e.g., not compatible with standard MoCAP workflows).
[0017] This paper describes mesh tracking-based dynamic 4D modeling for ML deformation training. The mesh tracking-based dynamic 4D modeling for ML deformation training involves: using a volumetric capture system for high-quality 4D scanning, using mesh tracking to establish temporal correspondences on sequences of 4D scanned face and full body meshes, using mesh registration to establish spatial correspondences between the 4D scanned face and full body meshes and a 3D CG physics simulator, and using machine learning to train surface deformations as increments to the physics simulator. Deformations can be predicted and synthesized for natural animation using a standard MoCAP animation workflow. Machine learning-based deformation synthesis and animation using a standard MoCAP animation workflow include using single or multi-view of a MoCAP actor. Figure 2 3D video as input, solve 3D model parameters (3D solution) for animation (excluding deformation), and predict 4D surface deformation from ML training based on the 3D model parameters obtained by 3D solution.
[0018] When modeling, there are facial and body processes. As described herein, in some embodiments, both the facial and body processes are photo-based, meaning the input is captured using a photo camera rather than a video camera. Using the photo camera input, the model is generated by capturing many different poses with muscle deformations (e.g., arms up, arms down, body twisted, arms to the side, legs straight). Depending on the pose, one or more shapes are determined that best approximate the pose (e.g., using matching techniques). These multiple shapes are then fused to make the muscle deformations more realistic. In some embodiments, the captured information is sparsely populated. From this sparse sensing, the system can reverse-map the model's motion. This sparse sensing can be mapped to a dense modeling map, resulting in multiple iterations. A modeling designer generates the model from a photo (e.g., a 3D scan) and attempts to emulate 4D (3D + time) animation. The animation team generates the animation from the sparse motion capture, but because the sensing is sparse, mapping can be difficult, resulting in numerous iterations. However, improvements to this implementation are possible.
[0019] The face and body are modeled based on blank shapes, which are based on photos, and the photos are 3D-based (i.e., no deformation information). The surface has no transition states, as each state is a sparse 3D scan. The sensing is also sparse, but the high-quality model is animated in real time.
[0020] There are many examples of mesh tracking technology, such as U.S. Patent No. 10,431,000, titled “ROBUST MESHTRACKING AND FUSION BY USING PART-BASED KEY FRAMES AND PRIORI MODEL,” filed on July 18, 2017. Another example is Wrap4D, which takes a sequence of textured 3D scans as input and generates a sequence of meshes with consistent topology as output.
[0021] During capture time, there is 4D capture (e.g., being able to see the face and / or body) and it is possible to see how the muscles move. For example, a target subject can be asked to move and the muscles will deform. For very complex cases, this is difficult for animators to do. Any complex muscle deformations are learned during the modeling phase. This enables compositing in the animation phase. This also enables modifications to be incorporated into the current MoCAP workflow. Using ML, if the motion is captured during the production phase, then the system densifies the motion (e.g., deformations) because the sensing is sparse. In some embodiments, the deformations are already known from the photo-video capture.
[0022] Figures 1A-1BIllustrated is a flow chart of a method for implementing volumetric capture and mesh tracking based machine learning 4D facial / body deformation training for time / cost efficient natural animation of game characters, in accordance with some embodiments.
[0023] In step 100, a high-quality 4D scan is generated using a volumetric capture system. As described in PCT patent application PCT / US2019 / 068151, a volumetric capture system can simultaneously capture both photos and videos to generate a high-quality 4D scan. A high-quality 4D scan includes a denser camera view for high-quality modeling. In some embodiments, rather than utilizing a volumetric capture system, another system for capturing 3D content and temporal information is utilized. For example, at least two separate 3D scans are acquired. Continuing with this example, the separate 3D scans can be captured and / or downloaded.
[0024] In step 102, static 3D modeling is performed. Once high-quality information has been captured for the scan, linear modeling is performed using the static 3D model. However, because 4D capture (photo and video capture and time) is performed, correspondence can be quickly established, which can be used to generate the character model. Static 3D modeling begins with a raw image, which is then cleaned, style / personality features are applied, and texturing is performed to produce high-quality, distortion-free frames. High-frequency details are also applied.
[0025] In step 104, rigging is performed. Rigging is a technique in computer animation where a character is represented by two parts: a surface representation used to draw the character (e.g., a mesh or skin) and a hierarchical collection of interconnected parts (e.g., bones that make up a skeleton). Rigging can be performed in any manner.
[0026] In step 108, dynamic 4D modeling based on mesh tracking is implemented for ML deformation training. Low-quality video can be used to improve mesh tracking for temporal correspondence. A delta between the character model and the 4D capture can be generated. This delta can be used for ML deformation training. Examples of incremental training techniques include: Implicit Part Network (IP-Net), which combines detailed implicit functions and parametric representations to reconstruct a 3D model of a person that remains controllable and accurate even in the presence of clothing. Given a sparse 3D point cloud sampled on the surface of a clothed person, an Implicit Part Network (IP-Net) is used to jointly predict the outer 3D surface of the clothed person, the inner body surface, and semantic correspondences to a parametric body model. The correspondences are then used to fit the body model to the inner surface, which is then non-rigidly deformed (under a parameterized body+displacement model) to the outer surface to capture clothing, face, and hair details. An exemplary IP-Net is further described in Bharat Lal Bhatnagar et al., “Combining Implicit Function Learning and Parametric Models for 3D Human Reconstruction” (Cornell University, 2020).
[0027] Once mesh tracking is achieved (e.g., correspondences between frames are established), incremental information can be determined. This incremental information enables training. Based on the trained knowledge, synthesis is possible during the MoCAP workflow.
[0028] In step 110 , MoCAP information is acquired. This can be acquired in any manner. For example, standard motion capture can be implemented, where the subject wears a special suit with markers. Unified face / body MoCap can also be implemented. By capturing the face and body together, the fit is more natural.
[0029] In step 112, ML 4D solving and morphing synthesis from 2D video to 4D animation is implemented. MoCap information can be used for 4D ML solving and morphing synthesis. In some embodiments, an inverse mapping is applied. Solving involves mapping the MoCap information to the model. The input is sparse, but a dense mapping is solved. ML using volumetric capture data can be used for implicit 4D solving.
[0030] In step 114, retargeting is applied, wherein the character model is applied to 4D animation with natural deformations. Retargeting includes facial retargeting and full body retargeting.
[0031] In step 116 , rendering including shading and relighting is performed to render the final video.
[0032] In some embodiments, fewer or additional steps are implemented. In some embodiments, the order of the steps is modified.
[0033] Unlike previous implementations that use dense camera setups close to the target subject's face, the system described herein focuses on motion and uses a camera setup far from the target subject to capture sparse motion (e.g., skeletal motion). Furthermore, with the system described herein, a larger field of view is possible, with more body and facial animation.
[0034] Figure 2 A block diagram of an exemplary computing device configured to implement a deformation training method according to some embodiments is illustrated. Computing device 200 can be used to acquire, store, compute, process, transmit, and / or display information, such as images and videos. Computing device 200 can implement any aspect of deformation training. Generally speaking, a hardware structure suitable for implementing computing device 200 includes a network interface 202, memory 204, processor 206, (one or more) I / O devices 208, bus 210, and storage device 212. The choice of processor is not critical, as long as a suitable processor with sufficient speed is selected. Memory 204 can be any conventional computer memory known in the art. Storage device 212 can include a hard drive, CDROM, CDRW, DVD, DVDRW, high-definition disk / drive, ultra-high-definition drive, flash memory card, or any other storage device. Computing device 200 can include one or more network interfaces 202. Examples of network interfaces include network cards connected to an Ethernet or other type of LAN. The I / O device(s) 208 can include one or more of the following: a keyboard, a mouse, a monitor, a screen, a printer, a modem, a touch screen, a button interface, and other devices. The deformation training application(s) 230 for implementing the deformation training method may be stored in the storage device 212 and the memory 204 and processed in the manner in which applications are typically processed. Figure 2 More or fewer of the components shown in can be included in the computing device 200. In some embodiments, deformation training hardware 220 is included. Figure 2 The computing device 200 in FIG. 1 includes an application 230 and hardware 220 for the deformation training method, but the deformation training method can be implemented on the computing device using hardware, firmware, software, or any combination thereof. For example, in some embodiments, the deformation training application 230 is programmed in a memory and executed using a processor. In another example, in some embodiments, the deformation training hardware 220 is programmed hardware logic including gates specifically designed to implement the deformation training method.
[0035] In some embodiments, (one or more) deformation training applications 230 include several applications and / or modules. In some embodiments, a module also includes one or more sub-modules. In some embodiments, fewer or additional modules can be included.
[0036] Examples of suitable computing devices include personal computers, laptops, computer workstations, servers, mainframe computers, handheld computers, personal digital assistants, cellular / mobile phones, smart devices, game consoles, digital cameras, digital video cameras, camera phones, smartphones, portable music players, tablet computers, mobile devices, video players, video disc recorders / players (e.g., DVD recorders / players, high-definition disc recorders / players, ultra-high-definition disc recorders / players), televisions, home entertainment systems, augmented reality devices, virtual reality devices, smart jewelry (e.g., smart watches), vehicles (e.g., self-driving vehicles), or any other suitable computing device.
[0037] To utilize the morphing training method described herein, a device such as a digital camera / camcorder / computer is used to acquire content, and then the same device or one or more additional devices analyze the content. The morphing training method can be implemented automatically with or without user assistance to perform morphing training.
[0038] In operation, the deformation training method provides a more accurate and efficient deformation and animation method.
[0039] Using mesh tracking on a dynamic 4D model, it is possible to generate correspondences, enabling ML training. Without correspondence information, it may be impossible to determine deformation information (e.g., how the shoulder muscles deform as they move). Using mesh tracking on a 4D volume capture, deformation information can be determined. Once ML occurs, there are deltas for the face and body, and this information can then be used for animation. Animators use characters to tell a story. But because detailed information can be too burdensome, it is saved aside for later use. Animators initially use "light" models (e.g., no detailed information).
[0040] Sparse sensing is used for machine learning-based deformation synthesis and animation using a standard MoCAP animation workflow. Sparse sensing enables a wider field of view, enabling the face and body to be captured together. Rather than using time-consuming handcrafted information to fill in the gaps of sparse sensing, photo-video volumetric capture is used during the modeling phase to learn surface dynamic deformations, which are then used during the animation phase. This allows game studios to use their standard MoCAP workflow, providing efficiency and quality improvements in many aspects of the process.
[0041] Some examples of machine learning 4D face / body deformation training based on volumetric capture and mesh tracking
[0042] 1. A method of programming in a non-transient state of a device, comprising:
[0043] Use mesh tracking to establish temporal correspondence between 4D scanned faces and full-body mesh sequences;
[0044] Using mesh registration to establish spatial correspondence between 4D scanned face and full body mesh sequences and a 3D computer graphics physics simulator; and
[0045] Using machine learning to train surface deformations as increments in a 3D computer graphics physics simulator.
[0046] 2. The method of clause 1, further comprising performing high-quality 4D scanning using a volume capture system.
[0047] 3. The method of clause 2, wherein the volumetric capture system is configured to simultaneously capture high-quality photographs and videos.
[0048] 4. The method of clause 1, further comprising acquiring multiple separate 3D scans.
[0049] 5. The method of clause 1, further comprising using standard motion capture animation to predict and synthesize deformations for natural animation.
[0050] 6. The method of clause 1, further comprising:
[0051] Use single or multiple views of motion-captured actors Figure 2 D video as input;
[0052] Solving 3D model parameters for animation; and
[0053] Based on the 3D model parameters obtained through 3D solution, 4D surface deformation is predicted from machine learning training.
[0054] 7. A device comprising:
[0055] Non-volatile memory used to store applications used for:
[0056] Use mesh tracking to establish temporal correspondence between 4D scanned faces and full-body mesh sequences;
[0057] Using mesh registration to establish spatial correspondence between 4D scanned face and full body mesh sequences and a 3D computer graphics physics simulator; and
[0058] Using machine learning to train surface deformations as increments in a 3D computer graphics physics simulator; and
[0059] A processor is coupled to the memory, the processor being configured to process the application.
[0060] 8. The apparatus of clause 7, wherein the application is further configured to perform high-quality 4D scanning using a volume capture system.
[0061] 9. The apparatus of clause 8, wherein the volumetric capture system is configured to simultaneously capture high-quality photographs and videos.
[0062] 10. The apparatus of clause 7, wherein the application is further configured to acquire a plurality of separate 3D scans.
[0063] 11. The apparatus of clause 7, wherein the application is further configured to use standard motion capture animation to predict and synthesize deformations for natural animation.
[0064] 12. The apparatus of clause 7, wherein the application is further configured to:
[0065] Use single or multiple views of motion-captured actors Figure 2 D video as input;
[0066] Solving 3D model parameters for animation; and
[0067] Based on the 3D model parameters obtained through 3D solution, 4D surface deformation is predicted from machine learning training.
[0068] 13. A system comprising:
[0069] Volume capture system for high-quality 4D scans; and
[0070] A computing device configured to:
[0071] Use mesh tracking to establish temporal correspondence between 4D scanned faces and full-body mesh sequences;
[0072] Using mesh registration to establish spatial correspondence between 4D scanned face and full body mesh sequences and a 3D computer graphics physics simulator; and
[0073] Using machine learning to train surface deformations as increments in a 3D computer graphics physics simulator.
[0074] 14. The system of clause 13, wherein the computing device is further configured to use standard motion capture animation to predict and synthesize deformations for natural animation.
[0075] 15. The system of clause 13, wherein the computing device is further configured to:
[0076] Use single or multiple views of motion-captured actors Figure 2 D video as input;
[0077] Solving 3D model parameters for animation; and
[0078] Based on the 3D model parameters obtained through 3D solution, 4D surface deformation is predicted from machine learning training.
[0079] 16. The system of clause 13, wherein the volumetric capture system is configured to simultaneously capture high-quality photographs and videos.
[0080] 17. A method of programming in a non-transitory state of a device, comprising:
[0081] Use single or multiple views of motion-captured actors Figure 2 D video as input;
[0082] Solving 3D model parameters for animation; and
[0083] Based on the 3D model parameters obtained through 3D solution, 4D surface deformation is predicted from machine learning training.
[0084] The present invention has been described with reference to specific embodiments, which incorporate details that facilitate understanding of the principles of construction and operation of the invention. Reference herein to specific embodiments and details thereof is not intended to limit the scope of the appended claims. It will be apparent to those skilled in the art that various other modifications may be made in the embodiments chosen for illustration without departing from the spirit and scope of the invention as defined by the claims.
Claims
1. A method of programming in a non-transitory memory of a device, comprising: High-quality 4D scans of the human body and the entire body are performed using a volumetric capture system configured to simultaneously capture photographs and videos; Use mesh tracking to establish temporal correspondence between 4D scanned faces and full-body mesh sequences; Using mesh registration to establish spatial correspondence between the 4D scanned face and full body mesh sequences and a 3D computer graphics physics simulator that applies assembly to the face and full body to form a surface representation and a hierarchical collection of interconnected parts; and We use machine learning to train surface deformation as a 3D computer graphics physics simulator of the face and full body incrementally with sequences of 4D scanned face and full body meshes, including combining implicit functions and parametric representations to reconstruct 3D models.
2. The method of claim 1 , further comprising acquiring a plurality of separate 3D scans.
3. The method of claim 1 further comprising using standard motion capture animation to predict and synthesize deformations for natural animation.
4. The method of claim 1 , further comprising: Use single-view or multi-view 2D video of a motion-captured actor as input; Solve 3D model parameters for animation; as well as Based on the 3D model parameters obtained through 3D solution, 4D surface deformation is predicted from machine learning training.
5. A device comprising: Non-transitory memory for storing applications for: High-quality 4D scans of the human body and the entire body are performed using a volumetric capture system configured to simultaneously capture photographs and videos; Use mesh tracking to establish temporal correspondence between 4D scanned faces and full-body mesh sequences; Using mesh registration to establish spatial correspondence between the 4D scanned face and full body mesh sequences and a 3D computer graphics physics simulator that applies assembly to the face and full body to form a surface representation and a hierarchical collection of interconnected parts; and Using machine learning to train surface deformations as 3D computer graphics physics simulators of faces and full bodies incrementally with sequences of 4D scanned face and full body meshes, including combining implicit functions and parametric representations to reconstruct 3D models; as well as A processor is coupled to the memory, and is configured to process the application.
6. The apparatus of claim 5, wherein the application is further configured to acquire a plurality of separate 3D scans.
7. The apparatus of claim 5, wherein the application is further configured to use standard motion capture animation to predict and synthesize deformations for natural animation.
8. The apparatus of claim 5, wherein the application is further configured to: Use single-view or multi-view 2D video of a motion-captured actor as input; Solving 3D model parameters for animation; and Based on the 3D model parameters obtained through 3D solution, 4D surface deformation is predicted from machine learning training.
9. A system comprising: a volumetric capture system for performing high-quality 4D scans of the human body and the entire body, wherein the volumetric capture system is configured to simultaneously capture photographs and videos; as well as A computing device configured to: Use mesh tracking to establish temporal correspondence between 4D scanned faces and full-body mesh sequences; Using mesh registration to establish spatial correspondence between the 4D scanned face and full body mesh sequences and a 3D computer graphics physics simulator that applies assembly to the face and full body to form a surface representation and a hierarchical collection of interconnected parts; and We use machine learning to train surface deformation as a 3D computer graphics physics simulator of the face and full body incrementally with sequences of 4D scanned face and full body meshes, including combining implicit functions and parametric representations to reconstruct 3D models.
10. The system of claim 9, wherein the computing device is further configured to use standard motion capture animation to predict and synthesize deformations for natural animation.
11. The system of claim 9, wherein the computing device is further configured to: Use single-view or multi-view 2D video of a motion-captured actor as input; Solving 3D model parameters for animation; and Based on the 3D model parameters obtained through 3D solution, 4D surface deformation is predicted from machine learning training.
Citation Information
Patent Citations
Robust mesh tracking and fusion by using part-based key frames and priori model
US10431000B2
Modelling of nonlinear soft-tissue dynamics for interactive avatars
WO2019207176A1