Automatic blending of facial expressions and full-body poses for creating dynamic digital human models using an integrated photo-video volumetric capture system and mesh tracking
The integrated photo-video volumetric capture system addresses the inefficiencies in creating digital human models by automatically generating high-fidelity poses and tracking muscle deformations, resulting in more accurate and cost-effective digital human models.
Patent Information
- Application Number
- JP2023560157
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-29
- Filing Date
- 2022-03-31
- Publication Date
- 2025-05-19
- Estimated Expiration
- 2042-03-31
AI Technical Summary
Current methods for creating digital human models in the entertainment industry are highly manual, time-consuming, and expensive, particularly in capturing realistic facial expressions and full-body poses.
An integrated photo-video volumetric capture system that simultaneously acquires 3D and 4D scans, enabling the detection of muscle deformations and automatic generation of high-fidelity extreme poses, along with mesh tracking for consistent form registration across poses.
This approach significantly reduces the manual effort required for animation, enhances the accuracy and efficiency of blending between poses, and allows for the creation of realistic digital human models with reduced production costs.
Smart Images

Figure 0007679490000001 
Figure 0007679490000002 
Figure 0007679490000003
Abstract
Description
Technical Field
[0001] The present invention relates to graphics for three-dimensional computer vision and the entertainment industry. More specifically, the present invention relates to acquiring and processing three-dimensional computer vision and graphics for the creation of movie, television, music, and game content.
[0002] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Provisional Application No. 63 / 169,323, filed on April 1, 2021, entitled "AUTOMATIC BLENDING OF HUMAN FACIAL EXPRESSION AND FULL - BODY POSE FOR DYNAMIC DIGITAL HUMAN MODEL CREATION USING INTEGRATED PHOTO - VIDEO VOLUMETRIC CAPTURE SYSTEM AND MESH - TRACKING" under 35 U.S.C. § 119(e), and hereby incorporates by reference the entire disclosure of the provisional application for all purposes.
Background Art
[0003] Virtual humans are highly manual, time - consuming, and expensive. The recent trend is to efficiently create realistic digital human models using multi - view camera 3D / 4D scanners rather than starting from scratch with hand - crafted computer - graphic (CG) artworks. Various 3D scanner studios (3Lateral, Avatta, TEN24, Pixel Light Effect, Eisko) and 4D scanner studios (4DViews, Microsoft, 8i, DGene) for human digitization based on camera capture exist worldwide.
[0004] The photo-based 3D scanner studio includes a plurality of arrays of high-definition photo cameras. Conventional 3D scanning techniques are usually used to create rigged models, which do not capture deformations, so animation requires manual work. The video-based 4D scanner (4D = 3D + time) studio includes a plurality of arrays of high frame rate machine vision cameras. This captures the movement of natural surfaces, but due to fixed videos and actions, it cannot create unprecedented expressions or body movements. The dummy actors need to perform many series of movements, and the workload of the actors is very large.
Summary of the Invention
[0005] The integrated photo-video volumetric capture system for 3D / 4D scanning acquires 3D scans and 4D scans by simultaneously acquiring images and videos. The volumetric capture system for high-quality 4D scanning and mesh tracking is used to establish the correspondence of the form of the entire series of meshes to be 4D scanned for the generation of corrected shapes used in shape completion and skeleton-driven deformation. The volumetric capture system assists in mesh tracking to maintain mesh registration (consistency of form), in addition to the ease of modeling extreme poses. It can identify the major upper and lower body joints, which are important for generating deformations and capturing similar ones using a wide range of motions for all types of movement across all categories of joints. By using the volumetric capture system and mesh tracking, form changes are tracked. Each captured pose has a similar form that makes blending between multiple poses easier and more accurate.
[0006] In one aspect, a non-transitory programed method of a device comprises using a volumetric capture system configured for 3D scanning and 4D scanning that includes simultaneously capturing photos and videos, wherein the 3D scanning and 4D scanning include detecting deformation of an actor's muscles and performing mesh generation based on the 3D scanning and 4D scanning. The 3D scanning and 4D scanning comprise a 3D scan used for generating automatically high-fidelity extreme poses and a 4D scan with high temporal resolution that enables automatic registration of a mesh of extreme poses for blending in mesh tracking. Generating the automatically high-fidelity extreme poses includes using 3D scans of the actor and the deformation of the actor's muscles to generate the automatically high-fidelity extreme poses. The 4D scanning and mesh tracking are used to establish a correspondence of a series of 4D scanned meshes for generating corrected shapes for shape completion and skeleton-driven deformation. The method further comprises identifying and targeting the joints and muscles of the actor by a volumetric capture system for 3D scanning and 4D scanning. The mesh generation includes muscle evaluation or projection based on 3D scanning, 4D scanning, and machine learning. Performing the mesh generation includes using 3D scanning and 4D scanning to generate a mesh in an extreme pose that includes muscle deformation. The method further comprises performing mesh tracking that tracks morphological changes so that each captured pose has a similar morphology for blending between poses.
[0007] In another aspect, the apparatus is a non-transitory memory for storing an application, the application using a volumetric capture system configured for 3D scanning and 4D scanning, including simultaneously capturing photos and videos, the 3D scanning and 4D scanning including detecting deformation of an actor's muscles, and the memory including performing mesh generation based on the 3D scanning and 4D scanning, and a processor coupled to the memory, the processor being configured to execute the application. The 3D scanning and 4D scanning comprise a 3D scan used to generate automatically high-fidelity extreme poses, and a 4D scan including high temporal resolution that enables automatic registration of the mesh of the extreme poses for blending in mesh tracking. Generating the automatically high-fidelity extreme poses includes using a 3D scan of the actor or the deformation of the actor's muscles to generate the automatically high-fidelity extreme poses. The 4D scanning and mesh tracking are used to establish a correspondence of a series of 4D scanned meshes for generating a corrected shape for shape completion and skeleton-driven deformation. The application is further configured for identifying and targeting joints and the actor's muscles by a volumetric capture system for 3D scanning and 4D scanning. The mesh generation includes 3D scanning, 4D scanning, and muscle evaluation or projection based on machine learning. Performing the mesh generation includes using 3D scanning and 4D scanning to generate a mesh in extreme poses including muscle deformation. The application is further configured such that each captured pose has a similar form for blending between poses for performing mesh tracking for tracking morphological changes.
[0008] In another aspect, the system includes simultaneously capturing photos and videos, and 3D and 4D scanning includes detecting the deformation of the actor's muscles. A volumetric capture system for 3D and 4D scanning, and a computing device configured to receive the photos and videos captured from the volumetric capture system, and perform mesh generation based on 3D scanning and 4D scanning. 3D scanning and 4D scanning include a 3D scan used to generate automatically high-fidelity extreme poses, and a 4D scan with high temporal resolution that enables automatically registering the extreme poses for blending to mesh tracking. Generating automatically high-fidelity extreme poses includes using 3D scans of the actor and the deformation of the actor's muscles to generate automatically high-fidelity extreme poses. 4D scanning and mesh tracking are used to establish a correspondence in the form of a series of 4D scanned meshes for generating a corrected shape for shape completion and skeleton-driven deformation. The volumetric capture system is further configured to identify and target the joints and the actor's muscles by the volumetric capture system for 3D scanning and 4D scanning. Mesh generation includes muscle evaluation or projection based on 3D scanning, 4D scanning, and machine learning. Performing mesh generation includes using 3D scanning and 4D scanning to generate a mesh in an extreme pose including muscle deformation. The volumetric capture system is further configured such that each captured pose has a similar form for blending between poses to perform mesh tracking for tracking morphological changes.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0010] The automatic blending system utilizes an integrated photo - video volumetric capture system for 3D / 4D scans that acquires 3D scans and 4D scans by simultaneously capturing images and videos. The 3D scan can be used to generate automatically high - fidelity extreme poses, and the 4D scan includes high temporal resolution that enables mesh tracking to automatically register a mesh of extreme poses for blending.
[0011] A volumetric capture system (photo - video based) for high - quality 4D scanning and mesh tracking can be used to establish a correspondence of form across a series of 4D - scanned meshes for generating a corrective shape, where the corrective shape is used for shape completion and skeleton driven deformation. The photo - video system, unlike hand - made shape modeling that has manual shape generation but aids in registration, and 3D - scanning - based approaches that aid in shape generation but do not aid in registration, aids in mesh tracking to maintain mesh registration (consistency of form) in addition to the ease of extreme - pose modeling.
[0012] The methods described herein are based on photo-video capture from a "Photo-Video Volumetric Capture System". Photo-video based capture is described in PCT patent application PCT / US2019 / 068151, entitled "PHOTO-VIDEO BASED SPATIAL-TEMPORAL VOLUMETRIC CAPTURE SYSTEM FOR DYNAMIC 4D HUMAN FACE AND BODY DIGITIZATION", filed on December 20, 2019, the entire disclosure of which is incorporated herein by reference for all purposes. As described in the above application, the photo-video capture system is capable of capturing high-fidelity textures in a short amount of time, with video being captured while photos are captured, and this video can be used to establish correspondences (e.g., transitions) between the scattered photos. This corresponding information can be used to perform mesh tracking.
[0013] It is possible to identify the major joints of the upper and lower body, which is important for generating and capturing deformations using a wide range of motions for all types of movements of all types of joints. The joints can be used in the deformation of muscles. For example, by knowing how the joints move and how the muscles near the joints deform, the skeleton / joint information can be used for the deformation of the muscles used in mesh generation. As a further example, it is possible to use the images and videos obtained by having a video of the muscle deformation, and a more accurate mesh of the muscle deformation can be generated.
[0014] By using a photo-video system and mesh tracking, morphological changes can be tracked. In this way, each captured pose has the same form, which makes blending between multiple poses easier and more accurate.
[0015] FIG. 1 shows a flowchart of a method for animating a subject using a volumetric capture system for photographs / videos according to some embodiments. At step 100, mesh creation / generation is performed using an integrated volumetric photograph / video system. Mesh generation includes modeling and registration of extreme poses for blending. As described, an integrated photograph / video volumetric capture system for 3D / 4D scans acquires 3D scans and 4D scans by simultaneously acquiring images and videos of the subject / actor. The 3D scan can be used to generate automatically high-fidelity extreme poses, and the 4D scan includes a high temporal resolution such that it can automatically register a mesh of extreme poses for blending to mesh tracking. At step 102, skeleton fitting is performed. Skeleton fitting can be performed by any method, such as based on the trajectories of relevant markers. At step 104, skin weight painting is performed. Skin weight painting can be performed by any method, such as determining and painting the weights of each segment of the skin accordingly. At step 104, animation is performed. Animation is performed by any method. Depending on the embodiment, each step can be performed manually, semi-automatically, and automatically. In some embodiments, fewer or additional steps are performed. In some embodiments, the order of the steps is changed.
[0016] Figure 2 shows an example of a mesh generated by combining a neutral pose and an extreme pose according to some embodiments. Any standard pose, such as standing with the arms down, the arms raised, or the arms extended horizontally, can be considered a neutral pose. An extreme pose is a pose such as when the subject moves between standard poses. Extreme poses are captured by targeting specific parts of the human muscles, thereby enabling the generation of extreme shapes during the game development process. By targeting all muscle groups of the human body and using a photo / video system and mesh tracking, the problem of maintaining capture and mesh registration in the process of graphic game development can be solved.
[0017] When developing a new video game, a model for the game is captured. Actors usually come to the studio only once to record specific movements and action performances. The studio uses a photo / video volumetric capture system to comprehensively capture the muscle deformations of all actors. Furthermore, by using the existing kinematic movements and types of deformations that occur in the human body, similar deformations can be made to the corresponding mesh. The system can deform the model in the same way as human movement / deformation using previously captured neutral poses and additionally captured poses. Additionally, the system can be trained using kinematic movements, deformations, and / or other knowledge and data.
[0018] Figure 3 is a diagram showing an example of the correlation between human anatomy and computer graphics according to some embodiments. In human anatomy, the movement of the musculoskeletal system involves receiving signals from a person's motor cortex. At that time, muscle deformation occurs, and the muscle pulls on the bone, enabling joint / bone movement. In addition, there is movement of the skin / fat. In the computer graphics mesh, particularly by performing joint / bone movement, the motion driver triggers the movement of the animated character. Then, mesh deformation (Skeletal Subspace Deformation (SSD)) occurs, followed by mesh deformation (Pose Space Deformation (PSD)). An obvious correlation can be seen between human anatomy and the mesh generated using computer graphics.
[0019] Figures 4A - B are diagrams showing examples of muscle movement in some embodiments. Body parts bend at joints as shown, such as the head that bends at the neck, the hand that bends at the wrist, the finger that bends at the finger joint, the leg that bends at the knee, and the foot that bends at the ankle. In some embodiments, the movement of all joints can be assigned to 12 categories. In some embodiments, by classifying the joint movement into categories, accurate muscle deformation can be generated based on the classified movement. For example, when a character bends the knee, specific muscles in the leg deform, but machine learning can be used to deform the correct muscles at the appropriate time. This muscle movement is the type of movement that includes the range of motion performed by an actor. The muscle movement is the target of capture. [Figures 4A - B, DeSaix, Peter, et al. "Anatomy & Physiology (OpenStax)." (2013). (Retrieved from https: / / openlibrary-repo.ecampusontario.ca / jspui / handle / 123456789 / 331)]
[0020] Figure 5 shows examples of groups of major muscles according to some embodiments. The upper and lower bodies each have four joints (excluding finger / toe joints). The joints of the upper body include the shoulder, elbow, neck, and hand, and the joints of the lower body include the torso, waist, knee, and ankle. Each joint has a corresponding group of muscles. As will be described, these corresponding groups of muscles deform as the character moves. The muscles of the lower and upper bodies are the main targets for capture when the actor moves.
[0021] Figure 6 is a diagram showing examples of joint-based movement types for mesh capture according to some embodiments. There are many different movement types, and the angular range of motion (from 0 to 180 degrees) varies for each of the joints of the upper and lower bodies. By including various movement types, the desired muscles can be captured and used later when generating the mesh.
[0022] Figure 7 is a diagram showing examples of joint-based movement types for mesh capture according to some embodiments. Two of the 12 movement types are shown (flexion / extension and internal rotation / external rotation). In some embodiments, the angular range of motion can be selected from 0 degrees, 90 degrees, and 180 degrees, and in some embodiments, the angular range of motion can be adjusted more finely, such as to a specific value of an angle or a fractional angle.
[0023] Figure 8 is a diagram showing examples of extreme poses according to some embodiments. Image 800 shows six types of movements, such as raising the arms horizontally, raising the arms from below the waist to above the head, and extending the arms forward. Image 802 shows four joints and the target muscles.
[0024] Figure 9 is a diagram illustrating automatic blend shape extraction according to some embodiments. The pose parameter 900 is combined with the face action unit 902 to result in a 4D tracked mesh 904. The method of automatic blend shape extraction uses 4D scans of facial movements, thereby facilitating the character creation process and reducing production costs. As a method for 4D face scans, U.S. Application No. 17 / 411,432, filed on August 25, 2021, titled "PRESERVING GEOMETRY DETAILS IN A SEQUENCE OF TRACKED MESHES", etc. can be used, and the entirety of this application is incorporated herein by reference for all purposes. This provides a high-quality 4D tracked mesh of a moving face as shown in 904, and the pose parameter 900 can also be obtained from the tracked 4D mesh. The user may use control points or bones for pose reproduction. [Figure 9, center figure is from P.Ekman, Wallace V. Friesen, Joseph C. Hager, “Facial action coding system: A technique for the measurement of facial movement>>Psychology 1978, 2002. ISBN 0-931835-01-1 1.]
[0025] The face action unit is interesting. Using a 4D tracked mesh that includes various different available expressions, a set of character-specific face action units can be automatically generated. This can be regarded as the decomposition of the 4D mesh into dynamic pose parameters and static action units, where only the action units are unknown. For the problem of segmentation, machine learning techniques can be utilized.
[0026] Figure 10 illustrates a flowchart for performing mesh generation according to some embodiments. At step 1000, for high-quality 3D / 4D scanning, a volumetric capture system can be utilized. As described in PCT patent application PCT / US2019 / 068151, a volumetric capture system can simultaneously acquire photos and videos for high-quality 3D / 4D scanning. High-quality 3D / 4D scanning includes denser camera views for high-quality modeling. In some embodiments, instead of utilizing a volumetric capture system, a system for acquiring other 3D content and time information is utilized. For example, at least two separate 3D scans are acquired. As a further example, the separate 3D scans can be captured and / or downloaded.
[0027] During the capture time period, joint and muscle movement and deformation are acquired. For example, specific muscles and specific muscle deformations are captured over time. During the capture time period, specific joints of the actor and the muscles corresponding to those joints can be targeted. For example, the target subject / actor is required to move, and the muscles deform. This deformation of the muscles is captured either statically or during movement. This information obtained from the movement and deformation can be used in the system's learning so that the system can perform any movement of the character using the joint and muscle information. In very complex situations, it is very difficult for an animator to do this. Any complex muscle deformation is learned during the modeling stage. This enables synthesis at the animation stage.
[0028] In step 1002, mesh generation is executed. When high-quality information is captured for scanning, mesh generation including registration for extreme pose modeling and blending is executed. Using 3D scan information, high-quality extreme poses can be automatically generated. For example, frames between keyframes can be appropriately generated using 4D scan information including frame information between keyframes. The high temporal resolution of the 4D scan information enables mesh tracking for automatically registering the mesh of the extreme pose for blending. As another example, the 4D scan enables mesh generation of muscle deformation over time. Similarly, using machine learning that also includes joint information in addition to the corresponding muscle and muscle deformation information, a mesh including muscle deformation information can be generated even when the movement is not acquired by the capture system. For example, assume that an actor is required to perform actions of standing and jumping vertically and running for capture, but the capture system fails to acquire the actor's action of jumping while running. However, based on the acquired information including muscle deformation between the actions of standing and jumping vertically and running, and using machine learning with joint knowledge and other physiological information, a mesh for jumping while running including detailed muscle deformation can be generated. In some embodiments, mesh generation includes 3D scanning, 4D scanning, and muscle evaluation or projection based on machine learning.
[0029] The major joints of the upper and lower body, which are important for deformation generation and deformation capture, can be identified using a wide range of motions for all types of motion across all categories of joints.
[0030] By using a volumetric capture system and mesh tracking, morphological changes can be tracked. Thus, each captured pose will have a similar form, making blending between multiple poses easier and more accurate. The targeted joints and muscles can be utilized when generating the mesh.
[0031] In some embodiments, mesh generation includes static mesh generation based on 3D scan information, and this mesh can be modified / animated by using 4D scan information. For example, as the mesh moves over time, additional mesh information can be established / generated from 4D scan information and / or video content of machine learning information. As described, the transition between each frame of the animated mesh can maintain the form so that mesh tracking and blending are smooth. In other words, a morphological correspondence is established across a series of 4D scanned meshes to generate the corrective shapes used for shape completion and skeleton-driven deformation.
[0032] In some embodiments, fewer or additional steps are performed. In some embodiments, the order of the steps is changed.
[0033] Figure 11 shows an example of a block diagram of a preferred computing device that executes an automatic blending method according to some embodiments. The computing device 1100 can be used for acquiring, storing, computing, processing, communicating, and / or displaying information such as images and videos. The computing device 1100 can execute automatic blending in any manner. Generally, the hardware configuration suitable for the execution of the computing device 1100 includes a network interface 1102, a memory 1104, a processor 1106, an I / O device 1108, a bus 1110, and a storage device 1112. As long as an appropriate processor with sufficient speed is selected, the selection of the processor is not important. The memory 1104 can be a general computer memory well-known in the art. The storage device 1112 can include a hard drive, a CDROM, a CDRW, a DVD, a DVDRW, a high-density disk / drive, an ultra-HD drive, a flash memory card, or any other storage device. The computing device 1100 can include one or more network interfaces 1102. As an example of a network interface, it can include a network card connected to Ethernet or other LANs. The I / O device 1108 can include one or more keyboards, mice, monitors, screens, printers, modems, touchscreens, button interfaces, and other devices. The automatic blending application 1130 used for the execution of the automatic blending method is stored and executed in the storage device 1112 and the memory 1104 so that the application is generally executed. The computing device 1100 can include more or fewer components than those shown in Figure 11. In some embodiments, automatic blending hardware 1120 is included.The computing device 1100 in FIG. 11 includes an application 1130 and hardware 1120 for an automatic blending method, which can be executed on a computing device that combines hardware, firmware, software, or any combination thereof. For example, in some embodiments, the automatic blending application 1130 is programmed in memory and executed using a processor. As another example, in some embodiments, the automatic blending hardware 1120 is programmed hardware logic that includes gates specially designed to execute the automatic blending method.
[0034] In some embodiments, the automatic blending application 1130 includes a number of applications and / or modules. In some embodiments, a module may similarly include one or more sub-modules. In some embodiments, fewer or additional modules may be included.
[0035] Suitable examples of computing devices include personal computers, laptop computers, computer workstations, servers, mainframe computers, handheld computers, personal digital assistants, cellular / mobile phones, smart appliances, gaming consoles, digital cameras, digital camcorders, camera phones, smartphones, portable music players, tablet computers, mobile devices, video players, video disc recorders / players (e.g., DVD recorders / players, high-definition disc recorders / players, ultra HD disc recorders / players), televisions, home entertainment systems, augmented reality devices, virtual reality devices, smart jewelry (e.g., smartwatches), vehicles (e.g., self-driving vehicles), or other suitable computing devices.
[0036] To utilize the automatic blending method described herein, a device such as a digital camera / camcorder / computer is used to acquire content, and the same device or one or more additional devices further analyze the content. The automatic blending method can be automatically executed by the user himself or herself or without the user's involvement to perform automatic blending.
[0037] By executing this automatic blending method, a more accurate and efficient method of automatic blending and animation is provided. The automatic blending method is different from manual shape modeling that supports registration but has manual shape generation, and an approach based on 3D scans that supports shape generation but does not support registration. In addition to the ease of modeling extreme poses, a photo / video system that supports mesh tracking to maintain mesh registration (form consistency) is utilized. By using the photo / video system and mesh tracking, morphological changes can be tracked. Thus, each captured pose has a similar form that makes blending between multiple poses easier and more accurate.
[0038] Some embodiments of automatic blending of human expressions and full-body poses for dynamic digital human model creation using an integrated photo / video volumetric capture system and mesh tracking (Item 1) A method programmed non-temporarily in a device, Using a volumetric capture system configured for 3D scanning and 4D scanning including simultaneously capturing photos and videos, wherein the 3D scanning and 4D scanning include steps of detecting deformation of an actor's muscles, Executing mesh generation based on the 3D scanning and 4D scanning, A method comprising. (Item 2) The 3D scanning and 4D scanning are a 3D scan that generates automatically high-fidelity extreme poses, a 4D scan including high temporal resolution that enables registration of a mesh of extreme poses for blending automatically to mesh tracking, The method according to item 1, comprising. (Item 3) Generating automatically high-fidelity extreme poses includes using a 3D scan of the actor and deformation of the actor's muscles to generate the automatically high-fidelity extreme poses, the method according to item 2. (Item 4) 4D scanning and mesh tracking are used to establish a correspondence of form across a series of 4D scanned meshes for generating a corrected shape for shape completion and skeleton-driven deformation, the method according to item 2. (Item 5) The method according to item 1, further comprising identifying and targeting the joints and muscles of the actor by the volumetric capture system for 3D scanning and 4D scanning. (Item 6) Mesh generation includes muscle evaluation or projection based on the 3D scanning, 4D scanning, and machine learning, the method according to item 1. (Item 7) Performing mesh generation includes using the 3D scanning and 4D scanning to generate a mesh in an extreme pose including muscle deformation, the method according to item 1. (Item 8) The method according to item 1, comprising performing mesh tracking to track a morphological change to enable each pose captured for blending between poses to have a similar form. (Item 9) A non-transitory memory for storing an application, the application is A volumetric capture system configured for 3D scanning and 4D scanning, including simultaneously capturing photos and videos, wherein the 3D scanning and 4D scanning include detecting deformation of an actor's muscles, and executing mesh generation based on the 3D scanning and 4D scanning, a non-transitory memory therefor, a processor coupled to the memory and configured to process the application, and an apparatus comprising the same. (Item 10) The 3D scanning and 4D scanning are a 3D scan used to automatically generate high-fidelity extreme poses, and a 4D scan including high temporal resolution that enables automatic registration of the extreme pose mesh for blending in mesh tracking, The apparatus according to item 9, including the same. (Item 11) Generating an automatically high-fidelity extreme pose includes using the 3D scan of the actor and the deformation of the actor's muscles to generate an automatically high-fidelity extreme pose, for the apparatus according to item 10. (Item 12) The 4D scanning and mesh tracking are used to establish a morphological correspondence across a series of 4D scanned meshes to generate a correction shape for shape completion and skeleton-driven deformation, for the apparatus according to item 10. (Item 13) The application is further configured to identify and target joints and the muscles of a precursor actor by the volumetric capture system for 3D scanning and 4D scanning, for the apparatus according to item 9. (Item 14) Mesh generation includes the 3D scanning, 4D scanning, and machine learning-based muscle estimation and projection, for the apparatus according to item 9. (Item 15) The execution of mesh generation is the device according to item 9, including the 3D scanning and 4D scanning for generating a mesh in an extreme pose including muscle deformation of the actor. (Item 16) The application is further configured such that each captured pose can have a similar form for blending between poses, for the device according to item 9, to perform mesh tracking for tracking morphological changes. (Item 17) A scanning volumetric capture system for 3D and 4D scanning, including simultaneously capturing photos and videos, and the 3D scanning and 4D scanning including detecting muscle deformation of the actor, receiving the photos and videos captured from the volumetric capture system, performing mesh generation based on the 3D scanning and 4D scanning a computing device configured to comprising a system. (Item 18) The 3D scanning and 4D scanning are a 3D scan used to automatically generate high-fidelity extreme poses, a 4D scan including high temporal resolution that enables automatic registration of meshes of extreme poses for blending in mesh tracking, the system according to item 17, including. (Item 19) Generating automatically high-fidelity extreme poses includes using a 3D scan of the actor and muscle deformation of the actor to generate the automatically high-fidelity extreme poses, for the system according to item 18. (Item 20) 4D scanning and mesh tracking are used to establish morphological correspondences across a series of 4D scanned meshes to generate corrected shapes for shape completion and skeleton-driven deformation, the system of item 18. (Item 21) The volumetric capture system further comprises the system of item 17 configured to identify and target joints and the actor's muscles by the volumetric capture system for 3D scanning and 4D scanning. (Item 22) Mesh generation comprises the system of item 17 including the 3D scanning, 4D scanning, and machine learning-based muscle evaluation or projection. (Item 23) Performing mesh generation includes using the 3D scanning and 4D scanning to generate a mesh in extreme poses including muscle deformation, the system of item 17. (Item 24) The volumetric capture system further comprises the system of item 17 configured such that each captured pose has a similar morphology for blending between each pose to perform mesh tracking for tracking morphological changes.
[0039] The present invention has been described from the perspective of specific embodiments incorporating details to facilitate understanding of the principles of the structure and operation of the invention. Such reference in this specification to specific embodiments and their details is not intended to limit the scope of the claims appended hereto. It will be readily apparent to those skilled in the art that various other changes can be made in the embodiments selected for illustration without departing from the spirit and scope of the invention as defined by the claims.
Claims
1. 1. A method non-transiently programmed into a device, comprising: using a volumetric capture system configured for 3D and 4D scanning including simultaneously capturing photographs and video, the 3D and 4D scanning including detecting muscle deformations of the actor; performing mesh generation based on the 3D and 4D scanning; A method for providing the above.
2. The 3D scanning and 4D scanning include: 3D scanning to generate automatic high-fidelity extreme poses; 4D scans with high temporal resolution that allow for mesh tracking to automatically register meshes to extreme poses for blending; The method of claim 1 , comprising:
3. The method of claim 2 , wherein generating automatic high-fidelity extreme poses comprises using a 3D scan of the actor and muscle deformations of the actor to generate the automatic high-fidelity extreme poses.
4. The method of claim 2 , wherein 4D scanning and mesh tracking are used to establish morphological correspondences across a series of 4D scanned meshes for generating corrective shapes for shape completion and skeleton-driven deformation.
5. 10. The method of claim 1, further comprising identifying and targeting the actor's joints and muscles by the volumetric capture system for 3D and 4D scanning.
6. The method of claim 1 , wherein mesh generation includes muscle estimation or projection based on the 3D scanning, 4D scanning, and machine learning.
7. The method of claim 1 , wherein performing mesh generation includes using the 3D and 4D scanning to generate meshes in extreme poses including muscle deformations.
8. The method of claim 1 , comprising performing mesh tracking to track morphological changes to enable each captured pose to have a similar morphology for blending between poses.
9. A non-transitory memory for storing an application, said application being using a volumetric capture system configured for 3D and 4D scanning, including simultaneously capturing photographs and video, the 3D and 4D scanning including detecting muscle deformations of the actor; and performing mesh generation based on the 3D scanning and 4D scanning; and a non-transient memory for a processor coupled to the memory and configured to process the application; An apparatus comprising:
10. 3D and 4D scanning are 3D scans that are used to automatically generate high-fidelity extreme poses; 4D scans containing high temporal resolution, enabling automatic registration of meshes at extreme poses for blending into mesh tracking; The apparatus of claim 9 , comprising:
11. 11. The apparatus of claim 10, wherein automatically generating high fidelity extreme poses comprises using a 3D scan of the actor and muscle deformations of the actor to automatically generate high fidelity extreme poses.
12. The apparatus of claim 10 , wherein 4D scanning and mesh tracking are used to establish morphological correspondences across a series of 4D scanned meshes to generate corrective shapes for shape completion and skeleton-driven deformation.
13. The apparatus of claim 9 , wherein the application is further configured to identify and target joints and precursor muscles by the volumetric capture system for 3D and 4D scanning.
14. The apparatus of claim 9 , wherein mesh generation includes muscle estimation and projection based on the 3D scanning, 4D scanning, and machine learning.
15. The apparatus of claim 9 , wherein performing mesh generation comprises the 3D scanning and 4D scanning to generate meshes in extreme poses including muscle deformations.
16. The apparatus of claim 9 , wherein the application is further configured to perform mesh tracking for tracking morphological changes such that each captured pose can have a similar morphology for blending between poses.
17. a scanning volumetric capture system for 3D and 4D scanning, including simultaneously capturing photographs and video, the 3D and 4D scanning including detecting muscle deformations of the actor; receiving the captured photographs and video from the volumetric capture system; Perform mesh generation based on the 3D and 4D scanning A computing device configured to A system comprising:
18. The 3D scanning and 4D scanning include: 3D scans that are used to automatically generate high-fidelity extreme poses; 4D scans containing high temporal resolution, enabling automatic registration of meshes at extreme poses for blending into mesh tracking; The system of claim 17 , comprising:
19. 20. The system of claim 18, wherein automatically generating high fidelity extreme poses comprises using a 3D scan of the actor and muscle deformations of the actor to generate the automatic high fidelity extreme poses.
20. 20. The system of claim 18, wherein 4D scanning and mesh tracking are used to establish morphological correspondences across a series of 4D scanned meshes to generate corrective shapes for shape completion and skeleton-driven deformation.
21. 20. The system of claim 17, wherein the volumetric capture system is further configured for identifying and targeting joints and muscles of the actor by the volumetric capture system for 3D scanning and 4D scanning.
22. 20. The system of claim 17, wherein mesh generation includes muscle estimation or projection based on the 3D scanning, 4D scanning, and machine learning.
23. 20. The system of claim 17, wherein performing mesh generation includes using the 3D and 4D scanning to generate meshes in extreme poses including muscle deformations.
24. 20. The system of claim 17, wherein the volumetric capture system is further configured to perform mesh tracking to track morphological changes such that each captured pose has a similar morphology for blending between each pose.
Citation Information
Patent Citations
Method and apparatus for motion capture of dynamic object
KR1020110070058A
Photo-video based spatial-temporal volumetric capture system
WO2020132631A1