Method of editing three-dimensional volumetric data

By separating the head portion and generating an SMPL model for 3D volumetric data, the method enables efficient editing of 3D volumetric data, addressing the challenges of temporal polymorphism and sequential editing.

WO2025135243A1PCT designated stage expired Publication Date: 2025-06-26KWANGWOON UNIVERSITY INDUSTRY ACADEMIC COLLABORATION FOUNDATION +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2023/021259
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2023-12-21
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Editing 3D volumetric data synthesized from multiple cameras is challenging due to its sequential nature and temporal polymorphism, making it difficult to modify and edit the data efficiently.

Method used

The method involves separating the head portion from the volumetric data, generating an SMPL model, transferring the head portion to the SMPL model, and then performing editing operations such as fitting, rigging, and retargeting to create a 3D edited model.

Benefits of technology

This approach facilitates efficient editing of 3D volumetric data by allowing for easier modification of clothing and rigging, reducing the time and cost associated with editing complex 3D mesh sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2023021259_26062025_PF_FP_ABST
    Figure KR2023021259_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method of editing three-dimensional volumetric data, in which three-dimensional mesh data configured in a sequence form of a three-dimensional model is edited by separating a head part from volumetric data, generating an SMPL model from the volumetric data, transferring the head part to the SMPL model to generate a three-dimensional model, and then performing an edit such as fitting, rigging, or retargeting. The method comprises the steps of: (a) receiving an input of a three-dimensional mesh data sequence configured by a series of consecutive frames; (b) estimating three-dimensional poses composed of joints and bones from the sequence of three-dimensional meshes; (c) separating a head part from the three-dimensional mesh of a key frame by using a three-dimensional pose; (d) estimating a three-dimensional mesh and pose data of an SMPL model for the key frame of the three-dimensional mesh sequence; (e) generating a three-dimensional basic model of the key frame by transferring the head part to the three-dimensional mesh of the SMPL model; (f) editing estimated three-dimensional poses of key frames and generating a sequence of three-dimensional poses of all the frames from the edited three-dimensional poses of the key frames; and (g) generating a three-dimensional editing model by editing the three-dimensional basic model of the key frame, and animating the three-dimensional editing model by reflecting the sequence of the three-dimensional poses of all the frames.
Need to check novelty before this filing date? Find Prior Art

Description

How to edit 3D volumetric data

[0001] The present invention relates to a method for editing 3D volumetric data, which comprises editing 3D mesh data composed in the form of a sequence of 3D models, separating a head portion from volumetric data, generating an SMPL model from the volumetric data, transferring the head portion to the SMPL model to generate a 3D model, and then performing editing such as fitting, rigging, and retargeting.

[0002] Using multiple point-of-view cameras, sequential shots are taken in time, and the captured multi-point images are merged in the same time order to create a single 3D mesh model. The 3D mesh models created in this way are sequenced in time. Sequences of 3D mesh models can vividly capture the appearance and movement of the subject being filmed, and are therefore widely utilized in various content production fields. Such data is generally referred to as 3D volumetric data [Non-patent Documents 1-4].

[0003] 3D volumetric data is synthesized into a 3D mesh using multiple cameras and is produced sequentially over time, making editing the captured subject nearly impossible. While 3D volumetric sequences offer the advantage of accurately recording the appearance and motion of the subject, they also present the disadvantage of being extremely difficult to edit and modify. Because 3D mesh data exists sequentially over time, modifying a single 3D mesh model requires modifying 3D mesh models across a large number of frames.

[0004] Typically, 3D mesh data synthesized photogrammetrically using multi-view images all have distinct mesh structures (or topologies). In other words, while a single object synthesized sequentially over time may appear to have the same shape, the mesh topology varies from frame to frame.

[0005] Therefore, consistently modifying a 3D mesh with such temporal polymorphism requires high cost and a long time.

[0006] (Non-patent Document 1) Guo, Kaiwen and Lincoln, Peter and Davidson, et al, The Relightables: Volumetric Performance Capture of Humans with Realistic Relighting, December 2019, ACM Trans. Graph., vol. 38. no. 6, p.19 https: / doi.org / 10.1145 / 3355089.3356571, doi = 10.1145 / 3355089.3356571

[0007] (Non-patent Document 2) Pietroszek, Krzysztof and Eckhardt, Christian, Volumetric Capture for Narrative Films, 26th ACM Symposium on Virtual Reality Software and Technology, 2020, https: / doi.org / 10.1145 / 3385956.3422116, doi = 10.1145 / 3385956.3422116

[0008] (Non-patent Document 3) Schreer, Oliver and Feldmann, Ingo and Ebner, et al, Advanced Volumetric Capture and Processing, SMPTE Motion Imaging Journal, 2019, vol. 128, no. 5, 18-24, doi=10.5594 / JMI.2019.2906835

[0009] (비특허문헌 4) Schreer, Oliver and Feldmann, Ingo and Renault, et al, Capture and 3D Video Processing of Volumetric Video, 2019 IEEE International Conference on Image Processing (ICIP), 2019, 4310-4314. doi=10.1109 / ICIP.2019.8803576

[0010] (비특허문헌 5) Cao, Zhe, et al. "OpenPose: realtime multi-person 2D pose estimation using Part Affinity Fields." arXiv preprint arXiv:1812.08008 (2018).

[0011] (비특허문헌 6) J. M. Singh and R. Ramachandra, "3D Face Morphing Attacks: Generation, Vulnerability and Detection," in IEEE Transactions on Biometrics, Behavior, and Identity Science, doi: 10.1109 / TBIOM.2023.3324684.

[0012] (비특허문헌 7) Choutas, Vasileios, et al. "Monocular expressive body regression through body-driven attention." Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part X 16. Springer International Publishing, 2020.

[0013] The purpose of the present invention is to solve the above-described problem, and to provide a method for editing 3D volumetric data, which edits 3D mesh data composed of a sequence of 3D models, separates a head portion from volumetric data, creates an SMPL model from the volumetric data, transfers the head portion to the SMPL model to create a 3D model, and then performs editing such as fitting, rigging, and retargeting.

[0014] In order to achieve the above object, the present invention relates to a method for editing 3D volumetric data, comprising the steps of: (a) receiving a 3D mesh data sequence composed of a series of consecutive frames; (b) estimating a 3D pose composed of joints and bones from the sequence of the 3D mesh; (c) separating a head part from a 3D mesh of a key frame using the 3D pose; (d) estimating a 3D mesh and pose data of an SMPL model for a key frame of the 3D mesh sequence; (e) transferring the head part to the 3D mesh of the SMPL model to generate a 3D basic model of the key frame; (f) editing the 3D poses of the estimated key frames and generating a sequence of 3D poses of the entire frame from the 3D poses of the edited key frames; and, (g) editing the 3D basic model of the key frame to generate a 3D edited model and animating the 3D edited model by reflecting the sequence of 3D poses of the entire frame.

[0015] As described above, according to the method for editing 3D volumetric data according to the present invention, by separating the head portion from the volumetric data and then transferring it to an SMPL model to create a 3D model and editing it, an effect of facilitating editing work such as costume fitting and rigging is obtained.

[0016] Figures 1a and 1b are block diagrams of the configuration of the entire system for implementing the present invention.

[0017] FIG. 2 is a flowchart illustrating a method for editing three-dimensional volumetric data according to an embodiment of the present invention.

[0018] FIG. 3 is a flowchart illustrating a method for estimating a pose of a three-dimensional mesh according to an embodiment of the present invention.

[0019] FIGS. 4A to 4C are exemplary screens showing a process for estimating a pose of a 3D mesh according to an embodiment of the present invention, wherein FIG. 4A is an exemplary screen for a 3D volumetric sequence, FIG. 4B is a projection image, and FIG. 4C is an exemplary screen for a 2D pose image.

[0020] FIG. 5a and FIG. 5b are exemplary diagrams illustrating a process for estimating a 3D pose in a 3D mesh according to an embodiment of the present invention. FIG. 5a is a projection image of an AABB box, and FIG. 5b is an exemplary diagram for pose error.

[0021] FIG. 6 is a detailed flowchart illustrating a step of separating a head from a three-dimensional mesh according to one embodiment of the present invention.

[0022] FIGS. 7A to 7D are exemplary images showing a process of separating a head from a 3D mesh according to one embodiment of the present invention, wherein FIG. 7A shows bone selection and direction vector calculation, FIG. 7B shows direction vector calculation from a point 1 / 3 of a bone to a vertex, FIG. 7C shows angle calculation between two direction vectors, and FIG. 7D shows head portion separation.

[0023] FIG. 8 is a flowchart illustrating a method for correcting a separated mesh according to one embodiment of the present invention.

[0024] FIGS. 9A to 9E are exemplary images showing a process for correcting a separated head portion according to an embodiment of the present invention, wherein FIG. 9A is a result of volumetric primary region separation, FIG. 9B is a result of vertex and face removal, FIG. 9C is a result of coordinate movement of a vertex, FIG. 9D is a mesh of a final corrected portion, and FIG. 9E is an exemplary image of a mesh of a final corrected head.

[0025] FIG. 10 is a diagram showing the input (2D image) and output (3D mesh in SMPL format) of the Expose deep learning model according to one embodiment of the present invention.

[0026] FIGS. 11A to 11E are exemplary screens for a process of creating a 3D editing model and motion according to one embodiment of the present invention, wherein FIG. 9A is an exemplary screen for a volumetric head part, FIG. 9B is an exemplary screen for a SMPL model torso, FIG. 9C is an exemplary screen for new model creation and clothing / shoe fitting, FIG. 9D is an exemplary screen for SMPL skeleton rigging, and FIG. 9E is an exemplary screen for body retargeting and animating.

[0027] FIGS. 12a and 12b are example images of an original 3D volumetric sequence and an edited 3D volumetric sequence according to one embodiment of the present invention.

[0028] Hereinafter, specific details for implementing the present invention will be described with reference to the drawings.

[0029] In addition, in describing the present invention, the same parts are given the same reference numerals, and their repeated description is omitted.

[0030]

[0031] First, examples of the configuration of the entire system for implementing the present invention will be described with reference to FIGS. 1a and 1b.

[0032] As shown in Fig. 1a, the editing method of 3D volumetric data according to the present invention (hereinafter referred to as the editing method) can be implemented by a program system on a computer terminal (10) that inputs and edits a 3D mesh sequence.

[0033] That is, the editing method can be implemented as a program system (30) on a computer terminal (10) such as a PC, smartphone, or tablet PC. In particular, the editing method is configured as a program system, and can be installed and executed on the computer terminal (10). The editing method provides a service for editing a 3D mesh sequence by utilizing the hardware or software resources of the computer terminal (10).

[0034] In addition, as another embodiment, as shown in FIG. 1b, the above editing method can be implemented by configuring a server-client system composed of an editing client (30a) and an editing server (30b) on a computer terminal (10).

[0035] Meanwhile, the editing client (30a) and editing server (30b) can be implemented according to a typical client-server configuration method. That is, the functions of the entire system can be divided according to the client's performance, the amount of communication with the server, etc. While described below as an editing system, it can be implemented in various divisions depending on the server-client configuration method.

[0036] Meanwhile, as another embodiment, the editing method may be implemented as a program, running on a general-purpose computer, or as a single electronic circuit, such as an ASIC (application-specific integrated circuit). Alternatively, it may be developed as a dedicated computer terminal dedicated solely to editing color-stable 3D mesh sequences. Other possible implementations are also possible.

[0037]

[0038] Next, a method for editing three-dimensional volumetric data according to an embodiment of the present invention will be described with reference to FIG. 2.

[0039] As shown in FIG. 2, the method for editing 3D volumetric data according to the present invention includes a step of receiving a 3D mesh sequence (S10), a step of estimating a pose of the 3D mesh sequence (S20), a step of separating a head portion based on pose information (S30), a step of generating an SMPL model from the 3D mesh sequence (S40), a step of generating a 3D base model (S50), and a step of editing using the 3D base model (S70). In addition, the editing step (S70) is further configured to include a step of fitting a garment to the 3D base model (S71), a step of performing rigging on the 3D base model to which the garment is fitted (S72), and a step of retargeting through rigging (S73). In addition, the method may further include a step of editing a sequence of 3D poses (S60).

[0040] In summary, when 3D volumetric data generated from a multi-view camera is input, the head is separated and a SMPL model is created from the 3D volumetric data. The head is then transferred to the SMPL model to create a body. This body is then dressed in clothing to create a new 3D base model with the same body shape and face as the original model. After inserting bone and joint information into this 3D base model, various source motions are input to generate a new 3D volumetric sequence.

[0041] First, 3D volumetric data, i.e., a 3D mesh sequence, is input (S10).

[0042] 3D mesh sequence or 3D volumetric data is 3D video data of a person, which is composed of a 3D mesh of multiple frames in successive time.

[0043] In particular, for a series of consecutive frames captured by re-view cameras, a 3D mesh model is generated for each frame, and a 3D mesh sequence is a sequence of frames of the generated 3D mesh model.

[0044] The 3D mesh of each frame corresponds to a model (or body model) that represents a person as a mesh.

[0045] Meanwhile, a series of frames can be divided into keyframes and intermediate frames. Keyframes are set by skipping a certain number of frames. For example, a keyframe can be set every ten consecutive frames. Intermediate frames represent the frames that exist between keyframes.

[0046]

[0047] Next, the pose of the 3D model, particularly the 3D skeleton information consisting of joints and bones (skeleton), is estimated from the 3D mesh (S20).

[0048] As shown in FIGS. 3 and 4a to 4c, first, in order to estimate the 3D pose of the 3D mesh, projection images (multi-view) are generated by viewing the 3D mesh from multiple directions (four directions including front, back, left, right, etc.) (S21). Next, the positions of 2D joints in the projection images are extracted using the OpenPose library (S22), and the approximate positions of 3D joints are generated by calculating intersection points in 3D (S23). Finally, a post-processing process is performed to extract the positions of high-precision 3D joints (S24).

[0049] Meanwhile, the pose of the 3D mesh is estimated for each frame.

[0050] First, the step (S21) of obtaining a projection image is described.

[0051] When estimating the positions of two-dimensional joints from images projected from multiple directions using OpenPose, the accuracy of joint positions estimated from images projected from the frontal direction can be the highest. Therefore, the spatial distribution of the three-dimensional coordinates of the points that make up the three-dimensional mesh is analyzed to find the frontal direction of the three-dimensional mesh, and the frontal direction is rotated so that it is parallel to the Z-axis. Principal Component Analysis (PCA) is used to find the frontal direction. PCA is used to find the principal components of distributed data.

[0052] Applying PCA to a 3D mesh yields 3D vectors for the x, y, and z axes, which can most simply represent the distribution of the 3D mesh. Since the distribution along the y-axis, which is the vertical direction of the object, is not necessary to find the front, the 3D mesh is projected onto the xz plane, and PCA is performed on this 2D plane. In PCA, the covariance matrix is ​​first found, and two eigenvectors for the matrix are obtained. The vector with the smallest eigenvalue among the two obtained eigenvectors indicates the front direction. Using the vector found through PCA, the front of the 3D mesh is rotated so that the z-axis becomes the z-axis.

[0053] After finding the front of the object, an AABB (Axis-aligned Bounding Box) is established to determine the projection plane in space. The process of projecting from a 3D to a 2D plane is to convert the world coordinate system to coordinates on the projection plane using the MVP (Model View Projection) matrix, which is a 4x4 matrix.

[0054] Next, the step (S22) of estimating a 2D pose from each projected 2D image is described.

[0055] Once four projection images are generated, a 2D skeleton is extracted using OpenPose [Non-patent Document 5].

[0056] OpenPose, a project presented at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2017, is a method developed at Carnegie Mellon University in the United States. Based on a Convolutional Neural Network (CNN), it is a library capable of extracting body, hand, and facial features of multiple people from photos in real time.

[0057] This project's key feature is its ability to quickly identify poses for multiple people. Before the release of OpenPose, estimating poses for multiple people primarily involved a top-down approach: detecting each person in a photo and repeatedly finding poses for each detected person.

[0058] OpenPose is a bottom-up approach that improves performance without repetitive processing. Bottom-up methods estimate all joints of a person, connect the positions of each joint, and then reconstruct the joint positions for each person. Typically, bottom-up methods face the challenge of determining which person a joint belongs to. To address this, OpenPose utilizes part affinity fields, which can infer the person to whom a body part belongs.

[0059] The results of extracting the skeleton using OpenPose are output as an image and JSON (JavaScript Object Notation) file.

[0060] Next, the 3D pose generation step (S23) and correction step (S24) through 3D intersection are described.

[0061] The process of reconstructing the 2D skeleton pixel coordinate system back to the 3D coordinate system calculates the extracted joint coordinates on four projection planes located in space. Connecting the matching coordinates on the four planes yields four coordinates that intersect in space. Figure 5a illustrates the extraction of the 3D joint of the left shoulder of a 3D body model.

[0062] Meanwhile, 2D pose estimation inevitably has errors, which cause projection lines to deviate from the intersection space. As illustrated in Figure 5b, the red projection line on the back can be confirmed to deviate from the intersection space when viewed from the front and side. Experimentally, the diameter of the intersection space is set to 3 cm. That is, after defining a 3D virtual sphere, if the virtual projection line does not pass through this space, the node along this virtual projection line is not included in the calculation that integrates the 3D nodes.

[0063] That is, the average point of the intersection space is set as the center, and only candidate coordinates within a pre-determined range (e.g., a sphere with a diameter l) from the center are selected, and other coordinates are excluded.

[0064] After defining points at each viewpoint for a 3D node using the candidate coordinates that were not removed, the average coordinate is calculated. The (x, z) coordinates are determined from above, and the y coordinate is determined from the side. The calculated (x, y, z) coordinates must match the (x, y) coordinates from the front. This process is illustrated in Figure 5b.

[0065] Figure 4c visually displays the stellate results for one frame on a 3D model.

[0066] Meanwhile, the pose is obtained from a specific frame of the 3D mesh sequence or from the 3D mesh of each keyframe.

[0067]

[0068] Next, the head portion is separated from the volumetric 3D mesh using pose information or joint information (S30). That is, the mesh area corresponding to the head portion is separated from the 3D mesh.

[0069] The separation process is divided into two stages. The first stage separates the head region using the direction vectors between the skeleton and mesh vertices. The second stage further refines the separation based on the results of the first stage.

[0070] First, the primary separation operation is explained.

[0071] As shown in Fig. 6, the process of separating the mesh area is performed by calculating the direction vector of the skeleton using the three-dimensional coordinates of the two joints that constitute each skeleton.

[0072] First, to separate the head and torso, a bone connecting the head and torso (the bone connecting the neck joint and the head joint) is selected (S31). Next, the vector from the neck joint to the head joint is calculated as the bone's direction vector (S32). Then, the direction vector from the starting point of the bone to each mesh vertex (hereinafter referred to as the vertex direction vector) is calculated (S33).

[0073] Next, the angle between the direction vector of the skeleton and the direction vector of the vertex is calculated (S34), and if the angle is within 90 degrees, the head part is separated by separating it into points of the corresponding bone (S35).

[0074] At this time, as shown in FIGS. 7a to 7d, the direction vector is calculated from the 1 / 3 point of the skeleton (1 / 3 point from the neck joint) to the mesh vertex.

[0075] In particular, the head separation task is performed on a 3D mesh of a specific frame of a 3D mesh sequence whose pose was previously obtained.

[0076]

[0077] Next, the secondary separation operation or correction operation is described.

[0078] Secondary region separation corrects the incorrectly separated error portions based on the first region separation results to obtain a mesh of a more finely separated region (or head region).

[0079] As shown in Fig. 8, if the head mesh separated according to the first separation process includes other regions (e.g., torso region), the corresponding vertices and faces of the other regions are removed.

[0080] In addition, if the shape of the head mesh is distorted and separated because the part corresponding to the head area that should be separated is excluded, the position and coordinates of the closest vertex (among the vertices of the head area) are detected from the vertex of the excluded area, and the detected vertex is moved to the excluded and distorted part.

[0081] Meanwhile, preferably, other areas or distorted sections can be selected and input by the operator. That is, the operator can review the results of the primary separation task and select and input other areas or distorted sections. Based on the operator's input, the correction process is automatically performed as described above.

[0082] Figures 9a to 9e illustrate exemplary images of each correction process. Figure 9a shows the result image of the first region separation operation, and Figure 9b illustrates the removal of vertices and faces from other regions. Furthermore, Figure 9c illustrates the movement of vertex coordinates, and Figure 9d illustrates the mesh of the final corrected portion. Furthermore, Figure 9e illustrates the mesh of the final corrected head.

[0083] Meanwhile, the head separation operation can be performed for each specific frame or for each 3D mesh of each keyframe. That is, the head part can be separated for each keyframe.

[0084]

[0085] Next, an SMPL model is generated from a 3D mesh sequence (S40).

[0086] A 3D model is created using the SMPL-X method by estimating body posture and shape, face, and hands from RGB images of 3D volumetric data using signal processing technology or deep learning networks.

[0087] That is, in a 3D mesh sequence, a 2D image (RGB image) is acquired from a 3D mesh of a specific frame (by projecting it onto a 2D plane). Then, a signal processing or deep learning model is used to capture a human body, face, and hands from the 2D image, and a 3D object in SMPL-X format is generated. If an algorithm or deep learning network that can generate a 2D image and SMPL mesh is used, an SMPL mesh model can be generated.

[0088] This process is shown in Fig. 10. As an example, as shown in Fig. 10, an SMPL model can be estimated from a 2D image by a deep learning model such as the expose model. ExPose (EXpressive POse and Shape rEgression) captures a human body, face, and hands from an RGB image of a human and generates a 3D Human in the SMPL-X format [Non-patent Document 7]. The network generates a body pose (θ b ), hand pose(θ h ), facial pose(θ f ), shape(β), and expression(ψ) are predicted.

[0089] SMPL (Skinned Multi-Person Linear Model) is a data format designed to precisely represent the human body as a three-dimensional mesh, and is widely used in the fields of artificial intelligence and graphics. SMPL-X is a model of SMPL that includes fingers.

[0090] The SMPL model uses an image to find a human body contained in the image, estimates the pose of the body, and then applies the estimated pose to a human body model that has already been defined, thereby outputting a human body model in the pose.

[0091] The SMPL model consists of shape parameters representing the body's appearance and pose parameters representing joints and skeletons. In other words, the SMPL model includes a three-dimensional mesh of the body's appearance and pose (skeleton and joint) information of the three-dimensional mesh.

[0092] In addition, after analyzing the features of the body in the image, the human body model is transformed to have features similar to those features, thereby ultimately creating a human body model in the form of a 3D mesh similar to the human included in the image.

[0093] 2D images do not contain all the three-dimensional information inherent in 3D objects or the human body. Therefore, 3D information generated through inference or prediction from 2D images is bound to contain errors.

[0094] In order to minimize errors in the process of converting such 2D images into 3D information (3D mesh), 3D information of each SMPL model is extracted from each 2D image of multiple viewpoints, and the most appropriate viewpoint is selected among them. In other words, the 2D images of multiple viewpoints have their respective accuracies (confidence) calculated. These are sorted in order of accuracies, and the 2D image with the highest accuracy is converted into 3D information (3D mesh). At this time, the 3D mesh of a specific frame of the volumetric is projected to various viewpoints to obtain 2D images of multiple viewpoints.

[0095] Meanwhile, SMPL models can be generated from a 3D mesh of a specific frame or each keyframe. That is, SMPL models can be generated from a specific frame and used continuously, or newly generated from a keyframe and used when necessary.

[0096]

[0097] Next, the head part separated from the 3D mesh is transferred to the head part of the SMPL model to create a 3D basic model (S50).

[0098] Using the previously separated head mesh, the volumetric head is transferred to the SMPL-X face to create a new 3D model (or 3D base model) to be edited.

[0099] Face transfer can be accomplished using conventional methods [Non-patent Document 6]. For example, face replacement, morphing, or deformation can be applied. Alternatively, deep learning technology can be used for transformation.

[0100] Meanwhile, preferably, for each keyframe, the head part of the corresponding keyframe is transferred to the SMPL model to create a 3D basic model of the corresponding frame.

[0101]

[0102] Next, the step (S60) of editing the pose sequence is described.

[0103] That is, it edits the 3D poses of keyframes or the sequence of those keyframes, and generates a sequence of 3D poses (of all frames) from the 3D poses of the edited keyframes.

[0104] Volumetric data is a sequence of 3D meshes composed of a series of frames. Through the pose estimation step (S20), the 3D pose of a specific frame can be estimated for the 3D mesh of that frame.

[0105] First, the 3D pose of each keyframe is estimated from each 3D mesh of a series of keyframes, thereby estimating the sequence of 3D poses of the series of keyframes. In other words, the 3D pose sequence of a series of keyframes can be estimated by estimating the 3D poses only for the key frames. In this case, the 3D pose of the intermediate frame can be estimated by interpolating the 3D pose of the key frame.

[0106] Next, the sequence of 3D poses of the estimated keyframes is edited to generate a new sequence of keyframes of 3D poses. In particular, the sequence of 3D poses of the estimated keyframes is edited. Editing the sequence of 3D poses consists of editing the 3D pose skeleton and editing the sequence.

[0107] Skeleton editing in a 3D pose involves editing the joints and bones of a 3D pose in a specific frame. Since a 3D pose consists of joints and bones, editing is possible by changing the positions of joints or bones.

[0108] Editing a sequence of 3D poses involves changing the order of 3D poses, inserting other 3D poses, or deleting existing 3D poses. For example, if there are keyframes kf1, kf2, kf3, kf4, ..., kfn, you can change the order of some of them, such as kf1, kf3, kf4, kf2, .... Alternatively, you can import and insert a sequence of 3D poses prepared in advance as a template. For example, by deleting the existing keyframe kf3 and inserting the sample sequence ks1, ks2, ks3, you can obtain a new sequence of kf1, kf2, ks1, ks2, ks3, kf4, ..., kfn.

[0109] Next, the 3D pose of the entire frame is estimated from the key frames of the new 3D pose, thereby generating a sequence of 3D poses of the entire frame (hereinafter referred to as the edited 3D pose sequence). That is, the 3D pose of the intermediate frames between the key frames is estimated by interpolation. The positions of the bones and joints of the intermediate frames are estimated from the bones and joints of the key frames by interpolation.

[0110]

[0111] Next, the step (S70) of editing using a 3D basic model is described.

[0112] Dress the 3D base model with clothes, accessories, shoes, etc., rig the skeleton of SMPL-X to add simple movements, and use the torso retargeting technology to move the rigged 3D model.

[0113] Figures 11a to 11e illustrate a process for editing a 3D basic model. In particular, Figures 11a to 11e illustrate a process for creating a 3D edited model and then animating it by reflecting a 3D pose sequence.

[0114] Specifically, as illustrated in FIG. 2, in the editing step (S70), a 3D edited model is first created by dressing the generated 3D basic model with new clothing, accessories, shoes, and other items of clothing (S71). Preferably, the 3D basic model can be edited for each keyframe to create a 3D edited model.

[0115] Next, rigging work is performed on the 3D editing model (S72). Since the 3D editing model has not been rigged separately, it cannot be moved. Therefore, rigging work is required for the 3D editing model. Preferably, rigging work can be performed on the 3D editing model for each keyframe.

[0116] A 3D editing model is created by fitting clothes to a 3D basic model, and the pose data of the SMPL model is used when rigging the 3D editing model.

[0117] Therefore, by rigging the SMPL-X skeleton to the newly created 3D editing model, a 3D model with an animable SMPL-X skeleton structure is finally created.

[0118] Next, the 3D model is animated by reflecting the sequence of 3D poses using motion retargeting (S73). That is, motion retargeting is performed by reflecting the 3D pose sequence edited in the previous force sequence editing step (S60).

[0119] Motion retargeting is a technique for applying motion data from one character to another. In other words, a sequence of 3D poses is 3D motion data, so a sequence of 3D poses can be applied to a 3D edited model for animation.

[0120] Meanwhile, a 3D basic model is created for each keyframe and edited to create a 3D edited model. Furthermore, the previously edited sequence of 3D poses can be reflected in the 3D edited model for animation. The edited sequence of 3D poses is the sequence from one keyframe to the next.

[0121] Additionally, the order of keyframes follows the order of the pose sequence. That is, based on the keyframe of the pose sequence, the animation is performed up to the next keyframe using the 3D editing model of the corresponding keyframe. Therefore, if the order of keyframes in the pose sequence is changed due to editing, the animation is performed using the 3D editing model of the keyframes according to the changed order.

[0122] Additionally, when new keyframes are inserted between keyframe k and keyframe k+1, the sequence of all frames up to the next keyframe (keyframe k+1) is animated using the three-dimensional editing model of keyframe k.

[0123] For such movements, a 3D mesh model is output over time. Using the output 3D model, the original synthesized 3D volumetric data can be replaced or inserted to create an edited 3D volumetric mesh sequence. At this point, adding new movements to an existing 3D volumetric sequence (or video) or modifying clothing, accessories, hairstyles, makeup, and other elements can create a variety of effects.

[0124] Figures 12a and 12b show the original 3D volumetric sequence and the edited and generated 3D volumetric sequence.

[0125] Finally, as shown in Figures 12a and 12b, a generated volumetric sequence can be obtained in addition to the original volumetric sequence. By combining and integrating these two, the volumetric sequence can be appropriately edited, allowing for editing and reprocessing of volumetric models that are otherwise difficult to edit and reprocess.

[0126] Above, the invention made by the inventor has been specifically described according to the above embodiments, but the present invention is not limited to the above embodiments, and it goes without saying that various modifications can be made without departing from the spirit thereof.

[0127] The present invention is applied to a technology for editing three-dimensional volumetric data, which can facilitate editing tasks such as clothing fitting and rigging by separating a head portion from volumetric data, transferring it to an SMPL model, and then creating and editing a three-dimensional model.

Claims

1. In the method of editing 3D volumetric data, (a) a step of receiving a three-dimensional mesh data sequence consisting of a series of consecutive frames; (b) a step of estimating a 3D pose composed of joints and bones from a sequence of the 3D mesh; (c) a step of separating the head part from the 3D mesh of the keyframe using the 3D pose; (d) a step of estimating the 3D mesh and pose data of the SMPL model for a specific frame of the 3D mesh sequence; (e) a step of transferring the head portion to the three-dimensional mesh of the SMPL model to create a three-dimensional basic model of the key frame; (f) a step of editing the 3D poses of the estimated key frames and generating a sequence of 3D poses of the entire frame from the 3D poses of the edited key frames; and, (g) A method for editing 3D volumetric data, characterized by comprising the step of editing a 3D basic model of the key frame to create a 3D edited model, and animating the 3D edited model by reflecting a sequence of 3D poses of the entire frame.

2. In paragraph 1, A method for editing three-dimensional volumetric data, characterized in that, in the step (c) above, a vector directed from the head joint to the neck joint is calculated as a direction vector of the skeleton, a direction vector of a vertex from a pre-determined point of the skeleton to a mesh vertex is calculated, and a head portion is separated by calculating an angle between the direction vector of the skeleton and the direction vector of the vertex.

3. In paragraph 2, In the step (c) above, after calculating the angle between the direction vectors to separate the head part, if the mesh of the head part includes other areas, the corresponding vertices and faces of the other included areas are removed. A method for editing three-dimensional volumetric data, characterized in that when a part corresponding to a head area to be separated is excluded, the position and coordinates of the closest vertex among the vertices of the head area are detected from the vertex of the excluded area, and the detected vertex is moved to the distorted part.

4. In paragraph 1, A method for editing 3D volumetric data, characterized in that in the step (f), a sequence of 3D poses of a series of key frames is estimated by estimating a 3D pose of each key frame from each 3D mesh of a series of key frames; a sequence of 3D poses of the estimated key frames is edited to generate a sequence of key frames of a new 3D pose; and a sequence of 3D poses of the entire frame is generated by estimating a 3D pose of the entire frame from the key frames of the new 3D pose.

5. In paragraph 4, In the step (f) above, the editing of the sequence of 3D poses is a method for editing 3D volumetric data, characterized in that the editing of the sequence of 3D poses of key frames is performed by changing the order of the 3D poses, inserting another 3D pose in the sequence, or deleting an existing 3D pose.

6. In paragraph 4, A method for editing 3D volumetric data, characterized in that in the step (f) above, a sequence of 3D poses of the entire frame is generated by estimating 3D poses of intermediate frames by interpolation from key frames of the 3D poses.

7. In paragraph 1, A method for editing 3D volumetric data, characterized in that in the step (g), a 3D editing model is created by fitting a garment to the 3D basic model, rigging the 3D editing model with a pose of an SMPL model, and animating the 3D editing model using motion retargeting.

8. In paragraph 7, A method for editing 3D volumetric data, characterized in that in the step (g) above, retargeting a 3D edited model of a specific keyframe using a 3D pose sequence from that frame to the next frame.

9. In paragraph 8, A method for editing 3D volumetric data, characterized in that in the step (g) above, the order of key frames of the 3D editing model follows the order of the pose sequence.

Citation Information

Patent Citations

  • Face image processing apparatus and computer program

    JP2011039869A

  • Method and apparatus for shape deforming surface based 3D human model

    KR1020100073174A

  • Method and apparatus for shape transferring of 3D model

    KR1020120071299A

  • Body shape and pose estimation via volumetric regressor for raw three dimensional scan models

    US20220101603A1

  • KR20230079256A