Method for encoding three-dimensional volume data
By estimating pose and shape data from 3D volume data and encoding only the residual data between the 3D estimation model and the original data, the method addresses the inefficiencies in existing 3D volume data encoding techniques, achieving reduced data amount and complexity.
Patent Information
- Application Number
- PCT/KR2023/021263
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-26
AI Technical Summary
Existing methods for encoding 3D volume data are inefficient due to the complexity and high capacity of 3D volumetric models, which makes it difficult to compress and store or transmit 3D volume data sequences effectively.
The proposed method estimates pose and shape data from 3D volume data, applies this data to a 3D template model to generate a 3D estimation model, and encodes only the residual data between the 3D estimation model and the original 3D volume data.
This approach significantly reduces the data amount and complexity of 3D volume models by encoding only the residual data, making it more efficient for storage and transmission.
Smart Images

Figure KR2023021263_26062025_PF_FP_ABST
Abstract
Description
Encoding method of three-dimensional volume data
[0001] The present invention relates to a method for encoding three-dimensional volume data, which comprises estimating pose and shape data from three-dimensional volume data, applying the pose and shape data to a three-dimensional template model to generate a three-dimensional estimation model, and encoding residual data from the three-dimensional volume data together with the three-dimensional estimation model.
[0002] In general, 3D volumetric data contains information about objects existing in three-dimensional space. Unlike 2D data, it includes height information, allowing for a more realistic representation of the shape of an object.
[0003] The scope of applications for 3D volumetric data is expected to expand further with technological advancements. Such data is used in computer graphics, medical imaging, and scientific simulations. For example, in the medical field, 3D images obtained through MRI or CT scans reveal detailed structures within the human body. These high-resolution data contain crucial medical information in each pixel (or voxel).
[0004] Processing and visualizing this massive 3D data requires high-performance computing resources. Therefore, advanced graphics processing technologies, efficient data storage and access methods, and complex algorithms are essential. Technologies such as GPU acceleration, parallel processing, and data compression are used for real-time visualization and analysis.
[0005] In particular, in the field of virtual humans, 3D volumetric data is essential for the creation and representation of virtual human models. Virtual human models are created based on 3D volumetric data, which can be used to express the virtual human model's appearance, facial expressions, and movements.
[0006] Additionally, 3D volumetric data is utilized in the production of virtual human content [Patent Document 1]. Using virtual human models, various virtual human content, such as movies, dramas, and advertisements, can be produced.
[0007] Typically, 3D volumetric models acquired directly through camera systems or similar means preserve the vivid shape and movement of the target object. In particular, such 3D volumetric models are composed of temporal sequences (multiple frames).
[0008] Compared to 2D video, 3D volumetric models are very complex and high-capacity because the volume data corresponding to each frame includes a normal map, material map, diffuse map, etc. in addition to the mesh (or point cloud) and texture. Moreover, 3D volumetric models are very complex temporally because the structure of the geometric mesh (or point cloud) that constitutes the 3D volume is different for each frame.
[0009] Because 3D volumetric data is so complex and high-volume, efficient encoding methods are needed to transmit or store sequences of 3D volumetric data. Because the meshes constituting each frame of 3D volumetric data can have very different geometric structures, applying temporal correlation to compression is difficult. Therefore, a different encoding method than conventional methods is needed.
[0010] (Patent Document 1) Korean Patent Document No. 10-2575567 (Published on September 7, 2023)
[0011] The purpose of the present invention is to solve the above-described problems, and to provide a method for encoding three-dimensional volume data, which estimates pose and shape data from three-dimensional volume data, applies the pose and shape data to a three-dimensional template model to generate a three-dimensional estimation model, and encodes residual data from the three-dimensional volume data together with the three-dimensional estimation model.
[0012] In order to achieve the above object, the present invention relates to a method for encoding three-dimensional volume data, comprising: (a) a step of inputting three-dimensional volume data; (c) a step of estimating a three-dimensional pose from the three-dimensional volume data; (d) a step of estimating a three-dimensional shape from the three-dimensional volume data; (e) a step of generating a three-dimensional estimation model by modifying a predefined three-dimensional template model using the estimated three-dimensional pose data (hereinafter referred to as pose estimation data) and the estimated three-dimensional shape data (hereinafter referred to as shape estimation data); (f) a step of calculating a residual of the three-dimensional volume data with respect to the three-dimensional estimation model; and, (h) a step of generating the pose and shape estimation data and the residual data as transmission data.
[0013] As described above, according to the encoding method of three-dimensional volume data according to the present invention, by encoding only residual data between a three-dimensional deformation model and three-dimensional volume data, the data amount of the three-dimensional volume model can be significantly reduced and the complexity can be reduced.
[0014] Figures 1a and 1b are block diagrams of the configuration of the entire system for implementing the present invention.
[0015] FIG. 2 is a flowchart illustrating a method for encoding three-dimensional volume data according to an embodiment of the present invention.
[0016] FIGS. 3A to 3F are images illustrating a process of encoding three-dimensional volume data according to an embodiment of the present invention, wherein FIG. 3A is an example image of a three-dimensional template model, FIG. 3B is three-dimensional volume data, FIG. 3C is an example image of pose and shape estimation data, FIG. 3D is an example image of three-dimensional volume model deformation, FIG. 3E is an example image of three-dimensional volume data comparison, and FIG. 3F is an example image of three-dimensional volume residual data.
[0017] Figure 4 is a flowchart illustrating a method for estimating a pose of a three-dimensional mesh according to one embodiment of the present invention.
[0018] FIGS. 5A to 5C are exemplary screens showing a process for estimating a pose of a 3D mesh according to an embodiment of the present invention, wherein FIG. 5A is an exemplary screen for a 3D volumetric image, FIG. 5B is a projection image, and FIG. 5C is an exemplary screen for a 2D pose image.
[0019] FIG. 6a and FIG. 6b are exemplary diagrams illustrating a process for estimating a 3D pose in a 3D mesh according to an embodiment of the present invention. FIG. 6a is a projection image of an AABB box, and FIG. 6b is an exemplary diagram for pose error.
[0020] Figures 7a to 7c are exemplary images showing a process of modifying a template model according to one embodiment of the present invention.
[0021] Figure 8 is a flowchart illustrating a method for decoding three-dimensional volume data according to one embodiment of the present invention.
[0022] Hereinafter, specific details for implementing the present invention will be described with reference to the drawings.
[0023] In addition, in describing the present invention, the same parts are given the same reference numerals, and their repeated description is omitted.
[0024]
[0025] First, examples of the configuration of the entire system for implementing the present invention will be described with reference to FIGS. 1a and 1b.
[0026] As shown in Fig. 1a, the coding method of three-dimensional volume data according to the present invention (hereinafter referred to as the coding method) can be implemented as a program system on a computer terminal (10) that receives three-dimensional volume data and codes it.
[0027] That is, the editing method can be implemented as a program system (30) on a computer terminal (10) such as a PC, smartphone, or tablet PC. In particular, the coding method is configured as a program system, and can be installed and executed on the computer terminal (10). The coding method provides a service for coding three-dimensional volume data by utilizing hardware or software resources of the computer terminal (10).
[0028] In addition, as another embodiment, as shown in FIG. 1b, the coding method can be implemented by configuring a server-client system composed of a coding client (30a) and a coding server (30b) on a computer terminal (10).
[0029] Meanwhile, the coding client (30a) and coding server (30b) can be implemented according to a typical client-server configuration method. That is, the functions of the entire system can be divided according to the client's performance, the amount of communication with the server, etc. While described below as a coding system, it can be implemented in various divisions depending on the server-client configuration method.
[0030] Meanwhile, as another embodiment, the coding method can be implemented as a program that operates on a general-purpose computer, or as a single electronic circuit, such as an ASIC (application-specific integrated circuit). Alternatively, the method can be developed as a dedicated computer terminal dedicated solely to editing color-stable 3D mesh sequences. Other possible implementations are also possible.
[0031]
[0032] Next, a method for encoding three-dimensional volume data according to an embodiment of the present invention will be described with reference to FIGS. 2 to 6a and 6b.
[0033] As shown in FIG. 2, the encoding method of three-dimensional volume data according to the present invention comprises a step of receiving three-dimensional volume data (S10), a step of quantizing the three-dimensional volume data (S20), a step of estimating a pose (S30), a step of estimating a shape (S40), a step of generating a three-dimensional deformation model from a three-dimensional template model (S50), a step of calculating a residual with the three-dimensional volume data (S60), and a step of generating transmission data (S80). In addition, the method may further comprise a step of minimizing the residual (S70).
[0034] Figures 3a to 3f illustrate a process of actually processing three-dimensional volume data according to the process of Figure 2.
[0035]
[0036] First, 3D volume data is input (S10).
[0037] 3D volumetric data represents a three-dimensional virtual human, and is structured in data formats such as meshes or point clouds. Specifically, 3D volumetric data refers to volumetric models directly acquired through camera systems or other means. However, 3D volumetric data is not limited to any specific method.
[0038]
[0039] Next, the 3D volume data is quantized (S20).
[0040] Since the mesh or point cloud that constitutes the 3D volume data can have infinite precision, a precision limit is imposed for encoding.
[0041] Preferably, the entire space of the three-dimensional volume data is divided into unit areas, and the three-dimensional volume data within each unit area is integrated and set as one representative value (e.g., mean / median value, etc.).
[0042]
[0043] Next, a 3D pose is estimated from the input 3D volume data to generate 3D pose estimation data (S30). The 3D pose (estimation) data is 3D skeleton data composed of joints and bones (skeleton).
[0044] That is, 3D pose data can be estimated directly from input 3D volume data, or 3D pose data can be estimated from a 2D image obtained from 3D volume data.
[0045] As an example, 3D volumetric data is projected onto a 2D plane to obtain a 2D image, and a 3D pose is estimated from the obtained 2D image. In this case, 3D pose estimation from a 2D body image can be achieved using artificial intelligence, such as a deep learning network, or a rule-based algorithm.
[0046] Meanwhile, 2D images do not contain all the 3D information of a 3D object (body). Therefore, errors are inevitable in 3D information generated through inference or prediction from 2D images. Therefore, pose data is estimated from 2D images of two or more viewpoints to correct or minimize errors. For example, pose data can be estimated from 2D images of each viewpoint, and the final value can be estimated by taking the average of overlapping pose data.
[0047] Additionally, as another embodiment, 3D pose estimation can be performed using 3D data as is. That is, 3D poses can be directly estimated from 3D volume data.
[0048] By estimating the 3D pose of 3D volume data, the motion of 3D volume data can be analyzed.
[0049] The following 3D pose estimation method is only one example and is not limited thereto.
[0050] As shown in Fig. 4, first, in order to estimate the 3D pose of the 3D volume data, projection images (multi-view) are generated by viewing the 3D volume data from multiple directions (four directions including front, back, left, and right) (S31). Next, the positions of the 2D joints in the projection images are extracted using the OpenPose library (S32), and the approximate positions of the 3D joints are generated by calculating the intersection point in 3D (S33). Finally, a post-processing step is performed to extract the positions of the high-precision 3D joints (S34).
[0051] Figures 5a to 5c illustrate a process of actually estimating 3D pose data according to the method of Figure 4.
[0052] First, the step (S31) of obtaining a projection image is described.
[0053] When estimating the positions of two-dimensional joints from images projected from multiple directions using OpenPose, the accuracy of joint positions estimated from images projected from the frontal direction can be the highest. Therefore, the spatial distribution of the three-dimensional coordinates of the points that make up the three-dimensional mesh is analyzed to find the frontal direction of the three-dimensional mesh, and the frontal direction is rotated so that it is parallel to the Z-axis. Principal Component Analysis (PCA) is used to find the frontal direction. PCA is used to find the principal components of distributed data.
[0054] Applying PCA to 3D volume data yields 3D vectors for the x, y, and z axes, which can most simply represent the distribution of the 3D volume data. Since the distribution along the y-axis, which is the vertical direction of the object, is not necessary to find the front, the 3D volume data is projected onto the xz plane, and PCA is performed on this 2D plane. In PCA, the covariance matrix is first found, and two eigenvectors for the matrix are obtained. The vector with the smaller eigenvalue among the two obtained eigenvectors indicates the front direction. Using the vector found through PCA, the front of the 3D volume data is rotated so that the z-axis becomes the z-axis.
[0055] After finding the front of the object, an AABB (Axis-aligned Bounding Box) is established to determine the projection plane in space. The process of projecting from a 3D to a 2D plane is to convert the world coordinate system to coordinates on the projection plane using the MVP (Model View Projection) matrix, which is a 4x4 matrix.
[0056] Next, the step (S32) of estimating a 2D pose from each projected 2D image is described.
[0057] Once four projection images are generated, a 2D skeleton is extracted using OpenPose [Non-patent Document 5].
[0058] OpenPose, a project presented at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2017, is a method developed at Carnegie Mellon University in the United States. Based on a Convolutional Neural Network (CNN), it is a library capable of extracting body, hand, and facial features of multiple people from photos in real time.
[0059] This project's key feature is its ability to quickly identify poses for multiple people. Before the release of OpenPose, estimating poses for multiple people primarily involved a top-down approach: detecting each person in a photo and repeatedly finding poses for each detected person.
[0060] OpenPose is a bottom-up approach that improves performance without repetitive processing. Bottom-up methods estimate all joints of a person, connect the positions of each joint, and then reconstruct the joint positions for each person. Typically, bottom-up methods face the challenge of determining which person a joint belongs to. To address this, OpenPose utilizes part affinity fields, which can infer the person to whom a body part belongs.
[0061] The results of extracting the skeleton using OpenPose are output as an image and JSON (JavaScript Object Notation) file.
[0062] Next, the 3D pose generation step (S33) and correction step (S34) through 3D intersection are described.
[0063] The process of reconstructing the 2D skeleton pixel coordinate system back to the 3D coordinate system calculates the extracted joint coordinates on four projection planes located in space. Connecting the matching coordinates on the four planes yields four coordinates that intersect in space. Figure 6a illustrates the extraction of the 3D joint of the left shoulder of a 3D body model.
[0064] Meanwhile, 2D pose estimation inevitably has errors, which can cause projection lines to deviate from the intersection space. As illustrated in Figure 6b, the red projection line at the back can be confirmed to deviate from the intersection space when viewed from the front and side. Experimentally, the diameter of the intersection space is preset to 3 cm, etc. In other words, after defining a 3D virtual sphere, if the virtual projection line does not pass through this space, the node along this virtual projection line is not included in the calculation that integrates the 3D nodes.
[0065] That is, the average point of the intersection space is set as the center, and only candidate coordinates within a pre-determined range (e.g., a sphere with a diameter l) from the center are selected, and other coordinates are excluded.
[0066] After defining points at each viewpoint for a 3D node using the candidate coordinates that were not removed, the average coordinate is calculated. The (x, z) coordinates are determined from above, and the y coordinate is determined from the side. The calculated (x, y, z) coordinates must match the (x, y) coordinates from the front. This process is illustrated in Figure 6b.
[0067] Figure 5c visually displays the stellate results for 3D volume data on a 3D model.
[0068]
[0069] Next, the 3D shape and clothing data of the body are estimated for the input 3D volume data to generate shape estimation data (S40).
[0070] That is, 3D shape data can be estimated directly from input 3D volume data, or 3D shape data can be estimated from a 2D image obtained from 3D volume data.
[0071] 3D shape data consists of 3D shape data of the body and clothing data.
[0072] The 3D shape data of the body is the data that determines whether the virtual human is thin or fat, short or tall, etc. Preferably, the 3D shape data of the body is the same shape data used in the SMPL model.
[0073] Additionally, the clothing data is data for a 3D human-worn clothing template model. The clothing template model is a template representing clothing, such as tops, bottoms, skirts, coats, and shoes. The clothing template model is pre-built. Preferably, the clothing template model can also include hats, hairstyles, and the like.
[0074] That is, using neural networks or other methods, the clothing template model is estimated from pre-built clothing template models based on 3D volumetric data or projected 2D images. Multiple clothing models can be estimated. For example, if a person is wearing a hat and shoes, multiple clothing models, including the hat, top, bottom, and shoes, will be estimated.
[0075] Additionally, the clothing template model can be adjusted by feature parameters such as (relative) length, size, and width. When the clothing template model is estimated, these specific parameters are also estimated, so that a clothing template model reflecting these features can be estimated.
[0076] In one embodiment, after mapping three-dimensional volume data onto a two-dimensional plane in space, shape data of the volume model is extracted from a two-dimensional image. That is, three-dimensional volume data is mapped onto at least one plane in space to obtain a two-dimensional image of at least one viewpoint. Then, shape data is estimated from at least one two-dimensional image.
[0077] 2D images do not contain all the three-dimensional information possessed by a 3D object (body). Therefore, 3D information generated through inference or prediction from 2D images inevitably contains errors. Therefore, shape data is estimated from 2D images taken from two or more viewpoints to compensate for or minimize these errors.
[0078] For example, 3D shape data can be extracted from each 2D image at multiple viewpoints, and the most appropriate viewpoint can be selected. In other words, the confidence level for each 2D image at multiple viewpoints is calculated, and 3D shape data is estimated from the 2D image with the highest confidence level.
[0079] As another example, shape data is estimated from a two-dimensional image at each point in time, and the final value is estimated using the average value of overlapping shape data, etc.
[0080] In particular, methods for estimating 3D shapes from 2D images (videos) can utilize artificial intelligence (AI) such as deep learning networks or rule-based algorithms. These 3D shape estimation methods are merely examples and are not limited thereto.
[0081] By estimating the shape of the three-dimensional volume data, the external shape of the three-dimensional volume data can be estimated.
[0082] Through the pose and volume estimation process (S30, S40), the appearance of a 3D volume model can be newly created or restored.
[0083] Morphological data can be in the form of latent variables generated through deep learning, or it can be data such as displacements that require adjustment for each region. In the former case, morphological data is estimated through deep learning, while in the latter case, it is estimated through a rule-based algorithm.
[0084]
[0085] Next, a 3D estimation model is created by transforming the 3D template model (S50).
[0086] A number of 3D template models (N) are defined in advance and stored (retained).
[0087] The 3D template model consists of a 3D body template model and a 3D clothing template model.
[0088] A 3D body template model is a model that represents a 3D virtual human, and can set various 3D virtual humans using pre-defined feature parameters. Preferably, the feature parameters of the 3D template model are parameters that represent the characteristics of the virtual human, and are divided into shape parameters and pose parameters. Shape parameters are variables that determine whether the virtual human is thin or fat, short or tall, etc. Pose parameters are variables that determine the movement or posture of the virtual human.
[0089] Additionally, the 3D clothing template model is a model that represents a 3D clothing, and can set a 3D clothing in a specific pose using pose data. In this case, the pose data uses previously estimated pose estimation data.
[0090] The pose and shape data estimated in the previous steps (S30, S40) are applied to the 3D template model to obtain a transformed 3D template model (or 3D body estimation model).
[0091] That is, pose and shape data are used to set the values of the characteristic parameters of a 3D body template model. By inputting these parameter values into the 3D body template model, a 3D body estimation model can be created. This 3D body estimation model will have a form very similar to the virtual human of the original 3D volume data.
[0092] Additionally, pose estimation data is applied to a 3D clothing template model to obtain a 3D clothing template model (or 3D clothing estimation model). That is, the 3D clothing template model can be expressed as a 3D clothing model with various poses depending on the pose. The 3D clothing model is estimated by applying the pose estimation data.
[0093] Additionally, a 3D estimation model is created by combining a 3D clothing estimation model with a 3D body estimation model. At this time, the clothing estimation model can be modified and combined according to the size or position of the 3D body estimation model.
[0094] In summary, the deformation of the 3D template model first deforms the 3D body template model by reflecting the body's pose estimation data and shape estimation data (S51), then deforms the clothing template model by reflecting the pose estimation data (S52), and then combines the deformed clothing template model with the deformed 3D body template model to generate a final 3D estimation model (S53).
[0095] Figures 7a to 7c illustrate three-dimensional volume data and a three-dimensional template model transformed therefrom. That is, Figure 7a is three-dimensional volume data, Figure 7b is a three-dimensional body template model transformed by body pose and shape estimation data, and Figure 7c is a model in which a clothing template model is combined with the body template model of Figure 7b.
[0096] Transforming a 3D template model can utilize a deep learning network or a rule-based algorithm. Preferably, a SMPL-X-based deep learning model is used to transform the 3D template model.
[0097] The SMPL (Skinned Multi-Person Linear) model is a type of data format created to precisely represent the human body as a 3D mesh, and is widely used in the fields of artificial intelligence and graphics. In addition, SMPL-X is a model that includes fingers in SMPL. The SMPL model searches for the human body included in a 2D image, estimates the body pose, and then applies the estimated pose to a predefined human body model to create a human body model that assumes the pose. In addition, after analyzing the features of the body in the image, the human body model is transformed to have features similar to those features. Through this, a 3D mesh-type human body model similar to the human included in the image can be ultimately created.
[0098]
[0099] Next, the residual of the 3D volume data for the 3D estimation model is calculated (S60).
[0100] That is, the difference, or residual, between 3D estimation models for 3D volumetric data is calculated. The residual, in this case, refers to the residual in 3D space (specifically, quantized space). For example, the residual can be calculated as the difference between the 3D location in the 3D volumetric data and the 3D location in the 3D estimation model.
[0101] The precision and quantity of the residual are determined depending on the size units into which the space is divided in the space quantization step (S20).
[0102]
[0103] Next, minimize the residual (S70).
[0104] If the residual of the 3D model does not reach a predetermined level (critical level), the pose and shape estimation steps (S30, S40) can be performed repeatedly. At this time, in the pose and shape estimation steps (S30, S40), the position and angle of the 2D plane (projection plane) defined in space are newly defined so that pose and shape estimation can be performed using new parameters.
[0105] That is, the position and angle of the two-dimensional projection plane or projection point are newly defined, and the pose and shape estimation steps (S30, S40) and subsequent steps (S50, S60) are repeated using the newly defined parameters.
[0106]
[0107] Next, transmission data is generated (S80).
[0108] The transmitted data includes pose estimation data, shape estimation data, and residual data of the 3D volume model. Additionally, if multiple 3D template models are used, information regarding which 3D template model was used must also be transmitted.
[0109] Pose estimation data, shape estimation data, and 3D volumetric model residual data can be spatially and temporally compressed and transmitted. Each data can be compressed and transmitted using a variety of conventional methods.
[0110]
[0111] Next, a method for decoding three-dimensional volume data according to an embodiment of the present invention will be described with reference to FIG. 8.
[0112] As shown in Fig. 8, first, 3D transmission data is received (S110).
[0113] 3D transmission data includes pose estimation data, shape estimation data, residual data of a 3D volume model, etc. Additionally, a 3D template model may be included.
[0114] Next, using the transmitted pose estimation data, a 3D template model (data shared in advance by the encoder and decoder) is animated and transformed into a volume model that performs the same motion as the transmitted pose (S120).
[0115] Next, the shape of the template model moving in the transmitted pose is delicately deformed using the transmitted shape estimation data (S130).
[0116] Next, residual data of a 3D volume model is added to a 3D template model whose appearance has been deformed, thereby restoring it to a model identical to the originally input 3D volume data.
[0117]
[0118] Above, the invention made by the inventor has been specifically described according to the above embodiments, but the present invention is not limited to the above embodiments, and it goes without saying that various modifications can be made without departing from the spirit thereof.
[0119] The present invention is applied to a technology for encoding a three-dimensional volume data, which can significantly reduce the amount of data of a three-dimensional volume model and reduce complexity by encoding only residual data between a three-dimensional deformation model and three-dimensional volume data.
Claims
1. In a method for encoding 3-dimensional volume data, (a) a step of inputting three-dimensional volume data; (c) a step of estimating a 3D pose from the 3D volume data; (d) a step of estimating a three-dimensional shape from the three-dimensional volume data; (e) a step of generating a 3D estimation model by modifying a predefined 3D template model using estimated 3D pose data (hereinafter referred to as pose estimation data) and estimated 3D shape data (hereinafter referred to as shape estimation data); (f) calculating the residual of the three-dimensional volume data for the three-dimensional estimation model; and, (h) A method for encoding three-dimensional volume data, characterized by including a step of generating the pose and shape estimation data and the residual data as transmission data.
2. In paragraph 1, A method for encoding three-dimensional volume data, characterized in that the method further comprises the step of (b) quantizing the three-dimensional volume data.
3. In paragraph 1, A method for encoding three-dimensional volume data, characterized in that in step (c) or step (d), the three-dimensional volume data is projected onto a two-dimensional plane to obtain a two-dimensional image, and three-dimensional pose or shape data is estimated from the two-dimensional image.
4. In paragraph 3, A method for encoding three-dimensional volume data, characterized in that in step (c) or step (d), the three-dimensional volume data is projected onto two or more two-dimensional planes to obtain two-dimensional images of two or more viewpoints, and pose or shape data is estimated from each of the two-dimensional images to correct or minimize errors.
5. In paragraph 1, The above 3D template model is composed of a 3D body template model expressing a 3D virtual human body and a 3D clothing template model expressing 3D clothing. The above shape estimation data consists of body shape estimation data and an estimated clothing template model. A method for encoding three-dimensional volume data, characterized in that in the step (e), the three-dimensional body template model is transformed by reflecting the pose estimation data and the body shape estimation data, the clothing template model is transformed by reflecting the pose estimation data, and the three-dimensional estimation model is generated by combining the transformed clothing template model with the transformed three-dimensional body template model.
6. In paragraph 1, A method for encoding three-dimensional volume data, characterized in that, in the step (f) above, a residual is calculated as the difference between the three-dimensional location in space of the three-dimensional volume data and the three-dimensional location in space of the three-dimensional estimation model.
7. In paragraph 3, The method is a three-dimensional volume data encoding method characterized in that (g) if the residual is below a predetermined threshold level, the projection position and angle are newly defined and steps (c) to (f) are repeated.
8. In paragraph 1, A method for encoding three-dimensional volume data, characterized in that in the step (h) above, when two or more three-dimensional template models are used, data on the three-dimensional template models used are included in the transmission data.
Citation Information
Patent Citations
Image processor and image processing program
JP2008015895A
Method and apparatus for shape transferring of 3D model
KR1020120071299A
Systems, methods, and devices for managing segmented media content
KR1020210102855A
Commercial vehicle cooling system
KR1020230109208A
A substrate processing apparatus
KR102294505B1