Mesh data transmission device, mesh data transmission method, mesh data reception device, and mesh data reception method
By pre-processing and encoding mesh data efficiently, the method addresses the challenges of transmitting and receiving mesh data in 3D applications, improving compression performance and reducing latency for VR, AR, and autonomous driving.
Patent Information
- Application Number
- PCT/KR2025/008910
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-25
- Filing Date
- 2025-06-25
- Publication Date
- 2026-01-02
AI Technical Summary
The sheer number of points in 3D space makes it difficult to efficiently transmit and receive mesh data, leading to high processing requirements and latency issues in applications like VR, AR, and autonomous driving.
A method involving pre-processing, encoding a base mesh, encoding displacement data, and encoding attribute data, including steps like reconstructing texture map arrangement, simplifying mesh data, and generating texture coordinates to improve compression efficiency.
Enhances the compression performance of dynamic mesh data, reducing latency and encoding/decoding complexity, enabling high-quality 3D services such as VR, AR, and autonomous driving.
Smart Images

Figure KR2025008910_02012026_PF_FP_ABST
Abstract
Description
Mesh data transmission device, mesh data transmission method, mesh data reception device, and mesh data reception method
[0001] The embodiments provide a method for providing 3D content to provide users with various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services.
[0002] Among 3D content, point cloud data and mesh data are collections of points in 3D space. However, the sheer number of points in 3D space makes it difficult to generate point cloud or mesh data.
[0003] That is, there is a problem that a lot of processing is required to transmit and receive 3D data with a large amount of points, such as point cloud data or mesh data.
[0004] The technical problem according to the embodiments is to provide a device and method for efficiently transmitting and receiving mesh data in order to solve the problems described above.
[0005] The technical problem according to the embodiments is to provide a device and method for resolving latency and encoding / decoding complexity of mesh data.
[0006] The technical problem according to the embodiments is to provide a device and method for efficiently performing encoding and decoding of a displacement vector.
[0007] However, the scope of the embodiments is not limited to the aforementioned technical tasks, and the scope of the embodiments may be expanded to other technical tasks that can be inferred by a person skilled in the art based on the entire contents of this document.
[0008] To achieve the above-described purpose and other advantages, an encoding method according to embodiments may include a step of pre-processing input mesh data, a step of encoding a base mesh of the pre-processed mesh data, a step of encoding displacement data of the pre-processed mesh data, and a step of encoding attribute data of the pre-processed mesh data.
[0009] According to embodiments, the pre-processing step may include, if the input mesh data includes a plurality of texture maps, a step of reconstructing the arrangement of the plurality of texture maps based on the similarity of the plurality of texture maps.
[0010] According to embodiments, the reconstructing step may reconstruct the arrangement of the plurality of texture maps by changing the order of texture map index information of the plurality of texture maps.
[0011] According to embodiments, the plurality of texture maps correspond to one frame, and signaling information for each vertex within the frame may include index information of a texture map referenced by the corresponding vertex and position information within the referenced texture map.
[0012] According to embodiments, the pre-processing step may further include a step of simplifying mesh data including the plurality of texture maps to generate simplified mesh data, a step of generating texture coordinates of each vertex of the simplified mesh data, and a step of performing fitting to make the simplified mesh data having the texture coordinates similar to the input mesh data after sub-dividing the simplified mesh data, thereby generating fitted sub-divided mesh data.
[0013] According to embodiments, the texture coordinate generation step may include a step of performing segmentation by combining polygons or vertices having similar characteristics based on characteristics of polygons or vertices constituting the simplified mesh data, a step of generating mesh patches of a current frame based on a set of polygons or vertices segmented in the step, and a step of generating texture coordinates of each vertex of the simplified mesh data by packing the mesh patches of the current frame onto a 2D image based on a mesh patch packing result of a previous frame in which mesh patch packing has been completed.
[0014] According to embodiments, the packing step may include a step of checking whether there is a mesh patch matching a current mesh patch of a current frame among mesh patches of the previous frame, and, if it is confirmed in the step of checking that the matching mesh patch is in the previous frame, a step of determining a mapping position of the current mesh patch on a 2D image based on a mapping position of a matched mesh patch in the previous frame, and a step of packing the current mesh patch at the determined mapping position on the 2D image.
[0015] According to embodiments, the mesh patch matching verification step may determine that a mesh patch matches the current mesh patch if, among the mesh patches of the previous frame, the direction of the current mesh patch is the same and the number of polygons and / or vertices constituting the mesh patch is similar.
[0016] According to embodiments, the step of encoding the attribute data may regenerate texture coordinates based on the texture coordinates of each vertex of the simplified mesh data and the texture coordinates of each vertex of the input mesh data to perform video encoding.
[0017] According to embodiments, the encoding device includes a memory and at least one processor connected to the memory, wherein the at least one processor can be configured to pre-process input mesh data, encode a base mesh of the pre-processed mesh data, encode displacement data of the pre-processed mesh data, and encode attribute data of the pre-processed mesh data.
[0018] According to embodiments, the at least one processor may, if the input mesh data includes a plurality of texture maps, reconstruct the arrangement of the plurality of texture maps based on the similarity of the plurality of texture maps.
[0019] According to embodiments, the at least one processor may simplify mesh data including the plurality of texture maps to generate simplified mesh data, generate texture coordinates of each vertex of the simplified mesh data, and perform fitting to make the simplified mesh data similar to the input mesh data by subdividing the simplified mesh data having the texture coordinates.
[0020] According to embodiments, the at least one processor may perform segmentation by combining polygons or vertices having similar characteristics based on characteristics of polygons or vertices constituting the simplified mesh data, generate mesh patches of the current frame based on a set of polygons or vertices segmented in the step, and pack the mesh patches of the current frame onto a 2D image based on a mesh patch packing result of a previous frame in which mesh patch packing has been completed, thereby generating texture coordinates of each vertex of the simplified mesh data.
[0021] According to embodiments, a computer-readable storage medium can store a bitstream generated by the encoding method.
[0022] According to embodiments, a transmission method may include: obtaining a bitstream for mesh data; generating the bitstream based on the steps of: pre-processing input original mesh data; encoding a base mesh of the pre-processed mesh data; encoding displacement data of the pre-processed mesh data; and encoding attribute data of the pre-processed mesh data; and transmitting data including the bitstream.
[0023] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can provide a quality 3D service.
[0024] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can achieve various video codec methods.
[0025] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can provide general-purpose 3D content such as autonomous driving services.
[0026] The mesh data transmission method and mesh data transmission device according to the embodiments can regenerate a texture map with high inter-frame image correlation by reflecting image similarity between frames when generating texture coordinates of a simplified mesh. This can improve the compression performance of the dynamic mesh of V-Mesh, and in particular, can improve the compression performance of the texture map video of the mesh.
[0027] The drawings are included to further understand the embodiments, and the drawings illustrate the embodiments together with the description related to the embodiments. For a better understanding of the various embodiments described below, reference should be made to the following description of the embodiments in conjunction with the following drawings, in which like reference numerals correspond to corresponding parts throughout the drawings.
[0028] Figures 1(a) and 1(b) are diagrams showing examples of an encoder and a decoder according to embodiments.
[0029] FIG. 2 illustrates a system for providing dynamic mesh content according to embodiments.
[0030] Figure 3 illustrates a V-MESH compression method according to embodiments.
[0031] Figure 4 illustrates pre-processing of V-MESH compression according to embodiments.
[0032] Figure 5 illustrates a mid-edge subdivision method according to embodiments.
[0033] Figure 6 shows a displacement generation process according to embodiments.
[0034] Figure 7 illustrates an encoding process of mesh data according to embodiments.
[0035] Figure 8 shows a lifting conversion process for displacement according to embodiments.
[0036] Figure 9 illustrates a process of packing transformation coefficients into a 2D image according to embodiments.
[0037] Figure 10 illustrates an attribute transfer process of a V-MESH compression method according to embodiments.
[0038] Figure 11 illustrates a decoding process of mesh data according to embodiments.
[0039] Fig. 12 is a drawing showing an example of a transmitting device according to embodiments.
[0040] Fig. 13 is a drawing showing an example of a receiving device according to embodiments.
[0041] FIG. 14 is a diagram showing an example of dynamic mesh data having multiple attribute information according to embodiments.
[0042] FIG. 15 is a drawing showing an example of a texture map video regenerated according to embodiments.
[0043] Fig. 16 is a drawing showing an example of a detailed block diagram of a parameterization unit according to embodiments.
[0044] FIG. 17 is a diagram showing an example of mesh patches constituting a simplified mesh according to embodiments being mapped onto a 2D image.
[0045] FIG. 18 is a diagram showing another example of a texture map video regenerated according to embodiments.
[0046] Fig. 19 is a flowchart showing an example of a mesh data encoding method according to embodiments.
[0047] Fig. 20 is a flowchart showing an example of a mesh data decoding method according to embodiments.
[0048] Preferred embodiments of the embodiments are described in detail, examples of which are illustrated in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to illustrate preferred embodiments of the embodiments, rather than merely show embodiments that can be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these details.
[0049] While most of the terms used in the examples are commonly used in the field, some terms were arbitrarily selected by the applicant, and their meanings are described in detail in the following descriptions as needed. Therefore, the examples should be understood based on the intended meaning of the terms, not simply their names or meanings.
[0050] With the recent development of 3D data modeling and rendering technology, research on creating and processing 3D data is being conducted in various fields such as Virtual Reality (VR), Augmented Reality (AR), autonomous driving, Computer-Aided Design (CAD) / Computer-Aided Manufacturing (CAM), and Geographic Information Systems (GIS). 3D data can be represented as point clouds, meshes, etc., depending on the representation format. Among these, a mesh is composed of geometric information expressing the coordinate values of each vertex (or point), connection information indicating the connection relationship between vertices, a texture map expressing the color information of the mesh surface as 2D image data, and texture coordinates indicating mapping information between the surface of the mesh and the texture map. In the present disclosure, a mesh is defined as a dynamic mesh if one or more of the elements that make up the mesh change over time, and a static mesh if they do not change. In other words, dynamic mesh data may refer to mesh data that has an object or movement.
[0051] Because dynamic mesh data has a large amount of data for elements that constitute the mesh compared to two-dimensional image data, technologies have been developed to efficiently compress this large amount of mesh data to store and transmit it.
[0052] Figures 1(a) and 1(b) illustrate a V-DMC-based encoder and decoder according to embodiments. In particular, Figure 1(a) illustrates an encoder, and Figure 1(b) illustrates a decoder.
[0053] The basic structure of the currently in-progress V-DMC (v-mesh) is as shown in Fig. 1(a) and Fig. 1(b). The encoder according to Fig. 1(a) and the decoder according to Fig. 1(b) perform the process of encoding and decoding media representing dynamic meshes using V3C (Visual Volumetric Video-based Coding) technology. The preprocessor converts the input dynamic mesh representation into several V3C components (base mesh, displacement set, 2D representation of attributes, and atlas). The original mesh is simplified into the base mesh. The base mesh can be encoded using any mesh codec. The displacement vector can be encoded into the V3C geometry video component using any video codec, either indicated by the profile or based on the SEI (supplemental enhancement information) message. For example, depending on the profile, the displacement vector (or displacement data) can be encoded via a video codec-based encoder, a zero-run length encoder, an arithmetic encoder, etc. The attribute data can include additional attributes. For example, texture or material information can be included as additional attributes, and can be encoded based on any video codec. The atlas data includes information on how to perform inverse reconstruction, and is provided to the V3C (or v-mesh) decoding and / or rendering system of the receiving device. For example, the atlas data can include a method for performing subdivision of the base mesh, a method for applying displacement vectors to subdivided mesh vertices, a method for applying attributes to the reconstructed mesh, etc.
[0054] An encoder according to embodiments may comprise a memory and at least one processor connected to the memory. The at least one processor may be configured to perform operations such as a pre-processor, an atlas encoding unit, a basemesh encoding unit, a displacement vector encoding unit, a video encoding unit, and a multiplexer.
[0055] The atlas encoding unit encodes an atlas of mesh data to generate an atlas bitstream. The basemesh encoding unit encodes a basemesh of mesh data to generate a basemesh bitstream. The displacement vector encoding unit encodes a displacement vector of mesh data to generate a displacement vector bitstream. The video encoding unit encodes an attribute of mesh data to generate an attribute bitstream. An encoder according to embodiments may generate parameter information (which may be referred to as signaling information, metadata, etc.) related to each encoding. An encoder according to embodiments may generate a compressed bitstream including parameter information, an atlas, a basemesh, displacement vectors, and / or attributes.
[0056] A decoder according to embodiments may comprise a memory and at least one processor connected to the memory. The at least one processor may be configured to perform operations such as a demultiplexer, an atlas decoding unit, a basemesh decoding unit, a displacement vector decoding unit, and a video decoding unit.
[0057] The atlas decoding unit decodes the atlas within the bitstream. The basemesh decoding unit decodes the basemesh within the bitstream. The displacement vector decoding unit decodes the displacement vector within the bitstream. The video decoding unit decodes the attribute within the bitstream. The decoder according to the embodiments may perform each decoding operation based on parameter information within the bitstream. In the decoder according to the embodiments, the basemesh processing unit restores the current basemesh from the decoded basemesh based on the atlas and / or parameter information. In the decoder according to the embodiments, the displacement processing unit restores the displacement vector by performing coordinate system transformation of the decoded displacement vector based on the atlas and / or parameter information. In the decoder according to the embodiments, the mesh restoration unit restores the final mesh by combining the restored basemesh and the restored displacement vector based on the atlas and / or parameter information. The restored mesh processing unit of the decoder according to the embodiments can generate and render a reconstructed dynamic mesh image by combining the decoded attribute (or texture map) with the restored final mesh. That is, the reconstructed dynamic mesh image can be displayed to the user.
[0058] Below, the operation of the V-DMC encoder and decoder of Fig. 1 is described in more detail.
[0059] FIG. 2 illustrates a system for providing dynamic mesh content according to embodiments.
[0060] The system of FIG. 2 includes a transmitting device (100) and a receiving device (110) according to embodiments. The transmitting device (100) may include a mesh video acquisition unit (101), a mesh video encoder (102), a file / segment encapsulator (103), and a transmitter (104). The receiving device (110) may include a receiving unit (111), a file / segment decapsulator (112), a mesh video decoder (113), and a renderer (114). Each component of FIG. 2 may correspond to hardware, software, a processor, and / or a combination thereof. Hereinafter, the mesh data transmitting device according to embodiments may be interpreted as a term referring to a 3D data transmitting device or transmitting device (100), or a mesh video encoder (hereinafter, referred to as an encoder) (102). The mesh data receiving device according to the embodiments may be interpreted as a term referring to a 3D data receiving device or receiving device (110), or a mesh video decoder (hereinafter, decoder) (113).
[0061] The system of FIG. 2 can perform video-based dynamic mesh compression and decompression.
[0062] Advances in 3D capture, modeling, and rendering have enabled users to consume diverse forms of 3D content, such as AR, XR, metaverse, and holograms, across multiple platforms and devices. 3D content increasingly represents objects with greater precision and realism, enabling users to enjoy immersive experiences. To achieve this, the creation and use of 3D models requires a significant amount of data. Among various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. Embodiments include a series of processing steps in a system that utilizes such mesh content.
[0063] First, the method of compressing dynamic mesh data starts from the V-PCC (Video-based point cloud compression) standard technology for point cloud data. Point cloud data is data that has color information at the coordinates (X, Y, Z) of a vertex (or point). In the present disclosure, the coordinates (i.e., position information) of a vertex are referred to as geometry information, the color information of a vertex is referred to as attribute information, and the geometry information and attribute information are referred to as vertex information or point cloud data. The vertex information to which connectivity information between vertices is added is referred to as mesh data. When creating content, it can be created in the form of mesh data from the beginning. Alternatively, it can be used by converting it into mesh data by adding connectivity information to point cloud data.
[0064] Currently, the MPEG standards body defines the data types of dynamic mesh data as the following two types.
[0065] Category 1: Mesh data with texture maps as color information.
[0066] Category 2: Mesh data with vertex colors as color information.
[0067] Mesh coding standards for Category 1 data are currently under development, and work on Category 2 data standards is also planned for the future. The overall process for providing mesh content services may include acquisition, encoding, transmission, decoding, rendering, and / or feedback, as shown in Figure 2.
[0068] To provide mesh content services, 3D data acquired through multiple cameras or specialized cameras can be processed into mesh data types through a series of processes and then converted into video. The generated mesh video is then transmitted through a series of processes, and the receiving end can then reprocess the received data into mesh video and render it. This allows mesh video to be presented to users, who can then interact with the mesh content according to their intended intent.
[0069] A mesh compression system may include a transmitting device (100) and a receiving device (110) as shown in FIG. 2. The transmitting device (100) may encode mesh video to output a bitstream, and transmit the bitstream to the receiving device (110) in the form of a file or streaming (streaming segment) via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0070] In the above transmitting device (100), the encoder may be called a mesh video / video / picture / frame encoding device, and in the receiving device (110), the decoder may be called a mesh video / video / picture / frame decoding device. The transmitter may be included in a mesh video encoder. The receiver may be included in a mesh video decoder. The renderer (114) may include a display unit, and the renderer and / or the display unit may be configured as separate devices or external components. The transmitting device (100) and the receiving device (110) may further include separate internal or external modules / units / components for a feedback process.
[0071] Mesh data represents the surface of an object as a number of polygons. Each polygon is defined by vertices in 3D space and connection information that describes how the vertices are connected. It can also contain vertex attributes such as vertex color and normal. Mapping information that allows the surface of the mesh to be mapped to a 2D planar area can also be included in the attributes of the mesh. The mapping can be described as a set of parameter coordinates, commonly called UV coordinates or texture coordinates, associated with the mesh vertices. Meshes contain 2D attribute maps, which can be used to store high-resolution attribute information such as texture, normal, and displacement. Here, displacement can be used interchangeably with displacement information or displacement vector.
[0072] The mesh video acquisition unit (101) may include processing 3D object data acquired through a camera, etc. into a mesh data type having the attributes described above through a series of processes and generating a video composed of such mesh data. The mesh video may have attributes of the mesh, such as vertices, polygons, connection information between vertices, colors, normals, etc., that may change over time. A mesh video having attributes and connection information that change over time in this way may be expressed as a dynamic mesh video.
[0073] A mesh video encoder (102) can encode an input mesh video into one or more video streams. One video can include multiple frames, and one frame can correspond to a still image / picture. In this document, a mesh video can include a mesh image / frame / picture, and a mesh video can be used interchangeably with a mesh image / frame / picture. The mesh video encoder (102) can perform a Video-based Dynamic Mesh (V-Mesh) Compression procedure. The mesh video encoder (102) can perform a series of procedures such as prediction, transformation, quantization, and entropy coding for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0074] The file / segment encapsulator (103) can encapsulate encoded mesh video data and / or mesh video-related metadata in the form of a file, etc. Here, the mesh video-related metadata may be received from a metadata processing unit, etc. The metadata processing unit may be included in the mesh video encoder (102) or may be configured as a separate component / module. The file / segment encapsulator (103) can encapsulate the corresponding data in a file format such as ISOBMFF, or process it in the form of other DASH segments, etc. The file / segment encapsulator (103) may include mesh video-related metadata in the file format according to an embodiment. The mesh video metadata may be included in boxes at various levels in the ISOBMFF file format, for example, or may be included as data in a separate track within the file. Depending on the embodiment, the file / segment encapsulator (103) may encapsulate the mesh video related metadata itself into a file.
[0075] The transmission processing unit can process encapsulated mesh video data for transmission according to the file format. The transmission processing unit can be included in the transmission unit (104) or can be configured as a separate component / module. The transmission processing unit can process mesh video data according to any transmission protocol. The processing for transmission can include processing for transmission through a broadcast network or processing for transmission through broadband. According to an embodiment, the transmission processing unit can receive not only mesh video data but also mesh video-related metadata from the metadata processing unit and process it for transmission.
[0076] The transmission unit (104) can transmit encoded video / image information or data output in the form of a bitstream to the reception unit (111) of the reception device (110) via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (104) can include an element for generating a media file through a predetermined file format and can include an element for transmission via a broadcasting / communication network. The reception unit (111) can extract the bitstream and transmit it to a decoding device.
[0077] The receiving unit (111) can receive mesh video data transmitted by a mesh data transmission device. Depending on the channel through which it is transmitted, the receiving unit (111) can receive mesh video data through a broadcast network, through a broadband, or through a digital storage medium.
[0078] The receiving processing unit can perform processing according to the transmission protocol on the received mesh video data. The receiving processing unit can be included in the receiving unit (111) or can be configured as a separate component / module. In order to correspond to the processing performed for transmission on the transmitting side, the receiving processing unit can perform the reverse process of the aforementioned transmission processing unit. The receiving processing unit can transfer the acquired mesh video data to the file / segment decapsulator (112) and transfer the acquired mesh video-related metadata to the metadata parser. The mesh video-related metadata acquired by the receiving processing unit can be in the form of a signaling table.
[0079] The file / segment decapsulator (112) can decapsulate mesh video data in the form of a file received from a receiving processing unit. The file / segment decapsulator (112) can decapsulate files according to ISOBMFF, etc., to obtain a mesh video bitstream or mesh video-related metadata (metadata bitstream). The obtained mesh video bitstream can be transmitted to the mesh video decoder (113), and the obtained mesh video-related metadata (metadata bitstream) can be transmitted to the metadata processing unit. The mesh video bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the mesh video decoder (113) or may be configured as a separate component / module. The mesh video-related metadata obtained by the file / segment decapsulator (112) may be in the form of a box or track within a file format. The file / segment decapsulator (112) may receive metadata required for decapsulation from the metadata processing unit, if necessary. The mesh video related metadata may be passed to the mesh video decoder (113) and used in the mesh video decoding procedure, or may be passed to the renderer (114) and used in the mesh video rendering procedure.
[0080] The mesh video decoder (113) can receive a bitstream and perform a reverse operation corresponding to the operation of the mesh video encoder (102) to decode the video / image. The decoded mesh video / image can be displayed through the display unit of the renderer (114). The user can view all or part of the rendered result through a VR / AR display or a general display.
[0081] The feedback process may include a process of transmitting various feedback information that may be acquired during the rendering / display process to the transmitter or to the decoder on the receiver. Interactivity may be provided in mesh video consumption through the feedback process. Depending on the embodiment, head orientation information, viewport information indicating the area that the user is currently viewing, etc. may be transmitted during the feedback process. Depending on the embodiment, the user may interact with things implemented in the VR / AR / MR / autonomous driving environment, in which case information related to the interaction may be transmitted to the transmitter or the service provider during the feedback process. Depending on the embodiment, the feedback process may not be performed.
[0082] Head orientation information can refer to information about the user's head position, angle, and movement. Based on this information, information about the area the user is currently viewing within the mesh video, i.e. viewport information, can be calculated.
[0083] Viewport information can be information about the area the user is currently viewing in the mesh video. This can be used to perform gaze analysis to determine how the user consumes the mesh video, which area of the mesh video they are gazing at, and for how long. Gaze analysis can be performed on the receiving side and transmitted to the transmitting side through a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.
[0084] Depending on the embodiment, the aforementioned feedback information may not only be transmitted to the transmitter but may also be consumed by the receiver. That is, the aforementioned feedback information may be utilized to perform decoding, rendering, and other processes on the receiver. For example, head orientation information and / or viewport information may be utilized to preferentially decode and render only the mesh video for the area currently being viewed by the user.
[0085] This document relates to embodiments of dynamic mesh video compression as described above. The method / embodiment disclosed in this document can be applied to the Video-based Dynamic Mesh Compression (V-Mesh) standard of the Moving Picture Experts Group (MPEG) or the next-generation video / image coding standard. Dynamic mesh video compression is a method for processing mesh connection information and attributes that change over time, and it can perform lossy and lossless compression for various applications such as real-time communication, storage, free-viewpoint video, and AR / VR.
[0086] The dynamic mesh video compression method described below is based on MPEG's V-Mesh method.
[0087] In this document, picture / frame can generally mean a unit representing one video of a specific time period.
[0088] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, only the pixel / pixel value of the chroma component, or only the pixel / pixel value of the depth component.
[0089] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0090] As described above, the encoding process of Fig. 2 is as follows.
[0091] That is, the video-based dynamic mesh compression (V-Mesh) compression method can provide a method of compressing dynamic mesh video data based on 2D video codecs such as HEVC (High Efficiency Video Coding) and VVC (Versatile Video Coding). The V-Mesh compression process receives the following data as input and performs compression.
[0092] Input mesh: Contains the 3D coordinates of the vertices that make up the mesh, normal information for each vertex, mapping information that maps the mesh surface to a 2D plane, and connection information between the vertices that make up the surface. The mesh surface can be expressed as triangles or more polygons, and connection information between the vertices that make up each surface is stored according to a set shape. The input mesh can be saved in the OBJ file format.
[0093] Attribute map: (Hereinafter, texture map is also used in the same meaning): Contains information about the attributes of the mesh (color, normal, displacement, etc.), and stores data in the form of mapping the surface of the mesh onto a 2D image. Mapping which part (surface or vertex) of the mesh each data of this attribute map corresponds to is based on the mapping information contained in the input mesh. Since the attribute map has data for each frame of the mesh video, it can also be expressed as an attribute map video. The attribute map in the V-Mesh compression method mainly contains the color information of the mesh, and is saved in an image file format (PNG, BMP, etc.).
[0094] Material Library File: Contains information about the material attributes used in a mesh, and in particular, information that links the input mesh to its corresponding attribute map. It is saved in the Wavefront Material Template Library (MTL) file format.
[0095] In the V-Mesh compression method, the following data and information can be generated through the compression process.
[0096] Base mesh: The input mesh is simplified (decimated) through a pre-processing process, thereby expressing the objects of the input mesh using the minimum number of vertices determined by the user's standards.
[0097] Displacement: This is displacement information used to express the input mesh as similarly as possible to the base mesh, and is expressed in the form of 3D coordinates.
[0098] Atlas information: This is the metadata required to reconstruct a mesh using base mesh, displacement, and attribute map information. Atlas information can be created and utilized as sub-mesh units (such as patches) that make up the mesh.
[0099] Referring to FIGS. 3 to 7, a method for encoding mesh position information (or vertex position information) is described, and referring to FIGS. 7 to 10, etc., a method for encoding attribute information (attribute map) by restoring mesh position information is described.
[0100] Figure 3 illustrates a V-MESH compression method according to embodiments.
[0101] Fig. 3 illustrates the encoding process of Fig. 2, and the encoding process may include a pre-processing process and an encoding process. The mesh video encoder (102) of Fig. 2 may include a pre-processor (200) and an encoder (201) as shown in Fig. 3. In addition, the transmitting device of Fig. 2 may be broadly referred to as an encoder, and the mesh video encoder (102) of Fig. 2 may be referred to as an encoder. The V-Mesh compression method may include a pre-processing process (Pre-processing, 200) and an encoding process (Encoding, 201) as shown in Fig. 3. The pre-processor (200) of Fig. 3 may be located in front of the encoder (201) of Fig. 3. The pre-processor (200) and the encoder (201) of Fig. 3 may be referred to as a single encoder.
[0102] The pre-processor (200) can receive a static of a dynamic mesh (M(i)) and / or an attribute map (A(i)). The pre-processor (200) can generate a base mesh (m(i)) and / or a displacement (d(i)) through pre-processing. The pre-processor (200) can receive feedback information from the encoder (201) and generate the base mesh and / or the displacement based on the feedback information.
[0103] The encoder (201) can receive a base mesh (m(i)), a displacement (d(i)), a static of a dynamic mesh (M(i)), and / or an attribute map (A(i)). In the present disclosure, at least one of the base mesh (m(i)), the displacement (d(i)), the static of a dynamic mesh (M(i)), and / or the attribute map (A(i)) can be referred to as mesh-related data. The encoder (201) can encode the mesh-related data to generate a compressed bitstream.
[0104] Figure 4 illustrates a pre-processing process of V-MESH compression according to embodiments.
[0105] Fig. 4 illustrates the configuration and operation of the preprocessor of Fig. 3. In Fig. 4, the input mesh may include a static of a dynamic mesh (M(i)) and / or an attribute map (A(i)). In addition, the input mesh may include three-dimensional coordinates of vertices constituting the mesh, normal information of each vertex, mapping information for mapping the mesh surface to a 2D plane, connection information between vertices constituting the surface, etc.
[0106] Fig. 4 shows a process of performing pre-processing on an input mesh. The pre-processing process (200) may largely include four steps: 1) GoF (Group of Frame) generation, 2) Mesh Decimation, 3) UV parameterization, and 4) Fitting subdivision surface (300). According to embodiments, GoF generation may be referred to as a GoF generation process or a GoF generation unit, mesh simplification may be referred to as a mesh simplification process or a mesh simplification unit, UV parameterization may be referred to as a UV parameterization process or a UV parameterization unit, and the fitting subdivision surface may be referred to as a fitting subdivision surface process or a fitting subdivision surface unit. The pre-processor (200) can generate displacement and / or base meshes from the received input mesh and transmit them to the encoder (201). The pre-processor (200) can transmit GoF information associated with GoF generation to the encoder (201).
[0107] Below, each step of Fig. 4 is described.
[0108] GoF Generation: This is the process of generating a reference structure for mesh data. If the number of vertices, the number of texture coordinates, the vertex connection information, and the texture coordinate connection information of the mesh of the previous frame and the current mesh are all the same, the previous frame can be set as the reference frame. That is, if only the vertex coordinate values are different between the current input mesh and the reference input mesh, the encoder (201) can perform inter frame encoding. Otherwise, intra frame encoding is performed for the corresponding frame.
[0109] Mesh Decimation: This process simplifies the input mesh to create a simplified mesh, or base mesh. Vertices to be removed from the original mesh are selected based on user-defined criteria, and the selected vertices and the triangles connected to them can be removed.
[0110] In the process of performing mesh simplification (Mesh decimation), the input mesh (voxelized), target triangle ratio (TTR), and minimum triangle component (CCCount) information are passed as input, and the simplified mesh (decimated mesh) can be obtained as output. In this process, connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.
[0111] UV parameterization: This is the process of mapping a 3D surface of a decimated mesh into a texture domain. Parameterization can be performed using the UVAtlas tool. This process generates mapping information, which indicates where each vertex of the decimated mesh can be mapped to on a 2D image. This mapping information is expressed and stored as texture coordinates, and through this process, the final base mesh is created.
[0112] OrthoAtlas technology generates texture coordinates using orthographic projection. OrthoAtlas technology sequentially generates patches and packs them. First, adjacent triangles are divided to generate Connected Components (CCs), and then the optimal CCs are merged using a cost function to generate patches. The cost function can measure the cost based on the degree of distortion that occurs when orthogonally projecting patches in each direction. By packing the patch that minimizes the cost function into the texture domain, the texture coordinates can be calculated. In the case of orthoAtlas technology, texture coordinates and texture connection information can be derived from the base mesh decoder without compressing them during the base mesh encoding process.
[0113] Fitting subdivision surface (300): This is a process of performing subdivision on a decimated mesh (i.e., a simplified mesh having texture coordinates). The displacement and base mesh generated through this process are output to the encoder (201). A user-defined method, such as a mid-edge method, may be applied as the subdivision method. A fitting process is performed so that the input mesh and the mesh on which the subdivision is performed are similar to each other. In the present disclosure, the mesh on which the fitting process is performed is referred to as a fitted subdivision mesh (or fitted subdivision mesh). This process is a process of performing fitting so that the mesh on which the subdivision is performed on the base mesh is similar to the surface of the input mesh. As a subdivision method, a user-defined method such as the mid-edge method (see Fig. 5), the loop method, and the LS3 method can be applied.
[0114] Figure 5 illustrates a mid-edge subdivision method according to embodiments.
[0115] Figure 5 illustrates the mid-edge method of the fitting subdivision surface described in Figure 4. Referring to Figure 5, an original mesh containing four vertices is subdivided to generate a sub-mesh. A sub-mesh can be generated by creating a new vertex in the middle of the edge between the vertices. Then, a fitting process is performed so that the input mesh and the sub-mesh become similar to each other, thereby generating a fitted sub-division mesh.
[0116] When a fitted subdivided mesh (hereinafter referred to as a fitted subdivided mesh) is generated, displacement is calculated using this result and a pre-compressed and decoded base mesh (hereinafter referred to as a reconstructed base mesh). That is, the reconstructed base mesh is subdivided in the same way as the fitting subdivision surface. The difference in position of each vertex between this result and the fitted subdivided mesh is the displacement for each vertex. Since the displacement represents the position difference in three-dimensional space, it is also expressed as a value in the (x, y, z) space of the Cartesian coordinate system. Depending on the user input parameters, the (x, y, z) coordinate values can be converted to (normal, tangential, bi-tangential) coordinate values of the local coordinate system.
[0117] Fig. 6 illustrates a displacement generation process according to embodiments. The displacement generation process of Fig. 6 may be performed in a pre-processor (200) or in an encoder (201).
[0118] Fig. 6 illustrates in detail the displacement calculation method of the fitting subdivision surface (300) as described in Fig. 5.
[0119] An encoder and / or pre-processor according to embodiments may include 1) a subdivision unit, 2) a local coordinate system calculation unit, and 3) a displacement calculation unit. The subdivision unit may perform subdivision on a restored base mesh to generate a subdivided restored base mesh. Here, the restoration of the base mesh may be performed in the pre-processor (200) or in the encoder (201). The local coordinate system calculation unit may receive a fitted subdivision mesh and a subdivided restored base mesh, and may convert a coordinate system of the mesh into a local coordinate system based on the fitted subdivision mesh and the subdivided restored base mesh. The local coordinate system calculation operation may be optional. The displacement calculation unit may calculate a positional difference between the fitted subdivision mesh and the subdivided restored base mesh. For example, a positional difference value between vertices of two input meshes may be generated. The vertex positional difference value becomes a displacement.
[0120] The mesh data transmission method and device according to the embodiments can encode mesh data as follows. Mesh data is a term including point cloud data. Point cloud data (which may be referred to as point cloud for short) according to the embodiments can refer to data including vertex coordinates (or geometry information) and color information (or attribute information). In addition, geometry images, attribute images, occupancy maps, and additional information (or patch information) generated through patch generation and packing based on vertex coordinates and color information are also referred to as point cloud data. Therefore, point cloud data including connection information can be referred to as mesh data. In this document, point cloud and mesh data can be used interchangeably.
[0121] The V-Mesh compression (reconstruction) method according to the embodiments may include intra frame encoding and inter frame encoding.
[0122] Based on the results of the GoF generation described above, intra-frame encoding or inter-frame encoding is performed. In the case of intra-encoding, the data to be compressed may be a base mesh, displacement, attribute map, etc. In the case of inter-encoding, the data to be compressed may be a displacement, attribute map, and a motion field between a reference base mesh and the current base mesh.
[0123] Figure 7 illustrates a V-DMC encoding process according to embodiments.
[0124] The elements of the transmitting device illustrated in FIG. 7 may be implemented by hardware, software, a processor connected to a memory, and / or a combination thereof. That is, the elements of the transmitting device illustrated in FIG. 7 may be implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not illustrated in the drawing. One or more processors may perform at least one or more of the operations and / or functions of the elements of the transmitting device illustrated in FIG. 7. In addition, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the transmitting device illustrated in FIG. 7. The execution order of each block in FIG. 7 may be changed, some blocks may be omitted, and some blocks may be newly added.
[0125] In the present disclosure, the operation process of a transmitter for compressing and transmitting dynamic mesh data using V-Mesh compression technology may be as shown in Fig. 7. The transmitter of Fig. 7 may support both an intra-frame encoding (or intra-encoding or intra-screen encoding) process and / or an inter-frame encoding (or inter-encoding or inter-screen encoding) process.
[0126] The encoding process of FIG. 7 details the encoding of the mesh video encoder (102) of FIG. 2. The encoder of FIG. 7 may include a pre-processor (200) and / or an encoder (201). The pre-processor (200) and encoder (201) of FIG. 7 may correspond to the pre-processor (200) and encoder (201) of FIG. 4.
[0127] The preprocessor (200) can receive an input mesh and perform the preprocessing described above. The preprocessing can generate a base mesh and / or a fitted subdivision mesh.
[0128] The quantizer (411) of the encoder (201) can quantize the base mesh and / or the fitted subdivided mesh.
[0129] According to embodiments, the base mesh quantized in the mesh quantization unit (411) may be output to a static mesh encoder (413) or a motion vector encoder (414) through a switching unit (412). According to embodiments, the base mesh is output to a motion vector encoder (414) through the switching unit (412) when inter-encoding is performed on the corresponding mesh frame, and is output to a static mesh encoder (413) through the switching unit (412) when intra-encoding is performed on the corresponding mesh frame. The motion vector encoder (414) may be referred to as a motion encoder.
[0130] For example, when performing intra encoding or intra frame encoding for the corresponding mesh frame, the base mesh can be compressed through a static mesh encoder (413). In this case, encoding can be performed on connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. That is, vertex coordinates, vertex connection information, texture coordinates, texture connection information, etc. of the mesh can be encoded in the static mesh encoder (413). The base mesh bitstream generated through encoding is transmitted to a multiplexer (not shown).
[0131] As another example, when performing inter-encoding (or inter-frame encoding) on the corresponding mesh frame, the motion vector encoder (414) may receive the current base mesh and the reference reconstructed base mesh (or the reconstructed quantized reference base mesh) as input, calculate a motion vector between the two meshes, and encode the value. In addition, the motion vector encoder (414) may perform connection information-based prediction using the previously encoded / decoded motion vector as a predictor, and entropy-encode a differential motion vector (or residual motion vector) obtained by subtracting the predicted motion vector from the current motion vector. According to embodiments, the motion vector encoding may be performed on a vertex basis or a subgroup basis. The motion vector bitstream generated through the motion vector encoding is transmitted to a multiplexer (not shown) as a base mesh bitstream. That is, in the case of intra-frame encoding, the static mesh bitstream is input to the multiplexer as the base mesh bitstream, and in the case of inter-frame encoding, the motion vector bitstream is input to the multiplexer as the base mesh bitstream.
[0132] In FIG. 7, the base mesh restoration unit (415) can receive a base mesh encoded by the static mesh encoder (413) or a motion vector encoded by the motion vector encoder (414) to generate a reconstructed base mesh. The base mesh restoration unit (415) performs restoration of the base mesh according to the encoding type (inter-screen encoding or intra-screen encoding) of the current mesh. For example, the base mesh restoration unit (415) can perform static mesh decoding on the base mesh encoded by the static mesh encoder (413) to restore the base mesh. At this time, quantization can be applied before static mesh decoding, and inverse quantization can be applied by the inverse quantization unit (416) after static mesh decoding. That is, when intra-screen encoding is performed, the inverse quantization unit (416) can perform inverse quantization on the quantized base mesh through the mesh quantization unit (411) to restore the current base mesh. As another example, the base mesh restoration unit (415) can restore the base mesh based on the restored quantized reference base mesh and the motion vector encoded by the motion vector encoder (414). That is, when inter-screen encoding is performed, the current base mesh can be generated by decoding the motion vector using the motion vector decoding method and then applying (i.e., adding) the decoded motion vector to the reference restoration base mesh. At this time, when the motion vector is not quantized, the motion vector restoration process is omitted and the current base mesh can be restored using the motion vector calculated by the motion vector encoder (414). The restored base mesh is output to the displacement vector calculation unit (417) and the mesh restoration unit (425).
[0133] According to embodiments, the displacement vector calculation unit (417) can perform mesh refinement on the restored base mesh. In addition, the displacement vector calculation unit (417) can calculate a displacement vector, which is a difference value in vertex positions between the restored base mesh that has been refined and the fitted subdivision (or subdivided) mesh generated by the pre-processor (200). That is, the displacement vector is a difference in positions between the vertices of the two meshes so that the fitted subdivision (or subdivided) mesh becomes similar to the original mesh. At this time, the displacement vector can be calculated as many times as the number of vertices of the subdivision mesh. That is, the displacement vector of the number of vertices of the subdivision (subdivided) mesh can be calculated through the displacement vector calculation unit (417).
[0134] The lifting transformation unit (418) can perform a lifting transformation on the input displacement vector to generate a lifting coefficient (or displacement vector transformation coefficient). The quantizer (419) can quantize the lifting coefficient, i.e., the displacement vector transformation coefficient.
[0135] In the present disclosure, the displacement vector or the quantized displacement vector transform coefficient can be encoded through a 2D video codec-based encoding method, and / or a zero run length encoding method, and / or an arithmetic encoding method, etc.
[0136] If an arithmetic encoding method is used, the displacement vector or the quantized displacement vector transform coefficient is encoded based on an arithmetic codec in an arithmetic encoding unit (421) after inter prediction in an inter prediction unit (420), and if a 2D video codec-based encoding method is used, the displacement vector or the quantized displacement vector transform coefficient is encoded based on a 2D video codec in a video encoding unit (423) after image packing in an image packing unit (422) and can be output as a displacement bitstream (i.e., a compressed displacement bitstream). For example, the image packing unit (422) can pack an image based on quantized lifting coefficients (i.e., displacement vector transform coefficients). The video encoding unit (423) can encode the packed image. That is, the quantized lifting coefficients are packed into one frame as a 2D image by the image packing unit (422), compressed by the video encoding unit (423), and output as a displacement bitstream (i.e., a compressed displacement bitstream).
[0137] The displacement vector restoration unit (424) may include a video decoder, an image unpacking unit, an inverse quantizer, and an inverse linear lifting unit. That is, the displacement vector restoration unit (424) performs decoding on an encoded displacement vector in the video decoder, performs image unpacking in the image unpacking unit, performs inverse quantization in the inverse quantizer, and then performs inverse transformation in the inverse linear lifting unit to restore the displacement vector. The restored displacement vector is output to the mesh restoration unit (425). The mesh restoration unit (425) restores the deformed mesh based on the base mesh restored by the base mesh restoration unit (415) and the displacement vector restored by the displacement vector restoration unit (424). That is, the mesh restoration unit (425) restores the reconstructed and deformed mesh through the restored displacement output from the displacement vector restoration unit (424) and the restored base mesh (or subdivided restored base mesh) output from the inverse quantization unit (416). The present disclosure refers to the reconstructed and deformed mesh as a restored deformed mesh. The restored mesh (or restored deformed mesh) has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.
[0138] The attribute transfer (426) receives an input mesh and / or an input attribute map, and regenerates an attribute map based on the restored deformed mesh. The attribute map refers to a texture map corresponding to attribute information among mesh data components, and in the present disclosure, the attribute map and the texture map may be used interchangeably. The push-pull padding (427) may pad data in the attribute map based on the push-pull method. The color space conversion unit (428) may convert the space of the color component of the attribute map. For example, the attribute map may be converted from an RGB color space to a YUV color space. The video encoding unit (429, also referred to as a video encoder) may encode the attribute map and output it as a compressed attribute bitstream.
[0139] According to embodiments, the atlas encoder (430) may encode atlas information (or atlas data) to generate a compressed atlas bitstream. Then, the atlas bitstream generated through atlas information encoding is transmitted to the multiplexer (431). In the present disclosure, the atlas may be information required in a mesh reconstruction process and may mean information such as tiles and patches. In addition, the atlas information may mean data required in processes such as 2D mapping for a 3D object, texture mapping-related information, mesh decoding, and mesh restoration, and may include additional information such as a segmentation method, a transformation method, a quantization method, and the position and size of a patch within an atlas frame. In the present disclosure, the atlas information may be encoded through Exp-Golomb coding of the atlas encoder (430), etc.
[0140] According to embodiments, the multiplexer (430) can multiplex an input compressed base mesh bitstream, a compressed displacement (or displacement vector) bitstream, a compressed attribute (or texture map) bitstream, and a compressed atlas bitstream to generate a single compressed bitstream. The multiplexed bitstream can be encapsulated into one or more tracks of a file.
[0141] According to embodiments, the bitstream or file multiplexed in the multiplexer (431) may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0142] In summary, for the base mesh, encoding can be performed in different ways depending on the base mesh type (INTRA type, INTER type, SKIP type). If the base mesh is INTRA type, it can be encoded using the static mesh encoding method. If the base mesh is INTER type, the motion field between the reference base mesh and the current base mesh can be encoded. If the current base mesh is SKIP type, the reference base mesh can be derived as the current base mesh.
[0143] After being encoded by the encoder, the decoded base mesh can be subdivided into a subdivided mesh through a subdivision process. Subdivision algorithms such as mid-point subdivision and loop subdivision can be used.
[0144] Static Base Mesh Encoding (Intra Base Mesh Encoding): When performing Intra encoding on the current base mesh, the base mesh generated during the preprocessing process can be encoded using static mesh compression technology after going through the quantization process. Static mesh compression applies MPEG EdgeBreaker (MEB) technology, and the vertex position information, mapping information (texture coordinates), vertex connection information, and normals of the base mesh are compressed.
[0145] Connection information can be encoded and compressed based on the edgebreaker algorithm. The edgebreaker algorithm sequentially traverses triangles according to rules, mapping symbols based on the characteristics of each triangle, and then encoding the corresponding symbols.
[0146] A technique for compressing vertex position information can encode the residual value, which is the difference between the current vertex and the predicted value, after obtaining the predicted value based on a prediction technique such as multiple parallelogram prediction.
[0147] A technique for compressing mapping information (texture coordinates) can encode the residual value, which is the difference between the current mapping information (texture coordinates) and the predicted value, after obtaining the predicted value based on a prediction technique such as stretch.
[0148] Techniques for compressing normals can encode residual values, which are the differences between the current normal and the predicted values, after obtaining predicted values based on prediction techniques such as delta coding, multiple parallelogram prediction, and cross product-based prediction.
[0149] Motion Field Encoding (Inter Basemesh Encoding): Inter basemesh encoding can be performed when a one-to-one correspondence exists between the reference mesh and the current input mesh, and only the vertex position information differs. When performing inter encoding, instead of compressing the base mesh, the difference between the vertices of the reference base mesh and the current base mesh, i.e. the motion field, can be calculated and this information can be encoded. The reference base mesh is the result of quantizing the already decoded base mesh data and is determined by the reference frame index.
[0150] The motion field can be encoded as a value, or the predicted motion field can be calculated by averaging the motion fields of the reconstructed vertices among the vertices connected to the current vertex, and the residual motion field, which is the difference between the predicted motion field value and the motion field value of the current vertex, can be encoded. This value can be encoded using entropy coding.
[0151] Displacement Encoding: After base mesh encoding, restoration and dequantization are performed to generate a Recon. base mesh, and the displacement between the result of performing subdivision on this and the fitted subdivided mesh can be calculated. For effective encoding, a data transform process such as wavelet transform can be applied to the displacement information, and Figure 8 shows the process of transforming displacement information using lifting transform in V-Mesh. The displacement vector transform coefficients generated through the transform process are quantized, and the quantized transform coefficients can be compressed through a video codec or through arithmetic encoding depending on the compression method.
[0152] When compressed through a video codec, it is packed into a 2D image as shown in Fig. 9. The transformation coefficients are organized into one block for each N^2(N*N) unit, and each block can be packed in z-scan order. The horizontal number of blocks is fixed to N, but the vertical number of blocks can be determined according to the number of vertices of the subdivided base mesh. Within one block, the transformation coefficients can be packed by sorting them with Morton code. The packed images generate a displacement video for each GoF unit, and this displacement video can be encoded using an existing video compression codec.
[0153] When compressed through arithmetic encoding, inter-frame prediction can be performed on the quantized displacement vector transform coefficients. When inter-frame prediction is performed on the current quantized displacement vector transform coefficients, the residual value, which is the difference between the current displacement vector transform coefficients and the reference displacement vector transform coefficients, can be encoded, and information about the reference target can be encoded. Depending on the displacement vector type, if it is an INTRA type, the quantized displacement vector transform coefficients can be arithmetic encoded, and if it is an INTER type, the residual value can be arithmetic encoded. When encoding, it can be performed based on Context Adaptive Binary Arithmetic Coding (CABAC). The CABAC process can first binarize the displacement vector data and map it to a bin string. The bin string can be a binarized output of 0 and 1, and each 0 or 1 can be a bin. Each bin can be arithmetic encoded using context information selected from a context model, and a process of updating the probability can be performed.
[0154] Figure 8 shows a lifting conversion process for displacement according to embodiments.
[0155] Figure 9 illustrates a process of packing transformation coefficients (or lifting coefficients) according to embodiments into a 2D image.
[0156] Figures 8 and 9 illustrate the process of converting the displacement of the encoding process of Figure 7 and the process of packing the conversion coefficients, respectively.
[0157] The encoding method according to the embodiments includes displacement encoding.
[0158] After base mesh encoding and / or motion field encoding, a reconstructed base mesh is generated through restoration and dequantization, and the displacement between the result of performing subdivision on the reconstructed base mesh and the fitted subdivided mesh generated through the fitting subdivision surface can be calculated (417 in Fig. 7). For effective encoding, a data transform process such as wavelet transform can be applied to the displacement information (418 in Fig. 7).
[0159] FIG. 8 shows a process of transforming displacement information using a lifting transform in the lifting transform unit (418) of FIG. 7. For example, a linear wavelet-based lifting transform may be performed. The transform coefficients generated through the transform process are quantized in a quantizer (419) and then packed into a 2D image as in FIG. 9 through an image packing unit (422). The transform coefficients are organized into one block for each 256 (=16×16) units, and each block can be packed in a z-scan order. The horizontal number of blocks is fixed to 16, but the vertical number of blocks can be determined according to the number of vertices of the subdivided base mesh. The transform coefficients can be packed by sorting them with a Morton code within one block. The packed images generate a displacement video for each GoF unit, and this displacement video can be encoded using an existing video compression codec in a video encoding unit (423, or video encoder).
[0160] Referring to FIG. 8, a base mesh (original) may include vertices and edges for LoD (Level Of Detail) 0. A first subdivision mesh generated by dividing (or subdividing) the base mesh includes vertices generated by further dividing (or subdividing) edges of the base mesh. The first subdivision mesh includes vertices for LoD0 and vertices for LoD1. LoD1 includes the subdivided vertices and the vertices of the base mesh (LoD0). The first subdivision mesh may be further divided (or subdivided) to generate a second subdivision mesh. The second subdivision mesh includes LoD2. LoD2 includes base mesh vertices (LoD0), LoD1 including vertices further divided (or subdivided) from LoD0, and vertices further divided (or subdivided) from LoD1. LoD is a Level of Detail (LoD) that indicates the degree of detail of mesh data content. As the level index increases, the distance between vertices becomes closer and the level of detail increases. In other words, the smaller the LoD value, the lower the detail of the mesh data content, and the larger the LoD value, the higher the detail of the mesh data content. LoD N contains the vertices included in the previous LoDN-1 as is. When a mesh (or vertex) is further divided through subdivision, the mesh can be encoded based on a prediction and / or update method by considering the previous vertices v1, v2, and the subdivided vertex v. Instead of directly encoding information about the current LoD N, a residual value between the previous LoD N-1 can be generated and the mesh can be encoded using the residual value to reduce the bitstream size. The prediction process refers to the operation of predicting the current vertex v using the previous vertices v1 and v2. Since adjacent subdivision meshes have similar data, this property can be utilized for efficient encoding.Current vertex position information is predicted as a residual of previous vertex position information, and the previous vertex position information is updated through the residual. In the present disclosure, vertex, apex, and point may be used with the same meaning. In addition, LoDs may be defined during the subdivision process of the base mesh. According to embodiments, the subdivision process of the base mesh may be performed in the pre-processor (200) or in a separate component / module.
[0161] Referring to FIG. 9, a vertex has a transformation coefficient (also called a lifting coefficient) generated through a lifting transformation. The transformation coefficient of a vertex related to a lifting transformation can be packed into an image by an image packing unit (422) and then encoded by a video encoding unit (423).
[0162] Figure 10 illustrates an attribute transfer process of a V-MESH compression method according to embodiments.
[0163] According to the embodiments, FIG. 10 shows the detailed operation of the attribute transfer (426) of FIG. 7.
[0164] Encoding according to embodiments includes attribute map encoding. According to embodiments, attribute map encoding may be performed in the video encoding unit (429) of FIG. 7.
[0165] According to embodiments, in the present disclosure, the encoder compresses information about the input mesh through base mesh encoding (i.e., intra encoding), motion field encoding (i.e., inter encoding), and displacement encoding. In the encoding process, the compressed input mesh is restored through base mesh decoding (intra frame), motion field decoding (inter frame), and displacement video decoding processes, and the restored result, the reconstructed deformed mesh (hereinafter referred to as Recon. deformed mesh), is used to compress the input attribute map as shown in FIG. 7. The reconstructed deformed mesh (Recon. deformed mesh) has position information of vertices, texture coordinates, and corresponding connection information, but does not have color information corresponding to the texture coordinates. Therefore, as shown in Fig. 10, in the V-Mesh compression method, a new attribute map having color information corresponding to the texture coordinates of the restored deformed mesh is regenerated through the attribute transfer process of the attribute transfer (426).
[0166] According to embodiments, attribute transfer (426) first checks whether each point P(u, v) of a 2D texture domain belongs to a texture triangle of a reconstructed deformed mesh, and if it exists in a texture triangle T, the barycentric coordinate of P(u, v) according to the triangle T ( , , ) is calculated. And the 3D vertex positions of triangle T and ( , , ) is used to compute the 3D coordinates M(x, y, z) of P(u, v). Find the vertex coordinates M'(x', y', z') and the triangle T' containing this vertex that corresponds to the most similar position to the calculated M(x, y, z) in the input mesh domain. Then, the center of mass coordinates of M'(x', y', z') in this triangle T' ( ', ', ') is calculated. The texture coordinates corresponding to the three vertices of Triangle T' and ( ', ', ') is used to calculate the texture coordinates (u', v'), and the color information corresponding to these coordinates is found in the input attribute map. The color information found in this way is immediately assigned to the (u, v) pixel location in the new attribute map. If P(u, v) does not belong to any triangle, the pixel at that location in the new attribute map can be filled with a color value using a padding algorithm, such as the push-pull algorithm of the push-pull padding (427).
[0167] A new attribute map generated through attribute transfer (426) is grouped into GoF units to form an attribute map video, which is compressed using a video codec of the video encoding unit (429).
[0168] Referring to Figure 10, the reference relationship between the input mesh, the input attribute map, the reconstructed deformed mesh, and the regenerated attribute map can be seen.
[0169] The decoding process of Fig. 2 can perform the reverse process of the corresponding process of the encoding process of Fig. 2. The specific decoding process is as follows.
[0170] Figure 11 illustrates a decoding process of V-Mesh technology according to embodiments.
[0171] Fig. 11 illustrates the configuration and operation of the mesh video decoder (113) of the receiving device of Fig. 2. In addition, Fig. 11 can restore mesh data by performing the reverse process of the encoding process of Fig. 7. In the present disclosure, the receiving device of Fig. 11 may be referred to as a mesh data receiving device or decoder or a decoder of the receiving device or a V-Mesh decoder or a dynamic mesh decoder.
[0172] The elements of the receiving device illustrated in FIG. 11 may be implemented by hardware, software, a processor connected to a memory, and / or a combination thereof. That is, the elements of the receiving device illustrated in FIG. 11 may be implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not illustrated in the drawing. One or more processors may perform at least one or more of the operations and / or functions of the elements of the receiving device illustrated in FIG. 11. In addition, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the receiving device illustrated in FIG. 11. The execution order of each block in FIG. 11 may be changed, some blocks may be omitted, and some blocks may be newly added.
[0173] Fig. 11 may largely include a demultiplexer (611), an atlas decoder (612), and a decoding unit (620).
[0174] According to embodiments, a bitstream (i.e., a compressed bitstream) of mesh data received by a receiver (not shown) may be demultiplexed into a base mesh bitstream (or a base mesh substream) after file / segment decapsulation, a displacement vector bitstream (or a displacement substream), a texture map bitstream (or an attribute map substream or a texture map substream), and / or an atlas bitstream in a demultiplexer (611). If the bitstream of mesh data is not encapsulated in a file form in the transmitter, the decapsulation process is omitted in the receiver. If the current mesh has inter-encoding applied, the base mesh bitstream may be a motion vector bitstream.
[0175] According to embodiments, an atlas bitstream is provided to an atlas decoder (612). The atlas decoder (612) can decode the atlas bitstream to restore atlas information. The restored atlas information can be used in a mesh decoding process. According to embodiments, the process in which the atlas information is used may be a segmentation process, a displacement vector restoration process, etc., and may include information such as tiles and patches.
[0176] According to embodiments, the atlas bitstream may be decoded through Exp-Golomb coding of the atlas decoder (612), etc. At this time, the atlas may be information required for the mesh reconstruction process, and may mean information such as tiles and patches. In addition, the atlas data may mean data required for the mesh decoding, mesh reconstruction, etc. processes, and may include a segmentation method, a transformation method, a quantization method, the location and size of a patch within the atlas frame, etc.
[0177] According to embodiments, the base mesh bitstream is provided to a motion vector decoder (623) via a switching unit (621) or to a static mesh decoder (622).
[0178] For example, if the current mesh has inter encoding applied, the base mesh bitstream, i.e., the motion vector bitstream, is received, demultiplexed, and then output to the motion vector decoder (623) through the switching unit (621). As another example, if the current mesh has intra encoding applied, the base mesh bitstream is received, demultiplexed, and then output to the static mesh decoder (622) through the switching unit (621). Here, the motion vector decoder (623) may be referred to as a motion decoder.
[0179] According to embodiments, the motion vector decoder (623) can perform decoding on a motion vector bitstream in units of vertices or subgroups.
[0180] According to embodiments, the motion vector decoder (623) can reconstruct the final motion vector by adding the differential motion vector (i.e., residual motion vector) decoded from the bitstream using the previously decoded motion vector as a predictor. That is, the motion vector decoder (623) decodes the differential motion vector (or residual motion vector) in units of vertices or subgroups (or subblocks) through the motion vector bitstream, and performs prediction based on connection information using the previously decoded motion vector as a predictor to decode the motion vector by adding it to the residual motion vector.
[0181] According to embodiments, the static mesh decoder (622) can decode the base mesh bitstream to restore connection information, vertex geometry information, texture coordinates (i.e., attribute geometry information), normal information, etc. of the base mesh. That is, the static mesh decoder (622) can restore a restored quantized base mesh, for example, connection information, vertex geometry information, vertex texture coordinates, etc. of the base mesh.
[0182] According to embodiments, the base mesh restoration unit (631) may restore the current base mesh based on the decoded motion vector or the decoded base mesh. For example, if the current mesh has inter-screen encoding applied, the base mesh restoration unit (631) may add the decoded (or restored) motion vector to the reference base mesh and then perform inverse quantization to generate a restored base mesh (i.e., the current base mesh). As another example, if the current mesh has intra-screen encoding applied, the base mesh restoration unit (631) may perform inverse quantization on the decoded (or restored) base mesh through the static mesh decoder (622) to generate a restored base mesh (i.e., the current base mesh). According to embodiments, inverse quantization may be omitted.
[0183] According to embodiments, the displacement sub-bitstream is provided to the arithmetic decoding unit (625) or the video decoding unit (627) through the switching unit (624) depending on the decoding method.
[0184] For example, if the decoding method is an arithmetic codec method, the arithmetic decoding unit (625) decodes the displacement sub-stream based on the arithmetic codec, and the inverse prediction unit (626) performs the inverse process of prediction on the decoded displacement information and then outputs it to the inverse quantization unit (629). As another example, if the decoding method is a 2D video codec method, the video decoding unit (627) decodes the displacement sub-stream based on the 2D video codec, and the image unpacking unit (628) unpacks the image of the decoded displacement video and then outputs it to the inverse quantization unit (629).
[0185] The displacement information provided from the above-mentioned inverse prediction unit (626) or image unpacking unit (628) is inversely quantized in the inverse quantization unit (629) and inversely transformed in the inverse linear lifting unit (630), and then restored as displacement information for each vertex (i.e., Recon. displacements).
[0186] According to embodiments, the mesh restoration unit (632) reconstructs and restores the deformed mesh (i.e., decoded mesh) using the restored displacement output from the inverse linear lifting unit (630) and the restored base mesh output from the base mesh restoration unit (631). That is, the dequantized restored base mesh is combined with the restored displacement information to generate the final decoded mesh. In the present disclosure, the final decoded mesh is referred to as a reconstructed deformed mesh.
[0187] According to embodiments, an attribute map sub-stream is decoded through a video decoding unit (633) corresponding to a video compression codec used in encoding, and then restored to a final attribute map (i.e., decoded attribute map) through a color conversion unit (634) through processes such as color format conversion and color space conversion.
[0188] According to embodiments, the restored decoded mesh and decoded attribute map can be utilized by the receiver as final mesh data that can be utilized by the user.
[0189] To summarize Fig. 11, the base mesh sub-stream can be decoded through a static mesh decoder (622) based on MEB (MPEG EdgeBreaker) technology depending on the basemesh type, for example, if it is an INTRA type, and as a result, connection information, vertex geometry information, vertex mapping information (texture coordinates), etc. of the base mesh can be restored.
[0190] In an encoder according to embodiments, when the texture parameterization method is orthoAtlas, the decoder can derive mapping information (texture coordinates) and attribute information (texture) connection information using vertex coordinates. The process of deriving mapping information (texture coordinates) and connection information can generate mapping information (texture coordinates) and attribute information (texture) connection information by calculating a homography transform of each face and then projecting the vertex based on the homography transform.
[0191] According to embodiments, when the base mesh type is INTER, motion information can be decoded through entropy decoding and inverse prediction processes. The decoded motion information is combined with a reference base mesh that has already been reconstructed and stored in a buffer to generate a reconstructed quantized base mesh for the current frame. An inverse quantization process can be performed on the reconstructed base mesh.
[0192] That is, the motion vector decoder derives the motion field of the base mesh of the current frame through motion estimation and compensation based on the base mesh in the reference frame when the mesh data in the bitstream is encoded based on inter prediction. The static mesh decoder decodes the base mesh when the mesh data in the bitstream is encoded based on intra prediction.
[0193] The displacement sub-stream is decoded into displacement video through the decoder of the video compression codec, for example, if it is compressed through a video codec, depending on the compression method used by the encoder, and then the image unpacking process is performed.
[0194] As another example, when compressed through arithmetic coding, the displacement sub-stream can be decoded into binarized syntax elements through arithmetic decoding, a contextual probability model (CPM) can be adaptively determined according to each bin of the syntax elements, and arithmetic decoding can be performed by predicting the occurrence probability of the bin through the CPM. The binarized syntax elements can be decoded through inverse binarization. Quantized displacement vector transform coefficients can be derived from the decoded syntax elements. As another example, when the displacement information type is INTER (when inter prediction is performed), an inverse inter prediction process is performed using reference information for the quantized displacement vector transform coefficients.
[0195] The quantized displacement vector transformation coefficient is restored to displacement information for each vertex through the inverse quantization, inverse transformation, and coordinate system transformation processes.
[0196] The restored base mesh and the restored displacement information are combined to generate the final decoded mesh. The attribute map substream is decoded by the decoder of the video compression codec used by the encoder, and then reconstructed into the final attribute map through processes such as color format conversion.
[0197] The restored decoded mesh and decoded attribute map can be utilized by the receiver as final mesh data that can be utilized by the user.
[0198] Fig. 12 illustrates a mesh data transmission device according to embodiments.
[0199] FIG. 12 corresponds to the transmitting device (100) or mesh video encoder (102) of FIG. 2, the encoder (preprocessor and encoder) of FIG. 3 or FIG. 7, and / or a transmitting encoding device corresponding thereto. Each component of FIG. 12 corresponds to hardware, software, a processor, and / or a combination thereof.
[0200] The operation process of a transmitter for compressing and transmitting dynamic mesh data using V-Mesh compression technology may be as shown in Fig. 12. The transmitter of Fig. 12 may perform an intra-frame encoding (or intra-encoding or intra-screen encoding) process and / or an inter-frame encoding (or inter-encoding or inter-screen encoding) process.
[0201] The pre-processor (811) receives the original mesh as input and generates a simplified mesh (decimated mesh) (or base mesh) and a fitted decimated mesh (or subdivision). Simplification can be performed based on the target number of vertices or target number of polygons that constitute the mesh. Parameterization, which generates texture coordinates and texture connection information per vertex, can be performed on the simplified mesh. For example, parameterization is a process of mapping a 3D surface to a texture domain for the decimated mesh. If parameterization is performed using the UVAtlas tool, mapping information is generated that can identify where each vertex of the decimated mesh can be mapped on a 2D image. The mapping information is expressed and stored as texture coordinates, and the final base mesh is generated through this process. In addition, the work of quantizing the mesh information in floating-point form into fixed-point form can be performed. This result can be output as a base mesh to a motion vector encoder (813) or a static mesh encoder (814) through a switching unit (812). The pre-processor (811) can perform mesh subdivision on the base mesh to generate additional vertices. Depending on the subdivision method, vertex connection information, texture coordinates, and texture coordinate connection information including the added vertices can be generated. The pre-processor (811) can generate a fitted subdivided mesh by adjusting the vertex positions so that the subdivided mesh becomes similar to the original mesh.
[0202] According to embodiments, the base mesh is output to a motion vector encoder (813) via a switching unit (812) when performing inter-encoding for the corresponding mesh frame, and is output to a static mesh encoder (814) via a switching unit (812) when performing intra-encoding for the corresponding mesh frame. The motion vector encoder (813) may be referred to as a motion encoder.
[0203] For example, when performing intra-encoding (or intra-frame encoding) on the corresponding mesh frame, the base mesh can be compressed through a static mesh encoder (814). In this case, encoding can be performed on connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. The base mesh bitstream generated through encoding is transmitted to a multiplexer (823).
[0204] As another example, when performing inter-encoding (or inter-frame encoding) for the corresponding mesh frame, the motion vector encoder (813) can receive a base mesh and a reference reconstructed base mesh (or a reconstructed quantized reference base mesh) as input, calculate a motion vector between the two meshes, and encode the value. In addition, the motion vector encoder (813) can perform prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and encode a residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated through encoding is transmitted to the multiplexer (823).
[0205] The base mesh restoration unit (815) can receive the base mesh encoded by the static mesh encoder (814) or the motion vector encoded by the motion vector encoder (813) and generate a reconstructed base mesh. For example, the base mesh restoration unit (815) can perform static mesh decoding on the base mesh encoded by the static mesh encoder (814) to restore the base mesh. At this time, quantization can be applied before the static mesh decoding, and inverse quantization can be applied after the static mesh decoding. As another example, the base mesh restoration unit (815) can restore the base mesh based on the reconstructed quantized reference base mesh and the motion vector encoded by the motion vector encoder (813). The reconstructed base mesh is output to the displacement calculation unit (816) and the mesh restoration unit (820).
[0206] The displacement calculation unit (816) can perform mesh refinement on the restored base mesh. The displacement calculation unit (816) can calculate a displacement vector, which is a difference value in the vertex positions between the restored base mesh that has been refined and the fitted subdivision (or refined) mesh generated by the pre-processor (811). At this time, the displacement vector can be calculated as many times as the number of vertices of the refined mesh. The displacement calculation unit (816) can convert the displacement vector calculated in the 3D Cartesian coordinate system into a local coordinate system based on the normal vector of each vertex.
[0207] The displacement vector video generation unit (817) may include a linear lifting unit, a quantizer, and an image packing unit. That is, in the displacement vector video generation unit (817), the linear lifting unit may transform the displacement vector for effective encoding. The transformation may be performed by a lifting transformation, a wavelet transformation, etc., according to embodiments. In addition, quantization may be performed on the transformed displacement vector value, i.e., the transform coefficient, in the quantizer. At this time, different quantization parameters may be applied to each axis of the transform coefficient, and the quantization parameters may be derived according to the agreement of the encoder / decoder. The displacement vector information that has undergone transformation and quantization may be packed into a 2D image in the image packing unit. The displacement vector video generation unit (817) may generate a displacement vector video by bundling the packed 2D images for each frame, and the displacement vector video may be generated for each GoF (Group of Frame) unit of the input mesh.
[0208] The displacement vector video encoder (818) can encode the generated displacement vector video using a video compression codec. The generated displacement vector video bitstream is transmitted to a multiplexer (823).
[0209] The displacement vector restoration unit (819) may include a video decoder, an image unpacking unit, an inverse quantizer, and an inverse linear lifting unit. That is, the displacement vector restoration unit (819) performs decoding on an encoded displacement vector in the video decoder, performs image unpacking in the image unpacking unit, performs inverse quantization in the inverse quantizer, and then performs inverse transformation in the inverse linear lifting unit to restore the displacement vector. The restored displacement vector is output to the mesh restoration unit (820). The mesh restoration unit (820) restores the deformed mesh based on the base mesh restored in the base mesh restoration unit (815) and the displacement vector restored in the displacement vector restoration unit (819). The restored mesh (or restored deformed mesh) has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.
[0210] The texture map video generation unit (821) can regenerate a texture map based on the texture map (or attribute map) of the original mesh and the restored deformed mesh output from the mesh restoration unit (820). According to embodiments, the texture map video generation unit (821) can assign color information per vertex of the texture map of the original mesh to the texture coordinates of the restored deformed mesh. According to embodiments, the texture map video generation unit (821) can generate a texture map video by grouping the regenerated texture maps by GoF unit for each frame.
[0211] The generated texture map video can be encoded using a video compression codec of the texture map video encoder (822). The texture map video bitstream generated through encoding is transmitted to a multiplexer (823).
[0212] A multiplexer (823) multiplexes a motion vector bitstream (e.g., in case of inter encoding), a base mesh bitstream (e.g., in case of intra encoding), a displacement vector bitstream, and a texture map bitstream into a single bitstream. The single bitstream can be transmitted to a receiver via a transmitter (824). Alternatively, the motion vector bitstream, the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream can be generated as a file with one or more track data or encapsulated into segments and transmitted to a receiver via the transmitter (824).
[0213] Referring to FIG. 12, a transmitting device (encoder) can encode a mesh in an intra-frame or inter-frame manner. A transmitting device according to intra-encoding can generate a base mesh, a displacement vector (or displacement), and a texture map (or attribute map). A transmitting device according to inter-encoding can generate a motion vector (or motion), a displacement vector (or displacement), and a texture map (or attribute map). The texture map obtained from the data input unit is generated and encoded based on the restored mesh. The displacement is generated and encoded through the difference in vertex positions between the base mesh and the divided (or subdivided or subdivided) mesh. More specifically, the displacement is the difference in positions between the fitted sub-divided mesh and the sub-divided restored base mesh, i.e., the difference in vertex positions between the two meshes. In addition, the base mesh is generated by simplifying and encoding the original mesh through pre-processing. Motion is generated as motion vectors for the mesh of the current frame based on the reference base mesh of the previous frame.
[0214] Fig. 13 illustrates a mesh data receiving device according to embodiments.
[0215] Fig. 13 corresponds to the receiving device (110) or mesh video decoder (113) of Fig. 2, the decoder of Fig. 11, and / or the receiving decoding device corresponding thereto. Each component of Fig. 13 corresponds to hardware, software, a processor, and / or a combination thereof. The receiving (decoding) operation of Fig. 13 may follow the reverse process of the corresponding process of the transmitting (encoding) operation of Fig. 12.
[0216] The bitstream of the mesh data received by the receiver (910) is demultiplexed into a compressed motion vector bitstream (e.g., inter decoding) or a base mesh bitstream (e.g., intra decoding), a displacement vector bitstream, and a texture map bitstream after file / segment decapsulation in the demultiplexer (911). For example, if the current mesh has inter-screen encoding (i.e., inter encoding) applied, the motion vector bitstream is received, demultiplexed, and then output to the motion vector decoder (913) via the switching unit (912). As another example, if the current mesh has intra-screen encoding (i.e., intra encoding) applied, the base mesh bitstream is received, demultiplexed, and then output to the static mesh decoder (914) via the switching unit (912). Here, the motion vector decoder (913) may be referred to as a motion decoder.
[0217] According to embodiments, if the current mesh has inter-screen encoding applied according to frame header information, the motion vector decoder (913) can perform decoding on the motion vector bitstream. According to embodiments, the motion vector decoder (913) can reconstruct the final motion vector by adding the previously decoded motion vector as a predictor to the residual motion vector decoded from the bitstream.
[0218] According to embodiments, if the current mesh has been encoded within the screen according to the frame header information, the static mesh decoder (914) can decode the base mesh bitstream to restore connection information, vertex geometry information, texture coordinates, normal information, etc. of the base mesh.
[0219] According to embodiments, the base mesh restoration unit (915) can restore the current base mesh based on the decoded motion vector or the decoded base mesh. For example, if the current mesh has inter-screen encoding applied, the base mesh restoration unit (915) can generate a restored base mesh by adding the decoded motion vector to the reference base mesh and then performing inverse quantization. As another example, if the current mesh has intra-screen encoding applied, the base mesh restoration unit (915) can generate a restored base mesh by performing inverse quantization on the base mesh decoded through the static mesh decoder (914).
[0220] According to embodiments, the displacement vector video decoder (917) may decode the displacement vector bitstream as a video bitstream using a video codec or decode it using an arithmetic codec. That is, if the displacement vector bitstream is encoded using a video codec, for example, depending on the encoding codec type, a depacking process may be performed after decoding using the video codec. As another example, if encoded using arithmetic coding, arithmetic decoding may be performed on the displacement vector bitstream through a displacement vector arithmetic decoding unit, and if inter-screen prediction is performed, a residual value may be added to a reference displacement vector transform coefficient through inter-screen prediction to generate a current displacement vector transform coefficient.
[0221] According to embodiments, the displacement vector restoration unit (918) extracts displacement vector transformation coefficients from the decoded displacement vector video, and restores the displacement vector by applying inverse quantization and inverse transformation processes to the extracted displacement vector transformation coefficients. To this end, the displacement vector restoration unit (918) may include an image unpacking unit, an inverse quantizer, and an inverse linear lifting unit. If the restored displacement vector is a value in a local coordinate system, a process of inverse transformation to a Cartesian coordinate system may be performed.
[0222] The mesh restoration unit (916) can generate additional vertices by performing subdivision on the restored base mesh. Through subdivision, vertex connection information including the added vertices, texture coordinates, and texture coordinate connection information can be generated. At this time, the mesh restoration unit (916) can generate a final restored mesh (or a restored deformed mesh) by combining the subdivided restored base mesh with the restored displacement vector.
[0223] According to embodiments, the texture map video decoder (919) can decode the texture map bitstream as a video bitstream using a video codec to restore the texture map. The restored texture map has color information for each vertex contained in the restored mesh, and the color value of each vertex can be obtained from the texture map using the texture coordinates of each vertex.
[0224] According to embodiments, the mesh restored by the mesh restoration unit (916) and the texture map restored by the texture map video decoder (919) are shown to the user through a rendering process in the mesh data renderer (920).
[0225] Referring to FIG. 13, a receiving device (decoder) can decode a mesh in an intra-frame or inter-frame manner. A receiving device according to intra-decoding can receive a base mesh, a displacement vector (or referred to as displacement), a texture map (or referred to as attribute map), and render mesh data based on the restored mesh and the restored texture map. A receiving device according to inter-decoding can receive a motion vector (or referred to as motion), a displacement vector (or referred to as displacement), a texture map (or referred to as attribute map), and render mesh data based on the restored mesh and the restored texture map.
[0226] A mesh data transmission device and method according to embodiments may pre-process mesh data, encode the pre-processed mesh data, and transmit a bitstream including the encoded mesh data. A point mesh data reception device and method according to embodiments may receive a bitstream including mesh data and decode the mesh data. The mesh data transmission and reception method / device according to embodiments may be abbreviated as method / device according to embodiments. The mesh data transmission and reception method / device according to embodiments may also be referred to as 3D data transmission and reception method / device or point cloud data transmission and reception method / device.
[0227] A mesh data transmission device and method according to embodiments may pre-process mesh data, encode the pre-processed mesh data, and transmit a bitstream including the encoded mesh data. A mesh data reception device and method according to embodiments may receive a bitstream including mesh data and decode the mesh data. A mesh data transmission / reception method / device according to embodiments may be abbreviated as method / device. A mesh data transmission / reception method / device according to embodiments may also be referred to as a 3D data transmission / reception method / device.
[0228] As described above, the transmitting device first performs pre-processing on the input original mesh as in FIG. 7 or FIG. 12. More specifically, as in FIG. 4, the mesh simplification unit of the pre-processor (200) generates a simplified mesh (decimated mesh) of the input original mesh, and the atlas parameterization (also called parameterization, Atlas parameterization, UV parameterization) unit generates texture coordinates for the vertices that constitute the simplified mesh. Then, the simplified mesh and texture coordinates are compressed and restored as base mesh data of the input original mesh.
[0229] At this time, the attribute transfer (426) of FIG. 7 or the texture map video generation unit (821) of FIG. 12 regenerates a new texture map based on the texture map of the original mesh and the texture coordinates of the mesh restored after encoding. Here, the texture coordinates of the restored mesh are the result of performing subdivision on the restored base mesh, and are values calculated based on the texture coordinates generated by the pre-processor (200). The texture map images regenerated based on the texture coordinates of the restored mesh are processed as video and compressed by an existing video codec (e.g., the video encoding unit (429) of FIG. 7 or the texture map video encoder (822) of FIG. 12). In this way, the texture coordinates generated by the pre-processor (200) affect the structure / shape of the regenerated texture map image, and consequently also affect the texture map video compression performance.
[0230] In this way, V-Mesh regenerates the texture map of the input original mesh as a texture map for the restored mesh during the encoding process, and then processes and compresses the regenerated texture map images into a video stream through a video encoder.
[0231] In the present disclosure, the original mesh (i.e., input mesh) input includes three-dimensional coordinates of vertices constituting the mesh, normal information of each vertex, mapping information for mapping the mesh surface to a 2D plane, connection information between the vertices constituting the surface, etc. The surface of the mesh can be expressed as a triangle or more polygons, and connection information between the vertices constituting each surface is stored according to a predetermined shape. The input mesh can be saved in the OBJ file format.
[0232] In the present disclosure, a texture map (or attribute map) contains information about the attributes of a mesh (color, normal, displacement, etc.), and stores data in the form of mapping the surface of the mesh onto a 2D image. Mapping which part (surface or vertex) of the mesh each data of this texture map corresponds to is based on the mapping information contained in the input mesh. Since the texture map contains data for each frame of the mesh video, it can also be expressed as a texture map (or attribute map) video. The texture map in the V-Mesh compression method mainly contains color information of the mesh, and is stored in an image file format (PNG, BMP, etc.).
[0233] At this time, the pre-processor and the attribute transfer (426) of FIG. 7 or the texture map video generation unit (821) of FIG. 12 may receive mesh data having one attribute information or may receive mesh data having multiple attribute information. In the present disclosure, mesh data having multiple attribute information may be referred to as mesh data having multiple attribute information or dynamic mesh data having multiple attribute information. That is, the input mesh may be mesh data having one attribute information or may be mesh data having multiple attribute information. In other words, the input mesh may be mesh data having one texture map or may be mesh data having multiple texture maps.
[0234] In the present disclosure, mesh data having one attribute information means that one texture map (e.g., one PNG image) is applied to one frame, and mesh data having multiple attribute information means that multiple texture maps (e.g., multiple PNG images) are applied to one frame.
[0235] That is, mesh data with a single attribute information means that all vertices within a frame reference a single texture map to express a visual attribute, while mesh data with multiple attribute information means that different vertices within a frame reference different texture maps to express a visual attribute. For example, assuming that a frame consists of a person, in the case of mesh data with a single attribute information, all vertices constituting the person reference a single texture map. In contrast, in the case of mesh data with multiple attribute information, the person is divided into multiple regions and a different texture map is referenced for each region. For example, in the case of multiple attributes, a person is divided into multiple regions (e.g., face region, torso region, arm region, leg region, and other regions), and the vertices of the face region, torso region, arm region, leg region, and other regions reference different texture maps. In other words, different parts of the mesh are mapped to different attribute images.
[0236] FIG. 14 is a diagram illustrating an example of dynamic mesh data having multiple attribute information according to embodiments. FIG. 14 is an example of a mesh within the Real World Textured Things (RWTT) dataset having multiple attributes, which is an example of an object (e.g., a person) captured using five texture images of 1k*1k resolution. The RWTT is a collection of publicly available textured 3D models generated using state-of-the-art commercially available photo-reconstruction tools. The purpose of this dataset is to provide a challenging benchmark for geometry processing algorithms targeting parameterized and textured 3D models derived from the real world.
[0237] In Fig. 14, "Original" in the upper left corner represents the original 3D mesh model, visually representing the mesh structure of the human body. This model is the original data before any actual textures are applied, and only the geometric structure of the model is included. In addition, "Voxelized (qp=12, qt=12, rotated / translated)" in Fig. 14 represents the original mesh in a voxelized state, showing the result of performing the pre-processing process (including quantization and position transformation) for compression.
[0238] In Fig. 14, the upper right portion is an example of a texture map corresponding to a frame related to the 3D mesh model, i.e., mesh data, in the upper left portion, showing that five texture maps are configured for the same 3D mesh model. Here, 5 is a number to aid understanding by those skilled in the art, and it is obvious that all texture maps greater than or equal to 2 belong to the present disclosure.
[0239] In Fig. 14, the lower right part shows an example in which multiple texture maps are concatenated to form a sequence. Fig. 14 is an example, and multiple texture maps may be concatenated horizontally or vertically to form a sequence. That is, referring to Fig. 14, a frame related to mesh data may include an object (e.g., a person or a building). For example, when the object is large in size or for high resolution, there may be multiple attribute information (i.e., texture maps) regarding one object and / or a bounding box regarding one object. For example, when the attribute is a texture map, there may be multiple texture maps regarding the original object or the voxelized object (qp=12, qt=12, rotated / translated). The multiple texture maps may be referred to as texture1, texture2, texture3, texture4, texture5, etc. In order to restore the object of Fig. 14, information from texture1 to texture5 may all be required.
[0240] However, if the data within multiple texture maps are listed and used sequentially without considering their characteristics, the video coding efficiency will be poor.
[0241] The present disclosure proposes a method for generating new texture coordinates so as to improve the compression performance of a texture map video using a video codec when using dynamic mesh data having multiple attribute information, and so as to maximize temporal consistency in a regenerated texture map video stream. In addition, by improving the compression performance of the texture map video through the present disclosure, the dynamic mesh compression performance can be improved. In addition, the present disclosure allows a V-DMC user to store and utilize a bitstream generated after compressing mesh data obtained using an encoder to which the present disclosure is applied, and to use fewer resources in storing and utilizing the bitstream at a receiver.
[0242] That is, the present disclosure relates to a method for improving video compression performance of data having a plurality of attribute information.
[0243] According to embodiments, the present disclosure enables the regeneration of attribute information (i.e., a texture map) with high inter-frame image correlation, i.e., maximally maintaining temporal consistency, by generating texture coordinates by reflecting inter-frame image similarity when generating texture coordinates.
[0244] According to embodiments, the present disclosure provides a method for reconstructing the arrangement of multi-attribute information during a pre-processing process of a pre-processor when mesh data having multi-attribute information is input as an input mesh. The pre-processing process of the present disclosure may be performed in the pre-processor (200) of FIG. 3, FIG. 4, FIG. 7, or FIG. 12.
[0245] According to embodiments, the pre-processor (200) of the present disclosure generates texture coordinates for a simplified input mesh, the attribute transfer (or texture map video generation unit) regenerates a texture map using the generated texture coordinates, and the video encoder compresses the regenerated texture map using a video codec.
[0246] According to embodiments, in the V-Mesh compression method, pre-processing is first performed on the input original mesh as shown in FIG. 4. More specifically, the pre-processor (200) creates a simplified mesh (decimated mesh) from the input original mesh, i.e., the input mesh, through a mesh simplification unit, and generates texture coordinates for vertices constituting the simplified mesh through an atlas parameterization (also called UV parameterization) unit. The simplified mesh and texture coordinates become base mesh data of the input original mesh and are compressed and restored.
[0247] According to embodiments, the mesh simplification unit may select vertices to be removed from an input mesh based on criteria defined by the user, and then remove the selected vertices and triangles connected to the selected vertices. In the process of performing mesh simplification, information on a voxelized input mesh, a target triangle ratio (TTR), and a minimum triangle component (CCCount) may be transmitted as input, and a simplified mesh may be obtained as output. In this process, connected triangle components smaller than the set minimum triangle component (CCCount) may be removed.
[0248] According to embodiments, the atlas parameterization unit maps a 3D surface to a texture domain for a simplified mesh. The present disclosure can perform parameterization using the UVAtlas tool. Through this process, mapping information is generated, indicating where each vertex of the simplified mesh can be mapped to on a 2D image. The mapping information is expressed and stored as texture coordinates, and through this process, a final base mesh is generated.
[0249] According to embodiments, when mesh data having multi-attribute information is input as an input mesh, the pre-processor (200) may additionally perform a process of reconstructing the arrangement of the multi-attribute information. At this time, the process of reconstructing the arrangement of the multi-attribute information may be performed in the mesh simplification unit, or may be performed by adding a separate block or module before the mesh simplification unit.
[0250] According to embodiments, the present disclosure can arrange texture maps by adjusting the indexes in consideration of the efficiency of video coding in the input form (a 1 x 6 concatenation form or a 2 x 3 rectangular form). For example, if 6 texture maps are input in the order of 123456, the 6 texture maps can be rearranged in the order of 231456. Here, 123456 is each texture map index for identifying 6 texture maps.
[0251] According to embodiments, the similarity can be measured through correlation or feature matching of all or part of each texture map (i.e., patch unit) and the reconstructing can be done right from the arrangement process. That is, the present disclosure can improve the performance of texture map video compression using a video codec when using dynamic mesh data having multi-attribute information, and can reconstruct the arrangement of multiple texture maps so that temporal consistency can be reflected as much as possible in the regenerated texture map video stream.
[0252] And, in the case of mesh data having multiple texture maps, since vertices within one frame reference different texture maps, signaling information (or metadata parameter) for each vertex within the frame includes, for each vertex, an index (or identifier) of a texture map referenced by the vertex and texture coordinate information (e.g., UV coordinate information) of the vertex within the texture map. That is, each vertex has signaling information (or metadata) including texture map index information that it should reference and position information within the texture map.
[0253] In the pre-processor (200) of the present disclosure, the rearranged multiple texture maps correspond to one frame. That is, multiple texture maps exist in one frame. For example, if five texture maps, each with a resolution of 1k*1k, are merged, a texture map, i.e., a texture image, having a resolution of 1k*5k is applied to one frame. That is, the present disclosure merges (or connects) multiple texture maps whose arrangements have been reconstructed and treats them as a texture map of a single frame. The present disclosure can reconstruct the arrangement of multiple texture maps by changing the order of texture map indexes.
[0254] By merging multiple texture maps input in this way or multiple texture maps whose arrangements are reconstructed based on similarity, temporal continuity can be secured and can be effectively utilized in video prediction-based encoding.
[0255] Fig. 15 shows an example of a texture map video regenerated according to the aforementioned V-Mesh method. That is, when generating texture coordinates for a simplified mesh in the atlas parameterization step of the pre-processor (200), the texture coordinates for each vertex are generated using only information of the current frame generated by merging multiple texture maps. That is, in Fig. 15, each frame may include multiple (e.g., 6 or more) texture maps. At this time, the multiple texture maps may be reconstructed in arrangement based on similarity, etc., and then mapped to the corresponding frame. Reconstruction of the arrangement of multiple texture maps may be performed by changing the order of texture map indices. This process may be performed on a frame-by-frame basis. And, if the base mesh with the texture coordinates generated in this way is compressed and restored again, and this is used to regenerate the texture map, the texture maps regenerated in the attribute transfer (see 426 in FIG. 7) or the texture map video generation unit (821) of FIG. 12 have low image correlation between frames, as shown in FIG. 15. Taking the mesh patch called a face as an example in FIG. 15, it can be seen that in the i-th frame it is located near the lower right (951), in the i+1-th frame it is located near the left center (952), in the i+2-th frame it is separated into two and located near the right center (953, 954), and in the i+3-th frame it is located near the upper right (955). In this way, the texture coordinates of the simplified mesh generated using only the current frame information in the parameterization step and the texture maps regenerated based on the original texture map have low image correlation between frames and show almost no temporal consistency.
[0256] In particular, when compressing a texture map video using a video codec as in FIG. 7 or FIG. 12, that is, when compressing a texture map video by applying an inter-screen encoding method, the accuracy of inter-screen prediction is low, resulting in a large amount of residual signal data to be encoded, which may result in a large number of compressed bitstreams. In addition, the larger the encoded bitstream size of the input mesh, the more resources and costs are required for system operations such as data transmission and storage.
[0257] That is, the V-Mesh compression method described so far simplifies the input dynamic mesh data and encodes the simplified mesh data using the static mesh compression method. At this time, the simplified mesh data includes geometric information about the vertices that make up the mesh, and texture coordinate information for obtaining color information of each vertex from multiple texture maps. V-Mesh generates texture coordinates for the simplified mesh by processing the texture coordinates of the multiple texture maps of the original mesh. At this time, the generated texture coordinates are processed to be more efficient in texture map video compression than the texture coordinates of the original mesh, but since the temporal consistency characteristics of the video are not reflected, there are limitations in achieving sufficient performance and efficiency when applying the inter-frame compression method for compression. In particular, as the input mesh content and its texture map resolution increase, and also as the target bitrate increases, the resolution and capacity of the compressed texture map video increase, so the aforementioned method inevitably has limitations in compressing, transmitting, and utilizing mesh content.
[0258] Therefore, in order to improve the performance of texture map video compression using a video codec, this paper proposes a method of generating new texture coordinates so that temporal consistency can be reflected to the maximum extent possible in the regenerated texture map video stream. That is, by using the method proposed in this paper when encoding a texture map video in a transmitting device, the performance of the texture map video compression can be improved, and thus, the performance of the dynamic mesh compression can be improved. In addition, a user can store, utilize, and transmit the generated bitstream after compressing the mesh data acquired using an encoder to which the present disclosure is applied, and a receiver can use fewer resources when storing and utilizing the bitstream. In other words, by regenerating and compressing a texture map (i.e., an attribute map) composed by merging multiple texture maps based on the method proposed in this disclosure, the cost required for using mesh content in media and communication systems, etc. can be reduced, and the scope of applications utilizing mesh content can be further expanded. In this case, as an embodiment, multiple texture maps are reconfigured based on similarity, etc., and then merged into a single texture map.
[0259] In particular, since the proportion of the texture map video bitstream in the mesh bitstream compressed with V-Mesh is very large, improving the texture map video compression performance using the method proposed in this disclosure can improve the compression performance of V-Mesh.
[0260] Fig. 16 shows an example of a detailed block diagram of a parameterization unit according to embodiments.
[0261] According to embodiments, the parameterization unit is included in the pre-processor (200) of FIG. 3, FIG. 4, FIG. 7 or FIG. 12. That is, in one embodiment, the parameterization unit of FIG. 16 is located between the mesh simplification unit and the fitting subdivision surface unit within the pre-processor (200).
[0262] The parameterization unit of FIG. 16 may include a polygon / vertex segmentation unit (11011), a mesh patch segmentation unit (11013), and a mesh patch packing unit (11015). Each component for parameterization of FIG. 16 corresponds to hardware, software, a processor, and / or a combination thereof.
[0263] First, the simplified mesh in the mesh simplification section of the pre-processor (200) is input to the polygon / vertex division section (11011) of the parameterization section.
[0264] In the polygon / vertex segmentation unit (11011) according to the embodiments, segmentation of polygons or vertices is performed based on the characteristics of polygons (in the shape of triangles or squares) or vertices that constitute the mesh for a decimated mesh. In this segmentation operation, the direction of the polygon formed by the connection of the vertices that constitute the mesh or the direction of each vertex is determined to be closest to which of the directions of the planes of the bounding box that encloses the mesh object. The direction can be determined by the normal vector of the polygon or the normal vector value of the vertex. In other words, the normal vector of the mesh polygon or the normal vector of the vertex is compared to which of the normal vectors of the six planes of the bounding box is most similar, and the direction of the plane that has the most similar normal vector is determined for the corresponding polygon or vertex. The normal vector of the polygon can be calculated using the position coordinate values of the vertices that constitute the polygon. The normal vector of a vertex can be used as is if the input mesh data includes normal information for each vertex. If normal information is not included, it can be calculated using the vertices neighboring the vertex. To this end, the method of calculating the normal value of each point when generating a patch for a point cloud in V-PCC can be applied. The direction of a polygon or vertex can be one of the six planes of the bounding box, or it can be one of the directions including directions added by the user.
[0265] According to embodiments, the mesh patch segmentation unit (11013) performs the task of dividing into mesh patches, which are sets of adjacent polygons or points with the same direction, based on the direction information of polygons or vertices obtained during the segmentation process of the polygon / vertex segmentation unit (11011). When this process is performed based on polygons, polygons with the same direction and adjacent to each other are calculated based on the current polygon, and the maximum number of adjacent polygons, i.e., the maximum number of polygons that can form a mesh patch, can be defined by the user. After the mesh patch segmentation process is performed for all polygons, polygons that are not included in the patch can be included in the patch with the most dominant direction among the directions of the adjacent polygon patches or included in the polygon patch with the most similar direction. When this process is performed based on vertices, a method of dividing a point cloud into patches can be applied when generating patches for a point cloud in V-PCC.
[0266] According to embodiments, the mesh patch packing unit (11015) performs the task of mapping the previously generated mesh patches onto a single 2D image as shown in Fig. 17. The mesh patch packing can be performed by applying the point cloud patch packing method of V-PCC.
[0267] FIG. 17 is a diagram showing an example of mesh patches constituting a simplified mesh according to embodiments being mapped onto a 2D image.
[0268] According to embodiments, the mesh patch packing unit (11015) can perform mapping by referring to the packing result of the previous frame when mapping the generated mesh patches onto a single 2D image, thereby outputting a simplified mesh having texture coordinates. That is, the mesh patch packing unit (11015) performs mapping by referring to the packing result of the previous frame when mapping the generated mesh patches onto a single 2D image, thereby generating mapping information that can identify where each vertex of the simplified mesh is mapped on the 2D image. This mapping information is expressed and stored as texture coordinates.
[0269] That is, the mesh patch packing unit (11015) can perform mesh patch packing that maps the generated mesh patches in order of their sizes onto a 2D image. For example, if it is determined that there is no frame in which mesh patch packing has been completed previously, mesh patch packing that maps the generated mesh patches in order of their sizes onto a 2D image as described above can be performed. According to embodiments, the mesh patch packing unit (11015) can search for a position in the raster scan order starting from the (0, 0) coordinate of the image so that each mesh patch can be mapped onto a 2D image of a size specified by the user. In addition, the mesh patch can be mapped onto a 2D image (i.e., a 2D frame) by rotating according to an angle specified by the user. At this time, a new patch cannot be mapped to a position that has already been filled by a previous patch within the 2D image (i.e., a 2D frame).
[0270] And, in order to generate texture coordinates so that the similarity between frames can be reflected, when packing mesh patches, it can be performed by referring to the packing result of the previous frame. That is, if it is confirmed that there is a frame in which mesh patch packing has been completed previously, it is determined whether there is a matching mesh patch among the mesh patches of the previous frame for each mesh patch of the current frame. At this time, the matching criteria may be used such as the direction of the mesh patch, the positional similarity of the polygons / vertexes constituting the mesh patch, etc. In other words, if the direction of the mesh patches is the same, the number of polygons / vertexes constituting the mesh patches is similar, and their positions in the three-dimensional world or the two-dimensional world are similar, the two mesh patches can be matched with each other. If there is a mesh patch matching the current mesh patch in the previous frame, the current mesh patch can be mapped to be located on the current 2D image by referring to the position where the matched mesh patch is mapped on the 2D image. The mesh patches can be mapped onto the current 2D image to have the same location as the matched mesh patches, and if the same location is occupied by another patch, it can be mapped to a point as close as possible to that location. The mapping is performed by giving priority to mesh patches that have found matched mesh patches in the previous frame, and the unmatched mesh patches are then mapped to appropriate locations in the remaining image space. Through this process, the mesh patches of the current frame can be mapped to similar locations as the mesh patches of the previous frame.
[0271] And, by repeating this process, when the mesh patch packing is all completed, the result of each mesh patch constituting the simplified mesh being mapped and located on a 2D image can be obtained, as shown on the right side of FIG. 17. And the positions of the vertices constituting each mesh patch on the 2D image can be obtained. The positions where each vertex is mapped on the 2D image have a 2D coordinate value, and in the method proposed in the present disclosure, these coordinate values are used as the texture coordinates of the vertices. These texture coordinates are compressed as information constituting the base mesh as the texture coordinates for each vertex of the simplified mesh. That is, the simplified mesh having the texture coordinates is input to the fitting subdivision surface unit, and the subdivision (i.e., subdivision) and fitting processes are performed, and at this time, the fitted subdivided mesh and the base mesh are output to the encoder (201). The detailed operation of the fitting subdivision surface unit will be described with reference to FIG. 4, and will be omitted here to avoid redundant description.
[0272] Additionally, the above texture coordinates serve as a reference for determining where in the 2D texture map the color information of each vertex will be stored when regenerating the texture map for the mesh restored from attribute transfer.
[0273] According to embodiments, the encoder (201) calculates displacement information (or displacement vector) based on the base mesh and the fitted subdivided mesh output from the pre-processor (200), and generates a reconstructed defomed mesh based on the calculated displacement information. For example, the reconstructed defomed mesh is obtained by adding the reconstructed displacement to the reconstructed base mesh that has been subdivided (or subdivided or subdivided). Details of generating the reconstructed defomed mesh in this document will be referred to the description of FIG. 7 or FIG. 12, and will be omitted here to avoid redundant description.
[0274] According to embodiments, the attribute transfer (426) of FIG. 7 or the texture map video generation unit (821) of FIG. 12 regenerates a texture map (or attribute map) based on the texture map of the original mesh and the restored deformed mesh, as described above.
[0275] That is, as the process described in Fig. 16 progresses, the texture coordinates output from the parameterization section of the pre-processor (200) become a reference for determining where in the 2D texture map the color information of each vertex will be stored when regenerating the texture map for the restored mesh in the attribute transfer step.
[0276] As described above, the present disclosure has packed similar mesh patches between frames into similar locations on a 2D image through the aforementioned process, so that similar mesh patches can have similar texture coordinates. As a result, a texture map regenerated based on the texture coordinates can have a similar shape between frames. Through this, the final generated texture map video can maintain temporal consistency as much as possible, as shown in FIG. 18, and can exhibit high compression performance when inter-frame prediction is applied during texture map video compression.
[0277] The present disclosure extends the above-described process to a texture composition method with multiple attribute information, thereby maximizing encoding efficiency by maintaining temporal continuity. Therefore, encoding efficiency can be maximized by reconstructing a mesh data layout with multiple attribute information as an input mesh, or by using it as is to generate the texture coordinates described above.
[0278] Fig. 18 shows another example of a texture map video regenerated according to embodiments. That is, Fig. 18 shows an example of the result of mapping mesh patches of the current frame onto a single 2D image by referring to the packing result of the previous frame.
[0279] That is, the present disclosure generates texture coordinates for each vertex using the packing result of the previous frame and the information of the current frame when generating texture coordinates for a simplified mesh in the atlas parameterization step of the pre-processor (200). Then, the base mesh having the texture coordinates generated in this way is compressed and then restored, and when this is used to regenerate the texture map, the texture maps regenerated in the attribute transfer (426 of FIG. 7) or the texture map video generation unit (821) of FIG. 12 have a high image correlation between frames, as shown in FIG. 18. Taking a mesh patch called a face as an example in FIG. 18, it can be seen that it is packed in almost the same position (i.e., near the lower right corner within the frame) (14051-14054) in the i-th frame, the i+1-th frame, the i+2-th frame, and the i+3-th frame. In this way, the texture coordinates of the simplified mesh generated using the previous frame information and the current frame information in the parameterization step and the regenerated texture maps based on the original texture map show a high image correlation between frames and maintain temporal consistency.
[0280] And, the texture map regenerated in the attribute transfer (see 426 of FIG. 7) or the texture map video generation unit (821) of FIG. 12 is encoded in the video encoder (see 429 of FIG. 7 or 822 of FIG. 12) and output as a compressed attribute map bitstream. In the present disclosure, the attribute map bitstream is used in the same meaning as the compressed texture map bitstream. At this time, the compressed attribute bitstream is multiplexed with other bitstreams in a multiplexer and then transmitted to the receiving device of FIG. 11 or FIG. 13, or is encapsulated into a file or segment and then transmitted to the receiving device of FIG. 11 or FIG. 13.
[0281] A detailed description of the receiving device of FIG. 11 or FIG. 13, which processes the compressed bitstream or file / segment received from the transmitting device to restore the mesh and attribute map, has been provided above, so a detailed description thereof will be omitted here to avoid redundant description.
[0282] The 3D data transmission device and method described so far are summarized as follows.
[0283] 1) For a simplified mesh, segmentation is performed by combining polygons or vertices with similar characteristics based on the characteristics of the polygons or vertices that make up the mesh.
[0284] 2) The segmented polygon or vertex set becomes one mesh patch, and a mesh patch packing process is performed to map the mesh patches generated from the input frame onto a 2D image.
[0285] 3) The mesh patch packing method applies a method of packing mesh patches of the current frame by referring to the packing result of the previous frame so that the packing similarity between frames can be reflected.
[0286] 4) Once the packing of all mesh patches onto the 2D image is complete, the mapped coordinate locations are used as texture coordinates for each vertex that makes up the patches.
[0287] 5) The texture map regenerated based on the texture coordinates obtained above can have a similar shape between frames, and thus the final generated texture map video can exhibit high compression performance when compressing a video using inter-screen prediction by maintaining temporal consistency as much as possible.
[0288] Fig. 19 is a flowchart illustrating an example of an encoding method according to embodiments. The encoding method according to embodiments may include a step of encoding a base mesh of mesh data (S31011), a step of encoding a displacement of mesh data (S31012), and a step of encoding an attribute of mesh data (S31013).
[0289] The encoding method of FIG. 19 may further include a step of pre-processing the mesh data before performing steps S31011 to S31013, i.e., before encoding the mesh data. The pre-processing of the present disclosure is performed in the pre-processor (200) of FIG. 4. FIG. 7 or FIG. 12 are diagrams showing examples of a transmitting device including the pre-processor (200) of FIG. 4.
[0290] According to embodiments, the pre-processing step may include a step of reconstructing the arrangement of multiple texture maps by measuring similarity through a method such as correlation or feature matching of all or part of each texture map (patch unit) when mesh data having multiple attribute information (i.e., texture maps) is input as an input mesh.
[0291] According to embodiments, the preprocessing step may include a mesh simplification step, a parameterization step, and a fitting subdivision surface step. The preprocessing may further include a GoF generation step. A detailed description of each step is provided in reference to the description of FIG. 4 described above.
[0292] According to embodiments, the parameterization step may include a polygon segmentation step, a mesh patch segmentation step, and a mesh patch packing step. For a detailed description of each step, refer to the description of FIGS. 16 to 18 described above. That is, when the polygon segmentation step, the mesh patch segmentation step, and the mesh patch packing step are performed as described in FIGS. 16 to 18, texture coordinates for the vertices of the simplified mesh are generated.
[0293] And, the texture coordinates for the vertices of the simplified mesh are encoded in the step (S31011) of encoding the base mesh of the mesh data by including them as base mesh data. The encoded base mesh is restored through the base mesh restoration unit. The mesh restoration unit generates a restored mesh by combining the restored base mesh and the displacement vector restored after encoding in the step (S31012) of encoding the displacement of the mesh data. That is, the base mesh restored by the base mesh restoration unit is combined with the displacement vector after subdivision, and the texture coordinates for the vertices additionally generated during the subdivision of the base mesh are calculated based on the texture coordinates of the restored base mesh. As a result, the texture coordinates of the mesh restored by the mesh restoration unit include the texture coordinates of the restored base mesh and the texture coordinates of the subdivided vertices calculated based on the texture coordinates.
[0294] The step (S31013) of encoding attributes of the above mesh data may perform encoding based on a video codec for attribute data (or texture map). To this end, the texture map video generation unit (or attribute transfer) processes the original texture map and regenerates it as a texture map for the restored mesh. Here, the vertex of the original mesh that is most similar to each vertex of the restored mesh is found, and the color information of the vertex is acquired from the original texture map and assigned to the 2D image of the texture map to be regenerated. The texture coordinates of the corresponding vertices of the restored mesh become the locations where the color information is assigned. In other words, a new texture map is regenerated based on the texture coordinates of the restored mesh. The regenerated texture maps for each frame are processed as a video, compressed through a texture map video encoder, and then transmitted. The texture map video generation unit (or attribute transfer) may be referred to as a texture map video generation step or an attribute transfer step.
[0295] At this time, since the texture coordinates generated in the pre-processing stage reflect the similarity between screens, the similarity between screens can be maintained in the texture map video regenerated through the texture map video generation unit (or attribute transfer), which can further improve the compression performance when encoding the texture map video using inter-screen encoding.
[0296] According to embodiments, the step (S31011) of encoding the base mesh of the mesh data may encode the base mesh through a static mesh encoder when performing intra-encoding or intra-frame encoding on the corresponding mesh frame. In this case, encoding may be performed on connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. The step (S31011) of encoding the base mesh of the mesh data may, when performing inter-encoding or inter-frame encoding on the corresponding mesh frame, use a motion vector encoder to calculate a motion vector between the base mesh and the reference restored base mesh (or the restored quantized reference base mesh) and encode the value. In addition, a previously encoded / decoded motion vector may be used as a predictor to perform connection information-based prediction, and a residual motion vector obtained by subtracting the predicted motion vector from the current motion vector may be encoded.
[0297] The step (S31012) of encoding the displacement of the above mesh data may perform video codec-based encoding on the displacement data or arithmetic codec-based encoding. The step (S31012) of encoding the displacement of the above mesh data may convert the coordinate system of the displacement data from a three-dimensional Cartesian coordinate system to a local coordinate system before encoding the displacement data.
[0298] The encoding method may further include a step of transmitting a bitstream including encoded basemesh, encoded displacement data, encoded attribute data and / or atlas data.
[0299] The encoding method of the present disclosure can be performed by an encoding device (encoder). The encoding device includes a memory and at least one processor connected to the memory, and the at least one processor can be configured to pre-process mesh data, encode a base mesh of the mesh data, encode displacement of the mesh data, and encode attributes of the mesh data.
[0300] The embodiments further include a computer-readable storage medium storing a bitstream generated by the method according to FIG. 19.
[0301] Embodiments further include a method comprising the steps of obtaining a bitstream for mesh data, the bitstream being generated based on the steps of encoding a basemesh of the mesh data, encoding a displacement of the mesh data, and encoding an attribute of the mesh data, and transmitting data including the bitstream.
[0302] Fig. 20 is a flowchart showing an example of a decoding method according to embodiments. The decoding method according to embodiments may include a step of decoding a base mesh in a bitstream (S32011), a step of decoding a displacement in the bitstream (S32012), and a step of decoding an attribute in the bitstream (S32013). The decoding step of Fig. 20 may further include a step of receiving a bitstream including a base mesh, displacement data, and attribute data, or a file in which a bitstream is encapsulated. The receiving step performs a decapsulation process to extract a bitstream when a file is received, and omit the decapsulation process when a bitstream is received.
[0303] The step (S32011) of decoding the base mesh in the bitstream may, if the current mesh has been subjected to inter-screen encoding, use the previously decoded motion vector as a predictor to restore the final motion vector by adding it to the residual motion vector decoded from the base mesh bitstream. The step (S32011) of decoding the base mesh in the bitstream may, if the current mesh has been subjected to intra-screen encoding, statically decode the base mesh bitstream to restore connection information, vertex geometry information, texture coordinates, normal information, etc. of the base mesh.
[0304] The step (S32012) of decoding displacement in the above bitstream performs decoding on the displacement bitstream based on the video codec if the displacement data is encoded based on a video codec, and performs decoding based on the arithmetic codec if the displacement data is encoded based on an arithmetic codec.
[0305] The step of decoding an attribute in the bitstream (S32013) restores attribute data by decoding the attribute bitstream based on a video codec. According to embodiments, the step of decoding an attribute (S32013) may restore attribute information (e.g., texture or texture map) of each vertex based on signaling information (or metadata) after the attribute bitstream is decoded. Here, the signaling information includes texture map index information for identifying a texture map to be referenced by each vertex and position information (e.g., UV coordinates) within the texture map.
[0306] The base mesh, displacement data, and attribute data decoded in steps S32011-S32013 can be rendered after undergoing post-processing such as mesh restoration and reconstruction.
[0307] The decoding method of the present disclosure can be performed by a decoding device (decoder). The decoding device includes a memory and at least one processor connected to the memory, and the at least one processor can be configured to decode a basemesh within a bitstream, decode a displacement within the bitstream, and decode an attribute within the bitstream.
[0308] In this way, the present disclosure can reflect the temporal consistency of a video by referring to the mesh patch packing result of the previous frame when packing a mesh patch onto a 2D image. That is, when dividing a simplified mesh into mesh patch units and mapping it onto a 2D image, by referring to the mesh patch packing result of the previous frame, the most similar mesh patches can be mapped to similar positions, so that the texture coordinates regenerated based on the mapped positions can maintain similarity between screens. As a result, the texture map images regenerated based on the texture coordinates can have similar structures and shapes between frames, and when compressing this using a video codec, inter-screen prediction can be effectively applied, thereby achieving improved compression performance.
[0309] Therefore, the texture map video compression performance can be improved, which is directly linked to the improved mesh content compression performance using V-Mesh. As mentioned above, since the video bitstream occupies a very large portion of the encoded mesh bitstream, improving this can lower the resource occupancy required by the mesh system and reduce usage costs, thereby enabling more efficient mesh system operation and expanding the range of applications utilizing mesh. In particular, the mesh patch packing method that reflects the mesh patch packing result of the previous frame has higher usability when the user directly creates and produces mesh content and transmits and utilizes it. For example, systems / platforms / services such as AR / hologram-based video conferencing systems that communicate in real time using 3D objects reflecting the user's appearance can be examples. In this way, the present disclosure can have the effect of enhancing the usability of V-Mesh.
[0310] As described above, the present disclosure can regenerate a texture map with high inter-frame image correlation (i.e., maximally reflecting temporal consistency) by reflecting inter-frame image similarity when generating texture coordinates of a simplified mesh. This can improve the compression performance of dynamic meshes of V-Mesh, and in particular, the compression performance of texture map videos of the mesh.
[0311] Each of the parts, modules, or units described above may be software, processors, or hardware parts that execute sequential execution processes stored in memory (or storage units). Each of the steps described in the embodiments described above may be performed by processors, software, or hardware parts. Each of the modules / blocks / units described in the embodiments described above may operate as a processor, software, or hardware. In addition, the methods presented in the embodiments may be implemented as code. This code may be written on a processor-readable storage medium and thus may be read by a processor provided by an apparatus.
[0312] Furthermore, throughout the specification, when a part is said to "include" a component, this does not exclude other components, unless otherwise specifically stated, but rather implies the inclusion of other components. Furthermore, terms such as "part" described in the specification mean a unit that processes at least one function or operation, which may be implemented using hardware, software, or a combination of hardware and software.
[0313] For convenience of explanation, this specification has been described separately in each drawing. However, it is also possible to design new embodiments by combining the embodiments described in each drawing. Furthermore, designing a computer-readable recording medium containing a program for executing the previously described embodiments, as required by those skilled in the art, is also within the scope of the embodiments.
[0314] The devices and methods according to the embodiments are not limited to the configurations and methods of the embodiments described above, but the embodiments may be configured by selectively combining all or part of each embodiment so that various modifications can be made.
[0315] Although preferred embodiments of the embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above, and various modifications may be made by those skilled in the art to which the present disclosure pertains without departing from the spirit or scope of the embodiments claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the embodiments.
[0316] The various components of the devices of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. The various components of the embodiments may be implemented by a single chip, for example, a single hardware circuit. The components according to the embodiments may be implemented by separate chips. At least one of the components of the devices of the embodiments may be configured with one or more processors capable of executing one or more programs, and the one or more programs may perform, or include instructions for performing, one or more of the operations / methods according to the embodiments. The executable instructions for performing the methods / operations of the devices of the embodiments may be stored in non-transitory CRMs or other computer program products configured to be executed by one or more processors, or may be stored in temporary CRMs or other computer program products configured to be executed by one or more processors. In addition, the memory according to the embodiments may be used as a concept including not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. Additionally, it may include implementations in the form of carrier waves, such as transmissions via the Internet. Furthermore, processor-readable recording media may be distributed across network-connected computer systems, allowing processor-readable code to be stored and executed in a distributed manner.
[0317] In this document, " / " and "," are interpreted as "and / or". For example, "A / B" is interpreted as "A and / or B", and "A, B" is interpreted as "A and / or B". Additionally, "A / B / C" means "at least one of A, B, and / or C". Also, "A, B, C" means "at least one of A, B, and / or C". Additionally, "or" in this document is interpreted as "and / or". For example, "A or B" can mean 1) "A" only, 2) "B" only, or 3) "A and B". In other words, "or" in this document can mean "additionally or alternatively".
[0318] Various elements of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. Various elements of the embodiments may be implemented on a single chip, such as a hardware circuit. In some embodiments, the embodiments may optionally be implemented on separate chips. In some embodiments, at least one of the elements of the embodiments may be implemented within one or more processors that include instructions for performing operations according to the embodiments.
[0319] Additionally, the operations according to the embodiments described in this document may be performed by a transceiver device including one or more memories and / or one or more processors according to the embodiments. One or more memories may store programs for processing / controlling the operations according to the embodiments, and one or more processors may control various operations described in this document. One or more processors may be referred to as a controller, etc. The operations according to the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in a processor or a memory.
[0320] Terms such as "first" and "second" may be used to describe various components of the embodiments. However, the various components according to the embodiments should not be interpreted in a limited manner by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal may be referred to as a "second user input signal." Similarly, a second user input signal may be referred to as a "first user input signal." The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although "first user input signal" and "second user input signal" are both user input signals, they do not necessarily mean the same user input signals unless the context clearly indicates otherwise.
[0321] The terminology used to describe the embodiments is for the purpose of describing particular embodiments and is not intended to be limiting of the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless the context clearly dictates otherwise. The expressions “and / or” are used to mean all possible combinations of the terms. The expression “comprises” or “includes” describes the presence of features, numbers, steps, elements, and / or components, but does not mean that additional features, numbers, steps, elements, and / or components are not included. Conditional expressions such as “if” or “when” used to describe the embodiments are not intended to be limited to only optional cases. When a specific condition is satisfied, a related action is performed in response to a specific condition, or a related definition is intended to be interpreted.
[0322] As described above, the relevant contents have been described in the best form for carrying out the embodiments.
[0323] As described above, the embodiments may be applied, in whole or in part, to mesh data transmission and reception devices and systems. Those skilled in the art will appreciate that various modifications and variations may be made to the embodiments within the scope of the embodiments. The embodiments may include modifications and variations, and such modifications and variations do not depart from the scope of the claims and their equivalents.
Claims
1. Step of pre-processing the input mesh data; A step of encoding a base mesh of the above pre-processed mesh data; A step of encoding displacement data of the above pre-processed mesh data; and A step of encoding attribute data of the above pre-processed mesh data; comprising; Encoding method.
2. In the first paragraph, the pre-processing step An encoding method comprising a step of reconstructing the arrangement of a plurality of texture maps based on the similarity of the plurality of texture maps, when the input mesh data includes a plurality of texture maps.
3. In paragraph 2, The above reconstruction step is an encoding method for reconstructing the arrangement of the plurality of texture maps by changing the order of texture map index information of the plurality of texture maps.
4. In paragraph 2, An encoding method in which the above plurality of texture maps correspond to one frame, and signaling information for each vertex within the frame includes index information of a texture map referenced by the corresponding vertex and position information within the referenced texture map.
5. In the second paragraph, the pre-processing step A step of generating simplified mesh data by simplifying mesh data including the above multiple texture maps, A step of generating texture coordinates of each vertex of the above simplified mesh data, and An encoding method further comprising the step of generating fitted subdivided mesh data by performing fitting to make the simplified mesh data having the above texture coordinates similar to the input mesh data after subdividing the simplified mesh data.
6. In the fifth paragraph, the texture coordinate generation step A step of performing segmentation by combining polygons or vertices having similar characteristics based on the polygon or vertex characteristics that constitute the above simplified mesh data, A step of generating mesh patches of the current frame based on a set of polygons or vertices segmented in the above step, and An encoding method comprising a step of generating texture coordinates of each vertex of the simplified mesh data by packing the mesh patches of the current frame onto a 2D image based on the mesh patch packing result of the previous frame in which mesh patch packing is completed.
7. In the 6th paragraph, the packing step A step of checking whether there is a mesh patch among the mesh patches of the previous frame that matches the current mesh patch of the current frame, and An encoding method comprising: if it is confirmed in the above step that a matching mesh patch exists in the previous frame, determining a mapping position of the current mesh patch on a 2D image based on the mapping position of the matched mesh patch in the previous frame, and packing the current mesh patch into the determined mapping position on the 2D image.
8. In the 7th paragraph, the mesh patch matching verification step An encoding method for determining that a mesh patch matches the current mesh patch if the mesh patches of the previous frame have the same direction as the current mesh patch and have similar numbers of polygons and / or vertices constituting the mesh patch.
9. In the 7th paragraph, the step of encoding the attribute data An encoding method for encoding a video by regenerating texture coordinates based on the texture coordinates of each vertex of the simplified mesh data and the texture coordinates of each vertex of the input mesh data.
10. Memory; and At least one processor connected to the memory; wherein the at least one processor comprises: Pre-processing the incoming mesh data; Encode the base mesh of the above pre-processed mesh data; Encoding displacement data of the above pre-processed mesh data; and Encoding attribute data of the above pre-processed mesh data; configured to do so; Encoding device.
11. In the 10th paragraph, the at least one processor, An encoding device that reconstructs the arrangement of a plurality of texture maps based on the similarity of the plurality of texture maps when the input mesh data includes a plurality of texture maps.
12. In the 11th paragraph, the at least one processor, Simplifying mesh data including the above multiple texture maps to generate simplified mesh data, Generate texture coordinates for each vertex of the above simplified mesh data, and An encoding device that subdivides simplified mesh data having the above texture coordinates and then performs fitting to make it similar to the input mesh data, thereby generating fitted subdivided mesh data.
13. In the 12th paragraph, the at least one processor, Segmentation is performed by combining polygons or vertices with similar characteristics based on the polygon or vertex characteristics that constitute the above simplified mesh data, Generate mesh patches of the current frame based on the set of polygons or vertices segmented in the above step, and An encoding device that generates texture coordinates of each vertex of the simplified mesh data by packing the mesh patches of the current frame onto a 2D image based on the mesh patch packing result of the previous frame in which mesh patch packing has been completed.
14. A computer-readable storage medium storing a bitstream generated by the method according to Article 10.
15. Step of obtaining bitstream for mesh data, The bitstream is generated based on the steps of: pre-processing input original mesh data; encoding a base mesh of the pre-processed mesh data; encoding displacement data of the pre-processed mesh data; and encoding attribute data of the pre-processed mesh data; and A method comprising the step of transmitting data including the bitstream.
Citation Information
Patent Citations
Method and Apparatus for Generation Texture Map, and Database Generation Method
KR1020160046399A
System and method of avoiding jamming signal
KR1020250054582A
Displacement fire extinguishing gas
KR1020250157542A
Controller, memory module and memory system
KR1020260008996A
Apparatus, method, and system of locating sound source and offseting sound source
KR102156922B1